Erbgut

01

Every DNA coding rule is a hypothesis.

We measure which ones pay off.

DNA can hold data for centuries in the space of a sugar cube. The problem is not writing it. It is that reading it back is expensive, and every pipeline today decides how to encode by following rules nobody has checked against their own sequencing channel.

Erbgut measures those rules, keeps the ones that earn their place, and learns from where the decoder actually fails, so a file comes back from fewer reads.

Scroll

02

A file becomes DNA.

03

Writing it is once.

03

Reading it back is the bill that comes back.

04

Every read comes back wrong. Differently wrong.

05

The field's answer: read more, take a vote.

05

67.2%

of strands come back exactly right.

six reads per strand, real Nanopore data

06

A small model learned what voting cannot fix.

06

88.1%

0.8M parameters. Same reads, same file, one extra step.

07

Then we turn it around.

07

The failures teach the encoder what to avoid.

08

Every rule, measured on your channel.

08

Each standard rule pays off on exactly one channel. Nobody measures that today.

09

88.1%

against 67.2%

exact strands at six reads, on 1,996 held-out real clusters

0.8M

parameters

trained in 13 minutes

229

strands exact, decoded

on the microcontroller itself, against 211 classic

Erbgut

Read the code Open the pitch deck

Real reads: Microsoft clustered Nanopore reads (MIT) and the DNAformer binned reads, Technion (CC BY 4.0).