DNA Oligo Encoding With Fountain Codes Under Sequencing Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA storage technologies face limitations in achieving high information density and reliability due to biochemical constraints, such as high GC content and homopolymer runs, which lead to sequencing errors and dropout issues, resulting in suboptimal utilization of the Shannon information capacity and challenges in perfect data retrieval.
Innovation Solution
The implementation of a DNA Fountain encoding algorithm that uses fountain codes to encode data into DNA oligos, screening sequences for biochemical constraints like GC content and homopolymer runs, and employing error correction mechanisms to ensure robust data retrieval, thereby approaching the theoretical information capacity of DNA storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If DNA sequences with high GC content and homopolymer runs are used for storage, then information density increases, but sequencing errors and dropout issues increase
Solution Approach 1:
The patent applies preliminary action by pre-screening DNA sequences against biochemical constraints (GC content, homopolymer runs, secondary structures) before synthesis. This proactive filtering prevents problematic sequences from being synthesized, thereby avoiding sequencing errors and dropout issues while maintaining high information density through efficient use of valid sequences.
2Reliability
If fountain codes are used to encode data into DNA oligos, then data retrieval reliability improves, but encoding complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the data encoding process into distinct modular steps: data segmentation into blocks, application of fountain codes to generate encoded oligos, and separate screening against biochemical constraints. This modular approach manages encoding complexity while achieving high data retrieval reliability through the erasure correction capabilities of fountain codes.
3Reliability
If screening sequences against biochemical constraints is implemented, then sequencing error rate decreases, but encoding time increases
Solution Approach 1:
The patent implements preliminary screening of DNA sequences against multiple biochemical constraints (GC content, homopolymer runs, secondary structures) before synthesis. This upfront filtering prevents the generation of problematic sequences, reducing sequencing errors while the efficient implementation of constraint checking minimizes the time penalty during the encoding phase.
Data Source
AI summary
Efficient encoding and decoding of data for storage in polymers is provided. In various embodiments, an input file is read. The input file is segmented into a plurality of segments. A plurality of packets is generated from the plurality of segments by applying a fountain code. Each of the plurality of packets is encoded as a sequence of monomers. The sequences of monomers are screened against at least one constraint. An oligomer is outputted corresponding to each sequence that passes the screening.


