DNA Oligo Encoding With Fountain Codes Under Sequencing Constraints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA storage technologies face limitations in achieving high information density and reliability due to biochemical constraints, such as high GC content and homopolymer runs, which lead to sequencing errors and dropout issues, resulting in suboptimal utilization of the Shannon information capacity and challenges in perfect data retrieval.

Innovation Solution

The implementation of a DNA Fountain encoding algorithm that uses fountain codes to encode data into DNA oligos, screening sequences for biochemical constraints like GC content and homopolymer runs, and employing error correction mechanisms to ensure robust data retrieval, thereby approaching the theoretical information capacity of DNA storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If DNA sequences with high GC content and homopolymer runs are used for storage, then information density increases, but sequencing errors and dropout issues increase

Engineering Contradiction:
Improveinformation densityVSAvoidsequencing accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-screening DNA sequences against biochemical constraints (GC content, homopolymer runs, secondary structures) before synthesis. This proactive filtering prevents problematic sequences from being synthesized, thereby avoiding sequencing errors and dropout issues while maintaining high information density through efficient use of valid sequences.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If fountain codes are used to encode data into DNA oligos, then data retrieval reliability improves, but encoding complexity increases

Engineering Contradiction:
Improvedata retrieval reliabilityVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the data encoding process into distinct modular steps: data segmentation into blocks, application of fountain codes to generate encoded oligos, and separate screening against biochemical constraints. This modular approach manages encoding complexity while achieving high data retrieval reliability through the erasure correction capabilities of fountain codes.

Inventive Principle:
Principle #1Segmentation

3Reliability

If screening sequences against biochemical constraints is implemented, then sequencing error rate decreases, but encoding time increases

Engineering Contradiction:
Improvesequencing error rateVSAvoidencoding time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary screening of DNA sequences against multiple biochemical constraints (GC content, homopolymer runs, secondary structures) before synthesis. This upfront filtering prevents the generation of problematic sequences, reducing sequencing errors while the efficient implementation of constraint checking minimizes the time penalty during the encoding phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10742233B2Efficient encoding of data for storage in polymers such as DNA
Publication Date: 2020.08.11 ERLICH LAB LLC
  • US10742233B2 patent drawing
  • US10742233B2 patent drawing
  • US10742233B2 patent drawing

AI summary

Efficient encoding and decoding of data for storage in polymers is provided. In various embodiments, an input file is read. The input file is segmented into a plurality of segments. A plurality of packets is generated from the plurality of segments by applying a fountain code. Each of the plurality of packets is encoded as a sequence of monomers. The sequences of monomers are screened against at least one constraint. An oligomer is outputted corresponding to each sequence that passes the screening.