Molecular Data Storage Using Composite DNA Letters for Higher Density
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA-based data storage systems face limitations in data storage capacity and density due to inherent information redundancy and chemical constraints, leading to inefficiencies in synthesis, storage, and sequencing processes.
Innovation Solution
The use of a composite letter alphabet approach, where each letter is defined by a predetermined mixture of molecular bases, leverages information redundancy to enhance data capacity, utilizing a coding scheme that extends the available alphabet and balances molecular strands without biases, allowing for higher information capacity and density.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional DNA-based data storage systems process large numbers of nominally identical molecules in parallel, then synthesis and sequencing can be performed efficiently, but significant information redundancy occurs and data storage capacity is limited
Solution Approach 1:
The patent changes the fundamental parameter of molecular representation from identical copies to composite molecules with varying base compositions. Each composite letter is represented by molecules having different proportions of A, C, G, and T bases at specific positions, allowing the alphabet size to exceed the number of base types. This parameter change eliminates information redundancy while maintaining parallel processing efficiency.
Solution Approach 2:
The patent introduces composite DNA letters formed by mixing multiple base types in predetermined ratios. Instead of using single base types (A, C, G, or T) to represent each letter, composite letters use combinations of bases with specific proportions, creating a richer information encoding scheme that increases data storage capacity without requiring additional molecules.
2Quantity of substance
If the alphabet size is extended beyond the number of basic molecular building blocks, then information capacity increases, but the complexity of the coding scheme increases
Solution Approach 1:
The patent uses probability vectors as parameters to define composite letters, where each vector contains the proportions of different bases. This mathematical parameterization provides a systematic and manageable way to extend the alphabet size beyond the number of base types, making the coding scheme complex yet structured and implementable.
Solution Approach 2:
The patent incorporates error correction codes that provide feedback mechanisms to detect and correct errors in composite letter representation. This feedback system manages the complexity by automatically handling decoding uncertainties, making the extended coding scheme more reliable and easier to implement despite the increased alphabet size.
3Quantity of substance
If composite letters with predetermined base mixtures are used, then data density increases to approximately 4.3 bits per synthesized position, but the precision required in base composition control increases
Solution Approach 1:
The patent transforms the discrete base selection problem into a continuous parameter space using probability vectors. By controlling base compositions as continuous proportions rather than discrete selections, the system achieves higher data density (4.3 bits per position) while making the precision requirements more manageable through statistical rather than absolute control.
Solution Approach 2:
The patent uses multiple copies of molecular sequences to represent each composite letter. By synthesizing many identical copies with the same base composition ratios, the system averages out synthesis variations and achieves the required compositional precision through statistical convergence, reducing the burden on individual molecule synthesis precision.
4Measurement precision
If molecular sequences are synthesized with balanced compositions without biases, then sequencing accuracy improves, but the synthesis process becomes more challenging
Solution Approach 1:
The patent changes the synthesis target from uniform base distribution to controlled non-uniform distributions defined by probability vectors. Each composite letter has a specific base composition profile that must be maintained across multiple copies, providing clear synthesis targets that improve sequencing accuracy while making the composition control requirements explicit and manageable.
Data Source
AI summary
A data storage system and method are provided, as well as systems and methods for fabrication, and writing and reading of data therein. The data storage system includes at least one population of molecular sequences including chains of basic molecular building-blocks, and defining at least one respective data-block encoding data in the data storage system. The data of the data-block is encoded in a sequence S=(π1, π2, . . . , πk . . . , πK-1, πK) of encoded letters {πk} associated with an alphabet Σ≡{σm}|m=1 to M, which are encoded according to the types of basic molecular building-blocks appearing at k respective location along storage segments of the molecular sequences of the population. The molecular sequences include a number Z of different types of basic molecular building-blocks {En}|n=1 to Z, while the alphabet Σ has a size M strictly greater than the number Z of types of building-blocks. Each alphabet letter σm is associated with a vector {Pmn}|n=1 to Z indicative of occurrences of basic molecular building-block En of type n in the alphabet letter σm. Accordingly each encoded letter πk at location k in the storage segments of molecular sequences of the data-block/population, is mapped to a corresponding alphabet letter σm by determining a match between the occurrence of basic molecular building-blocks of different types at that locations k of the molecular sequences of the population, with the vector {Pmn}|n=1 to Z associated with the alphabet letter σm. In some implementations the component Pmn of the vector {Pmn}m|n=1 to Z associated with alphabet letter σm is indicative of a probability that a basic molecular building-block En of type n, 1≤n≤Z, appears at the location k of the storage segment of a molecular strand of the at least one population in case the letter πk encoded at that location k corresponds to the alphabet letter σm.


