Composite DNA Alphabet Encoding for Higher Data Storage Density
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA-based data storage systems face limitations in data storage capacity and density due to inherent information redundancy in DNA synthesis and sequencing technologies, leading to significant redundancy and inefficiencies.
Innovation Solution
A composite letter alphabet approach is employed, leveraging information redundancy by defining each letter as a mixture of molecular bases, allowing for higher information capacity and improved data storage density through the use of composite DNA letters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional DNA-based data storage systems are used, then data can be stored with long-term stability, but data storage capacity and density are limited due to inherent information redundancy
Solution Approach 1:
The patent changes the fundamental parameter of DNA encoding by using composite letters composed of multiple nucleotide bases (e.g., AG, CT, TG) instead of single bases. This parameter change increases the information capacity per synthesized position from the traditional 2 bits (4 possible bases) to approximately 4.3 bits per position, directly addressing the contradiction by improving data storage capacity while reducing information redundancy through more efficient encoding
Solution Approach 2:
The patent applies composite materials principle by creating composite DNA letters that combine multiple nucleotide bases into single encoding units. These composite letters (e.g., AG, CT, TG, AC) function as composite information carriers, allowing higher information density to be stored in the same physical DNA space, thereby resolving the contradiction between storage capacity and redundancy
2Quantity of substance
If more DNA sequences are synthesized to increase storage capacity, then data density improves, but synthesis costs and system complexity increase
Solution Approach 1:
By changing the encoding parameter from single nucleotides to composite letters, the system achieves higher storage density without proportionally increasing the number of DNA molecules required. The composite encoding scheme allows more information to be packed into fewer synthesized positions, reducing system complexity while improving data storage density
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A data storage system and method are provided, as well as systems and methods for fabrication, and writing and reading of data therein. The data storage system includes at least one population of molecular sequences including chains of basic molecular building -blocks, and defining at least one respective data-block encoding data in the data storage system. The data of the data-block is encoded in a sequence S = (π 1, π 2,..., π k..., π κ-1, π κ) of encoded letters {π k} associated with an alphabet ∑ ≡ { σm } |m= 1 to M, which are encoded according to the types of basic molecular building -blocks appearing at k respective location along storage segments of the molecular sequences of the population. The molecular sequences include a number Z of different types of basic molecular building -blocks {En}|n=1 to z, while the alphabet ∑ has a size M strictly greater than the number Z of types of building -blocks. Each alphabet letter om is associated with a vector {Pmn}|n=1 to z indicative of occurrences of basic molecular building-block En of type n in the alphabet letter σm. Accordingly each encoded letter π κ at location k in the storage segments of molecular sequences of the data-block/population, is mapped to a corresponding alphabet letter om by determining a match between the occurrence of basic molecular building-blocks of different types at that locations k of the molecular sequences of the population, with the vector {Pm n}|n=1 toz associated with the alphabet letter σm. In some implementations the component Pm n of the vector { Pn}m|n=1 to z associated with alphabet letter σm is indicative of a probability that a basic molecular building-block En of type n, 1 ≤ n ≤ Z, appears at the location k of the storage segment of a molecular strand of the at least one population in case the letter π κ encoded at that location k corresponds to the alphabet letter σm.