Degenerate-Base DNA Data Storage for Higher Information Density
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA digital data storage methods are hindered by high costs per unit data storage, limiting their practical implementation due to the high cost of DNA storage and the theoretical limit of information density using only four DNA bases (A, C, G, and T).
Innovation Solution
The use of degenerate bases or mixed bases, which are additional characters beyond the standard four, to encode data, allowing for a higher information capacity by encoding beyond the 2.0 bit/nt limit, achieved by synthesizing and mixing different nucleotide bases on a substrate according to specific ratios, effectively increasing the information capacity to 3.37 bit/nt and reducing DNA length by half.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If only four standard DNA bases (A, C, G, T) are used for encoding, then the storage system is simple to implement, but the information density is limited to 2.0 bit/nt
Solution Approach 1:
The patent changes the fundamental parameter of base composition by introducing degenerate bases (combinations of 2-4 bases) in addition to the four standard bases. This transforms the encoding system from using only 4 discrete base types to using 15 encoding characters (4 standard + 11 degenerate bases), thereby increasing information density from 2.0 bit/nt to 3.37 bit/nt while maintaining DNA's inherent simplicity
Solution Approach 2:
The patent creates composite base structures by combining multiple nucleotide bases into degenerate bases (e.g., mixing A and T in specific ratios to create a degenerate base). These composite bases function as single encoding units during synthesis but provide multiple information states, enabling higher information density without fundamentally changing the DNA storage mechanism
2Quantity of substance
If longer DNA sequences are used to achieve required storage capacity, then more data can be stored, but the synthesis cost and time increase significantly
Solution Approach 1:
The patent segments the encoding process by dividing data into blocks that are encoded using degenerate bases. This allows parallel synthesis of multiple DNA sequences with the same encoding length, reducing overall synthesis time while maintaining high storage capacity. The segmented approach enables modular processing and scaling
Solution Approach 2:
By changing the information density parameter from 2.0 bit/nt to 3.37 bit/nt through degenerate base encoding, the patent reduces the required DNA sequence length by approximately half for the same storage capacity. This directly reduces synthesis time and cost while achieving the required storage capacity
3Length of stationary object
If higher information density is achieved using degenerate bases, then DNA length is reduced, but the synthesis process becomes more complex
Solution Approach 1:
The patent modifies the synthesis process by introducing controlled variability in base composition through degenerate bases. Instead of synthesizing identical sequences, the system synthesizes sequences with controlled base ratio distributions. This parameter change enables shorter DNA lengths (reducing manufacturing complexity) while achieving the same information storage through statistical encoding
Solution Approach 2:
The patent uses copying strategies where multiple DNA molecules with varying degenerate base compositions are synthesized in parallel. These copies collectively represent the encoded information, with the ensemble of molecules providing the required information density without requiring any single molecule to be excessively long or complex
Data Source
AI summary
Disclosed is a storage method of DNA digital data, including: encoding a plurality of bit data to a plurality of base sequences including at least one degenerate base; and synthesizing at least two types of bases constituting the at least one degenerate base on a substrate based on a mixing ratio.


