Reed-Solomon Error Protection for Nucleic Acid Data Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid digital data storage methods are costly and error-prone, requiring high-accuracy sequencing due to high density data storage, which is inefficient and expensive, especially for nanopore sequencing.
Innovation Solution
The system employs error protection and correction schemes, such as Reed-Solomon codes, and efficient encoding methods using combinatorial arrangements of nucleic acid molecules, allowing for faster and more accurate reading with tolerance to errors, and optimized data structures for access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If base-by-base nucleic acid synthesis is used to store data at high density, then data storage density is improved, but sequencing cost and error rate increase
Solution Approach 1:
The patent segments data into multiple layers, where each layer is encoded using a different encoding scheme (e.g., base-by-base for some layers, combinatorial for others). This allows high-density storage in certain layers while using more robust encoding in other layers to maintain overall sequencing accuracy and reduce error rates.
Solution Approach 2:
The patent changes encoding parameters across different data layers by using variable encoding schemes with different redundancy levels. Some layers use higher redundancy encoding to tolerate sequencing errors, while others use denser encoding where errors are less critical, thereby optimizing the balance between storage density and sequencing reliability.
2Quantity of substance
If base-by-base nucleic acid synthesis is used for high density storage, then bits-per-base increases, but sequencing cost and time increase
Solution Approach 1:
The patent divides data into multiple layers that can be sequenced independently or in parallel. By segmenting the data storage into layers with different encoding densities, the system can process less dense layers faster while maintaining high overall storage capacity, thereby reducing total sequencing time.
Solution Approach 2:
The patent applies partial action by using different sequencing depths for different layers. Layers with higher redundancy encoding require less sequencing depth to achieve accurate readout, allowing the system to allocate sequencing resources efficiently and reduce overall sequencing time while maintaining data integrity.
3Quantity of substance
If base-by-base nucleic acid synthesis is used, then data storage density is improved, but error tolerance decreases
Solution Approach 1:
The patent segments data into multiple layers with different error tolerance characteristics. By distributing data across layers with varying encoding redundancy, the system ensures that errors in high-density layers do not catastrophically affect the entire dataset, as other layers provide error correction capability.
Solution Approach 2:
The patent applies beforehand cushioning by incorporating error correction codes and redundancy encoding in certain layers before sequencing occurs. This preemptive error protection allows the system to tolerate sequencing errors that inevitably occur during high-density base-by-base synthesis and reading.
4Measurement precision
If high accuracy sequencing is required for high density data, then data retrieval accuracy is improved, but cost increases
Solution Approach 1:
The patent segments data into layers with different accuracy requirements. By identifying which layers require high accuracy and which can tolerate lower accuracy, the system can apply cost-effective sequencing methods to appropriate layers, reducing overall sequencing cost while maintaining sufficient data retrieval accuracy for the application.
Solution Approach 2:
The patent changes the required measurement precision parameter across different data layers based on their encoding schemes and error correction capabilities. Layers with strong error correction can be sequenced at lower precision thresholds, reducing the cost of high-accuracy sequencing while maintaining overall data retrieval accuracy through the combined information from multiple layers.
Data Source
AI summary
The systems, devices, and methods described herein provide scalable methods for writing data to and reading data from nucleic acid molecules. The present disclosure covers four primary areas of interest: (1) accurately and quickly reading information stored in nucleic acid molecules, (2) partitioning data to efficiently encode data in nucleic acid molecules, (3) error protection and correction when encoding data in nucleic acid molecules, and (4) data structures to provide efficient access to information stored in nucleic acid molecules.


