Reed-Solomon Error Protection for Nucleic Acid Data Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current nucleic acid digital data storage methods are costly and error-prone, requiring high-accuracy sequencing due to high density data storage, which is inefficient and expensive, especially for nanopore sequencing.

Innovation Solution

The system employs error protection and correction schemes, such as Reed-Solomon codes, and efficient encoding methods using combinatorial arrangements of nucleic acid molecules, allowing for faster and more accurate reading with tolerance to errors, and optimized data structures for access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If base-by-base nucleic acid synthesis is used to store data at high density, then data storage density is improved, but sequencing cost and error rate increase

Engineering Contradiction:
Improvedata storage densityVSAvoidsequencing accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments data into multiple layers, where each layer is encoded using a different encoding scheme (e.g., base-by-base for some layers, combinatorial for others). This allows high-density storage in certain layers while using more robust encoding in other layers to maintain overall sequencing accuracy and reduce error rates.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes encoding parameters across different data layers by using variable encoding schemes with different redundancy levels. Some layers use higher redundancy encoding to tolerate sequencing errors, while others use denser encoding where errors are less critical, thereby optimizing the balance between storage density and sequencing reliability.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If base-by-base nucleic acid synthesis is used for high density storage, then bits-per-base increases, but sequencing cost and time increase

Engineering Contradiction:
Improvebits-per-baseVSAvoidsequencing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent divides data into multiple layers that can be sequenced independently or in parallel. By segmenting the data storage into layers with different encoding densities, the system can process less dense layers faster while maintaining high overall storage capacity, thereby reducing total sequencing time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by using different sequencing depths for different layers. Layers with higher redundancy encoding require less sequencing depth to achieve accurate readout, allowing the system to allocate sequencing resources efficiently and reduce overall sequencing time while maintaining data integrity.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If base-by-base nucleic acid synthesis is used, then data storage density is improved, but error tolerance decreases

Engineering Contradiction:
Improvedata storage densityVSAvoidsequencing errors
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent segments data into multiple layers with different error tolerance characteristics. By distributing data across layers with varying encoding redundancy, the system ensures that errors in high-density layers do not catastrophically affect the entire dataset, as other layers provide error correction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies beforehand cushioning by incorporating error correction codes and redundancy encoding in certain layers before sequencing occurs. This preemptive error protection allows the system to tolerate sequencing errors that inevitably occur during high-density base-by-base synthesis and reading.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

4Measurement precision

If high accuracy sequencing is required for high density data, then data retrieval accuracy is improved, but cost increases

Engineering Contradiction:
Improvedata retrieval accuracyVSAvoidsequencing cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent segments data into layers with different accuracy requirements. By identifying which layers require high accuracy and which can tolerate lower accuracy, the system can apply cost-effective sequencing methods to appropriate layers, reducing overall sequencing cost while maintaining sufficient data retrieval accuracy for the application.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the required measurement precision parameter across different data layers based on their encoding schemes and error correction capabilities. Layers with strong error correction can be sequenced at lower precision thresholds, reducing the cost of high-accuracy sequencing while maintaining overall data retrieval accuracy through the combined information from multiple layers.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12437841B2Systems and methods for storing and reading nucleic acid-based data with error protection
Publication Date: 2025.10.07 BIOMEMORY AMERICA LLC
  • US12437841B2 patent drawing
  • US12437841B2 patent drawing
  • US12437841B2 patent drawing

AI summary

The systems, devices, and methods described herein provide scalable methods for writing data to and reading data from nucleic acid molecules. The present disclosure covers four primary areas of interest: (1) accurately and quickly reading information stored in nucleic acid molecules, (2) partitioning data to efficiently encode data in nucleic acid molecules, (3) error protection and correction when encoding data in nucleic acid molecules, and (4) data structures to provide efficient access to information stored in nucleic acid molecules.