DNA Data Encoding with Fountain Codes for Error-Tolerant Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing systems face challenges in efficiently encoding and decoding data for storage and transmission, particularly in formats that allow for error correction and recovery, especially in distributed systems.

Innovation Solution

A computing system utilizing a fountain code process to segment data into blocks, generate seed data, and synthesize polynucleotide strands for encoding and decoding data in genetic materials like DNA/RNA, incorporating error correction mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Duration of action of stationary object

If data is encoded into genetic materials for storage, then data retention duration is improved, but device complexity and manufacturing difficulty increase

Engineering Contradiction:
Improvedata retention durationVSAvoidencoding system complexity
Core Design Contradiction:
Duration of action of stationary objectVSDevice complexity

Solution Approach 1:

The patent segments data into multiple data blocks, each independently encoded into separate polynucleotide strands using fountain codes. This segmentation allows the data to be distributed across multiple genetic material units, improving durability and retrieval flexibility while managing the complexity of genetic material encoding through modular processing.

Inventive Principle:
Principle #1Segmentation

2Reliability

If error correction mechanisms are implemented in data encoding, then data reliability is improved, but manufacturing precision requirements increase

Engineering Contradiction:
Improvedata integrityVSAvoidsynthesis accuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent implements fountain code encoding that proactively generates redundant data packets before storage. These redundant packets serve as a cushion against potential data loss or corruption during synthesis and storage, allowing error correction without requiring extremely high manufacturing precision during the encoding process itself.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The patent utilizes the four chemical bases of DNA (A, C, G, T) as encoding parameters, where each base can represent two bits of information. This parameter system provides inherent redundancy and error detection capabilities, as the biological system naturally handles variations and mutations, reducing the burden on manufacturing precision while maintaining data reliability.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If data is segmented and encoded with metadata for each block, then data recovery capability is improved, but processing time increases

Engineering Contradiction:
Improvedata recovery capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary segmentation and metadata attachment during the encoding phase. Each data block is pre-tagged with metadata indicating its position, size, and relationships, enabling rapid identification and retrieval during decoding without requiring complex real-time analysis, thus reducing processing time during data recovery operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250390761A1Apparatus and methods for embedding data in genetic material
Publication Date: 2025.12.25 CUSTOMARRAY INC
  • US20250390761A1 patent drawing
  • US20250390761A1 patent drawing
  • US20250390761A1 patent drawing

AI summary

Methods, systems, and apparatuses to encode data for storage in genetic materials. For example, a computing system may segment user data into a plurality of data blocks and generate seed data characterizing a plurality of fountain code seeds. Additionally, the computing system may, for each data block, implement a set of operations that generate one or more data packets. In some instances, the set of operations may include, for each of the plurality of fountain code seeds, determining a bit value and corresponding metaCode value and determining which of the fountain code seeds has a metaCode value of the bit value that matches a value of the bit position identified in the metadata. Moreover, the computing system may, for each data packet, cause an implementation of a second set of operations that synthesize a polynucleotide strand in accordance with at least bit values of the corresponding data packet.