Polynucleotide Data Storage Random Access via Primer-Guided Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage technologies are unable to keep pace with exponentially growing data volumes, and existing DNA data storage methods face challenges in coding and random access, limiting the efficient retrieval of digital data encoded by polynucleotides.

Innovation Solution

A framework for encoding digital data using polynucleotides that includes segmenting data into segments, associating each segment with unique group identifiers, and using universal and group-specific primers to enable simultaneous amplification and sequencing for random access, allowing selective retrieval of requested data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional data storage technologies are used, then current storage capacity is maintained, but they cannot keep pace with exponentially growing amounts of data

Engineering Contradiction:
Improvedata storage capacityVSAvoiddata growth rate
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent replaces conventional mechanical/electronic storage systems with a biochemical system using polynucleotides (DNA/RNA) as the storage medium. This substitution enables exponentially higher data density, allowing the system to keep pace with rapidly growing data volumes that traditional storage technologies cannot handle.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If DNA data storage methods are used, then high information density is achieved, but challenges in coding and random access limit efficient retrieval

Engineering Contradiction:
Improveinformation densityVSAvoidrandom access efficiency
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent segments digital data into multiple segments before encoding them into polynucleotides. Each segment is associated with unique group identifiers, allowing the system to retrieve specific segments through targeted amplification and sequencing operations, thereby enabling efficient random access to high-density stored data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces group identifiers as intermediary elements that link digital data segments to their polynucleotide encodings. These identifiers serve as mediators that enable selective retrieval through primer-based amplification, solving the random access problem in high-density DNA storage systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If universal sequences are used for whole pool amplification, then all polynucleotides can be copied efficiently, but selective access to specific data requires additional group identifier regions

Engineering Contradiction:
Improveamplification efficiencyVSAvoidsequence structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a nested sequence structure where group identifier regions are embedded within universal sequences. This nesting allows the system to maintain efficient whole-pool amplification capabilities while simultaneously enabling selective access to specific data segments through the embedded group identifiers, without requiring separate structural elements.

Inventive Principle:
Principle #7Nested doll (Nesting)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach minimizes inefficiencies in data retrieval by enabling rapid, efficient random access to digital data stored in polynucleotides, overcoming the limitations of destructive sequencing methods and ensuring high data integrity.

Implementation Method 1

Random-access via PCR or other methods selects only those files that need to be sequenced

Methodology Applied
Scientific EffectPCR amplification:

Implementation Method 2

the polynucleotides have universal sequences that correspond to primers that can be used to amplify and replicate or copy the whole pool of polynucleotides in a storage container

Methodology Applied
Scientific EffectPolynucleotide replication:

Implementation Method 3

nucleotide sequencing is used to facilitate random access of the selected sequences

Methodology Applied
Scientific EffectNucleotide sequencing:

Data Source

PatentUS20250356950A1Whole pool amplification and in-sequencer random-access of data encoded by polynucleotides
Publication Date: 2025.11.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250356950A1 patent drawing
  • US20250356950A1 patent drawing
  • US20250356950A1 patent drawing

AI summary

This disclosure describes an efficient method to copy all polynucleotides encoding digital data of digital files in a polynucleotide storage container while maintaining random access capabilities over a collection of files or data items in the container. The disclosure further describes a process whereby random-access and sequencing of the polynucleotides are combined in a single step.