A DNA storage readout method based on multiple hidden reference sequences

By employing multiple hidden reference sequences and a multi-stage forward-backward algorithm, the problem of high-complexity de novo assembly in DNA storage is solved, achieving highly reliable data recovery under low coverage and reducing the complexity and cost of data recovery in DNA storage.

CN120596016BActive Publication Date: 2026-06-26TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2025-05-27
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing DNA storage technologies face highly complex de novo assembly problems when recovering large fragments of data. In particular, third-generation nanopore sequencing technology has a high sequencing error rate, while second-generation high-throughput sequencing technology suffers from read fragmentation, which increases operational complexity. Traditional methods fail to locate errors in the event of insertion/deletion.

Method used

We employ a multiple hidden reference sequence approach, identify low-error reads through sliding correlation, construct a multi-stage forward-backward algorithm and soft-decision decoding, and combine decoding feedback reference sequences to correct insertion/deletion errors, thus transforming it into a low-complexity resequencing problem.

Benefits of technology

Error-free recovery of original data under low sequencing coverage, phased correction of insertion/deletion errors, providing highly reliable data readout, and reducing the complexity and cost of data recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596016B_ABST
    Figure CN120596016B_ABST
Patent Text Reader

Abstract

The application discloses a DNA storage reading method based on multiple hidden reference sequences, which constructs multiple hidden reference sequences to realize reliable recovery of large fragment DNA from a read containing insertion / deletion errors; the constructed hidden reference sequence changes the reading of large fragment DNA storage from scratch into a low complexity resequencing problem; wherein, the watermark reference sequence quickly identifies low error reads, the skeleton reference sequence identifies reads containing insertion / deletion errors, and the decoding feedback reference sequence identifies reads for filling low coverage areas; the forward-backward algorithm of each read corrects the insertion / deletion errors in stages to realize reliable data recovery. The application realizes reliable data reading from reads containing insertion / deletion errors, and has the advantages that the constructed multiple hidden reference sequences identify reads with complementary characteristics of error rates in stages, the forward-backward algorithm of each read corrects the insertion / deletion errors in stages, and gradual reading is realized.
Need to check novelty before this filing date? Find Prior Art