A DNA storage readout method based on multiple hidden reference sequences
By employing multiple hidden reference sequences and a multi-stage forward-backward algorithm, the problem of high-complexity de novo assembly in DNA storage is solved, achieving highly reliable data recovery under low coverage and reducing the complexity and cost of data recovery in DNA storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2025-05-27
- Publication Date
- 2026-06-26
AI Technical Summary
Existing DNA storage technologies face highly complex de novo assembly problems when recovering large fragments of data. In particular, third-generation nanopore sequencing technology has a high sequencing error rate, while second-generation high-throughput sequencing technology suffers from read fragmentation, which increases operational complexity. Traditional methods fail to locate errors in the event of insertion/deletion.
We employ a multiple hidden reference sequence approach, identify low-error reads through sliding correlation, construct a multi-stage forward-backward algorithm and soft-decision decoding, and combine decoding feedback reference sequences to correct insertion/deletion errors, thus transforming it into a low-complexity resequencing problem.
Error-free recovery of original data under low sequencing coverage, phased correction of insertion/deletion errors, providing highly reliable data readout, and reducing the complexity and cost of data recovery.
Smart Images

Figure CN120596016B_ABST