UMI Tagging for Single-Stranded DNA Library Preparation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for preparing single-stranded DNA libraries for sequencing, such as those used in cancer detection, face challenges due to amplification biases and errors during PCR and sequencing, which obscure rare mutations, and fail to preserve duplex and connectivity information from double-stranded DNA fragments.
Innovation Solution
A method involving tagging both strands of double-stranded DNA fragments with unique sequence tags, allowing for identification and analysis of complementary strands, thereby correcting for errors and maintaining connectivity information through the use of partition-specific barcodes or UMIs during library preparation and sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If PCR amplification is used during library preparation, then the quantity of DNA is increased, but amplification biases and errors are introduced that reduce measurement precision
Solution Approach 1:
The patent applies preliminary action by tagging each dsDNA fragment with a unique molecular identifier (UMI) before PCR amplification occurs. This pre-tagging allows subsequent bioinformatic correction of amplification errors by comparing reads with the same UMI, thereby resolving the contradiction between needing sufficient DNA quantity through amplification and maintaining measurement precision despite amplification biases.
2Ease of manufacture
If single-stranded DNA library preparation is used, then the library can be prepared for sequencing, but connectivity information from the original double-stranded DNA fragments is lost
Solution Approach 1:
The patent performs preliminary tagging of both strands of dsDNA with identical UMIs before denaturation into single strands. This preliminary action ensures that even though the DNA is later converted to single-stranded form for library preparation, the connectivity information is preserved through the shared UMI that links complementary reads back to their original dsDNA fragment.
Solution Approach 2:
The patent uses the UMI as a digital copy or identifier that is replicated on both strands of the dsDNA fragment. This copying mechanism allows the connectivity information to be maintained in the sequence data without physically maintaining the double-stranded structure throughout the library preparation and sequencing process.
3Productivity
If standard sequencing protocols are used, then sequencing can be performed, but process-based errors obscure rare mutations
Solution Approach 1:
The patent implements feedback through the UMI system, where reads sharing the same UMI form a family that can be used to correct each other's errors. By comparing multiple reads with identical UMIs and consensus-building, the method provides feedback that eliminates process-based errors while maintaining high sequencing throughput, thereby resolving the contradiction between productivity and measurement precision.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances the accuracy of sequencing by correcting for amplification and sequencing errors, and preserves the connectivity information of DNA fragments, improving the detection of rare mutations and cancer-related variants.
Implementation Method 1
heating the droplets to denature the dsDNA or chemically denaturing the dsDNA to produce single-strand DNA (ssDNA) fragments
Implementation Method 2
chemically denaturing the dsDNA to produce single-strand DNA (ssDNA) fragments
Implementation Method 3
ligating the unique sequence tags to 3′ ends of the ssDNA fragments
Data Source
AI summary
Aspects of the invention relate to methods and compositions for preparing and analyzing a single-stranded sequencing library from a double-stranded DNA (e.g., double-stranded cfDNA) sample. In some embodiments, the sample includes double-stranded DNA (dsDNA) molecules, and damaged dsDNA (e.g., nicked dsDNA) molecules. In some embodiments, the sample includes single-stranded DNA (ssDNA) molecules. The subject methods facilitate the collection of information, including strand-pairing and connectivity information, from dsDNA, ssDNA and damaged DNA (e.g., nicked DNA) molecules in a sample, thereby providing enhanced diagnostic information as compared to sequencing libraries that are prepared using conventional methods.


