Sequencing Library Preparation for PCR-Error-Resistant RNA-Seq
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing RNA-seq methods rely on unique molecular identifiers (UMIs) that are susceptible to PCR errors, leading to inaccurate detection of PCR duplicates and reduced accuracy in RNA-seq analyses.
Innovation Solution
The method involves random fragmentation of nucleic acids and the use of short random sequences (N-mers) to create unique molecular identities in polynucleotides, resistant to PCR errors, by combining random cleavage locations and N-mers to ensure accurate detection of PCR duplicates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If unique molecular identifiers (UMIs) are used to detect PCR duplicates, then PCR duplicate detection capability is improved, but accuracy of RNA-seq analysis deteriorates due to PCR errors changing UMI sequences
Solution Approach 1:
The patent divides the UMI sequence into multiple segments or components. Instead of using a single continuous UMI sequence that is vulnerable to PCR errors, the identification tag is segmented into multiple shorter sequences distributed throughout the polynucleotide. This segmentation reduces the impact of any single PCR error on the overall ability to identify and remove PCR duplicates, thereby resolving the contradiction between maintaining duplicate detection capability and ensuring analysis accuracy.
Solution Approach 2:
The patent changes the parameters of the identification tag by using multiple shorter sequences instead of one long sequence, and by positioning them at specific locations within the polynucleotide. This parameter change allows the system to maintain robust PCR duplicate detection while becoming more resistant to the effects of PCR errors, thus resolving the accuracy-reliability contradiction.
2Reliability
If multiple sources of molecular diversity are combined to create unique identities, then resistance to PCR errors is improved, but complexity of library preparation deteriorates
Solution Approach 1:
The patent merges multiple sources of molecular diversity into a unified identification system. By combining random fragmentation patterns with multiple distributed identification tags, the system creates a robust unique identity for each polynucleotide that is resistant to PCR errors. This merging approach achieves high reliability without proportionally increasing complexity, as the multiple diversity sources work together synergistically rather than additively.
Solution Approach 2:
The patent implements a multi-functional identification system where the same structural features serve multiple purposes: random fragmentation provides both diversity and a basis for identification, while the distributed tags serve both as identifiers and as error-correcting elements. This multi-functionality reduces the need for separate components, thereby limiting the increase in preparation complexity while achieving high PCR error resistance.
3Measurement precision
If random fragmentation and short random sequences are used to create unique molecular identities, then accuracy of PCR duplicate detection is improved, but cost of sequencing increases
Solution Approach 1:
The patent uses multiple shorter identification sequences instead of one extremely long unique identifier. This partial action approach provides sufficient uniqueness and accuracy for PCR duplicate detection without requiring excessive sequencing depth or length. The multiple shorter tags collectively provide the necessary information to accurately identify duplicates while keeping the total sequencing burden manageable, thus resolving the contradiction between detection accuracy and sequencing cost.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach provides accurate and cost-effective sequencing libraries, especially for single-cell analyses, with reduced sequencing costs and labor, by creating libraries with distinct molecular identities that are resistant to PCR errors, enabling reliable RNA-seq data for personalized medicine.
Implementation Method 1
treating the RNA/DNA hybrids with enzymes that bind with RNA/DNA hybrids at random locations and integrate exogenous sequences at the randomly bound locations
Implementation Method 2
The fragments are reverse transcribed to make complementary DNA. Reverse transcription is preferably performed in the presence of oligos with short random N-mers (i.e., molecular diversity enhancers)
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This invention relates to systems and methods for making libraries of molecularly distinct polynucleotides. In particular, methods of the invention involve randomly fragmenting nucleic acids (e.g., RNA) to create fragments with cleaved ends at random cleavage locations. Preferably, methods also include reverse transcribing the fragments of RNA in the presence of molecular diversity enhancers (i.e., short random sequences), thereby creating polynucleotides with the molecular diversity enhancers copied therein. The result is a library of polynucleotides that are uniquely identifiable based on combinations of the random cleavage locations and molecular diversity enhancers.