UMI Ligation for PCR Duplicate Detection in RNA Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In RNA sequencing, accurate gene expression measurements are hindered by PCR duplicate artifacts, making it difficult to distinguish between unique cDNA molecules and PCR duplicates, which increases sequencing costs and reduces data quality due to the generation of unusable data reads.
Innovation Solution
The method involves ligating an adaptor with an indexing primer binding site, an indexing site, an identifier site, and a target sequence primer binding site to nucleic acid fragments, allowing for the detection and removal of duplicate sequencing reads by identifying and separating sequencing reads with a duplicate identifier site and target sequence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If paired end sequencing is used to distinguish PCR duplicates, then duplicate detection accuracy is improved, but sequencing costs increase
Solution Approach 1:
The patent extracts the duplicate detection function from the sequencing process itself by using unique molecular identifiers (UMIs) that are ligated to nucleic acid fragments before amplification. These UMIs are then detected separately during data analysis, allowing duplicate identification without requiring paired-end sequencing. This separates the detection function from the main sequencing workflow, reducing costs while maintaining accuracy.
Solution Approach 2:
The patent introduces unique molecular identifiers (UMIs) as intermediary elements that are ligated to nucleic acid fragments before PCR amplification. These UMIs serve as mediators that track individual molecules through the amplification process, enabling duplicate detection by comparing UMI sequences rather than relying on sequencing read positions. This intermediary approach provides accurate duplicate detection without the cost of paired-end sequencing.
2Measurement precision
If paired end sequencing is performed to identify PCR duplicates, then data quality is improved, but productivity decreases due to instrument inefficiency
Solution Approach 1:
The patent extracts the duplicate identification requirement from the sequencing instrumentation by implementing UMI-based detection in the data analysis phase. This allows standard high-throughput sequencers to operate at full capacity without the need for resource-intensive paired-end sequencing, thereby maintaining productivity while ensuring data quality through accurate duplicate detection.
Solution Approach 2:
By introducing UMIs as intermediary markers ligated before amplification, the patent enables duplicate detection through simple sequence comparison in bioinformatics pipelines. This approach maintains high sequencing throughput because it does not require additional sequencing cycles or specialized instrument configurations, unlike paired-end sequencing.
3Ease of manufacture
If sequencing reads with identical starting positions are assumed to be duplicates, then duplicate removal is simplified, but measurement precision deteriorates due to false positives
Solution Approach 1:
The patent uses unique molecular identifiers (UMIs) as intermediary markers that are ligated to each nucleic acid fragment before amplification. These UMIs serve as unique tags for individual molecules, allowing accurate duplicate identification by comparing UMI sequences rather than relying on sequencing read starting positions. This resolves the false positive problem while maintaining ease of implementation through straightforward bioinformatics comparison.
Solution Approach 2:
The patent changes the parameter used for duplicate identification from sequencing read starting positions to unique molecular identifier (UMI) sequences. This parameter change eliminates false positives caused by random identical starting positions while maintaining simple duplicate removal logic through direct UMI sequence comparison in data analysis pipelines.
Data Source
AI summary
The present invention provides methods, compositions and kits for detecting duplicate sequencing reads. In some embodiments, the duplicate sequencing reads are removed.


