Sequencing Library Preparation for PCR-Error-Resistant RNA-Seq

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing RNA-seq methods rely on unique molecular identifiers (UMIs) that are susceptible to PCR errors, leading to inaccurate detection of PCR duplicates and reduced accuracy in RNA-seq analyses.

Innovation Solution

The method involves random fragmentation of nucleic acids and the use of short random sequences (N-mers) to create unique molecular identities in polynucleotides, resistant to PCR errors, by combining random cleavage locations and N-mers to ensure accurate detection of PCR duplicates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If unique molecular identifiers (UMIs) are used to detect PCR duplicates, then PCR duplicate detection capability is improved, but accuracy of RNA-seq analysis deteriorates due to PCR errors changing UMI sequences

Engineering Contradiction:
ImprovePCR duplicate detection capabilityVSAvoidaccuracy of RNA-seq analysis
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent divides the UMI sequence into multiple segments or components. Instead of using a single continuous UMI sequence that is vulnerable to PCR errors, the identification tag is segmented into multiple shorter sequences distributed throughout the polynucleotide. This segmentation reduces the impact of any single PCR error on the overall ability to identify and remove PCR duplicates, thereby resolving the contradiction between maintaining duplicate detection capability and ensuring analysis accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the identification tag by using multiple shorter sequences instead of one long sequence, and by positioning them at specific locations within the polynucleotide. This parameter change allows the system to maintain robust PCR duplicate detection while becoming more resistant to the effects of PCR errors, thus resolving the accuracy-reliability contradiction.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple sources of molecular diversity are combined to create unique identities, then resistance to PCR errors is improved, but complexity of library preparation deteriorates

Engineering Contradiction:
Improveresistance to PCR errorsVSAvoidcomplexity of library preparation
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple sources of molecular diversity into a unified identification system. By combining random fragmentation patterns with multiple distributed identification tags, the system creates a robust unique identity for each polynucleotide that is resistant to PCR errors. This merging approach achieves high reliability without proportionally increasing complexity, as the multiple diversity sources work together synergistically rather than additively.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a multi-functional identification system where the same structural features serve multiple purposes: random fragmentation provides both diversity and a basis for identification, while the distributed tags serve both as identifiers and as error-correcting elements. This multi-functionality reduces the need for separate components, thereby limiting the increase in preparation complexity while achieving high PCR error resistance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If random fragmentation and short random sequences are used to create unique molecular identities, then accuracy of PCR duplicate detection is improved, but cost of sequencing increases

Engineering Contradiction:
Improveaccuracy of PCR duplicate detectionVSAvoidsequencing cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses multiple shorter identification sequences instead of one extremely long unique identifier. This partial action approach provides sufficient uniqueness and accuracy for PCR duplicate detection without requiring excessive sequencing depth or length. The multiple shorter tags collectively provide the necessary information to accurately identify duplicates while keeping the total sequencing burden manageable, thus resolving the contradiction between detection accuracy and sequencing cost.

Inventive Principle:
Principle #16Partial or excessive action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach provides accurate and cost-effective sequencing libraries, especially for single-cell analyses, with reduced sequencing costs and labor, by creating libraries with distinct molecular identities that are resistant to PCR errors, enabling reliable RNA-seq data for personalized medicine.

Implementation Method 1

treating the RNA/DNA hybrids with enzymes that bind with RNA/DNA hybrids at random locations and integrate exogenous sequences at the randomly bound locations

Methodology Applied
Scientific EffectTransposase integration: Enzyme

Implementation Method 2

The fragments are reverse transcribed to make complementary DNA. Reverse transcription is preferably performed in the presence of oligos with short random N-mers (i.e., molecular diversity enhancers)

Methodology Applied
Scientific EffectReverse transcription: Enzyme

Data Source

PatentEP4240870B1Systems and methods for making sequencing libraries
Publication Date: 2025.12.31 ILLUMINA INC
  • EP4240870B1 patent drawingFigure 1
  • EP4240870B1 patent drawingFigure 2
  • EP4240870B1 patent drawingFigure 3

AI summary

This invention relates to systems and methods for making libraries of molecularly distinct polynucleotides. In particular, methods of the invention involve randomly fragmenting nucleic acids (e.g., RNA) to create fragments with cleaved ends at random cleavage locations. Preferably, methods also include reverse transcribing the fragments of RNA in the presence of molecular diversity enhancers (i.e., short random sequences), thereby creating polynucleotides with the molecular diversity enhancers copied therein. The result is a library of polynucleotides that are uniquely identifiable based on combinations of the random cleavage locations and molecular diversity enhancers.