Sequencing Library Preparation Using Intrinsic Molecular Identifiers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for sequencing and quantifying mRNA in single cells are time-consuming and require additional steps and resources, such as the use of synthetic unique molecular identifiers (UMIs), which can consume sequencing real estate and complicate the process.

Innovation Solution

The use of intrinsic molecular identifiers (IMIs) generated by random cleavage or priming of nucleic acids, providing a unique label within the nucleic acid molecule, allowing for deduplication of sequence reads and accurate quantification of mRNA transcripts without the need for UMIs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If synthetic unique molecular identifiers (UMIs) are attached to all nucleic acid molecules, then quantitative measurement of nucleic acids is enabled, but the process becomes more complex and consumes additional sequencing capacity

Engineering Contradiction:
Improvequantitative measurement accuracyVSAvoidlibrary preparation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The invention extracts and utilizes the intrinsic sequence information already present in each nucleic acid molecule as its unique identifier. Instead of adding external UMIs, the method takes out a segment of the molecule's own sequence (typically 5-20 base pairs) to serve as the identifier, thereby eliminating the need for additional synthetic tags and reducing library preparation complexity while maintaining quantitative measurement accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Each nucleic acid molecule serves itself by using its own intrinsic sequence as the unique identifier. The molecule's natural sequence variation provides the distinguishing feature needed for deduplication, eliminating the need for external labeling services and reducing the overall complexity of the quantification system

Inventive Principle:
Principle #25Self-service

2Reliability

If synthetic UMIs are used for deduplication, then sequence read deduplication is achieved, but sequencing real estate is consumed and resources are increased

Engineering Contradiction:
Improvededuplication accuracyVSAvoidsequencing capacity requirement
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The method extracts a segment of intrinsic sequence from each nucleic acid molecule to serve as the identifier. This extracted sequence segment replaces the need for separate UMI sequences, thereby conserving sequencing capacity while maintaining the ability to uniquely identify and deduplicate reads from the same original molecule

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention merges the function of the unique identifier with the nucleic acid sequence itself. The intrinsic sequence segment serves both as part of the molecule's identity and as the deduplication key, combining multiple functions into a single element and reducing the total amount of sequencing data needed

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If traditional sequencing library preparation methods are used, then nucleic acids can be sequenced, but the process takes days or weeks to complete

Engineering Contradiction:
Improvenucleic acid quantificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

By extracting intrinsic sequence information directly from the nucleic acid molecules without requiring conversion to cDNA or addition of synthetic UMIs, the method streamlines the library preparation process. This direct utilization of intrinsic sequences reduces the number of processing steps and enables faster turnaround from sample to quantification result

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The method performs the identification and labeling function during the initial library preparation steps by utilizing intrinsic sequences that are already present in the nucleic acids. This preliminary action eliminates the need for subsequent complex processing steps that would otherwise be required to add and manage synthetic UMIs, thereby reducing overall processing time

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables rapid, sensitive, and specific quantification of mRNA transcripts in single cells by using intrinsic labels, reducing processing time and resource consumption while maintaining accuracy and scalability.

Implementation Method 1

random cleavage or priming of sample nucleic acids, at a random cutting or priming site

Methodology Applied
Scientific EffectRandom cleavage:

Implementation Method 2

The transposase cleaves the DNA (e.g., at a random site) to generate the identifier sequence

Methodology Applied
Scientific EffectTransposase cleavage: Enzyme

Implementation Method 3

Adaptors or PCR handles are attached at the random sites and library preparation steps, such as amplification or sequencing proceed by annealing primers to those PCR handles

Methodology Applied
Scientific EffectHybridization:

Implementation Method 4

Amplification provides sequencing libraries in which a sequencing primer anneals in the adaptor or PCR handle

Methodology Applied
Scientific EffectPCR amplification: Enzyme

Data Source

PatentUS12534721B2Quantitative detection and analysis of moleculesip
Publication Date: 2026.01.27 ILLUMINA INC
  • US12534721B2 patent drawing
  • US12534721B2 patent drawing
  • US12534721B2 patent drawing

AI summary

The invention provides systems and methods for making sequencing libraries that are useful for quantitatively analyzing nucleic acids in a sample. Sample nucleic acids are randomly cleaved at, and PCR handled are attached to, a random cut site. The nucleic acid is amplified into a sequencing library in which a sequencing primer generates a sequence read from adjacent the random cut site. The sequence reads can be mapped to a reference, but they will also include a unique identifier sequence that comes from within the nucleic acid molecule being analyzed, i.e., an intrinsic molecular identifier (IMI). The IMI is unique for each molecule and can thus be used to deduplicate sequence reads originating from the same molecule.