Sequencing Library Preparation Using Intrinsic Molecular Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for sequencing and quantifying mRNA in single cells are time-consuming and require additional steps and resources, such as the use of synthetic unique molecular identifiers (UMIs), which can consume sequencing real estate and complicate the process.
Innovation Solution
The use of intrinsic molecular identifiers (IMIs) generated by random cleavage or priming of nucleic acids, providing a unique label within the nucleic acid molecule, allowing for deduplication of sequence reads and accurate quantification of mRNA transcripts without the need for UMIs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If synthetic unique molecular identifiers (UMIs) are attached to all nucleic acid molecules, then quantitative measurement of nucleic acids is enabled, but the process becomes more complex and consumes additional sequencing capacity
Solution Approach 1:
The invention extracts and utilizes the intrinsic sequence information already present in each nucleic acid molecule as its unique identifier. Instead of adding external UMIs, the method takes out a segment of the molecule's own sequence (typically 5-20 base pairs) to serve as the identifier, thereby eliminating the need for additional synthetic tags and reducing library preparation complexity while maintaining quantitative measurement accuracy
Solution Approach 2:
Each nucleic acid molecule serves itself by using its own intrinsic sequence as the unique identifier. The molecule's natural sequence variation provides the distinguishing feature needed for deduplication, eliminating the need for external labeling services and reducing the overall complexity of the quantification system
2Reliability
If synthetic UMIs are used for deduplication, then sequence read deduplication is achieved, but sequencing real estate is consumed and resources are increased
Solution Approach 1:
The method extracts a segment of intrinsic sequence from each nucleic acid molecule to serve as the identifier. This extracted sequence segment replaces the need for separate UMI sequences, thereby conserving sequencing capacity while maintaining the ability to uniquely identify and deduplicate reads from the same original molecule
Solution Approach 2:
The invention merges the function of the unique identifier with the nucleic acid sequence itself. The intrinsic sequence segment serves both as part of the molecule's identity and as the deduplication key, combining multiple functions into a single element and reducing the total amount of sequencing data needed
3Measurement precision
If traditional sequencing library preparation methods are used, then nucleic acids can be sequenced, but the process takes days or weeks to complete
Solution Approach 1:
By extracting intrinsic sequence information directly from the nucleic acid molecules without requiring conversion to cDNA or addition of synthetic UMIs, the method streamlines the library preparation process. This direct utilization of intrinsic sequences reduces the number of processing steps and enables faster turnaround from sample to quantification result
Solution Approach 2:
The method performs the identification and labeling function during the initial library preparation steps by utilizing intrinsic sequences that are already present in the nucleic acids. This preliminary action eliminates the need for subsequent complex processing steps that would otherwise be required to add and manage synthetic UMIs, thereby reducing overall processing time
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables rapid, sensitive, and specific quantification of mRNA transcripts in single cells by using intrinsic labels, reducing processing time and resource consumption while maintaining accuracy and scalability.
Implementation Method 1
random cleavage or priming of sample nucleic acids, at a random cutting or priming site
Implementation Method 2
The transposase cleaves the DNA (e.g., at a random site) to generate the identifier sequence
Implementation Method 3
Adaptors or PCR handles are attached at the random sites and library preparation steps, such as amplification or sequencing proceed by annealing primers to those PCR handles
Implementation Method 4
Amplification provides sequencing libraries in which a sequencing primer anneals in the adaptor or PCR handle
Data Source
AI summary
The invention provides systems and methods for making sequencing libraries that are useful for quantitatively analyzing nucleic acids in a sample. Sample nucleic acids are randomly cleaved at, and PCR handled are attached to, a random cut site. The nucleic acid is amplified into a sequencing library in which a sequencing primer generates a sequence read from adjacent the random cut site. The sequence reads can be mapped to a reference, but they will also include a unique identifier sequence that comes from within the nucleic acid molecule being analyzed, i.e., an intrinsic molecular identifier (IMI). The IMI is unique for each molecule and can thus be used to deduplicate sequence reads originating from the same molecule.


