Molecular Index Tag Sequencing for Error-Resolved Nucleic Acid Reads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for next-generation sequencing face challenges in distinguishing amplification errors from real SNPs or mutations, especially in complex samples like mammalian cDNA or genomic samples, due to the introduction of errors during sample preparation and base calling, which are not effectively addressed by current tagging methods that require large numbers of unique identifiers and result in reduced read lengths and high costs.
Innovation Solution
Utilizing Molecular Index Tags (MITs) to tag nucleic acid molecules, forming a reaction mixture with a specific ratio of MITs to sample nucleic acids, and attaching MITs to sample nucleic acid segments to create a library for sequencing, allowing differentiation of errors from real differences through sequencing and amplification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If unique identifiers are used to tag nucleic acid molecules, then amplification errors can be identified, but the cost increases and read lengths are reduced
Solution Approach 1:
The unique identifier is segmented into multiple components: a platform-specific sequence and a user-defined index. This segmentation allows the system to maintain error identification capability while reducing the overall length and cost of the tag, as the user only needs to provide a short index rather than a complete unique identifier.
Solution Approach 2:
The platform-specific sequence serves multiple functions: it enables error identification, facilitates data pooling from multiple samples, and works across different sequencing platforms. This multi-functionality reduces the need for additional dedicated tags, thereby reducing overall tag length and manufacturing cost while maintaining reliability.
2Quantity of substance
If more unique identifiers are generated, then more nucleic acid molecules can be tagged, but the cost increases and the amount of sample sequence that can be read decreases
Solution Approach 1:
By dividing the unique identifier into a fixed platform-specific sequence and a variable user index, the system can tag many more molecules without proportionally increasing read length consumption. The user index can be kept short (e.g., 4-12 bases) while the platform-specific sequence handles the bulk of the identification functionality.
Solution Approach 2:
The system changes the parameter of identifier structure from completely unique random sequences to a hybrid structure with fixed and variable portions. This parameter change allows efficient tagging of large numbers of molecules while consuming less sequencing read length, as the fixed portion is read once and reused across multiple samples.
3Reliability
If existing tagging methods are used, then amplification errors can be distinguished, but they are not effective for complex samples like mammalian cDNA or genomic samples
Solution Approach 1:
The tagging system is designed to be universally applicable across different sample types including simple templates, cDNA libraries, and complex genomic samples. The platform-specific sequence combined with user indexes provides a flexible framework that adapts to various sample complexities while maintaining error differentiation capability through consensus sequencing of identically tagged molecules.
Solution Approach 2:
The system dynamically adapts to different sample types by allowing users to adjust the index length and composition based on the complexity of their sample. For complex samples like mammalian cDNA or genomic DNA, users can employ longer or more diverse indexes to handle the increased diversity, while the core error identification mechanism remains the same.
Data Source
AI summary
The present disclosure provides methods and compositions for sequencing nucleic acid molecules and identifying individual sample nucleic acid molecules using Molecular Index Tags (MITs). Furthermore, reaction mixtures, kits, and adapter libraries are provided.


