Molecular Index Tag Sequencing for Error-Resolved Nucleic Acid Reads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for next-generation sequencing face challenges in distinguishing amplification errors from real SNPs or mutations, especially in complex samples like mammalian cDNA or genomic samples, due to the introduction of errors during sample preparation and base calling, which are not effectively addressed by current tagging methods that require large numbers of unique identifiers and result in reduced read lengths and high costs.

Innovation Solution

Utilizing Molecular Index Tags (MITs) to tag nucleic acid molecules, forming a reaction mixture with a specific ratio of MITs to sample nucleic acids, and attaching MITs to sample nucleic acid segments to create a library for sequencing, allowing differentiation of errors from real differences through sequencing and amplification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If unique identifiers are used to tag nucleic acid molecules, then amplification errors can be identified, but the cost increases and read lengths are reduced

Engineering Contradiction:
Improveerror identification accuracyVSAvoidmanufacturing cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The unique identifier is segmented into multiple components: a platform-specific sequence and a user-defined index. This segmentation allows the system to maintain error identification capability while reducing the overall length and cost of the tag, as the user only needs to provide a short index rather than a complete unique identifier.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The platform-specific sequence serves multiple functions: it enables error identification, facilitates data pooling from multiple samples, and works across different sequencing platforms. This multi-functionality reduces the need for additional dedicated tags, thereby reducing overall tag length and manufacturing cost while maintaining reliability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If more unique identifiers are generated, then more nucleic acid molecules can be tagged, but the cost increases and the amount of sample sequence that can be read decreases

Engineering Contradiction:
Improvenumber of taggable moleculesVSAvoidread length
Core Design Contradiction:
Quantity of substanceVSLength of moving object

Solution Approach 1:

By dividing the unique identifier into a fixed platform-specific sequence and a variable user index, the system can tag many more molecules without proportionally increasing read length consumption. The user index can be kept short (e.g., 4-12 bases) while the platform-specific sequence handles the bulk of the identification functionality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of identifier structure from completely unique random sequences to a hybrid structure with fixed and variable portions. This parameter change allows efficient tagging of large numbers of molecules while consuming less sequencing read length, as the fixed portion is read once and reused across multiple samples.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If existing tagging methods are used, then amplification errors can be distinguished, but they are not effective for complex samples like mammalian cDNA or genomic samples

Engineering Contradiction:
Improveerror differentiation capabilityVSAvoidapplicability to complex samples
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The tagging system is designed to be universally applicable across different sample types including simple templates, cDNA libraries, and complex genomic samples. The platform-specific sequence combined with user indexes provides a flexible framework that adapts to various sample complexities while maintaining error differentiation capability through consensus sequencing of identically tagged molecules.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adapts to different sample types by allowing users to adjust the index length and composition based on the complexity of their sample. For complex samples like mammalian cDNA or genomic DNA, users can employ longer or more diverse indexes to handle the increased diversity, while the core error identification mechanism remains the same.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12571034B2Compositions and methods for identifying nucleic acid molecules
Publication Date: 2026.03.10 NATERA INC
  • US12571034B2 patent drawing
  • US12571034B2 patent drawing
  • US12571034B2 patent drawing

AI summary

The present disclosure provides methods and compositions for sequencing nucleic acid molecules and identifying individual sample nucleic acid molecules using Molecular Index Tags (MITs). Furthermore, reaction mixtures, kits, and adapter libraries are provided.