Variable-Length Nonrandom UMIs for Low-Frequency Variant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing technologies face challenges in accurately detecting sequences of low allele frequency, such as fetal cell-free DNA and circulating tumor DNA, due to errors and noise from various sources, which conventional deep sequencing cannot overcome.
Innovation Solution
The use of variable-length, nonrandom unique molecular indices (vNRUMIs) in sequencing adapters, which are designed to identify individual nucleic acid molecules by forming a set of vNRUMIs with an edit distance of at least two, allowing for error suppression and accurate sequencing of low-frequency sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If deep sequencing is performed to detect low allele frequency sequences, then sequencing depth increases, but sequencing errors and noise from various sources also increase
Solution Approach 1:
The patent applies preliminary action by incorporating unique molecular indices (UMIs) into the sequencing adapters before the sequencing process begins. These UMIs are randomly generated sequences that are ligated to each DNA fragment during library preparation, creating a unique identifier for each original molecule. This preliminary tagging allows subsequent bioinformatic processing to group reads by their UMI, enabling consensus building and error suppression while maintaining the benefits of deep sequencing coverage.
2Measurement precision
If conventional sequencing adapters are used, then the sequencing process is simple, but the detection accuracy of low-frequency variants is insufficient
Solution Approach 1:
The patent applies segmentation by dividing the adapter structure into distinct functional segments: a universal binding region for DNA fragment attachment and a variable UMI region for unique identification. The adapter is segmented such that different adapters can share the universal binding region while having different UMI sequences, allowing modular design and simplifying the overall system complexity despite the enhanced functionality.
Solution Approach 2:
The patent applies parameter changes by varying the UMI sequence parameters (nucleotide composition, length, randomness) to optimize detection accuracy. The UMIs are designed with specific probabilistic properties to ensure uniqueness while maintaining randomness, and the patent explores different UMI length parameters (e.g., 8-12 nucleotides) to balance uniqueness with sequencing error tolerance, thereby improving low-frequency variant detection without excessive complexity.
Data Source
AI summary
The disclosed embodiments concern methods, systems and computer program products for determining sequences of interest using unique molecular indexes (UMIs) that are uniquely associable with individual polynucleotide fragments, including sequences with low allele frequencies or long sequence length. In some implementations, the UMIs include variable-length nonrandom UMIs (vNRUMIs). Methods and systems for making and using sequencing adapters comprising vNRUMIs are also provided.


