Variable-Length Oligonucleotide Labels for Low-Frequency Mutation Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing (NGS) methods face challenges in detecting low allelic frequency mutations due to inherent error and amplification biases, leading to inefficient sequencing and increased costs when using molecular barcodes, and existing error correction methods fail to distinguish true variants from sequencing errors.
Innovation Solution
A method involving the use of a pool of oligonucleotides with varying lengths to label nucleic acids, altering the read start and stop coordinates during sequencing, allowing for the differentiation of true mutations from sequencing errors by introducing synthetic breakpoints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If molecular barcodes with fixed length and high diversity are used to label nucleic acids, then the ability to detect low allelic frequency mutations is improved, but the cost and time required for barcode synthesis and pooling increases significantly
Solution Approach 1:
The patent changes the parameter of barcode length from fixed to variable, using barcodes ranging from 50-200 bases. This variable length approach, combined with controlled diversity (10-1000 different barcodes), reduces the complexity of synthesis and pooling while maintaining sufficient unique identifiers for detecting low allelic frequency mutations
Solution Approach 2:
Instead of using extremely high diversity (>100,000 combinations) as in conventional methods, the patent uses a partial approach with 10-1000 barcodes of variable length, which provides sufficient diversity for the application while significantly reducing synthesis complexity and cost
2Ease of manufacture
If fixed length molecular barcodes are used, then synthesis is simpler, but sequencing efficiency decreases because NGS/Illumina phasing calculations cannot be made
Solution Approach 1:
The patent changes the barcode length parameter from fixed to variable (50-200 bases), which enables NGS/Illumina phasing calculations to be performed while still maintaining relatively simple synthesis procedures. The variable length provides the necessary information for sequencing efficiency without requiring excessively complex barcode designs
3Reliability
If the sample is split into multiple replicate processing steps for error identification, then error detection capability is improved, but costs and complexity increase and assay sensitivity decreases
Solution Approach 1:
The patent applies preliminary action by incorporating unique variable length barcodes onto individual nucleic acid molecules before the main sequencing process. This pre-labeling allows error detection through consensus calling of barcode-associated reads without requiring multiple replicate processing steps, thereby reducing complexity while maintaining error detection capability
Solution Approach 2:
The patent uses barcode copying where each nucleic acid molecule is tagged with a unique barcode that is then copied and amplified along with the target molecule. This allows the original molecule's identity to be tracked through multiple amplification cycles, enabling error detection without splitting the sample into replicates
4Measurement precision
If high diversity barcodes are used to ensure unique labeling of each nucleic acid, then mutation detection accuracy is improved, but the process becomes costly and time-consuming
Solution Approach 1:
The patent changes the barcode parameters to variable length (50-200 bases) with moderate diversity (10-1000 variants). This provides sufficient unique labeling capability while significantly reducing synthesis and pooling time compared to conventional high diversity approaches, as fewer distinct barcode sequences need to be synthesized and combined
Data Source
AI summary
A method of labelling a nucleic acid of interest (NAOI) is provided. In some embodiments, the method may comprise contacting a sample comprising the nucleic acid of interest with a pool of oligonucleotides, the pool comprising oligonucleotides having at least 5 different lengths; and attaching an oligonucleotide from the pool on to one or each end of the nucleic acid of interest, wherein attachment of an oligonucleotide moves the read start and/or stop coordinate when the labelled NAOI is sequenced.


