Variable-Length Barcodes for Low-Frequency Mutation Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing nucleic acid sequencing methods face challenges in detecting mutations at low allele frequencies due to inherent errors and biases, which are not adequately addressed by current molecular barcoding techniques, leading to inaccurate identification of mutations and increased costs.
Innovation Solution
A method involving the use of a pool of oligonucleotides with varying lengths to differentially label nucleic acids, altering the read start and stop coordinates during sequencing, allowing for improved distinction between different starting molecules and accurate identification of mutations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If molecular barcodes of fixed length are used to label nucleic acids, then the number of possible tags increases, but sequencing efficiency decreases and phasing calculations cannot be made
Solution Approach 1:
The patent changes the parameter of barcode structure from fixed length to variable length. By using oligonucleotides with different lengths (e.g., 1-20 nucleotides) instead of uniform fixed-length barcodes, the system maintains high diversity while enabling effective phasing calculations during sequencing, thus resolving the contradiction between tag quantity and sequencing efficiency
Solution Approach 2:
The patent introduces length as an additional dimension for barcode differentiation. Instead of relying solely on sequence composition diversity, the system uses length variation as another dimension to create unique identifiers, which improves both diversity and sequencing performance by enabling accurate phasing
2Quantity of substance
If separate barcode synthesis reactions are performed to generate diverse barcodes, then tag diversity increases, but the process becomes costly and time-consuming
Solution Approach 1:
The patent merges the barcode synthesis step with the library preparation process. By incorporating variable-length oligonucleotides directly into the library preparation workflow rather than performing separate synthesis and pooling reactions, the system achieves high tag diversity while reducing the number of separate steps, time required, and associated costs
3Quantity of substance
If PCR amplification is used to increase nucleic acid quantity, then sufficient material for sequencing is obtained, but errors are introduced and propagated through the reaction
Solution Approach 1:
The patent applies preliminary error correction by using unique molecular identifiers (UMIs) with variable lengths attached to each original nucleic acid molecule before PCR amplification. This allows bioinformatic tools to trace back amplified reads to their original templates and correct for PCR-introduced errors, maintaining reliability while enabling necessary amplification
Solution Approach 2:
The patent uses variable-length UMIs as molecular copies or fingerprints of the original nucleic acid molecules. By attaching these unique length-based identifiers before amplification, the system creates a record of the original molecule that can be used to distinguish true variants from PCR errors during data analysis
Data Source
AI summary
A method of sequencing a nucleic acid of interest (NAOI) is provided. In some embodiments, the method may comprise providing a sample from a patient, the sample having cell-free DNA molecules comprising NAOIs, amplifying the NAOIs, ligating adapters to the NAOIs, and sequencing the NAOIs. A method of labelling a NAOIs is also provided. In some embodiments, the method may comprise providing a sample from a patient having the NAOIs, amplifying the NAOIs, contacting the amplified NAOIs with a pool of oligonucleotides, attaching oligonucleotides to each of the amplified NAOIs to label the NAOIs. The method may further comprise sequencing the labelled NAOIs and grouping the resulting sequencing reads.


