Error-Correctible Molecular Barcodes for Sequencing Duplicate Discrimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Short-read sequencing platforms face challenges in distinguishing PCR duplicates from unique molecules at high coverage levels, leading to 'noise' and confounded quantification of sequencing data, particularly in transcriptome analysis where extensive correction for duplicate molecules is required.
Innovation Solution
A population of nucleic acid adaptors with at least 50,000 different molecular barcode sequences, designed for error correction, is used to tag DNA fragments, allowing for accurate sequencing and analysis even if barcode sequence reads contain errors, enabling discrimination of over a billion molecules and facilitating allele calling, copy number analysis, and gene expression estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high coverage sequencing is performed to confirm genetic variants, then detection sensitivity is improved, but the ability to distinguish PCR duplicates from unique molecules deteriorates
Solution Approach 1:
The patent applies preliminary action by incorporating unique molecular barcodes onto DNA fragments before PCR amplification. This allows each original molecule to be tagged with a unique identifier prior to duplication, enabling later distinction between true biological variants and PCR artifacts. The barcoding step is performed in advance, before the problematic PCR amplification that creates duplicates.
Solution Approach 2:
The patent introduces molecular barcodes as an intermediary element between the DNA fragment and the sequencing read. These barcodes serve as mediators that carry information about the original molecule's identity through the PCR amplification process, allowing bioinformatic tools to trace each read back to its unique source molecule and filter out duplicates.
2Quantity of substance
If random nucleotide barcodes are used to label DNA fragments, then barcode diversity and coverage are improved, but error detection and correction capabilities deteriorate
Solution Approach 1:
The patent applies parameter changes by transitioning from completely random nucleotide sequences to barcodes with structured composition. The barcodes incorporate specific features such as error correction codes, defined lengths, and controlled nucleotide distributions. This changes the parameters of the barcode design from pure randomness to structured randomness, maintaining diversity while adding error detection and correction capabilities.
Solution Approach 2:
The patent applies composite materials by creating barcodes that combine multiple functional elements within a single sequence. The barcodes integrate unique identifiers, error correction codes, and quality control features into a composite structure. This composite design allows the barcode to simultaneously provide diversity for molecule discrimination and reliability for error correction.
Data Source
AI summary
A population of nucleic acid adaptors is provided. In some embodiments, the population contains at least 50,000 different molecular barcode sequences, where the barcode sequences are double-stranded and at least 90% of the barcode sequences have an edit distance of at least 2. In certain cases, the adaptor may have an end in which the top and bottom strands are not complementary (i.e., may be in the form of a Y-adaptor). In some embodiments and depending on the how the adaptor is going to be employed, the other end of the adaptor may have a ligatable end or may be a transposon end sequence.


