Consensus Sequence Determination Using Anchor Segments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-throughput DNA sequencing methods generate vast amounts of data with high error rates, making it difficult to accurately identify majority and minority species, especially when they differ minimally, and to determine their prevalence.
Innovation Solution
A computer-implemented method that uses an anchor segment of known sequence to evaluate the accuracy of sequencing reads, assign them to accepted or rejected classes, and iteratively poll nucleobases to determine a consensus sequence for the adjacent unknown segment, allowing for the correction and refinement of sequencing errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If high-throughput sequencing methods are used to generate vast amounts of sequencing data, then productivity increases, but sequencing error rate increases
Solution Approach 1:
The patent combines multiple sequencing reads that cover the same nucleic acid target into a consensus sequence. By merging information from multiple reads, the method leverages the high throughput advantage while compensating for individual read errors through collective validation, thus maintaining both productivity and reliability.
Solution Approach 2:
The patent employs iterative consensus building where each read is evaluated against the emerging consensus sequence. Reads that deviate significantly from the consensus are identified as potential errors and excluded, while supporting reads reinforce the consensus. This feedback mechanism continuously refines the sequence accuracy as more data is processed.
2Ease of operation
If sequence alignment-based methods are used to analyze sequencing reads, then ease of operation is improved, but measurement precision deteriorates due to high error rates
Solution Approach 1:
The patent performs preliminary filtering of sequencing reads based on quality metrics and alignment confidence scores before conducting species identification. By pre-screening reads to exclude those with high error probabilities, the method maintains operational simplicity while significantly improving the precision of downstream analysis.
Solution Approach 2:
The patent applies different quality thresholds and analysis strategies to different regions of the sequencing data. High-confidence regions are analyzed with standard alignment methods, while low-confidence regions undergo more rigorous validation or are excluded from species identification, thus optimizing both ease of operation and measurement precision locally.
3Productivity
If majority species identification is prioritized, then productivity is improved, but reliability of minority species identification deteriorates
Solution Approach 1:
The patent applies a tiered analysis approach where reads are first processed for majority species identification using streamlined methods, then the same processed data is re-analyzed with more sensitive parameters for minority species detection. This partial re-analysis of already-processed data enables minority species identification without requiring complete re-processing, thus maintaining productivity while improving reliability.
Data Source
AI summary
The invention provides methods of determining a consensus sequence from multiple raw sequencing reads of a nucleic acid target. The nucleic acid target includes an anchor segment of known sequence and an adjacent segment of unknown sequence. The anchor segment provides a means to assess the quality of a raw target sequencing read. Raw target sequencing reads meeting or exceeding a threshold are assigned to an accepted class. The consensus sequence of the adjacent segment can be determined from raw target sequencing reads in the accepted class. Successive polling steps determine successive consensus nucleobases in a nascent sequence of the adjacent segment. Raw target sequencing reads can be removed or reintroduced from the accepted class depending on their correspondence to the most recently determined consensus nucleobase and/or the nascent sequence.


