Consensus Sequence Determination Using Anchor Segments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-throughput DNA sequencing methods generate vast amounts of data with high error rates, making it difficult to accurately identify majority and minority species, especially when they differ minimally, and to determine their prevalence.

Innovation Solution

A computer-implemented method that uses an anchor segment of known sequence to evaluate the accuracy of sequencing reads, assign them to accepted or rejected classes, and iteratively poll nucleobases to determine a consensus sequence for the adjacent unknown segment, allowing for the correction and refinement of sequencing errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If high-throughput sequencing methods are used to generate vast amounts of sequencing data, then productivity increases, but sequencing error rate increases

Engineering Contradiction:
Improvesequencing throughputVSAvoidsequencing accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent combines multiple sequencing reads that cover the same nucleic acid target into a consensus sequence. By merging information from multiple reads, the method leverages the high throughput advantage while compensating for individual read errors through collective validation, thus maintaining both productivity and reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent employs iterative consensus building where each read is evaluated against the emerging consensus sequence. Reads that deviate significantly from the consensus are identified as potential errors and excluded, while supporting reads reinforce the consensus. This feedback mechanism continuously refines the sequence accuracy as more data is processed.

Inventive Principle:
Principle #23Feedback

2Ease of operation

If sequence alignment-based methods are used to analyze sequencing reads, then ease of operation is improved, but measurement precision deteriorates due to high error rates

Engineering Contradiction:
Improvesequence analysis simplicityVSAvoidspecies identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent performs preliminary filtering of sequencing reads based on quality metrics and alignment confidence scores before conducting species identification. By pre-screening reads to exclude those with high error probabilities, the method maintains operational simplicity while significantly improving the precision of downstream analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different quality thresholds and analysis strategies to different regions of the sequencing data. High-confidence regions are analyzed with standard alignment methods, while low-confidence regions undergo more rigorous validation or are excluded from species identification, thus optimizing both ease of operation and measurement precision locally.

Inventive Principle:
Principle #3Local quality

3Productivity

If majority species identification is prioritized, then productivity is improved, but reliability of minority species identification deteriorates

Engineering Contradiction:
Improveidentification speedVSAvoidminority species detection accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies a tiered analysis approach where reads are first processed for majority species identification using streamlined methods, then the same processed data is re-analyzed with more sensitive parameters for minority species detection. This partial re-analysis of already-processed data enables minority species identification without requiring complete re-processing, thus maintaining productivity while improving reliability.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11862299B2Algorithms for sequence determinations
Publication Date: 2024.01.02 GEN PROBE INC
  • US11862299B2 patent drawing
  • US11862299B2 patent drawing
  • US11862299B2 patent drawing

AI summary

The invention provides methods of determining a consensus sequence from multiple raw sequencing reads of a nucleic acid target. The nucleic acid target includes an anchor segment of known sequence and an adjacent segment of unknown sequence. The anchor segment provides a means to assess the quality of a raw target sequencing read. Raw target sequencing reads meeting or exceeding a threshold are assigned to an accepted class. The consensus sequence of the adjacent segment can be determined from raw target sequencing reads in the accepted class. Successive polling steps determine successive consensus nucleobases in a nascent sequence of the adjacent segment. Raw target sequencing reads can be removed or reintroduced from the accepted class depending on their correspondence to the most recently determined consensus nucleobase and/or the nascent sequence.