Consensus Sequence Generation for Homologous DNA Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-Generation Sequencing technologies face challenges in accurately sequencing homologous nucleotidic sequences due to high error rates and limitations in generating long reads, which complicates the assembly of consensus sequences from DNA molecules.

Innovation Solution

A system and method for generating consensus sequences by aligning and correcting sequence data, determining diversity status, and classifying reads to distinguish between main positions with and without diversity, thereby improving the accuracy of consensus sequence generation from Next-Generation Sequencing data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If Next-Generation Sequencing technologies are used to sequence homologous nucleotidic sequences, then sequencing speed and throughput are improved, but accuracy deteriorates due to high error rates and limitations in generating long reads

Engineering Contradiction:
Improvesequencing speed and throughputVSAvoidsequencing accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the sequencing data processing into distinct stages: initial alignment of reads to reference sequences, identification of diverse main positions, classification into groups based on diversity status, and iterative correction processes. This segmentation allows each stage to optimize for its specific function while collectively achieving high accuracy consensus sequences from high-throughput NGS data

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where consensus sequences generated in one iteration are used to improve subsequent iterations. The system continuously refines accuracy by using previously generated consensus information to guide further correction and classification of reads, creating a self-improving cycle that resolves the accuracy-throughput tradeoff

Inventive Principle:
Principle #23Feedback

2Length of moving object

If PacBio Sequencing System is used to generate long reads, then read length is improved, but accuracy deteriorates due to high error rates

Engineering Contradiction:
Improveread lengthVSAvoidsequencing accuracy
Core Design Contradiction:
Length of moving objectVSMeasurement precision

Solution Approach 1:

The patent merges multiple long reads with high error rates into consensus sequences by aligning them and identifying positions where they agree. By combining information from multiple independent reads, the system achieves high accuracy at each position while maintaining the long read length advantage, effectively merging the strengths of multiple erroneous sequences into one accurate consensus

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If consensus sequences are generated from homologous nucleotidic sequences, then accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improveconsensus sequence accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by focusing computational resources only on main positions that exhibit diversity among reads. Rather than uniformly processing all positions, the system identifies and concentrates effort on variable positions while efficiently handling invariant positions, thereby reducing overall computational complexity while maintaining accuracy where it matters most

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs partial correction actions by iteratively correcting only the necessary portions of sequence data that contain errors or diversity. Rather than reprocessing entire datasets repeatedly, the system applies corrections selectively to specific regions and positions that require improvement, reducing computational overhead while achieving convergence to accurate consensus sequences

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10937523B2Methods, systems and computer readable storage media for generating accurate nucleotide sequences
Publication Date: 2021.03.02 EMORY UNIVERSITY
  • US10937523B2 patent drawing
  • US10937523B2 patent drawing
  • US10937523B2 patent drawing

AI summary

Methods, systems and computer-readable storage media relate to generating one or more consensus sequences. The methods may include determining a group of one or more reads with each main position without diversity from a group of one or more aligned reads based on diversity status of each main position, each group including sequence data disposed at a plurality of main positions and a plurality of secondary position regions disposed adjacent to the main positions. The methods may also include determining legitimate sequence data from each second position region having one or more nucleotides for each group of one or more reads without diversity; and generating a consensus sequence including sequence data disposed at each main position without diversity and legitimate sequence data disposed at each secondary position region for each group of one or more reads with each main position without diversity.