Nucleic Acid Sequencing Read Error Correction with Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sequencing technologies suffer from high error rates, particularly in lower-cost methods like sequencing by flow, which confound downstream analyses by mistaking sequencing errors for true genetic variants, leading to inaccurate variant calling.

Innovation Solution

A machine learning-based sequencing error suppression algorithm is trained on a subset of sequencing data to identify and correct recurrent sequencing errors by distinguishing them from true genetic variants, using a neural network to predict the reference sequence and correct mismatches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If lower-cost sequencing techniques (sequencing by flow) are used, then cost is reduced and throughput is increased, but sequencing error rate increases leading to inaccurate variant calling

Engineering Contradiction:
Improvesequencing throughputVSAvoidsequencing accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an error correction module as an intermediary component between the sequencing device and variant calling analysis. This module uses machine learning models trained on sequencing data to identify and correct sequencing errors before variant calling, thereby mediating the trade-off between using high-throughput lower-cost sequencing methods and maintaining accurate variant detection

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary error correction processing to sequencing reads before they are used for variant calling. By training machine learning models on sequencing data and using them to correct errors in advance, the system prepares the data to maintain accuracy despite using higher-error-rate sequencing methods

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If sequencing errors are corrected by filtering mismatches, then false positives are reduced, but true genetic variants may be suppressed

Engineering Contradiction:
Improvevariant calling accuracyVSAvoidability to distinguish variants from errors
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs dynamic machine learning models that adapt to the specific sequencing data being analyzed. The models are trained on the actual sequencing reads and can dynamically adjust their error correction decisions based on patterns learned from the data, allowing them to distinguish between sequencing errors and true genetic variants rather than applying static filtering rules

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the sequencing data by using machine learning models to transform raw sequencing reads into corrected reads. The models learn optimal correction parameters from training data and apply these transformations to correct errors while preserving true variants, effectively changing the data parameters to improve accuracy

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250253012A1Error Correction of Nucleic Acid Sequencing Reads
Publication Date: 2025.08.07 ULTIMA GENOMICS INC
  • US20250253012A1 patent drawing
  • US20250253012A1 patent drawing
  • US20250253012A1 patent drawing

AI summary

Error correction of nucleic acid sequencing reads is described. A machine learning model may be trained using aligned sequencing reads generated during a nucleic acid sequencing event, the aligned sequencing reads aligned to a reference sequence. The trained machine learning model may output predicted reference sequences for the aligned sequencing reads that have a position of mismatch with the reference sequence. The aligned sequencing reads may be selectively corrected according to whether the machine learning model predicts that the aligned sequencing reads are variants or sequencing errors based on the predicted reference sequences relative to the reference sequence.