Hierarchical Neural Network for High-Throughput Sequencing Basecalling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing techniques generate massive amounts of data, leading to challenges in processing and accuracy due to fluorescent signal decay, crosstalk, and cluster phasing, which affect the quality of basecalling results.

Innovation Solution

A hierarchical processing network structure using multiple network blocks with base networks and sequence filters to generate sequencing quality indicators and filter out low-quality signals, improving the accuracy of basecalling by enhancing data quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If next generation sequencing techniques are used to increase throughput, then productivity is improved, but measurement precision deteriorates due to fluorescent signal decay and crosstalk

Engineering Contradiction:
Improvesequencing throughputVSAvoidbasecalling accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces machine learning models as an intermediary between raw fluorescent signal data and basecalling results. The ML models process and interpret the degraded signals, compensating for signal decay and crosstalk effects, thereby maintaining high basecalling accuracy despite the challenges of high-throughput sequencing

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the basecalling problem from direct signal interpretation to a machine learning classification problem. By changing the parameters and approach of signal analysis using trained neural networks, the system can accurately determine nucleotide sequences even when traditional signal quality metrics deteriorate

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If machine learning models are used for basecalling, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvebasecalling accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by training machine learning models in advance on large datasets before actual sequencing analysis. The pre-trained models can then rapidly process sequencing data without requiring complex real-time computations, reducing the complexity burden during actual basecalling operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating trained machine learning model copies that can be deployed across multiple processing units. This allows parallel processing of sequencing data, improving efficiency while distributing the computational complexity across multiple simpler units rather than requiring one highly complex system

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method significantly improves the accuracy and quality of basecalling results by filtering out low-quality signals and enhancing data processing efficiency, potentially matching or surpassing the quality of traditional low-throughput sequencing techniques.

Implementation Method 1

The added fluorescent dye is excited by absorbing light energy from the laser

Methodology Applied
Scientific EffectAbsorption (EM radiation): Absorption (EM radiation)

Implementation Method 2

The fluorescents signals are collected and analyzed to predict the sequences of the nucleic acid

Methodology Applied
Scientific EffectFluorescence: Fluorescence

Data Source

PatentUS20240013861A1Methods and systems for enhancing nucleic acid sequencing quality in high-throughput sequencing processes with machine learning
Publication Date: 2024.01.11 GENESENSE TECH INC
  • US20240013861A1 patent drawing
  • US20240013861A1 patent drawing
  • US20240013861A1 patent drawing

AI summary

This disclosure provides an improved filtering techniques that can provide higher sequencing accuracy for processing high-throughput sequencing data. The filtering structure uses a hierarchical network structure comprising one or more network blocks for obtaining high quality sequences. Each network block comprises a base network module (or simply a base network) and a sequence filter. The base network generates one or more sequencing quality indicators. The sequencing quality indicators can represent qualities of accuracy of the basecalling of individual bases in a sequence, the quality of one or more sequences individually, or the quality of a group of sequences as a whole. The sequence filter generates the filtered results based on the various filtering strategies based on the one or more sequencing quality indicators.