DNA Sequencing Image Analysis Pipeline for High-Throughput Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Nucleic acid sequencing technologies face challenges in efficiently analyzing and processing the vast amounts of data generated from sequencing-by-synthesis methods, particularly in terms of speed and accuracy, due to the large quantity of sequence information produced.

Innovation Solution

The methods and systems described involve calculating predictor values for base calls, using quality tables generated by Phred scoring, and applying algorithms like End Anchored Maximal Scoring Segments (EAMSS) to evaluate base call quality, as well as performing phasing corrections to improve sequencing accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If massively parallel sequencing reactions are performed simultaneously to generate large amounts of sequence information, then the quantity of sequence information produced is improved, but the speed and quality of analysis of sequence data deteriorates

Engineering Contradiction:
Improvequantity of sequence informationVSAvoidspeed of analysis
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments the analysis process into distinct stages: image processing, base calling, and quality scoring. Each stage handles specific tasks independently, allowing parallel processing of multiple sequencing reactions without creating a bottleneck in the analysis pipeline. This segmentation enables the system to maintain high productivity even as the quantity of sequence information increases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by performing quality scoring during the base calling stage rather than as a separate subsequent step. Quality metrics are calculated and assigned to each base call as it is made, allowing downstream analysis to proceed with pre-evaluated data quality information. This eliminates the need for separate quality assessment passes and accelerates overall analysis speed.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If massively parallel sequencing reactions are performed simultaneously to generate large amounts of sequence information, then the quantity of sequence information produced is improved, but the quality of analysis of sequence data deteriorates

Engineering Contradiction:
Improvequantity of sequence informationVSAvoidquality of analysis
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where quality scores are continuously monitored and used to adjust base calling parameters. The quality scoring system provides real-time feedback on the reliability of each base call, allowing the system to maintain high measurement precision even when analyzing large volumes of data from massively parallel sequencing reactions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

By performing quality assessment during the base calling stage itself, the system ensures that quality metrics are embedded in the data from the outset. This preliminary quality evaluation prevents degradation of analysis quality that would occur if quality assessment were deferred to a separate later stage, maintaining high measurement precision across all sequence data regardless of volume.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If traditional analysis methods are used to process sequence data, then the existing infrastructure can handle the processing, but the speed and efficiency of data analysis deteriorates

Engineering Contradiction:
Improveinfrastructure requirementsVSAvoidefficiency of data analysis
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent implements dynamic processing where the analysis pipeline adapts to the volume and characteristics of incoming data. The system dynamically adjusts processing parameters and resource allocation based on real-time data flow, enabling efficient handling of variable data loads without requiring proportional increases in infrastructure complexity. This dynamic approach maintains high analysis efficiency while working within existing computational resources.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3940082A1Methods and systems for analyzing image data
Publication Date: 2022.01.19 ILLUMINA INC
  • EP3940082A1 patent drawingFigure 1A~1B
  • EP3940082A1 patent drawingFigure 2A~2C
  • EP3940082A1 patent drawingFigure 3~4

AI summary

Methods and systems for analysis of image data generated from various reference points. Particularly, the methods and systems provided are useful for real time analysis of image and sequence data generated during DNA sequencing methodologies.