Genomic Data Analyzer for NGS Variant Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing (NGS) technologies face challenges in achieving high specificity and sensitivity in genomic analysis due to biases from sequencing technology, DNA enrichment methods, and genomic data structures, particularly in regions with homopolymer and heteropolymer repeats, which require manual customization and are costly and inefficient for routine clinical applications.

Innovation Solution

A genomic data analyzer system that automatically detects experimental and genomic context-specific biases by configuring sequence alignment and variant calling modules to refine data processing, using probability distributions to accurately estimate repeat pattern lengths and minimize experimental errors, enabling robust variant calling across diverse sequencing setups.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual customization is used for analyzing homopolymer and heteropolymer repeat regions, then measurement precision is improved, but device complexity and time consumption increase

Engineering Contradiction:
Improvevariant detection accuracyVSAvoidworkflow complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-calculating expected probability distributions for various repeat pattern lengths using control samples before analyzing patient data. This pre-computation of reference distributions enables automated comparison and eliminates the need for manual customization during actual variant detection, while maintaining high precision through statistically robust expected value comparisons.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service by automatically detecting experimental biases and genomic context-specific biases without human intervention. The bias detection module autonomously analyzes control samples to characterize sequencing and enrichment biases, then uses these characteristics to automatically adjust probability distribution calculations for accurate variant calling in repeat regions, eliminating manual workflow customization.

Inventive Principle:
Principle #25Self-service

2Productivity

If automated bias detection is implemented, then productivity is improved, but device complexity increases

Engineering Contradiction:
Improveworkflow automation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex automated workflow into distinct functional modules: a bias detection module that analyzes control samples to characterize sequencing and enrichment biases, a probability distribution calculation module that computes expected distributions for various repeat pattern lengths, and a variant detection module that compares patient data against these distributions. This modular segmentation manages system complexity by making each module independently testable and maintainable while achieving high overall productivity through automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses parameter changes by dynamically adjusting probability distribution parameters based on detected biases. The bias detection module quantifies sequencing and enrichment biases as numerical parameters, which are then used to modify the expected probability distribution calculations. This automatic parameter adjustment enables the system to adapt to different experimental conditions and genomic contexts without manual reconfiguration, significantly improving productivity across diverse applications.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If probability distributions are used to estimate repeat pattern lengths, then measurement precision is improved, but loss of time increases due to computational complexity

Engineering Contradiction:
Improverepeat pattern length estimation accuracyVSAvoidcomputational processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies preliminary action by pre-computing expected probability distributions for various repeat pattern lengths using control samples before patient data analysis. These pre-calculated reference distributions are stored and reused during variant detection, eliminating the need to perform complex computational simulations during actual patient sample analysis. This approach maintains high measurement precision through statistically robust distributions while significantly reducing computational processing time for clinical applications.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240428888A1Methods for detecting variants in next-generation sequencing genomic data
Publication Date: 2024.12.26 SOPHIA GENETICS SA
  • US20240428888A1 patent drawing
  • US20240428888A1 patent drawing
  • US20240428888A1 patent drawing

AI summary

A genomic data analyzer may be configured to detect and characterize, with a variant calling module, genomic variant scenarios on sequencing reads from an enriched patient genomic sample comprising a combination of a first repeat pattern and a second repeat pattern, such as repeats of homopolymer (single nucleotide) and/or heteropolymer (multiple nucleotide) basic motifs. The variant calling module may estimate the probability distribution of the length of the first repeat pattern and the probability distribution of the repeat pattern length measurements in patient data to the distribution of the repeat pattern length measurements in control data, in order to remove biases possibly induced by the next generation sequencing laboratory setup both in control and patient data. The variant calling module may further measure, read by read, the joint probability distribution for the first and the second repeat patterns lengths, and compare it with the expected joint probability distribution for various genomic variant scenarios for the patient, each variant scenario being characterized by a first length of the first repeat pattern and a second length of the second repeat pattern, to select the most likely patient genomic variant scenario as the scenario for which the measured joint probability distribution best matches the expected joint probability distribution.