NGS Repeat Variant Calling Without Control Sample Bias

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current next-generation sequencing (NGS) technologies face challenges in achieving optimal specificity and sensitivity for detecting genomic variants, particularly in homopolymer and heteropolymer regions, leading to false positives and negatives, and require manual configuration and optimization that is costly and not scalable.

Innovation Solution

A method for automating genomic data processing that includes identifying reference repeat patterns, measuring distribution of repeat lengths, and using best-fit models to estimate allele variants, with confidence levels, to improve variant calling accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If NGS technology is used for high throughput sequencing, then productivity increases, but measurement precision deteriorates due to biases in homopolymer and heteropolymer regions

Engineering Contradiction:
Improvesequencing throughputVSAvoidvariant detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary statistical modeling step that mediates between the biased NGS measurements and the true variant status. By comparing observed read distributions against expected distributions under different variant hypotheses, the system resolves sequencing biases without sacrificing throughput. This intermediary analysis layer transforms biased raw data into accurate variant calls.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter space by moving from direct base-calling to statistical distribution analysis. Instead of relying on absolute read counts or simple alignment scores, the system analyzes the distribution of read positions, depths, and quality scores across multiple samples. This parameter transformation enables accurate variant detection despite NGS sequencing biases.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If target enrichment is performed to focus on specific genomic regions, then manufacturing precision improves, but object-generated harmful factors increase due to PCR amplification biases

Engineering Contradiction:
Improvetarget region coverageVSAvoidPCR amplification bias
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The patent applies preliminary action by performing target enrichment and PCR amplification before sequencing, which allows the subsequent statistical model to account for and correct the introduced biases. The method incorporates expected bias patterns from enrichment and amplification into the statistical framework, enabling accurate variant calling despite these preliminary processing steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback by comparing observed read distributions against expected distributions that incorporate known biases from target enrichment and PCR amplification. The statistical model iteratively refines variant calls by feedback from the discrepancy between observed and expected patterns, ultimately resolving the harmful effects of amplification bias.

Inventive Principle:
Principle #23Feedback

3Productivity

If multiple samples are multiplexed in NGS experiments, then productivity increases, but reliability decreases due to increased complexity in data analysis

Engineering Contradiction:
Improvenumber of samples processedVSAvoidvariant calling accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges data from multiple multiplexed samples into a unified statistical analysis framework. By combining read data across samples and analyzing joint distributions, the system maintains high reliability even as productivity increases. The merged analysis leverages information from all samples to improve variant detection confidence and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The statistical modeling framework is designed with universality to handle any number of multiplexed samples through the same analytical pipeline. The method functions universally across different sample sizes, sequencing depths, and experimental conditions, maintaining reliability regardless of the scale of multiplexing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If manual workflow setup is used for NGS analysis, then adaptability improves, but loss of time increases due to case-per-case optimization

Engineering Contradiction:
Improveworkflow customizationVSAvoidworkflow setup time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the statistical modeling framework to automatically adapt to different datasets and experimental conditions without requiring manual reconfiguration. The system self-adjusts parameters and selects appropriate models based on the input data characteristics, eliminating time-consuming manual setup while maintaining adaptability to diverse genomic analysis needs.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12633377B2Methods for detecting variants in next-generation sequencing genomic data
Publication Date: 2026.05.19 SOPHIA GENETIS SA
  • US12633377B2 patent drawing
  • US12633377B2 patent drawing
  • US12633377B2 patent drawing

AI summary

A genomic data analyzer may be configured to detect and characterize, with a variant calling module, genomic variants from next generation sequencing reads out of a pool of enriched genomic patient samples without suffering from next generation sequencing workflow biases such as those introduced by sequencing errors in particular in repeat patterns regions of the human genome such as homopolymers or heteropolymers. The variant calling module may estimate the probability distribution of the length of the repeat pattern for each patient sample and cross-analyze it against other samples in a single experimental pool to identify best-fit variant models for each pair of samples. The variant calling module may further group samples according to their matching best-fit variant models and identify which group of patient samples carries the wild type reference without the need for control data in the pool. The variant calling module may subsequently characterize the homozygous or heterozygous repeat patterns variants for each patient sample with improved specificity and accuracy even in the presence of next generation sequencing biases.