NGS Variant Annotation Using Haplotype Reconstruction Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current NGS workflows face challenges in efficiently identifying genomic variants across diverse sequencing technologies and databases due to varying alignment and variant calling parameters, leading to non-standardized representations that hinder automated and cost-effective data processing for clinical applications.

Innovation Solution

A method and genomic data analyzer that reconstructs patient and database haplotypes within an extended genomic region to identify matching variants by comparing reconstructed strings, allowing for efficient matching of variants with different representations, even if they are not exact matches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If NGS workflows use diverse sequencing technologies and databases with varying alignment and variant calling parameters, then the ability to detect variants across different platforms is improved, but the data processing complexity and lack of standardization worsen

Engineering Contradiction:
Improvecompatibility across sequencing technologiesVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies homogeneity by standardizing variant representations through a common data model and schema. All variant calls from different NGS platforms are normalized to a unified format with consistent fields for genomic position, reference allele, alternate allele, and quality metrics. This homogeneous representation enables automated processing across diverse sequencing technologies without requiring platform-specific handling, thus reducing data processing complexity while maintaining versatility.

Inventive Principle:
Principle #33Homogeneity

Solution Approach 2:

The patent introduces an intermediary layer in the form of a standardized variant annotation system that sits between raw NGS data from different platforms and the final analysis. This intermediary layer includes a common data model, standardized annotation pipelines, and a unified reference database framework that mediates between diverse input formats and output requirements, enabling automated processing without losing the ability to handle technology-specific variations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If NGS workflows use diverse sequencing technologies and databases with varying alignment and variant calling parameters, then the ability to detect variants across different platforms is improved, but the cost of data processing increases

Engineering Contradiction:
Improvecompatibility across sequencing technologiesVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent applies parameter changes by transforming variant data from diverse platforms into a standardized parameter set with fixed data types, length constraints, and value ranges. The common data model defines specific parameters for genomic coordinates, allele frequencies, quality scores, and filtering thresholds that can be uniformly processed. This parameter standardization enables efficient automated filtering and analysis pipelines that reduce computational resource consumption compared to handling unstandardized, platform-specific data formats.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If exact matching is required for variant identification, then the precision of variant detection is improved, but the ability to identify variants with different representations worsens

Engineering Contradiction:
Improvevariant detection accuracyVSAvoidmatching flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by performing variant normalization and canonical representation transformation before the matching process. All input variants from different platforms are pre-processed to convert them into a standard format using defined rules for handling indels, MNPs, and complex variants. This preliminary standardization ensures that subsequent exact matching operations can successfully identify variants even when they were originally represented differently across platforms, thus maintaining both precision and flexibility.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent inverts the traditional approach by not trying to make the matching algorithm flexible to handle various representations, but rather by making all representations conform to a single standard before matching. Instead of implementing complex fuzzy matching logic, the system reverses the problem by standardizing the input data first, then applying simple exact matching, which achieves both high precision and broad adaptability.

Inventive Principle:
Principle #13The other way round (Inversion)

4Measurement precision

If manual processing is used for variant identification, then the accuracy of complex variant matching is improved, but the productivity and automation level worsen

Engineering Contradiction:
Improvevariant matching accuracyVSAvoiddata processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies self-service by implementing automated pipelines that perform variant normalization, annotation, and matching without requiring manual intervention. The system includes self-contained algorithms for handling complex variants, automatic quality filtering, and standardized reporting generation. This automation maintains high accuracy through rigorously tested algorithms while dramatically increasing processing throughput compared to manual methods, enabling the system to handle large volumes of NGS data efficiently.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12573472B2Methods for detecting variants in next- generation sequencing genomic data
Publication Date: 2026.03.10 SOPHIA GENETICS SA
  • US12573472B2 patent drawing
  • US12573472B2 patent drawing
  • US12573472B2 patent drawing

AI summary

A genomic data analyzer workflow may be configured to identify, with a variant annotation module, subsets of patient variants which match at least one medical reference variant database entry, even if the variant calling information in genomic data analyzer workflow and the database use different variant representations of SNP, MNP, INDELS and DELINS. In particular, database variants which are included into a subset of patient variants may be identified even if they do not exactly match the corresponding strings. The variant annotation module may be adapted to apply a branch-and-bound-like algorithm to efficiently process all possible subsets of patient variants in a genomic region.