Variant Annotation Matching Across Different NGS Representations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current NGS workflows lack an efficient and automated method to identify genomic variants in medical reference databases independently of variant calling representations and sequencing technologies, leading to computational inefficiencies and manual processing challenges in large-scale genomic analysis.
Innovation Solution
A method and genomic data analyzer that reconstructs patient and database haplotypes over extended genomic regions to identify matching variants by concatenating genomic information strings, allowing for efficient matching of variants with different representations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional manual methods are used to identify variants in medical reference databases, then accuracy can be maintained, but productivity is low and manual processing is required
Solution Approach 1:
The patent introduces an intermediary normalization process that converts variants from different representations into a common format before comparison. This mediator layer enables automated matching between patient variants and database entries without requiring manual intervention, while maintaining accuracy through systematic transformation rules
Solution Approach 2:
The patent transforms variants by changing their representation parameters - converting between different coordinate systems, reference genomes, and variant formats. This parameter transformation enables automated comparison and matching while preserving the biological meaning of the variants
2Adaptability or versatility
If NGS workflows process multiple samples with different sequencing technologies, then versatility is improved, but device complexity and processing difficulty increase
Solution Approach 1:
The patent creates a universal variant normalization framework that handles multiple sequencing technologies and variant representations through a single automated pipeline. This multi-functional system processes variants from different NGS platforms, target enrichment methods, and sequencing chemistries using the same normalization and matching algorithms
Solution Approach 2:
The patent segments the complex variant matching process into distinct modular steps: normalization, transformation, and comparison. Each module handles specific aspects of variant processing independently, making the overall workflow more manageable and adaptable to different sequencing technologies
3Measurement precision
If exact matching methods are used for variant identification, then measurement precision is high, but adaptability to different variant representations is poor
Solution Approach 1:
The patent performs preliminary normalization and transformation of variants before the actual matching process. By pre-converting all variants to a common representation format, the system maintains high matching accuracy while being able to handle diverse input representations from different sequencing technologies and databases
Data Source
Figure 1a~1b
Figure 2
Figure 3
AI summary
A genomic data analyzer workflow may be configured to identify, with a variant annotation module, subsets of patient variants which match at least one medical reference variant database entry, even if the variant calling information in genomic data analyzer workflow and the database use different variant representations of SNP, MNP, INDELS and DELINS. In particular, database variants which are included into a subset of patient variants may be identified even if they do not exactly match the corresponding strings. The variant annotation module may be adapted to apply a branch-and-bound-like algorithm to efficiently process all possible subsets of patient variants in a genomic region.