Rare Disease Phenotype–Variant Matching with Hungarian Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for diagnosing rare human diseases through DNA sequencing face challenges in accurately matching phenotype descriptions with pathogenic variants, often leading to incorrect interpretations due to reliance on incomplete or incorrect phenotype information.
Innovation Solution
A method and system that utilize a Hungarian matching algorithm to perform one-to-one genotype-phenotype matching by segmenting and prioritizing phenotypes and genotypes based on metadata, using Human Phenotype Ontology and variant prioritization techniques to calculate a similarity score, ensuring accurate matching through rank-based overlapping and Hungarian matching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional methods (Whole Genome Sequencing, Whole Exome Sequencing, gene panel sequencing) are used to identify variants, then the coverage of genetic analysis is improved, but the accuracy of phenotype-genotype matching deteriorates due to reliance on incomplete or incorrect phenotype information
Solution Approach 1:
The patent introduces a mediator mechanism that uses phenotype information from multiple sources (including other organisms) as an intermediary to bridge the gap between genomic variants and human phenotypes. This intermediary approach allows the system to leverage available phenotypic data without directly relying on potentially incorrect human phenotype annotations, thereby improving matching accuracy while maintaining comprehensive genetic coverage
Solution Approach 2:
The patent replaces conventional direct lookup methods (mechanical system) with a computational matching system that uses algorithms to compare and rank variant-phenotype associations. This substitution of the mechanical lookup approach with an intelligent matching system enables more accurate phenotype-genotype correspondence by considering multiple factors beyond simple database lookup
2Productivity
If direct look-up of subject phenotypes in resources such as OMIM is performed to filter variant lists, then the speed of variant identification is improved, but the reliability of phenotype interpretation deteriorates due to risk of incorrect phenotype interpretation
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously refines phenotype interpretations by comparing results against multiple data sources and using iterative ranking processes. The system provides feedback loops that allow re-evaluation of phenotype assignments based on variant priorities and phenotypic similarity scores, thereby maintaining both speed and reliability through iterative refinement rather than single-pass lookup
Solution Approach 2:
The patent changes the parameters used for phenotype matching from binary presence/absence in databases to continuous similarity scores and ranked lists. By transforming the matching process into a gradient-based ranking system that considers multiple parameters (phenotypic similarity, variant priority, inheritance patterns), the system achieves both rapid processing and high reliability through nuanced parameter comparison rather than simple database lookup
3Quantity of substance
If phenotype information of other organisms is used to extrapolate human gene phenotype associations, then the comprehensiveness of phenotype data is improved, but the accuracy of matching deteriorates due to species-specific differences
Solution Approach 1:
The patent applies local quality by treating different organism types differently in the matching process. Instead of uniform treatment of all phenotypic data, the system weights and prioritizes human phenotype data locally while using phenotypic information from other organisms as supplementary evidence. This localized differentiation ensures that human-specific phenotypic characteristics receive higher priority in the final matching, maintaining accuracy while still benefiting from the completeness provided by cross-species data
Data Source
Figure 1
Figure 2A
Figure 2B~2C
AI summary
Diagnosis of rare human diseases using DNA sequencing is a fast growing area of research. Conventional methods carries a risk of incorrect phenotype interpretation. However, obtaining a correct genotype and phenotype matching is challenging. A system for matching phenotype descriptions and pathogenic variants provides a one to one mapping of the phenotype and genotypes of a plurality of subjects under test. Initially, a plurality of phenotypes and a plurality of genome sequences are segmented based on metadata. A phenotype driven gene prioritization and a variant prioritization is applied on the segmented data method. A similarity score is calculated between the phenotype driven gene prioritization output and the variant prioritization output. The similarity score is further utilized to obtain a one to one matching of the plurality of phenotypes and the plurality of genotype sequences of the plurality of subjects under test.