Phenotypic Fit Analysis for Gene Prioritization Beyond Genotype Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing diagnostic genetic testing methods struggle to accurately analyze the relationship between a subject's clinical symptoms and genetic profile, often missing individuals with milder diseases or disorders with variable expressivity due to reliance on perfect genotype matches.

Innovation Solution

A method and system that combines curated datasets and real-world data to rank the likelihood of a gene-phenotype match using gene-phenotype associations, subject-gene similarity subscores, and machine learning models, independent of genotype information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If diagnostic genetic testing relies on perfect genotype matches, then genotype accuracy is improved, but the ability to detect individuals with milder diseases or variable expressivity deteriorates

Engineering Contradiction:
Improvegenotype match accuracyVSAvoiddisease detection capability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the analysis into two independent parts: genotype analysis and phenotype analysis. The phenotype analysis is further segmented into multiple subscores (HRSS, Jaccard, word2vec, doc2vec) that can be calculated independently. This allows the system to maintain high genotype match accuracy while separately optimizing phenotype-based disease detection, thereby resolving the contradiction between genotype precision and disease detection reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phenotype information as an intermediary between genotype and disease diagnosis. Instead of directly matching genotypes to diseases, the system uses phenotype subscores as a mediator that bridges the gap between genetic data and clinical presentation. This intermediary allows detection of milder diseases that may not have clear genotype matches, resolving the contradiction by enabling phenotype-driven diagnosis complementing genotype-based approaches.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual curation of gene-phenotype associations is used, then association accuracy is improved, but processing time and labor requirements deteriorate

Engineering Contradiction:
Improvegene-phenotype association accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service through automated literature mining and NLP processing. The system automatically retrieves, formats, and processes scientific literature to extract gene-phenotype associations without requiring manual curation. The NLP models (word2vec, doc2vec) automatically compute similarity subscores, enabling the system to process large datasets independently and rapidly, thereby maintaining high association accuracy while dramatically reducing processing time and labor requirements.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual curation process with automated computational systems. Instead of human experts manually reviewing and curating gene-phenotype associations, the system uses automated literature mining pipelines, NLP algorithms, and machine learning models to perform the same function. This substitution maintains or improves association accuracy while eliminating the time-consuming manual process, directly resolving the contradiction between precision and processing time.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If multiple data sources are integrated, then comprehensive analysis capability is improved, but system complexity deteriorates

Engineering Contradiction:
Improvedata integration capabilityVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the data integration system into distinct modules: literature mining module, NLP processing module, subscore calculation module (HRSS, Jaccard, word2vec, doc2vec), and machine learning module. Each module processes specific data types and performs specific functions. This segmentation allows the system to handle multiple data sources (scientific literature, clinical data, gene databases) independently and systematically, improving comprehensive analysis capability while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal processing framework that can handle multiple data sources through a common set of tools and algorithms. The NLP models and subscore calculations serve multiple purposes across different data types. The machine learning model integrates all subscores to produce final disease probability predictions. This multi-functional universal system improves adaptability to various data sources while avoiding the complexity of separate specialized systems for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260057139A1Methods and systems for phenotypic fit analysis
Publication Date: 2026.02.26 GENEDX LLC
  • US20260057139A1 patent drawing
  • US20260057139A1 patent drawing
  • US20260057139A1 patent drawing

AI summary

The present disclosure provides a method for performing phenotypic fit analysis. The method comprises computer processing an input dataset to produce a set of genes and a set of gene-phenotype associations (GPAs) associated with the set of genes. The method further comprises determining, for a subject, a plurality of subject-gene similarity subscores, based at least in part on the GPAs associated with the set of genes. The method further comprises determining a predicted likelihood of association between the subject and at least a subset of the set of genes, based at least in part on the plurality of subject-gene similarity subscores.