Genetic Variant Classification via Population Probabilistic Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current classification systems for genetic variants, such as those related to BRCA1 and BRCA2 genes, often leave patients with uncertain variants (VUS) in a state of clinical limbo, lacking clear guidance for management and treatment, due to indeterminate classifications that do not provide sufficient information for healthcare providers.
Innovation Solution
A computer-based system analyzes data from large patient populations to assign statistical probabilities and weights to genetic variants, comparing personal and family health histories to composite control cohorts to determine the likelihood of variants being deleterious or benign, allowing for reclassification and improved clinical management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a 5-tier classification system is used for genetic variants, then variants can be categorized into deleterious, suspected deleterious, VUS, favor polymorphism, and polymorphism, but patients with VUS variants remain in a state of clinical uncertainty without clear management guidance
Solution Approach 1:
The patent changes the parameter of classification from static 5-tier categories to dynamic probabilistic scores based on population data. By calculating the proportion of disease carriers versus non-carriers in the population with the same variant, the system transforms uncertain VUS classifications into quantifiable risk probabilities, providing clinical guidance while maintaining classification precision.
Solution Approach 2:
The system incorporates feedback loops where population sequencing data continuously refines variant classification. As more population data is collected, the probabilistic scores for variants are updated, allowing VUS variants to be reclassified based on emerging evidence from the population, thereby reducing clinical uncertainty over time.
2Reliability
If population-scale DNA sequencing is conducted to improve variant classification, then more accurate disease risk assessments can be obtained, but the complexity of analyzing and interpreting large-scale genetic data increases
Solution Approach 1:
The patent introduces an intermediary computational framework that mediates between raw population sequencing data and clinical variant classification. The system uses population frequency data as an intermediary layer to translate complex genetic variation patterns into simplified probabilistic scores, making large-scale data analysis manageable while improving assessment reliability.
Solution Approach 2:
The analysis system is segmented into modular components: population data collection, variant frequency calculation, probabilistic scoring, and clinical interpretation. This segmentation allows each module to process specific aspects of the data independently, reducing overall system complexity while maintaining high reliability in disease risk assessment.
3Measurement precision
If statistical probabilities from large patient populations are used to reclassify variants, then more accurate clinical management decisions can be made, but the time and computational resources required for analysis increase
Solution Approach 1:
The system performs preliminary actions by pre-calculating and storing population variant frequencies and probabilistic scores in databases. When a patient's variant needs classification, the system retrieves pre-computed population data and applies it directly, avoiding time-consuming re-analysis of entire population datasets while maintaining high classification accuracy.
Solution Approach 2:
The patent transforms the time-intensive process of individual variant analysis into efficient parameter-based lookup operations. By changing from detailed sequence-by-sequence analysis to parameter-based probabilistic scoring using pre-computed population statistics, the system achieves accurate reclassification with significantly reduced time and computational resources.
Data Source
AI summary
A computer-implemented method is discussed that includes identifying, by a computer server system, stored electronic data that represents genetic sequencing for one or more genes for individuals in a population of patients who have submitted to genetic sequencing; generating, for each of multiple individuals and from the stored electronic data, probability data for the individuals and probability or weighting data, or both, for relatives of the individuals, the probability data representing likelihoods that a particular person corresponding to the probability data carries a deleterious mutation in a particular gene; and generating a score for a genetic variant, wherein the score is a function of probability or weighting data, or both, for the individuals and for relatives of the individuals, and the score represent a composite probability that a certain variant is a deleterious or benign variant.


