Genomic Variant Interpretation via Pooled Allele Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The interpretation of DNA variants from sequence-based tests in clinical laboratories is inefficient due to the complexity of tests, the need for extensive literature review, and the challenge of scaling with increasing test volumes, leading to delays in patient treatment and misclassification of variants due to ethnic bias and lack of contextual genomic analysis.
Innovation Solution
A knowledge-based system that evaluates genomic variants using expert-curation and ontology-based structured information, providing automated classification and integrating phenotype information to streamline variant interpretation and identify suitable patients for clinical trials.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual literature review and expert interpretation are used for variant analysis, then interpretation accuracy can be maintained, but the time required for test results increases significantly
Solution Approach 1:
The system performs preliminary automated annotation of variants using pooled allele statistics and literature mining before expert review. Bibliographies are generated in advance for each variant, and literature is pre-curated and structured in advance, so that when a variant is observed, the interpretation work has already been partially completed, significantly reducing the time experts need to spend while maintaining accuracy
Solution Approach 2:
The system creates structured copies of literature information and variant data in standardized formats. Curated bibliographies are copied and associated with variants, and pooled allele statistics are copied from large datasets to individual variant assessments. This allows rapid retrieval and comparison without re-reading original sources, accelerating the interpretation process
2Adaptability or versatility
If the number of genes assayed per test increases to improve diagnostic capability, then test comprehensiveness improves, but the complexity of variant interpretation increases
Solution Approach 1:
The system segments the interpretation task by processing each variant independently with its own automated annotation pipeline, then aggregating results. Each variant receives its own bibliography and pooled statistics assessment, allowing parallel processing of multiple variants without compounding complexity. The large panel of genes is broken down into individual variant assessments that can be handled systematically
Solution Approach 2:
The system implements a universal interpretation framework that handles all variants across all genes using the same pooled allele statistics approach and literature mining pipeline. This multi-functional system applies consistent methods regardless of gene or variant type, providing scalable interpretation capability that grows with test complexity without requiring proportionally increased expert resources
3Measurement precision
If pooled allele statistics from multiple populations are used for variant assessment, then accuracy and reduction of ethnic bias improve, but data processing complexity increases
Solution Approach 1:
The system merges allele frequency data from multiple population cohorts into pooled statistics that can be directly applied to variant assessment. By combining datasets and calculating pooled allele frequencies and observation counts, the system reduces ethnic bias and improves accuracy. The merging is done systematically using statistical methods that account for population differences, transforming complex multi-population data into simplified pooled metrics for interpretation
Data Source
AI summary
Disclosed herein are system, method, and computer program product embodiments for building a community database of allele counts. An embodiment operates by receiving human variant datasets derived from samples generated by distinct users, wherein the users consented to share pooled variant observations with other users; determining that a plurality of variant observations meet the inclusion criteria for a pool; and calculating one or more anonymized allele statistics from the pool.


