Complex Trait Classification with Rare Variant Gene Burden Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to identify and classify the genetic patterns underlying complex human traits and disorders, particularly those influenced by rare variants with low minor allele frequency, which are difficult to detect and analyze due to their rarity and lack of statistical significance in traditional genome-wide association studies.
Innovation Solution
A computational classification model is trained using genetic data from cohorts with and without complex disorders to identify a minimal set of genes with variant patterns, utilizing a penalized linear classification model like LASSO to determine an individual's propensity for a disorder based on aggregated variant burden scores, including rare variants, and optionally incorporating health record data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional genome-wide association studies are used to identify genetic variants, then common variants can be detected, but rare variants with low minor allele frequency cannot be statistically significant
Solution Approach 1:
The patent combines multiple rare variants within the same gene into an aggregate burden score. Instead of evaluating each rare variant individually (which lacks statistical power), the method aggregates the effects of multiple rare variants across a gene, thereby achieving statistical significance while maintaining detection accuracy for rare variants.
2Area of stationary object
If genome-wide association studies examine the entire genome, then comprehensive genetic coverage is achieved, but the complexity and computational burden increase
Solution Approach 1:
The patent extracts and focuses analysis on specific genes that are pre-identified as potentially relevant to the complex trait or disorder. Rather than analyzing the entire genome, the method extracts and evaluates only the variant burden within selected genes, thereby maintaining comprehensive genetic coverage of relevant regions while significantly reducing computational complexity.
3Reliability
If a large number of cohort individuals are used in traditional GWAS, then statistical power increases, but the required sample size becomes prohibitively large for rare variants
Solution Approach 1:
By merging multiple rare variants into an aggregate burden score within each gene, the patent increases statistical power without requiring a proportionally large increase in cohort sample size. The aggregation approach allows the method to achieve reliable detection of rare variant contributions with more manageable cohort sizes compared to traditional GWAS.
4Measurement precision
If computational models incorporate multiple genes and variants, then classification accuracy improves, but the model complexity increases
Solution Approach 1:
The patent segments the complex genetic analysis into discrete gene-level burden scores. Instead of analyzing all variants simultaneously across the genome (which would be computationally intensive), the method divides the problem into manageable segments (individual genes), calculates burden scores for each segment, and then integrates these segments into the final classification model, thereby improving accuracy while controlling complexity.
Data Source
AI summary
Processes to identify a subset of trait-related genes and classify individuals are described. Generally, systems generate classification models which are used to identify the subset of trait-related genes and classify individuals. The classification models are also used in various applications, including developing research tools, performing diagnostics, and treating individuals.


