IBD Cluster-Based Variant Origin Characterization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for assessing populations and identifying variants of interest lack effective tools to characterize the origins, migration patterns, and historical geographic locations of genetic variants, which are crucial for understanding associated phenotypes and targeting at-risk populations.
Innovation Solution
A method involving DNA dataset analysis to identify Identity-by-Descent (IBD) clusters, annotate them with genealogical data, and use community-specific models to predict individual assignments to communities, thereby characterizing variants and determining their origins and distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional allele frequency modeling methods are used to study population structure, then historical migration patterns can be elucidated, but the ability to characterize specific genetic variants and their origins is limited
Solution Approach 1:
The patent segments the population genetics analysis into distinct components: IBD segment detection, cluster formation based on shared segments, and community assignment. This segmentation allows the system to characterize specific genetic variants and their population origins without requiring a complete overhaul of traditional methods, thereby reducing information loss while managing complexity.
Solution Approach 2:
The patent introduces IBD (identity-by-descent) segments as an intermediary mechanism to connect individuals to their ancestral populations. By using shared IBD segments as a mediator, the system can infer population structure and variant origins without directly observing historical migration events, thus preserving information while avoiding excessive methodological complexity.
2Measurement precision
If IBD analysis is performed to identify familial relationships and population structure, then insights into population structure can be gained, but the ability to predict individual community membership is insufficient
Solution Approach 1:
The patent implements a feedback mechanism where IBD analysis results inform cluster formation, which in turn refines community assignments. The system uses the detected IBD segments to identify clusters of related individuals, then uses these clusters to improve community membership predictions. This feedback loop enhances measurement precision while maintaining operational ease through automated iterative refinement.
Solution Approach 2:
The patent performs preliminary IBD analysis and cluster formation before final community assignment. By pre-identifying clusters of individuals who share recent common ancestry through IBD segments, the system prepares structured population groups that make the subsequent community assignment process more straightforward and accurate, thus improving precision without compromising ease of operation.
3Loss of information
If genetic data is analyzed to characterize variants of interest, then insights into phenotype etiology can be obtained, but the characterization of variant origins and distributions is incomplete
Solution Approach 1:
The patent merges multiple analysis streams: IBD segment detection, cluster formation, community assignment, and variant characterization. By combining these previously separate analytical steps into an integrated workflow, the system recovers information about variant origins and distributions that would be lost in isolated analyses, while improving overall productivity through streamlined processing.
Solution Approach 2:
The patent creates a universal framework that can characterize any genetic variant of interest by leveraging shared IBD segments and community structure. The same IBD-based clustering approach works across different variants and populations, making the variant characterization process more efficient while comprehensively capturing origin and distribution information through the multi-functional community assignment system.
Data Source
AI summary
Disclosed are techniques for characterizing variants of interest and predicting assignments of individuals to communities based on obtained genetic information. To characterize a variant, DNA datasets of reference individuals are accessed and used to generate a cluster with additional individuals. Reference individuals carry a variant at a genetic locus and the additional individuals share IBD with reference individuals. Statistics of genealogical data of the cluster are generated. A result summarizing the characterization of the variant is generated based on the statistics. To determine if an individual belongs to a community, a subset of the individual's haplotypes are inputted into a community-specific model. The model is trained using the training samples that each include haplotypes of reference individuals and a label identifying whether the reference individual belongs to the community. Based on the output of the model, it is determined whether the individual is a member of the community.


