Ethnicity Subregion Assignment Using Clustered Inheritance Segments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing genetic research methods struggle to accurately assign individuals to ethnicities and communities of origin, particularly for groups lacking historical records and relying on short matching-data segments, which are often filtered due to noise, limiting the ability to trace ancestry and connect with ethnic roots.
Innovation Solution
A computer-implemented method using unsupervised clustering, such as the Louvain method, to generate and assign individuals to ethnicities based on short matching-data segments, optimizing parameters like match segment lengths and total match lengths to reduce noise, and employing a reference panel with metadata for accurate categorization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If short matching-data segments are used for ethnicity assignment, then the ability to trace ancestry and connect with ethnic roots is improved, but noise in the data increases leading to inaccurate classifications
Solution Approach 1:
The patent combines multiple short matching-data segments into aggregated match scores and cumulative length metrics. By merging numerous small DNA segment matches into composite reliability scores, the system能够利用原本会被过滤的短段匹配数据,同时保持分类准确性。
Solution Approach 2:
The patent changes the parameters used for ethnicity assignment from traditional long-segment requirements to optimized parameters that include cumulative length of matches, number of matches, and match scores. This parameter transformation allows short segments to contribute meaningfully to ethnicity determination while maintaining reliability through mathematical aggregation.
2Reliability
If short matching-data segments are filtered out due to noise, then data accuracy is improved, but the ability to assign ethnicity to individuals with limited historical records deteriorates
Solution Approach 1:
The patent converts the previously harmful noise from short matching segments into beneficial signal by developing scoring mechanisms that identify and weight genuine distant relationships. What was once considered error (short segment noise) is now transformed into useful information for assigning ethnicity to individuals with limited records, turning a weakness into a strength.
Solution Approach 2:
The patent introduces intermediary computational layers including match scoring algorithms, cumulative length calculations, and threshold-based filtering that mediate between raw short segment data and final ethnicity assignments. These intermediaries process and validate short segment matches, allowing them to be used reliably without directly propagating noise to the final classification.
3Reliability
If traditional ethnicity assignment methods are used, then data accuracy is maintained, but the precision and meaningfulness of connections to ethnic roots deteriorates
Solution Approach 1:
The patent segments the ethnicity assignment process into multiple independent evaluation dimensions: match score, cumulative length, number of matches, and threshold comparisons. This segmentation allows each dimension to be optimized independently and combined to produce highly precise assignments, moving beyond traditional single-criterion methods to achieve both accuracy and precision.
Solution Approach 2:
The patent adds multiple new dimensions to the ethnicity assignment problem beyond traditional methods. Instead of relying solely on long segment matches, the system evaluates cumulative length across many segments, total number of matches, and weighted scores. This dimensional expansion transforms a 1D problem into a multi-dimensional classification system, dramatically improving precision.
Data Source
AI summary
A computing device may receive an inheritance dataset of a target named entity. The device may access a plurality of clusters associated with a region, each cluster comprising inheritance data for a plurality of reference panel named entities. The device may determine that the inheritance dataset of the target named entity has at least a threshold amount of inheritance sequences that are classified to the region. The device may compare, for each cluster, the inheritance dataset of the target named entity to the reference panel named entities in the cluster to identify similarities and shared inheritance segments between the target named entity and the reference panel named entities. The device may determine, for each cluster, a metric based on the inheritance segments shared. The device may assign the target named entity to one or more ethnicities based on the comparison between the metric and the threshold specific to the cluster.


