Consensus Ancestry Classification from Comprehensive Tumor Profiling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing genomic profiling methods, such as single-gene and multi-gene panels, fail to provide comprehensive coverage, leading to missed clinically significant alterations and require invasive sample depletion, while whole-genome sequencing is limited by high costs and complex datasets. Self-identified race and ethnicity data lacks consistency and accuracy in reflecting genetic backgrounds, hindering precision medicine and healthcare equity for diverse populations.
Innovation Solution
A consensus-based classification technique using expanded reference datasets and multiple algorithms, including k-nearest neighbors, principal component correlation, and ADMIXTURE analysis, to infer genetically inferred ancestry from comprehensive genomic profiling, ensuring accurate and robust ancestry classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If comprehensive genomic profiling is performed using expanded reference datasets and multiple classification algorithms, then measurement precision of genetic ancestry inference is improved, but device complexity increases
Solution Approach 1:
The patent segments the ancestry inference process into multiple independent classification algorithms (k-nearest neighbors, principal component correlation, and ADMIXTURE analysis), each handling specific aspects of the analysis. This segmentation allows each algorithm to specialize in particular computational tasks, improving overall measurement precision while maintaining manageable complexity through modular design
Solution Approach 2:
The patent creates a composite classification system that integrates multiple different algorithms (k-nearest neighbors, principal component correlation, and ADMIXTURE analysis) working together. This composite approach combines the strengths of each individual algorithm, achieving superior measurement precision in genetic ancestry inference that exceeds what any single algorithm could provide alone
2Reliability
If multiple classification algorithms are used to determine consensus ancestry calls, then reliability of ancestry classification is improved, but loss of time in processing increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing the genomic data into standardized formats and pre-computing reference profiles before the actual classification process. This preliminary preparation enables the multiple classification algorithms to operate more efficiently, reducing the time penalty associated with using multiple algorithms while maintaining high reliability through consensus calling
3Ease of operation
If self-identified race and ethnicity data is used for ancestry classification, then ease of operation is improved, but measurement precision deteriorates
Solution Approach 1:
The patent introduces an intermediary computational layer that processes self-identified race and ethnicity data through multiple classification algorithms. This intermediary transformation converts subjective self-identification into objective genetic ancestry classifications by comparing genomic profiles against reference datasets, thereby improving measurement precision while maintaining the ease of initial data collection
Data Source
AI summary
The disclosure relates to comprehensive genomic profiling (CGP) and to consensus-based classification techniques for determining genetically inferred ancestry from CGP of tumor DNA. Aspects are directed towards accessing reference and subject sequencing files and identifying genomic variants using a hybrid variant tool. The reference variant file is consolidated into a datastore formatted file that is queried to perform joint variant calling to generate a final reference variant file. The final reference variant file and the subject variant file are merged. On the merged variant file, principal component (PC) analysis is performed, and the PCs are used by a first and second classification process to generate a first and second ancestry call. The merged variant file is input into a third classification process to generate a third ancestry call. A consensus genetically inferred ancestry (GIA) call is predicted based on the first, the second, and the third ancestry calls.


