Genomic Mixture Modeling for Disease Subgroup Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for categorizing patients based on disease subgroups using genomic data are challenging due to the complex genomic landscape of diseases, making it difficult to identify actionable and prognostic features for different subgroups.
Innovation Solution
The development of methods and systems that process genomic data from patients to identify disease subgroups by generating candidate best fit models, selecting the best fit model based on fit statistics, and applying it to patient data to determine subgroups and associated genomic profiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If comprehensive genomic profiling is performed on all patients, then the completeness of genomic data is improved, but the complexity of data analysis and computational resources required worsen
Solution Approach 1:
The patent segments the large-scale genomic data analysis problem into smaller, manageable components by dividing patients into subgroups based on shared genomic alterations. This segmentation reduces the complexity of analyzing all patient data simultaneously while maintaining comprehensive profiling capabilities, as each subgroup can be analyzed with targeted approaches rather than requiring uniform analysis of the entire cohort.
2Measurement precision
If the number of disease subgroups is increased to capture more genomic diversity, then the precision of subgroup classification is improved, but the sample size required for each subgroup worsens
Solution Approach 1:
The patent transitions from analyzing genomic data in a single dimension to multiple dimensions by considering various genomic alteration types (mutations, copy number variations, structural variants) and their combinations. This multidimensional approach enables precise subgroup classification based on complex genomic profiles while maintaining adequate sample sizes, as patients are classified across multiple genomic dimensions rather than requiring large samples for each individual alteration type.
3Measurement precision
If iterative model generation is performed to identify the optimal number of subgroups, then the accuracy of subgroup identification is improved, but the computational time required worsens
Solution Approach 1:
The patent applies preliminary action by using unsupervised learning algorithms to pre-identify potential disease subgroups and genomic patterns before performing iterative model generation. This preliminary clustering and pattern recognition reduces the search space for subsequent iterative optimization, enabling accurate subgroup identification while minimizing computational time by avoiding exhaustive exploration of all possible subgroup configurations.
Data Source
AI summary
Methods for identifying disease subgroups are described. The methods may comprise, for example, receiving subject data for a plurality of subjects diagnosed with the disease; creating a plurality of candidate best fit latent class or mixture models by: i) providing an estimate of a number of subgroups; ii) generating a set of models, each model of the set comprising the same estimate of the number of subgroups; iii) selecting a candidate best fit model from the set; and iv) repeating (i)-(iii) at least once using a different estimate of the number of subgroups to obtain a plurality of candidate best fit models; selecting a best fit model from the plurality of candidate best fit models based on a fit statistic; and applying the best fit model to the subject data to identify a number of subgroups for the disease and an associated genomic profile for each subgroup.


