Prediction Model Incorporating Data Groupings
Through principal component analysis and Bayesian method, a weighted prediction model is constructed, which solves the problem of reduced prediction accuracy caused by the increase in data set samples and achieves more efficient prediction performance.
Patent Information
- Application Number
- CN202011104920.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-15
- Filing Date
- 2020-10-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-17
AI Technical Summary
As the number of samples of data sets increases, the incremental heterogeneity in the data structure leads to the reduction of the accuracy of traditional prediction algorithms, and it is impossible to effectively process data sets with significant differences between samples.
Principal component analysis was used to group samples and predictive models were constructed based on different groups. The population hierarchical structure probability of the test sample was calculated by Bayesian method, using these probabilities as weights, and the prediction results of multiple prediction models were weighted and summed as the final decision.
Improve prediction accuracy on data sets with different sample groups, and more effectively utilize the endogenous structure in the data set, and improve the performance of the prediction model.
Smart Images

Figure CN112669908B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of U.S. Provisional Application No. 62 / 915,459, filed on October 15, 2019. Background of the Invention
[0003] The present disclosure generally relates to the prediction of results, and particularly to prediction models incorporating data packets.
[0004] Accurate prediction models have important guiding significance in many fields. For example, in the medical field, the best recommendations related to cancer screening (e.g., the frequency of implementing screening and / or which screening tests to implement) can be proposed based on the cancer risk of a specific patient. Moreover, if a patient has a specific disease, the optimal treatment plan can be selected based on the prediction results.
[0005] Traditionally, techniques such as linear or logistic regression are used to generate predictions based on one or more independent variables. In traditional methods, a research team designs a study to test a specific hypothesis that a specific variable (or set of variables) is related to a specific outcome, and then collects a sample size sufficient to test the hypothesis, where the size is predetermined based on the expected effect size, potential confounding variables to be controlled, etc.
[0006] Recently, machine learning has made personalized prediction possible, especially when faced with a large number of potentially relevant variables. Machine learning classifiers are typically given a large number of "training" samples, where both the variables and the outcomes in the dataset are known. Known training procedures are used to train the classifier to optimize the objective function. Generally, the training of machine learning classifiers is a dynamic process, and as new samples are added to the training dataset, the classifier is retrained to utilize the new information. Summary of the Invention
[0007] As the number of samples in the dataset increases, the differences in data structures among the samples become more and more obvious. This increasing heterogeneity leads to a decrease in the accuracy of prediction algorithms that assume "the entire training dataset is a homogeneous population". For example, strong predictor variables for some groups may contribute little to other samples.
[0008] Certain embodiments of the claimed invention relate to techniques applicable to prediction for population stratification. The method of principal component analysis is used to group samples according to the data structure, and prediction models are constructed based on different groups. For test samples, the probability of belonging to different groups is calculated based on the Bayesian method according to their population stratification structure, and this probability is used as a weight to perform a weighted sum of the prediction results of multiple prediction models as the final decision.
[0009] The techniques described herein can be applied to any dataset where there are differences between sample groups. While the examples described herein relate to disease prediction using genomic data, similar techniques can also be applied in other contexts. For example, in the healthcare field, the data can include biomarkers other than genomic data (e.g., blood chemistry data; medical imaging data; biometric parameters such as heart rate or blood pressure; family medical history; behavioral parameters (such as diet or exercise), and the prediction can involve diagnosis (e.g., the presence or absence of a particular disease), the likelihood of developing a disease, the expected response to a particular treatment course, etc. The techniques described herein can also be applied to other fields such as finance (e.g., predicting future investment returns or the likelihood of loan default), insurance (e.g., predicting the possible value of future claims of an insured), etc.
[0010] The following detailed description and the accompanying drawings will provide a better understanding of the nature and advantages of the claimed invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A flowchart showing a process for predicting the likelihood of an outcome according to an embodiment of the present invention.
[0012] Figure 2 Shows a flowchart of a process for grouping a training set that can be used with the process of Figure 1 in some embodiments of the present invention.
[0013] Figure 3 Shows a flowchart of a process for calculating a prediction result that can be used with the process of Figure 1 in some embodiments of the present invention.
[0014] Figures 4A - 4D Four figures showing the results of applying the process according to an embodiment of the present invention to a simulated dataset.
[0015] Figure 5 Is a bar graph showing the results of applying the process according to an embodiment of the present invention to a simulated dataset.
[0016] Figure 6 Is a graph showing the receiver operating characteristic (ROC) curve of Alzheimer's disease data using the process according to an embodiment of the present invention and global logistic regression analysis.
[0017] Figure 7 Is a graph showing the ROC curve of schizophrenia data using the process according to an embodiment of the present invention and global logistic regression analysis. DETAILED DESCRIPTION
[0018] To provide an understanding of the various features of the claimed invention, embodiments are described in which genomic data is used to predict the likelihood that an individual will develop a particular disease. However, it should be understood that the same techniques can be applied to other types of data, and the invention is not limited to genomic data, disease prediction, or the healthcare field.
[0019] Prism Vote process
[0020] Figure 1 FIG. 7 shows a flowchart of a process 100 for predicting the likelihood of an outcome according to an embodiment of the present invention. Process 100 can be implemented using a suitably designed computer system.
[0021] In block 102, a training set of data samples is identified. The training set contains N individual data samples. For each data sample x i , which contains p independent variables {x ij} (for j = 1, …, p) and whose dependent variable (sample disease status) y i is known. For example, the set of variables {x ij} can represent p different single nucleotide polymorphisms (SNPs). For each SNP, the variable x ij takes on a value of 0, 1, or 2, corresponding to the number of minor alleles in the genotype. For example, if G is the minor allele and the observed genotype is GG, the SNP value is encoded as 2. If the observed genotype is CC, the SNP value is encoded as 0. In disease prediction, the dependent variable y i can indicate the presence (y i = 1) or absence (y i = 0) of the disease. In cases such as predicting variable physical characteristics (e.g., blood glucose level or cholesterol level), the outcome y i can be a variable with continuous values. Depending on the specific information represented in the data sample x i , other coding schemes can be used.
[0022] In block 104, the training set data samples are divided into a plurality of groups. The number of groups can be selected based on the sample size (i.e., the number of samples N) and the minimum number of samples per group (C). In some embodiments, the number of groups (K) can be selected within the range 2 ≤ K ≤ N / C, depending on factors such as the number of independent variables. Machine learning classifiers may require an even larger training set to produce a reliable prediction model, especially if the number of variables is large. Examples of techniques that can be used to optimize the number of groups for a given training data set are described below.
[0023] Figure 2Displays a flowchart of process 200 for splitting or grouping a training set (which can be implemented at block 104 of process 100). Process 200 involves using a matrix representation of training data and elements of principal component analysis to define similarity.
[0024] At block 202, from the training sample matrix X. In some embodiments, each row of matrix X may correspond to a data sample x i and each column corresponds to an independent variable. Thus, for a training data set of N samples (each sample having p variables), X is an N×p matrix. Depending on the specific combination of variables, it may be necessary to perform a standardization operation on each column so that all variables are within a similar numerical range.
[0025] At block 204, the eigenvalues and eigenvectors of matrix X can be calculated. Specifically, the eigenvalues λ j and eigenvectors v j (for j = 1,…N) can be calculated according to the following:
[0026] XX′v j = λ j v j (1)
[0027] where X′ is the matrix transpose of X. Assume the eigenvalues are sorted in descending order of magnitude.
[0028] At block 206, select the first q eigenvectors v with the largest eigenvalues j to stratify the training sample set. The specific value of q can be determined using a scree plot or similar techniques.
[0029] At block 208, calculate the weighted average of the eigenvectors according to the following q eigenvectors: the g vector:
[0030]
[0031] where
[0032]
[0033] The g vector is an N-dimensional vector that summarizes the distribution of the N data samples along the first q eigenvectors, and each element in g corresponds to a different data sample.
[0034] At block 210, the g vector can be used to partition the training data into K groups. Specifically, the components of the g vector can be sorted by magnitude, and the sample groups can be partitioned using the quantiles of g. For example, if K = 2, the data samples corresponding to the first N / 2 components of the (sorted) g can be assigned to one group, and the remaining ones to another group. If K = 20, each quantile can include N / 20 data samples. (Since there are eigenvalue corresponding to each data sample, this also means that part of the eigenvector is assigned to one group)
[0035] In some embodiments, the same number of data samples (e.g., N / K) can be assigned to each group. If N / K is not an integer, rounding techniques can be used. In other cases, the partitioning can be unequal, e.g., following the natural clustering of the data samples as shown in the principal component plot. Different groups can include different numbers of data samples without limitation as long as each group includes at least a minimum number of data samples to support the training of the prediction model for that group.
[0036] In the case where the variables {x ij} represent genomic variations (e.g., where the variables represent SNPs), the eigenvectors can be interpreted as ancestral directions. Objects with high variation among the first q eigenvectors are genetically closer and grouped together. In this context, process 200 can be understood as a method for decomposing and integrating heritability as a spectrum. In other cases, the interpretation can be different, but the general approach is to form clusters within a set of data samples, where the clusters reflect the endogenous structure in the dataset.
[0037] At block 212, the center of each group can be calculated as the mean of the top q eigenvectors within that group. Specifically, for the k-th group (where k = 1, …, K), the center can be defined as the q-dimensional vector c k
[0038]
[0039] where i0 corresponds to the first data sample of the group, v ij is the i-th dimension of the j-th eigenvector, and j = 1, …, q. As described below, the vector c k can be used to calculate the probability that a new test sample belongs to the k-th group.
[0040] Referring again to Figure 1 , after the training data has been partitioned into several groups, at block 106, prediction models are independently trained within the groups. In some embodiments, the same prediction model is used for all groups, but since the training datasets are different, the same independent variables will have different effect size estimates, so the prediction models for different groups will produce different prediction results.
[0041] For example, the prediction model can be a linear regression model that predicts the result as a variable y of continuous values. The linear regression model for a training set of N data samples with p variables can be expressed as:
[0042] Y = Xβ + ε (5)
[0043] where are the dependent variable observations, is the independent variable observation matrix, and are the parameters to be estimated. {ε1, ε2,..., ε N} are independent and identically distributed, with a mean of 0 and a variance of σ 2 . Techniques for calculating the parameters of, for example, a linear regression model from training data are known in the art and can be applied in the context of process 100.
[0044] According to block 106 of process 100, the linear regression model of equation (5) is applied to each group separately. That is, there are K models of this form, instead of a single model:
[0045] Y k = X k β k + ε k (6)
[0046] where and k = 1,..., K.
[0047] Each component of ε k is independently and identically distributed, following a normal distribution with a mean of 0 and a variance of (in this example, the variance is assumed to be the same for all k).
[0048] Although linear regression is used as an example, it should be understood that process 100 is not limited to any specific prediction model. Other prediction models can be used, such as logistic regression models, non-linear models, support vector machine (SVM) models, deep learning models, provided that a sufficient sample size is available for training each group's prediction model independently. Depending on the specific prediction model, training includes calculating linear regression (e.g., using equation 6), applying machine learning algorithms to train a deep learning classifier, or any other technique for training a specific prediction model based on the available training data.
[0049] The trained prediction model can be used to make predictions. For example, in block 108 of process 100, "test" sample(s) can be obtained. As used herein, the test sample s can be one that was not used in training the prediction model and for which the relevant variables x s = {x sjAny data sample of}. In some cases, the dependent variable (y s ) of the test sample s is unknown (e.g., when using a trained prediction model in clinical practice to predict a patient's outcome). In other cases, the outcome y s may be known (e.g., when testing a trained prediction model to evaluate its performance).
[0050] In block 110, the final prediction result of the test sample can be calculated based on the prediction model of each group and the probability that the test sample belongs to that group. The prediction model of each group will predict the dependent variable of the test sample, and finally, the weights based on the probability that the test sample belongs to each group will be used to combine the predictions of each group.
[0051] Figure 3 Shows a flowchart of process 300 for calculating the prediction result that can be used in block 110. In block 302, for each group k, each group prediction (y k ) is calculated based on the assumption that the test sample belongs to group k. For example, the (known) variables {x sj} associated with the test sample s can be provided as the input to the prediction model trained for each group, and the prediction model can be applied to calculate the predicted result.
[0052] In block 304, for each group k, the probability of each group that the test sample s belongs to that group can be determined. In some embodiments, the probability that the test sample belongs to a particular group can be calculated based on the distance to the center of that group and Bayesian methods.
[0053] For example, the center of the group can be defined according to equation (4) above. The distance to the center can be calculated as follows. First, the eigenvector of the test set including the test sample s is calculated using the matrix XX' including the test sample s. (The matrix XX' can also include some or all of the training samples). Thus, the eigenvector v s of the test sample s can be determined. The distance between v s and the center c k (defined according to equation (4) above) of group k can be calculated as:
[0054]
[0055] where
[0056]
[0057] and i = 1, …, n k . The off-diagonal terms are 0 because the eigenvectors are orthogonal. Under the assumption that the sample s belongs to group k, observing a variable x s= {x sj} The (tail) probability of the sample s is:
[0058]
[0059] The closer the sample is to the center of the group, the greater the tail probability. Pr(x s | s ∈ k) can be used as a similarity measure of the sample s to the group k. Equation (9) can be used in Bayesian analysis to determine the probability that the sample s belongs to the k-th group given the variable x s .
[0060] In block 306, the prediction result of the sample s is calculated based on each set of prediction results y k determined in block 302 and the probability that the sample s belongs to the k-th group. (Note that the training of the prediction model is assumed not to include the test sample s). The Bayesian model can be used to calculate the prediction. For example, assume that the prediction model has two results: y = 1 corresponds to a positive disease, and y = 0 corresponds to a negative disease. The predicted probability of the disease for the object s is the combined prediction from each group:
[0061]
[0062] where Pr(y s = 1 | s ∈ k, x s ) is the result of the prediction model from the k-th group. The probability Pr(s ∈ k | x s ) can be given by:
[0063]
[0064] where Pr(s ∈ k) = n k / N is the proportion of the k-th group training sample size to the total training sample size, which is also the prior probability that the data sample belongs to the group k. Pr(x s | s ∈ k) is the probability of observing x s when the data sample belongs to the group k, as defined in Equation (9).
[0065] Process 100 is illustrative, and variations and modifications are possible. For example, any type of prediction model can be used, including but not limited to linear regression models. (It is assumed that the same type of prediction model and the same set of input variables are used for all groups). Process 100 represents a class of processes where the training data set is grouped and prediction models for different groups are trained, and where the prediction for the test sample is made by combining the predictions from different groups according to the probability of the test sample in a specific group. This process is referred to as the "Prism Vote" process in this article.
[0066] The Prism Vote process, such as process 100, can be implemented by a computer and can operate on data sets of any size with any number of variables. A set of predictive models and associated parameters (e.g., eigenvalues, eigenvectors, group centers) can be generated, for example, by implementing blocks 102-106, and the predictive models and parameters can be stored for later use. Obtaining a test sample and calculating a predictive result for the test sample (e.g., blocks 108 and 110) can be implemented at any time after the predictive model has been trained and stored, and the same set of predictive models can be applied to any number of test samples. In addition, the trained predictive models and associated parameters can be provided to a computer system other than the system used to train the predictive models. For example, a Prism Vote process for predicting a disease can be trained by a research team, and the trained predictive models and associated parameters can be assigned to a clinician (e.g., in a laboratory). The clinician can apply the trained predictive models and associated parameters to data collected from individual patients, e.g., by using a computer to calculate a predictive result for a given patient. The predictive result can be used to inform treatment decisions or other care recommendations (e.g., dietary or lifestyle changes).
[0067] In some embodiments, if, for each group, the probability that a test sample belongs to that group (e.g., as determined from Equation 11) is less than a threshold (e.g., 0.05), the test sample can be classified as an outlier. For outliers, a "global" predictive model can be trained using all the training samples (without other partitioning of the groups), and the global predictive model can be used to determine the predictive result for the outlier. In the case where a test sample is identified as an outlier, the predictive result report can be annotated to indicate that the test sample is an outlier.
[0068] Performance of the Prism Vote Process
[0069] The Prism Vote method, such as process 100, develops predictive models applicable to different groups of population-stratified data. However, the partitioning of the data set reduces the amount of training samples for the predictive models within each group, and the reduction in the sample size may lead to less reliable predictive results. To understand when the Prism Vote process can be expected to provide more reliable predictions than a "global" predictive model (i.e., a single predictive model trained using all the training data), the expected prediction error (EPE) of different methods can be considered. For illustrative purposes, assume that the predictive model is a linear regression model as described above. Similar analysis can be applied to other predictive models.
[0070] For the "global" linear regression model trained on all the training data samples The EPE can be expressed as:
[0071]
[0072] For the K-group Prism Vote process with a linear regression model for each group, the EPE can be expressed as:
[0073]
[0074] where is the prediction from the regression model for the k-th group of the sample X, and w k (X) is the weight of the prediction for the k-th group of the sample X.
[0075] Assuming a linear regression model, when EPE PV (x s ) < EPE LR (x s ), it is expected that the Prism Vote process is more accurate than the global model prediction for this test sample x s . The derivation of the inequality is summarized as PVI(x s ) > 0, where PVI(x s ) is given by:
[0076]
[0077] where is the least squares estimate from the k-th group according to process 100, is the least squares estimate using the global linear regression model (trained on all samples). The smallest PVI value in the test dataset is taken to measure whether to implement Prism Vote for the entire test data.
[0078]
[0079] When equation (15) yields PVI > 0, for all test samples, it is expected that using the Prism Vote process that independently trains each group with a linear regression model can provide better prediction performance than a single linear regression model trained on all training samples.
[0080] In addition, when linear regression is used as the prediction model, the optimal total number of groups K can be calculated as:
[0081]
[0082] It should be understood that K values other than those indicated by equation (16) can also be selected, even if the performance is suboptimal. In addition, if the prediction model is not a linear regression model, similar logic can be used to define the conditions under which the expected Prism Vote process outperforms a single model and / or to determine the optimal number of groups.
[0083] Example 1
[0084] For illustration of the selection of the number of groups, four data sets were simulated for the study. In each data set, the "true" total number of groups was taken from (1, 2, 3, or 4), and it was assumed that each group followed a different linear regression prediction model to generate the data. The Prism Vote model was trained on each data set with different values of K (K = 1, 2, 3, 4, 5).
[0085] Figures 4A - 4D Figures 410, 420, 430, and 440 corresponding to the four data sets are shown. In Figure 4A , the true number of groups is 1 (traditional linear regression); in Figure 4B , there are 2 groups; in Figure 4C , there are 3 groups; in Figure 4D , there are 4 groups. For each data set, the PVI (lines 411, 421, 431, 441) defined by the above equation (14) and the mean squared error (MSE) measurement -1*MSE (lines 412, 422, 432, 442) are shown as a function of the number of groups K used in the Prism Vote process. In each case, the PVI is maximized when K is chosen to be the true number of groups, and the value of K that produces the maximum PVI also provides the minimum MSE, as indicated by the above equation (16).
[0086] Example 2
[0087] To illustrate the performance of Prism Vote, a simulation study has been carried out using two populations with different phenotypes. Data for five different scenarios were generated. Each scenario used the same set of predictors (variables) and linear regression model, but differed in the effect differences (expressed as the mean β difference) of various independent variables between the two populations. In scenario 1, there was no difference in the effect sizes between the two populations (mean β difference was 0); in scenarios 2 - 4, there were increasing differences in the effect sizes (mean β differences were 0.18, 0.4, 0.67); in scenario 5, the effects were completely different (mean β difference was 1). The resulting PVIs for scenarios 1 - 5 were 0.27, 0.10, 0.76, 1.46, and 3.23. For each scenario, the mean squared error (MSE) of the global linear regression model (trained on all data samples) and the Prism Vote process with K = 2 was calculated.
[0088] Figure 5 is a bar chart comparing the predicted MSEs in each scenario. For each scenario, the results of the global (traditional) linear regression model are on the left, and the results of the Prism Vote process are on the right. From Figure 5As can be seen, when there is a difference in effect size, the Prism Vote process can provide a reduced MSE (better performance) compared to the conventional process.
[0089] Example 3
[0090] The Prism Vote process using a logistic regression prediction model for each group has been applied to two genome-wide datasets related to Alzheimer's disease and schizophrenia, respectively. A global logistic regression prediction model has also been applied to the same datasets for comparison. Each trained model has been applied to test data with known results in order to evaluate sensitivity and specificity.
[0091] Figure 6 is a graph showing the receiver operating characteristic (ROC) curves of Alzheimer's disease data for the Prism Vote process (line 602) and global logistic regression (line 604). In the case of the Prism Vote process, the average area under the curve (AUC) in 5-group cross-validation (5GCV) reached 74.36%, representing a 3.5% improvement over conventional logistic regression.
[0092] Figure 7 is a graph showing the ROC curves of schizophrenia data for the Prism Vote process (line 702) and global logistic regression (line 704). In the case of the Prism Vote process, the 5GCV average AUC is 68.2%, representing a 3.1% improvement compared to conventional logistic regression.
[0093] These examples show that the Prism Vote process of the type described herein can be used to improve the accuracy of predictions. It should be understood that these embodiments are illustrative rather than restrictive. Performance can depend on the particular set of variables and outcomes being modeled, as well as on the prediction model used, the number of groups, and the size of the dataset.
[0094] Computer system implementation
[0095] Data analysis and computational operations of the type described herein can be implemented in a computer system that is typically of conventional design, such as a desktop computer, a tablet computer, a mobile device (e.g., a smart phone), etc. Such a system can include one or more processors that execute program code (e.g., a general-purpose microprocessor that can serve as a central processing unit (CPU) and / or a specialized processor such as a graphics processing unit (GPU) that can provide enhanced parallel processing capabilities); a memory and other storage devices that store the program code and data; user input devices (e.g., a keyboard, a pointing device such as a mouse or a touchpad, a microphone); user output devices (e.g., a display device, a speaker, a printer); combined input / output devices (e.g., a touch screen display); signal input / output ports; a network communication interface (e.g., a wired network interface such as an Ethernet interface and / or a wireless network communication interface such as Wi-Fi); etc. Computer programs incorporating the various features of the claimed invention can be encoded and stored on various computer-readable storage media; suitable media include magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, and other non-transitory media. (It should be understood that the "storage" of data is different from the propagation of data using transitory media such as carrier waves). The computer-readable medium encoded with the program code can be packaged together with a compatible computer system or other electronic device, or the program code can be provided separately from the electronic device (e.g., downloaded via the Internet or as a separately packaged computer-readable storage medium).
[0096] As described above, the training of the prediction model and the application of the trained prediction model to the training data can be performed at different times and / or by different computer systems or the same computer system. Additionally, when new training data is available, the training portion of the PrismVote process can be repeated from time to time.
[0097] Additional embodiments
[0098] Although the invention has been described with reference to specific embodiments, those skilled in the art will understand that changes and modifications can be made. All of the processes described above are illustrative and can be modified. Processing operations described as separate boxes can be combined, the order of operations can be modified to the extent permitted by logic, the processing operations described above can be changed or omitted, and additional processing operations not specifically described can be added. Specific definitions and data formats can be modified as needed.
[0099] In various embodiments, any type of predictive model and any number of groupings can be used to implement the Prism Vote process, provided that sufficient data is available for training each grouped predictive model. The Prism Vote process can group data based on, for example, the endogenous structure of a dataset such as described above, thereby enhancing predictive performance for a particular type.
[0100] Furthermore, while the foregoing embodiments refer to disease prediction using genomic data, it should be understood that this is merely an example application of the Prism Vote process. For example, genomic data can be used to predict any phenotypic trait, including disease (or its absence), response to treatment, expected physiological characteristics (such as blood sugar or cholesterol levels), effectiveness of preventive measures, etc. Similarly, the independent variables are not limited to genomic data. In a healthcare context, any quantifiable information about an individual related to a medical condition can be used as a variable. Examples include medical imaging data, blood test results, stress test results, etc.
[0101] The applicability of the Prism Vote process is not limited to the healthcare field either. For example, in the financial sector, similar techniques can be used to predict the performance of publicly traded stocks based on multiple variables, where the dataset can be heterogeneous in terms of factors such as industrial sector, dependence on weather or commodities.
[0102] Accordingly, although the invention has been described with reference to specific embodiments, it should be understood that the invention is intended to cover all modifications and equivalents within the scope of the appended claims.
Claims
1. A computer-implemented method for non-diagnostic purposes of predicting the likelihood of a result based on a set of variables, the method comprising: Identifying a training set of data samples, wherein each data sample in the training set includes a plurality of variables and a known result, the plurality of variables indicating the presence or absence of each of a plurality of single nucleotide polymorphisms (SNPs) in a subject's genome, and the known result indicating the phenotypic characteristics of the subject; Partitioning the data samples of the training set into a plurality of groups based on a measure of similarity of the data samples; Training a prediction model for each group, wherein the prediction model predicts the likelihood of a result based on the variables, and wherein the training of the prediction model is performed independently for each group; For a patient, obtaining a test sample for which the variables are known; And Predicting the phenotypic characteristics of the patient based on the test sample, wherein predicting the phenotypic characteristics of the patient includes: For each group, using the prediction model of the group to determine the probability of the result; For each group, determining the probability that the test sample belongs to the group; and Calculating a predicted result of the test sample based on the probabilities of each group weighted by the probability that the test sample belongs to the group, wherein the predicted result represents the predicted phenotypic characteristics of the patient.
2. The method according to claim 1, wherein partitioning the data samples of the training set includes: Establishing a matrix of the training set of data samples; Calculating a set of eigenvalues and a set of eigenvectors from the matrix; Sorting the eigenvectors based on the respective magnitudes of the eigenvalues; And Using the sorted eigenvectors to partition the data samples of the training set.
3. The method according to claim 2, wherein using the sorted eigenvectors to partition the data samples of the training set includes: Selecting a subset of the sorted eigenvectors as significant eigenvectors; Calculating a weighted average vector of the significant eigenvectors, wherein the weighted average vector uses weights determined according to the eigenvalues; Sorting the components of the weighted average vector; And Assigning each data sample from the training set to one of the groups using quantiles of the weighted average vector.
4. The method according to claim 1, further comprising: Calculating a center of each of the plurality of groups.
5. The method according to claim 4, wherein for each group, determining the probability that the test sample belongs to the group includes calculating a distance metric between the test sample and the center of the group.
6. The method according to claim 1, wherein the predicted result of the test sample is calculated based on a Bayesian model.
7. The method according to claim 1, wherein the prediction model for each group is a generalized linear model.
8. The method according to claim 1, wherein the phenotypic characteristics correspond to physiological characteristics.
9. A computer system, comprising: A memory; And A processor connected to the memory and configured to: A training set of identified data samples, wherein each data sample in the training set includes a plurality of variables and a known result, the plurality of variables indicating the presence or absence of each of a plurality of single nucleotide polymorphisms (SNPs) in a subject's genome, and the known result indicating the phenotypic characteristics of the subject; Partitioning the data samples of the training set into a plurality of groups based on a measure of similarity of the data samples; Training a prediction model for each group, wherein the prediction model predicts the likelihood of a result based on the variables, and wherein the training of the prediction model is performed independently for each group; For a patient, obtaining a test sample for which the variables are known; And Predicting the phenotypic characteristics of the patient based on the test sample, wherein predicting the phenotypic characteristics of the patient includes: For each group, using the prediction model of the group to determine the probability of the result; For each group, determining the probability that the test sample belongs to the group; and Calculating a prediction result of the test sample based on the probabilities of each group weighted by the probability that the test sample belongs to the group, wherein the prediction result represents the predicted phenotypic characteristics of the patient.
10. The computer system of claim 9, wherein the processor is further configured to partition the data samples of the training set, which includes: Establishing a matrix of the training set of data samples; Calculating a set of eigenvalues and a set of eigenvectors from the matrix; Sorting the eigenvectors based on the respective magnitudes of the eigenvalues; And Using the sorted eigenvectors to partition the data samples of the training set.
11. The computer system of claim 10, wherein the processor is further configured to use the sorted eigenvectors to partition the data samples of the training set, which includes: Selecting a subset of the sorted eigenvectors as significant eigenvectors; Calculating a weighted average vector of the significant eigenvectors, wherein the weighted average vector uses weights determined according to the eigenvalues; Sorting the components of the weighted average vector; And Assigning each data sample from the training set to one of the groups using quantiles of the weighted average vector.
12. The computer system of claim 9, wherein the processor is further configured to: Calculate a center of each of the plurality of groups, wherein for each group, determining the probability that the test sample belongs to the group includes calculating a distance metric between the test sample and the center of the group.
13. The computer system of claim 9, wherein the prediction result of the test sample is calculated based on a Bayesian model.
14. The computer system of claim 9, wherein the prediction model for each group is a generalized linear model.
15. The computer system of claim 9, wherein the phenotypic characteristics correspond to one or more of a physiological characteristic, the presence or absence of a disease, or a response to treatment.
16. A computer-readable storage medium having program code instructions stored therein, the program code instructions causing the computer system to implement the following method when executed by a processor of the computer system, the method comprising: Identifying a training set of data samples, wherein each data sample in the training set includes a plurality of variables and a known result, the plurality of variables indicating the presence or absence of each of a plurality of single nucleotide polymorphisms (SNPs) in a subject's genome, and the known result indicating the presence or absence of a disease; Partitioning the data samples of the training set into a plurality of groups based on a measure of similarity of the data samples; Training a prediction model for each group, wherein the prediction model predicts the likelihood of a result based on the variables, and wherein the training of the prediction model is performed independently for each group; For a patient, obtaining a test sample for which the variables are known; And Predicting whether the patient has a disease based on the test sample, wherein predicting whether the patient has a disease includes: For each group, using the prediction model of the group to determine the likelihood of the result; For each group, determining the probability that the test sample belongs to the group; and Calculating a prediction result for the test sample based on the probabilities of each group weighted by the probability that the test sample belongs to the group, wherein the prediction result predicts whether the patient has the disease.
17. The computer-readable storage medium of claim 16, wherein partitioning the data samples of the training set includes: Establishing a matrix of the training set of data samples; Calculating a set of eigenvalues and a set of eigenvectors from the matrix; Sorting the eigenvectors based on the respective magnitudes of the eigenvalues; And Using the sorted eigenvectors to partition the data samples of the training set.
18. The computer-readable storage medium of claim 17, wherein using the sorted eigenvectors to partition the data samples of the training set includes: Selecting a subset of the sorted eigenvectors as significant eigenvectors; Calculating a weighted average vector of the significant eigenvectors, wherein the weighted average vector uses weights determined according to the eigenvalues; Sorting the components of the weighted average vector; And Assigning each data sample from the training set to one of the groups using quantiles of the weighted average vector.
19. The computer-readable storage medium of claim 16, further comprising: Calculating a center for each of the plurality of groups; Wherein for each group, determining the probability that the test sample belongs to the group includes calculating a distance metric between the test sample and the center of the group.
20. The computer-readable storage medium of claim 16, wherein the prediction result for the test sample is calculated based on a Bayesian model.
21. The computer-readable storage medium of claim 16, wherein the prediction model for each group is a generalized linear model.