A multi-omics integrated whole genome prediction method
By employing a multi-omics ensemble approach, we constructed genome-phenotype, metabolomics-phenotype, and transcriptomics-phenotype models and performed 10-fold cross-validation, which solved the problem of insufficient prediction accuracy for complex traits in existing technologies and enabled high-precision prediction of rice hybrid phenotypes.
Patent Information
- Application Number
- CN202510165371.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Existing genome-wide selection methods are insufficient to effectively capture complex gene interactions and their downstream regulation, resulting in insufficient accuracy in predicting complex traits. Furthermore, multi-omics integration methods have failed to significantly improve predictive power.
A multi-omics ensemble approach was adopted, including acquiring parental genome, metabolome, and transcriptome data, filtering out low-frequency and high-deletion-rate SNPs, constructing genome, metabolome, and transcriptome matrices after standardization, training genome-phenotype, metabolome-phenotype, and transcriptome-phenotype models with 10-fold cross-validation, and making predictions using linear regression models.
It significantly improved the prediction accuracy of rice hybrid phenotypes and enhanced the predictive power of multi-omics, especially the predictive power of genome-metabolomics, genome-transcriptomics and genome-metabolomics-transcriptomics integration, which increased by 48.4%, 33.3% and 45.8% respectively.
Smart Images

Figure CN120108492B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of plant breeding, in particular to a whole genome prediction method based on multi-omics integration. BACKGROUND
[0002] The whole genome selection is abbreviated as GS. The GS breeding is to construct a statistical model by using the correlation between the genetic markers of the whole genome and the phenotype, and then to predict the individuals with unknown phenotype. Compared with the traditional plant breeding method, the GS breeding can improve the prediction and genetic gain. However, even if the complete gene sequence is available, the GS is still difficult to capture the complex interaction between genes and the downstream regulation, resulting in the bottleneck of the prediction accuracy of complex traits. At present, the low-cost high-throughput molecular technology develops rapidly, and the metabolites and transcripts can be accurately quantified. Since the metabolome and the transcriptome are between the genotype and the phenotype information flow, and contain rich site interaction information, the integration of these data is a way to break through the bottleneck of the GS prediction.
[0003] The existing multi-omics integration method, such as directly combining the genotype matrix, the metabolite matrix and the transcript matrix according to the column into a large matrix to input into the model for training and predicting the phenotype, has made certain progress. However, for some traits, the prediction ability has not been improved, and is only comparable to the prediction ability of a single omics, or even has been reduced. SUMMARY
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present application is to provide a whole genome prediction method based on multi-omics integration, so as to improve the prediction accuracy of the phenotype of rice hybrid.
[0005] In order to achieve the above-mentioned purpose, the present application provides the following scheme:
[0006] A whole genome prediction method based on multi-omics integration, comprising:
[0007] obtaining parent genomic data, parent metabolomic data, parent transcriptomic data and hybrid phenotype data;
[0008] filtering out the SNPs with the minimum allele frequency lower than 0.05 and the deletion rate greater than 0.2 in the parent genomic data to obtain preprocessed genomic data, and respectively performing standardization processing on the parent metabolomic data and the parent transcriptomic data to obtain preprocessed metabolomic data and preprocessed transcriptomic data;
[0009] speculating a hybrid genomic matrix, a hybrid metabolomic matrix and a hybrid transcriptomic matrix according to the preprocessed genomic data, the preprocessed metabolomic data and the preprocessed transcriptomic data, respectively;
[0010] The hybrid phenotype data, the hybrid genome matrix, the hybrid metabolome matrix and the hybrid transcriptome matrix are integrated and randomly divided into a training set and a test set;
[0011] The training set is used to train pre-constructed genome-phenotype prediction models, metabolome-phenotype prediction models and transcriptome-phenotype prediction models according to ten-fold cross-validation to obtain a training prediction matrix, and the test set is predicted by using the trained genome-phenotype prediction models, metabolome-phenotype prediction models and transcriptome-phenotype prediction models to obtain a test prediction matrix;
[0012] The training prediction matrix is input into a linear regression model for training, and the linear regression model is predicted according to the test prediction matrix to obtain a final hybrid phenotype prediction value.
[0013] Preferably, it further comprises:
[0014] The determination coefficient between the final predicted phenotype value and the observed phenotype value of the test set is calculated to obtain an evaluation index;
[0015] The mean of the evaluation index of several times of training is obtained to obtain a final prediction ability of the model.
[0016] Preferably, the ratio of the training set to the test set is 9:1.
[0017] Preferably, the expression of the hybrid genome matrix is:
[0018]
[0019] The expression of the hybrid metabolome matrix is:
[0020]
[0021] The expression of the hybrid transcriptome matrix is:
[0022]
[0023] G_m, M_m and T_m are the hybrid male parent genome matrix, the hybrid male parent metabolome matrix and the hybrid male parent transcriptome matrix, respectively; G_f, M_f and T_f are the hybrid female parent genome matrix, the hybrid female parent metabolome matrix and the hybrid female parent transcriptome matrix, respectively. Preferably, the model expression of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model is:
[0024] y train = 1 μ + O α + ε
[0025] wherein, i,j = 1,2,...,...n; y train is the phenotype observation vector of the hybrid in the training set; μ is the fixed effect; O is any one of the genomic matrix, the metabolomic matrix and the transcriptomic matrix of the hybrid in the training set; ε is the residual vector; h is the bandwidth parameter; p is any one of the number of SNPs, the number of metabolites and the number of transcripts; α is any one of the genomic random effect, the metabolomic random effect and the transcriptomic random effect; K is any one of the genomic Gaussian kernel, the metabolomic Gaussian kernel and the transcriptomic Gaussian kernel; is any one of the genomic variance, the metabolomic variance and the transcriptomic variance; D ij is the Euclidean distance between the i-th sample and the j-th sample of the hybrid in the training set; n is the number of hybrids in the training set.
[0026] Preferably, the expression of the training prediction matrix is:
[0027]
[0028] the expression of the test prediction matrix is:
[0029]
[0030] wherein, and respectively represent the predicted values of the genomic-phenotype prediction model, the metabolomic-phenotype prediction model and the transcriptomic-phenotype prediction model on the training set; and respectively represent the average of 10 predicted values of the genomic-phenotype prediction model, the metabolomic-phenotype prediction model and the transcriptomic-phenotype prediction model on the test set; k represents the k-fold data used as the test data of the model when training the model by cross-validation.
[0031] Preferably, the expression of the linear regression model is:
[0032]
[0033] wherein, q ∈ {1,2,3}; is the i-th phenotype observation of the hybrid based on the training set; b is the intercept term; w1, w2, w3 are respectively the first regression coefficient, the second regression coefficient and the third regression coefficient; ε is the error term; I A is the indicator function; is the i-th element in the first column.
[0034] The present application discloses the following technical effects:
[0035] The present application provides a whole genome prediction method based on multi-omics integration, which solves the problem that the prior art only improves the prediction of a single omics by constructing a genome-phenotype prediction model, a metabolome-phenotype prediction model, a transcriptome-phenotype prediction model, and a linear regression model and training the model using ten-fold cross-validation, and realizes the improvement of the prediction of the model. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0037] Figure 1 A whole genome prediction process diagram based on multi-omics integration is provided for the embodiments of the present application.
[0038] Figure 2 A prediction power box plot of 9 prediction models for 4 traits of rice hybrids is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0040] The purpose of the present application is to provide a whole genome prediction method based on multi-omics integration to improve the prediction accuracy of the phenotype of rice hybrids.
[0041] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0042] Figure 1 A whole genome prediction process diagram based on multi-omics integration is provided for the embodiments of the present application, as shown in Figure 1 The present application provides a whole genome prediction method based on multi-omics integration (MOEGS), which comprises:
[0043] Step 100: obtaining parent genome data, parent metabolome data, parent transcriptome data, and hybrid phenotype data;
[0044] Step 200: filtering out SNPs (single nucleotide polymorphisms) with minor allele frequency less than 0.05 and deletion rate greater than 0.2 in the parent genome data to obtain preprocessed genome data, and respectively normalizing the parent metabolome data and the parent transcriptome data to obtain preprocessed metabolome data and preprocessed transcriptome data;
[0045] Step 300: respectively inferring a hybrid genome matrix, a hybrid metabolome matrix and a hybrid transcriptome matrix according to the preprocessed genome data, the preprocessed metabolome data and the preprocessed transcriptome data;
[0046] Step 400: integrating the hybrid phenotype data, the hybrid genome matrix, the hybrid metabolome matrix and the hybrid transcriptome matrix and randomly dividing them into a training set and a test set;
[0047] Step 500: training the pre-constructed genome-phenotype prediction model, metabolome-phenotype prediction model and transcriptome-phenotype prediction model according to ten-fold cross-validation using the training set to obtain a training prediction matrix, and predicting the test set using the trained genome-phenotype prediction model, metabolome-phenotype prediction model and transcriptome-phenotype prediction model to obtain a test prediction matrix;
[0048] Step 600: training a linear regression model by inputting the training prediction matrix as a feature, and predicting the linear regression model according to the test prediction matrix to obtain a final hybrid phenotype prediction value.
[0049] Further, it further comprises:
[0050] calculating the determination coefficient between the final predicted phenotype value and the observed phenotype value of the test set to obtain an evaluation index;
[0051] obtaining the mean of the evaluation index of several times of training to obtain the final prediction ability of the model.
[0052] Preferably, the ratio of the training set and the test set is 9:1.
[0053] Further, the expression of the hybrid genome matrix is:
[0054]
[0055] The expression of the hybrid metabolome matrix is:
[0056]
[0057] The expression of the hybrid transcriptome matrix is:
[0058]
[0059] wherein G_m, M_m, T_m are hybrid father genome matrix, hybrid father metabolome matrix, hybrid father transcriptome matrix respectively; G_f, M_f, T_f are hybrid mother genome matrix, hybrid mother metabolome matrix, hybrid mother transcriptome matrix respectively.
[0060] Specifically, the model expression of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model is:
[0061] y train = 1 μ + Oa + e;
[0062] wherein, i,j = 1, 2, …, … n; y train is the hybrid phenotype observation value vector of the training set; μ is the fixed effect; O is any one of the hybrid genome matrix, the hybrid metabolome matrix and the hybrid transcriptome matrix in the training set; e is the residual error vector; h is the bandwidth parameter; p is any one of the number of SNPs, the number of metabolites and the number of transcripts; a is any one of the genome random effect, the metabolome random effect and the transcriptome random effect; K is any one of the genome Gaussian kernel, the metabolome Gaussian kernel and the transcriptome Gaussian kernel; is any one of the genome variance, the metabolome variance and the transcriptome variance; D ij is the Euclidean distance between the i th sample and the j th sample of the hybrid in the training set; n is the number of hybrids in the training set.
[0063] Further, the expression of the training prediction matrix is:
[0064]
[0065] The expression of the test prediction matrix is:
[0066]
[0067] wherein, and respectively represent the predicted values of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model on the training set; and respectively represent the mean of 10 predicted values of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model on the test set; k represents the k-fold data taken as the test data of the model when training the model by cross-validation.
[0068] Specifically, the expression of the linear regression model is:
[0069]
[0070] wherein q e {1, 2, 3}; is the ith phenotype observation of the hybrid based on the training set; b is the intercept term; w1, w2, w3 are the first regression coefficient, the second regression coefficient and the third regression coefficient, respectively; ε is the error term; I A is the indicator function; The ith element in the first column.
[0071] Specifically, q represents the value of the indicator function I A When q = 1
[0072]
[0073] Only and are the prediction values of the genome-phenotype prediction model and the metabolome-phenotype prediction model based on the training set on the phenotype of the hybrid, so it is a genome-metabolome integration. When q = 2
[0074]
[0075] Only and are the prediction values of the genome-phenotype prediction model and the transcriptome-phenotype prediction model based on the training set on the phenotype of the hybrid, so it is a genome-transcriptome integration. When q = 3
[0076]
[0077] and are the prediction values of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model based on the training set on the phenotype of the hybrid, so it is a genome-metabolome-transcriptome integration.
[0078] Preferably, the genomic-phenotype prediction model, the metabolome-phenotype prediction model, and the transcriptome-phenotype prediction model are trained based on the training set and the test set is predicted by using ten-fold cross-validation, and a training prediction matrix and a test prediction matrix are obtained, and the specific process is as follows: the training set is randomly divided into 10 parts, in each iteration, 1 part is taken as the test data, and the remaining 9 parts are taken as the training data, then the parameters estimated by using the training data are used to predict the hybrid phenotype value of the test data, and the hybrid phenotype value of the test set is predicted. After 10 iterations, 3 hybrid phenotype prediction value vectors based on the genomic-phenotype prediction model, the metabolome-phenotype prediction model, and the transcriptome-phenotype prediction model on the training set are obtained, and 3 hybrid phenotype prediction value vectors based on the genomic-phenotype prediction model, the metabolome-phenotype prediction model, and the transcriptome-phenotype prediction model on the test set are obtained, and finally the hybrid phenotype prediction value vectors based on the training set and the test set are combined to obtain the final training prediction matrix and the test prediction matrix.
[0079] Optionally, the training prediction matrix is input as a feature into a linear regression model for training, and the linear regression model is predicted based on the test prediction matrix to obtain the final hybrid phenotype prediction value, and the specific process is as follows: the training prediction matrix is taken as a new feature, and the hybrid phenotype observation value vector in the training set is taken as a label and input into the linear regression model for training, the trained linear regression model is used to predict on the test prediction matrix, and thus the final hybrid phenotype prediction value is obtained.
[0080] Specifically, the method provided in the embodiment is described in detail as follows:
[0081] 1) Data acquisition
[0082] The genomic, metabolome, and transcriptome data of the parents and the hybrid phenotype data are acquired, and the genomic, metabolome, and transcriptome of the hybrid are inferred from the genomic, metabolome, and transcriptome of the parents. In this embodiment, a set of published rice data set is used for testing, including 210 parents, 278 hybrids, 1619 bins of the parents from 270820 SNPs, 1000 metabolites, 24994 gene expressions, and four yield-related traits: single plant yield, tiller number, grain number per panicle, and 1000-grain weight.
[0083] 2) Data preprocessing
[0084] The genomic, metabolome, and transcriptome data of the parents are preprocessed. For the genomic data, the SNPs with a minimum allele frequency lower than 0.05 and a missing rate greater than 0.2 are filtered out, and for the metabolome and transcriptome data, standardization processing is performed.
[0085] 3) Speculating hybrid genome, metabolome and transcriptome
[0086] Suppose the hybrid genome, metabolome and transcriptome matrices of the paternal parent are G_m, M_m and T_m, respectively, and the hybrid genome, metabolome and transcriptome matrices of the maternal parent are G_f, M_f and T_f, respectively. Then the hybrid genome, metabolome and transcriptome matrices are:
[0087]
[0088] 4) Construction of omics-phenotype prediction model
[0089] Dataset division: The dataset is randomly divided into 10 parts using ten-fold cross-validation, and one part is taken as the test set, and the remaining 9 parts are taken as the training set. Based on the training set, the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model are constructed. The model expression is as follows:
[0090] y train = 1 μ + Oa + e
[0091] Where y train is the hybrid phenotype observation value vector based on the training set, and μ is the fixed effect.
[0092] 5) Training of omics-phenotype prediction model
[0093] The above model is trained using ten-fold cross-validation, and the test set is predicted. When the model is trained, a prediction matrix based on the training set is obtained:
[0094]
[0095] And a prediction matrix based on the test set is obtained:
[0096]
[0097] 6) Secondary training
[0098] The predicted values on the training set are input into the linear regression model for training, and the predicted values on the test set are predicted to obtain the final hybrid phenotype prediction value. The model expression is:
[0099]
[0100] The trained linear regression model is used to predict the phenotype value of the test set, and the final predicted phenotype value is:
[0101]
[0102] Where q e {1,2,3}, is the phenotypic prediction value of the s-th sample of hybrid in the test set, s = 1, 2,... m, m is the number of hybrids in the test set, is the estimated value of the intercept term of the linear regression model obtained by estimating with the training set, and are the regression coefficients of the linear regression obtained by estimating with the training set, I A is the indicator function, q = 1 represents the genome-metabolome integration, q = 2 represents the genome-transcriptome integration, q = 3 represents the genome-metabolome-transcriptome integration.
[0103] 7) Model evaluation
[0104] The coefficient of determination (R ) between the predicted phenotype values obtained by the linear regression model test and the observed phenotype values (y 2 ) of the test set is used as the evaluation index of the model. After each fold of the divided data set is trained as the test set, R 2 of 10 times of training can be obtained. The R 2 of 10 times of training is averaged to obtain the R 2 of 10-fold cross-validation. In order to ensure the randomness and reliability of the results, this process is repeated 20 times, and the average of the R 2 of 20 times is taken as the final prediction ability of the model.
[0105] Reference Figure 2, 4 traits are single plant yield, tiller number, grain number per panicle and 1000-grain weight; 9 kinds of prediction models are G, M, T, GM, GT, GMT, G+M, G+T and G+M+T, which respectively represent genome prediction, metabolome prediction, transcriptome prediction, integrated genome-metabolome prediction, integrated genome-transcriptome prediction, integrated genome-metabolome-transcriptome prediction, genome-metabolome integrated prediction, genome-transcriptome integrated prediction and genome-metabolome-transcriptome integrated prediction. In each box plot, different lower case letters at the bottom indicate significant differences between different models. The single plant yield, tiller number, grain number per panicle and 1000-grain weight of rice hybrids are predicted by using the above method; the cross-validation results show that, compared with genome prediction, MOEGS significantly improves the prediction of traits. Genome-metabolome integration improves the prediction of single plant yield, tiller number, grain number per panicle and 1000-grain weight by 48.4%, 11.8%, 17.1% and 1.5% respectively; genome-transcriptome integration improves by 33.3%, 12.1%, 16.0% and 2.5% respectively; and genome-metabolome-transcriptome integration improves by 45.8%, 10.4%, 17.8% and 2.7% respectively. In addition, the prediction of MOEGS is higher than that of the integrated multi-omics; compared with genome-metabolome integration, genome-metabolome integration improves the prediction of single plant yield, tiller number and grain number per panicle by 12.9%, 9.6% and 3.7% respectively; compared with genome-transcriptome integration, genome-transcriptome integration improves the prediction of single plant yield, tiller number, grain number per panicle and 1000-grain weight by 21.3%, 40.9%, 9.7% and 1.8% respectively; compared with genome-metabolome-transcriptome integration, genome-metabolome-transcriptome integration improves the prediction of single plant yield, tiller number, grain number per panicle and 1000-grain weight by 32.6%, 33.5%, 10.3% and 1.5% respectively.
[0106] The beneficial effects of the present application are as follows:
[0107] The present application improves the multi-omics prediction of the model and improves the prediction accuracy of the phenotype of rice hybrids by constructing genome-phenotype prediction model, metabolome-phenotype prediction model, transcriptome-phenotype prediction model and linear regression model and training the model using ten-fold cross-validation.
[0108] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.
[0109] The principles and implementation manners of the present application are described by using specific examples in the present application, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for the general technical personnel in the art, the specific implementation manners and application ranges will be changed according to the idea of the present application. In conclusion, the content of the present specification should not be understood as the limitation of the present application.
Claims
1. A multi-omics ensemble-based whole-genome prediction method, characterized in that, The method comprises the following steps: obtaining parent genome data, parent metabolome data, parent transcriptome data and hybrid phenotypic data; filtering out SNPs with minor allele frequency less than 0.05 and deletion rate greater than 0.2 in the parent genome data to obtain pre-processed genome data, and performing standardization processing on the parent metabolome data and the parent transcriptome data respectively to obtain pre-processed metabolome data and pre-processed transcriptome data; speculating hybrid genome matrix, hybrid metabolome matrix and hybrid transcriptome matrix according to the pre-processed genome data, the pre-processed metabolome data and the pre-processed transcriptome data respectively; integrating the hybrid phenotypic data, the hybrid genome matrix, the hybrid metabolome matrix and the hybrid transcriptome matrix and randomly dividing them into a training set and a test set; training pre-constructed genome-phenotype prediction model, metabolome-phenotype prediction model and transcriptome-phenotype prediction model according to ten-fold cross-validation using the training set to obtain a training prediction matrix, and predicting the test set using the trained genome-phenotype prediction model, metabolome-phenotype prediction model and transcriptome-phenotype prediction model to obtain a test prediction matrix; training a linear regression model by inputting the training prediction matrix as a feature, and predicting the linear regression model according to the test prediction matrix to obtain a final hybrid phenotypic prediction value.
2. The method of claim 1, wherein, Further comprising: calculating the determination coefficient between the final predicted phenotypic value and the observed phenotypic value of the test set to obtain an evaluation index; obtaining the mean of the evaluation index of several times of training to obtain the final prediction ability of the model.
3. The method of claim 1, wherein the method is based on multi-omics integration. The ratio of the training set to the test set is 9:
1. 4.The whole genome prediction method based on multi-omics integration of claim 1, wherein, The expression of the hybrid genome matrix is: The expression of the hybrid metabolome matrix is: The expression of the hybrid transcriptome matrix is: Wherein, G_m, M_m, T_m are the hybrid paternal genome matrix, the hybrid paternal metabolome matrix and the hybrid paternal transcriptome matrix respectively; G_f, M_f, T_f are the hybrid maternal genome matrix, the hybrid maternal metabolome matrix and the hybrid maternal transcriptome matrix respectively.
5. The method of claim 4, wherein the method is a multi-omics integrated whole genome prediction method. The model expression of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model is: y train = 1 μ + Oa + ε; wherein, y train is the hybrid phenotype observation vector of the training set; μ is the fixed effect; O is any one of the hybrid genomic matrix, the hybrid metabolomic matrix, and the hybrid transcriptomic matrix in the training set; ε is the residual vector; h is the bandwidth parameter; p is any one of the number of SNPs, the number of metabolites, and the number of transcripts; α is any one of the genomic random effect, the metabolomic random effect, and the transcriptomic random effect; K is any one of the genomic Gaussian kernel, the metabolomic Gaussian kernel, and the transcriptomic Gaussian kernel; is any one of the genomic variance, the metabolomic variance, and the transcriptomic variance; D ij is the Euclidean distance between the i-th sample and the j-th sample of the hybrid in the training set; n is the number of hybrids in the training set.
6. The method of claim 5, wherein the method is a multi-omics integrated whole genome prediction method. The expression of the training prediction matrix is: The expression of the test prediction matrix is: wherein, and respectively represent the predicted values of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model on the training set; and respectively represent the mean of 10 predicted values of the genome-phenotype prediction model, the metabolome-phenotype prediction model and the transcriptome-phenotype prediction model on the test set; k represents the k-fold data used as the test data of the model when training the model with cross-validation.
7. The method of claim 6, wherein the method is a multi-omics integrated whole genome prediction method. The expression of the linear regression model is: where q e {1,2,3}; ytrain i is the ith phenotypic observation of the hybrid based on the training set; b is an intercept term; w1, w2, w3 are the first, second, and third regression coefficients, respectively; and ε is an error term; I A is an indicator function; is the ith element in the first column.
Citation Information
Patent Citations
Hybrid seed prediction method based on Bayesian model integrating parent phenotypes
CN113053459A
Multi-view GBLUP method for integrating multiple types of data for phenotype prediction
CN117153247A