Plant genome prediction method, product and equipment based on multi-dimensional data selection index

By using a plant genome prediction method based on multidimensional data selection indices, and integrating multidimensional data through inner and outer layer cross-validation and selection index strategies, the problem of poor accuracy in multidimensional data prediction in existing technologies is solved, and a more efficient and accurate breeding process is achieved.

CN121768481APending Publication Date: 2026-03-31YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing genomic selection models suffer from poor prediction accuracy, high computational burden, and high risk of overfitting when integrating multidimensional data, especially in predicting low heritability phenotypes.

Method used

A plant genome prediction method based on multidimensional data selection index is adopted. The model is constructed through inner and outer layer cross-validation, the weight vectors of different omics are calculated, and the multidimensional data is integrated by using the selection index strategy to reduce dimensionality explosion and computational burden, thereby improving prediction accuracy.

Benefits of technology

It significantly improves the prediction accuracy of multidimensional data, reduces computational complexity, reduces the risk of overfitting, and enhances the accuracy and efficiency of the breeding process, especially improving the prediction accuracy by more than 10% in the prediction of traits with low heritability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768481A_ABST
    Figure CN121768481A_ABST
Patent Text Reader

Abstract

The invention discloses a plant genome prediction method, product and equipment based on a multi-dimensional data selection index, and belongs to the technical field of plant breeding. The method comprises the following steps: S1, data preparation: acquiring species data of a target species, and ensuring that multiple omics data of each individual are in one-to-one correspondence with phenotypes; s2, model construction: carrying out model training by adopting five-fold cross validation comprising two cycles of an inner layer and an outer layer so as to solve the problems of overfitting and generalization ability evaluation; s3, weight estimation: obtaining weight vectors of different omics based on a selection index strategy; s4, prediction calculation: using training set data to obtain phenotype prediction values of the prediction set by each group, assigning corresponding weights to obtain selection indexes combined with all group data, and repeating for a plurality of times to cover the whole data set; and S5, accuracy evaluation: calculating a decision coefficient of a selection index and a true value, and verifying a model effect by adopting a plurality of methods. According to the invention, the prediction precision of multiple omics can be obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of plant breeding technology, specifically relating to a plant genome prediction method, product, and equipment based on multidimensional data selection index. Background Technology

[0002] Genomic selection (GS) plays a crucial role in plant breeding. GS constructs statistical models using molecular marker data from the entire genome and phenotypic data from training populations to estimate the effect of each detected marker. It then calculates the phenotypic value or estimated genomic breeding value for each individual in the test population, and selects individuals based on these values. However, genomic prediction has limitations in integrating complex epistatic interactions and downstream regulation, and often performs poorly in predicting low-heritability phenotypes. In recent years, the rapid advancements in technologies such as next-generation sequencing and metabolomics platforms have generated massive amounts of omics data from crops. These data have been used to dissect the genetic basis of traits with unprecedented precision and provide relevant information for decision-making to achieve breeding goals. Therefore, the efficient integration and utilization of multidimensional data, including genomics, transcriptomics and metabolomics, phenotyping and environmental data, can lead to higher prediction accuracy for complex agronomic traits, significantly accelerating crop improvement breeding processes and playing a vital role in promoting precision breeding and agricultural development.

[0003] With continuous improvement and optimization of genomics (GS) statistical models, their stability and variety have increased. However, existing technologies still have some problems and shortcomings. Currently, most GS models focus on prediction and selection for single omics, neglecting the development of multidimensional data prediction models. Even if some models can be used for multidimensional data prediction, the ways different methods integrate multidimensional data vary greatly. This difference not only causes difficulties for breeders' analysis but also significantly affects prediction accuracy. For example, directly merging genomic, transcriptomic, and metabolomic data column-wise may lead to dimensionality explosion, greatly increasing the computational burden. Therefore, there is an urgent need to develop a multidimensional data genomic prediction model with higher prediction accuracy, lower computational requirements, more stable performance, and a unified multidimensional data integration method. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the aforementioned related technologies.

[0005] Therefore, the purpose of this invention is to provide a plant genome prediction method, product, and device based on a multidimensional data selection index, which can significantly improve the prediction accuracy of multidimensional data.

[0006] To solve the above-mentioned technical problems, the present invention is implemented as follows: This invention provides a plant genome prediction method based on a multidimensional data selection index, the method comprising: S1. Data preparation: Obtain species data for the target species and ensure that the multi-omics data of each individual corresponds one-to-one with the phenotype; S2. Model Building: Five-fold cross-validation with two loops (inner and outer layers) is used for model training to address overfitting and generalization ability assessment issues. S3. Weight estimation: Obtain the weight vectors of different omics based on the index selection strategy; S4. Prediction calculation: Using the training set data, obtain the phenotypic prediction values ​​of each omics group on the prediction set, assign corresponding weights, and obtain the selection index that combines all omics data. Repeat this several times to cover the entire dataset. S5. Accuracy assessment: Calculate the coefficient of determination between the selection index and the true value, and use several methods to verify the model's effectiveness.

[0007] In addition, the plant genome prediction method based on multidimensional data selection index according to the present invention may also have the following additional technical features: In some implementations, the species data in step S1 includes genomic, transcriptomic, metabolomic, and phenotypic data.

[0008] In some implementations, step S3 includes: constructing a multidimensional data model of quantitative traits based on a selection index strategy, and solving for the weight vectors of different omics by combining observable parameters, including phenotypic variance, through the variance-covariance matrix of multi-omics predicted values, the covariance matrix of trait breeding values ​​and predicted values.

[0009] In some implementations, the weight vector is obtained as follows: , Where P is the variance-covariance matrix of the predicted values ​​of m omics data; G is the covariance matrix of the predicted values ​​of m omics data.

[0010] In some of these implementations, , , in: Let y be the predicted value of the m-th omics data, where y is the phenotypic value of the quantitative trait; A Let be the covariance matrix between the trait breeding value and the m predicted values.

[0011] A Let be the covariance matrix between the trait breeding value and the m predicted values.

[0012] In some implementations, the selection index that combines all omics data is obtained by assigning corresponding weights in step S4 by linearly weighting the multi-omics prediction values ​​using a weight vector b.

[0013] In some of these implementations, the linear weighting method is as follows: , in: Let y be the predicted value of the k-th omics data, where y is the phenotypic value of the quantitative trait; b k represents the weight of the k-th omics data.

[0014] In some implementations, step S5 uses eight methods—GBLUP, BayesB, SVM, RR, PLS, RKHS, XGBoost, and LightGBM—to evaluate the accuracy of genome prediction models based on selection indices from quantitative trait multidimensional data.

[0015] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the plant genome prediction method based on a multidimensional data selection index as described in any of the preceding embodiments.

[0016] This invention also provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the plant genome prediction method based on multidimensional data selection index as described in any of the preceding embodiments.

[0017] Compared with the prior art, the present invention has at least the following beneficial effects: In this embodiment of the invention, the plant genome prediction method based on multidimensional data selection index can improve the accuracy of multidimensional data joint prediction; by calculating the weights of different omics, the prediction accuracy in the breeding process is improved, making genome selection more precise and efficient. In this embodiment of the invention, the plant genome prediction method based on multidimensional data selection index is provided. By calculating the weight of each omics through the selection index model, the multidimensional data information is accurately allocated, avoiding the dimensionality explosion and computational burden caused by directly merging data, and unifying the multidimensional data integration method. In this embodiment of the invention, the plant genome prediction method based on multidimensional data selection index provides hyperparameter optimization and generalization ability assessment in the inner and outer loops, respectively, which effectively reduces the risk of overfitting and improves the reliability of model evaluation. In this embodiment of the invention, the plant genome prediction method based on multidimensional data selection index eliminates the genetic dependence during the derivation process and can solve the covariance matrix related parameters using only directly observable data such as phenotypic variance, thereby reducing computational complexity. In this embodiment of the invention, the plant genome prediction method based on multidimensional data selection index provides an average improvement of more than 10% in prediction accuracy compared to conventional multidimensional data models for complex agronomic traits such as low heritability, using various algorithms such as PLS, SVM, and XGBoost, thereby accelerating the crop breeding process.

[0018] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of a genome prediction process based on selection index of multidimensional quantitative trait data, as disclosed in an embodiment of the present invention. Figure 2 The predictive power of 210 rice inbred lines with 4 traits under different omics under an embodiment of the present invention; Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific examples and application scenarios.

[0022] Please see Figure 1 As shown, in some embodiments of the present invention, a plant genome prediction method based on a multidimensional data selection index is provided. The data required by this method includes genomic data, transcriptomic data, metabolomic data, and phenotypic data of the target species.

[0023] Cross-validation is used to randomly divide the training and prediction sets. To effectively reduce the risk of overfitting and provide more reliable model evaluation results, this scheme uses 5-fold nested cross-validation to build the model, including two cross-validation loops: an inner loop and an outer loop. The specific process is as follows: Outer loop: The entire dataset is divided into 5 folds. In each fold, a subset is selected as the test set, and the remaining 4 folds are used as the training set. The purpose of the outer loop is to evaluate the model's generalization ability and ensure that the test set is not used in training to avoid information leakage. Inner loop: On the outer training set, another 5-fold cross-validation is performed to find the optimal combination of hyperparameters. The inner loop divides the outer training set into 5 subsets to calculate the predicted values ​​for the entire training set.

[0024] An exponential strategy is chosen to estimate the weights of different omics on the inner dataset. The construction of the exponential model for quantitative trait multidimensional data is as follows: Let y be the phenotypic value of a quantitative trait. The predicted value of y from the k-th omics data is denoted as... Where k = 1, 2, ..., m, m is the total number of omics data, and here m = 3. There are genomic data (k = 1), transcriptomic data (k = 2), and metabolomic data (k = 3).

[0025] The selection index for all m omics datasets is: , P is the variance-covariance matrix of the predicted values ​​of m omics data.

[0026] , A is the covariance matrix between the trait breeding values ​​and the m omics prediction values. The index weights are obtained based on the following formula: , Matrix P can be obtained from the predicted values ​​of multiple omics datasets. The covariance matrix G cannot be directly observed and requires additional derivation.

[0027] To derive the k-th element of matrix G The prediction accuracy of the k-th omics data needs to be defined as: , in, It is the square root of the heritability of the trait under study. Predictive accuracy can be expressed as: , in, It is the phenotypic variance of the trait. It is the k-th diagonal element of matrix P.

[0028] and Both parameters can be obtained from the data. Prediction accuracy enables this invention to solve for... The calculation formula is as follows: , because , and Both can be estimated from the data, and this covariance can be easily obtained. In the derivation process, heritability has been eliminated, therefore the model only contains phenotypic variance. )kick in.

[0029] This method utilizes 5-fold nested cross-validation. In each inner loop, 5-fold cross-validation is used to estimate the phenotypic predictions for the entire training set using data from each omics. Then, the weights of each omics are calculated based on the selection index method. Simultaneously, the phenotypic predictions for the entire prediction set are estimated using data from each omics in the training set. Different weights are assigned to the predictions of different omics, resulting in the selection index of the prediction set combining all omics data. This process is repeated 5 times to obtain the selection index of the entire dataset combining all omics data. Finally, the coefficient of determination of the true values ​​and the selection index is calculated as the model's prediction accuracy.

[0030] Finally, eight methods—GBLUP, BayesB, SVM, RR, PLS, RKHS, XGBoost, and LightGBM—were used to evaluate the predictive accuracy of the selection index model based on quantitative trait multidimensional data analysis.

[0031] Example 1:

[0032] This embodiment provides a genome prediction method based on a selection index of quantitative trait multidimensional data to improve the prediction accuracy of genome prediction models. The method includes the following steps: Step 101: Provide genomic, transcriptomic, metabolomic, and phenotypic data. For each individual, the multi-omics data must correspond one-to-one with the phenotype. Transcriptomic and metabolomic data need to be standardized in advance.

[0033] Step 102: A 5-fold nested cross-validation was used to build the model, consisting of two cross-validation loops: an inner loop and an outer loop. The specific process is as follows: Outer loop: The entire dataset is divided into 5 folds. In each fold, a subset is selected as the test set, and the remaining 4 folds are used as the training set. The purpose of the outer loop is to evaluate the model's generalization ability and ensure that the test set is not used in the training process to avoid information leakage. Inner loop: On the outer training set, another 5-fold cross-validation is performed to find the optimal combination of hyperparameters. The inner loop divides the outer training set into 5 subsets to calculate the predictions for the entire training set.

[0034] Step 103: Let y be the phenotypic value of the quantitative trait. The predicted value of y from the k-th omics data is denoted as... Where k = 1, 2, ..., m, m is the total number of omics data, and here m = 3. We have genomic data (k = 1), transcriptomic data (k = 2), and metabolomic data (k = 3).

[0035] The selection index for all m omics datasets is: , in, It is a vector of predicted values ​​for traits from m omics datasets. It is an exponential weight vector derived from m omics datasets.

[0036] , P is the variance-covariance matrix of the predicted values ​​of m omics data.

[0037] A is the covariance matrix between the trait breeding value and the m predicted values. The index weights are obtained based on the following formula: , Matrix P can be obtained from the predicted values ​​of multiple omics data.

[0038] Step 104: In each inner loop, use 5x cross-validation to estimate the phenotypic prediction of the entire training set using each omics data. Then, calculate the weight of each omics based on the selection index method. At the same time, use the training set data to estimate the phenotypic prediction of the entire prediction set using each omics data. Then, assign different weights to the prediction values ​​of different omics to obtain the selection index that combines all omics data.

[0039] Step 105: This process is repeated 5 times to obtain the selection index for the entire dataset. Then, the determination coefficients of the true values ​​and selection indices are calculated to evaluate the model's predictive accuracy. Finally, eight methods—GBLUP, BayesB, SVM, RR, PLS, RKHS, XGBoost, and LightGBM—are used to evaluate the accuracy of the genome prediction model based on the selection index of quantitative trait multidimensional data.

[0040] Example 2: This embodiment provides a further detailed description of the preceding embodiments.

[0041] Regarding data acquisition: This embodiment utilizes a publicly available rice dataset, which includes 210 rice intercalation molecules (RILs) generated from hybridization of two rice varieties (Zhenshan 97 and Minghui 63). Genomic data includes 1619 bins identified from 270,820 SNPs generated after sequencing the 210 RILs. Metabolomics data includes 1000 metabolites, of which 317 are from germinating seeds and the remaining 683 are from flag leaves. Four agronomic traits were examined: yield per plant (YD), tiller number per plant (TP), grain number per panicle (GN), and 1000 grain weight (KGW).

[0042] 1. Data preprocessing: The transcriptome and metabolome data need to be standardized in advance.

[0043] 2. Nested Cross-Validation Model Construction: A 5-fold nested cross-validation model was used, consisting of two cross-validation loops: an inner loop and an outer loop. The specific process is as follows: Outer Loop: The entire dataset is divided into 5 folds. Within each fold, a subset is selected as the test set, and the remaining 4 folds are used as the training set. The purpose of the outer loop is to evaluate the model's generalization ability and ensure that the test set is not used in the training process to avoid information leakage. Inner Loop: On the outer training set, another 5-fold cross-validation is performed to find the optimal hyperparameter combination. The inner loop divides the outer training set into 5 subsets to calculate the prediction values ​​for the entire training set.

[0044] 3. Integration of Selection Index Model: Let y be the phenotypic value of a quantitative trait, and let the predicted value of y from the k-th omics data be denoted as... Where k = 1, 2, ..., m, and m is the total number of omics data, here m = 3. There are genomic data (k = 1), transcriptomic data (k = 2), and metabolomic data (k = 3). The selection index for all m omics datasets is: , in, It is a vector of predicted values ​​for traits from m omics datasets. It is an exponential weight vector derived from m omics datasets.

[0045] , P is the variance-covariance matrix of the predicted values ​​of m omics data.

[0046] , A is the covariance matrix between the trait breeding value and the m predicted values. The index weights are obtained based on the following formula: , Matrix P can be obtained from the predicted values ​​of multiple omics data.

[0047] 4. Prediction using the multidimensional data selection index model: In each inner loop, 5x cross-validation is used to estimate the phenotypic prediction of the entire training set using data from each omics. Then, the weight of each omics is calculated based on the selection index method. Simultaneously, the phenotypic prediction of the entire prediction set is estimated using the training set data from each omics. Different weights are then assigned to the prediction values ​​of different omics to obtain the final phenotypic of the prediction set. This process is repeated 5 times to obtain the phenotypic prediction of the entire dataset. Finally, the coefficient of determination between the true value and the predicted value is calculated as the prediction accuracy of the model.

[0048] 5. Comparison of different algorithms: Constructing multidimensional prediction models based on selection indices using eight methods. Genomic prediction models based on selection indices from quantitative trait multidimensional data were constructed using eight methods: GBLUP, BayesB, SVM, RR, PLS, RKHS, XGBoost, and LightGBM.

[0049] Analysis based on the results of 210 rice inbred lines is as follows: Figure 2 The results show that, compared with genomics, conventional multidimensional data modeling and selection index-based multidimensional mathematical models significantly improve the predictive power of traits. The prediction accuracy of the multidimensional data selection index-based strategy is significantly better than that of the conventional multidimensional data strategy in PLS, SVM, RR, XGBoost, and LightGBM models. However, in GBLUP, BayesB, and RKHS models, the multidimensional data selection index-based model does not show a significant difference compared to the conventional multidimensional data model. Overall, however, for the traits YD, TP, GN, and KGW, the prediction accuracy of the multidimensional data selection index model is improved by an average of 11.13%, 20.59%, 10.26%, and 20.63% compared to the conventional multidimensional data prediction model, respectively.

[0050] This invention develops a multidimensional data prediction framework based on a selection index. The key feature of this framework lies in its innovative quantitative trait multidimensional data selection index strategy. It calculates the weights of different omics based on training set data, achieving a more accurate allocation of information from different omics. It also overcomes the limitations of prediction frameworks that rely solely on genomic data, demonstrating a significant advantage in improving prediction accuracy. The multidimensional data selection index model of this invention can be used not only for genomic, transcriptomic, and metabolomic data, but also for high-throughput phenotypic and environmental data, exhibiting broad applicability.

[0051] All parts of this invention not described in detail herein can be referred to in the prior art or are known to those skilled in the art. This embodiment does not limit these aspects and will not describe them in detail here.

[0052] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A method for predicting plant genomes based on multi-dimensional data selection indices, characterized by, The method comprises: S1. Data preparation: obtaining species data of target species, ensuring that the multi-omics data of each individual corresponds to the phenotype one-to-one; S2. Model construction: using 5-fold cross-validation containing two cycles of inner and outer layers for model training to solve overfitting and generalization ability evaluation problems; S3. Weight estimation: obtaining the weight vector of different omics based on the selection index strategy; S4. Prediction calculation: obtaining the phenotype prediction value of each omics on the prediction set using the training set data, and obtaining the selection index combined with all omics data after assigning the corresponding weight, and repeating several times to cover the entire data set; S5. Accuracy evaluation: calculate the determination coefficient of the selection index and the true value, and verify the model effect by using several methods.

2. The plant genome prediction method based on multi-dimensional data selection index according to claim 1, characterized in that, The species data in step S1 includes genome, transcriptome, metabolome and phenotype data.

3. The plant genome prediction method based on multi-dimensional data selection index according to claim 1, characterized in that, The content of step S3 includes: constructing a quantitative trait multi-dimensional data selection index model based on the selection index strategy, obtaining the weight vector of different omics through the variance-covariance matrix of multi-omics prediction values, the covariance matrix of trait breeding value and prediction value, and combining the observable parameters including phenotype variance.

4. The plant genome prediction method based on multi-dimensional data selection index according to claim 3, characterized in that, The index weight vector is obtained in the following manner: , Wherein, P is the variance-covariance matrix of m omics data prediction values; G is the covariance matrix of m omics data prediction values.

5. The plant genome prediction method based on multi-dimensional data selection index according to claim 4, characterized in that, , , wherein: is the predicted value for the mth omics data pair y, y being the phenotypic value of a quantitative trait; A is the covariance matrix between the trait breeding values and the m predicted values.

6. The plant genome prediction method based on multi-dimensional data selection index according to claim 4, characterized in that, The way to obtain the selection index combined with all omics data after assigning the corresponding weight in step S4 is: linearly weighting the multi-omics prediction values through the index weight vector b.

7. The plant genome prediction method based on multi-dimensional data selection index according to claim 6, characterized in that, The linear weighting method is: , wherein: is the predicted value for the kth omics data pair y, y is the phenotypic value of a quantitative trait; b k is the weight for the kth omics data.

8. The plant genome prediction method based on multi-dimensional data selection index according to claim 1, characterized in that, In step S5, GBLUP, BayesB, SVM, RR, PLS, RKHS, XGBoost and LightGBM are used to evaluate the accuracy of the genome prediction model based on quantitative trait multi-dimensional data selection index.

9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the plant genome prediction method based on multi-dimensional data selection index in any one of claims 1-8.

10. A computer device comprising: The memory, the processor and the computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to realize the steps of the plant genome prediction method based on multi-dimensional data selection index in any one of claims 1-8.