Plant phenotype multi-precision information joint prediction method
Through a multi-precision information joint prediction method, combined with SNP, gene expression and phenotype data, the model is constructed using the multi-omics prediction module, and the interpretability analysis is performed using the SHAP algorithm, which solves the shortcomings of plant phenotype prediction methods in the existing technology in high-dimensional data processing and multi-omics data integration, and achieves high accuracy and interpretability plant phenotype prediction.
Patent Information
- Application Number
- CN202510506112.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing plant phenotype prediction methods process high-dimensional, sparse and noise-rich biological data, it is difficult to achieve high prediction accuracy, and lack the ability to effectively integrate multi-omics and cross-platform data, resulting in insufficient interpretability of the model.
A joint prediction method of plant phenotype multi-precision information is adopted to realize joint prediction of multiomic information through steps such as data acquisition and preprocessing, model construction and training, and interpretability analysis. Specific steps include SNP data preprocessing, gene expression data processing, phenotype data processing, model construction using genome prediction module, transcriptome prediction module and comprehensive prediction module, and interpretability analysis using SHAP algorithm.
It improves the prediction accuracy in high-dimensional data and complex interaction scenarios, realizes effective integration of multi-omic data, provides reliable biological explanations, and improves the technical support for breeding and the sustainable development capabilities of agriculture.
Smart Images

Figure CN120048358A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of plant phenotype prediction, and particularly to a method for jointly predicting multi-precision information of plant phenotypes. Background Art
[0002] In recent years, with the rapid development of machine learning algorithms, a new generation of non-linear models has gradually been introduced into the field of plant trait prediction. Among them, support vector machines capture non-linear relationships with kernel functions, and ensemble methods such as random forests and gradient boosting trees improve generalization performance by constructing multiple decision trees. Deep learning models are favored for their powerful feature extraction and representation capabilities.
[0003] However, current plant phenotype prediction still faces multiple challenges: traditional statistical models are interpretable, but the linear assumption is difficult to insight into complex molecular mechanisms and environmental interactions; machine learning methods can handle non-linear relationships, but in biological data scenarios with high-dimensional, sparse, and noisy data, they have high requirements for data scale and are prone to overfitting; although deep learning performs outstandingly in feature extraction, it often requires huge computing resources, and the interpretability of the model is insufficient. The "black box" property makes it difficult to mine and verify biological significance; at the same time, most existing prediction methods lack the ability to effectively integrate multi-omics and cross-platform data, and it is difficult to fully utilize the opportunities brought by the explosion of modern biological data.
[0004] Therefore, it is of great significance to develop a new generation of plant phenotype prediction methods. This method needs to have the following characteristics: it can achieve higher prediction accuracy when dealing with high-dimensional data and complex interactions, show the potential to integrate multi-source information in the joint analysis of multi-omics data, and at the same time provide reliable biological explanations. These innovations will provide stronger technical support for modern breeding and also make important contributions to global food security and sustainable agricultural development. Summary of the Invention
[0005] Therefore, it is of great significance to develop a new generation of plant phenotype prediction methods. This method needs to have the following characteristics: it can achieve higher prediction accuracy when dealing with high-dimensional data and complex interactions, show the potential to integrate multi-source information in the joint analysis of multi-omics data, and at the same time provide reliable biological explanations. These innovations will provide stronger technical support for modern breeding and also make important contributions to global food security and sustainable agricultural development.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a method for jointly predicting multi-precision information of plant phenotypes, including the following steps: S1. Data acquisition and preprocessing, and the preprocessing steps include SNP data preprocessing, gene expression data processing, and phenotype data processing; S2. Model construction and training, realizing the joint prediction of multi-omics information through the genomic prediction module, transcriptomic prediction module, and comprehensive prediction module; S3. Interpretability analysis, using the SHAP algorithm to quantify the contribution of each SNP and gene expression feature to the model prediction.
[0007] Preferably, the specific operation steps of the SNP data preprocessing include filtering the VCF file according to a preset threshold, using a sliding window with a preset window size and step length to perform dimensionality reduction on SNPs; finally obtaining a ten-thousand-dimensional feature vector; the specific operation steps of the gene expression data processing include converting the original transcriptomic data into the RPKM form to correct the differences in sequencing depth between different samples; the specific operation steps of the phenotypic data processing include identifying and processing outliers through the box plot method, and performing standardization processing to convert the phenotypic data into a standard normal distribution with a mean of 0 and a standard deviation of 1.
[0008] Preferably, in S2, the genomic prediction module is used to predict the crop phenotype through the additive coding form of gene single nucleotide polymorphism (SNP).
[0009] Preferably, in S2, the transcriptomic prediction module is used to predict the crop phenotype through reads per kilobase per million reads (RPKM).
[0010] Preferably, in S2, the comprehensive prediction module in S2 is used to perform phenotypic prediction by integrating the SNP and RPKM results.
[0011] Preferably, in S3, the SHAP algorithm is used to quantify the contribution of each SNP and gene expression feature to the model prediction, perform a genome-wide association analysis on the phenotype predicted by the model and the actually measured phenotype, compare the GWAS results of the mixed phenotype and the true phenotype, and identify the prediction-specific loci.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Systematization and completeness of data processing 1) Multi-omics preprocessing process: Establish a standardized quality control and dimensionality reduction strategy for SNP, gene expression, and phenotypic data to ensure the accuracy of the input data from the source.
[0013] 2) Strict quality control: By setting key parameters (MAF, missing rate, sequencing depth, etc.) and sliding window dimensionality reduction, accurately eliminate noise and redundancy.
[0014] 3) Multi-source data integration: Unified standardization processing enables subsequent models to fully utilize genomic and transcriptomic information, providing high-quality features for phenotypic prediction.
[0015] 2. Efficiency and stability of model construction 1) Combination of Convolutional Neural Network (CNN) and Transformer: CNN is good at mining local sequence features, while Transformer is responsible for capturing long-range dependencies and overall structures, enhancing the modeling ability for high-dimensional and complex data.
[0016] 2) Hyperparameter optimization: Achieve fine-tuning of the model through techniques such as parameter search, and obtain the best balance between generalization performance and running efficiency.
[0017] 3) Flexibility and robustness: Compared with traditional linear or general machine learning methods, deep learning models perform better in complex gene interaction scenarios and maintain stable prediction performance.
[0018] 3. Interpretability of prediction results 1) Quantification of feature contribution: Use SHAP (SHapley Additive exPlanations) to evaluate the importance of SNP and gene expression features, and intuitively present the key driving factors of the model.
[0019] 2) Functional annotation: Use tools such as SNPEFF to interpret the biological meaning of significant SNP sites and quickly locate gene regions closely related to phenotypes.
[0020] 4. Multilevel biological insights Integrated analysis of Genome-Wide Association Studies (GWAS): Synchronously incorporate the model prediction values and real phenotypes into GWAS to discover more potential functional sites.
[0021] 5. Practicality and scalability 1) Wide range of applicable scenarios: This method can be applied to different species and various plant traits, and can be extended to other biological fields.
[0022] 2) Standardized process: Data processing, model training, and result interpretation are universal and easy to be quickly applied in actual breeding or research.
[0023] Modular framework: New omics data and analysis functions can be flexibly added to meet the future needs of large-scale multi-omics research. Description of the drawings
[0024] Figure 1 is the flow chart of a method for jointly predicting multi-precision information of plant phenotypes in the present invention.
[0025] Figure 2 is the training schematic diagram of three prediction modules in the present invention.
[0026] Figure 3It is the prediction accuracy graph of the SNP feature model of the present invention.
[0027] Figure 4 It is the prediction accuracy graph of the transcriptome feature model of the present invention Figure 5 It is the prediction accuracy graph of the feature combination model of the present invention.
[0028] Figure 6 It is the importance ranking graph after calculating the SHAP value for the SNP features of the plant height trait based on the genomic prediction module.
[0029] Figure 7 It is the importance ranking graph after calculating the SHAP value for the transcriptome features of the plant height trait based on the transcriptome prediction module.
[0030] Figure 8 It is the GWAS analysis result graph of the true plant height phenotype.
[0031] Figure 9 It is the GWAS analysis result graph combining the predicted phenotype and the true phenotype. Detailed implementation manners
[0032] Please refer to Figure 1 As shown, the present invention relates to a method for jointly predicting multi-precision information of plant phenotypes. In this implementation, 529 cultivated rice germplasm materials with both phenotype and genomic sequencing data are selected, and 287 of these materials also have transcriptome sequencing data. Additionally, 4195 rice materials with genomic sequencing data are collected for phenotype prediction and GWAS analysis. Specifically, it includes the following steps: S1. Data acquisition and preprocessing. The preprocessing steps include SNP data preprocessing, gene expression data processing, and phenotype data processing; SNP data preprocessing Use the GATK 4.6.0.0 software to perform strict quality control and filtering on the original genomic sequencing data. The filtering parameters include: 1. QD (Quality by Depth) < 2.0 2. FS (Fisher Strand) > 60.0 3. MQ (Mapping Quality) < 40.0 4. MQRankSum < -12.5 5. ReadPosRankSum < -8.0 After filtering, use PLINK v2.00a6LM to convert the obtained SNP data into an additive coding form of 0, 1, 2 for subsequent modeling analysis.
[0033] Transcriptome data preprocessing: Standardize the original transcriptome data into the form of RPKM (Reads Per Kilobase per Million mapped reads). The RPKM standardization process comprehensively considers gene length, sequencing depth, and comparability among different samples to ensure that the expression values accurately reflect the differences in genes among different materials.
[0034] S2. Model construction and training, achieving the joint prediction of multi-omics information through the genomic prediction module, transcriptome prediction module, and comprehensive prediction module; As Figure 2 shown, in this embodiment, three prediction modules are developed, namely the genomic prediction module, the transcriptome prediction module, and the comprehensive prediction module. All modules use ten-fold cross-validation (10-fold CV) to evaluate the prediction performance and generalization ability, and the early stopping method is used to avoid overfitting.
[0035] Genomic prediction module: 1) Extract the SNP additive coding matrix from 529 materials as input features.
[0036] 2) Use the sliding window method to screen 10,000 SNP sites from the features.
[0037] 3) Construct the genomic prediction module and perform hyperparameter optimization.
[0038] 4) Use the ten-fold cross-validation method to evaluate the model performance.
[0039] 5) Record the model prediction accuracy.
[0040] Transcriptome prediction module: 1) Matrixize the RPKM expression data of 287 materials to form gene expression feature inputs.
[0041] 2) Use the Sklearn KBest method to screen the top 10,000 important features from the transcriptome features.
[0042] 3) Construct the transcriptome prediction module and perform hyperparameter optimization.
[0043] 4) Use the ten-fold cross-validation method to evaluate the model performance.
[0044] 5) Record the model prediction accuracy.
[0045] Comprehensive prediction module: 1) For 287 materials with both SNP and transcriptome data, splice the SNP and transcriptome features to form a joint feature set.
[0046] 2) Use the SklearnKBest method to screen the top ten thousand important features from the combined features.
[0047] 3) Build a comprehensive prediction module and perform hyperparameter optimization.
[0048] 4) Adopt the ten-fold cross-validation method to evaluate the model performance.
[0049] 5) Record the prediction accuracy of the combined feature model.
[0050] S3. Interpretability analysis: Use the SHAP algorithm to quantify the contribution of each SNP and gene expression feature to the model prediction; 1) SHAP calculation: In the genome and the comprehensive prediction model, use SHAP to evaluate the contribution of each SNP locus.
[0051] 2) Sorting and annotation: Sort the SNPs in descending order according to the SHAP scores, and use SNPEFF to perform functional annotation on the genomic vcf file for subsequent identification of genes or regions significantly related to the phenotype.
[0052] 3) Information integration: Collect the annotation results through a custom script, integrate the functional information of multiple significant SNPs, and provide data support for in-depth interpretation.
[0053] Analyze the importance of the transcriptome: 1) In the transcriptome or the comprehensive prediction model, quantify the relative importance of gene expression features. In the present invention, the relative importance is calculated by SHAP calculation.
[0054] 2) Sort the expression features according to the SHAP values, focus on screening the key genes with high rankings, and lay a foundation for subsequent functional research or exploration of biological mechanisms.
[0055] GWAS analysis and locus discovery After the SNP model is trained as described above, apply it to 4195 rice materials containing only SNP data to predict their phenotypes.
[0056] Perform a genome-wide association study (GWAS) on the predicted phenotypes and the known phenotype data together, and compare with the known phenotype data GWAS.
[0057] The prediction accuracy of the SNP feature model is as Figure 3As shown in the figure, in the SNP features, the prediction accuracy of the TransformerGP model is significantly higher than that of other models, with an average value reaching 57.77%, which is higher than that of the models GradientBoosting (12.27%), RandomForest (8.85%), LightGBM (4.86%), XGBoost (7.22%), KNeighborsRegressor (9.49%), Lasso (11.02%), ElasticNet (11.27%), LinearRegression (13.31%), Ridge (13.26%), RRBLUP (5.36%), DeepGS (12.21%), and DNNGP (11.58%) respectively. Moreover, the prediction accuracies for Heading_date and Plant_height reach 71.9% and 80.5% respectively.
[0058] The prediction accuracy of the transcriptome feature model is as Figure 4 shown. In the transcriptome features, the prediction accuracy of the TransformerGP model is significantly higher than that of other models, with an average value reaching 57.97%, which is higher than that of the models GradientBoosting (15.53%), RandomForest (7.11%), LightGBM (4.26%), XGBoost (6.54%), KNeighborsRegressor (24.59%), Lasso (19.92%), ElasticNet (20.41%), LinearRegression (16.62%), Ridge (16.62%), DeepGS (22.28%), and DNNGP (10.65%) respectively. Moreover, the prediction accuracies for Heading_date and Plant_height reach 79.1% and 80.1% respectively.
[0059] The prediction accuracy of the feature combination model is as Figure 5 shown. In the present invention, the prediction accuracies of the genomic prediction module, the transcriptome prediction module, and the comprehensive prediction module are compared, and it is found that the multi-omics prediction accuracy is improved compared with that of the single-omics.
[0060] As Figure 6 shown, the importance ranking results after calculating the SHAP values for the SNP features of the plant height trait based on the genomic prediction module are presented. This result intuitively reflects the importance degree of each SNP feature's contribution to the yield trait, providing an important clue for analyzing the genetic basis of the yield trait.
[0061] As Figure 7As shown, it presents the importance ranking results after calculating the SHAP values for the transcriptome features of plant height traits based on the transcriptome prediction module. This result clearly reveals the relative contributions of each transcriptome feature to the yield traits, providing an important basis for deeply exploring the molecular mechanism of yield traits.
[0062] The following table shows the annotation information of the important SNPs of plant height traits screened by the genomic prediction module, where the annotation results are generated based on the SNPEFF tool. SNPEFF annotation provides detailed functional information for these SNPs, including gene positions, types of effects (such as synonymous mutations, missense mutations or splice site variations), and possible functional impact levels. This information provides important support for deeply analyzing the biological functions of key SNPs and their mechanisms of action in yield traits.
[0063] Table 1 Annotation information table of important SNPs of plant height traits screened by the genomic prediction module
[0064] As Figure 8 and Figure 9 shown, Figure 8 it presents the GWAS analysis results of the true plant height phenotype, Figure 9 while this is the GWAS analysis result after combining the predicted phenotype with the true phenotype. By adopting the TransformerGP algorithm, it can effectively increase the sample size of GWAS, thereby significantly improving the statistical power of the analysis and discovering more significantly associated genetic loci.
[0065] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary engineering and technical personnel in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for joint prediction of plant phenotype multi-precision information, characterized in that: The following steps are involved: S1. Data acquisition and preprocessing. The preprocessing steps include SNP data preprocessing, gene expression data processing and phenotype data processing; S2, model construction and training, to achieve joint prediction of multi-omics information through genome prediction module, transcriptome prediction module and comprehensive prediction module; S3. Interpretability analysis, using the SHAP algorithm to quantify the contribution of each SNP and gene expression feature to model prediction.
2. A method for joint prediction of plant phenotype multi-precision information according to claim 1, characterized in that: The specific operation steps of the SNP data preprocessing include filtering the VCF file according to a preset threshold, using a sliding window with a preset window size and step length to reduce the dimension of the SNP; finally obtaining a 10,000-dimensional feature vector; the specific operation steps of the gene expression data processing include converting the original transcriptome data into RPKM format to correct the differences in sequencing depth between different samples; the specific operation steps of the phenotypic data processing include identifying and processing outliers through the box plot method, and using standardization to convert the phenotypic data into a standard normal distribution with a mean of 0 and a standard deviation of 1.
3. A method for joint prediction of plant phenotype multi-precision information according to claim 1, characterized in that: In S2, the genome prediction module is used to predict crop phenotypes through the additive coding form of gene single nucleotide polymorphisms (SNPs).
4. A method for joint prediction of plant phenotype multi-precision information according to claim 1, characterized in that: In S2, the transcriptome prediction module is used to predict crop phenotypes using reads per kilobase per million (RPKM).
5. A method for joint prediction of plant phenotype multi-precision information according to claim 1, characterized in that: In S2, the comprehensive prediction module is used to perform phenotype prediction by integrating SNP and RPKM results.
6. A method for joint prediction of plant phenotype multi-precision information according to claim 1, characterized in that: In S3, the SHAP algorithm is used to quantify the contribution of each SNP and gene expression feature to the model prediction, and the phenotype predicted by the model is subjected to genome-wide association analysis with the actual measured phenotype, and the GWAS results of the mixed phenotype and the true phenotype are compared to identify prediction-specific sites.
Citation Information
Patent Citations
Plant phenotype multi-omics prediction method and system
CN118983005A
Wheat abiotic stress character SNP prediction method based on automatic CNN model
CN119446290A
Cited By
Medicinal plant phenotype prediction method based on gene action mode
CN121122390A