Soybean oil and protein content prediction and SNP (Single Nucleotide Polymorphism) analysis method and system

By combining LD analysis, GWAS, SVR model and SHAP interpretability analysis, the problems of high-dimensional soybean oil and protein content, overfitting, insufficient interpretability and data redundancy in small sample data processing were solved, and accurate prediction and explanatory improvement were achieved.

CN120048354APending Publication Date: 2025-05-27SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY +2
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411878411.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, when processing high-dimensional and small sample data of soybean oil and protein content, overfitting problems are prone to occur, model interpretation is insufficient, it is difficult to effectively capture the nonlinear relationships in the data, and there are problems with high false positive rates and data redundancy.

Method used

Feature selection and dimension screening were performed by combining linkage imbalance (LD) analysis, genome-wide association analysis (GWAS), support vector regression (SVR) models and SHAP interpretive analysis, and the encoded data introduced weights based on base dimensions, and predictive and explanatory analysis were performed.

Benefits of technology

Accurate prediction of soybean oil and protein content in high-dimensional small sample data is achieved, which improves the interpretability and operability of the model, reduces data redundancy, and identifies SNP sites significantly related to the target phenotype, which improves the ability to understand the relationship between genotype and phenotype.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048354A_ABST
    Figure CN120048354A_ABST
Patent Text Reader

Abstract

The invention discloses a soybean oil and protein content prediction and SNP (Single Nucleotide Polymorphism) analysis method and system. The method comprises the following steps: acquiring a soybean sample data set; performing feature selection and dimension screening on the preprocessed soybean sample data set to obtain a screened first data set; encoding the first data set, and introducing a weight based on the base dimension to obtain encoded data; the coded data are input into a prediction model, and the oil content and the protein content are predicted; and performing explanatory analysis on the prediction model to obtain influence data of the SNP sites on phenotypes. According to the method, the LD analysis, the GWAS analysis, the SVR model and the SHAP explanatory analysis are combined, so that accurate prediction of the soybean oil content and the protein content in high-dimensional small sample data is realized, and meanwhile, the explanatory property and operability of the model are improved. The data redundancy can be effectively reduced, the SNP sites significantly related to the target phenotype can be identified, and the ability of understanding the relationship between the genotype and the phenotype is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image automated processing, and particularly to a method and system for predicting soybean oil and protein contents and their SNP analysis. Background Art

[0002] Currently, the prediction of soybean oil and protein contents and the analysis related to single nucleotide polymorphisms (SNPs) are important research directions in the field of modern agricultural biotechnology. These studies aim to improve crop quality, increase yields, and better understand how genetic variation affects important agronomic traits of soybeans through genomics means.

[0003] However, the genomic data of soybeans usually has a high dimension (millions of SNP loci), but the available sample size is limited, usually only a few hundred. Existing machine learning models, such as deep learning models, are prone to overfitting problems when facing this kind of data, making it difficult for the models to effectively capture the non-linear relationships in the data. And traditional machine learning models, especially deep learning models, are often regarded as "black box" models, making it difficult to explain which features (such as specific SNP loci) have a greater impact on the prediction results. This limits the application of the models in agricultural genomics, especially in scenarios where the relationship between genes and phenotypes needs to be understood; for capturing the complex non-linear relationships between genes and phenotypes, some correlation measurement methods, such as the Pearson correlation coefficient (PCC), mainly reflect linear relationships and cannot effectively reveal the complex non-linear associations between genes and phenotypes; for the problems of high false positive rates and data redundancy, in genome-wide association studies (GWAS), due to the sample population structure and complex genetic background, the false positive results are relatively high, and there is a large amount of redundant information in the SNP data; at the same time, the redundant information in large-scale SNP data and the model complexity lead to low computational efficiency. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method and system for predicting soybean oil and protein contents and their SNP analysis to solve the problems existing in the prior art, such as the challenges in dealing with high-dimensional and small-sample data, insufficient model interpretability, and many deficiencies in the analysis ability of complex quantitative traits, which bring limitations to the prediction and analysis of genomic data.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In the first aspect, the present invention provides a method for predicting soybean oil and protein contents and their SNP analysis, including:

[0008] Obtain a soybean sample data set;

[0009] Perform feature selection and dimensionality screening on the preprocessed soybean sample dataset to obtain the first screened dataset;

[0010] Encode the first dataset, introducing weights based on the base dimension to obtain encoded data;

[0011] Input the encoded data into a prediction model to predict the oil content and protein content;

[0012] Perform interpretive analysis on the prediction model to obtain data on the impact of SNP sites on phenotypes.

[0013] As a preferred embodiment of the method for predicting soybean oil and protein content and its SNP analysis according to the present invention, wherein: the feature selection includes:

[0014] Perform LD analysis on the preprocessed soybean sample dataset to reduce data redundancy and obtain preliminarily screened data.

[0015] As a preferred embodiment of the method for predicting soybean oil and protein content and its SNP analysis according to the present invention, wherein: the dimensionality screening includes:

[0016] Perform GWAS analysis on the preliminarily screened data using a mixed linear model, and obtain the first dataset through P-value filtering.

[0017] As a preferred embodiment of the method for predicting soybean oil and protein content and its SNP analysis according to the present invention, wherein encoding the first dataset and introducing weights based on the base dimension includes:

[0018] Encode the four bases of the SNP site as a four-dimensional vector;

[0019] Based on one-hot encoding, introduce a counting weight in each dimension to reflect the occurrence frequency of a specific base;

[0020] If there is a missing or heterologous situation, it is encoded as zero.

[0021] As a preferred embodiment of the method for predicting soybean oil and protein content and its SNP analysis according to the present invention, wherein inputting the encoded data into a prediction model to predict the oil content and protein content includes:

[0022] Input the encoded data into an SVR model;

[0023] Optimize the model parameters through grid search;

[0024] Use the oil content and protein content as output labels to obtain estimated values of the oil content and protein content.

[0025] As a preferred embodiment of the soybean oil, protein content prediction, and its SNP analysis method of the present invention, wherein: performing an interpretive analysis on the prediction model, including:

[0026] Performing a global interpretive analysis on the prediction model using the SHAP method;

[0027] Calculating the marginal contribution of each SNP locus to the prediction result to obtain the average contribution of each SNP locus to the prediction result.

[0028] As a preferred embodiment of the soybean oil, protein content prediction, and its SNP analysis method of the present invention, wherein: performing an interpretive analysis on the prediction model, including:

[0029] Performing a local interpretive analysis on the prediction model using the SHAP method;

[0030] For a single sample, identifying the positive and negative impacts of specific SNP loci on the prediction result of the sample.

[0031] In a second aspect, the present invention provides a soybean oil, protein content prediction, and its SNP analysis system, including:

[0032] An acquisition module for acquiring a soybean sample data set;

[0033] A screening module for performing feature selection and dimensionality screening on the preprocessed soybean sample data set to obtain a first screened data set;

[0034] An encoding module for encoding the first data set, the encoding introducing weights based on the base dimension to obtain encoded data;

[0035] A prediction module for inputting the encoded data into a prediction model to predict the oil content and protein content;

[0036] An analysis module for performing an interpretive analysis on the prediction model to obtain data on the impact of SNP loci on the phenotype.

[0037] In a third aspect, the present invention provides an electronic device, including:

[0038] A memory and a processor;

[0039] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the soybean oil, protein content prediction, and its SNP analysis method are implemented.

[0040] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions, which when executed by a processor, implement the steps of the method for predicting soybean oil and protein content and its SNP analysis.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows: By combining LD analysis, GWAS analysis, SVR model and SHAP interpretability analysis, the present invention realizes accurate prediction of soybean oil and protein content in high-dimensional small-sample data, while improving the interpretability and operability of the model. It can not only effectively reduce data redundancy, but also identify SNP loci significantly related to the target phenotype, improve the understanding ability of the relationship between genotype and phenotype, and provide an important reference for soybean breeding. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 It is a schematic diagram of the overall process of the method for predicting soybean oil and protein content and its SNP analysis according to an embodiment of the present invention;

[0044] Figure 2 It is a Q-Q plot, Manhattan circular plot, and bar chart of the GWAS results before and after LD screening of oil content in the method for predicting soybean oil and protein content and its SNP analysis according to an embodiment of the present invention;

[0045] Figure 3 It is a Q-Q plot, Manhattan circular plot, and bar chart of the GWAS results before and after LD screening of protein content in the method for predicting soybean oil and protein content and its SNP analysis according to an embodiment of the present invention;

[0046] Figure 4 It is a Q-Q plot, Manhattan circular plot, and bar chart of the GWAS results before and after LD screening of water-soluble protein content in the method for predicting soybean oil and protein content and its SNP analysis according to an embodiment of the present invention;

[0047] Figure 5 It is a summary chart of the SHAP analysis of the SVR model for oil, protein, and water-soluble protein content in the method for predicting soybean oil and protein content and its SNP analysis according to an embodiment of the present invention;

[0048] Figure 6Scatter plot of SHAP values of SNP sites in the prediction of soybean oil, protein content, and its SNP analysis method for oil, protein, and water-soluble protein content in an embodiment of the present invention. Detailed implementation manners

[0049] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0050] Embodiment 1

[0051] Refer to Figure 1 , which is an embodiment of the present invention, and provides a method for predicting soybean oil and protein content and its SNP analysis, including:

[0052] S100: Obtain a soybean sample dataset;

[0053] S200: Perform feature selection and dimensionality screening on the preprocessed soybean sample dataset to obtain a first screened dataset;

[0054] S300: Encode the first dataset, introducing weights based on the base dimension to obtain encoded data;

[0055] S400: Input the encoded data into a prediction model to predict the oil content and protein content;

[0056] S500: Perform interpretability analysis on the prediction model to obtain data on the impact of SNP sites on phenotypes.

[0057] It should be noted that existing machine learning models (such as deep learning) when dealing with small-sample, high-dimensional data such as soybean oil, since the data dimension far exceeds the number of samples, the model is prone to capturing noise in the data rather than real signals, resulting in insufficient generalization performance and a high risk of overfitting; traditional methods are difficult to effectively extract non-linear feature relationships in high-dimensional sparse data; in the feature engineering of genomic data, SNPs usually use one-hot encoding as input features, but there are certain limitations in the encoding process, including: existing encoding methods usually only consider genotypes (such as homozygotes, heterozygotes), and do not fully consider specific base types (such as A, T, C, G), only encoding homozygotes, or only encoding homozygotes of 4 bases, without fully considering the diversity of heterozygotes and their potential impact on genotype-phenotype. This encoding method may lead to information loss, affecting the accurate capture of the genotype-phenotype relationship by the model, thereby reducing the accuracy and reliability of the prediction.

[0058] Therefore, through steps S100 - S500, by optimizing the screening steps and encoding while considering base types, homozygous and heterozygous information simultaneously, the prediction performance of the model is improved, and support is provided for subsequent model interpretability analysis; select appropriate machine learning models and global and local analysis methods to identify SNP sites that contribute significantly to the phenotype prediction results, enhance the interpretability of the model, thus solving the "black box" problem of the model in existing methods and providing more transparent analysis results.

[0059] Example 2

[0060] Refer to Figure 1 , which is an embodiment of the present invention. Based on the above embodiment, a method for predicting soybean oil and protein content and its SNP analysis is provided.

[0061] In the embodiment of the present application, in step S100, the soybean sample dataset is obtained from the Soybean Improvement Center of scientific research schools, including soybean samples from China, the United States, Japan, and Brazil. Three phenotypic traits are measured for each sample: oil content, protein content, and water - soluble protein content.

[0062] In an alternative embodiment, the soybean sample dataset in step S100 can also be downloaded from a public database, such as GenBank, etc.;

[0063] In another alternative embodiment, the soybean sample dataset in step S100 can also be directly collected through experiments to construct a custom dataset.

[0064] Furthermore, after obtaining the dataset, the quality of the dataset is improved through pre - processing.

[0065] Specifically, pre - processing can be achieved by removing samples and sites with high missing rates, filtering out samples and SNP sites with a missing rate exceeding a certain proportion (such as 20%), to ensure the integrity of the remaining data. Secondly, filter low - allele - frequency sites: remove SNP sites with an allele frequency less than a set threshold (such as 0.05) to reduce insignificant sites.

[0066] In the embodiment of the present application, in step S200, feature selection includes:

[0067] Perform LD analysis on the pre - processed soybean sample dataset to reduce data redundancy and obtain preliminarily screened data.

[0068] Specifically, use linkage disequilibrium (LD) analysis, set a correlation threshold such as: R 2 = 0.2, which can retain relatively independent SNP sites.

[0069] In the embodiment of the present application, the dimension screening in step S200 includes:

[0070] Perform GWAS analysis on the preliminarily screened data using a mixed linear model, and obtain the first data set through P-value filtering.

[0071] Specifically, using GWAS (Genome-Wide Association Study), SNP loci that are significantly associated with the target phenotype in different environments can be screened out; through P-value filtering, for example, using P < 0.01 as the threshold, irrelevant SNP loci can be excluded.

[0072] In an alternative embodiment, the feature selection or dimension screening in step S200 can only select GWAS analysis;

[0073] In another alternative embodiment, the feature selection or dimension screening in step S200 can also be achieved by means such as principal component analysis.

[0074] It should be noted that in the present application, it is preferred to first perform LD analysis to obtain the preliminarily screened data, which can reduce the SNP sites to tens of thousands; then perform GWAS analysis, and through P-value filtering, the first data set can reduce the SNP sites to hundreds, so as to improve the data quality and significantly enhance the feature quality of the data and the prediction ability of the subsequent model.

[0075] In the embodiment of the present application, in step S300, the first data set is encoded, and weights are introduced based on the base dimension, including the following steps A1 - A3:

[0076] A1: Encode the four bases of the SNP locus into a four-dimensional vector;

[0077] A2: Based on one-hot encoding, introduce a count weight in each dimension to reflect the occurrence frequency of a specific base;

[0078] A3: If there is a missing or heterologous situation, it is encoded as zero.

[0079] Exemplarily, the above scheme can be expressed as: the encoding of homozygous A is (2, 0, 0, 0), the encoding of heterozygous AT is (1, 1, 0, 0), and the encoding of homozygous T is (0, 2, 0, 0).

[0080] In an alternative embodiment, the encoding of the first data set in step S300 can also be achieved through homologous encoding, that is, only homologous sites are encoded, and the site information significantly associated with the phenotype is retained;

[0081] In another alternative embodiment, encoding the first data set in step S300 can also be achieved through genotype encoding, that is: encoding the homologous, heterologous, and missing situations of SNP sites as 0, 1, and 2 respectively.

[0082] It should be noted that in this application, it is preferred to use improved one-hot encoding, that is: encoding by introducing weights based on the base dimension can enhance the sensitivity of the model to base features while retaining base type information, improve prediction performance, and lay a foundation for model interpretability analysis.

[0083] In the embodiment of this application, in step S400, inputting the encoded data into the prediction model to predict the oil content and protein content includes the following steps B1 - B3:

[0084] B1: Input the encoded data into the SVR model;

[0085] B2: Optimize the model parameters through grid search;

[0086] B3: Use the oil content and protein content as output labels to obtain estimated values of the oil content and protein content.

[0087] In an alternative embodiment, inputting the encoded data into the prediction model in step S400 can also be performed by inputting it into a random forest model (RF) or extreme gradient boosting (XGBoost) for prediction.

[0088] In another alternative embodiment, inputting the encoded data into the prediction model in step S400 can also be performed by inputting it into a deep learning model (DualCNN, DNNGP, etc.) for prediction.

[0089] It should be noted that in this application, among the above-mentioned achievable models, the SVR model is preferred. Its (RBF kernel) has the best prediction performance and can show good prediction effects and generalization capabilities in small-sample high-dimensional data. Compared with other models, its PCC (Pearson correlation coefficient) exceeds 0.6, and the R 2 value is close to 0.4, showing good performance.

[0090] In the embodiment of this application, in step S500, performing interpretability analysis on the prediction model includes:

[0091] Using the SHAP method to perform global interpretability analysis on the prediction model;

[0092] Calculating the marginal contribution of each SNP site to the prediction result to obtain the average contribution of each SNP site to the prediction result.

[0093] In the embodiment of this application, in step S500, performing interpretability analysis on the prediction model includes:

[0094] The SHAP method is used to perform local interpretability analysis on the prediction model;

[0095] For a single sample, identify the positive and negative impacts of specific SNP loci on the prediction result of the sample.

[0096] Exemplarily, for the model predicting soybean oil content, through SHAP analysis, SNP loci located on specific chromosomes (such as 8 and 20) can be identified as being significantly correlated with oil content. These loci remain after LD screening and GWAS analysis and show strong positive and negative impacts in the SHAP interpretation.

[0097] In an alternative embodiment, the interpretability analysis of the prediction model in step S500 can also be locally interpreted through LIME (Local Interpretable Model-agnostic Explanations). For each sample, LIME will create a local linear model or other easily understandable model and calculate the features that contribute the most to a specific prediction result.

[0098] In another alternative embodiment, the interpretability analysis of the prediction model in step S500 can also be globally interpreted through Partial Dependence Plots (PDPs), which can reflect the average relationship between features and response variables across the entire dataset.

[0099] It should be noted that in this application, the SHAP method is preferred. It calculates the contribution of each SNP locus to the prediction result based on the Shapley value in game theory, and can explain the output of the model from both global and local perspectives, revealing which SNP loci have a significant impact on the phenotype. Compared with the other analysis methods, SHAP not only supports local interpretation, that is, a detailed explanation of how a single sample affects the prediction result; it also supports global interpretation by summarizing the local interpretations of all samples to understand the distribution of feature importance across the entire dataset; at the same time, the SHAP values are additive, meaning that the sum of the SHAP values of each feature is equal to the difference between the model prediction value and the baseline value, which helps to intuitively understand how SNP loci work together to form the final prediction.

[0100] Example 3

[0101] Refer to Figures 2 - 6 and Table 1 - Table 5. Based on the previous embodiment, this embodiment provides a comparative effect case of a method for predicting soybean oil and protein content and its SNP analysis to illustrate the feasibility and beneficial effects of our solution.

[0102] Obtain a soybean sample dataset; a soybean SNP dataset collected from the Soybean Improvement Center of Nanjing Agricultural University, including 219 soybean samples from 2015 to 2017 from China, the United States, Japan, and Brazil.

[0103] Perform preprocessing to filter out samples and SNP loci with a missing rate exceeding 20%; remove SNP loci with an allele frequency less than 0.05, and finally screen out valid SNP loci, as shown in Table 1:

[0104] Table 1: Number of SNP loci with p < 0.01 for each fold of different phenotypes for three consecutive years before and after LD analysis and filtering with different p-values

[0105]

[0106] Perform LD analysis on the valid SNP loci for each year, and set the correlation threshold to R 2 = 0.2, and retain relatively independent SNP loci;

[0107] Perform GWAS analysis on the continuous three-year data of soybean oil content, protein content, and water-soluble protein content respectively using a mixed linear model. Select the significance threshold p-value = 0.01 to further screen the SNP loci, and screen out SNP loci significantly related to the target phenotype. Through P-value filtering, with P < 0.01 as the threshold, exclude irrelevant SNP loci. It can be found that through LD+GWAS analysis, the number of SNP loci can be further reduced from tens of thousands to hundreds, as Figures 2 - 4 shown. As a comparison, all SNPs without LD filtering after quality control are screened with p-value = 0.001. After this screening method, thousands of SNPs are obtained, which is an order of magnitude more than the SNPs obtained by LD+GWAS.

[0108] Among them, Figure 2Quantile - quantile plots and Manhattan plots for the GWAS analysis of oil content from 2015 to 2017, comparison before and after LD filtering. Comparison of QQ plots and Manhattan plots of GWAS analysis of oil content from 2015 to 2017 before and after LD screening. In the figure: (A) QQ plot of GWAS analysis for all SNP loci after quality control: This panel shows the comparison between the observed values of oil content and the negative logarithm of the expected p - values over three years, with different years represented by curves of different colors. The red dashed line represents the expected line where the observed and expected values coincide under the null hypothesis. (B) Circular Manhattan plot for all SNP loci after quality control: The outermost circle shows the SNP density on the chromosome, followed by three concentric circles representing the data for 2015, 2016, and 2017 respectively. Blue and red represent the significance thresholds of -Log(p) = 4 and 6 respectively. Yellow markers indicate SNP loci exceeding the blue threshold, and red highlights indicate SNP loci exceeding the red threshold. (C) Multiple Manhattan bar plot for all SNP loci after quality control: The GWAS results for different years are shown in different colors. (D) QQ plot of GWAS analysis for SNP loci after LD screening: These charts show the change in data distribution after LD screening, indicating that the observed values are closer to the expected values, suggesting that the data has been processed more precisely. (E) Circular Manhattan plot for SNP loci after LD screening: After LD screening, these charts reveal a unique distribution of significant peaks, especially on chromosomes 8 and 20, indicating that important genetic markers related to oil content are retained after screening. (F) Multiple Manhattan bar plot for SNP loci after LD screening: The GWAS results for different years are shown in different colors.

[0109] As Figure 3 Shown are the QQ plots and Manhattan plots of the GWAS analysis of the protein content trait before and after LD screening from 2015 to 2017; in the figure: (A) QQ plot of GWAS analysis for all SNP loci after quality control. (B) Circular Manhattan plot for all SNP loci after quality control. (C) Multiple Manhattan bar plot for all SNP loci after quality control. (D) QQ plot of GWAS analysis for SNP loci after LD screening. (E) Circular Manhattan plot for SNP loci after LD screening. (F) Multiple Manhattan bar plot for SNP loci after LD screening.

[0110] As Figure 4QQ plots and Manhattan plots of GWAS analysis of the water-soluble protein content trait before and after LD screening from 2015 to 2017 are shown. In the figure, (A) QQ plot of GWAS analysis after quality control of all SNP sites. (B) Circular Manhattan plot after quality control of all SNP sites. (C) Multiple Manhattan bar plot after quality control of all SNP sites. (D) QQ plot of GWAS analysis of SNP sites after LD screening. (E) Circular Manhattan plot of SNP sites after LD screening. (F) Multiple Manhattan bar plot of SNP sites after LD screening.

[0111] Subsequently, one-hot encoding was performed on the four bases (A, T, C, G), and counting weights were introduced in each dimension to reflect the occurrence frequency of specific bases.

[0112] To verify the effect of the improved one-hot encoding, genotype encoding and homologous encoding were simultaneously performed, and three types of encoded data were obtained respectively.

[0113] The encoded data was input into the prediction model to predict the oil content and protein content;

[0114] The five-fold cross-validation method was used to evaluate the model performance, and multiple machine learning models were tried, including: Support Vector Regression (SVR), Random Forest (RF), Extreme Gradient Boosting (XGBoost), deep learning models (DualCNN, DNNGP). In the training of the deep learning model, to address the overfitting problem, an early stopping strategy was introduced, and the training was terminated when the performance of the test set no longer improved.

[0115] Use PCC (Pearson correlation coefficient) and R 2 (Coefficient of determination) as the model performance evaluation index. The model parameters were optimized through grid search, and the results are shown in Table 2 - Table 4. The experimental results show that the SVR model performs best in small sample high-dimensional data, and its PCC and R 2 are both better than other models.

[0116] During the model training process, the five-fold cross-validation method was adopted to ensure the broad representativeness and reliability of the evaluation results. When conducting model interpretability analysis, the one-fold model with the best performance on the test set was selected, and the results are shown in Table 4, so as to ensure that the explanations given by SHAP have high reference value and practicality.

[0117] Table 2: PCC of phenotype prediction under different models and encoding methods

[0118]

[0119] Table 3: R of phenotype prediction under different models and encoding methods 2

[0120]

[0121] Table 4: Prediction accuracy of the best model for SVR training on the test set and the entire dataset

[0122]

[0123] As Figure 5 shown, the impacts of the top 40 important features are summarized, and the SHAP value impacts and their distributions of these features in all samples are presented. Due to the base-type encoding, each locus is extended from one-dimensional encoding to four-dimensional. Two of the dimensions correspond to the two possible bases at this locus, and the other two dimensions are constantly 0, which has no practical significance. In addition, among the two meaningful dimensions, when it is 1 or 2, it indicates that the base represented by this dimension appears at this locus. When the base represented by this dimension does not appear, it is also filled with 0, and the 0 value at this position also has no practical significance. When calculating the SHAP value, the 0 value will be wrongly regarded as a feature state, so these meaningless SHAP values are filtered. The three figures are the summary figures of oil, protein, and water-soluble protein contents in sequence.

[0124] As Figure 6 shown, the SHAP value results are visualized by scatter plots in the way of using a type Manhattan plot. The SHAP values of the two bases at each locus are added as the importance of this locus, and then the scatter plot visualization of the importance of all loci is carried out. The three figures are the scatter plots of SHAP value SNP loci for oil, protein, and water-soluble protein contents in sequence.

[0125] Table 5 shows the SNP loci with significant SHAP values during the oil content analysis. The oil content of soybean samples in 2015 was predicted by the SVR model. After model training, SHAP value analysis was used. The results showed that the significant loci on chromosome 8 and chromosome 20 during GWAS analysis were also retained, and other loci were also mined. For example, there are significant loci on chromosome 13 and chromosome 18. The gene Glyma.13G241700 on chromosome 13 and the gene Glyma.18G167300 on chromosome 18 have been proven in the literature to affect the oil content of soybeans. For the two SNP loci S13_30222011 and S18_3596879 analyzed by SHAP, these two genes can be found through LD block analysis, etc. However, these two genes cannot be found through GWAS analysis. Therefore, the effectiveness of the SHAP analysis method is proved. The other significant SNP loci discovered by SHAP further narrow the scope for further mining genes related to oil content and provide potential research loci.

[0126] Table 5: Significant loci screened by SHAP value analysis of oil content

[0127] rs chr ps p_Oil p_Protein p_WSPC shap_0 shap_1 shap 35 S03_5815646 3 5815646 0.006682 0.003152 0.004819 0.004894 0.003323 0.008217 45 S03_34163406 3 34163406 0.006735 0.009562 0.003117 0.003403 0.003877 0.007280 73 S05_17593383 5 17593383 0.000027 0.004030 0.005009 0.005029 0.003157 0.008186 97 S06_37014425 6 37014425 0.000132 0.000744 0.000772 0.002261 0.005406 0.007667 111 S07_1718405 7 1718405 0.002630 0.008981 0.005970 0.008016 0.009165 0.017181 132 S08_7712671 8 7712671 0.000022 0.000251 0.000005 0.006492 0.007363 0.013854 187 S08_17263065 8 17263065 0.002286 0.006969 0.002525 0.004611 0.004144 0.008755 188 S08_17389293 8 17389293 0.002557 0.005803 0.002501 0.005235 0.004813 0.010048 230 S11_18222993 11 18222993 0.009433 0.002321 0.007936 0.005545 0.004487 0.010032 286 S13_30222011 13 30222011 0.003825 0.003443 0.003542 0.003306 0.004700 0.008006 316 S16_6031823 16 6031823 0.001277 0.003924 0.000452 0.004499 0.006104 0.010603 328 S17_18319099 17 18319099 0.000440 0.004250 0.007896 0.006074 0.007605 0.013679 331 S18_3500701 18 3500701 0.003999 0.001048 0.003548 0.003433 0.004139 0.007572 332 S18_3583959 18 3583959 0.008237 0.000203 0.007666 0.003867 0.004185 0.008052 333 S18_3596879 18 3596879 0.001559 0.002936 0.009682 0.006168 0.004772 0.010941 337 S18_16425324 18 16425324 0.000713 0.006449 0.007436 0.004420 0.004625 0.009045 338 S18_16463046 18 16463046 0.000779 0.008293 0.008769 0.004737 0.005079 0.009816 379 S20_664552 20 664552 0.000083 0.001942 0.007686 0.008339 0.007253 0.015592 387 S20_15927298 20 15927298 0.000338 0.004325 0.003731 0.005098 0.002000 0.007097 459 S20_45600576 20 45600576 0.002201 0.009443 0.002207 0.004373 0.004191 0.008565 460 S20_45631477 20 45631477 0.001948 0.003036 0.002052 0.004391 0.004721 0.009112 461 S20_45685177 20 45685177 0.004002 0.009728 0.005329 0.004737 0.005116 0.009853

[0128] From the above comparison scenarios, it can be seen that the present invention improves the prediction accuracy. By using the Support Vector Regression (SVR) model, the prediction accuracy of soybean oil and protein content is significantly improved when dealing with small-sample high-dimensional data. Compared with traditional machine learning models, SVR has stronger generalization ability on small-sample data sets, and its Pearson correlation coefficient (PCC) can reach up to 0.674 at most, and the R 2 is all above 0.34. This high-precision prediction can more accurately assist researchers in phenotypic analysis and breeding selection; the present invention enhances the model interpretability: by introducing the SHAP (Shapley value) interpretability analysis method, the "black box" problem of traditional machine learning models is solved. SHAP analysis can reveal the contribution of each SNP locus to the prediction result, enabling researchers to understand the association between genes and phenotypes from both global and local perspectives. This feature makes this method not only limited to prediction but also able to provide valuable gene information for breeding; at the same time, the present invention effectively reduces data redundancy: the present invention uses LD (linkage disequilibrium) analysis and GWAS (genome-wide association study) to perform dimensionality screening on SNP data, thereby reducing the influence of redundant data. Through these screening steps, the key SNP loci related to soybean oil and protein content are finally retained, improving the efficiency and accuracy of the model and reducing the false positive rate; the present invention is also applicable to complex traits controlled by multiple genes. The oil and protein content of soybeans are complex quantitative traits controlled by multiple genes, and it combines the advantages of machine learning models and GWAS, can better handle such complex traits controlled by multiple genes, provides a more comprehensive means of genomic data analysis, and overcomes the limitations of traditional GWAS in dealing with complex traits; the present invention improves the data processing efficiency: through LD analysis and GWAS screening, the dimensionality of the input data is reduced, improving the computational efficiency of machine learning models. When facing high-dimensional data with millions of SNP loci, the solution proposed by the present invention can significantly shorten the computational time without reducing the prediction performance, making it more suitable for application in actual production environments. This method is not only applicable to soybeans but can also be extended to genomic data analysis of other crops and animals, especially having broad application prospects in the scenario of small-sample high-dimensional data. It provides an efficient and interpretable technical solution for phenotypic prediction and breeding improvement in modern agricultural genomics.

[0129] Example 4

[0130] The above is a schematic solution of a method for predicting soybean oil and protein content and its SNP analysis in this embodiment. It should be noted that the technical solution of the system for predicting soybean oil and protein content and its SNP analysis belongs to the same concept as the above-mentioned method for predicting soybean oil and protein content and its SNP analysis. For the details not described in detail in the technical solution of the system for predicting soybean oil and protein content and its SNP analysis in this embodiment, reference can be made to the description of the technical solution of the above-mentioned method for predicting soybean oil and protein content and its SNP analysis.

[0131] This embodiment also provides a system for a method for predicting soybean oil and protein content and its SNP analysis, including:

[0132] An acquisition module, configured to acquire a soybean sample data set;

[0133] A screening module, configured to perform feature selection and dimensionality screening on the preprocessed soybean sample data set to obtain a first screened data set;

[0134] An encoding module, configured to encode the first data set, introducing weights based on the base dimension to obtain encoded data;

[0135] A prediction module, configured to input the encoded data into a prediction model to predict the oil content and protein content;

[0136] An analysis module, configured to perform interpretive analysis on the prediction model to obtain data on the influence of SNP sites on phenotypes.

[0137] This embodiment also provides an electronic device applicable to the situation of predicting soybean oil and protein content and its SNP analysis, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for predicting soybean oil and protein content and its SNP analysis proposed in the above embodiment.

[0138] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for predicting soybean oil and protein content and its SNP analysis proposed in the above embodiment.

[0139] The storage medium proposed in this embodiment and the method for predicting soybean oil and protein content and its SNP analysis proposed in the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment can be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0140] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and the necessary general-purpose hardware, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0141] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A soybean oil, protein content prediction and SNP analysis method thereof, characterized in that, include: Get the soybean sample dataset; Performing feature selection and dimension screening on the preprocessed soybean sample data set to obtain a screened first data set; Encoding the first data set, wherein the encoding introduces weights based on a base dimension to obtain encoded data; The encoded data are input into the prediction model to predict the oil content and protein content; The prediction model is subjected to an explanatory analysis to obtain data on the impact of the SNP site on the phenotype.

2. The soybean oil and protein content prediction and SNP analysis method according to claim 1, characterized in that: The feature selection comprises: LD analysis was performed on the preprocessed soybean sample data set to reduce data redundancy and obtain preliminary screening data.

3. The soybean oil and protein content prediction and SNP analysis method according to claim 2, characterized in that: The dimension screening includes: The preliminary screening data were subjected to GWAS analysis using a mixed linear model and filtered by P value to obtain a first data set.

4. The soybean oil and protein content prediction and SNP analysis method according to claim 3, characterized in that: The first data set is encoded, and the encoding introduces weights based on a base dimension, including: Encode the four bases of the SNP site into a four-dimensional vector; Based on one-hot encoding, count weights are introduced in each dimension to reflect the frequency of occurrence of specific bases; If missing or heterozygous, it is coded as zero.

5. The soybean oil and protein content prediction and SNP analysis method according to claim 4, characterized in that: The encoded data is input into the prediction model to predict the oil content and protein content, including: Input the encoded data into the SVR model; Optimize model parameters through grid search; The oil content and protein content are taken as output labels to obtain the estimated values ​​of oil content and protein content.

6. The soybean oil and protein content prediction and SNP analysis method according to claim 5, characterized in that: The prediction model is subjected to an explanatory analysis, including: The SHAP method was used to conduct a global interpretative analysis of the prediction model; The marginal contribution of each SNP site to the prediction result was calculated to obtain the average contribution of each SNP site to the prediction result.

7. The soybean oil and protein content prediction and SNP analysis method according to claim 6, characterized in that: The prediction model is subjected to an explanatory analysis, including: The SHAP method was used to conduct local interpretability analysis of the prediction model; For a single sample, identify the positive and negative impact of a specific SNP site on the prediction result of the sample.

8. A system for predicting the soybean oil and protein content and the SNP analysis method thereof according to any one of claims 1 to 7, characterized in that: include: An acquisition module is used to obtain a soybean sample data set; A screening module, used for performing feature selection and dimension screening on the preprocessed soybean sample data set to obtain a screened first data set; An encoding module, used for encoding the first data set, wherein the encoding introduces weights based on a base dimension to obtain encoded data; A prediction module, used for inputting the encoded data into a prediction model to predict the oil content and protein content; The analysis module is used to perform an explanatory analysis on the prediction model to obtain the data on the impact of the SNP site on the phenotype.

9. An electronic device, comprising: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the soybean oil, protein content prediction and SNP analysis method thereof according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the soybean oil and protein content prediction and SNP analysis method thereof as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Corn phenotype prediction method and system

    CN113593635A

  • Gene-to-phenotype prediction method and system based on combination of XGBoost feature selection and deep learning

    CN116597894A

  • 4mC locus recognition algorithm based on fusion of pruning pre-training model and artificial feature coding

    CN117216656A

  • Biological phenotype prediction and gene locus screening method based on deep learning

    CN119007807A

  • Optimal Soybean Loci

    US20150128308A1