Machine learning-based method, system and application for salt-tolerant soybean genome-wide selection
By employing genome-wide selection technology based on machine learning, and utilizing a convolutional neural network model with CA attention mechanism and multiple residual modules, the problem of insufficient screening accuracy for salt tolerance traits in soybeans was solved, achieving efficient salt-tolerant soybean breeding and significantly improving breeding efficiency and prediction accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for soybean salt-alkali tolerance breeding suffer from problems such as an imperfect identification technology system, low breeding efficiency, and insufficient phenotypic prediction accuracy. Traditional methods are difficult to effectively improve the screening accuracy of soybean salt tolerance traits.
Using machine learning-based genome-wide selection technology, a soybean salt stress phenotype prediction model was constructed by combining genome-wide association analysis and a convolutional neural network model (ResCANet) with CA attention mechanism and multiple residual module. SNP sites that are significantly associated with the phenotype were screened to predict soybean salt tolerance traits.
It significantly improves the prediction accuracy of soybean salt tolerance traits, reduces field breeding costs, and enhances breeding efficiency. It can more accurately identify phenotypic genetic variations, thereby improving breeding efficiency.
Smart Images

Figure CN119889448B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of salt-tolerant soybean whole genome selection, and particularly relates to a machine learning-based whole genome selection technology for salt-tolerance traits of soybean and application thereof. BACKGROUND
[0002] Soybean (Glycine max) is the fourth largest crop in the world and the largest oil crop in terms of planting area. Stabilizing or increasing domestic soybean supply is of great strategic significance to maintaining national grain and oil security and sustainable agricultural development. Therefore, developing salt-tolerant soybean varieties, improving soybean yield and quality, improving land utilization, increasing farmers' income and agricultural economic benefits, and promoting sustainable development of China's agriculture all have important theoretical and practical significance.
[0003] At present, the research basis of soybean salt tolerance is relatively weak, and there are outstanding problems such as imperfect salt-tolerant identification technology system and low breeding efficiency. Traditional breeding mode mainly relies on specific phenotype observation and breeder's experience, combined with simple biological statistical analysis, accompanied by a series of problems such as randomness, low efficiency, weak comprehensive evaluation ability and long breeding cycle. Whole genome selection technology is a new selection breeding strategy for predicting individual phenotypes by genotypes, which can effectively reduce the high field cost brought by traditional phenotype selection and improve breeding efficiency.
[0004] However, early genome selection technology uses fewer genetic variation sites, and uses traditional statistical methods such as regression method, and the prediction accuracy of phenotype still needs to be improved. With the development of high-throughput sequencing and artificial intelligence, whole genome resequencing and machine learning methods are gradually applied to crop genome selection breeding. The existing genome selection method mainly uses whole SNP data analysis, ignoring the important influence of some key SNPs on phenotype, which may lead to unsatisfactory prediction results. The existing machine learning prediction model may have certain limitations in feature extraction and fusion of SNP data, thereby affecting the prediction accuracy.
[0005] Therefore, how to improve the screening accuracy of salt-tolerant soybean has become a problem to be solved in the field of breeding. SUMMARY
[0006] In order to solve the above problems in the prior art, the application provides a machine learning-based whole genome selection technology for salt-tolerant soybean.
[0007] According to the first embodiment of the application, a machine learning-based whole genome selection method for salt-tolerant soybean is provided, which comprises the following steps:
[0008] 1) Data acquisition: obtaining soybean population materials and their whole genome resequencing data; determining the leaf scorch index, Na + , K+ Ca 2+ Cl - content, etc. The SNPs were strictly controlled, and individuals with high deletion rates were deleted. The deletion rate, heterozygosity, and minor allele frequency (MAF) of the SNP sites were checked and screened in detail.
[0009] 2) Genome-wide association study (GWAS) of salt tolerance phenotype of soybean: Plink was used to analyze the population structure and estimate the genetic relationship between samples. The genotype and phenotype were associated by a mixed linear model. The genetic correlation matrix and genetic relationship matrix of the population structure were combined to reduce potential confounding effects, and high-quality SNP sites significantly associated with the phenotype were obtained.
[0010] 3) Model establishment: select SNP sites significantly associated with the phenotype as input features of the model. The soybean resequencing dataset is divided into a training set and a test set. After training the training set using machine learning methods, the model is established. The ResCANet model is constructed by using the coordinate attention (CA) attention mechanism and the multiple residual module.
[0011] (Coordinate Attention) attention mechanism and the multiple residual module.
[0012] 4) Model accuracy evaluation: input the SNP sites significantly associated with the salt stress phenotype and the salt stress phenotype into the linear regression (LR) and machine learning SVR model for prediction. Compare the prediction accuracy and mean squared error (MSE) of the LR, SVR, and ResCANet models. The accuracy of the ResCANet prediction model is significantly improved compared with the LR and SVR models, and the MSE is significantly reduced.
[0013] The second embodiment of the application provides a machine learning-based salt-tolerant soybean whole genome selection system in the first embodiment, the system comprising the following modules:
[0014] 1) Genotype acquisition module: acquire the leaf scorch, Na + K + Ca 2+ Cl - content phenotypes of the soybean germplasm resource population after salt treatment, obtain the resequencing data of the soybean germplasm resource, and screen high-quality SNP sites by whole genome association analysis of the salt stress phenotype;
[0015] 2) Model building module: for the soybean salt stress phenotype and genotype dataset, five-fold cross-validation is used, and for each fold, the data is divided into five subsets, four of which are used for training, and the remaining one is used for testing, and the Pearson correlation coefficient PCC is calculated; the CA attention mechanism module and the double residual module are introduced to build the soybean salt stress machine learning prediction model ResCANet;
[0016] 3) Soybean population phenotype prediction module of salt stress phenotype to be detected: input the soybean population genotype data SNP to be predicted, use the soybean salt stress phenotype machine learning ResCANet model to predict the salt stress phenotype of each individual in the soybean population; according to the predicted phenotype, select salt-tolerant soybean varieties for field phenotype experiment.
[0017] The third embodiment of the application provides an application of the machine learning-based salt-tolerant soybean whole genome selection technology in the previous embodiment in screening and identifying salt-tolerant soybeans, comprising the following steps:
[0018] 1) Obtain the soybean sample to be detected, and perform whole genome resequencing on it;
[0019] 2) Obtain the whole genome SNP site of the sample to be detected;
[0020] 3) Strictly quality control the SNP, delete individuals with high deletion rate, and perform detailed inspection and screening on the deletion rate, heterozygosity and sub-MAF of the SNP site.
[0021] 4) Select the SNP site significantly related to the phenotype and input it into the machine learning ResCANet prediction model to predict the phenotype of soybean salt stress.
[0022] The fourth embodiment of the application provides an application of the machine learning-based salt-tolerant soybean whole genome selection technology in the previous embodiment in soybean variety breeding or assisted soybean breeding, comprising the following steps:
[0023] 1) Obtain the soybean sample to be detected, and perform whole genome resequencing on it;
[0024] 2) Obtain the whole genome SNP site of the sample to be detected;
[0025] 3) Strictly quality control the SNP, delete individuals with high deletion rate, and perform detailed inspection and screening on the deletion rate, heterozygosity and sub-MAF of the SNP site.
[0026] 4) Select the SNP site significantly related to the phenotype and input it into the machine learning ResCANet prediction model to predict the phenotype of soybean salt stress.
[0027] 5) a determination module for combining salt-tolerant soybean selection and breeding, according to the prediction results of the phenotypic data, selecting soybean varieties with salt-tolerant potential for laboratory and field salt-tolerant selection.
[0028] The machine learning-based salt-tolerant soybean whole genome selection technology provided by the embodiments of the present application first investigates the genotypes and salt stress phenotypes of the soybean population used for modeling, constructs a model using the convolutional neural network of the CA attention mechanism and multiple residual modules of machine learning, and predicts the salt stress phenotype of the soybean sample to be tested using the model to guide the selection and breeding of salt-tolerant soybeans.
[0029] The beneficial effects of the present example are that the soybean salt stress phenotype can be effectively predicted, the field breeding cost is reduced, and the salt-tolerant soybean breeding efficiency is significantly improved. The advantages of the prediction model are:
[0030] 1. By selecting SNPs significantly related to the phenotype as input features of the model, the model can focus on important variations while reducing the number of features, improving the training efficiency and performance of the model.
[0031] 2. Introducing a multiple attention mechanism module, by assigning multiple attention weights to different genetic variation sites, the model can simultaneously focus on multiple features of key sites at different levels. This method can effectively capture complex genetic associations and improve the model's ability to distinguish the importance of different genetic variation sites, thereby more accurately identifying genetic variations related to the phenotype.
[0032] 3. Introducing the CA attention mechanism, which can adaptively adjust the weight of each feature channel according to its importance. This mechanism enhances the feature response of important channels and suppresses irrelevant or redundant channels, thereby helping the model to focus more effectively on key genetic features and improving the expression and learning ability of genetic data. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 A flowchart of the soybean salt-tolerant phenotype prediction method based on whole genome association analysis and machine learning of the present application
[0034] Figure 2 A Manhattan plot for GWAS analysis of leaf scorch index of soybean after salt stress in the embodiments
[0035] Figure 3 A Manhattan plot for GWAS analysis of Na + , K + , Cl - and Ca 2+ ion content of soybean after salt stress in the embodiments
[0036] Figure 4A convolutional neural network framework diagram based on CA attention mechanism and multiple residual modules for the prediction model ResCANet in the embodiment
[0037] Figure 5 A comparison diagram of the prediction effect of the ResCANet prediction model and the SVR and linear regression model in the embodiment
[0038] Figure 6 A soybean recombinant inbred line parent phenotype real scene diagram after salt stress in the embodiment
[0039] Figure 7 A comparison diagram between the predicted value and the true value obtained by using the model ResCANet for one test sample in the embodiment DETAILED DESCRIPTION
[0040] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings Figures 1-6 , and the present application will be further described in detail.
[0041] The present application relates to the field of soybean salt-tolerant phenotype prediction based on genome-wide association analysis, and particularly relates to a convolutional neural network prediction method combining CA attention mechanism and residual module, as shown in Figure 1 .
[0042] The specific scheme provided by the present application is as follows:
[0043] Embodiment 1. Construction of a soybean salt stress phenotype prediction model based on machine learning
[0044] 1) Sample phenotype and genotype collection
[0045] 409 soybean natural populations were collected, and after the first three-leaf leaf of soybean unfolded, the soybean was treated with 150mM NaCl for 14 days. The number of leaf scorch, Na + , K + , Ca 2+ , Cl - content of the soybean grown under salt treatment and control conditions were identified. The leaf scorch index and relative ion content of each variety were calculated. Relative value = salt treatment / control x 100%
[0046] The whole genome resequencing data was aligned to the soybean Wm82.a4 version reference genome, and the whole genome SNP data of the soybean population was extracted. The obtained SNP data was strictly quality controlled to check the integrity and reliability of the data. Specifically, it included retaining double allelic variation SNP sites, removing sites with heterozygosity greater than 30%, filtering individuals with deletion rate higher than 2%, removing samples with SNP deletion greater than 5%, removing SNP sites with deletion rate higher than 10% in samples, and removing SNP sites with minor allele frequency less than 0.05 (MAF < 0.05).
[0047] 2) Whole genome association analysis of soybean salt stress phenotype
[0048] After SNP quality control, linkage disequilibrium analysis, population structure analysis and kinship analysis were used to reveal the distribution of genetic variation, genetic relationship between samples and its influence on phenotype, so as to improve the accuracy of model prediction and optimize breeding strategy. Plink was used to analyze the population structure and estimate the kinship between samples, and the genotype and phenotype were analyzed by mixed linear model for GWAS analysis, and 1e-3 threshold was used to screen significant SNPs, and the GWAS results were compared with linkage disequilibrium (LD). The results of whole genome association analysis of salt stress phenotype are shown in Figure 2 and Figure 3 . If a group of sites shows significant association in GWAS and shows correlation in LD, it can be considered that the reliability of the analysis result is high. Using this method, 7011 SNP sites were screened out which were significantly associated with leaf scorch index, 17799 SNP sites were significantly associated with Na + , 4692 SNP sites were significantly associated with K + , and 10515 SNP sites were significantly associated with Ca 2+ .
[0049] 3) Construction of prediction model ResCANet
[0050] The SNP sites significantly associated with phenotypes and soybean salt stress phenotypes are used as input to build the prediction model. The model uses the CA mechanism to focus on key SNP sites and enhances feature extraction and modeling capabilities through multiple residual modules. The CA attention mechanism can significantly improve the predictive ability and interpretability of the model in genomic selection and genetic research by assigning more precise weights to different SNP sites. In the analysis of genomic data, SNP sites often need to be fused at different levels to capture potential interactions. The two residual modules can help the model identify complex interactions between different SNP sites by deeper information fusion, thereby improving the model's ability to model complex gene-phenotype relationships. The first residual module is mainly responsible for local feature extraction and preliminary fusion, while the second residual module further integrates these features to capture higher-level abstract representations. This two-level structure not only helps the model learn complex relationships between SNP sites more effectively, but also speeds up the training process and improves the model's robustness and accuracy in large-scale, high-dimensional genetic data. In the ResCANet prediction model, each residual block is composed of three convolution branches with ReLU activation functions and a feature fusion layer. The fused features are first extracted in depth by a one-dimensional convolution layer containing 10 filters, then the key features are weighted and enhanced by the channel attention mechanism; then, the features are corrected by the batch normalization layer, and the essential information is further refined by the max pooling layer; finally, the flattened features are sent to a fully connected layer containing 32 neurons, and under the protection of a 0.4 proportion Dropout regularization, a single neuron fully connected layer outputs the prediction results, as shown in Figure 4 .
[0051] Five-fold cross-validation is used, and for each fold, the data is divided into five subsets, four of which are used for training, and the remaining one is used for testing, and the Pearson correlation coefficient (PCC) is calculated. Finally, by calculating the average PCC value of the five folds, the comprehensive performance evaluation of the model is obtained.
[0052] The model uses the Mean Absolute Error (MAE) as the loss function and uses the stochastic gradient descent (SGD) optimizer for training, with a learning rate of 1×10 -3 and a weight decay coefficient of 1×10 -6, the training period (epoch) is set to 200. In order to accelerate the convergence, the training data is divided into multiple small batches, and the batch size is set to 16. In addition, the model is evaluated by five-fold cross-validation on the data set, and the data is divided into five subsets, and each subset is used as the test set in turn, and the rest is used as the training set. This method not only makes full use of limited data, reduces the bias caused by the randomness of data division, but also effectively reduces the risk of model overfitting. Through multiple rounds of validation, cross-validation helps to optimize hyperparameters and model structure, further improving the stability and generalization ability of the model, ensuring the reliability and accuracy of the evaluation results. The correlation coefficient PCC is calculated as follows:
[0053]
[0054] PCC is used to measure the linear correlation between two variables, and its value ranges from -1 to 1. When PCC>0, it indicates that the two variables are positively correlated, that is, the increase of one variable is accompanied by the increase of the other variable; when PCC=0, it indicates that there is no linear relationship between the two variables, that is, there is no any predictable linear dependence between them.
[0055] 4) ResCANet prediction model evaluation
[0056] The SNP sites significantly associated with salt stress phenotype and salt stress phenotype are input into the linear regression (Linear Regression) and machine learning SVR model for prediction. After 200 simulations, the model prediction accuracy is compared with the ResCANet prediction model prediction accuracy, and it is found that the average accuracy of the ResCANet model is 0.88, the average accuracy of the SVR model is 0.55, and the average accuracy of the LR model is 0.43. The MSE of the ResCANet model is 0.006, which is significantly lower than the other two models. The comparison of the prediction effect of the ResCANet prediction model with the SVR and linear regression model prediction effect is shown in Figure 5 . It shows that the accuracy of the ResCANet prediction model is the highest, and it is suitable for salt stress phenotype prediction.
[0057] Example 2, prediction of salt stress phenotype of soybean natural population resequencing data
[0058] 1) Collection of samples to be tested
[0059] We collected a soybean recombinant inbred line (RIL) population containing 177 individuals, and the two parents of the population had obvious differences in leaf scorch number after salt treatment Figure 6). The RIL population was re-sequenced, and the whole genome alignment was performed to extract the whole genome SNP sites of all the tested materials. According to the SNP sites identified in the first embodiment which were significantly associated with the soybean salt stress phenotype, the SNP sites of the tested samples were extracted.
[0060] 2) Predicting the salt stress phenotype of the tested sample
[0061] The extracted SNP sites were input into the ResCANet prediction model to obtain the predicted soybean salt stress phenotype value of each tested sample.
[0062] 3) Evaluating the accuracy of the predicted soybean salt stress phenotype
[0063] The tested materials were planted in a greenhouse (12 hours light / 12 hours dark), and after the first trifoliate leaf unfolded, the soybean materials were treated with 150 mM NaCl for 14 days, and the leaf scorch index and relative ion content were determined. The Pearson correlation coefficient between the predicted value and the observed value was calculated using the ResCANet method, and it was found that the average correlation coefficient between the predicted phenotype and the observed phenotype reached 0.855, as shown in Figure 7 It is proved that the whole genome selection technology significantly improves the selection efficiency and significantly reduces the field workload.
Claims
1. A machine learning-based method for selecting salt-tolerant soybean genome-wide, characterized in that... The method includes the following steps: 1) Obtain the genotypes and salt-treated phenotypes of the dataset used for modeling, and then process them; 2) Strict quality control was performed on SNP sites and deletion samples; 3) Perform genome-wide association analysis on effective variant sites to obtain high-quality variant sites; 4) High-quality mutation site data from the training set were imported into a convolutional neural network with CA attention mechanism and multiple residual modules to build a model, which was named ResCANet. 5) Use the test set to verify the model's prediction performance and determine the optimal model; 6) Obtain resequencing data of soybean populations with salt stress phenotypes to be tested, extract high-quality variant sites, and use the optimal model to predict salt stress phenotypes.
2. The method according to claim 1, characterized in that... It also includes the step of selecting salt-tolerant soybean varieties for field phenotypic experiments based on the predicted soybean salt stress phenotype.
3. The method according to claim 1, characterized in that... Step 1) specifically includes: 1-1) Collect natural soybean populations and treat soybeans with 150mM NaCl for 14 days after the first trifoliate compound leaf of soybeans unfolds. 1-2) Identify the number of scorched leaves and Na+ in soybeans grown under salt treatment and control conditions. + Content, K + Content, Ca 2+ Content, Cl - content; 1-3) Calculate the leaf scorch index and relative ion content for each variety. Relative value = ion content of each soybean under salt treatment conditions / ion content of each soybean under control conditions x 100%; 1-4) Perform whole-genome resequencing on the natural soybean population, align the resequencing data to the soybean Wm82.a4 version reference genome, and extract whole-genome SNP data of the soybean population.
4. The method according to claim 1, characterized in that... The strict quality control mentioned in step 2) specifically includes: Biallelic variant SNP sites were retained, sites with heterozygosity greater than 30% were removed, individuals with deletion rates greater than 2% were filtered out, samples with SNP deletions greater than 5% were removed, SNP sites with deletion rates greater than 10% in the samples were removed, and SNP sites with minor allele frequencies (MAF) less than 0.05 were removed.
5. The method according to claim 1, characterized in that... The genome-wide association analysis described in step 3) specifically involves: using Plink to analyze the population structure and estimate the kinship between samples; simultaneously, using a mixed linear model to perform association analysis on genotypes and phenotypes; and combining the genetic correlation matrix and kinship matrix of the population structure to reduce potential confounding effects and obtain high-quality variant sites.
6. The method according to claim 1, characterized in that... Step 4) describes establishing a model that utilizes the CA mechanism to focus on key SNP sites and enhances feature extraction and modeling capabilities through a multiple residual module; wherein: By employing the CA attention mechanism to assign more refined weights to different SNP sites, the predictive power and interpretability of models in genome selection and genetic research are significantly improved. Employing a dual residual module enables deeper information fusion, helping the model identify complex interactions between different SNP sites and thus improving its ability to model complex gene-phenotype relationships. The first residual module is responsible for extracting and initially fusing local features, while the second residual module further integrates these features to capture higher-level abstract representations. This dual structure not only helps the model learn complex relationships between SNP sites more effectively but also accelerates the training process and improves the model's robustness and accuracy in large-scale, high-dimensional genomic data. In the ResCANet prediction model, each residual block consists of three convolutional branches carrying ReLU activation functions and a feature fusion layer. The fused features are first extracted through a one-dimensional convolutional layer with 10 filters, and then key features are weighted and enhanced through a channel attention mechanism. Next, the features are corrected for data distribution through a batch normalization layer, and the essential information is further extracted by a max pooling layer. Finally, the flattened features are fed into a fully connected layer with 32 neurons, and under the protection of 0.4 Dropout regularization, the prediction result is output through a single-neuron fully connected layer. The model uses Mean Absolute Error (MAE) as the loss function and is trained using a Stochastic Gradient Descent (SGD) optimizer with a learning rate of 1×10⁻⁶. -3 The weight decay coefficient is 1×10 -6 The training epoch was set to 200. To accelerate convergence, the training data was divided into multiple mini-batches with a batch size of 16. In addition, the model was evaluated using five-fold cross-validation, which divided the data into five subsets, using each subset as the test set and the rest as the training set, and calculating the Pearson correlation coefficient (PCC). Finally, the overall performance evaluation of the model was obtained by calculating the average PCC value of the five-fold cross-validation.
7. The method according to claim 1, characterized in that... Genotypic data for soybean varieties, including salt-tolerant and salt-sensitive varieties.
8. A machine learning-based whole-genome selection system for salt-tolerant soybeans, characterized in that... The system includes the following modules: 1) Genotype Acquisition Module: Obtains the number of scorched leaves and Na+ chromatograms of soybean germplasm resource populations after salt treatment. + Content, K + Content, Ca 2+ Content, Cl - Content, obtained soybean germplasm resource resequencing data; used salt stress phenotype genome-wide association analysis to screen high-quality SNP loci; 2) Model building module: For the soybean salt stress phenotype and genotype datasets, five-fold cross-validation is used. For each fold, the data is divided into five subsets, four of which are used for training and the remaining subset is used for testing. The Pearson correlation coefficient PCC is calculated. The CA attention mechanism module and the two residual module are introduced to build the soybean salt stress machine learning prediction model ResCANet. 3) Soybean population phenotype prediction module for salt stress phenotype to be tested: Input the genotype data SNP of the soybean population to be predicted, use the ResCANet model of soybean salt stress phenotype machine learning to predict the salt stress phenotype of each individual in the soybean population; select salt-tolerant soybean varieties for field phenotype experiments based on the predicted phenotype.
9. The application of the machine learning-based whole-genome selection method for salt-tolerant soybeans as described in any one of claims 1-7 in the screening and identification of salt-tolerant soybeans.
10. The application of the machine learning-based whole-genome selection method for salt-tolerant soybeans as described in any one of claims 1-7 in soybean variety breeding or assisted soybean breeding.
Citation Information
Patent Citations
Method for efficiently predicting salt tolerance of chrysanthemum based on whole genome
CN116732222A
Method for screening salt-tolerant characters of tomatoes
CN118887998A