SNP (Single Nucleotide Polymorphism) molecular marker set obviously related to lodging property of soybeans and application thereof

By using genome-wide association analysis and model construction, and utilizing a set of SNP molecular markers that are significantly associated with soybean lodging, the problem of low efficiency in traditional breeding methods has been solved, enabling early screening and efficient breeding of soybean lodging resistance.

CN121362847APending Publication Date: 2026-01-20THE SHENNONG LABORATORY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511436516.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Traditional breeding methods are inefficient in identifying lodging resistance traits in soybeans, cannot perform early screening, and are costly, making it difficult to achieve efficient breeding.

Method used

Twenty-two SNP molecular markers significantly associated with soybean lodging were identified using genome-wide association analysis. A genome selection model was constructed by combining LASSO regression and RRBLUP models, and significant SNP markers were used as fixed effects for early prediction and selection of lodging-resistant lines.

Benefits of technology

It significantly improved the accuracy of soybean lodging prediction and breeding efficiency, reduced costs, enabled early screening of superior individuals, and shortened the breeding generation interval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121362847A_ABST
    Figure CN121362847A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of soybean breeding, and relates to soybean lodging research. The invention provides an SNP (Single Nucleotide Polymorphism) molecular marker set remarkably related to lodging resistance of soybeans and application of the SNP molecular marker set. The marker set can be directly used for genome selection modeling and molecular marker-assisted selection of the lodging resistance of the soybeans. More importantly, the gene is taken as a fixed effect to be incorporated into a prediction model, so that the prediction performance of a soybean lodging genome selection model is remarkably improved, and lodging-resistant strains can be selected and predicted in early-stage individuals, so that the genetic improvement generation interval is effectively shortened, the selection accuracy and the breeding efficiency are improved, and the breeding cost is reduced. Finally, the lodging-resistant breeding process of the soybeans is promoted, and the breeding cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of soybean breeding and relates to the study of soybean lodging. Background Technology

[0002] Soybeans Glycine max L. Merrill, native to China, is one of the world's most important grain, oilseed, and cash crops, and is widely cultivated in China and around the world. Therefore, increasing soybean yield is crucial. In my country, the southern, Huang-Huai, and northern regions, where soybean yields are relatively high, also have correspondingly high planting densities (186,000–256,000 plants / hm²). 2 While increasing planting density can increase the photosynthetic production area, reduce production costs, and improve agricultural economic efficiency, it also increases the risk of lodging (LG). Lodging can lead to severe yield reduction in soybeans (reductions of 11-32%), hinder mechanized harvesting, and cause significant economic losses to agricultural production.

[0003] Soybean lodging is a complex trait caused by multiple factors, including meteorological conditions, environmental conditions, cultivation practices, and genetic background. Although its mechanism is complex, genetic factors play a dominant role, with a heritability as high as 0.53–0.93. Developing new high-yielding, dense-planting-tolerant, and lodging-resistant soybean germplasm is one of the effective ways to increase soybean yield. However, traditional breeding relies on the observation and selection of natural variations. From hybridization to trait stability, multiple generations of trait screening and stability testing are often required, a process that consumes a significant amount of time and effort, resulting in low breeding efficiency. Especially for traits such as lodging resistance, accurate identification is often only possible in the later stages of growth, making early screening impossible. Therefore, identifying single nucleotide polymorphism (SNP) sites related to soybean lodging resistance can provide reliable molecular markers for genetic improvement of soybean lodging resistance, marker-assisted selection, and genome-assisted selection breeding, thereby reducing breeding workload, saving costs, and improving breeding efficiency. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, one objective of this invention is to provide a set of SNP molecular markers significantly associated with soybean lodging resistance and their applications. This marker set can be directly used for genomic selection modeling and marker-assisted selection of soybean lodging resistance. More importantly, incorporating these markers as fixed effects into the prediction model not only significantly improves the predictive performance of the soybean lodging resistance genomic selection model but also enables the selection and prediction of lodging-resistant lines in early individuals. This effectively shortens the generation interval for genetic improvement, increases selection accuracy and breeding efficiency, and ultimately promotes the soybean lodging resistance breeding process while reducing breeding costs.

[0005] The technical solution of this invention is implemented as follows: In the first aspect of the present application, a SNP molecular marker set significantly associated with lodging of soybean is provided, which is mainly produced by the following steps: using multi-environment soybean lodging index data and 117k chip data for whole genome association analysis, 22 quantitative trait loci (QTL) significantly associated with lodging of soybean are identified, and the most significant SNP combinations of the QTL are collected to form the marker set.

[0006] Specifically, the SNP molecular marker set comprises the following 22 markers: Gm01_50178213, Gm03_3146453, Gm06_21136539, Gm06_21196104, Gm06_21512079, Gm06_23082606, Gm06_24462733, Gm06_27756888, Gm06_29537534, Gm06_34566910, Gm06_34744229, Gm06_38592928, Gm09_46773948, Gm11_25341100, Gm14_7160350, Gm18_48341696, Gm18_54496089, Gm19_1524582, Gm19_1603096, Gm19_44839696, Gm19_45054985 and Gm19_45924944.

[0007] Further, the SNP molecular marker set sites are as follows: Gm06_21136539, Gm06_21196104, Gm06_21512079, Gm06_23082606, Gm06_24462733, Gm06_27756888, Gm06_29537534, Gm06_34566910, Gm06_34744229, Gm06_38592928, Gm09_46773948, Gm11_25341100, Gm18_48341696, Gm18_54496089 and Gm19_1603096. In another aspect, the use of the SNP molecular marker set significantly associated with lodging of soybean as a template for designing primers or probes is also provided, which can be realized by existing biological technology means and the requirements of designing primers.

[0008] In the third aspect, the primers or probes for detecting the above-mentioned SNP molecular marker set are claimed.

[0009] In the fourth aspect, the use of the SNP molecular marker set significantly associated with lodging of soybean in multi-marker combination prediction of lodging index is claimed.

[0010] The SNP molecular marker set of the application can be directly used for LASSO genomic selection model construction, can effectively reduce the genotyping cost, the model constructed comprehensively integrates multiple significant markers, has strong generalization ability, can improve the prediction stability and overall selection efficiency, and has important application value in rapid screening of large-scale offspring population. The prediction model is constructed by using LASSO regression, using the glmnet package, using the genotype matrix of the SNP molecular marker set as the feature, using the multiple environment lodging index BLUP as the response variable, fitting the model on the training set, determining the optimal regularization parameter (lambda) through 5-fold cross-validation, and applying the model to the genotype data of the test set to obtain the estimated breeding value.

[0011] The SNP molecular marker set can be added as a fixed effect in the construction of the genomic selection model, which can improve the prediction accuracy and selection efficiency. The steps are: encoding the whole genome marker data into a numerical matrix, using the rrblup package in R language, calling the RRBLUP method to construct a genotype selection model, the model is based on a mixed linear model framework, the whole genome SNP marker is taken as a random effect, and the SNP molecular marker set of claim 1 or claim 2 is taken as a fixed effect, which is included in the model to predict the lodging index.

[0012] The application has the following beneficial effects: 1. The traditional genomic selection method such as RRBLUP, etc. defaults that all markers follow the same distribution, combines whole genome association analysis or linkage analysis, etc. to analyze the genetic basis of the target trait from the genetic point of view, and adds significant SNPs as fixed effects to the prediction model. The genotype model can not only distinguish the size of significant SNP effects, but also optimize the model bias, so as to more accurately estimate the marker effect. The application provides a soybean lodging significant related SNP molecular marker combination which can be integrated as a fixed effect into a whole genome prediction model. This combined strategy can effectively improve the prediction accuracy of soybean lodging, guide efficient breeding decision-making of offspring, accelerate breeding progress, improve breeding efficiency, and screen out lodging-resistant plants, thereby laying a foundation for breeding high-yield and lodging-resistant varieties.

[0013] 2. Although sequencing technology is developing, high-throughput sequencing of large-scale populations is costly. In large-scale early and middle-stage offspring population screening, customizing a low-density chip containing lodging significant related sites, or even other trait related sites, is beneficial to early rapid screening of excellent individuals or evaluation of parents, saving breeding time and reducing cost. The markers in the marker set can also be used to develop KASP (Kompetitive Allele-Specific PCR) markers for molecular marker-assisted selection, realizing early and low-cost genetic evaluation and individual selection. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are only some of the embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative effort based on these drawings are within the scope of protection of the present application.

[0015] Figure 1 Figure 5 is the distribution of lodging index in five environments; wherein LG, lodging index; 20JZ, Jingzhou in 20; 22SZ, Suzhou in 22; 23SZ, Suzhou in 23; 23JN, Jining in 23; 23XX, Xinxiang in 23.

[0016] Figure 2 Figure 6 is the correlation analysis of lodging index in five environments; wherein 20JZ, Jingzhou in 20; 22SZ, Suzhou in 22; 23SZ, Suzhou in 23; 23JN, Jining in 23; 23XX, Xinxiang in 23.

[0017] Figure 3 Figure 7 is the selection efficiency of different environments lodging index based on different marker sets and five selection ratios; wherein the horizontal axis is the selection ratio, and the threshold line is the selection efficiency of corresponding environmental phenotype selection; 20JZ, Jingzhou in 20; 22SZ, Suzhou in 22; 23SZ, Suzhou in 23; 23JN, Jining in 23; 23XX, Xinxiang in 23.

[0018] Figure 4 Figure 8 is the selection effect of different environments lodging index based on different marker sets and five selection ratios; wherein the horizontal axis is the selection ratio; 20JZ, Jingzhou in 20; 22SZ, Suzhou in 22; 23SZ, Suzhou in 23; 23JN, Jining in 23; 23XX, Xinxiang in 23. DETAILED DESCRIPTION

[0019] The technical solutions of the present application will be described clearly and completely below in combination with the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0020] The experimental methods used in the following experimental examples are conventional methods unless otherwise specified; the materials, reagents, etc. used are reagents and materials available from commercial channels unless otherwise specified. EMBODIMENT

[0021] The embodiment provided by the application is a soybean lodging significantly related marker set identification and whole genome selection. The embodiment is based on five environment lodging index data and multi-environment lodging index best linear unbiased prediction, 22 QTLs are obtained by using whole genome association analysis, and the most significant SNP set is used as a key marker set to perform genome selection prediction or as a fixed effect to add to a prediction model to predict the lodging index, and the specific process is as follows: 1 Material and method 1.1 Test material, phenotype identification and analysis In this study, a NAM (Nested Association Mapping) population with 'Zhongdou 41' as the common maternal parent was used, and 35 paternal parents were breeding varieties and local varieties of different ecological types in 18 provinces (Song Jian et al. 2024, Crop Science). The NAM population was planted in the summer of 2020 (F8) in the Science and Technology Industrial Park of Yangtze University (112.06°E, 30.37°N), with a plant spacing of 0.1 m and a row spacing of 0.5 m, and the planting density was 200,000 plants / hm 2 . In order to analyze the lodging resistance of different planting densities, the plant spacing and row spacing were adjusted, and the summer of 2022 (F 11 ) was planted in the Wanbei Comprehensive Experimental Station of Anhui Agricultural University (117.10°E, 33.69°N), and the planting density was 125,000 plants / hm 2 . In the summer of 2023 (F 12 ), some single plants were selected for plot tests in Wanbei Experimental Station, Xinxiang Experimental Base of Chinese Academy of Agricultural Sciences (113.78°E, 35.15°N), and Zhenlong Family Farm in Houcheng District, Jining City (116.48°E, 35.54°N), with a plant spacing of 0.08 m and a row spacing of 0.4 m, and the planting density was 312,500 plants / hm 2 . In the middle and late stages of soybean growth and development, all plants in the test plot were observed, and the lodging index of the plants was investigated, with a score range from 1 (all plants are straight) to 4 (all plants are prostrate), with an interval of 1 (1, no lodging: 0%; 2, light lodging: 0% x ≤30%; 3, moderate lodging: 30% x ≤60%; 4, severe lodging: 60% x ).

[0022] In this study, the distribution and correlation of the lodging index data of the five environments were statistically analyzed, and the best linear unbiased prediction (BLUP, Best Linear Unbiased Prediction) value and single-repetition heritability of the multi-environment lodging index were calculated, and the formula is as follows: H2 is the heritability; Vg is the genetic variance; Ve is the residual variance; and L is the number of environments.

[0023] 1.2 Genetic analysis Genotype identification was completed based on the single nucleotide polymorphism chip, and a total of 159,035 SNP markers were identified (Song et al. 2024, Chinese Journal of Crop Science). A total of 612 samples were extracted from the chip data, and the genotype data was converted to vcf format, and the non-polymorphic data was removed, containing 117,409 SNP markers. Using PLINK software to remove markers with less than 90% completeness (— geno0.05), and minor allele frequency less than 0.01 (— maf 0.01), leaving 98,502 SNPs. Using Beagle5.4 to fill in the missing genotypes. Based on the phenotype and genotype data, EMMAX (Efficient Mixed-Model Association Expedited) was used to locate QTL for different environments and multi-environment BLUP values, and the upstream and downstream of the significant site were taken as the QTL interval.

[0024] 1.3 Single marker and multi-marker combination prediction In this study, regression method was used to take the most significant SNP genotype of QTL and the phenotype of the corresponding environment to construct a linear model to evaluate its breeding value, evaluate the prediction accuracy of single marker and calculate the selection efficiency of single marker.

[0025] Single marker selection efficiency: In the formula, is the number of excellent haplotype samples, X i is the number of samples that are not prone to lodging (lodging index is 1) or multi-environment BLUP value is less than 0.

[0026] The significant correlation marker prediction model was constructed using LASSO (Least Absolute Shrinkage and Selection Operator) regression. Using the glmnet package, the genotype matrix of the most significant SNP marker of all QTL was used as the feature, and the five environments and multi-environment BLUP lodging index were used as the response variable, and the model was fitted on the training set. The optimal regularization parameter (λ) was determined by 5-fold cross-validation, and the model was applied to the genotype data of the test set to obtain the estimated breeding value.

[0027] 2.4 Genome selection The whole genome marker data were coded into a numerical matrix, and the RRBLUP method was used to construct the genomic selection model based on the mixed linear model framework using the rrblup package in R language. In addition, in order to evaluate the effect of the significant SNP marker set, the marker set was included in the above model as a fixed effect to predict the lodging index.

[0028] 2.5 Cross-validation In this study, single marker regression prediction, multi-marker combination regression prediction and genomic selection were all used 5-fold cross-validation. The marker set and the corresponding environment lodging index were randomly divided into 5 equal parts, and one of them was used as the validation set, and the remaining 5 parts were used as the training set. The prediction accuracy of the model was evaluated by calculating the Pearson correlation coefficient between the test set estimated breeding value and the test set phenotype value. This process used the same 20 random seeds and was iterated 100 times, and the average value was taken as the prediction accuracy.

[0029] In order to evaluate the genetic gain of the model for breeding, the selection efficiency and selection effect of the multi-environment BLUP value of different models were calculated. The selection proportion was set to 0.1, 0.15, 0.2, 0.25, 0.3. The selection efficiency and selection effect formulae are as follows: Selection efficiency: In the formula, SI is the selection proportion, n is the number of test set samples, SI * n is the number of selected individuals, X i is the number of samples that are not lodging (lodging index is 1) or multi-environment BLUP value is less than 0.

[0030] Selection effect: In the formula, is the mean of the test set phenotype lodging index, and the estimated breeding value is sorted in the positive direction. The first SI * n individuals are selected as the selection group, is the mean of the lodging index of the selection group.

[0031] 3 Results and analysis 3.1 Phenotype analysis The variation range of average lodging index of 612 materials planted in Jingzhou in 2020, Suzhou in 2022, Suzhou in 2023, Jining in 2023 and Xinxiang in 2023 was between 1.79 and 3.39, and the multi-environment BLUP was between -1.15 and 0.99. Among them, the average lodging score of Jingzhou in 2020 was the lowest, and the proportion of non-lodging and light lodging lines was 75.5%; while in Xinxiang in 2023, severe lodging occurred, and the proportion of severe lodging was as high as 72.6% Figure 1 ). The results of analysis showed that the genetic force in multiple environments reached 0.78, and the lodging index among five environments showed significant correlation Figure 2 .

[0032] 3.2 Genetic analysis and marker prediction Taking -lg(7 / 98502) as the threshold, EMMAX detected 22 lodging index-related QTLs in five environments, which were distributed on eight chromosomes (Chr01, Chr03, Chr06, Chr9, Chr11, Chr14, Chr18, Chr19) (Table 1).

[0033] Table 1 Lodging index QTLs located by EMMAX in five environments Note: LG, lodging index; 20JZ, Jingzhou in 2020; 22SZ, Suzhou in 2022; 23SZ, Suzhou in 2023; 23JN, Jining in 2023; 23XX, Xinxiang in 2023 The 22 significant SNPs were derived from the results of whole-genome association analysis of lodging index in five environments and multi-environment BLUP, and the-lg P was between 4.31 and 8.75, and the difference between REF type and ALT type haplotype was between 0.07 and 1.67 (Table 2).

[0034] Table 2 The most significant SNP sites in 22 lodging index QTLs Note: LG, lodging index; 20JZ, Jingzhou in 2020; 22SZ, Suzhou in 2022; 23SZ, Suzhou in 2023; 23JN, Jining in 2023; 23XX, Xinxiang in 2023 The absolute values ​​of the prediction accuracy of single markers for single environments ranged from 0.01 to 0.31. The selection efficiencies of BLUP phenotypic selection for multiple environments in Jingzhou (2020), Suzhou (2022), Suzhou (2023), Jining (2023), and Xinxiang (2023) were 55%, 19%, 36%, 33%, 8%, and 50%, respectively (Table 3). Although the lodging index of single markers fluctuated across different environments, 15 markers (Gm06_21136539, Gm06_21196104, Gm06_21512079, Gm06_23082606, Gm06_24462733, Gm06_27756888, Gm06_29537534, and Gm06_345669) showed high accuracy. 10. Gm06_34744229, Gm06_38592928, Gm09_46773948, Gm11_25341100, Gm18_48341696, Gm18_54496089, and Gm19_1603096 all improved the selection efficiency of lodging index for different environments, with improvements ranging from 5% to 425% (Table 3). Single-marker predictions for corresponding environments showed low accuracy and large variability; therefore, a combination of 22 markers was used for LASSO regression prediction, achieving a prediction accuracy between 0.28 and 0.51, a significant improvement compared to single-marker predictions (Table 3).

[0035] Table 3 Selection efficiency of single-marker and phenotypic selection in different environments Note: 20JZ, Jingzhou in 2020; 22SZ, Suzhou in 2022; 23SZ, Suzhou in 2023; 23JN, Jining in 2023; 23XX, Xinxiang in 2023. 3.32 Genome Selection In the RRBLUP prediction results, the accuracy of whole-genome markers in predicting lodging indices in different environments and multi-environment BLUP ranged from 0.32 to 0.54. The prediction ability for data from Suzhou in 2023 was the weakest, while the BLUP prediction was the strongest, exceeding the average prediction accuracy of 0.10 for the significant SNP marker set in multiple environments (Table 4). Meanwhile, to measure the effect of the significant SNP marker set as a fixed effect, whole-genome markers were used as random effects, and 22 significant SNPs were used as fixed effects. The average prediction accuracy of whole-genome markers ranged from 0.38 to 0.64, an improvement of 10.6% to 27.0% (Table 4).

[0036] Table 4 shows the average prediction accuracy of models based on different label sets for five environmental lodging indices and multi-environment BLUP. Note: 20JZ, Jingzhou in 2020; 22SZ, Suzhou in 2022; 23SZ, Suzhou in 2023; 23JN, Jining in 2023; 23XX, Xinxiang in 2023. The selection efficiency of different models under different selection ratios was compared. The selection efficiency of different models was significantly higher than that of phenotypic selection (P<0.05) Figure 3 Overall, the selection effect of the significantly associated lodging marker was similar to that of the whole genome marker. The model with the whole genome selection marker as a random effect and the significantly associated lodging marker as a fixed effect had the highest selection efficiency and selection effect on the non-lodging line in different environments, indicating that the model was robust and had good selection effect on the lodging-resistant line (P<0.05) Figure 4

[0037] The above results show that the significantly associated lodging marker set is directly used for genome selection modeling, the model has strong generalization ability, and shows good selection efficiency. Integrating the significantly associated SNP marker set of lodging into the genome prediction model as a fixed effect can significantly improve the prediction accuracy of soybean lodging resistance, and shows high accuracy and practicability in individual selection.

[0038] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.​

Claims

1. A set of SNP molecular markers significantly associated with soybean lodging, characterized in that: The SNP molecular marker set sites are as follows: Gm01_50178213, Gm03_3146453, Gm06_21136539, Gm06_21196104, Gm06_21512079, Gm06_23082606, Gm06_24462733, Gm06_27756888, Gm06_29537534, Gm06_34566910, Gm06_34744229, Gm06_38592928, Gm09_46773948, Gm11_25341100, Gm14_7160350, Gm18_48341696, Gm18_54496089, Gm19_1524582, Gm19_1603096, Gm19_44839696, Gm19_45054985 and Gm19_45924944.

2. The set of SNP molecular markers of claim 1 which are significantly associated with soybean lodging. The SNP molecular marker set sites are as follows: Gm06_21136539, Gm06_21196104, Gm06_21512079, Gm06_23082606, Gm06_24462733, Gm06_27756888, Gm06_29537534, Gm06_34566910, Gm06_34744229, Gm06_38592928, Gm09_46773948, Gm11_25341100, Gm18_48341696, Gm18_54496089 and Gm19_1603096.

3. The set of SNP molecular markers of claim 1 or 2, which are significantly associated with soybean lodging. The SNP molecular marker set belongs to the soybean Wm82.a2.v1 version.

4. Use of the SNP molecular marker set significantly related to lodging of soybean in claim 1 as a template for designing primers or probes.

5. Primers or probes for detecting the SNP molecular marker set in claim 1 or claim 2.

6. Use of the SNP molecular marker set significantly related to lodging of soybean in claim 1 or 2 in multi-marker combination to predict lodging index.

7. Use according to claim 6, characterized in that: The prediction model is constructed by LASSO regression, using the glmnet package, taking the genotype matrix of the SNP molecular marker set in claim 1 or claim 2 as the feature, and the multi-environment lodging index BLUP as the response variable, fitting the model on the training set, determining the optimal regularization parameter (λ) by 5-fold cross-validation, and applying the model to the genotype data of the test set to obtain the estimated breeding value.

8. Use according to claim 6, characterized in that: The whole genome marker data is encoded into a numerical matrix, and the RRBLUP method is called by using the rrblup package in R language to construct a genotype selection model, which is based on a mixed linear model framework, taking the whole genome SNP marker as a random effect and the SNP molecular marker set in claim 1 or claim 2 as a fixed effect, and incorporating them into the model to predict the lodging index. The whole genome marker data is encoded into a numerical matrix, and the RRBLUP method is called by using the rrblup package in R language to construct a genotype selection model, which is based on a mixed linear model framework, taking the whole genome SNP marker as a random effect and the SNP molecular marker set in claim 1 or claim 2 as a fixed effect, and incorporating them into the model to predict the lodging index.

9. Use of the set of SNP molecular markers of claim 1 or 2 that are significantly associated with soybean lodging in the selection of lodging resistant varieties.