SNP molecular markers, primer sets and their applications for predicting soybean seed hardness
Through whole-genome resequencing technology, SNP molecular markers in the interval of chromosome 18, 55510613kb±50kb, solved the problem of predicting soybean grain hardness, achieved efficient screening of soybean varieties with high hardness, and improved breeding effect.
Patent Information
- Application Number
- CN202510771652.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The prior art is difficult to effectively predict the hardness of soybean grains, resulting in poor selectivity of grains in soybean variety breeding, affecting seed germination, seedling growth and yield formation.
Through whole genome resequencing technology, the SNP molecular marker located in the interval of 55510613kb±50kb of chromosome 18 was screened, and the marker was used for PCR amplification, and the C/T polymorphism of the base at 520 bp was judged, and the grain hardness was predicted.
Accurate prediction of soybean grain hardness is achieved, selectiveness during breeding is improved, soybean varieties with high hardness are screened out, and seed hardness is increased by 0.95kgf, with an accuracy of 75%.
Smart Images

Figure CN120272645B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soybean breeding, and in particular to a SNP molecular marker, a primer set and an application thereof for predicting soybean seed hardness. Background Art
[0002] Soybeans are an important dual-purpose grain and oil crop, and cultivating high-yield soybean varieties is key to the development of the soybean industry. With the advancement and improvement of molecular biology techniques, fine-tuning gene regulation has enabled the aggregation and efficient utilization of superior alleles, making it a crucial technology for soybean breeding and a necessary means of improving soybean breeding capabilities in the future. The accumulation of superior alleles and the development of corresponding molecular markers are prerequisites for achieving fine-tuning gene regulation and aggregation.
[0003] Soybean seed hardness is a key soybean quality trait, impacting the quality and processing of natto and the flavor of edible soybeans. Seed hardness also affects water absorption and gas exchange, which in turn influences seed germination and seedling growth, playing a significant role in yield formation. Therefore, discovering molecular markers for superior soybean seed hardness alleles will not only provide technical support for the aggregation of superior soybean alleles, but also provide clear guidance for improving soybean quality using fine-grained gene regulation techniques.
[0004] In recent years, with the development of sequencing technology, researchers have gained a more comprehensive understanding of the soybean genome. Whole-genome association analysis is an advanced method for studying biological genomes. It is a research method that searches for genotypes related to biological phenotypes by typing large-scale population DNA samples with high-density genetic markers such as SNPs or CNVs throughout the genome. Lam et al. resequenced the genomes of 17 wild soybeans and 14 cultivated soybeans and discovered more than 6.3 million SNPs. Currently, association analysis has been widely used in plant research, such as soybean seed protein content, rice amino acid composition, and black wheat aluminum toxicity resistance. In recent years, the use of whole-genome association analysis to develop functional markers for target traits has become one of the hot topics in molecular biology research. Molecular marker-assisted selection can significantly improve the genetic improvement process of soybean varieties.
[0005] To date, quantitative trait loci (QTLs) for soybean kernel hardness have been mapped through linkage analysis. To date, more than 10 QTLs for kernel hardness have been registered in the soybean genome database across all 20 chromosomes. However, the vast majority of these QTLs are micro-effect loci and have not been validated, with a significant portion being duplicated, making them ineffective in predicting kernel hardness. Summary of the Invention
[0006] In order to solve the above technical problems, the present invention provides a SNP molecular marker, primer set and application for predicting soybean seed hardness. Through whole genome resequencing technology, SNP markers are screened to be significantly linked to target traits, which can be used to predict soybean seed hardness by molecular markers, thereby being able to be used to screen soybean varieties with high grain hardness.
[0007] To achieve the above objectives, the technical solutions of the present invention are specifically as follows.
[0008] In a first aspect, the present invention provides a SNP molecular marker for predicting soybean seed hardness. The nucleotide sequence of the SNP molecular marker is shown in SEQ ID NO: 1. The base at the 520 bp position of the nucleotide sequence shown in SEQ ID NO: 1 has a C / T polymorphism.
[0009] The SNP molecular markers in the present invention are located in the interval of 55510613kb±50kb on soybean chromosome 18, which are ideal marker intervals for regulating soybean seed hardness. The contribution rates of grain hardness at positions 55510613 are 9.10%~10.34%, and the additive effects are 0.60kgf~0.69kgf. The SNP molecular markers in the present invention can accurately predict that the grain hardness of soybean varieties with C at the 520bp position is higher than the hardness of soybean grains with T at the 520bp position, indicating that the molecular markers in the present invention can accurately predict soybean varieties with high grain hardness.
[0010] The second aspect of the present invention provides a primer set for amplifying the SNP molecular marker for predicting soybean seed hardness, the primer set comprising a forward primer and a reverse primer; the nucleotide sequence of the forward primer is shown in SEQ ID NO: 2; the nucleotide sequence of the reverse primer is shown in SEQ ID NO: 3.
[0011] The third aspect of the present invention provides a kit, which comprises the primer set.
[0012] In another preferred embodiment, the kit further comprises reagents for PCR amplification.
[0013] A fourth aspect of the present invention provides an application of the SNP molecular marker in soybean seed hardness prediction or soybean molecular-assisted screening and breeding.
[0014] A fifth aspect of the present invention provides an application of the primer set in soybean seed hardness prediction or soybean molecular-assisted screening and breeding.
[0015] A sixth aspect of the present invention provides a method for predicting soybean seed hardness, comprising the following steps:
[0016] Extracting DNA from the sample to be tested;
[0017] Using the DNA of the sample to be tested as a template, PCR amplification is performed using the primer set to obtain a PCR product;
[0018] If the 520th bp of the nucleotide sequence shown in SEQ ID NO: 1 is C, it is a soybean seed variety with high grain hardness;
[0019] If the 520th bp position of the nucleotide sequence shown in SEQ ID NO: 1 is T, it is a soybean seed variety with low hardness.
[0020] Compared with the prior art, the present invention has the following beneficial effects.
[0021] The present invention provides a SNP marker, primer set and application for predicting soybean seed hardness. Through whole genome resequencing technology, the SNP marker was screened to be significantly linked to the target trait. The SNP marker was located at position 55510613 of soybean chromosome 18, and the high-generation breeding lines of soybeans from 2022 to 2024 were measured. The results showed that the SNP marker of the present invention was used for selection in the high-generation breeding population, and the soybean hardness in the three years from 2022 to 2024 could be screened. It was found that the seed hardness of the soybean high-generation breeding line with the 520th position of the SNP marker C was higher than 12.38kgf, 13.86kgf and 14.48kgf, with an accuracy of 75%. Compared with the line with the 520th position of T, the seed hardness was increased by an average of 0.95kgf, indicating that the SNP marker of the present invention can be used to assist in the selection of high-hardness soybean seeds. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is the seed hardness distribution map of 350 associated groups from 2022 to 2024; in the figure, A is the seed hardness distribution map in 2022, B is the seed hardness distribution map in 2023, and C is the seed hardness distribution map in 2024.
[0023] Figure 2 This is a group structure diagram of the associated group.
[0024] Figure 3 Manhattan plot of the association analysis results of soybean seed hardness in 2022 for 350 natural populations.
[0025] Figure 4 Manhattan plot of the association analysis results of soybean seed hardness in 2023 for 350 natural populations.
[0026] Figure 5 Manhattan plot of the association analysis results of soybean kernel hardness in 2024 for 350 natural populations.
[0027] Figure 6 Box plot of phenotypic differences in soybean seed hardness corresponding to molecular markers for 350 natural populations over 3 years; in the figure, (a) is the Box plot of phenotypic differences in soybean seed hardness in 2022, (b) is the Box plot of phenotypic differences in soybean seed hardness in 2023, and (c) is the Box plot of phenotypic differences in soybean seed hardness in 2024. Detailed implementation method
[0028] To enable those skilled in the art to better understand and implement the technical solution of the present invention, the present invention will be further described below in conjunction with specific embodiments and drawings.
[0029] In the description of the present invention, unless otherwise specified, the reagents used are commercially available, and the methods used are conventional techniques in the art.
[0030] Example 1: Construction of a soybean seed hardness association population and determination of traits.
[0031] In the present invention, 3000 germplasm resources in the soybean germplasm resource library are used, and their sources cover most of the main soybean production areas at high latitudes. 3000 resources are planted in the field. After the seeds are completely mature, the seeds are harvested, and 20 representative seeds are randomly selected from each variety for measurement. The seed hardness of 3000 resources conforms to a normal distribution within the population, and the genetic diversity index is 3.90. 350 resources are extracted from them, and the genetic diversity index of seed hardness is still 3.90. These 350 resources are used as the association population, and the operation steps are as follows.
[0032] (1) According to the average value of population seed hardness denoted as X and the standard deviation denoted as δ, it is divided into 10 groups. Group 1 is <X - 2δ, group 10 is ≥X + 2δ, and each intermediate group differs by 0.5δ. The genetic diversity of each trait is evaluated using Shannon's information index H', H' = -ΣP i lnP i , where Pi represents the frequency of occurrence of the i-th variation. By calculating, the genetic diversity of seed hardness of 3000 resources is 3.90.
[0033] (2) Randomly extract 350 resources from each group, calculate the genetic diversity index H' of the 350 resources. When it is equal to 3.90, these 350 resources are determined as the association population. As Figure 1 shown, the seed hardness of 350 resources shows a normal distribution.
[0034] Example 2: Genome-wide association analysis of soybean seed hardness.
[0035] (1) The DNA of 350 individual leaves of the associated population was extracted using the CTAB method. The DNA concentration was detected using a Thermo nanodrop 2000, and the DNA purity and integrity were detected using 1wt% agarose gel electrophoresis.
[0036] (2) The whole genome resequencing technology of Annoroad Gene Co., Ltd. was used to perform whole genome sequencing on 350 resources. The specific operations are as follows.
[0037] Enzyme digestion plan: Enzyme digestion prediction software was used to predict the enzyme digestion of the published soybean reference genome. The endonucleases RsaI and HaeIII were selected to digest the genomes of each qualified sample, and SLAF fragments with a genomic fragment range of 364bp~414bp were selected.
[0038] Sequencing process: The resulting SLAF fragments were treated with Klenow fragment 3′→5′ exo-, NEB, and deoxyadenosine triphosphate at 37°C for 3′-end A addition. Dual-index sequencing adapters were then ligated, followed by PCR amplification, purification, pooling, and gel excision to select target fragments. After the library passed quality control, it was sequenced using an Illumina HiSeq™. To assess the accuracy of the library construction, the genome of Williams 82 soybean strain G. max Wm82.a2.v1 was used as a control and subjected to the same treatment for library construction and sequencing. A total of 112 Gb of reads were obtained, with an average sequencing Q30 of 92.20%.
[0039] The nucleotide sequences of the upstream primer and the downstream primer of the amplification primer set for PCR amplification are shown in SEQ ID NO. 4 and SEQ ID NO. 5.
[0040] SEQ ID NO. 4: 5'-AATGATACGGCGACCACCGA-3'.
[0041] SEQ ID NO. 5: 5'-CAAGCAGAAGACGGCATACG-3'.
[0042] (3) Based on the positioning results of the sequencing reads on the reference genome, GATK performed local realignment, GATK variant detection, samtools variant detection, and the intersection of variant sites obtained by GATK and samtools was taken to ensure the accuracy of the detected SNPs. The intersection of SNP markers obtained by the two methods was used as the final reliable SNP marker dataset, and a total of 3,306,713 population SNPs were obtained.
[0043] (4) Phylogenetic trees are used to represent the evolutionary relationships between species. Based on the closeness of the relationship between the organisms, the organisms are placed on a branching tree diagram to concisely represent the evolutionary process and kinship of the organisms. Based on SNPs, the population evolutionary tree of the sample was constructed using the MEGA5 software and the neighbor-joining algorithm.
[0044] (5) Population genetic structure analysis can provide information about the origin and composition of individual ancestry and is an important tool for genetic relationship analysis. Based on SNPs, the population structure of the sample was analyzed using admixture software. The clustering was performed by assuming that the number of clusters K of the sample was 1 to 16. The results are as follows: Figure 2 The clustering results were cross-validated, and the optimal number of clusters K was determined to be 8 based on the valley value of the cross-validation error rate.
[0045] (6) Based on SNPs, principal component analysis was performed using TASSEL5 software to obtain the principal component clustering of the samples. PCA analysis can help determine which samples are relatively close and which samples are relatively distant, which can assist in analysis.
[0046] (7) Plink software can be used to estimate the kinship between two individuals in a natural population. Kinship itself is defined as the relative value of the genetic similarity between two specific materials and the genetic similarity between any materials. Therefore, when the kinship value between two materials is less than 0, it is directly defined as 0.
[0047] (8) Based on the SNP molecular marker data, genetic structure data, Kinship matrix data and methionine content data of the associated population, the compressed mixed linear model of GAPIT software was used to perform genome-wide association analysis. X is the genotype and Y is the phenotype. Finally, each SNP site can get an association result. The results are as follows: Figures 3 to 5 As shown in the figure, with −log10(p) ≥ 7.82 as the screening criterion, each point in the figure represents a SNP site, and the solid line is the negative logarithm of 0.01 / SNP number. Points above the solid line indicate that the corresponding SNP marker is significantly associated with kernel hardness. The results in the figure show that a SNP marker significantly associated with kernel hardness was obtained at position 55510613 of chromosome 18, and the base there has a C / T polymorphism. Detailed information is shown in Table 1.
[0048] Table 1: SNP information significantly associated with soybean seed hardness
[0049]
[0050] Example 3: Application of SNP markers significantly associated with soybean seed hardness.
[0051] A SNP marker closely linked to soybean seed hardness is located at position 55510613 on chromosome 18 and is designated qHS18-1. The fragment was amplified by PCR using qHS18-1 primers using genomic DNA from the target material as a template. The nucleotide sequence of qHS18-1 is shown in SEQ ID NO. 1.
[0052] SEQ ID NO. 1: .
[0053] The nucleotide sequences of the amplification primers are shown in SEQ ID NO.2 and SEQ ID NO.3.
[0054] SEQ ID NO. 2: 5'-CTGATCCTCAAAAGGGCCATTT-3'.
[0055] SEQ ID NO. 3: 5'-TGCAGTTAAAGGATTGGAGCA-3'.
[0056] The specific steps for using the above-mentioned SNP molecular markers to assist in determining the grain hardness of the offspring of the variety are as follows.
[0057] 1. Use the CTAB method to extract the genomic DNA of the material to be identified.
[0058] 1) Take fresh soybean leaves to be tested, add liquid nitrogen and grind into powder. Take 50 mg and place in a 1.5 mL centrifuge tube.
[0059] 2) Add 0.6 mL of preheated CTAB extract, mix by inversion several times, incubate in a 65°C water bath for one hour, mixing every 15 minutes, and centrifuge at 12,000 rpm for 15 minutes.
[0060] 3) Add 0.6 mL of a 24:1 volume ratio of chloroform and isoamyl alcohol, mix thoroughly by inverting 10 times, and centrifuge at 10,000 rpm for 15 minutes.
[0061] 4) Transfer the supernatant to another empty centrifuge tube and re-extract with a mixture of chloroform and isoamyl alcohol in a volume ratio of 24:1. Then add 50 μL of 10 mg / mL ribonuclease and incubate at 23°C for 30 min.
[0062] 5) Add an equal volume of -20℃ pre-cooled isopropanol, freeze at -20℃ for 30 minutes, centrifuge at 5000 rpm for 10 minutes, and remove the supernatant.
[0063] 6) Wash twice with 70wt% ethanol. After drying, dissolve in sterile water to obtain genomic template DNA. Store the genomic template DNA in a 4°C refrigerator until ready for use.
[0064] 7) Use 0.8 wt% agarose to check the DNA concentration and dilute to the working concentration for PCR amplification.
[0065] 2. Use SNP marker primers to perform PCR amplification to obtain the amplified product.
[0066] 1) PCR amplification system: A total volume of 20 μL includes 3 μL of 50 ng genomic template DNA, 10 μL of Fast Start Taq enzyme-dye mixture, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.
[0067] 2) PCR amplification conditions: 30 cycles of initial denaturation at 94°C for 30 s, denaturation at 94°C for 30 s, annealing at 57°C for 30 s, and extension at 72°C for 1 min; final extension at 72°C for 10 min.
[0068] 3. Determine the hardness of the grain based on the sequence comparison results.
[0069] The 55510613kb±50kb interval of chromosome 18 of the present invention is a marker interval for regulating soybean seed hardness, among which the seed hardness contribution rate of SNP at position 55510613 is 9.10%~10.34%, and the additive effect is 0.60kgf~0.69kgf. Using this marker to select in high-generation breeding populations, 75% of soybean varieties in 2022, 2023 and 2024 can be screened. The seed hardness of the line with position 520 as C is 12.38kgf, 13.86kgf and 14.48kgf respectively. The average seed hardness is 0.95kgf higher than that of the line with position 520 as T. Figure 6 As shown, it is shown that the use of the SNP molecular markers in the present invention can greatly reduce the selection cost, improve the accuracy of selecting high-hardness soybean seeds, and provide a new approach for seed improvement technology.
[0070] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A use of a SNP molecular marker for screening soybean varieties with high grain hardness, characterized in that: The screening of soybean varieties with high grain hardness is carried out according to the following steps: Extracting DNA from the sample to be tested; Using the DNA of the sample to be tested as a template, PCR amplification is performed using a SNP molecular marker amplification primer set to obtain a PCR product; The PCR product is sequenced and compared with a SNP molecular marker to determine soybean seed hardness; wherein the nucleotide sequence of the SNP molecular marker is as shown in SEQ ID NO: 1, and the base at 520 bp of the nucleotide sequence shown in SEQ ID NO: 1 has a C / T polymorphism; The soybean variety with C at the 520th bp position of the nucleotide sequence shown in SEQ ID NO: 1 has a higher seed hardness than the soybean variety with T at the 520th bp position; The amplification primer set of the SNP molecular marker includes a forward primer and a reverse primer; The nucleotide sequence of the forward primer is shown in SEQ ID NO: 2; The nucleotide sequence of the reverse primer is shown in SEQ ID NO:
3.
2. Use of the SNP molecular marker according to claim 1 for screening soybean varieties with high grain hardness, characterized in that: The PCR amplification system was as follows: a total volume of 20 μL, including 3 μL of 50 ng genomic template DNA, 10 μL of Quick Start Taq enzyme dye mixture, 2 μL of each 10 pmol primer and 3 μL of ddH2O.
3. Use of the SNP molecular marker according to claim 1 for screening soybean varieties with high grain hardness, characterized in that: PCR amplification conditions: pre-denaturation at 94°C for 30 s, denaturation at 94°C for 30 s, annealing at 57°C for 30 s, extension at 72°C for 1 min; 30 cycles; final extension at 72°C for 10 min.
Citation Information
Patent Citations
Specific CAPS molecular marker snp545 for identifying hypoallergenic specific soybean variety, method and application thereof
CN110951907A
Soybean hundred-grain weight and size related SNP (Single Nucleotide Polymorphism) site, molecular marker, amplification primer and application thereof
CN116479164A