Application of SNP molecular markers in identifying soybean seed hardness

By screening SNP markers on soybean chromosome 09 using whole genome resequencing technology and using PCR and sequencing technology to determine grain hardness, the problem of accuracy in predicting soybean grain hardness was solved, and the accuracy of breeding and grain hardness were improved.

CN120519556BActive Publication Date: 2025-10-03YANBIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511013018.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-03
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing technologies for QTL localization of soybean grain hardness have limited genetic background and diversity of genetic materials, resulting in insufficient comprehensiveness and accuracy of QTL localization, making it difficult to achieve accurate prediction of soybean grain hardness.

Method used

Using SNP molecular markers, the SNP marker at position 27001292 on soybean chromosome 09 was screened out through whole genome resequencing technology. PCR amplification and sequencing technology were used to determine grain hardness, and the G/A polymorphism at the 151bp position of the nucleotide sequence was used to determine the grain hardness.

Benefits of technology

It achieves accurate prediction of soybean kernel hardness, improves the accuracy of selection in high-generation breeding populations, enhances kernel hardness, and reduces selection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120519556B_ABST
    Figure CN120519556B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of soybean breeding technology, and in particular to the use of a SNP molecular marker in identifying the hardness of soybean seeds. The marker is located at position 27001292 of soybean chromosome 9 and is a co-dominant marker. It is reliable and easy to use, which provides great convenience for soybean quality improvement breeding. The present invention uses the marker to assist in the selection of soybean high-generation breeding lines. The results show that the marker can be used to select in high-generation breeding populations, and the accuracy can reach 68.21%, 87.04% and 88.89% in 2022, 2023 and 2024. The seed hardness of the line with G at position 151 is higher than 13.5 kgf, which is an average increase of 1.51 kgf compared to the line with A at position 151. This shows that the SNP marker in the present invention can be used to assist in the selection of high-hardness soybean seeds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soybean breeding, and in particular to use of a SNP molecular marker in identifying soybean seed hardness. Background Art

[0002] Soybean seed hardness is a key soybean quality trait, impacting the quality and processing of natto and the flavor of edible soybeans. Seed hardness also affects water absorption and gas exchange, which in turn influences seed germination and seedling growth, playing a significant role in yield formation. Therefore, discovering molecular markers for superior soybean seed hardness alleles will not only provide technical support for the aggregation of superior soybean alleles, but also provide clear guidance for improving soybean quality using fine-grained gene regulation techniques.

[0003] Genome-wide association analysis (GSA) is an advanced method for studying biological genomes. It involves genotyping large population DNA samples with high-density genome-wide genetic markers, such as single nucleotide polymorphisms (SNPs) or convolutional nuclei (CNVs), to identify genotypes associated with biological phenotypes. Lam et al. resequenced the genomes of 17 wild soybeans and 14 cultivated soybeans and discovered over 6.3 million SNPs. To date, quantitative trait loci (QTLs) for soybean kernel hardness have been mapped using linkage analysis. To date, more than 10 QTLs for kernel hardness have been registered in the soybean genome database across all 20 chromosomes. However, the genetic background and diversity of the genetic material used for QTL mapping are often limited, which can affect the comprehensiveness and accuracy of QTL mapping. If the parents of the genetic material have little difference in kernel hardness-related genes or insufficient genetic diversity, some important QTLs may not be detected.

[0004] SNP molecular markers are DNA sequence polymorphisms caused by variations in a single nucleotide in the genome, and they offer greater accuracy than QTLs. Existing technologies primarily use SNP molecular markers to predict soybean 1000-grain weight, plant height, and drought resistance. Therefore, to accurately predict soybean kernel hardness, further research on SNP molecular markers is needed to provide a new approach for predicting soybean hardness. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a use of a SNP molecular marker in identifying the hardness of soybean seeds.

[0006] The technical solutions of the present invention are specifically as follows.

[0007] The present invention provides a use of a SNP molecular marker in identifying soybean seed hardness. The identification of soybean seed hardness is performed according to the following steps:

[0008] Extracting DNA from the sample to be tested;

[0009] Using the DNA of the sample to be tested as a template, PCR amplification is performed using a SNP molecular marker amplification primer set to obtain a PCR product;

[0010] The PCR product is sequenced and compared with a SNP molecular marker to determine the soybean seed hardness; wherein the nucleotide sequence of the SNP molecular marker is as shown in SEQ ID NO: 1, and the base at 151 bp of the nucleotide sequence shown in SEQ ID NO: 1 has a G / A polymorphism;

[0011] If the 151st bp of the nucleotide sequence shown in SEQ ID NO: 1 is G, it is a soybean seed variety with high hardness;

[0012] If the 151st bp of the nucleotide sequence shown in SEQ ID NO: 1 is A, it is a soybean seed variety with low hardness;

[0013] The amplification primer set of the SNP molecular marker includes a forward primer and a reverse primer;

[0014] The nucleotide sequence of the forward primer is shown in SEQ ID NO: 2;

[0015] The nucleotide sequence of the reverse primer is shown in SEQ ID NO: 3.

[0016] The SNP molecular markers in the present invention are located in the interval of 27001292kb±50kb on chromosome 09, which are ideal marker intervals for regulating soybean seed hardness. Among them, the contribution rate of SNP at position 27001292 to grain hardness is 8.10%~12.04%, and the additive effect is 1.36kgf~1.61kgf. The SNP molecular markers in the present invention can predict that the grain hardness of soybean varieties with G at 151bp is higher than that of soybean varieties with A at 151bp.

[0017] In another preferred embodiment, the PCR amplification system is:

[0018] The total volume was 20 μL, including 3 μL of 50 ng genomic template DNA, 10 μL Quick Taq HS DyeMix, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

[0019] In another preferred embodiment, the PCR amplification conditions are: 94°C pre-denaturation for 30 s, 94°C denaturation for 30 s, 57°C annealing for 30 s, 72°C extension for 1 min; 30 cycles; and 72°C final extension for 10 min.

[0020] Compared with the prior art, the present invention has the following beneficial effects.

[0021] The present invention provides the use of SNP molecular markers in identifying the hardness of soybean seeds. Through whole genome resequencing technology, the SNP marker was screened to be significantly linked to the target trait. The SNP marker was located at position 27001292 of soybean chromosome 09, and the high-generation breeding lines of soybeans from 2022 to 2024 were measured. The results showed that by using the SNP markers of the present invention for selection in high-generation breeding populations, the accuracy of screening in 2022, 2023 and 2024 can reach 68.21%, 87.04% and 88.89% respectively. The seed hardness of the line with G at position 151 is higher than 13.5 kgf. Compared with the seeds of the line with G at position 151, the average seed hardness is increased by 1.51 kgf, indicating that the SNP markers of the present invention can be used to assist in the selection of high-hardness soybean seeds. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is the seed hardness distribution map of 350 associated groups from 2022 to 2024; in the figure, A is the seed hardness distribution map in 2022, B is the seed hardness distribution map in 2023, and C is the seed hardness distribution map in 2024.

[0023] Figure 2 This is a group structure diagram of the associated group.

[0024] Figure 3 Manhattan plot of the association analysis results of soybean seed hardness in 2022 for 350 natural populations.

[0025] Figure 4 Manhattan plot of the association analysis results of soybean kernel hardness in 2023 for 350 natural populations.

[0026] Figure 5 Manhattan plot of the association analysis results of soybean kernel hardness in 2024 for 350 natural populations.

[0027] Figure 6 This is a box plot of the soybean seed hardness phenotypic differences corresponding to the molecular markers of soybean seed hardness in 350 natural populations over three years; (a) is the box plot of the soybean seed hardness phenotypic differences in 2022, (b) is the box plot of the soybean seed hardness phenotypic differences in 2023, and (c) is the box plot of the soybean seed hardness phenotypic differences in 2024. DETAILED DESCRIPTION

[0028] To enable those skilled in the art to better understand and implement the technical solution of the present invention, the present invention will be further described below in conjunction with specific embodiments and the accompanying drawings. In the description of the present invention, unless otherwise specified, the reagents used are commercially available, and the methods used are conventional techniques in the art.

[0029] Example 1: Construction of a soybean seed hardness association population and trait determination.

[0030] In the present invention, 3000 soybean germplasm resources north of 40°N are collected and measured. They are all collected, evaluated by the research team of cultivated soybean germplasm resources of the Soybean Research Institute of Jilin Academy of Agricultural Sciences, and stored in the germplasm resource bank of Jilin Academy of Agricultural Sciences.

[0031] In this example, 3000 germplasm resources in the soybean germplasm resource bank are used, and their source areas cover most of the main soybean production areas at high latitudes. 3000 resources are planted in the field. After the seeds are completely mature, the seeds are harvested, and 20 representative seeds are randomly selected from each variety for measurement. The seed hardness of the 3000 resources conforms to a normal distribution within the population, and the genetic diversity index is 3.90. 350 resources are selected from them, and the genetic diversity index of seed hardness is still 3.90. These 350 resources are used as the association population. The operation steps are as follows.

[0032] (1) According to the average value of the population seed hardness denoted as X and the standard deviation denoted as δ, the above germplasm resources are divided into 10 groups. Group 1 is <X - ²δ, and group 10 is ≥X + ²δ, with a difference of 0.5δ between each intermediate group. The genetic diversity of each trait is evaluated using Shannon's information index, i.e., H', and H' = -ΣP ,

[0034] , Figure 1 , i , i , , i ,

[0033] , ,

[0036] , ,

[0035] lnP i where P i represents the frequency of occurrence of the i-th variation. By calculating, the genetic diversity of the seed hardness of 3,000 resources is 3.90.

[0033] (2) Randomly select 35 resources from each group and calculate the genetic diversity index H' of the 350 resources. When it is equal to 3.90, these 350 resources are determined as the association population. As Figure 1 shown, the seed hardness of this population呈正态分布.

[0034] Example 2: Genome-wide association analysis of soybean seed hardness.

[0035] (1) Use the CTAB method to extract the DNA of single plant leaves of 350 resources in the association population, detect the DNA concentration with Thermo nanodrop 2000, and detect the DNA purity and integrity with 1wt% agarose electrophoresis.

[0036] It should be noted that the Chinese phrase "呈正态分布" in the original text is translated as "呈正态分布" in the English translation because the specific English expression for this is not provided in the instructions. If a more accurate English expression is required, it can be adjusted according to the actual situation.(2) The whole genome resequencing technology of Annoroad Gene Co., Ltd. was used to perform whole genome sequencing on 350 resources. The specific operations are as follows.

[0037] Enzyme digestion plan: Enzyme digestion prediction software was used to predict the digestion of the published soybean reference genome. The endonucleases RsaI and HaeIII were selected to digest the genomes of each qualified sample, and SLAF fragments with a genomic fragment range of 364 bp to 414 bp were selected.

[0038] Sequencing process: The resulting SLAF fragments were treated with Klenow fragment (3′→5′ exo–) (NEB) and deoxyadenosine triphosphate at 37°C for 3′-end amplification. Dual-index sequencing adapters were then ligated, followed by PCR amplification, purification, pooling, and gel excision to select target fragments. After the library passed quality control, it was sequenced using an Illumina HiSeq™. To assess the accuracy of the library construction, the genome of Williams 82 soybean strain G. max Wm82.a2.v1 was used as a control and subjected to the same treatment for library construction and sequencing. A total of 112 G reads were obtained, with an average sequencing quality of 92.20%.

[0039] The nucleotide sequences of the upstream primer and the downstream primer of the PCR amplification primer set are shown in SEQ ID NO.4 and SEQ ID NO.5.

[0040] SEQ ID NO. 4: 5'-AATGATACGGCGACCACCGA-3'.

[0041] SEQ ID NO. 5: 5'-CAAGCAGAAGACGGCATACG-3'.

[0042] Based on the mapping of sequencing reads to the reference genome, GATK performed local realignment, GATK variant detection, and samtools variant detection. The intersection of variant sites obtained by GATK and samtools was used to ensure the accuracy of the detected SNPs. The intersection of SNP markers obtained by the two methods was used as the final reliable SNP marker dataset, resulting in a total of 3,306,713 population SNPs.

[0043] (3) Phylogenetic trees are used to represent the evolutionary relationships between species. Based on the closeness of the relationship between the organisms, the organisms are placed on a branching tree diagram to concisely represent the evolutionary process and kinship of the organisms. Based on SNPs, the population evolutionary tree of the sample was constructed using the MEGA5 software and the neighbor-joining algorithm.

[0044] (4) Population genetic structure analysis can provide information about the origin and composition of individual ancestry and is an important tool for genetic relationship analysis. Based on SNPs, the population structure of the sample was analyzed using admixture software. The clustering was performed by assuming that the number of clusters K of the sample was 1 to 16. The results are as follows: Figure 2 The clustering results were cross-validated, and the optimal number of clusters K was determined to be 8 based on the valley value of the cross-validation error rate.

[0045] (5) Based on SNPs, principal component analysis was performed using TASSEL5 software to obtain the principal component clustering of the samples. PCA analysis can be used to determine which samples are relatively close and which samples are relatively distant, assisting in evolutionary analysis.

[0046] (6) Plink software can be used to estimate the kinship between two individuals in a natural population. Kinship itself is defined as the relative value of the genetic similarity between two specific materials and the genetic similarity between any materials. Therefore, when the kinship value between two materials is less than 0, it is directly defined as 0.

[0047] (7) Based on the SNP molecular marker data, genetic structure data, Kinship matrix data and methionine content data of the association population, the compressed mixed linear model of GAPIT software was used to perform genome-wide association analysis. X is the genotype and Y is the phenotype. Finally, each SNP site can get an association result, such as Figures 3 to 5 As shown in the figure, with −log10(p) ≥ 8.51 as the screening criterion, a SNP marker significantly associated with grain hardness was obtained at position 27001292 of chromosome 09. The detailed information is shown in Table 1.

[0048] Table 1 SNP information significantly associated with soybean seed hardness

[0049]

[0050] Example 3: Application of SNP markers significantly associated with soybean seed hardness.

[0051] A SNP marker closely linked to soybean seed hardness is located at position 27001292 on chromosome 09 and is designated qHS09-1. The fragment is amplified by PCR using qHS09-1 primers and genomic DNA from the material to be identified as a template. The nucleotide sequence of qHS09-1 is shown in SEQ ID NO. 1.

[0052] SEQ ID NO.1: AACCTCAAAGCCCCTCTCTGGGAACAAGTCTCAAGGTCAGTTCCTTCTTTTGTTTCCCCTCTTTTCAACTAAAACCCCTTCTATCATTCCTTAATTAGTAATAATAATGTTACACATAAATGTATAATGTATAGTCTGCTGAACAATTTTGTTTTTGAATGAATTAATTAGAGGTTTGTATGTATGTG TATATATTCAGGAAATTATCGGAGCTAGGTTATAATAGGAGTGCGAAGAAGTGCAAGGAGAAATTCGAGAACATTTACAAGTACCATAGGAGAACTAAAGAAGGTCGTTTTGGAAAATCAAACGGTGCAAAAACTTACCGATTTTTCGAGCAATTGGAAGCTTTAGACGGAAACCACTCACTTCTTCCTCCG.

[0053] The nucleotide sequences of the amplification primers are shown in SEQ ID NO.2 and SEQ ID NO.3.

[0054] SEQ ID NO. 2: 5'-AACCTCAAAGCCCCTCTCTG-3'.

[0055] SEQ ID NO. 3: 5'-CGGAGGAAGAAGTGAGTGGT-3'.

[0056] The specific steps for using the above-mentioned SNP molecular markers to assist in determining the grain hardness of the offspring of the variety are as follows.

[0057] 1. Use the CTAB method to extract the genomic DNA of the material to be identified.

[0058] 1) Grind fresh soybean leaves into powder using liquid nitrogen. Place 50 mg of the powder into a 1.5 mL centrifuge tube.

[0059] 2) Add 0.6 mL of preheated CTAB extract, mix by inversion several times, incubate in a 65°C water bath for one hour, mixing every 15 minutes, and centrifuge at 12,000 rpm for 15 minutes.

[0060] 3) Add 0.6 mL of a 24:1 volume ratio of chloroform and isoamyl alcohol, mix thoroughly by inverting 10 times, and centrifuge at 10,000 rpm for 15 min.

[0061] 4) Transfer the supernatant to another empty centrifuge tube and re-extract with a mixture of chloroform and isoamyl alcohol in a volume ratio of 24:1. Then add 50 μL of 10 mg / mL ribonuclease and incubate at 23°C for 30 minutes.

[0062] 5) Add an equal volume of -20℃ pre-cooled isopropanol, freeze at -20℃ for 30 minutes, centrifuge at 5000 rpm for 10 minutes, and remove the supernatant.

[0063] 6) Wash twice with 70wt% ethanol, blow dry, and dissolve in sterile water to obtain genomic template DNA. Store the genomic template DNA in a 4°C refrigerator for later use.

[0064] 7) Use 0.8 wt% agarose to check the DNA concentration and dilute to the working concentration for PCR amplification.

[0065] 2. Use SNP marker primers to perform PCR amplification to obtain the amplified product.

[0066] 1) PCR amplification system: A total volume of 20 μL includes 3 μL of 50 ng genomic template DNA, 10 μL Quick TaqHS DyeMix, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

[0067] 2) PCR amplification conditions: 30 cycles of initial denaturation at 94°C for 30 s, denaturation at 94°C for 30 s, annealing at 57°C for 30 s, and extension at 72°C for 1 min; final extension at 72°C for 10 min.

[0068] 3. Determine the hardness of the grain based on the sequence comparison results.

[0069] The interval of 27001292Kb±50Kb of chromosome 09 of the present invention is an ideal marker interval for regulating soybean seed hardness, as shown in Table 1. The contribution rate of the 27001292 SNP to seed hardness is 8.10%~12.04%, and the additive effect is 1.36kgf~1.61kgf. Using this marker to select in high-generation breeding populations, the accuracy of screening can reach 68.21%, 87.04% and 88.89% in 2022, 2023 and 2024. The seed hardness of the strain with G at position 151 is higher than 13.5kgf. Compared with the strain with A at position 151, the average seed hardness is increased by 1.51kgf. Figure 6 As shown, it is shown that the use of the SNP molecular markers in the present invention can greatly reduce the selection cost, improve the accuracy of selecting high-hardness soybean seeds, and provide a new approach for seed improvement technology.

[0070] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A use of a SNP molecular marker in identifying soybean seed hardness, characterized in that: The identification of soybean seed hardness is carried out according to the following steps: Extracting DNA from the sample to be tested; Using the DNA of the sample to be tested as a template, PCR amplification is performed using a SNP molecular marker amplification primer set to obtain a PCR product; The PCR product is sequenced and compared with a SNP molecular marker to determine the soybean seed hardness; wherein the nucleotide sequence of the SNP molecular marker is as shown in SEQ ID NO: 1, and the base at 151 bp of the nucleotide sequence shown in SEQ ID NO: 1 has a G / A polymorphism; The soybean variety with G at the 151st bp position of the nucleotide sequence shown in SEQ ID NO: 1 has a higher seed hardness than the soybean variety with A at the 151st bp position; The amplification primer set of the SNP molecular marker includes a forward primer and a reverse primer; The nucleotide sequence of the forward primer is shown in SEQ ID NO: 2; The nucleotide sequence of the reverse primer is shown in SEQ ID NO:

3.

2. Use of the SNP molecular marker according to claim 1 in identifying soybean seed hardness, characterized in that: The PCR amplification system is: The total volume was 20 μL, including 3 μL of 50 ng genomic template DNA, 10 μL Quick Taq HS DyeMix, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

3. Use of the SNP molecular marker according to claim 1 in identifying soybean seed hardness, characterized in that: The PCR amplification conditions were as follows: pre-denaturation at 94°C for 30 s, denaturation at 94°C for 30 s, annealing at 57°C for 30 s, and extension at 72°C for 1 min; 30 cycles; and final extension at 72°C for 10 min.

Citation Information

Patent Citations

  • SNP (Single Nucleotide Polymorphism) molecular marker related to soybean seed oil content and application of SNP molecular marker

    CN116926234A

  • KR1016609510000B1