SNP (Single Nucleotide Polymorphism) molecular marker for predicting hardness of soybean seeds, primer group and application

Through whole-genome resequencing technology, the SNP marker in the interval of chromosome 18, 55510613kb±50kb was screened, solving the problem of difficult prediction of soybean grain hardness, and achieving efficient screening of high-hardness soybean varieties, improving breeding accuracy and seed hardness.

CN120272645AActive Publication Date: 2025-07-08JILIN ACAD OF AGRI SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510771652.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the hardness of soybean grains, resulting in poor selectivity of grains in soybean variety breeding, affecting seed germination and seedling growth, and thus affecting yield formation.

Method used

The SNP molecular marker located in the interval of 55510613kb±50kb of chromosome 18 was screened through whole genome resequencing technology, and the marker was used for PCR amplification, and the polymorphism (C/T) of the base at 520 bp was judged to predict the hardness of the grain, and primer sets and kit assisted breeding were provided.

Benefits of technology

Accurate prediction of soybean grain hardness is achieved, selectiveness during breeding process is improved, high-hardness soybean varieties are screened out, and seed hardness is increased by 0.95kgf, with an accuracy of 75%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120272645A_ABST
    Figure CN120272645A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of soybean breeding, and provides an SNP (Single Nucleotide Polymorphism) molecular marker for predicting the hardness of soybean seeds, a primer group and application. The marker is located at the 55510613 site of the soybean No.18 chromosome, belongs to a codominant marker, is reliable and convenient to use, and provides great convenience for soybean quality improvement breeding work. The marker is used for carrying out auxiliary selection on a soybean advanced breeding line, and the result shows that when the marker is used for selecting an advanced breeding group, the seed hardness of 75% of groups with the 520th site C in 2022, 2023 and 2024 can be respectively 12.38 kgf, 13.86 kgf and 14.48 kgf, the accuracy can reach 75%, and the seed hardness can be averagely improved to 0.95 kgf. This shows that the marker is practical and effective for assisted selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soybean breeding, and particularly relates to an SNP molecular marker, a primer set for predicting the hardness of soybean grains, and applications thereof. Background Art

[0002] Soybean is an important crop for both oil and food. Cultivating high-yield soybean varieties is the key to the development of the soybean industry. With the development and improvement of molecular biology techniques, gene fine-tuning techniques can achieve the aggregation and efficient utilization of excellent alleles, which are the most crucial techniques for soybean variety breeding and also a necessary means to enhance the breeding ability of soybean varieties in the future. The accumulation of excellent alleles and the development of corresponding molecular markers are the prerequisites for realizing the fine regulation and aggregation of genes.

[0003] The hardness of soybean grains is an important quality trait of soybeans, which affects the quality and processing of natto and the taste quality of vegetable soybeans. At the same time, the hardness of seeds also affects seed water absorption and gas exchange, thereby affecting seed germination and seedling growth, and plays an important role in yield formation. Therefore, exploring molecular markers for excellent alleles of soybean grain hardness can not only provide technical reserves for the aggregation of excellent alleles of soybeans, but also provide clear guidance for improving soybean quality using gene fine-tuning techniques.

[0004] In recent years, with the development of sequencing technology, researchers have a more comprehensive understanding of the soybean genome. Genome-wide association study (GWAS) is an advanced method for studying biological genomes. It is a research method that genotypes large-scale population DNA samples with high-density genetic markers such as SNPs or CNVs across the whole genome to find genotypes related to biological phenotypes. Lam et al. re-sequenced the genomes of 17 wild soybeans and 14 cultivated soybeans and found more than 6.3 million SNPs. Currently, association analysis has been widely used in plant research, such as soybean seed protein content, rice amino acid composition, and aluminum toxicity resistance in triticale. In recent years, developing functional markers for target traits using genome-wide association analysis has become one of the hotspots in molecular biology research. Molecular marker-assisted selection can significantly improve the genetic improvement process of soybean varieties.

[0005] So far, through linkage analysis, quantitative trait locus (QTL) mapping studies have been conducted on genes related to soybean grain hardness. Up to now, more than 10 grain hardness QTLs have been registered in the soybean genome database on all 20 chromosomes. However, the vast majority of QTLs are minor-effect loci and have not been verified. A large part of them are duplicate mappings and cannot effectively predict the hardness of soybean grains. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides an SNP molecular marker, a primer set and an application for predicting the hardness of soybean grains. Through whole-genome resequencing technology, SNP markers significantly linked to target traits are screened, which can be used for molecular markers to predict the hardness of soybean grains, so as to screen soybean varieties with high grain hardness.

[0007] To achieve the above object, the technical solution of the present invention is as follows.

[0008] The first aspect of the present invention provides an SNP molecular marker for predicting the hardness of soybean grains. The nucleotide sequence of the SNP molecular marker is shown in SEQ ID NO: 1. At the 520th bp of the nucleotide sequence shown in SEQ ID NO: 1, there is a C / T polymorphism at this base.

[0009] The SNP molecular marker in the present invention is located in the interval of 55510613 kb ± 50 kb on chromosome 18 of soybean, which is an ideal marker interval for regulating the hardness of soybean grains. The contribution rate of grain hardness at the 55510613th position is 9.10% - 10.34%, and the additive effect is 0.60 kgf - 0.69 kgf. Through the SNP molecular marker in the present invention, it can accurately predict that the grain hardness of soybean varieties with C at the 520th bp is higher than that of soybean grains with T at the 520th bp, indicating that through the molecular marker in the present invention, soybean varieties with high grain hardness can be accurately predicted.

[0010] The second aspect of the present invention provides a primer set for amplifying the SNP molecular marker for predicting the hardness of soybean grains. The primer set includes a forward primer and a reverse primer; the nucleotide sequence of the forward primer is shown in SEQ ID NO: 2; the nucleotide sequence of the reverse primer is shown in SEQ ID NO: 3.

[0011] The third aspect of the present invention provides a kit, and the kit contains the primer set.

[0012] In another preferred embodiment, the kit further includes reagents for PCR amplification.

[0013] The fourth aspect of the present invention provides an application of the SNP molecular marker in predicting the hardness of soybean grains or soybean molecular-assisted screening breeding.

[0014] The fifth aspect of the present invention provides an application of the primer set in predicting the hardness of soybean grains or soybean molecular-assisted screening breeding.

[0015] The sixth aspect of the present invention provides a method for predicting the hardness of soybean grains, including the following steps: Extract the DNA of the sample to be tested; Using the DNA of the sample to be tested as a template, PCR amplification is carried out using the primer set described above to obtain a PCR product; If the 520th bp of the nucleotide sequence shown in SEQ ID NO: 1 is C, it is a high-grain hardness soybean seed variety; If the 520th bp of the nucleotide sequence shown in SEQ ID NO: 1 is T, it is a low-hardness soybean seed variety.

[0016] Compared with the prior art, the present invention has the following beneficial effects.

[0017] The present invention provides an SNP marker, a primer set and an application for predicting the hardness of soybean grains. Through whole-genome resequencing technology, an SNP marker significantly linked to the target trait is screened. The SNP marker is located at position 55510613 of chromosome 18 of soybean. And the high-generation breeding lines of soybeans from 2022 to 2024 were measured. The results show that by using the SNP marker of the present invention for selection in the high-generation breeding population, the hardness of soybeans from 2022 to 2024 can be screened. It is found that the seed hardness of the lines with the 520th position of the SNP marker of the high-generation breeding lines of soybeans being C is higher than 12.38 kgf, 13.86 kgf and 14.48 kgf, and the accuracy can reach 75%. Compared with the lines with the 520th position being T, the average seed hardness is increased by 0.95 kgf. This shows that the SNP marker in the present invention can be used for assisting in the selection of high-hardness soybean grains. Description of the Drawings

[0018] Figure 1 It is the seed hardness distribution map of 350 associated populations from 2022 to 2024; in the figure, A is the seed hardness distribution map in 2022, B is the seed hardness distribution map in 2023, and C is the seed hardness distribution map in 2024.

[0019] Figure 2 It is the population structure diagram of the associated population.

[0020] Figure 3 It is the Manhattan plot of the association analysis results of the soybean grain hardness of 350 natural populations in 2022.

[0021] Figure 4 It is the Manhattan plot of the association analysis results of the soybean grain hardness of 350 natural populations in 2023.

[0022] Figure 5 It is the Manhattan plot of the association analysis results of the soybean grain hardness of 350 natural populations in 2024.

[0023] Figure 6Box plot of phenotypic differences in soybean seed hardness corresponding to molecular markers for 350 natural populations over 3 years; in the figure, (a) is the Box plot of phenotypic differences in soybean seed hardness in 2022, (b) is the Box plot of phenotypic differences in soybean seed hardness in 2023, and (c) is the Box plot of phenotypic differences in soybean seed hardness in 2024. Detailed implementation manners

[0024] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be further described below in conjunction with specific embodiments and the accompanying drawings.

[0025] In the description of the present invention, unless otherwise specified, the reagents used are commercially available, and the methods used are conventional techniques in the art.

[0026] Example 1: Construction of a soybean seed hardness association population and determination of traits.

[0027] In the present invention, 3000 germplasm resources in the soybean germplasm resource bank are used, and their sources cover most of the main soybean production areas at high latitudes. 3000 resources are planted in the field. After the seeds are completely mature, the seeds are harvested, and 20 representative seeds are randomly selected from each variety for measurement. The seed hardness of 3000 resources conforms to a normal distribution within the population, and the genetic diversity index is 3.90. 350 resources are selected from them, and the genetic diversity index of seed hardness is still 3.90. These 350 resources are used as the association population, and the operation steps are as follows.

[0028] (1) According to the average value of the population seed hardness denoted as X and the standard deviation denoted as δ, they are divided into 10 groups. Group 1 is <X - 2δ, group 10 is ≥X + 2δ, and the difference between each intermediate group is 0.5δ. The genetic diversity of each trait is evaluated using Shannon's information index H'. H' = -ΣP i lnP i , where Pi represents the frequency of occurrence of the i-th variation. By calculating, the genetic diversity of the seed hardness of 3000 resources is 3.90.

[0029] (2) Randomly select 350 resources from each group, calculate the genetic diversity index H' of the 350 resources. When it is equal to 3.90, these 350 resources are determined as the association population. As Figure 1 shown, the seed hardness of 350 resources shows a normal distribution.

[0030] Example 2: Genome-wide association analysis of soybean seed hardness.

[0031] (1) Extract the DNA of single plant leaves of 350 resources in the association population by the CTAB method, detect the DNA concentration with a Thermo nanodrop 2000, and detect the DNA purity and integrity by 1wt% agarose electrophoresis.

[0032] (2) The whole-genome sequencing of 350 resources was carried out using the whole-genome resequencing technology of Annoroad Gene Co., Ltd., and the specific operations are as follows.

[0033] Restriction enzyme digestion scheme: The published soybean reference genome was predicted for restriction enzyme digestion using restriction enzyme prediction software. The endonucleases RsaI and HaeIII were selected to digest the genomes of qualified samples, and SLAF fragments with a genomic fragment range of 364 bp - 414 bp were selected.

[0034] Sequencing process: The obtained SLAF fragments were treated with Klenow fragment 3′→5′ exo–, NEB, and deoxyadenosine triphosphate at 37°C for 3′-end A addition, ligated with Dual-index sequencing adapters, then subjected to PCR amplification, purification, sample mixing, and gel cutting to select the target fragments. After the library quality inspection was qualified, sequencing was performed using Illumina HiSeq TM. To evaluate the accuracy of the library construction experiment, the G.max Wm82.a2.v1 genome of Williams 82 soybean was selected as a control for the same treatment to participate in library construction and sequencing. A total of 112G of reads data was obtained, and the average sequencing Q30 was 92.20%.

[0035] The nucleotide sequences of the upstream primer and downstream primer of the amplification primer set for PCR amplification are shown in SEQ ID NO.4 and SEQ ID NO.5.

[0036] SEQ ID NO.4: 5'-AATGATACGGCGACCACCGA-3'.

[0037] SEQ ID NO.5: 5'-CAAGCAGAAGACGGCATACG-3'.

[0038] (3) According to the positioning results of the sequencing Reads on the reference genome, steps such as local realignment by GATK, variant detection by GATK, and variant detection by samtools were carried out, and the intersection variant sites obtained by the two methods of GATK and samtools were taken to ensure the accuracy of the detected SNPs. The intersection of the SNP markers obtained by the two methods was used as the final reliable SNP marker dataset, and a total of 3,306,713 population SNPs were obtained.

[0039] (4) A phylogenetic tree is used to represent the evolutionary relationships between species. According to the degree of relatedness between different organisms, various organisms are placed on a branched tree-like chart to concisely represent the evolutionary history and phylogenetic relationships of organisms. Based on SNPs, a population phylogenetic tree of the samples was constructed using the neighbor-joining algorithm in MEGA5 software.

[0040] (5) Population genetic structure analysis can provide information on the ancestry and composition of individuals and is an important tool for analyzing genetic relationships. Based on SNPs, the population structure of the samples was analyzed using the admixture software. Assuming the number of subpopulations K of the samples was 1-16 respectively, clustering was performed, and the results are as Figure 2 shown. The clustering results were cross-validated, and the optimal number of subpopulations K was determined to be 8 according to the valley value of the cross-validation error rate.

[0041] (6) Based on SNPs, principal component analysis was performed using the TASSEL5 software to obtain the principal component clustering of the samples. Through PCA analysis, it is possible to know which samples are relatively close and which samples are relatively distant, which can assist in the analysis.

[0042] (7) The plink software can be used to estimate the genetic relationship between two individuals in a natural population. The genetic relationship itself is the relative value of the genetic similarity between two specific materials and the genetic similarity between any materials. Therefore, when the genetic relationship value between two materials is less than 0 in the results, it is directly defined as 0.

[0043] (8) Based on the SNP molecular marker data, genetic structure data, Kinship matrix data, and methionine content data of the association population, a genome-wide association analysis was performed using the compressed mixed linear model of the GAPIT software. X is the genotype and Y is the phenotype. Finally, an association result can be obtained for each SNP locus, and the results are as Figures 3 to 5 shown. Using −log10(p)≥7.82 as the screening criterion, one point in the figure represents one SNP locus, and the solid line is the negative logarithm of 0.01 / number of SNPs. Points above the solid line indicate that the corresponding SNP marker is significantly associated with kernel hardness. From the results in the figure, it can be seen that an SNP marker significantly associated with kernel hardness was obtained at position 55510613 on chromosome 18, and the base at this position has C / T polymorphism. The detailed information is shown in Table 1.

[0044] Table 1: Information on SNPs significantly associated with soybean kernel hardness

[0045] Example 3: Application of SNP markers significantly associated with soybean kernel hardness.

[0046] The SNP marker tightly linked to soybean kernel hardness is at position 55510613 on chromosome 18, named qHS18-1, which is a fragment obtained by PCR amplification using the genomic DNA of the material to be identified as a template and the qHS18-1 primer. Among them, the nucleotide sequence of qHS18-1 is as shown in SEQ ID NO. 1.

[0047] SEQ ID NO. 1: CTGATCCTCAAAAGGGCCATTTTCTCCCTTTCCTTTTCTAACTTTTTTATTAGTTTAATCTAGCTATATGTATTGTTGGTGTAAAATATTTTCATTCAAGCATCCAATACATCATATTAATGAATATGATAATGTGGTGATATAGTGAATTCTTATTAGATGTCTGCATAAAAAAGTTCTACATTGGTGCATATAATTTTTTATTTTATCTCTAGAACTTCCTATTTCTTAGGACTTTTTTCTTTACCAAGAATGATAGGAGCACAAAAGGGTTGAACAAGTGAGATATGAGATGAAATGAAAACGTGGTATTATGGTCCTCTTTTGTTTACAACTTGATCATAGTATAAGTATTAGGTAACATATGTCACTTCTCTAGACCTTGCCAACTAATTTTGTGAACCTACACTTTATGGTCTTGCTTTTGTCACCACGTTGAGTCAGAGACATGAGTTTGAAGCAAGCAAGCTAAGTAGGTACTAGGTAGTGAGTTCCTTTCTAATCAAAAGGACACTCCAGCTGCTGCATGCATCTTCATAAAACCTTGATTGCTCCAATCCTTTAACTGCA。

[0048] Among them, the nucleotide sequences of the amplification primers are shown in SEQ ID NO.2 and SEQ ID NO.3.

[0049] SEQ ID NO.2: 5’-CTGATCCTCAAAAGGGCCATTT-3’.

[0050] SEQ ID NO.3: 5’-TGCAGTTAAAGGATTGGAGCA-3’.

[0051] The specific steps for using the above SNP molecular marker to assist in judging the kernel hardness of variety offspring are as follows.

[0052] 1. Extract the genomic DNA of the material to be identified using the CTAB method.

[0053] 1) Take fresh leaves of the soybean to be detected, add liquid nitrogen and grind them into powder, and take 50 mg and put it into a 1.5 mL centrifuge tube.

[0054] 2) Add 0.6 mL of preheated CTAB extraction solution, invert and mix well several times, incubate in a water bath at 65 °C for one hour, mix once every 15 min, and centrifuge at 12000 rpm for 15 min.

[0055] 3) Add 0.6 mL of a mixed solution of chloroform and isoamyl alcohol with a volume ratio of 24:1, invert and mix well 10 times, and centrifuge at 10000 rpm for 15 min.

[0056] 4) Take the supernatant solution and transfer it to another empty centrifuge tube, re-extract it once with a mixed solution of chloroform and isoamyl alcohol with a volume ratio of 24:1, then add 50 μL of ribonuclease at 10 mg / mL, and place it at 23 °C for 30 min.

[0057] 5) Add an equal volume of isopropanol pre-cooled at -20 °C, place it in a -20 °C refrigerator for 30 min, and centrifuge at 5000 rpm for 10 min to remove the supernatant.

[0058] 6) Wash twice with 70 wt% ethanol. After drying, dissolve it with sterilized water to obtain genomic template DNA, and store the genomic template DNA in a 4 °C refrigerator for later use.

[0059] 7) Detect the DNA concentration with 0.8 wt% agarose, and dilute it to the working concentration for PCR amplification.

[0060] 2. Use SNP-labeled primers for PCR amplification to obtain amplification products.

[0061] 1) PCR amplification system: The total volume is 20 μL, including 3 μL of 50 ng genomic template DNA, 10 μL of FastStart Taq enzyme dye mixture, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

[0062] 2) PCR amplification conditions: Pre-denature at 94 °C for 30 s, denature at 94 °C for 30 s, anneal at 57 °C for 30 s, extend at 72 °C for 1 min; cycle 30 times; final extension at 72 °C for 10 min.

[0063] 3. Judge the kernel hardness according to the sequence alignment results.

[0064] The interval of 55510613 kb ± 50 kb on chromosome 18 of the present invention is the marker interval regulating the hardness of soybean seeds. Among them, the contribution rate of the SNP at position 55510613 to the seed hardness is 9.10% - 10.34%, and the additive effect is 0.60 kgf - 0.69 kgf. Using this marker for selection in a high-generation breeding population, it is possible to screen strains with the 520th position being C in 75% of soybean varieties in 2022, 2023, and 2024, and the seed hardness is 12.38 kgf, 13.86 kgf, and 14.48 kgf respectively. The average seed hardness is increased by 0.95 kgf compared to the strain with the 520th position being T. Specifically, as Figure 6 shown, it shows that using the SNP molecular marker in the present invention can greatly reduce the selection cost, improve the accuracy of selecting high-hardness soybean seeds, and provide a new method for seed improvement technology.

[0065] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. Use of an SNP molecular marker in identifying the high or low hardness of soybean seeds, characterized in that, The identification of the high or low hardness of soybean grains is carried out according to the following steps: Extract the DNA of the sample to be tested; Using the DNA of the sample to be tested as a template, perform PCR amplification with the amplification primer set of SNP molecular markers to obtain PCR products; Sequence the PCR products, and after comparing with the SNP molecular markers, judge the hardness of soybean grains; wherein, the nucleotide sequence of the SNP molecular marker is as shown in SEQ ID NO: 1, and at the 520bp of the nucleotide sequence shown in SEQ ID NO: 1, there is a C / T polymorphism at this base; If the base at the 520bp of the nucleotide sequence shown in SEQ ID NO: 1 is C, it is a soybean seed variety with high grain hardness; If the base at the 520bp of the nucleotide sequence shown in SEQ ID NO: 1 is T, it is a soybean seed variety with low hardness; The amplification primer set of the SNP molecular marker includes a forward primer and a reverse primer; The nucleotide sequence of the forward primer is as shown in SEQ ID NO: 2; The nucleotide sequence of the reverse primer is as shown in SEQ ID NO:

3.

2. Use of the SNP molecular marker according to claim 1 in identifying the high or low hardness of soybean grains, characterized in that, The PCR amplification system is: the total volume is 20 μL, including 3 μL of 50 ng genomic template DNA, 10 μL of fast-start Taq enzyme dye mixture, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

3. Use of the SNP molecular marker according to claim 1 in identifying the high or low hardness of soybean grains, characterized in that PCR amplification conditions: pre-denaturation at 94 °C for 30 s, denaturation at 94 °C for 30 s, annealing at 57 °C for 30 s, extension at 72 °C for 1 min; cycle 30 times; final extension at 72 °C for 10 min.

Citation Information

Patent Citations

  • Specific CAPS molecular marker snp545 for identifying hypoallergenic specific soybean variety, method and application thereof

    CN110951907A

  • Creation method of special soybean protein germplasm for processing

    CN113317196A

  • Soybean hundred-grain weight and size related SNP (Single Nucleotide Polymorphism) site, molecular marker, amplification primer and application thereof

    CN116479164A

  • SNP (Single Nucleotide Polymorphism) molecular marker for detecting soybean quality gene and application of SNP molecular marker

    CN120060553A

  • Markers For Aphid Resistant Germplasm In Soybean Plants

    US20090241214A1