Single nucleotide mutation site snp, kasp marker significantly associated with soybean hundred-grain weight and application thereof
By developing SNP sites and KASP markers related to soybean 100-seed weight and utilizing competitive allele-specific PCR technology, early molecular-assisted selection of soybean 100-seed weight trait was achieved, solving the problems of time-consuming and low-precision traditional breeding methods and improving breeding efficiency and economic benefits.
Patent Information
- Application Number
- CN202510255061.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Traditional soybean breeding, which selects based on the 100-seed weight of offspring plants, is time-consuming, labor-intensive, and easily affected by external conditions, resulting in low selection accuracy and difficulty in improving breeding efficiency.
We developed single nucleotide mutation sites (SNPs) that are significantly associated with the 100-seed weight of soybeans, and designed KASP markers based on these SNPs. Using competitive allele-specific PCR and fluorescence detection techniques, we amplified soybean DNA by PCR with specific primers to achieve early molecular-assisted selection.
It enables early and precise screening of the 100-seed weight trait in soybeans, reduces the burden of breeding work, improves breeding efficiency, significantly accelerates the breeding process, and has economic benefits.
Smart Images

Figure CN120060539B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular genetics breeding and provides a single nucleotide mutation site (SNP) and KASP marker that are significantly associated with soybean 100-seed weight and their application. These markers can be used for early molecular-assisted selection of soybean 100-seed weight traits to improve breeding efficiency. Background Technology
[0002] Soybeans, as a major source of high-quality protein and edible oil, are crucial to global food security. They are also an important industrial raw material. Although China is the origin of soybeans and a major consumer market, its yield per unit area is lower than the world's advanced levels. Increasing soybean production will not only enhance self-sufficiency and reduce import dependence, but also help ensure national food security and promote sustainable agricultural development.
[0003] Traditional soybean breeding typically relies on individual plant selection based on the 100-seed weight of offspring. This method is not only time-consuming and labor-intensive but also easily affected by external conditions, reducing selection accuracy. To improve the selection efficiency based on 100-seed weight, constructing specific molecular markers to assist selection by leveraging base differences in target genes is considered the most effective strategy. These molecular markers have demonstrated multiple advantages in crop genetic improvement, including early screening, independence from environmental factors, precision, speed, and efficiency, and have become an important technical tool. In particular, Kompetitive Allele-Specific PCR (KASP) is a novel single nucleotide polymorphism (SNP) genotyping method based on allele-specific amplification technology and highly sensitive fluorescence detection. The core of this technology lies in designing two forward primers and one universal reverse primer targeting specific allele SNP locations. Each forward primer contains a specific sequence that can be linked to a specific fluorescent marker. By using these forward primers with specific fluorescent labels in conjunction with universal reverse primers, PCR amplification is performed on the sample DNA, and the variation of alleles can be presented through different fluorescent signals.
[0004] Soybean 100-seed weight is a complex trait regulated by multiple quantitative trait loci (QTLs), involving the combined effects of factors such as the number of seeds per plant, 100-seed weight, number of pods per plant, and number of nodes. This trait exhibits high genetic stability but is also significantly influenced by external environmental factors. In recent years, research on the association between quantitative traits such as the number of seeds per plant, 100-seed weight, number of pods per plant, and number of nodes per plant and yield has become a hot topic in scientific research both domestically and internationally. Studies have shown that soybean yield is closely related to the number of pods per plant, the number of seeds per pod, and the weight of seeds per plant. Generally, higher yields can be achieved when there are high yields, a large number of seeds per plant, a large number of pods per plant, and a moderate 100-seed weight. These QTLs are mainly identified through linkage analysis or genome-wide association studies (GWAS). As a highly efficient gene mapping technique, GWAS can quickly and accurately identify SNP loci significantly associated with soybean 100-seed weight. Therefore, by utilizing SNPs significantly correlated with soybean 100-seed weight and developing closely related KASP markers for early selection in the breeding process (i.e., in the early generations), the burden of breeding work can be effectively reduced, the breeding process accelerated, and significant economic benefits brought about. Discovering SNPs significantly correlated with soybean 100-seed weight and developing corresponding KASP molecular markers for early molecular selection in breeding is of paramount importance for improving breeding efficiency. Summary of the Invention
[0005] The purpose of this invention is to identify single nucleotide mutation sites (SNPs) that are significantly associated with the 100-seed weight of soybeans, and to develop KASP molecular markers and their primer pairs based on the SNP information, so as to provide a molecular-assisted selection technique for early identification and screening of this trait.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] As a first aspect, a SNP molecular marker associated with the 100-seed weight trait of soybean is provided. The SNP S15_40068386, which is significantly associated with the 100-seed weight of soybean, is located at position 40068386bp on chromosome 15 of the soybean genome (version number Glycine max Wm82.a2.v1). It has undergone a C to T substitution. The phenotype corresponding to the CC genotype is a soybean variety with a large 100-seed weight, and the phenotype corresponding to the TT genotype is a soybean variety with a small 100-seed weight.
[0008] Secondly, the application of the aforementioned SNP molecular markers related to the soybean 100-seed weight trait in the identification of the soybean 100-seed weight trait is provided.
[0009] The specific application is as follows: detect the base type of the SNP molecular marker related to the 100-seed weight trait of soybean in the soybean sample to be tested. If the detection result shows that the base type is T, it is determined that the soybean variety has a small 100-seed weight; if the detection result is C, it is determined that the soybean variety has a large 100-seed weight.
[0010] As a third aspect, the application of the aforementioned SNP molecular markers related to the soybean 100-seed weight trait in genetic breeding is provided.
[0011] As a fourth aspect, a KASP-specific primer for the SNP molecular marker associated with the 100-seed weight trait of soybean is provided, consisting of three primers, including two specific primers designed for base differences at the target site, namely upstream primer F1 (SEQ ID NO.1) and upstream primer F2 (SEQ ID NO.2), and a universal downstream primer R (SEQ ID NO.3). The 3' ends of these two specific primers correspond to the variant bases of the alleles, while the 5' ends are connected to specific FAM and HEX fluorescent tag sequences required for the KASP reaction provided by Chengdu Hanchen Guangyi Biotechnology Co., Ltd.
[0012] KASP tags the upstream primer F1 sequence as follows:
[0013] 5'-GAAGGTGACCAAGTTCATGCTAAAATGCAAGCAGAACCAAACCAC-3' (SEQ ID NO. 1);
[0014] KASP tags the upstream primer F2 sequence as follows:
[0015] 5'-GAAGGTCGGAGTCAACGGATTAAAATGCAAGCAGAACCAAACCAT-3' (SEQ ID NO. 2);
[0016] The KASP-tagged downstream primer R sequence is as follows:
[0017] 5'-AACAGTGACTCAAACCAAACCTTG-3' (SEQ ID NO. 3).
[0018] In the process of synthesizing the above KASP molecular marker primers, carboxyfluorescein FAM was added to the 5' end of the forward primer F1 as a fluorescent signal marker (first 21 positions of the primer sequence); while hexachlorofluorescein aminophosphate HEX was added to the 5' end of the forward primer F2 as a fluorescent signal marker (also shown as the first 21 positions of the primer sequence).
[0019] Fifthly, the application of the aforementioned KASP-specific primers in identifying the 100-seed weight trait in soybean is provided. Specifically, a method using the developed KASP molecular markers for identification or screening is employed to target SNP loci significantly associated with soybean 100-seed weight. This involves detecting the deoxyribonucleotide genotype at position 40,068,386 bp on soybean chromosome 15 to determine whether it is TT or CC. The TT genotype corresponds to a lower 100-seed weight, while the CC genotype corresponds to a higher 100-seed weight. This marker-assisted technique contributes to the genetic improvement of the soybean 100-seed weight trait.
[0020] In the above method, the KASP primer set consists of upstream primer F1, upstream primer F2, and downstream primer R. A 384-well microplate PCR reaction system was constructed using a Matrix Arrayer 3250 reaction plate preparer, and the PCR amplification process was performed using a MatrixCycler 2010 high-throughput water bath thermal cycler. After amplification, fluorescence signals were scanned using a Matrix Scanner 2100 high-speed fluorescence scanner, followed by genotyping analysis using the accompanying Matrix Master software. If the initial genotyping results were unsatisfactory, additional amplification steps were performed, with the genotyping effect checked after every 5 cycles until the standard for complete genotyping was achieved.
[0021] The specific operating steps are as follows:
[0022] (1) Extraction of genomic DNA from soybean plants;
[0023] (2) The genomic DNA of the biological sample was amplified by PCR using the PCR-specific amplification primers described above to obtain the amplification product fragment. The PCR amplification product fragment was then subjected to KASP genotyping detection. If the detection result showed that the base type was T, the soybean variety was determined to have a smaller 100-seed weight; if the detection result was C, the soybean variety was determined to have a larger 100-seed weight.
[0024] The molecular marker primers were uniformly added to the same PCR reaction system, and three blank controls were set up using ultrapure water instead of sample template DNA. Subsequently, the DNA of soybean germplasm resources was amplified using the Gene Matrix high-throughput genotyping system.
[0025] 2 μl reaction system: Soybean sample DNA template, 5 ng / μl, 1 μl; 2x Master Mix for ASPCR V1, 1 μl; KASPAssay Mix, F1:F2:R = 1:1:3, 0.02 μl. Reaction conditions included: 95℃ pre-denaturation for 10 min; 95℃ denaturation for 20 sec, annealing at 61–55℃ for 40 sec, decreasing by 0.6℃ per cycle, for 10 cycles; 95℃ denaturation for 20 sec, annealing at 55℃ for 40 sec, for 30 cycles.
[0026] After the reaction, fluorescence scanning was performed on a Matrix Scanner 2100 high-speed fluorescence scanner, and genotyping analysis was performed using the accompanying Matrix Master software. The molecular marker primers clearly distinguished between two genotypes: red dots near the X-axis represented individuals carrying the C allele variant, with genotype CC; while blue dots near the Y-axis represented individuals carrying the T allele variant, with genotype TT.
[0027] (3) Select the desired soybean single plant or line with 100 grains weight based on the genotype in different segregating generations.
[0028] The beneficial effects of this invention are as follows:
[0029] (1) The SNP significantly associated with 100-seed weight in soybean, S15_40068386, was derived from whole-genome resequencing information and related GWAS loci of 270 cultivated soybean varieties. The SNP locus S15_40068386, significantly associated with 100-seed weight in soybean, was detected at position 40068386 bp on chromosome 15 of the soybean genome (version number Glycine max Wm82.a2.v1). Based on the desired selection of CC or TT genotype soybean low-generation breeding materials, this provides technical support for marker-assisted breeding of the 100-seed weight trait in soybean.
[0030] (2) This invention identifies a SNP locus located on soybean chromosome 15 that affects the 100-seed weight of soybean, and develops a KASP molecular marker based on this locus. This marker can directly and specifically identify and detect the C or T base at the SNP locus. This KASP molecular marker has significant application potential and can be used for early selection of soybean 100-seed weight and for molecular-assisted breeding.
[0031] (3) Using KASP molecular marker primers, 185 soybean samples were amplified and genotyped on the Gene Matrix high-throughput genotyping system. The results showed that the molecular marker primers could clearly distinguish between two genotypes: red dots near the X-axis represented individuals carrying the C allele variant (genotype CC), totaling 157 samples with an average 100-seed weight of 20.05 g; blue dots near the Y-axis represented individuals carrying the T allele variant (genotype TT), totaling 26 samples with an average 100-seed weight of 18.56 g. Dots near the origin of the X and Y axes represented blank controls; red crosses marked samples that failed to be genotyped, totaling 2 samples. Statistical analysis showed a significant difference in the 100-seed weight between the two genotypes of soybean samples, demonstrating the practical application value of the markers described in this invention. Attached Figure Description
[0032] Figure 1 These are Manhattan and QQ-plots showing the association analysis results between genotypes and 100-seed weight phenotypes at 207 SNP loci in soybean; where (a) is the Manhattan plot of the association analysis results and (b) is the QQ-plot of the association analysis results.
[0033] Figure 2 This is a graph showing the genotyping results of different soybean varieties using KASP markers; (a) shows the genotyping results of soybean materials SPBX001-SPBX093 (excluding SPBX052), and (b) shows the genotyping results of soybean materials SPBX097-SPBX189; the black squares near the origin represent blank controls without template DNA; the blue dots near the Y-axis and the red dots near the X-axis represent soybean varieties carrying T allelic variants and soybean varieties carrying C allelic variants, respectively; the red crosses represent materials that were not successfully genotyped. Detailed Implementation
[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods consistent with some aspects of this application as detailed in the appended claims.
[0035] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0036] The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0037] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0038] The present invention will be further described below with reference to the accompanying drawings and embodiments. Unless otherwise specified, the methods used are conventional methods.
[0039] Example 1: Identification of nucleotide mutation sites (SNPs) significantly associated with 100-seed weight in soybeans
[0040] The SNP S15_40068386, which is significantly associated with 100-seed weight of soybean in this invention, was obtained from whole-genome resequencing information and related GWAS loci of 270 cultivated soybean varieties. The method for obtaining it includes the following steps:
[0041] (1) Sample collection and acquisition: 270 core soybean germplasm resources were sown in the experimental field of the Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences in 2018 and subjected to routine field management. Simultaneously, the 100-seed weight data of these germplasm resources were recorded. 100-seed weight determination: 100 well-developed seeds were randomly selected, accurately weighed to 0.01 grams, and converted to the weight at a moisture content of 13%.
[0042] (2) SNP detection: Young leaves were collected during the growth stage of soybean V4, and high-quality soybean genomic DNA was extracted using the CTAB method for genome resequencing. Sequencing yielded a total of 8100 Gb of high-quality clean data, averaging 30 Gb per sample, with a sequencing depth of approximately 30-fold. The sequencing data were aligned to the soybean reference genome (version Glycine maxWm82.a2.v1) using BWA software, and duplicate reads were removed using PICARD software. High-quality SNP data were then obtained using GATK software. Finally, functional annotation of the SNP detection results was performed using ANNORVAR software.
[0043] (3) Genome-wide association analysis: Genome-wide association analysis was performed on the obtained SNP marker sites and the measured 100-grain weight phenotype information. The analysis software was TASSEL, and a mixed linear model was used for the analysis.
[0044] (4) Acquisition of 20K marker loci: To obtain 20K marker loci, genotype deletion rate was required to be less than 20%, heterozygous genotype ratio not exceeding 30%, maximum allele frequency less than 95%, and minimum allele frequency greater than 5% in 270 cultivated soybean samples. GWAS marker loci were selected first, followed by marker loci in gene coding regions, with one marker selected every 25kb; if no markers of the above two types were found in a continuous 75kb region, markers in non-coding regions were selected. Finally, 20,648 SNP marker loci were screened, of which 17,588 markers were located on functional genes, covering 31% of soybean coding genes. The SNP molecular markers were evenly distributed with an average spacing of 46kb.
[0045] (5) Obtaining 207 SNP sites evenly distributed on soybean chromosomes: 61 SNP sites were selected from published literature or patents, and 146 SNP sites were selected from 20K sites. A total of 207 SNP sites can be evenly distributed on soybean chromosomes, with a physical distance of about 5 Mbp between sites.
[0046] (6) Gene-phenotype association analysis: The genotypes of 207 SNP loci were associated with the 100-seed weight phenotype. The analysis software was TASSEL, and a mixed linear model was used for the analysis. The SNP locus S15_40068386, which is significantly associated with soybean 100-seed weight, was detected. It is located at 40068386 bp on chromosome 15 of the soybean genome (version number Glycine max Wm82.a2.v1).
[0047] Example 2: Development of KASP-labeled specific primers
[0048] Using the Primer-BLAST function of NCBI (https: / / www.ncbi.nlm.nih.gov / ), three primers were designed based on the nucleotide sequences before and after the S15_40068386 site: upstream primer F1 (SEQ ID NO.1), upstream primer F2 (SEQ ID NO.2), and downstream primer R (SEQ ID NO.3). F1 and F2 contain FAM and HEX fluorescent linker sequences, respectively (shown in the first 21 positions of the primer sequences), as follows:
[0049] KASP tags the upstream primer F1 sequence as follows:
[0050] 5'-GAAGGTGACCAAGTTCATGCTAAAATGCAAGCAGAACCAAACCAC-3' (SEQ ID NO. 1);
[0051] KASP tags the upstream primer F2 sequence as follows:
[0052] 5'-GAAGGTCGGAGTCAACGGATTAAAATGCAAGCAGAACCAAACCAT-3' (SEQ ID NO. 2);
[0053] KASP-tagged downstream primer R 5'-AACAGTGACTCAAACCAAACCTTG-3' (SEQ ID NO.3).
[0054] Example 3: Genotyping of SNP loci in 185 different soybean varieties and its application
[0055] Genomic DNA was extracted from 185 soybean samples of different varieties. Using the genomic DNA as a template, the DNA from the soybean samples was amplified using KASP-labeled primers on the Gene Matrix high-throughput genotyping system. The amplification system consisted of 2 μl reaction volumes: soybean sample DNA template, 5 ng / μl, 1 μl; 2x Master Mix for ASPCR V, 1 μl; KASPassay Mix, F1:F2:R = 1:1:3, 0.02 μl. Reaction conditions included: 95℃ pre-denaturation for 10 min; 95℃ denaturation for 20 sec, annealing at 61–55℃ for 40 sec, decreasing by 0.6℃ per cycle, for 10 cycles; and 95℃ denaturation for 20 sec, annealing at 55℃ for 40 sec, for 30 cycles.
[0056] After the reaction, fluorescence scanning was performed on a Matrix Scanner 2100 high-speed fluorescence scanner, and genotyping analysis was performed using the accompanying Matrix Master software. The results are as follows: Figure 2 Using KASP molecular marker primers, 185 soybean samples were amplified and genotyped on the GeneMatrix high-throughput genotyping system. The results showed that the molecular marker primers could clearly distinguish between two genotypes: red dots near the X-axis represented individuals carrying the C allele variant (genotype CC), totaling 157 samples with an average 100-seed weight of 20.05 g; blue dots near the Y-axis represented individuals carrying the T allele variant (genotype TT), totaling 26 samples with an average 100-seed weight of 18.56 g (see Tables 1 and 2). Dots near the origin of the X and Y axes represented blank controls (see Tables 1 and 2). Figure 2 The red cross marks indicate samples that failed to be successfully genotyped, totaling 2 samples. Statistical analysis showed a significant difference in the 100-seed weight between the two soybean genotypes (see Table 2), demonstrating the practical application value of the markers described in this invention.
[0057] Table 1. 100-seed weight of 185 different soybean varieties
[0058]
[0059]
[0060]
[0061] Table 2. Comparison of mean values among 185 soybean variety groups
[0062]
[0063] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0064] The above are merely optional embodiments of this disclosure and are not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. The application of a SNP molecular marker associated with the 100-seed weight trait of soybean in the identification of the 100-seed weight trait of soybean, characterized in that, The SNP molecular marker is located at position 40068386 bp on chromosome 15 of the soybean genome in version Glycine max Wm82.a2.v1, with polymorphisms of C or T. The CC genotype corresponds to the soybean trait of high 100-seed weight, while the TT genotype corresponds to the soybean trait of low 100-seed weight.
2. The application according to claim 1, characterized in that, The specific application is as follows: The base type of the SNP molecular marker related to the 100-seed weight trait in the soybean sample to be tested is detected. If the test result shows that the base type is T, the soybean variety is determined to have a smaller 100-seed weight; if the test result is C, the soybean variety is determined to have a larger 100-seed weight.
3. The application of a SNP molecular marker associated with the 100-seed weight trait in the genetic breeding of soybean 100-seed weight, characterized in that, The SNP molecular marker is located at position 40068386 bp on chromosome 15 of the soybean genome in version Glycine max Wm82.a2.v1, with polymorphisms of C or T. The CC genotype corresponds to the soybean trait of high 100-seed weight, while the TT genotype corresponds to the soybean trait of low 100-seed weight.
4. The application of a KASP-specific primer for a soybean 100-seed weight-related SNP molecular marker in identifying the soybean 100-seed weight trait, characterized in that, The primers include: The upstream primer F1 has the following nucleotide sequence as shown in SEQ ID NO.1: 5'- GAAGGTGACCAAGTTCATGCTAAAATGCAAGCAGAACCAAACCAC-3'; The upstream primer F2 has the following nucleotide sequence as shown in SEQ ID NO.2: 5'-GAAGGTCGGAGTCAACGGATTAAAATGCAAGCAGAACCAAACCAT-3'; The downstream primer R has the following nucleotide sequence as shown in SEQ ID NO.3: 5'-AACAGTGACTCAAACCAAACCTTG-3'.
5. The application according to claim 4, characterized in that, Includes the following steps: (1) Extract genomic DNA from the soybean sample to be tested; (2) Using the genomic DNA of the soybean sample to be tested as a template, PCR amplification reaction was performed using the KASP-specific primers for the SNP molecular markers related to the 100-seed weight trait of soybean to obtain the amplified product fragment. (3) KASP genotyping was performed on the PCR amplification product fragment. If the test result showed that the base type was T, the soybean variety was judged to have a smaller 100-seed weight; if the test result was C, the soybean variety was judged to have a larger 100-seed weight.
Citation Information
Patent Citations
Soybean whole genome SNP locus combination, gene chip and application
CN112575116A
Single nucleotide mutation site SNP and KASP markers significantly associated with soybean protein content and application thereof
CN112877467A