SNP loci, molecular markers, amplification primers and their applications related to soybean 100-grain weight

By locating the SNP site on soybean chromosome 14 and designing corresponding molecular markers and amplification primers, the problem of inaccurate soybean 100-grain relocation in the existing technology was solved, efficient molecular marker-assisted breeding was achieved, and the efficiency of soybean yield and quality improvement was improved.

CN116287421BActive Publication Date: 2025-09-26JILIN GOVERNOR FOUND MODERN AGRI TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310516747.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-09
Publication Date
2025-09-26
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately locate SNP sites related to soybean 100-grain weight, resulting in poor results in molecular marker-assisted breeding and inability to be widely applied to different cultivated soybean populations, affecting soybean yield and quality improvement.

Method used

SNP sites located at 20, 241, and 144 bp on soybean chromosome 14 were developed, with polymorphisms of A or T. Corresponding molecular markers and amplification primers were designed, and SNP markers significantly linked to target traits were screened through whole genome resequencing technology for use in molecular marker-assisted selection breeding.

Benefits of technology

It has achieved precise positioning of soybean 100-grain weight, improved the accuracy and efficiency of breeding, significantly accelerated the process of high yield and quality improvement, reduced selection costs, and improved the efficiency of quality improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116287421B_ABST
    Figure CN116287421B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of soybean molecular biology and molecular breeding technology, and specifically relates to a SNP site, molecular marker, amplification primer and its application related to soybean 100-grain weight. The SNP site is located at 20,241,144bp of soybean chromosome 14, and the polymorphism is A or T. The 100-grain weight of plants with the polymorphism A of the SNP site is higher than that of plants with the polymorphism T. The nucleotide sequence of the molecular marker is shown in SEQ ID NO.1, and the nucleotide sequence of the amplification primer of the molecular marker is shown in SEQ ID NO.2-3. The contribution rate of the 100-grain weight of the SNP site of the present invention is 21.47-22.02%, and the additive effect is 4.16-4.60g. A corresponding molecular marker was developed based on this SNP site. Using this marker, selection was carried out in the core germplasm resource sample group of Northeast China. It was found that 71.47% of the varieties with the 248th position of the molecular marker as A in 2020 and 2022 had a 100-grain weight higher than 17.64g and 19.03g, with an accuracy of 71.47%. This greatly reduced the selection cost, improved the efficiency of quality improvement, and accelerated the improvement process of high yield and excellent traits of soybeans.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of soybean molecular biology and molecular breeding, and particularly relates to SNP sites, molecular markers, amplification primers and applications thereof related to soybean 100-grain weight. Background Art

[0002] Soybeans are the world's most important dual-purpose grain and oil crop, and the development of high-yield soybean varieties is key to the development of my country's soybean industry. From a genetic perspective, soybean yield is an extremely complex, integrated trait, composed of multiple specific quantitative traits. Conventional breeding methods are inefficient and limited in their ability to increase soybean yield. However, with the development and improvement of molecular biology techniques, fine-tuning gene regulation technology can achieve the aggregation and efficient utilization of superior alleles. This is the most critical technology for the selection of breakthrough varieties and a necessary means to enhance soybean breeding capabilities in the future. The accumulation of superior alleles and the development of corresponding molecular markers are prerequisites for achieving fine-tuned gene regulation and aggregation.

[0003] Soybean 100-kernel weight is a key indicator of soybean yield. Therefore, identifying molecular markers for superior alleles for soybean 100-kernel weight not only provides a technical reserve for the aggregation of superior alleles for high-yield soybeans, but also provides clear guidance for improving soybean quality through precise gene regulation technology.

[0004] In recent years, with the development of sequencing technology, researchers have gained a more comprehensive understanding of the soybean genome. Genome-Wide Association Studies (GWAS) is an advanced method for studying biological genomes. It is a research method that searches for genotypes related to biological phenotypes by typing large-scale population DNA samples with high-density genetic markers (such as SNPs or CNVs). Lam et al. resequenced the genomes of 17 wild soybeans and 14 cultivated soybeans and discovered more than 6.3 million SNPs. Currently, association analysis has been widely used in plant research, such as soybean seed protein content, rice amino acid composition, and black wheat aluminum toxicity resistance. In recent years, the use of GWAS to develop functional markers for target traits has become one of the hot topics in molecular biology research. Molecular marker-assisted selection can significantly improve the genetic improvement process of soybean varieties.

[0005] To date, quantitative trait loci (QTLs) mapping of soybean 100-grain weight-related genes has been conducted through linkage analysis. To date, over 304 QTLs for 100-grain weight across all 20 chromosomes have been registered in the soybean genome database. However, the vast majority of QTLs are minor and unverified, with a significant portion being duplicated. Furthermore, genetic maps are constructed using relatively backward, low-density molecular markers such as SSR, AFLP, RFLP, and RAPD, resulting in inaccurate loci. Furthermore, most molecular markers currently implicated in functional loci are derived from recombinant inbred lines or populations constructed from a single parent. When applied to natural populations such as hybrids and landraces, these markers are often unsuitable for molecular marker-assisted breeding and are unable to account for the genetic contribution of the loci.

[0006] Therefore, it is necessary to find a marker that can more precisely locate the interval and be widely applied to screen for SNP (single nucleotide polymorphism) molecular markers associated with soybean seed 100-grain weight in hybrid parents of different cultivated soybean populations, and then apply them to the genetic improvement of soybean yield and quality. Through whole-genome resequencing technology, SNP markers significantly linked to target traits can be screened and used in molecular marker-assisted selection breeding, significantly improving the aggregation of superior alleles in soybean and genetically controlling soybean seed size for quality improvement. Summary of the Invention

[0007] One of the objectives of the present invention is to provide a SNP site associated with soybean 100-grain weight. The SNP site is located at 20,241,144bp on soybean chromosome 14 (the reference genome is G.max Wm82.a2.v1), and the polymorphism is A or T. The 100-grain weight of plants with the polymorphism A at the SNP site is higher than that of plants with the polymorphism T. This site has not been reported in previous studies and related patents.

[0008] A second object of the present invention is to provide a molecular marker containing the SNP site. The nucleotide sequence of the molecular marker is shown in SEQ ID NO.1. The degenerate base W at the 248th bp of the sequence is A or T. This region has not been reported in previous studies and related patents. It is a co-dominant marker, which is reliable and easy to use, and provides great convenience for soybean high-yield and quality improvement breeding.

[0009] The third object of the present invention is to provide amplification primers for the molecular markers, the nucleotide sequences of the amplification primers are shown in SEQ ID NO. 2-3.

[0010] A fourth object of the present invention is to provide the use of the SNP site, the molecular marker or the amplification primer in screening or identifying high-grain-yield soybean varieties.

[0011] The fifth object of the present invention is to provide the application of the SNP site, the molecular marker or the amplification primer in soybean molecular breeding, cultivation of transgenic soybeans, and identification of soybean germplasm resources.

[0012] A sixth object of the present invention is to provide a method for identifying different soybean 100-grain weight lines, comprising the following steps:

[0013] The genomic DNA of the soybean to be tested is used as a template and the amplification primers are used for PCR amplification. If the 248th bp of the amplified product is A, it is a soybean line with high 100-grain weight; if the 248th bp of the amplified product is T, it is a soybean line with low 100-grain weight.

[0014] The present invention has the following beneficial effects:

[0015] The SNP locus of the present invention contributes 21.47-22.02% to 100-grain weight, with an additive effect of 4.16-4.60g. A corresponding molecular marker was developed based on the SNP locus. Using this marker, selection in advanced breeding populations revealed that 71.47% of lines with an A at the 248th bp in 2020 and 2022 had 100-grain weights exceeding 17.64g and 19.03g, respectively, with an accuracy of 71.47%. This significantly reduces selection costs, improves the efficiency of quality improvement, and accelerates the improvement of soybean high yields and excellent traits. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a population structure diagram of an association population based on SNPs, obtained using admixture software. Clustering was performed using a K value of 1 to 16. The clustering results were cross-validated, and the optimal number of clusters was determined to be 8 based on the trough value of the cross-validation error rate.

[0017] Figure 2 The 100-grain weight (100SW) of 350 samples from a linked population over two years (E1: 2020, E2: 2022) is shown on the horizontal axis, while the vertical axis represents the number of individuals sampled. The results indicate that 100-grain weight of soybean seeds has a normal distribution and is a quantitative trait.

[0018] Figure 3 This is the Manhattan plot of the MLM association analysis results of 100-grain weight of soybeans from 350 natural populations over two years. The vertical axis is the negative logarithm of the p-value (-log 10(p)), with the chromosome on the abscissa and each point representing a SNP locus; the dashed line is the negative logarithm of 1 / SNP number. Points above the dashed line indicate that the corresponding SNP markers are significantly correlated with 100-seed weight. The point indicated by the arrow on the dashed line of chromosome 14 corresponds to the 20,241,144 bp SNP.

[0019] Figure 4 It is a box plot of the differences in 100-seed weight of soybeans corresponding to molecular markers for 350 natural populations over 2 years. Specific implementation mode

[0020] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but it should not be construed as a limitation of the present invention. Unless otherwise specified, the technical means used in the following embodiments are conventional means well-known to those skilled in the art. The materials, reagents, etc. used in the following embodiments can be obtained from commercial channels unless otherwise specified.

[0021] In this study, 3000 soybean germplasm resources north of 40°N in China were collected and measured. They were all collected, evaluated by the research team of cultivated soybean germplasm resources of the Soybean Research Institute of Jilin Academy of Agricultural Sciences, and stored in the germplasm resource bank of Jilin Academy of Agricultural Sciences.

[0022] Example 1: Construction of an association population for 100-seed weight of soybean seeds and trait determination

[0023] In this example, 3000 germplasm resources in the soybean germplasm resource bank were used. Their sources cover most of the main soybean production areas at high latitudes in China, including Heilongjiang Province, Jilin Province, Liaoning Province, Inner Mongolia Autonomous Region, Xinjiang Uygur Autonomous Region, etc. 3000 resources were planted in the field. After the seeds were completely mature, the seeds were harvested, and 20 representative seeds were randomly selected from each variety for measurement. The 100-seed weights of the 3000 resources conformed to a normal distribution within the population, and the genetic diversity index was 3.90. 350 resources were selected from them, and the genetic diversity index of 100-seed weight was still 3.90. These 350 resources were used as the association population. The operation steps are as follows:

[0024] They were divided into 10 groups according to the population average 100-seed weight (X) and standard deviation (δ). Group 1 is <X - 2δ, group 10 is ≥X + 2δ, and each intermediate group differs by 0.5δ. The genetic diversity of each trait was evaluated using Shannon's information index (H'), H' = -ΣP i lnP i , P i represents the frequency of occurrence of the i-th variation. By calculating, the genetic diversity of the 100-seed weights of 3000 resources was 3.90. <{\

[0025] In each group, 35 resources were randomly selected and the genetic diversity index H' of the 350 resources was calculated. When it was equal to 3.90, the 350 resources were determined to be an associated group, and the 100-grain weight of the group was normally distributed ( Figure 1 ).

[0026] Example 2: Genome-wide association analysis of soybean 100-grain weight

[0027] (1) DNA from leaves of 350 individual plants of the associated population was extracted using the CTAB method (TaKaRa kit Code No. 9768). DNA concentration was determined using a Thermo nanodrop 2000, and DNA purity and integrity were determined using 1% agarose gel electrophoresis.

[0028] (2) 350 resources were sequenced using the whole genome resequencing technology of Annoroad Gene Co., Ltd. The specific operations are as follows:

[0029] 1) Enzyme digestion scheme: Enzyme digestion prediction software was used to predict the digestion of the published soybean reference genome. The endonucleases RsaI and HaeIII were selected to digest the genomes of each qualified sample, and SLAF fragments with a genomic fragment range of 364-414 bp were selected.

[0030] 2) Sequencing Process: The resulting SLAF fragments were 3′-end-addition treated with Klenow Fragment (3′→5′exo–) (NEB) and dATP at 37°C, ligated with dual-index sequencing adapters, and amplified by PCR (PCR primers: F: 5'-AATGATACGGCGACCACCGA-3'; R: 5'-CAAGCAGAAGACGGCATACG-3'). Purification was performed using Agencourt AMPureXP beads (Beckman Coulter, High Wycombe, UK). The samples were then pooled and gel-cleaved to select the target fragments. After quality control, the library was sequenced using an Illumina HiSeq™. To assess the accuracy of the library construction, soybean ('Williams 82': G.max Wm82.a2.v1) was used as a control and subjected to the same treatment for library construction and sequencing.

[0031] 3) Based on the alignment of the sequencing reads on the reference genome, we performed local realignment using GATK, variant detection using GATK, and variant detection using samtools. The intersection of variant sites obtained by GATK and samtools was used to ensure the accuracy of the detected SNPs. The intersection of SNP markers obtained by the two methods was used as the final reliable SNP marker dataset, resulting in a total of 3,306,713 population SNPs.

[0032] (3) Phylogenetic trees are used to represent the evolutionary relationships between species. Based on the closeness of the relationships between organisms, each organism is placed on a branching tree diagram, which concisely represents the evolutionary history and kinship of the organisms. Based on SNPs, the population evolutionary tree of the sample was constructed using the MEGA5 software and the neighbor-joining algorithm.

[0033] (4) Population genetic structure analysis can provide information about the origin and composition of individual ancestry and is an important tool for genetic relationship analysis. Based on SNPs, the population structure of samples can be analyzed using admixture software ( Figure 2 ), assuming the number of sample clusters (K value) is 1-16, clustering is performed. The clustering results are cross-validated, and the optimal number of clusters is determined to be 8 based on the valley value of the cross-validation error rate.

[0034] (5) Based on SNPs, principal component analysis (PCA) was performed using TASSEL5 software to obtain the principal component clustering of the samples. PCA analysis can determine which samples are relatively close and which samples are relatively distant, which can assist in evolutionary analysis.

[0035] (6) Plink software can be used to estimate the relative kinship between two individuals in a natural population. Kinship itself is defined as the relative value of the genetic similarity between two specific materials and the genetic similarity between any materials. Therefore, when the result shows that the kinship value between two materials is less than 0, it is directly defined as 0.

[0036] (7) Based on the SNP molecular marker data, genetic structure data, Kinship matrix data and 100-grain weight data of the association population, a genome-wide association study (GWAS) was performed using the mixed linear model (MLM) of the GAPIT software. X is the genotype and Y is the phenotype. Finally, an association result can be obtained for each SNP site ( Figure 3 ), with -log10 (p) ≥ 6.58 was used as the screening criterion, and SNP markers (A / T) significantly associated with 100-grain weight were obtained at 20, 241, and 144 bp on chromosome 14. The detailed information is shown in Table 1.

[0037] Table 1 Information on significantly associated SNPs for soybean 100-grain weight (100SW)

[0038]

[0039] Example 3: Application of SNP markers significantly associated with soybean 100-grain weight

[0040] A SNP marker tightly linked to soybean 100-grain weight is located at 20,241,144 bp (A / T) on chromosome 14, designated HSW14-1. The fragment was amplified by PCR using genomic DNA from the material to be identified as a template and primers HSW14-1. The nucleotide sequence of HSW14-1 is shown in SEQ ID NO. 1, in which the degenerate base W at bp 248 is either A or T.

[0041] Among them, the amplification primers are:

[0042] HSW14-1-F: 5'-ACTTTCCAACGGTGCATGAT-3' (SEQ ID NO. 2);

[0043] HSW14-1-R: 5'-CAAACTCCACCTCCTTTGC-3' (SEQ ID NO. 3).

[0044] The specific steps for using the above-mentioned SNP molecular markers to assist in determining the 100-grain weight of the offspring of a variety are as follows:

[0045] (1) Extract genomic DNA of the material to be identified using the CTAB method

[0046] 1) Take fresh soybean leaves, add liquid nitrogen and grind into powder. Take an appropriate amount and place it in a 1.5 mL centrifuge tube.

[0047] 2) Add 0.6 mL of preheated CTAB extract, mix by inversion several times, incubate in a water bath at 65°C for one hour, mix every 15 minutes, and centrifuge at 12,000 rpm for 15 minutes.

[0048] 3) Add 0.6 mL of a 24:1 (v / v) chloroform:isoamyl alcohol solution, mix thoroughly by inverting 5-10 times, and centrifuge at 10,000 rpm for 15 min.

[0049] 4) Transfer the supernatant to another empty centrifuge tube and re-extract with a 24:1 (V / V) chloroform:isoamyl alcohol solution. Then add 50 μL of RNase (10 mg / mL) and incubate at room temperature for 30 min.

[0050] 5) Add an equal volume of -20°C pre-cooled isopropanol, freeze at -20°C for 30 min, and centrifuge at 5000 rpm for 10 min to remove the supernatant.

[0051] 6) Wash twice with 70% ethanol, blow dry, and dissolve in sterile water to obtain genomic template DNA. Store the genomic template DNA in a 4°C refrigerator for later use.

[0052] 7) Use 0.8% agarose to detect the concentration of DNA and dilute to the working concentration for PCR amplification.

[0053] 2. Use SNP marker primers to perform PCR amplification to obtain the amplified product.

[0054] 1) PCR amplification system: The total volume is 20 μL, including 3 μL of 10-50 ng genomic template DNA, 10 μL QuickTaq HS DyeMix, 2 μL of each 10 pmol primer, and 3 μL of ddH2O.

[0055] 2) PCR amplification conditions: 30 cycles of pre-denaturation at 94°C for 30 s, denaturation at 94°C for 30 s, annealing at 57°C for 30 s, and extension at 72°C for 1 min; final extension at 72°C for 10 min.

[0056] 3. Determine the seed weight based on sequence alignment results

[0057] The amplified products were sequenced and analyzed. The average 100-grain weight of the subpopulation of lines with A at position 248 from the 5' end of the amplified product was significantly higher than that of the subpopulation of lines with T at the same position ( Figure 4 In 2020 and 2022, 71.47% of the lines with the 248th position being A had 100-grain weights higher than 17.64g and 19.03g, respectively. This indicates that this marker is effective for assisted selection.

[0058] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0059] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A molecular marker related to soybean 100-grain weight, characterized in that: The nucleotide sequence of the molecular marker is shown in SEQ ID NO. 1, and the degenerate base W at 163 bp of the sequence is A or T.

2. The amplification primer for the molecular marker according to claim 1, characterized in that The nucleotide sequence of the amplification primer is shown in SEQ ID NO. 2-3.

3. Use of the molecular markers described in claim 1 or the amplification primers described in claim 2 in screening or identifying high-grain soybean variety resources.

4. Use of the molecular marker according to claim 1 or the amplification primer according to claim 2 in soybean molecular breeding, cultivation of transgenic soybeans, and identification of soybean germplasm resources.

5. A method for identifying different soybean 100-grain weight lines, characterized in that: The following steps are involved: The genomic DNA of the soybean to be tested is used as a template and PCR amplification is performed using the amplification primers described in claim 2. If the 163rd bp of the amplified product is A, it is a soybean line with high 100-grain weight; if the 163rd bp of the amplified product is T, it is a soybean line with low 100-grain weight.

Citation Information

Patent Citations

  • QTL related to soybean hundred-grain weight, molecular marker, amplification primer and application

    CN114107550A

  • CAPS marker related to soybean hundred-grain weight and application of CAPS marker

    CN115976265A