A haplotype molecular marker related to reproductive growth period of soybean and application thereof
By detecting haplotype molecular markers related to the reproductive growth stage of soybean, especially nucleotide G/A at position 2020 of SEQ ID NO.1, the time-consuming and laborious problem of detecting the reproductive growth stage of soybean in traditional methods has been solved, realizing a rapid and accurate breeding method and improving the accuracy and adaptability of soybean breeding.
Patent Information
- Application Number
- CN202510800259.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Traditional methods are difficult to use quickly and accurately to detect the reproductive growth stage of soybeans, which limits the precision of soybean breeding and the scope of planting.
Develop a detection method based on haplotype molecular markers to identify the reproductive growth period of soybean by detecting the polymorphism or genotype of SNPs, especially nucleotide G/A at position 2020 of SEQ ID NO.1, and use it in breeding to select specific genotypes to extend or shorten the reproductive growth period.
It enables rapid and accurate identification of the reproductive growth period, improves the precision of breeding, and allows for the cultivation of soybean varieties adapted to different photoperiods and regions, thus shortening the breeding cycle.
Smart Images

Figure CN120648838B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a haplotype molecular marker related to the reproductive growth stage of soybean and its application. Background Technology
[0002] Soybeans are an important source of vegetable oil and protein for humans. As a typical short-day crop, soybeans are highly sensitive to changes in photoperiod, thus their flowering and ripening periods are influenced by photoperiod. This characteristic leads to geographical limitations on the cultivation of different varieties, restricting their planting range. Therefore, the response of the growth period to photoperiod has become one of the important indicators for assessing the growth and development status of soybeans. The growth period traits of soybeans include flowering period, reproductive growth period, and ripening period, and these traits are closely related to soybean quality, yield, adaptability, and planting range.
[0003] Currently, the length of the reproductive growth period of soybean varies significantly among different soybean varieties, and this period is closely related to soybean quality, yield, adaptability, and planting range. Therefore, accurate and rapid detection and identification of the soybean reproductive growth period is of great significance. Traditional methods mainly rely on field surveys of the soybean growth period, but this method depends on surveying the entire growth cycle of the material, which is time-consuming and labor-intensive, limiting the application of rapid identification of the soybean reproductive growth period. Therefore, developing a method for identifying the soybean reproductive growth period based on haplotype molecular markers is of great importance. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a haplotype molecular marker related to the reproductive growth stage of soybean and its application.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows.
[0006] Application, wherein the application is P1 or P2;
[0007] The P1 refers to the application of substances that detect SNP polymorphisms or genotypes in the identification or auxiliary identification of the reproductive growth stage of soybeans, or in the preparation of products for the identification or auxiliary identification of the reproductive growth stage of soybeans.
[0008] The P2 refers to the application of substances used to detect SNP polymorphisms or genotypes in soybean breeding or the preparation of soybean breeding products.
[0009] The SNP is the 2020th nucleotide of SEQ ID NO.1 in the sequence listing, which is G or A.
[0010] The product contains the substance for detecting the polymorphism or genotype of soybean genomic SNPs as described in claim 1, and is any one of the following G1)-G3):
[0011] G1) Products that detect single nucleotide polymorphisms or genotypes related to the reproductive growth stage of soybeans;
[0012] G2) Products used for identification or auxiliary identification of the reproductive growth stage of soybeans;
[0013] G3) is a product used in soybean breeding.
[0014] A method for identifying or assisting in the identification of the reproductive growth stage of soybeans, comprising detecting the genotype of the SNP described in claim 1 in the soybean to be tested, and identifying or assisting in the identification of the reproductive growth stage of soybeans based on the genotype of the soybean to be tested: the reproductive growth stage of soybeans with the SNP genotype GG is higher than or candidate higher than the reproductive growth stage of soybeans with the genotype AA, wherein the SNP genotype GG is a homozygous type with nucleotide G at position 2020 of SEQ ID NO.1 in the sequence listing; and the SNP genotype AA is a homozygous type with nucleotide A at position 2020 of SEQ ID NO.1 in the sequence listing.
[0015] The application of the above methods in soybean breeding.
[0016] According to the above applications, products, and methods, the breeding objective includes cultivating or selecting soybeans with a longer or shorter reproductive growth period.
[0017] A method for soybean breeding, the method comprising replacing the 2020th nucleotide of SEQ ID NO.1 in the genome of a soybean with an SNP genotype of AA with G, to obtain a target soybean with a reproductive growth period higher than that of the soybean with the genotype of AA and a haplotype of GG; wherein the SNP genotype of AA is homozygous for the 2020th nucleotide of SEQ ID NO.1 in the sequence listing being A; and the SNP genotype of GG is homozygous for the 2020th nucleotide of SEQ ID NO.1 in the sequence listing being G.
[0018] The beneficial effects of adopting the above technical solution are as follows: By detecting the haplotype molecular marker (nucleotide G / A at position 2020 of SEQ ID NO.1) that is significantly related to the reproductive growth period of soybean, the present invention can quickly and accurately identify the reproductive growth period trait of soybean. This molecular marker provides a clear genetic target for soybean breeding. By selecting specific genotypes (such as GG type which can prolong the reproductive growth period and AA type which can shorten the reproductive growth period), soybean varieties adapted to different photoperiods and regional requirements can be directionally bred, shortening the breeding cycle and improving the breeding accuracy. Attached Figure Description
[0019] Figure 1 This is a frequency distribution map of the reproductive growth phase of a RIL population;
[0020] Figure 2 It is a frequency distribution map of the reproductive growth period in natural populations;
[0021] Figure 3 These are QTL loci associated with reproductive growth traits during the 3-year reproductive growth period in the RIL population;
[0022] Figure 4 These are Manhattan plots and QQ plots of average data from a 3-year reproductive growth period;
[0023] Figure 5 This is an LD block diagram of chromosome 4 during its reproductive growth phase;
[0024] Figure 6 It is a haplotype analysis of the reproductive growth phenotype;
[0025] Figure 7 These are photos of the planting site of RIL (Rich Ink Litter) populations and natural populations. Detailed Implementation
[0026] The following embodiments illustrate the present invention in detail. All raw materials and equipment used in the present invention are commercially available products and can be directly obtained through market purchase. Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods.
[0027] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0028] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0031] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1, Materials and Methods
[0033] 1.1 Test Materials
[0034] This study used a high-generation F10 recombinant inbred line (RIL) population constructed from a cross between Jidou 12 and wild soybean Y9, and a natural population constructed from 290 accessions with a broad genetic base, as experimental materials for studying the flowering and maturity stages of soybean. The RIL population contained 180 families; the parent Jidou 12 was a nationally approved variety bred by the Hebei Academy of Agricultural Sciences, and Y9 was a wild species. The natural population materials came from a wide range of sources, including 16 provinces in China (200, 68.97%), 6 states in the United States (83, 38.62%), and several other countries (7, 2.38%), exhibiting rich genetic diversity. Figure 7 As shown, the RIL population and the natural population were planted at the Dishan Experimental Station in Shijiazhuang City, Hebei Province (38°04′N, 114°28′E) from 2021 to 2023. The experiment was a randomized block design, with two replicates of materials planted in three rows of 3 meters each, with a row spacing of 0.5 meters. The plant spacing for the RIL population was 0.2 meters, and the plant spacing for the natural population was 0.1 meters. Field management followed local conventional management methods.
[0035] 1.2 Field Survey
[0036] First, the flowering and maturity periods of the RIL population and the natural population were investigated. Flowering Days (FD) were recorded at stage R1 (the number of days from sowing to flowering in 50% of the soybeans). Maturity Days (MD) were recorded at stage R8 (95% of the pods in 50% of the plants had turned to mature color and the seeds inside the pods were visibly swaying). Reproductive Days (RD) were recorded as stage R8 minus stage R1 (RD = R8 – R1).
[0037] 1.3 DNA Extraction
[0038] Genome samples were extracted from soybean leaves using the Kangwei DNA kit. The specific method is as follows:
[0039] (1) Pick about 100-500mg of fresh, unfolded leaves from soybean plants and put them into a 2ml centrifuge tube. Cool with liquid nitrogen and then grind quickly with a plastic grinding rod to ensure that the sample is fully pulverized.
[0040] (2) Add 400 μL of Buffer LP1 and 6 μL of RNase A (10 mg / mL) to the sample centrifuge tube, and vortex the sample for 1 minute to ensure thorough mixing. Let stand at room temperature for 20-30 minutes to allow the soybean leaf tissue to fully lyse.
[0041] (3) Add 130 μL of Buffer LP2 to the lysed sample, mix well, and centrifuge at 12,000 rpm for 5 minutes in a low-temperature high-speed centrifuge. After centrifugation, transfer the supernatant to a newly labeled 2 ml centrifuge tube.
[0042] (4) After centrifugation, transfer the supernatant to a new centrifuge tube. Then add 1.5 times the amount of Buffer LP3 (containing anhydrous ethanol) to the original volume and mix thoroughly.
[0043] (5) Pour the mixture into the adsorption column, centrifuge at 12000 rpm for 1 minute, and discard the filtrate.
[0044] (6) Add 500 μL of Buffer GW (containing anhydrous ethanol) reagent to the adsorption column, centrifuge at 12,000 rpm for 1 min, discard the filtrate again, and put the adsorption column back into the centrifuge tube. Repeat this process once.
[0045] (7) Discard the waste liquid in the tube, put the adsorption column back into the collection tube and let it air for 2 minutes to remove residual ethanol components.
[0046] (8) Transfer the adsorption column to a brand new 1.5ml centrifuge tube and let it stand at room temperature for 10-20 minutes to completely dry the adsorption membrane.
[0047] (9) Add 50-100 μL of Buffer GE reagent to the center of the adsorption membrane of the dried adsorption column and let it stand for several minutes.
[0048] (10) The adsorption column was placed in a centrifuge tube and centrifuged at 12,000 rpm for 1 min. The DNA in the tube was retained, and the DNA extraction quality was detected by 0.8% agarose gel electrophoresis. At the same time, the DNA was quantitatively detected by UV spectrophotometer. The qualified DNA was stored at -20℃ for later use.
[0049] 1.4 Genotyping
[0050] The RIL population established by Jidou 12 and Ye 9 was sequenced using GBS. The specific steps are as follows: (1) Place different samples and adapters with different barcodes in pairs on a plate; (2) Perform enzymatic digestion using ApeKI restriction endonuclease; (3) Use T4 ligase to connect the adapters to the sticky ends of the fragments caused by enzyme digestion; (4) Mix the samples with different barcodes in a pool, and then pass them through a fragment length screening column to filter out adapters that have not yet reacted; (5) Add PCR primers and perform PCR amplification.
[0051] 290 natural populations were resequencing using the following steps:
[0052] (1) Resequencing library construction, using The DNALibraryPrep Kit for ILM constructs resequencing libraries from quality-tested DNA using the following steps:
[0053] ① Take 200 ng of quantified and quality-tested DNA and place it in a 0.2 mL PCR tube. Add 4 μL of [unspecified ingredient] to the tube. End Repair Buffer and 2.6μL Add water to 20 μL of End Repair Enzyme and incubate at 37°C for 20 min in an ABI 9700 PCR instrument, followed by denaturation at 72°C for 20 min to complete DNA fragmentation, end repair, and A-tailing.
[0054] ② Add 2μL Ultra DNA Ligase, 8μL Ultra DNA Ligase Buffer and 4 μL Add water to the adapter for ILM to a final volume of 40 μL, and place on an ABI 9700 PCR instrument at 22°C for 60 min to complete the ligation of the sequencing adapter. Then, add 48 μL of GenoPrep DNA Clean Beads to the ligation product to purify it. After purification, use 0.68+0.2 magnetic beads to screen for fragments, retaining ligation products with insert fragments between 300-350 bp.
[0055] ③ Add 10 μL of sequencing adapter with barcode sequence and 10 μL of [unclear text - possibly a typo, should be removed] to the PCR tube from the previous step. PCR MasterMix was prepared and water was added to a final volume of 20 μL. Amplification was performed using an ABI 9700 PCR instrument. The amplification program was as follows: 98℃ pre-denaturation for 2 min, 98℃ denaturation for 30 s, 65℃ annealing for 30 s, 72℃ extension for 40 s, 5 cycles, and 72℃ extension for 4 min.
[0056] ④ Add 20 μL of GenoPrep DNA Clean Beads to the second round of PCR products, place on a magnetic rack until the solution is clear, discard the supernatant and add 100 μL of 80% ethanol to wash the magnetic beads, then add 35 μL of 10 mM Tris-HCl to obtain the purified DNA library.
[0057] (2) After library construction, preliminary quantification was performed using Qubit 2.0, and the effective concentration of the library was accurately quantified using qPCR to ensure library quality. After the library quality test was passed, sequencing was performed using the BGI MGI-2000 / MGI-T7 sequencing platform in PE150 mode.
[0058] (3) The quality control of SNP sites was performed using the vcftools software, and the obtained VCFs were used for genome-wide association study (GWAS).
[0059] 1.5 QTL mapping and genome-wide association analysis
[0060] QTL mapping for the RIL population was first performed using QTLICIMaping V4.2 software to construct a genetic map of the RIL population based on GBS sequencing results. The reference genome was Williams 82.a2.v1, with a total map length of 6,626.06 cM, including 6,288 specific-length amplified fragment markers covering 20 linkage groups (LGs), with an average distance of 1.81 cM between markers. QTL analysis of soybean reproductive growth traits was then conducted using the ICIM-ADD method in QTLICIMaping V4.2 software. QTLs were identified based on a significance threshold of LOD > 2.5. The naming convention for QTLs was: q + FD + chromosome number + QTL number.
[0061] GWAS was performed using TASSEL 2.1 software based on mixed linear models (MLM), combining natural population phenotypic and genotypic data. A significance threshold of -log10(P) ≥ 6 was set for all SNP loci in the GWAS for comparative analysis. Quantile-Quantile scatter plots were generated using ggplot2. Manhattan plots were generated using the qqman package in R. Figure 4 (As shown).
[0062] 1.6 Candidate gene prediction and haplotype analysis
[0063] All genes within a 50kb range before and after common mapping sites were extracted from the Soybase database (https: / / www.soybase.org / sbt / ). GO analysis of these genes was performed using the GO enrichment tool in Soybase. Candidate genes regulating reproductive stage traits were predicted based on soybean genome annotation data. Haploview 4.2 software was used to perform haplotype block analysis on the 50kb candidate regions.
[0064] Example 2, Results and Analysis
[0065] 2.1 Phenotypic analysis during the reproductive growth period
[0066] Phenotypic identification results of the RIL population from 2021 to 2023 showed that the reproductive growth period was between 40 and 65 days. The coefficient of variation for the reproductive growth period was 6.15%, with a skewness of 0.02 and a kurtosis of 0.21. The reproductive growth period of the natural population was between 40 and 90 days, with a coefficient of variation of 9.37%, a skewness of -0.01, and a kurtosis of 0.36. Figure 1 , Figure 2 As shown, the reproductive growth periods of both populations are normally distributed.
[0067] 2.2 Location and analysis of QTLs during the reproductive growth phase of RIL populations
[0068] Based on the flowering and maturity phenotypes of the RIL population, the length of the reproductive growth period was calculated. Using the ICIM method, a total of 14 QTLs associated with the reproductive growth period were detected. These QTLs were distributed on 10 different chromosomes. Figure 3 (Table 1). qRD4-2 on chromosome 4, qRD7-2 on chromosome 7, and qRD16 on chromosome 16 were repeatedly detected in both environments, indicating that they are three stably expressed QTLs, with qRD16 being the previously reported FT5a gene. qRD8 on chromosome 8, qRD13 on chromosome 13, and qRD17 on chromosome 17 all overlapped with or were adjacent to known reproductive-related loci. The LOD values of qRD3, qRD10, qRD13, and qRD17 ranged from 2.52 to 4.08, and their phenotypic contribution rates ranged from 2.58% to 4.07%. These loci were low-efficacy QTLs detected in one year with LOD and PVE (%) values not exceeding 5.
[0069] Table 4. QTL mapping results of reproductive growth traits in RIL populations
[0070]
[0071] 2.3 GWAS during the reproductive growth phase of natural populations
[0072] GWAS analysis of natural populations identified 88 SNP loci with a LOD value ≥ 6 associated with reproductive growth. Of these, 30 SNP loci had an LOD value ≥ 6.50, distributed across 15 chromosomes. Figure 4 (See Table 2). Chr10_45248537 is adjacent to the previously reported E2 gene, at a distance of 121 kb. Four QTNs (Chr06_6230952, Chr09_43772632, Chr15_38546523, and Chr18_24364093) are all located within 96–242 kb of known SNPs controlling soybean growth period. Furthermore, 15 new SNPs were identified that significantly differ from major QTLs or reported QTNs; the molecular mechanisms of these new sites require further investigation.
[0073] Table 2. GWAS results of reproductive growth traits in natural populations.
[0074]
[0075]
[0076] 2.4 Analysis of common loci for traits during the reproductive growth period
[0077] Linkage analysis and GWAS association analysis each have their limitations; the former suffers from low resolution, while the latter may have false positives. Therefore, combining these two methods allows the advantages of each to compensate for the limitations of the other, resulting in more accurate and precise QTL localization. This study compared the QTL effect regions detected by linkage analysis with significant sites associated with GWAS. The results showed that two SNP sites overlapped with stable QTL (qRD4-1, qMD19-2) regions: Chr04_16576138 and Chr19_45573720, located on chromosomes 4 and 19, respectively. Chr19_45573720 is located near a known SNP controlling soybean growth period, while Chr04_16576138 is a previously unreported, co-located site.
[0078] 2.5 Candidate gene prediction and haplotype analysis of co-location sites
[0079] This study performed candidate gene prediction and haplotype analysis on common loci of reproductive period-related traits in the two populations. Figure 5 qRD4-1 (located between 15855447bp and 16974748bp) in the RIL population and Chr04_16576138 identified in the natural population are loci significantly associated with reproductive growth phase. qMD19-2 (located between 45523455bp and 45322411bp) and Chr19_45573720 located in the natural population are loci significantly associated with maturity phase. Since the common loci on chromosome 19 regulating soybean maturity are located near previously identified SNPs controlling soybean growth phase, this study only performed candidate gene prediction and haplotype analysis on the common loci on chromosome 4 associated with reproductive growth phase. Haplotype analysis within a 50 kb range around the Chr04_16576138 site identified three candidate genes: Glyma.04g125400, Glyma.04g125500, and Glyma.04g125600. The encoding products of these candidate genes mainly include GRF zinc finger binding protein, histone lysine N-methyltransferase, histone methyltransferase, and zinc ion binding protein. However, further screening and functional validation of the predicted candidate genes are needed to determine their actual roles in regulating soybean growth stages.
[0080] The predicted full-length Glyma.04G125500 gene is 4444 bp. Further analysis revealed a key mutation in the coding sequence (CDS) of the Glyma.04G125500 gene: a synonymous mutation at position 2020 bp, where a base changed from G to A. These mutations collectively produced two distinct haplotypes: Hap1 and Hap2. Figure 6 ).
[0081] A total of 323 materials were analyzed. Some materials had poor sequencing quality (e.g., genes were not homozygous, the sequencing data had too high heterozygosity or deletion rate, and poor quality materials were deleted according to quality control), and some materials did not amplify effective bands. After removing the problematic materials, a total of 161 Hap1 materials (GG) and 35 Hap2 materials (AA) were finally detected.
[0082] Table 7. Germogenic phenotypic data of Hap1 and Hap2 haplotypes.
[0083]
[0084]
[0085]
[0086] Analysis of 161 Hap1 samples and 35 Hap2 samples, such as Figure 6 As shown in Tables 7 and 8, combined with the phenotypic data, the mean reproductive growth period for haplotype Hap1 materials was 67.74 days, and the mean reproductive growth period for haplotype Hap2 materials was 63.89 days. The reproductive growth period for haplotype Hap1 materials was significantly longer than that for haplotype Hap2 materials, and the difference in reproductive growth period between the two haplotypes reached a significant level (P<0.05).
[0087] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these examples without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0088] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0089] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. Use, characterized in that, The application is P1 or P2; The P1 is the use of a substance for detecting the polymorphism or genotype of SNP in identifying or assisting in identifying the reproductive growth period of soybean or in preparing a product for identifying or assisting in identifying the reproductive growth period of soybean; The P2 is the use of a substance for detecting the polymorphism or genotype of SNP in soybean breeding or in preparing a product for soybean breeding; The purpose of the soybean breeding is to breed or select soybean with longer reproductive growth period or shorter reproductive growth period; The SNP is the 2020th nucleotide of SEQ ID NO. 1 in the sequence listing, which is G or A.
2. A product characterized by, The product contains the substance for detecting the polymorphism or genotype of SNP of soybean genome in claim 1, and is any one of the following G1) to G3): G1) a product for detecting the single nucleotide polymorphism or genotype related to the reproductive growth period of soybean; G2) a product for identifying or assisting in identifying the reproductive growth period of soybean; G3) a product for soybean breeding; The purpose of the soybean breeding is to breed or select soybean with longer reproductive growth period or shorter reproductive growth period.
3. A method for identifying or aiding in the identification of the reproductive growth stage of soybean, characterized in that, By detecting the genotype of the SNP in claim 1 in the soybean to be tested, the reproductive growth period of soybean is identified or assisted in identifying according to the genotype of the soybean to be tested: the reproductive growth period of soybean with the SNP genotype of GG is higher or is a candidate for being higher than that of the soybean with the genotype of AA, the SNP genotype of GG is the homozygous type of the 2020th nucleotide of SEQ ID NO. 1 in the sequence listing being G, and the SNP genotype of AA is the homozygous type of the 2020th nucleotide of SEQ ID NO. 1 in the sequence listing being A.
4. The use of the method in claim 3 in soybean breeding; the purpose of the soybean breeding is to breed or select soybean with longer reproductive growth period or shorter reproductive growth period.
Citation Information
Patent Citations
Method for screening soybean varieties with different growth stages and plant heights and special kit thereof
CN107858445A
Application of soybean biostar gene TOF7 in regulation and control of soybean growth period and yield
CN113249389A