SNP (Single Nucleotide Polymorphism) site related to hundred-grain weight of soybean and application of SNP site
By using the SNP site at position 39283583 bp on chromosome 11 of the soybean genome and a specific KASP marker, the problems of a large QTL interval for 100-seed weight and a small number of functional genes in soybean were solved, achieving efficient molecular marker-assisted breeding and improving the accuracy and efficiency of soybean breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF CEREAL & OIL CROPS HEBEI ACAD OF AGRI & FORESTRY SCI
- Filing Date
- 2026-03-10
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the confidence intervals of QTLs related to 100-seed weight in soybeans are relatively large, and differences in genetic background make it difficult to apply to molecular-assisted breeding. Furthermore, the number of functional genes identified is limited, which affects the efficiency and effectiveness of soybean breeding.
A genotyping method for soybean breeding was developed by combining a specific KASP marker with an SNP locus (containing G and A alleles) at position 39283583 bp on chromosome 11 of the soybean genome. This method was used to screen for high 100-seed weight lines.
The QTL interval was successfully narrowed down to 340 kb, and the Glyma.11G239000 gene, which is highly associated with 100-seed weight, was identified. This provides an efficient molecular marker-assisted breeding tool and improves the accuracy and efficiency of soybean breeding.
Smart Images

Figure CN122012786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and in particular to a SNP site related to the 100-seed weight of soybean and its application. Background Technology
[0002] Soybean is an important economic crop and a major source of plant oil and plant protein for humans. With the continued growth of the global population and the improvement of living standards, global demand for soybeans is steadily increasing. Therefore, increasing soybean yield remains a key objective in soybean breeding. Seed weight is not only a crucial component affecting soybean yield but also significantly influences seed appearance and overall quality. Therefore, identifying quantitative trait loci (QTLs) and functional genes associated with seed weight is essential for the genetic improvement of soybean yield potential.
[0003] In soybean, over 300 QTL loci associated with grain weight have been identified through linkage and genome-wide association studies. However, due to factors such as low marker density in early QTL mapping and population size, the confidence intervals of many QTLs are relatively large. Furthermore, differences in genetic background make it difficult to apply many QTLs to molecular-assisted breeding. Fine mapping is an important approach to narrowing down QTL intervals and identifying functional genes.
[0004] In addition, several functional genes related to 100-seed weight have been identified using GWAS and transcriptome analysis. GmSWEET10a / b controls seed size by regulating sucrose transport from the seed coat to the cotyledons. GmSW17 encodes a ubiquitin-specific protease that interacts with GmSGF11 and GmENY2 to form a deubiquitinase module. This module affects H2Bub levels and negatively regulates GmDP-E2F-1 expression, thereby inhibiting G1-to-S conversion and affecting seed size. GmJAZ3 regulates GmCKXs gene expression through interaction with GmRR18a / GmMYC2a, participating in the regulation of seed size and grain weight. Furthermore, genes such as GmPP2C-1, GmCYP78A72, GmKIX8-1, and GmST05 also participate in the regulation of soybean seed weight. However, compared to the number of QTLs, the number of finely mapped QTLs and identified functional genes is relatively limited.
[0005] Therefore, discovering new soybean 100-seed weight-related SNPs with high phenotypic explanatory power, strong stability, and adaptability to germplasm from multiple ecological regions is of great significance for promoting the breeding of high-yield and high-quality soybean varieties and enhancing the competitiveness of the soybean industry. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a SNP site related to the 100-seed weight of soybean and its application.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows.
[0008] A SNP locus associated with 100-seed weight in soybean is located at physical location 39283583 bp on chromosome 11 of the soybean genome, with reference genome: Wm82.a2v1. This locus contains two allele types: G and A. Soybeans with different genotypes have different phenotypes of 100-seed weight.
[0009] Specifically, the upstream and downstream sequences of the SNP are shown below (corresponding to SEQ ID NO.4 in the sequence listing):
[0010] TCTGCGTTGTGATGTCACTGGGAAGATTCTTATGAAGTGCTTTCTTTCTGGAATGCCTGATTTAAAGTTGGGTCTGAATGACAAGATTGGTCTTGAGAAAGAAGCACAACTTAAGTCCCGTCCTRCTAAA AGGTAGCCTTCTGTTGCACTCATTCCCCTTTTGGCCACCAAGTTTTTTTTTTCCTTTAAATATTGTCATTTCAAACTTTAGTCTTTATGCTCTTCTGTTATGTTAGCATTGATA; where, R=[G / A].
[0011] The present invention also includes an application of the SNP site of claim 1, wherein the application is any of the following:
[0012] (1) Application of the SNP sites in soybean breeding;
[0013] (2) The application of the SNP sites in the preparation of products for soybean breeding;
[0014] (3) The application of the SNP sites in identifying or assisting in the identification of soybean 100-seed weight trait;
[0015] (4) The application of the SNP sites in the preparation of products for identification or auxiliary identification of the 100-seed weight trait of soybeans;
[0016] (5) Application of the SNP sites in the screening or breeding of soybean single plants, lines, varieties or cultivars based on the soybean 100-seed weight trait.
[0017] The present invention also includes a method for identifying or assisting in the identification of the 100-seed weight trait of soybeans, using a specific genotype based on the SNP locus to identify or assist in the identification of the 100-seed weight trait of soybeans, wherein the 100-seed weight of soybeans with the A-type gene is higher than or candidate to be higher than the 100-seed weight of soybeans with the G-type gene.
[0018] As a preferred embodiment of the present invention, the method for detecting whether the genotype of the SNP site in the soybean genome is G or A is as follows: the soybean genomic DNA to be tested is amplified by PCR using a specific KASP marker, and the genotype of the soybean to be tested is determined after fluorescence scanning.
[0019] As a preferred embodiment of the present invention, the specific KASP marker includes the downstream primer R shown in SEQ ID NO: 1, and the upstream primers Chr17-KASP-FAM and Chr17-KASP-VIC shown in SEQ ID NO: 2 and 3.
[0020] As a preferred technical solution of the present invention, the specific KASP label is used to perform PCR amplification on the soybean genomic DNA to be tested to obtain PCR products. Then, the fluorescence signal is converted into an analyzable value, and the fluorescence scanning results are displayed graphically. If FAM fluorescence is present and distributed near the y-axis, the genotype of the SNP site in the soybean genome to be tested is G; if VIC fluorescence is present and distributed near the x-axis, the genotype of the SNP site in the soybean genome to be tested is A.
[0021] The present invention also includes a specific KASP-labeled primer combination for detecting the SNP site, including downstream primer R as shown in SEQ ID NO: 1, and upstream primers Chr17-KASP-FAM and Chr17-KASP-VIC as shown in SEQ ID NO: 2 and 3.
[0022] The present invention also includes a reagent or kit for identifying or assisting in the identification of the 100-seed weight trait of soybean, the reagent or kit being used to detect the genotype of the SNP locus, and including at least the specific KASP marker primer combination.
[0023] The present invention also includes the use of the method, the specific KASP-labeled primer combination, the reagent or kit in any one of the following (1)-(6):
[0024] (1) Application in identifying or assisting in the identification of soybean 100-seed weight correlation traits;
[0025] (2) Application in the preparation of products for the identification or auxiliary identification of the weight correlation of 100 soybean seeds;
[0026] (3) Application in screening or assisting in the screening of high-grain-weight soybean varieties;
[0027] (4) Application in breeding or assisted breeding of soybean 100-seed weight related traits;
[0028] (5) Application in the preparation of products for breeding or assisted breeding of soybean 100-seed weight correlation traits;
[0029] (6) Application in identifying or assisting in the identification of soybean yield.
[0030] The present invention also includes a soybean breeding method, which first detects the genotype of the SNP site described in claim 1 in the soybean genome, and selects homozygous soybeans with the SNP site being A as parents for breeding.
[0031] The beneficial effects of the above technical solution are as follows: 100-grain weight is a key factor determining soybean yield. While there are many reports on 100-grain weight-related QTLs, relatively few have been developed for use in marker-assisted selection due to the large confidence intervals of most loci. In this invention, the research group first used a population of recombinant inbred lines (RILs) constructed from Jidou 17 and Zhonghuang 13 as parents to perform QTL mapping for 100-grain weight. One major-effect QTL, qHSW_11, was stably detected over three years, explaining 9.51-15.13% of the phenotypic variation. Further, the BC1F4 population was constructed, narrowing the confidence interval of qHSW_11 to 304.8 kb, predicting a total of 40 genes. Through parental differential sequence analysis, gene function annotation, and expression pattern analysis, four candidate genes were screened within this QTL region. Among them, the Glyma.11G239000 gene coding sequence showed an 8-bp insertion / deletion and six SNPs between the two parents. Further haplotype analysis in the SoyOmics database revealed that only different haplotypes of Glyma.11G239000 were highly correlated with 100-seed weight. Finally, using 247 soybean accessions as materials, genotyping was performed using the KASP marker developed based on sequence differences in the Glyma.11G239000 gene. Combined with multi-year phenotypic data, it was found that the 100-seed weight of the G / G haplotype of the Glyma.11G239000 gene (Hap02, ZH13) was significantly higher than that of the A / A haplotype (Hap01, JD17). Therefore, this KASP marker can be used for selecting high-100-seed-weight lines in soybean breeding. These findings enhance our understanding of seed weight regulation, lay the foundation for cloning soybean 100-seed-weight-related genes, and provide a practical tool for high-yield soybean breeding. Attached Figure Description
[0032] Figure 1 A bar chart showing the phenotypic distribution of 100 grain weights in different years for the population.
[0033] Figure 2 This is a schematic diagram of the qHSW_11 site and its fine localization.
[0034] Figure 3 This is a schematic diagram of candidate gene transcription level analysis.
[0035] Figure 4This is a schematic diagram of haplotype analysis of the Glyma.11G239000 gene.
[0036] Figure 5 This is a schematic diagram illustrating the KASP marker genotyping and phenotypic differences per 100 grains in 247 resource populations.
[0037] Figure 6 This is a population genetic linkage map.
[0038] Figure 7 This is a schematic diagram of the 100-grain weight phenotype of the parent and BC1F1.
[0039] Figure 8 This is a schematic diagram of the redistribution of 100 BC1F2 grains.
[0040] Figure 9 A schematic diagram of haplotype analysis of the Glyma.11G238800 gene. Detailed Implementation
[0041] The following embodiments illustrate the present invention in detail. All raw materials and equipment used in the present invention are conventional commercially available products and can be directly obtained through market purchase. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available. It should be understood that, as used in this specification and appended claims, the term "comprising" indicates the presence of the described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that the term "and / or" as used in this specification and appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0042] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [the described condition or event]," or "in response to detection." Furthermore, in the description of this specification and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. References to "one embodiment" or "some embodiments" described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in yet other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms “including,” “comprising,” “having,” and variations thereof all mean “including but not limited to,” unless otherwise specifically emphasized.
[0043] Previous Examples
[0044] In summary, in this invention, the inventors' research group conducted QTL mapping analysis on recombinant inbred line populations for three consecutive years to identify major QTLs related to 100-seed weight. Subsequently, through secondary population construction and analysis, the major stable QTLs were finely mapped. Candidate genes within the predicted intervals were then analyzed using sequence difference analysis, functional annotation analysis, and haplotype analysis. Finally, a KASP marker closely related to seed weight was developed and validated in soybean resource materials. This study lays an important foundation for the cloning of soybean seed weight-related genes and provides an important tool for high-yield marker-assisted breeding.
[0045] In this study, we continuously detected the major QTL locus qHSW_11, located on chromosome 11, associated with 100-seed weight, in a recurrent intercalary lymphoid transplantation (RIL) population over several years. To finely map this locus, we constructed the BC1F4 population using Zhonghuang 13 as the recurrent parent through marker-assisted selection. We then developed four InDel markers and seven KASP markers within the QTL region, constructing a high-density genetic linkage map. Combining the generation of backcross populations, individual genotyping, and phenotyping, we confirmed the role of qHSW_11 in controlling 100-seed weight, narrowing its region to approximately 340 kb of chromosome area, containing 40 annotated genes. To date, over 300 QTLs associated with soybean seed weight have been reported (SoyBase.org). This study identified five QTL loci associated with 100-seed weight on chromosomes 5, 6, 11, 18, and 19. Compared with previous studies, the qHSW_5 locus identified in this study overlaps with the qSWA1-1 locus located in the RIL population (Sun et al., 2012) and the QTN Gm05_38411981 locus located by GWAS (Cao et al., 2022). Within the qHSW_18 region, it overlaps with two markers identified by GWAS: Chr18_52898145 (Bhat et al., 2025) and Chr18_53141473 (Zhang et al., 2024). The SSR marker Satt229, located at 47,167,396 bp on chromosome 19, has also been reported to overlap with qHSW_19 identified in this study (Csanádi et al., 2001). The qHSW_6 region identified in this study is relatively large, and several QTLs associated with 100-grain weight have been reported in this region, including swHCC2-2 (Han et al., 2012) and qSW-06_1 (Kulkarni et al., 2017). The qHSW_11 locus identified in this study overlaps with the swHMB1-1 (and swHCB1-1) loci located by Han et al. (2012) in two RIL populations. However, the functional genes within this locus are still unknown.
[0046] In early studies, the low marker density of genetic maps led to large QTL intervals, posing a challenge to accurately predicting functional genes and developing precise molecular markers. Fine mapping of QTL intervals is an important approach to addressing this problem. Some studies have used residual heterozygous lines (RHLs) of the target fragment to obtain secondary segregating populations, successfully refining the intervals of certain loci. For example, Chen et al. (2023) constructed an F3 population of 1,800 families using 24 recombinant inbred lines, narrowing the interval of qSW16.1 to 33 kb. Similarly, Xia et al. used RHLs to construct a secondary population of 13,761 families, shortening the E1 interval to approximately 20 kb, ultimately successfully cloning the gene (Xia et al., 2012). Other studies have also used backcrossing methods to generate secondary segregating populations. For example, Zhang et al. (2024) used the BC3F3 population to finely map the rice flag leaf angle-related locus QFLANG-4B. Zhao et al. constructed a BC3F2 population of 6,000 individuals and successfully fine-mapped the rapeseed OILA5 locus to a 43 kb interval (Zhao et al., 2023). To narrow down the range of the stable major QTL qHSW_11, we selected a secondary population constructed by crossing the recombinant inbred line JZ 80 with Zhonghuang 13. The qHSW_11 chromosome segment of JZ80 originated from Jidou 17, while the chromosome segments of the other 100-seed weight QTLs (qHSW_05, qHSW_06, qHSW_18, qHSW_19) identified in this study all originated from Zhonghuang 13. This selection helps to eliminate interference from other loci. As expected, the 100-seed weight phenotype of the BC1F2 population, containing 214 families, showed a bimodal distribution. This indicates that phenotypic variation within the population is controlled by a single gene, suggesting that this population is suitable for fine-mapping analysis of the qHSW_11 locus. We used the BC1F2 population to determine that this locus is located within a 1.3 Mb interval. Subsequently, BC1F4 was constructed from 1,008 families that maintained heterozygosity in the target fragment, and the locus was finally finely located within a 340 kb interval.
[0047] Example 1: Soybean materials and field trials
[0048] A recombinant inbred line population (RILs) comprising 102 families was constructed by crossing Jidou 17 (18 g) and Zhonghuang 13 (24 g), which showed significant differences in 100-seed weight. The F2 generation of the RILs was passed down to the F6 generation using a single-seed method. The F6:8 to F6:10 ratios were used for 100-seed weight phenotypic identification and QTL mapping analysis. Additionally, 247 soybean germplasm resources from across China were used for 100-seed weight phenotypic identification and functional verification of related KASP molecular markers (Table S9, see Example 12).
[0049] The RIL population was sown for three consecutive years (2016 (F6:8), 2017 (F6:9), and 2018 (F6:10)) at the Tishang Experimental Station of the Institute of Grain and Oil Crops, Hebei Academy of Agricultural and Forestry Sciences. 247 soybean germplasm resources were planted at the Tishang Experimental Station in 2017 and 2018. Each family or resource was sown in three rows with a 3-meter row length in the field, with three randomized replicates. Row spacing was 50 cm and plant spacing was 10 cm.
[0050] Example 2: Trait Statistics and Data Analysis
[0051] For RIL populations and soybean resource materials, 10 plants with consistent growth were selected from the middle row of each plot to harvest mature seeds, which were then dried. For each sample, 100 mature dried seeds were randomly selected (using a seed counting plate), and their weight was measured using an electronic scale. The average of three technical replicates was used as the weight of 100 seeds. Frequency distribution plots were generated using Graphpad Prism 5.0 (http: / / www.graphpad.com / ). Statistical analysis was performed using SPSS Statistics 17.0. Analysis of variance was performed using phenotypic data from 2016 to 2018. Genotype and environment were fixed as factors to detect heritability. The test was used to compare means. The generalized heritability of a single environment was then calculated using the following formula: H = σ²G / (σ²G + (σ²e / r)); where σ²G is the genotypic variance; σ²e is the error variance; and r is the number of replicates.
[0052] Example 3: Construction of genetic linkage maps and QTL mapping
[0053] SNP genotyping of RILs was performed using the GBS method (Chen et al., 2022). Sequencing was conducted using an Illumina 2500 platform (Illumina, USA) by the Novogene Institute of Bioinformatics, Beijing. Burrows-WheelerAligner (BWA), SAM tools, and a custom Perl script were used to identify SNPs in the RIL population, and alignment and annotation were performed using the ANNOVAR software tool. Chi-square (χ²) was calculated for all SNPs. 2A segregational aberration test was used to detect markers. Markers with a segregational aberration test p < 0.001 or containing abnormal bases were filtered out. If TASSEL 5.0 results showed that more than 20% of individuals had missing genotypes, all markers were removed. Parental allele assignment and imputation were performed using the Window LD function of FsFHap integrated into TASSEL 5.0. After filtering, redundant markers were classified based on segregation patterns in the RIL population using the BIN function in IciMapping 4.1. bin was used as the genetic marker for constructing linkage maps, which were statistically computed on a MAP basis using IciMapping. Additive QTLs were detected using the CIM method with a p-value (PIN) of 0.01 for the input variables, based on the BIP model of QTL IciMapping. The LOD threshold for assessing the statistical significance of QTL effects was determined by 1000 permutations at a significance level of 0.05.
[0054] Example 4: Parental resequencing and molecular marker development
[0055] DNA was extracted from JD17 and ZH13 and sequenced using the Illumina NovaSeq 6000 platform. High-quality reads were aligned with “Glycine max Wm82.a2v1”. Molecular markers were developed using SNPs and InDels between Jidou 17 and Zhonghuang 13 within the qHSW-11 region. InDel primer sequences were designed using Primer 3, and KASP primer sequences were designed using “Design KASPprimer with Primer3 v2.5.0” (https: / / junli.netlify.app / apps / design-primers-with-primer3 / ). A total of 4 InDel markers and 7 KASP markers with polymorphism between JD17 and ZH13 were obtained, and these markers were used for fine mapping. For the INDEL marker, the PCR program was as follows: denaturation at 94°C for 5 minutes, followed by 33 cycles of 94°C for 30 seconds, 55°C for 30 seconds, and 72°C for 30 seconds; and a final extension at 72°C for 5 minutes. The 10 μL PCR system included: 5 μL 2×Taq PCR StarMix (Kangwei Biotechnology Beijing), 1 μL DNA template (approximately 50-100 ng), 1 μL of each InDel primer, and 3 μL of double-distilled water. PCR products were analyzed using 6% (w / v) agarose gel electrophoresis. The PCR program for the KASP marker included: 94°C for 10 minutes, 94°C for 30 seconds, 10 cycles of 65-55°C (decreasing by 1°C per cycle) for 40 seconds, 95°C for 30 seconds, 58°C for 30 seconds, and 72°C for 5 minutes. The total PCR volume was 8 μL, containing 4 μL of 2×157 Taq PCRStarMix (JasonGen, Beijing), 1 μL of DNA template (approximately 100-200 ng), 0.1 μL of each KASP primer, and 3 μL of double-distilled water. KASP PCR products were analyzed using a Bio-Rad CFX384 Touch2 instrument.
[0056] Example 5: Fine positioning of qHSW-11
[0057] To verify and further refine the genetic region of the major QTL qHSW-11, a secondary population was constructed by crossing the recombinant inbred line JZ 80 with ZH 13 from the RIL population. JZ80 had a 100-seed weight of 18g and carried the 17 allele of *Dioscorea opposita* within the qHSW-11 region, while other 100-seed weight QTL regions carried the 13 allele of *Zhonghuang*. In 2022, three BC1F1 plants were obtained. One of these plants was selected to develop the BC1F2 population, containing 290 lines, in 2023 to verify the chromosomal region of qHSW-11. Heterozygous plants of the qHSW-11 chromosomal region were identified and harvested after two consecutive generations (BC1F2 and BC1F3), and the BC1F4 population was finally constructed in 2024. BC1F1 individual plants, BC1F2, and BC1F4 populations were sown at the Dishan experimental station in the summers of 2022, 2023, and 2024, respectively. BC1F3 was sown at the Sanya Experimental Station (Sanya City, Hainan Province) in the winter of 2023.
[0058] Genomic DNA was extracted from young leaves of a plant population using the CTAB method. Based on QTL mapping results (… Figure 1 Recombinant plants within the BC1F2 and BC1F4 populations were identified using flanking markers, and the difference in 100-seed weight between recombinant homozygous and non-recombinant homozygous lines was analyzed. Student's test was used to analyze significant differences. A p-value < 0.01 indicated a significant difference in 100-seed weight among the tested progeny. Based on the progeny testing of all recombinants, qHSW-11 was narrowed down to a very short interval.
[0059] Example 6 RNA extraction and RT-qPCR
[0060] The expression levels of potential candidate genes in seeds of JD17, ZH13, and JZ80 soybean varieties were compared using RT-qPCR. These soybean varieties were planted at the Dishan Experimental Station in Shijiazhuang, Hebei Province in 2024. Seeds were then sampled 20 days after flowering. Total RNA was extracted using a plant RNA extraction kit (Tiangen Biotech). First-strand cDNA was synthesized using PrimeScript™ RT MasterMix (Perfect Real Time) (Novizan Biotech). Gene-specific primers were designed on the NCBI website, and sequences were synthesized. qRT-PCR reactions were performed using AceQ qPCR SYBR Green Master Mix (Novizan Biotech) on an ABI Step One Plus Realtime PCR system (Applied Biosystems, Foster City, CA, USA) according to the manufacturer's protocol. The qRT-PCR amplification conditions were 95 °C for 30 seconds, followed by 40 cycles of 95 °C for 10 seconds and 58 °C for 30 seconds. GmACTN 11 (Glyma.12g020500, GenBank accession number NM_001254696.2) was used as a reference gene (Hu et al., 2009) to standardize the relative expression levels of the test gene. Through 2 -△△CT The relative expression levels were calculated using the method (Livak and Schmittgen 2001). Each sample had three biological replicates and three technical replicates.
[0061] Example 7 Haplotype Analysis
[0062] Haplotype analysis of candidate genes was performed using the "HapSnap" function of the SoyOmic database (https: / / ngdc.cncb.ac.cn / soyomics / haplotype / ). 100-grain weight data from different soybean resources were downloaded from the SoyOmic database. Differences in 100-grain weight among different haplotypes were statistically analyzed using phenotypic data from the same year and location.
[0063] Example 8: Phenotypic identification of parental and recombinant inbred line populations
[0064] The 100-seed weight phenotype showed highly significant differences between the two parents, with ZH13 exhibiting a significantly higher 100-seed weight than JD17 (Table S4). Within the RIL population, the 100-seed weight phenotype showed extensive variation across all environments (Table S4). Figure 1 The variation range was 14.47 ~ 24.29 g. The population's 100-grain weight showed an approximately normal distribution, indicating that the population's 100-grain weight was controlled by multiple genes. Figure 1 Broadly defined heritability (H) 2The value was 0.67, indicating that the variation in 100-grain weight is mainly controlled by genotype (Table S4).
[0065] Table S4. Phenotypic Statistics of Parents and Population
[0066]
[0067] Example 9: Linkage Map Construction and QTL Analysis
[0068] We obtained 1,094,333 SNPs through sequencing alignment. After alignment between 102 lines from the two parents and the RIL population, a bin map was constructed using 1,261,526 markers across 20 chromosomes. This map covered 1,130 high-quality bins, with a total length of 2,416.41 cM and an average linkage group length of 120.82 cM. Chromosome 2 had the most polymorphic markers (89), with a total map length of 133.36 cM; chromosome 11 had the fewest polymorphic markers (33), with a total map length of 116.47 cM. The average distance between adjacent SNPs was 2.26 cM, ranging from 48 cM on chromosome 16 to 3.59 cM on chromosome 4 (Table S5). Figure 6 ).
[0069] Table S5 Population genetic map information
[0070]
[0071] Five QTLs for 100-grain weight were identified over three years, all of which were derived from Zhonghuang 13. These genes are located on five chromosomes, with a LOD of 2.76–4.78, explaining 7.05%–19.77% of the phenotypic variation (PVE). Among them, qHSW_11, located on chromosome 11, was consistently detectable over the three years, with PVE values ranging from 9.51% to 15.13% (Table 1). Figure 2 a) The qHSW_19 located on chromosome 19 had PVE values of 18.45% to 19.77% detected in both 2017 and 2018.
[0072] Table 1. QTL locus information for 100-grain weight under various conditions.
[0073]
[0074] Example 10: Fine positioning of qHSW_11
[0075] To narrow the confidence interval of the stable major-effect QTL qHSW_11, we constructed a secondary population by crossing Zhonghuang 13 with JZ80. The 100-seed weight of the three BC1F1 plants ranged from 22.38 to 23.50 g, significantly greater than that of the JZ80 plants (19 g). Figure 7 This indicates that the difference in 100-grain weight between Zhonghuang 13 and JZ80 may be controlled by a completely dominant gene, with the Zhonghuang 13 gene being dominant. The BC1F2 population contains 213 families, with 100-grain weights ranging from 19 to 23 g. The 100-grain weights in the BC1F2 population all exhibit a bimodal distribution. Figure 8 This suggests that the segregation of the 100-grain weight trait in this population may be controlled by a single gene.
[0076] In the BC1F2 population, we used seven developed markers for genotyping, obtaining a total of nine recombinant types. The ratio of large-grain families (158 families with a 100-grain weight of 21–26 g) to small-grain families (58 families with a 100-grain weight of 16–20 g) was approximately 3:1. This supports the hypothesis that this trait is controlled by a single, completely dominant locus. Therefore, in subsequent fine mapping analysis, homozygous ZH13 and heterozygous types were considered to be of the same type. The 100-grain weight of recombinant types R4–R7 was not significantly different from that of Jidou 17, but significantly lower than that of ZH13 (R1), R2–R3, and R8 types (…). Figure 2 b). Therefore, we narrowed down the qHSW_11 site to within 1.30 Mb, precisely locating it between the Indel-2 and Indel-4 markers on chromosome 11.
[0077] To further pinpoint qHSW_11, we constructed a BC1F4 population of 1,008 plants and screened them using KASP 1 and Indel 4 flanking markers. We then used seven developed KASP markers from this region to genotype recombinant individuals, identifying 16 recombinant types (…). Figure 2 c). Notably, compared to Zhonghuang 13, recombinant R4-R8, R14, R15, and R16 exhibited significantly lower 100-grain weights, while R2, R3, and R9-R13 showed no significant differences. Ultimately, we successfully narrowed down qHSW_11 to the interval between markers KASP 4 and KASP 6, approximately 340 kb (reference genome: W82 a2v1), ranging from 33.27 to 33.57 Mb. Figure 2 c).
[0078] Example 11: Candidate gene analysis within the qHSW 11 interval
[0079] According to the genome annotation of Wm82.a2v1 in Phytozome 13 (https: / / phytozome.jgi.doe.gov / pz / portal.html), 40 annotated genes are contained in this 340 kb region. To further identify candidate genes, we analyzed differentially expressed genes between Jidou 17 and Zhonghuang 13 within the qHSW_11 region. Within the qHSW_11 genome region, a total of 5 genes (Glyma.11G238800, Glyma.11G239000, Glyma.11G239300, Glyma.11G240600, Glyma.11G240700) exhibit sequence variations in exon regions, including 12 SNPs and 4 InDels. These variations result in 5 synonymous variants, 7 non-synonymous variants, 2 non-frameshift insertions, and 2 frameshift deletions.
[0080] RT-qPCR analysis showed that Glyma.11G240700 was almost not expressed in seeds. Glyma.11G240600, Glyma.11G238800, and Glyma.11G239300 were all expressed during seed development, but there was no significant difference in expression levels among the parents. Figure 3 The expression levels of Glyma.11G239000 in Jidou 17 and JZ80 were significantly lower than those in Zhonghuang 13; the expression levels of Glyma.11G238800 in Jidou 17 and JZ80 were significantly higher than those in Zhonghuang 13. Therefore, these two genes were initially predicted as candidate genes. Based on the functional annotation of Arabidopsis homologous genes, Glyma.11G239000 encodes an EIN3 family protein. Compared with ZH13, an 8 bp insertion was found in the coding sequence (CDS) region encoding the EIN3 DNA-binding domain in JD17. In addition, four SNPs were identified in the EIN3 DNA-binding region and two SNPs were identified in other CDS regions. Figure 4 ).
[0081] Example 12 Candidate gene haplotype analysis
[0082] Genetic diversity of Glyma.11G239000 and Glyma.11G238800 was assessed using the "HapSnap" function of the SoyOmics database (Liu et al., 2020). Glyma.11G238800 was classified into two major haplotypes (MAF 5%) among 4,384 soybean accessions. 100-seed weight data over two years were obtained from the SoyOmics database for 571 soybean varieties from the same location (Liu et al., 2020), including 257 Hap1 varieties (Jidou 17) and 314 Hap2 varieties (Zhonghuang 13). Analysis of the two major haplotypes showed no significant difference in 100-seed weight between the two major haplotypes. Figure 9 ).
[0083] Haplotype analysis revealed that the Glyma.11G239000 gene was classified into four haplotypes (MAF>5%) in 5,214 soybean accessions (126 wild soybean accessions, 1,908 local soybean accessions, and 3,180 cultivated soybean accessions). Haplotype 3 was mainly found in wild soybean, haplotype 2 had the largest proportion in local varieties, and haplotype 1 had a relatively large proportion in cultivated varieties. Figure 4 b). 100-seed weights over two years were obtained from the SoyOmics database for 641 soybean varieties from the same location (Liu et al., 2020), including 227 varieties of Hap1, 335 varieties of Hap2, and 53 varieties of Hap4. Hap3 comprised 26 varieties, which were not statistically analyzed due to MAF being less than 5%. Analysis of the three major haplotypes showed that varieties carrying Hap2 (Zhonghuang 13) exhibited higher 100-seed weights (18.07 g in 2014 and 18.15 g in 2015), while varieties carrying Hap1 (Jidou 17, 16.69 g in 2014; 17.02 g in 2015) and Hap4 (15.89 g in 2014 and 15.76 g in 2015) showed lower 100-seed weights. Figure 4 c). These results indicate that Glyma.11G239000 is associated with the 100-seed weight of soybean.
[0084] Example 13 Development and Validation of Gene-Specific Markers
[0085] To apply the above research in breeding programs, we constructed a specific KASP marker (KASP-Glyma.11G239000) based on the G / A mutation site at 54 bp (physical location 39,283,583 bp) in the Glyma.11G239000 coding region. Specifically, it includes:
[0086] Chr17-KASP-R: CCAAAAGGGGGAAATGAGTGC; SEQ ID NO.1;
[0087] Chr17-KASP-FAM: GAAGGTGACCAAGTTCATGCTACAACTTAAGTCCCGTCCTG; SEQ ID NO.2;
[0088] Chr17-KASP-VIC: GAAGGTCGGAGTCAACGGATTACAACTTAAGTCCCGTCCTA; SEQ ID NO.3.
[0089] To verify the effectiveness of the KASP marker, we tested it in 247 soybean germplasm resources with a 100-seed weight of 6.00–42.50 g. In this population, 137 and 110 soybean germplasm accessions, respectively, had homozygous genotypes G / G and A / A (…). Figure 5 a, Table S9). In the phenotypic data of 2017, the 100-grain weight ranged from 7.18 g to 28.81 g and from 10.83 g to 42.50 g for the two genotypes, with average values of 17.45 g and 19.51 g, respectively. Figure 5 b). In 2018, the 100-grain weights of the two genotypes ranged from 6.00 to 26.35 g and from 7.85 to 40.05 g, respectively, with average values of 15.09 g and 16.56 g, respectively. Figure 5 (b) One-way ANOVA showed that the 100-seed weight of soybean materials with the A / A genotype was significantly higher than that of soybean materials with the G / G genotype. These results indicate that the locus-specific marker can effectively distinguish between low and high 100-seed weight lines and can be effectively applied to marker-assisted breeding.
[0090] Table S9. Different haplotypes and phenotypes of Glyma.11G239000 in 247 materials.
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0098] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A SNP site associated with 100-seed weight of soybean, characterized in that: The SNP locus is located at physical location 39283583 bp on chromosome 11 of the soybean genome, with reference genome: Wm82.a2v1. This locus contains two allele types: G and A. Soybeans with different genotypes have different phenotypes of 100-seed weight.
2. An application of the SNP site according to claim 1, characterized in that: The application is any one of the following: (1) Application of the SNP sites in soybean breeding; (2) The application of the SNP sites in the preparation of products for soybean breeding; (3) The application of the SNP sites in identifying or assisting in the identification of soybean 100-seed weight trait; (4) The application of the SNP sites in the preparation of products for identification or auxiliary identification of the 100-seed weight trait of soybeans; (5) Application of the SNP sites in the screening or breeding of soybean single plants, lines, varieties or cultivars based on the soybean 100-seed weight trait.
3. A method for identifying or assisting in the identification of the 100-seed weight trait of soybeans, characterized in that: The soybean 100-seed weight trait is identified or assisted in by using a specific genotype based on the SNP site described in claim 1. The 100-seed weight of soybeans with the A-type gene is higher than or can be higher than that of soybeans with the G-type gene.
4. The method for identifying or assisting in the identification of the 100-seed weight trait of soybeans according to claim 3, characterized in that: The method for detecting whether the genotype of the SNP site in the soybean genome is G or A is as follows: PCR amplification of the soybean genomic DNA to be tested is performed using a specific KASP marker, and the genotype of the soybean to be tested is determined after fluorescence scanning.
5. The method for identifying or assisting in the identification of the 100-seed weight trait of soybeans according to claim 4, characterized in that: The specific KASP marker includes the downstream primer R as shown in SEQ ID NO: 1, and the upstream primers Chr17-KASP-FAM and Chr17-KASP-VIC as shown in SEQ ID NO: 2 and 3.
6. The method for identifying or assisting in the identification of the 100-seed weight trait of soybeans according to claim 5, characterized in that: The soybean genomic DNA to be tested was amplified by PCR using the specific KASP label described in claim 5 to obtain PCR products. The fluorescence signal was then converted into an analyzable value, and the fluorescence scanning results were displayed graphically. If FAM fluorescence was present and distributed near the y-axis, the genotype of the SNP site in the soybean genome to be tested was type G; if VIC fluorescence was present and distributed near the x-axis, the genotype of the SNP site in the soybean genome to be tested was type A.
7. A specific KASP-labeled primer combination, characterized in that: The primers used to detect the SNP sites described in claim 1 include the downstream primer R shown in SEQ ID NO: 1, and the upstream primers Chr17-KASP-FAM and Chr17-KASP-VIC shown in SEQ ID NO: 2 and 3.
8. A reagent or kit for identifying or assisting in the identification of the 100-seed weight trait of soybeans, characterized in that: The reagent or kit is used to detect the SNP locus genotype of claim 1, and includes at least the specific KASP marker primer combination as described in claim 7.
9. The use of the method of any one of claims 3-6, the specific KASP-labeled primer combination of claim 7, or the reagent or kit of claim 8 in any one of the following (1)-(6): (1) Application in identifying or assisting in the identification of soybean 100-seed weight correlation traits; (2) Application in the preparation of products for the identification or auxiliary identification of the weight correlation of 100 soybean seeds; (3) Application in screening or assisting in the screening of high-grain-weight soybean varieties; (4) Application in breeding or assisted breeding of soybean 100-seed weight related traits; (5) Application in the preparation of products for breeding or assisted breeding of soybean 100-seed weight correlation traits; (6) Application in identifying or assisting in the identification of soybean yield.
10. A soybean breeding method, characterized by: First, the genotype of the SNP site described in claim 1 in the soybean genome is detected, and homozygous soybeans with the SNP site being A are selected as parents for breeding.