A molecular marker closely linked to soybean seed protein content QTLqPro14 and its application

By identifying the QTL locus qPro14 on soybean chromosome 14 through genome-wide association analysis and developing PARMS markers, the problem of the negative correlation between soybean seed protein content and oil content was solved, enabling efficient screening and breeding of high-protein soybeans.

CN120905431BActive Publication Date: 2026-07-24INST OF FOOD CROPS HUBEI ACAD OF AGRI SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF FOOD CROPS HUBEI ACAD OF AGRI SCI
Filing Date
2025-09-02
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, soybean seed protein content is negatively correlated with oil content and yield, making it difficult to breed high-protein soybeans. Furthermore, the fine mapping of QTLs for soybean seed protein content and the validation of candidate genes are insufficient, making it difficult to achieve targeted breeding of high-protein soybeans.

Method used

By constructing an association population based on 768 core soybean germplasm resources, genome-wide association analysis was used to identify the QTL site qPro14 on soybean chromosome 14, and a PARMS marker closely linked to it was developed for efficient screening of soybean materials with high protein content.

Benefits of technology

This method enables efficient screening of soybean seed protein content, explains 0.32% of phenotypic variation, is simple to operate and low in cost, is suitable for high-throughput screening of large-scale breeding populations, and provides a basis for the creation of new high-protein soybean germplasm.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application belongs to the technical field of molecular biology and genetic breeding, and discloses a molecular marker closely linked to a soybean seed protein content QTL qPro14 and application. A seed protein content QTL qPro14 is identified on chromosome 14 of soybean by using whole genome association analysis. A significant SNP associated with the QTL is located at the 1,360,542th base of chromosome 14 of a reference genome Glycine_max_v2.1, and can explain 0.32% of phenotypic variation. A PARMS marker developed by using the SNP is clear in genotyping and simple in operation, and is suitable for molecular breeding of soybean seed protein content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular biology and genetic breeding technology, specifically relating to a molecular marker closely linked to soybean seed protein content QTL qPro14 and its application. Background Technology

[0002] Soybean (Glycine max (L) Merr.) is one of the most important sources of high-quality plant protein and edible oil in my country, playing an irreplaceable role in improving people's dietary structure and promoting the development of animal husbandry. Soybean protein possesses excellent physicochemical properties such as solubility, emulsification, water-holding capacity, oil-holding capacity, gelling properties, and foaming properties, making it an ideal food processing aid. The protein content in soybean seeds varies widely. In recent years, forward genetic studies of soybean seed protein have identified many quantitative trait loci (QTLs) that regulate variations in seed protein content. Exploring the regulatory mechanisms of soybean seed storage proteins can also help improve protein content. However, because soybean seed protein content is negatively correlated with oil content and yield, breeding high-protein soybeans presents certain challenges. To overcome the limitations imposed by this negative correlation on soybean protein content improvement, a deeper understanding of the synthetic pathways and genetic regulatory mechanisms of soybean seed proteins is needed. Therefore, analyzing the genetic basis of soybean seed protein content, identifying superior high-protein alleles in soybean germplasm resources, and creating new high-protein soybean germplasm are of great significance for the targeted breeding of high-protein soybean varieties.

[0003] Soybean seed protein content is a quantitative trait controlled by multiple genes, influenced by genotype, environment, and the interaction between genotype and environment (Patil et al. 2018). Most QTL mapping studies on soybean seed protein content are based on biparental mapping populations constructed from two soybean materials. As of June 2024, the Soybase database (https: / / www.soybase.org) published 241 QTLs related to soybean seed protein content, distributed across 20 chromosomes (Figure 1). However, few QTLs have been finely mapped and their candidate gene functions have been validated (Table 1). Among all identified soybean seed protein content QTLs, two loci, qPRO20 and qPRO15, on chromosomes 20 and 15, have been widely detected due to their large additive effects and strong stability (Warrington et al. 2015). The fine mapping of these two loci and the identification of key candidate genes have been a hot topic in the genetic anatomy of soybean seed protein. Thirty years after its initial detection, Fliege et al. (2022) successfully fine-mapped qPRO20 to a ~77.8 kb region and discovered insertions / deletions highly correlated with seed protein content in the CCT domain of Glyma.20G085100. They then validated the function of the CCT domain using RNA interference (RNAi) transgenic soybeans. Almost simultaneously, Goettel et al. (2022) also revealed Glyma.20G085100 as a candidate gene for qPRO20 using GWAS. Regarding another major potential site, qPRO15, Zhang et al. (2020) mapped qPRO15 to a 4 Mbp region using linkage mapping. Subsequent association analysis of this region in a natural population further identified GmSWEET39 (Glyma.15G049200) as a key candidate gene for qPRO15. Meanwhile, two GWAS studies analyzing the protein and oil content of soybean seeds also revealed the key role of Glyma.15G049200 (Miao et al 2020; Wang et al 2020).

[0004] With the development of next-generation sequencing technology and the high-quality assembly of the soybean reference genome, GWAS has become a powerful method for mining protein QTLs in the past decade, greatly improving the efficiency and accuracy of QTL mapping (Gupta et al. 2017; Patil et al. 2017). Hwang et al. (2014) conducted association analysis on 55,159 SNPs in 298 soybean germplasm resources, and found 40 significantly associated SNPs in 17 genomic regions highly associated with seed protein content. Zhang et al. (2014) used GWAS to find significant SNP clusters associated with seed protein on chromosomes 4, 10, 15, and 19 in 192 soybean germplasm resources. Multiple GWAS studies have detected significantly associated sites with soybean seed protein content on chromosomes 20 and 15, and have used the advantage of high mapping accuracy to help identify candidate genes for these two sites in later stages (Vaughn et al. 2014, Bandillo et al. 2015, Sonah 2015). Besides the two major loci on chromosomes 20 and 15, other frequently mapped important QTLs have also yielded new insights through genome-wide association studies. Many GWAS studies have shown significant associations between SNPs on chromosome 8 near qPRO8 and soybean seed protein content (Li et al 2018; Shook et al 2021). Another repeatedly mapped QTL, qPRO5 on chromosome 5, has also been identified multiple times by GWAS over the past decade (Hwang et al 2014; Vaughn et al 2014; Sonah et al 2015; Li et al. 2018). Duan et al. (2022) conducted a GWAS study on soybean seed thickness and found that GmST05 (Glyma.05G244100), in addition to controlling seed size, also affects seed protein and oil content. In addition, a recent GWAS study on seed thickness found that ST1 (Glyma.08G109100) is located near a known protein QTL on chromosome 8 and has an impact on seed morphology and seed oil content (Li et al 2022).

[0005] This invention is based on an associated population constructed from 768 core soybean germplasm resources from home and abroad. Combining population genotype data and seed protein content phenotypic data, genome-wide association analysis was used to identify a major QTL site qPro14 that regulates variation in soybean seed protein content. A PARMS marker closely linked to qPro14 was also developed, which can be used to assist in soybean plant type breeding. Summary of the Invention

[0006] The purpose of this invention is to provide a reagent for detecting bases at positions 1,360,542 of soybean chromosome 14 and its application in soybean seed protein content screening breeding.

[0007] Another object of the present invention is to provide the application of a reagent for detecting bases at positions 1,360,542 of soybean chromosome 14 in the preparation of a soybean seed protein content screening kit.

[0008] The final objective of this invention is to provide a method for screening and breeding soybean seed protein content.

[0009] To achieve the above objectives, the present invention adopts the following technical measures:

[0010] Obtaining a molecular marker tightly linked to soybean seed protein content QTL qPro14:

[0011] (1) Population Construction and Phenotypic Identification: 768 soybean accessions with broad genetic diversity from 23 provinces in my country were selected as core resources, and soybean-related populations were constructed based on these accessions. The 2024 soybeans were planted at the Hefei Experimental Base of the Anhui Academy of Agricultural Sciences (2024HF). The field trial adopted a randomized block design with three replicates; two rows were planted, with 20 plants per family per row, 2m long and 0.5m apart. After maturity, the grains were harvested, and the protein content (%) was determined using near-infrared spectroscopy. The average of the three replicates under each environment was taken as the phenotypic value of the material under that environment.

[0012] (2) Genotyping analysis: Using the BGI T7 sequencing platform, whole-genome resequencing was performed on 768 materials from the associated population. Genotyping was performed on the 768 materials using resequencing technology. The average sequencing depth was ~20×. SNPs with a deletion rate >10% and a minimum allele frequency <0.05 were filtered out. Finally, 6,339,330 high-quality SNPs were retained for whole-genome association analysis.

[0013] (3) Genome-wide association analysis: A mixed linear model (MLM) was used in GEMMAX software for association analysis, with a significance threshold set to P ≤ 1 / n (n being the number of SNPs, 6,339,330). The results showed a stable associated QTL locus, qPro14, on chromosome 14, which explained 0.32% of the phenotypic variation in the associated population. Its peak SNP marker was named S14_1360542, located at base 1360542 on chromosome 14 of the soybean genome reference genome Glycine_max_v2.1, with an allele of T / G. Materials containing the high grain protein content allele had an average grain protein content 2.34% higher than materials containing the low grain protein content allele.

[0014] (4) PARMS marker development: Specific primers were designed based on the upstream and downstream sequences of the S14_1360542 site to construct a PARMS detection system. The primer sequences are as follows:

[0015] PARMS14:TTATTAAAAGAACATTTTTC、

[0016] PARMS14P1:GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT,

[0017] PARMS14P2:GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATTTTG.

[0018] The scope of protection of this invention includes:

[0019] Application of reagents for detecting the genotype of soybean chromosome 14 at positions 1,360,542 in soybean seed protein content breeding.

[0020] Application of reagents for detecting bases at positions 1,360,542 on soybean chromosome 14 in the preparation of a soybean seed protein content screening kit.

[0021] In the above-described applications, if the base at position 1,360,542 of soybean chromosome 14 is detected to be T, then the soybean is determined to be a high-seed protein content material.

[0022] If the base at position 1,360,542 of soybean chromosome 14 is detected as G, then the soybean is determined to be a low-seed protein content material.

[0023] In the above applications, the preferred reagent is a primer.

[0024] The primers described above are preferably PARMS detection primers, and more preferably the primers provided by the present invention: PARMS14:TTATTAAAAGAACATTTTTC, PARMS14P1:GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT and PARMS14P2:GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.

[0025] A method for screening soybean seed protein content in breeding includes detecting bases at positions 1,360,542 of soybean chromosome 14 using conventional methods in the art. These conventional methods include, but are not limited to, sequencing, TaqMan probe method, AS-PCR method, molecular beacon method, high-resolution melting curve method, CAPS method, SNaPshot method, KASP method, PARMS method, gene chip method, or mass spectrometry.

[0026] The version number of the soybean reference genome used in this invention is Glycine_max_v2.1, and the URL is https: / / ensembl.gramene.org / Glycine_max / .

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] (1) The QTL qPro14 identified in this invention can explain 0.32% of the phenotypic variation in grain protein content in the associated population, and has high breeding application value.

[0029] (2) The developed PARMS markers are easy to operate, low in cost, and have clear typing, making them suitable for high-throughput screening of large-scale breeding populations. Detailed Implementation

[0030] Unless otherwise specified, the technical solutions described in this invention are all conventional techniques in the field; the reagents or materials described, unless otherwise specified, are all from commercial sources. The version number of the soybean reference genome used in this invention is Glycine_max_v2.1, and the URL is https: / / ensembl.gramene.org / Glycine_max / .

[0031] Example 1:

[0032] SNP molecular markers significantly associated with soybean seed protein content QTL qPro14:

[0033] Test materials: 768 soybean accessions with broad genetic diversity from 23 provinces in my country were used as core resources to construct a soybean-related population.

[0034] (1) Identification of soybean grain protein content: 2024 soybean was planted at the Hefei Experimental Base (2024HF) of Anhui Academy of Agricultural Sciences. The field experiment adopted a randomized block design with 3 replicates. The plants were planted in 2-row plots, with 20 plants per family per row, 2m long and 0.5m apart. After maturity, the grains were harvested and the grain protein content (%) was determined by near-infrared spectroscopy. The average of the 3 replicates under each environment was taken as the phenotypic value of the material under that environment.

[0035] (2) Genotyping analysis: Using the BGI T7 sequencing platform, whole-genome resequencing was performed on 768 materials from the associated population. Genotyping was performed on the 768 materials using resequencing technology. The average sequencing depth was ~20×. SNPs with a deletion rate >10% and a minimum allele frequency <0.05 were filtered out. Finally, 6,339,330 high-quality SNPs were retained for whole-genome association analysis.

[0036] (3) Genome-wide association analysis: Combining population genotype and phenotypic data, the mixed linear model (MLM) in GEMMAX software was used for association analysis, and the significance threshold was set to P ≤ 1 / n (n is the number of SNPs, 6339330).

[0037] (4) Obtaining qPro14 and its significantly associated SNP markers: Association analysis results showed that a stable associated QTL site qPro14 was found on chromosome 14, which was significantly associated in both environments and could explain 0.32% of the phenotypic variation. Its peak SNP marker was named S14_1360542, located at base 1360542 on chromosome 14 of the soybean Glycine_max_v2.1 reference genome, with an allele of T / G and its flanking sequence as: 5'-TTGTATGAGCTTAGTCTCGGTATGCAAATTGCAAAGTACTAGAGATATTT[T / G]ATTTATACATGAAAAATGTTCTTTTAATAAATAACTTGAAAAGCATTTAC-3'. Under the 2024HF environment, the average grain protein content of the material containing the high grain protein content allele was 2.34% higher than that of the material containing the low grain protein content allele.

[0038] Example 2:

[0039] Development of a PARMS marker closely linked to soybean seed protein content:

[0040] Based on the nucleotide sequences preceding and following the peak SNP marker S14_1360542, which is significantly associated with qPro14, the PARMS marker detection primer sequences were obtained according to primer design principles as follows:

[0041] PARMS14:TTATTAAAAGAACATTTTTC、

[0042] PARMS14P1: GAAGGTGACCAAGTTCATGCT GCAAAGTACTAGAGATATTTT、

[0043] PARMS14P2: GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.

[0044] The underlined part is the fluorescent connector.

[0045] The method for detecting the genotype of the soybean qPro14 locus using the above PARMS primer set is as follows:

[0046] (1) Extract genomic DNA from the soybeans to be tested.

[0047] (2) Preparation of the reaction system. The reaction system consisted of 5 μL, including 2.5 μL of 2 × PARMS PCR reaction mix (a product of Wuhan Jingtai Biotechnology Co., Ltd.), aqueous solutions of primers PARMS14, PARMS14P1, and PARMS14P2, DNA, and water. In the reaction system, the concentrations of primers PARMS14P1 and PARMS14P2 were both 150 nM, and the concentration of primer PARMS14 was 400 nM.

[0048] (3) Add 5 μL of paraffin oil to the reaction system (to prevent sample evaporation), and then perform PCR amplification.

[0049] The reaction program was as follows: 95℃ for 15 min; 95℃ for 20 s, 65℃ for 1 min, decreasing by 0.8℃ per cycle until reaching 57℃, for 10 cycles; 95℃ for 20 s, 57℃ for 1 min, for 32 cycles.

[0050] (4) After completing step (3), the signal is read on the TECAN Infinite M1000 and then the following judgment is made: if blue is displayed, the corresponding soybean is or is suspected to be a high-seed protein content soybean; if green is displayed, the corresponding soybean is or is suspected to be a low-seed protein content soybean.

[0051] Using the primers described above, the sequence amplified from the high-seed protein content material Yu Shu Xian 2 is:

[0052] 5'-GCAAAGTACTAGAGATATTT T ATTTATACATGAAAAATGTTCTTTTAATAA-3'

[0053] The amplification product sequence of the low-seed protein content material Jiyu 166 is:

[0054] 5'-GCAAAGTACTAGAGATATTT G ATTTATACATGAAAAATGTTCTTTTAATAA-3'

[0055] Example 3:

[0056] Universality of PARMS markers in the selection of soybean seed protein content traits:

[0057] The PARMS primer set designed in Example 2 was used to detect the genotype and genetic effect of the qPro14 locus in the soybean to be tested. 256 soybean varieties (lines) from both domestic and international sources were used for the test. Following the method described in Example 1, field identification was conducted in 2024 at the Baishiyi Base of the Chongqing Academy of Agricultural Sciences (2024CQ), and the seed protein content was examined.

[0058] The results showed that among the 256 soybean materials, 212 had the TT genotype, with an average grain protein content of 47.03%; and 44 had the GG genotype, with an average grain protein content of 46.00%. The difference in grain protein content between the TT and GG genotypes was highly significant (P-value = 1.19e-3). This result indicates that the qPro14 locus is segregated in the 256 domestic and international soybean materials and has a stable and reliable genetic effect.

[0059] The above results indicate that the prepared PARMS molecular marker qPro14 has a significant genetic effect on the seed protein content of soybeans and has a good screening effect.

Claims

1. Application of a reagent for detecting bases at positions 1,360,542 of soybean chromosome 14 in soybean seed protein content screening breeding. If the reagent detects a homozygote with a base T at position 1,360,542 of soybean chromosome 14, the soybean is determined to be a high-seed protein content material. If the reagent detects a homozygote with a base G at position 1,360,542 of soybean chromosome 14, the soybean is determined to be a low-seed protein content material. The version number of the soybean reference genome is Glycine_max_v2.

1.

2. Application of the reagent for detecting bases at positions 1,360,542 of soybean chromosome 14 in the preparation of a soybean seed protein content screening kit. If the reagent detects a homozygote with a base T at position 1,360,542 of soybean chromosome 14, the soybean is determined to be a high-seed protein content material. If the reagent detects a homozygote with a base G at position 1,360,542 of soybean chromosome 14, the soybean is determined to be a low-seed protein content material. The version number of the soybean reference genome is Glycine_max_v2.

1.

3. The application according to claim 1 or 2, characterized in that: The reagent is a primer.

4. The application according to claim 3, characterized in that: The primers are PARMS14:TTATTAAAAGAACATTTTTC, PARMS14P1:GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT and PARMS14P2:GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.

5. A method for screening soybean seed protein content in breeding, comprising detecting the genotype of the 1,360,542nd base of soybean chromosome 14, wherein the method is sequencing, TaqMan probe method, AS-PCR method, molecular beacon method, high-resolution melting curve method, CAPS method, SNaPshot method, KASP method, PARMS method, gene chip method, or mass spectrometry. If a homozygous genotype of T at the 1,360,542nd base of soybean chromosome 14 is detected, the soybean is determined to be a high seed protein content material; if a homozygous genotype of G at the 1,360,542nd base of soybean chromosome 14 is detected, the soybean is determined to be a low seed protein content material. The version number of the soybean reference genome is Glycine_max_v2.1.