Molecular marker closely linked with soybean grain protein content QTL qPro14 and application of molecular marker
By identifying the QTL site qPro14 on soybean chromosome 14 and developing PARMS markers, the problem of the negative correlation between protein content and oil content in high-protein soybean breeding was solved, realizing an efficient and low-cost breeding screening method and improving the ability to analyze soybean grain protein content.
Patent Information
- Application Number
- CN202511244825.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-02
AI Technical Summary
In existing technologies, breeding high-protein soybeans faces the challenge of a negative correlation between protein content and oil content, and there is insufficient fine mapping of QTLs for soybean seed protein content and validation of candidate genes, making high-protein soybean breeding difficult.
By constructing an association population based on 768 core soybean germplasm resources, genome-wide association analysis was used to identify the QTL site qPro14 on soybean chromosome 14, and a PARMS marker closely linked to it was developed for efficient screening of soybean materials with high protein content.
This method enables efficient screening of soybean seed protein content, explains 0.32% of phenotypic variation, is simple to operate and low in cost, and is suitable for high-throughput screening of large-scale breeding populations, providing technical support for high-protein soybean breeding.
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of molecular biology and genetic breeding, and particularly relates to a molecular marker closely linked to a soybean seed protein content QTL qPro14 and application thereof. BACKGROUND
[0002] Soybean (Glycine max (L) Merr.) is the most important source of high-quality plant protein and edible oil in China, and plays an irreplaceable role in improving people's dietary structure and promoting the development of animal husbandry. Soybean protein has good solubility, emulsifying property, water holding capacity, oil holding capacity, gelation, foaming property and other physicochemical properties, and is also an ideal food processing aid. The protein content in soybean seeds varies widely. In recent years, many quantitative trait loci (QTL) that regulate the variation of seed protein content have been identified in the positive genetics of soybean seed protein. Exploring the regulatory mechanism of soybean seed storage protein also helps to improve the protein content. However, since the protein content in soybean seeds is negatively correlated with oil content and yield, the breeding practice of cultivating high-protein soybeans is challenging. In order to overcome the limitations of this negative correlation on the improvement of soybean protein content, it is necessary to have a deeper understanding of the synthesis pathway and genetic regulatory mechanism of soybean seed protein. Therefore, it is of great significance to analyze the genetic basis of soybean seed protein content, to mine high-protein excellent alleles in soybean germplasm resources, and to create new germplasm of high-protein soybeans for the directional breeding of high-protein soybean varieties.
[0003] Soybean seed protein content is a quantitative trait controlled by multiple genes, which is affected by genotypes, environments, and the interaction between genotypes and environments (Patil et al 2018). Most of the QTL mapping studies for soybean seed protein content were based on two soybean materials to construct biparental mapping populations. By June 2024, Soybase database (https: / / www.soybase.org) announced 241 QTLs associated with soybean seed protein content, which were distributed on 20 chromosomes. However, there were still few QTLs that were finely mapped and verified by candidate gene function. Among all the identified QTLs for soybean seed protein content, two loci on chromosome 20 and 15, qPRO20 and qPRO15, were widely detected due to their large additive effects and strong stability (Warrington et al. 2015). The fine mapping of these two loci and the identification of key candidate genes have been the hotspots of soybean seed protein genetic dissection. Until 30 years after the first detection, Fliege et al. (2022) successfully fine-mapped qPRO20 to a ~77.8 kb interval and found a high correlation between seed protein content and insertion / deletion in the CCT domain of Glyma.20G085100, and verified the function of CCT by RNA interference (RNAi) transgenic soybean. Almost at the same time, Goettel et al. (2022) also revealed Glyma.20G085100 as the candidate gene of qPRO20 by GWAS. For another major hot spot qPRO15, Zhang et al. (2020) mapped qPRO15 to a 4 Mbp interval by linkage mapping, and then performed association analysis of the candidate segment in the interval in a natural population, further determining GmSWEET39 (Glyma.15G049200) as the key candidate gene of qPRO15. Meanwhile, two GWAS studies on soybean seed protein and oil content also revealed the key role of Glyma.15G049200 (Miao et al 2020; Wang et al 2020).
[0004] With the development of next-generation sequencing technology and high-quality assembly of soybean reference genome, GWAS has become a powerful method to mine protein QTL in the past decade, greatly improving the efficiency and accuracy of QTL mapping (Gupta et al 2017; Patil et al. 2017). Hwang et al. (2014) performed association analysis on 298 soybean germplasm resources using 55,159 SNPs, and found 40 significant SNPs associated with 17 highly related genomic regions of seed protein content. Zhang et al. (2014) found significant SNP clusters associated with seed protein on chromosomes 4, 10, 15 and 19 in 192 soybean germplasm resources by GWAS. Multiple GWAS studies detected significant association loci with soybean seed protein content on chromosomes 20 and 15, and used the advantage of high positioning accuracy to identify candidate genes at these two loci in the later stage (Vaughn et al 2014, Bandillo et al 2015, Sonah 2015). In addition to the two major loci on chromosomes 20 and 15, other important QTLs frequently located were also obtained more new insights by whole genome association. Many GWAS showed significant association with soybean seed protein content in SNPs near qPRO8 on chromosome 8 (Li et al 2018; Shook et al 2021). Another repeatedly located QTL, qPRO5 on chromosome 5, was also identified by GWAS in the past decade (Hwang et al 2014; Vaughn et al 2014; Sonah et al 2015; Li et al. 2018). Duan et al. (2022) conducted GWAS on soybean seed thickness and found GmST05 (Glyma.05G244100), which not only controls seed size but also affects seed protein and oil. In addition, a recent GWAS on seed thickness found that ST1 (Glyma.08G109100) is located near the confirmed protein QTL on chromosome 8, and has an impact on seed shape change and seed oil content (Li et al 2022).
[0005] The present application is based on an association population constructed from 768 domestic and foreign soybean core germplasm resources, combined with population genotype data and seed protein content phenotype data, and a major QTL site qPro14 regulating soybean seed protein content variation is identified by whole genome association analysis, and a PARMS marker closely linked to it is developed, which can be used to assist soybean plant type breeding. SUMMARY
[0006] The application aims to provide application of a reagent for detecting base No. 1,360,542 of chromosome No. 14 of soybean in screening breeding of soybean seed protein content.
[0007] The application further aims to provide application of the reagent for detecting base No. 1,360,542 of chromosome No. 14 of soybean in preparation of a screening kit for soybean seed protein content.
[0008] The application further aims to provide a screening breeding method for soybean seed protein content.
[0009] In order to achieve the above-mentioned purposes, the application adopts the following technical measures:
[0010] Obtaining of a molecular marker closely linked to the QTL qPro14 of soybean seed protein content:
[0011] (1) Population construction and phenotype identification: 768 soybean materials with extensive genetic diversity from 23 provinces in China are selected as core resources, and a soybean association population is constructed based on the materials. The 2024 is planted in Hefei Experimental Base of Anhui Academy of Agricultural Sciences (2024HF), and a randomized block design is used in field test, and three repetitions are arranged; 2-row zones are planted, and one row of each family is planted with a total of 20 plants, the row length is 2m, and the row spacing is 0.5m. After maturation, the seeds are harvested, and the near-infrared spectroscopy method is used to determine the seed protein content (Protein Content, %), and the average value of three repetitions in each environment is taken as the phenotype value of the material in the environment.
[0012] (2) Genotype analysis: the whole genome resequencing of the 768 materials of the association population is performed by using the Huada T7 sequencing platform, the genotyping of the 768 materials is performed by using the resequencing technology, the average sequencing depth is about 20x, the SNP sites with a missing rate of >10% and a minimum allele frequency of <0.05 are filtered, and finally 6,339,330 high-quality SNPs are reserved for whole genome association analysis.
[0013] (3) Whole genome association analysis: the mixed linear model (MLM) in the GEMMAX software is used for association analysis, and the significance threshold is set to P≤1 / n (n is the number of SNPs, 6339330). A stable associated QTL site qPro14 is found on chromosome 14, which can explain the correlation population phenotype variation rate of 0.32%. The peak SNP marker is named S14_1360542, which is located at base No. 1360542 of chromosome No. 14 of the Glycine_max_v2.1 reference genome of soybean, the allele is T / G, and the average seed protein content of the material containing the high seed protein content allele is 2.34% higher than that of the material containing the low seed protein content allele.
[0014] (4) PARMS marker development: specific primers were designed according to the sequence upstream and downstream of S14_1360542 site to construct a PARMS detection system. The primer sequences are as follows:
[0015] PARMS14: TTATTAAAAGAACATTTTTC,
[0016] PARMS14P1: GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT, PARMS14P2: GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.
[0017] The protection scope of the present application includes:
[0018] The reagent for detecting the genotype of base No. 1,360,542 on chromosome No. 14 of soybean is applied in breeding of soybean seed protein content.
[0019] The reagent for detecting base No. 1,360,542 on chromosome No. 14 of soybean is applied in preparation of a soybean seed protein content screening kit.
[0020] The above-mentioned application, if the base No. 1,360,542 on chromosome 14 of soybean is detected as T, it is determined that the soybean is a high seed protein content type material.
[0021] The above-mentioned application, if the base No. 1,360,542 on chromosome 14 of soybean is detected as G, it is determined that the soybean is a low seed protein content type material.
[0022] The above-mentioned application, the reagent is preferably a primer.
[0023] The above-mentioned primer is preferably a PARMS detection primer, and more preferably the primer provided by the present application: PARMS14: TTATTAAAAGAACATTTTTC, PARMS14P1: GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT, and PARMS14P2: GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.
[0024] A soybean seed protein content screening and breeding method, comprising detecting base 1,360,542 of chromosome 14 of soybean by using conventional schemes in the art, which include but are not limited to sequencing method, TaqMan probe method, AS-PCR method, molecular beacon method, high-resolution melting curve method, CAPS method, SNaPshot method, KASP method, PARMS method, gene chip method or mass spectrometry method.
[0025] The version number of the soybean reference genome used in the application is Glycine_max_v2.1, and the website is https: / / ensembl.gr amene.org / Glycine_max / .
[0026] Compared with the prior art, the application has the following advantages:
[0027] (1) The QTL qPro14 identified in the application can explain 0.32% of the seed protein content phenotype variation of the association population, and has high breeding application value.
[0028] (2) The developed PARMS marker is simple to operate, low in cost, clear in typing, and suitable for high-throughput screening of large-scale breeding populations. DETAILED DESCRIPTION
[0029] The technical solutions described in the application are conventional technologies in the art if not specifically stated; the reagents or materials described are from commercial channels if not specifically stated. The version number of the soybean reference genome used in the application is Glycine_max_v2.1, and the website is https: / / ensembl.gr amene.org / Glycine_max / .
[0030] Example 1:
[0031] SNP molecular marker significantly associated with soybean seed protein content QTL qPro14:
[0032] Test materials: 768 materials from 23 provinces in China with wide genetic diversity as the core resources of the soybean association population.
[0033] (1) Soybean population seed protein content identification: 2024 was planted in the Hefei Experimental Base of Anhui Academy of Agricultural Sciences (2024HF), and the field test adopted a randomized block design with 3 repetitions; 2-row zones were planted, and each family was planted in one row with a total of 20 plants, the row length was 2m, and the row spacing was 0.5m. After maturation, the seeds were harvested, and the near-infrared spectroscopy method was used to determine the seed protein content (Protein Content, %), and the average value of 3 repetitions in each environment was taken as the phenotype value of the material in the environment.
[0034] (2) Genotype analysis: 768 samples of the association population were subjected to whole genome resequencing using Huada T7 sequencing platform. Genotyping was performed on the 768 samples using resequencing technology. The average sequencing depth was ~20x. SNP sites with a deletion rate >10% and a minimum allele frequency <0.05 were filtered. Finally, 6,339,330 high-quality SNPs were retained for whole genome association analysis.
[0035] (3) Whole genome association analysis: The mixed linear model (MLM) in the GEMMAX software was used for association analysis in combination with population genotype data and phenotype data. The significance threshold was set to P≤1 / n (n is the number of SNPs, 6339330).
[0036] (4) Obtaining qPro14 and its significantly associated SNP marker: The association analysis results showed that a stable associated QTL site qPro14 was found on chromosome 14, which was significantly associated in 2 environments and could explain 0.32% of the phenotypic variation rate. The peak SNP marker was named S14_1360542, located at the 1360542th base of chromosome 14 of the Glycine max v2.1 reference genome, with an allele of C / T, and the flanking sequence was: 5'-TTGTATGAGCTTAGTCTCGGTATGCAAATTGCAAAGTACTAGAGATATTT[T / G]ATTTATACATGAAAAATGTTCTTTTAATAAATAACTTGAAAAGCATTTAC-3'. In the 2024HF environment, the average seed protein content of the material containing the high seed protein content allele was 2.34% higher than that of the material containing the low seed protein content allele.
[0037] Example 2:
[0038] Development of a PARMS marker closely linked to soybean seed protein content:
[0039] According to the nucleotide sequences of the positions before and after the peak SNP marker S14_1360542 significantly associated with qPro14, the PARMS marker detection primer sequences were obtained according to the primer design principle as follows:
[0040] PARMS14: TTATTAAAAGAACATTTTTC,
[0041] PARMS14P1: GAAGGTGACCAAGTTCATGCT GCAAAGTACTAGAGATATTTT,
[0042] PARMS14P2: GAAGGTCGGAGTCAACGGATT GCAAAGTACTAGAGATATTTG.
[0043] Underlined is the fluorescent linker.
[0044] The method for detecting the genotype of the soybean qPro14 locus to be tested by using the above-mentioned PARMS primer set is as follows:
[0045] (1) Extract the genomic DNA of the soybean to be tested.
[0046] (2) Prepare the reaction system. The reaction system is 5 μL, including 2.5 μL 2xPARMS PCR reaction mix (a product of Wuhan Jingpeibio Technology Co., Ltd.), primer PARMS12, primer PARMS12P1, primer PARMS12P2 aqueous solution, DNA and water. In the reaction system, the concentrations of primer PARMS12P1 and primer PARMS12P2 are both 150 nM, and the concentration of primer PARMS12 is 400 nM.
[0047] (3) Add 5 μL of paraffin oil (to prevent sample evaporation) to the reaction system, and then perform PCR amplification.
[0048] The reaction program is: 95℃ for 15 min; 95℃ for 20 s, 65℃ for 1 min, decrease by 0.8℃ per cycle until 57℃, 10 cycles; 95℃ for 20 s, 57℃ for 1 min, 32 cycles.
[0049] (4) After completing step (3), perform signal reading on TECAN Infinite M1000, and then make the following judgment: if blue is displayed, the corresponding soybean is or is suspected to be a high-seed-protein-content soybean; if green is displayed, the corresponding soybean is or is suspected to be a low-seed-protein-content soybean.
[0050] The sequence amplified in the high-seed-protein-content material Yushu Xian No. 2 by using the above-mentioned primer is:
[0051] 5'-GCAAAGTACTAGAGATATTT T ATTTATACATGAAAAATGTTCTTTTAATAA-3'
[0052] The sequence of the amplification product of the low-seed-protein-content material Jiyu 166 is:
[0053] 5'-GCAAAGTACTAGAGATATTT G ATTTATACATGAAAAATGTTCTTTTAATAA-3'
[0054] Example 3:
[0055] Universality of PARMS marker in selection of soybean seed protein content trait:
[0056] The PARMS primer set designed in Example 2 was used to detect the genotype and genetic effect of the qPro14 locus in the soybean to be tested. The soybean to be tested was 256 soybean varieties (lines) at home and abroad. Field identification was carried out in 2024 at Baishiying base of Chongqing Academy of Agricultural Sciences (2024CQ) according to the method used in Example 1, and the seed protein content was investigated.
[0057] The results showed that among the above 256 soybean materials, there were 212 materials with genotype TT, and the average seed protein content was 47.03%; there were 44 materials with genotype GG, and the average seed protein content was 46.00%. The difference in seed protein content between TT and GG genotypes reached a very significant level (P-value = 1.19e-3). The results showed that the genotype of qPro14 locus in 256 soybean materials at home and abroad was separated, and had stable and reliable genetic effect.
[0058] The above results show that the prepared PARMS molecular marker qPro14 has a greater genetic effect on the seed protein content of soybean and has a good screening effect.
Claims
1. Use of a reagent for detecting base 1,360,542 of chromosome 14 of soybean in screening breeding of soybean seed protein content.
2. Use of a reagent for detecting base 1,360,542 of chromosome 14 of soybean in preparation of a kit for screening soybean seed protein content.
3. Use according to claim 1 or 2, characterized in that: If the reagent detects that base 1,360,542 of chromosome 14 of soybean is T, it is determined that the soybean is a high seed protein content type material.
4. Use according to claim 1 or 2, characterized in that: If the reagent detects that base 1,360,542 of chromosome 14 of soybean is G, it is determined that the soybean is a low seed protein content type material.
5. Use according to claim 1 or 2, characterized in that: The reagent is a primer.
6. Use according to claim 5, characterized in that: The primer is PARMS14: TTATTAAAAGAACATTTTTC, PARMS14P1: GAAGGTGACCAAGTTCATGCTGCAAAGTACTAGAGATATTTT, and PARMS14P2: GAAGGTCGGAGTCAACGGATTGCAAAGTACTAGAGATATTTG.
7. A method for screening breeding of soybean seed protein content, comprising detecting the genotype of base 1,360,542 of chromosome 14 of soybean, and the method comprises sequencing, TaqMan probe, AS-PCR, molecular beacon, high resolution melting curve, CAPS, SNaPshot, KASP, PARMS, gene chip or mass spectrometry.
Citation Information
Patent Citations
Single nucleotide mutation site SNP and KASP markers significantly associated with soybean protein content and application thereof
CN112877467A
Soybean grain protein content related molecular marker located on soybean chromosome 14 and application of molecular marker
CN117305501A
SNP (Single Nucleotide Polymorphism) molecular marker for detecting soybean quality gene and application of SNP molecular marker
CN120060553A
Soybean polymorphisms and methods of genotyping
WO2008153804A2