SNP sites associated with soybean protein content and uses thereof
By detecting the SNP genotype at position 35708082 on chromosome 19 of the soybean genome, and using quantitative real-time PCR and high-resolution melting curve analysis, the problem of low efficiency in traditional breeding methods was solved, enabling efficient screening and identification of high-protein soybean varieties, thus improving soybean breeding efficiency and protein content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEAST AGRICULTURAL UNIVERSITY
- Filing Date
- 2025-05-15
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies are insufficient to rapidly and effectively increase the protein content of soybean seeds. Traditional breeding methods are inefficient, and the protein content of soybeans is significantly affected by the environment. Furthermore, there is a lack of theoretical support for molecular marker-assisted breeding.
By detecting the SNP genotype at position 35708082 on chromosome 19 of the soybean genome, high-protein soybean varieties were screened and identified using quantitative real-time PCR and high-resolution melting curve analysis, and breeding methods for high-protein soybean varieties were developed.
This method enables rapid and accurate identification of soybean protein content, screening of high-protein varieties, improving the efficiency of soybean breeding and the protein content of soybean seeds, and providing a theoretical basis for molecular-assisted breeding.
Smart Images

Figure BDA0005404145390000081 
Figure BDA0005404145390000091 
Figure HDA0005404145400000011
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to SNP sites related to soybean protein content and their applications. Background Technology
[0002] Soybeans are one of China's important food crops and a significant source of protein. They are rich in plant-based protein, calcium, phosphorus, iron, carotene, soy lecithin, protease inhibitors, and plant cholesterol. Therefore, regular consumption of soybeans can not only improve the body's immune function but also protect the cardiovascular system. Protein and oil are two of the most important characteristics of soybeans, with protein accounting for approximately 40% of the total weight of soybean seeds. Soybean protein is rich in amino acids, including many essential amino acids, making it a highly nutritious and high-quality plant-based protein. Besides being consumed, soybean protein can also be used as a raw material in animal feed, medicine, and health products, thus the demand for soybean protein is substantial.
[0003] Soybean protein content is a quantitative trait, easily influenced by the environment. Numerous studies both domestically and internationally have investigated QTL mapping for soybean protein content and its components. Research indicates that soybean yield, protein content, and protein production in different cold environments of Northeast China are significantly affected by varietal genotype and production environment. QTLs controlling quality traits are primarily distributed across all 20 chromosomes, and because soybean protein content is a quantitative trait easily regulated by the environment, it possesses a complex genetic basis. Therefore, using marker-assisted breeding to explore genetic loci related to soybean protein content and elucidate the genetic mechanisms influencing protein content variations can provide more theoretical support for molecular-assisted breeding of soybean quality traits.
[0004] Given the limited planting area and difficulty in increasing yield per unit area, increasing the protein content of soybean seeds is a feasible way to increase total protein production. While increasing soybean protein content using traditional breeding methods is relatively slow, with the gradual development of bioscience and technology, traditional breeding can be combined with molecular breeding. By detecting single nucleotide polymorphisms, superior soybean varieties with high protein content can be screened. Summary of the Invention
[0005] One object of the present invention is to provide a novel use for a substance for detecting the genotype of a soybean SNP locus.
[0006] This invention provides the application of a substance for detecting the genotype of a soybean SNP locus in any of the following (a1)-(a8):
[0007] (a1) To identify or assist in the identification of the protein content of the soybean to be tested;
[0008] (a2) Prepare products for identification or auxiliary identification of the soybean protein content to be tested;
[0009] (a3) Screening or assisted screening of high-protein soybean varieties;
[0010] (a4) Prepare products for screening or assisting in the screening of high-protein soybean varieties;
[0011] (a5) Soybean variety improvement;
[0012] (a6) Prepare products for soybean variety improvement;
[0013] (a7) Soybean breeding;
[0014] (a8) Prepare products for soybean breeding.
[0015] Another object of the present invention is to provide a product that functions as any one of the following (c1)-(c4):
[0016] (c1) Identify or assist in the identification of the protein content of the soybean to be tested;
[0017] (c2) Screening or assisted screening of high-protein soybean varieties;
[0018] (c3) Soybean variety improvement;
[0019] (c4) Soybean breeding.
[0020] The product provided by this invention includes a substance for detecting the genotype of a soybean SNP locus.
[0021] Furthermore, the product also includes other reagents for quantitative PCR and high-resolution melting curve analysis, such as DNA polymerase and fluorescent dyes.
[0022] In some embodiments, the DNA polymerase is EasyTaq polymerase. The fluorescent dye is EvaGreen.
[0023] Furthermore, the product also includes positive controls (such as genomic DNA of Hap1 genotype soybean, genomic DNA of Hap2 genotype soybean, and genomic DNA of Hap3 genotype soybean) and negative controls.
[0024] The application of the above-mentioned product in any of the following (c1)-(c4) is also within the scope of protection of this invention:
[0025] (c1) Identify or assist in the identification of the protein content of the soybean to be tested;
[0026] (c2) Screening or assisted screening of high-protein soybean varieties;
[0027] (c3) Soybean variety improvement;
[0028] (c4) Soybean breeding.
[0029] The SNP site described above is located at position 35708082 on chromosome 19 of the soybean genome.
[0030] The substance described above for detecting the genotype of the soybean SNP locus to be tested can be any one of the following (b1)-(b3):
[0031] (b1) PCR primers for amplifying soybean genomic DNA fragments including the SNP sites;
[0032] (b2) PCR reagents containing the PCR primers described in (b1);
[0033] (b3) A kit containing the PCR primers described in (b1) or the PCR reagents described in (b2).
[0034] In a specific embodiment of the present invention, the PCR primers consist of single-stranded DNA as shown in Sequence 1 and single-stranded DNA as shown in Sequence 2.
[0035] Another objective of this invention is to provide a method for identifying or assisting in the identification of the protein content in soybeans to be tested.
[0036] The method for identifying or assisting in the identification of soybean protein content provided by the present invention includes the following steps: detecting the genotype of the soybean SNP site to be tested, and identifying or assisting in the identification of soybean protein content based on the genotype of the soybean SNP site to be tested; the SNP site is located at deoxyribonucleotide position 35708082 on chromosome 19 of the soybean genome.
[0037] In the above methods for identifying or assisting in the identification of the protein content of soybeans to be tested, the genotype is Hap1 genotype, Hap2 genotype or Hap3 genotype;
[0038] The Hap1 genotype is a homozygous type with C deoxyribonucleic acid at position 35708082 on chromosome 19 of the soybean genome.
[0039] The Hap2 genotype is a homozygous type with deoxyribonucleic acid A at position 35708082 on chromosome 19 of the soybean genome.
[0040] The Hap3 genotype is a heterozygous type with C and A deoxyribonucleotides at position 35708082 on chromosome 19 of the soybean genome.
[0041] In the above-mentioned methods for identifying or assisting in identifying the protein content of soybeans to be tested, the method for identifying or assisting in identifying the protein content of soybeans based on the genotype of the SNP locus to be tested may be that the protein content of soybeans with the Hap1 genotype is higher or is a candidate higher than that of soybeans with the Hap2 genotype; the protein content of soybeans with the Hap2 genotype is higher or is a candidate higher than that of soybeans with the Hap3 genotype.
[0042] In the above-mentioned methods for identifying or assisting in the identification of the protein content of soybeans to be tested, the method for detecting the genotype of the SNP site in the soybean to be tested can be direct sequencing or sequencing of PCR products containing the SNP site.
[0043] In this invention, the sequencing method is not limited and can be a sequencing method classified by various classification methods, such as first-generation sequencing (i.e., Sanger sequencing), second-generation sequencing (i.e., NGS sequencing), and third-generation sequencing (i.e., long-read sequencing) classified by technology generation, as well as genome sequencing and transcriptome sequencing classified by application field, and single-end sequencing, paired-end sequencing, etc. classified by sequencing method.
[0044] Furthermore, the method for detecting the genotype of the soybean SNP site to be tested includes the following steps: using the soybean genomic DNA to be tested as a template, performing PCR amplification using the above-mentioned PCR primers to obtain a melting curve, and determining the genotype of the soybean SNP site to be tested based on the melting curve.
[0045] Furthermore, the method for determining the genotype of the soybean SNP site to be tested based on the melting curve is as follows: if the melting curve of the soybean to be tested is the same as that of the Hap1 genotype soybean, then the soybean genotype to be tested is the Hap1 genotype; if the melting curve of the soybean to be tested is the same as that of the Hap2 genotype soybean, then the soybean genotype to be tested is the Hap2 genotype; if the melting curve of the soybean to be tested is the same as that of the Hap3 genotype soybean, then the soybean genotype to be tested is the Hap3 genotype.
[0046] Furthermore, the PCR amplification reaction system (10 μL) is as follows: 0.4 μL genomic DNA (10 ng / μL), 0.4 μL forward primer (single-stranded DNA molecule shown in Sequence 1), 0.4 μL reverse primer (single-stranded DNA molecule shown in Sequence 2), 5 μL EasyTaq polymerase, 1 μL EvaGreen, and 2.8 μL ddH2O. The final concentration of both the forward and reverse primers in the reaction system is 10 μM.
[0047] The PCR amplification reaction conditions are as follows: 94℃ pre-denaturation for 5 min; 94℃ denaturation for 30 s, 50℃ annealing for 30 s, 72℃ extension for 30 s, steps two to four are set for 32 cycles, 72℃ extension for 7 min, then the temperature is increased by 0.2℃ per minute, and the temperature is increased to 85℃ to absorb fluorescence and obtain the melting curve.
[0048] Another objective of this invention is to provide a method for screening or assisting in the screening of high-protein soybean varieties.
[0049] The method for screening or assisting in screening high-protein soybean varieties provided by the present invention includes the following steps: selecting soybean varieties with the Hap1 genotype; wherein the Hap1 genotype is a homozygous type with deoxyribonucleotide C at position 35708082 on chromosome 19 of the soybean genome.
[0050] Finally, another objective of this invention is to provide a method for improving soybean varieties or breeding soybeans.
[0051] The method for soybean variety improvement or soybean breeding provided by the present invention includes the following steps: selecting soybean varieties with the Hap1 genotype as parents for breeding.
[0052] In the above-mentioned methods for improving soybean varieties or breeding soybeans, the indicators for variety improvement or breeding include protein content.
[0053] The purpose of the variety improvement or breeding includes developing high-protein soybean varieties. High protein means a protein content greater than that of the parent varieties.
[0054] The protein content mentioned above refers to the seed protein content. This seed protein content can be determined using a Fox Grain Analyzer.
[0055] The high-protein soybean variety mentioned above can be a soybean variety with a seed protein content greater than or equal to 39.4%.
[0056] The reference genome version number of any of the soybean genomes mentioned above is Glycine max Wm82.a2.v1.
[0057] The soybean mentioned above can be any soybean germplasm resource, variety, strain or single plant.
[0058] In some implementations, the soybeans are recombinant inbred lines obtained by using Kenfeng 14, Kenfeng 15, Kenfeng 19 and Heinong 48 as parents, and by preparing double cross combinations (Kenfeng 14 × Kenfeng 15) × (Heinong 48 × Kenfeng 19) using the single-seed transfer method.
[0059] This invention provides a SNP locus associated with soybean protein content, located at position 35708082 on chromosome 19 of the soybean genome, with polymorphisms of C or A. Based on this SNP locus, the invention also developed the HRM melting curve method for identifying soybean protein content and validated it in four-way recombinant inbred lines. Validation results show that the SNP locus discovered in this invention is significantly correlated with soybean seed protein content. Soybean materials with the Hap1 genotype (SNP locus C) have significantly higher seed protein content than those with the Hap2 genotype (SNP locus A). Furthermore, soybean materials with the Hap2 genotype (SNP locus A) have significantly higher seed protein content than those with both SNP loci (C and A) in the Hap3 genotype. This invention provides a theoretical basis for molecular-assisted breeding of soybean protein content traits and is of great significance for cultivating new high-protein soybean varieties. Attached Figure Description
[0060] Figure 1 Melting curves were used to identify different genotypes of SNP loci using HRM. The top graph shows the melting curves for fluorescence intensity of the three haplotypes; the bottom graph shows the solubility curves for the three haplotypes.
[0061] Figure 2 A comparative analysis of protein content in three haplotype soybean materials from a recombinant self-pollinated population. Detailed Implementation
[0062] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0063] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0064] The following embodiments of Kenfeng 14 (corresponding to kenfeng14 in the literature), Kenfeng 15 (corresponding to kenfeng15 in the literature), Kenfeng 19 (corresponding to kenfeng19 in the literature), and Heinong 48 (corresponding to heinong48 in the literature) are all recorded in the literature "[1] Li, Wen-Xia., Wang, Ping., Wang, Ping., Zhao, Hengxing., Sun, Xu.. QTL for MainStem Node Number and Its Response to Plant Densities in 144 Soybean FW-RILs. Frontiers in plant science, 2021, 12."
[0065] The construction method of the four-way recombinant inbred line (FW-RIL) material in the following examples refers to the method in the literature "Ning Hailong et al., Construction of population genetic map of soybean four-way recombinant inbred line, Soybean Science, October 2015".
[0066] Example 1: Obtaining SNP sites related to soybean protein content
[0067] I. Construction of a soybean protein content-related population and determination of its traits
[0068] 1. Test materials
[0069] The test materials included four-way recombinant inbred lines (FW-RIL) and related materials.
[0070] Four-way recombinant inbred line (FW-RIL) materials: Kenfeng 14, Kenfeng 15, Kenfeng 19 and Heinong 48 were used as parents to prepare two hybrid combinations (Kenfeng 14 × Kenfeng 15) × (Heinong 48 × Kenfeng 19). The F1 generation of the two hybrid combinations was crossed to form the F2 generation. The F2 generation was self-crossed for 7 generations, and homozygous four-way recombinant inbred lines (FW-RIL) with each genotype were obtained by single seed propagation.
[0071] Related materials: A resource bank constructed from natural populations of 455 high-quality soybean germplasm resources, including 4 local varieties, 387 domestic varieties and 44 foreign varieties.
[0072] 2. Experimental Methods
[0073] FW-RIL was planted in the following 14 environments from 2018 to 2021: In 2018, it was planted in Acheng City, Heilongjiang Province (E1-E2, E126.95°, N45.52°); in 2019, it was planted in Acheng City, Heilongjiang Province (E3-E4, E126.95°, N45.52°); in 2019, it was planted in Shuangyashan City, Heilongjiang Province (E5-E6, E131.15°, N46.64°); and in 2020, it was planted in... It was planted in Acheng City, Heilongjiang Province (E7-E8, E126.95°, N45.52°) in 2020, in Shuangyashan City, Heilongjiang Province (E9-E10, E131.15°, N46.64°) in 2021, and in Xiangyang Town, Heilongjiang Province (E13-E14, E126.68°, N45.72°) in 2021.
[0074] Germplasm resource groups were planted in the following two environments between 2018 and 2020: Harbin, Heilongjiang Province (E1) in 2018, Shuangyashan, Heilongjiang Province (E2) in 2019, Harbin (E3) in 2019, and Harbin (E4) in 2020.
[0075] After maturity, a suitable amount of seeds from each variety were selected and their protein content was determined using a Fowles cereal analyzer. Ten replicate measurements were performed on each sample, and the average value was used as the phenotypic value for QTL and GWAS mapping. Based on the measured protein content data, the mean, coefficient of variation, kurtosis, and skewness of the protein content of the four-way recombinant inbred lines were calculated using SAS 9.2 and Excel 2013. Normality tests and analysis of variance were performed. The absolute values of kurtosis and skewness in the 14 environments of FW-RIL were all close to 0, indicating that the protein content followed a normal distribution.
[0076] II. QTL and GWAS Joint Mapping of Protein Content Traits
[0077] Based on the linkage map constructed in previous studies, additive QTLs were located in the FW-RIL population using the GAPL software, employing both interval mapping (IM-ADD) and inclusive composite interval mapping (ICIM-ADD) methods. The scan step size was set to 1.00 cM, and the LOD threshold was set to 2.50. The PIN value for the ICIM-ADD method was set to 0.001. Over 4 years and 14 environments, 56 major-effect QTLs with a contribution rate greater than 10% of protein content-related phenotypes were detected in the four-dimensional population. Based on population structure and LOD results, GWAS analysis was performed using the R language software package mrMLM.GUI. This package includes five multi-site methods: mrMLM, FASTmrMLM, FASTmrEMMA, pLARmEB, and pKWmEB to locate QTNs. In the first stage, except for FASTmrEMMA, the critical P-value was set to 0.005, while the critical P-values for the other methods were set to 0.01. In the final stage, the critical LOD value for significant QTNs was set to 3. The phylogenetic matrix used in the analysis was also calculated using this R software, and a total of 333 QTNs associated with protein content were detected.
[0078] III. Identification of SNP sites related to soybean protein content
[0079] The 333 QTNs identified by association analysis in the germplasm population were compared with the 56 QTLs identified by linkage analysis in the recombinant inbred line population FW-RIL. Among these, 44 QTN loci were found within the genomic regions of 7 QTLs repeatedly mapped under multiple methods and conditions. Potential candidate genes were searched at 43kb intervals on either side of the QTN loci based on the LD decay distance (86kb). Based on resequencing results, inter-parental gene sequence variation analysis was performed to screen for genes with amino acid sequence differences caused by promoter or exon variations. Genes related to protein content were screened based on gene functional annotation, ultimately identifying a SNP locus associated with soybean protein content. This SNP locus is located at position 35708082 on chromosome 19 of the soybean genome (reference genome version: Glycine max Wm82.a2.v1), with polymorphisms of C or A. This locus is significantly associated with soybean protein content.
[0080] Example 2: High-resolution melting curve (HRM) method for identifying soybean protein content based on SNP sites.
[0081] 1. Primer design
[0082] To identify the genotype of the SNP locus significantly associated with soybean protein content obtained in Example 1, a molecular marker was developed based on the SNP locus at 35708082 bp on chromosome 19 of the soybean genome. Primer pairs were designed using PrimerPremier5 software to make the target product 147 bp in length. The primer sequences are as follows:
[0083] Forward primer sequence: 5'-ACTTCTTGTCACTACAGCATT-3' (Sequence 1).
[0084] Reverse primer sequence: 5'-CACACTTTTGGAAGCCTTGG-3' (sequence 2).
[0085] 2. High-resolution melting curve (HRM) method for identifying soybean protein content
[0086] Using soybean genomic DNA as a template, PCR amplification was performed using the primer pair designed in step 1.
[0087] The PCR reaction system (10 μL) consisted of: 0.4 μL genomic DNA (10 ng / μL), 0.4 μL forward primer, 0.4 μL reverse primer, 5 μL EasyTaq polymerase (colorless) (Kangwei Century Biotechnology Co., Ltd., Cat: CW2965M), 1 μL EvaGreen (Biotium, Cat: 31000), and 2.8 μL ddH2O. The final concentration of both the forward and reverse primers in the PCR reaction system was 10 μM.
[0088] The PCR reaction conditions are as follows: 94℃ pre-denaturation for 5 min; 94℃ denaturation for 30 s, 50℃ annealing for 30 s, 72℃ extension for 30 s, steps two to four are set for 32 cycles, 72℃ extension for 7 min, then the temperature is increased by 0.2℃ per minute to 85℃ to absorb fluorescence and obtain the melting curve.
[0089] Using Roche HRM analysis was performed using a 96 real-time quantitative PCR instrument. If the melting curve of the soybean being tested is the same as that of the Hap1 genotype soybean, then the soybean genotype is Hap1 (CC homozygous); if the melting curve is the same as that of the Hap2 genotype soybean, then the soybean genotype is Hap2 (AA genotype); if the melting curve is the same as that of the Hap3 genotype soybean, then the soybean genotype is Hap3 (CA heterozygous). Specifically, Hap1 genotype soybeans have an SNP of C, Hap2 genotype soybeans have an SNP of A, and Hap3 genotype soybeans have both SNPs of C and A. The melting curves of Hap1, Hap2, and Hap3 genotype soybeans are shown below. Figure 1 As shown.
[0090] Example 3: Application of the High-Resolution Melting Profile (HRM) Method for Identifying Soybean Protein Content
[0091] Test materials: 89 soybean materials derived from a four-way recombinant inbred line (FW-RIL) population.
[0092] Experimental methods: The genotypes of different SNP loci in the test materials were identified using the method described in Example 2. The protein content in the seeds of the test materials was determined using a Fox Grain Analyzer.
[0093] The results of the detection of SNP loci genotypes and seed protein content of the tested materials are shown in Table 1. The results showed that among the 89 materials, 66 were soybean materials with Hap1 genotype (CC genotype), 16 were soybean materials with Hap2 genotype (AA genotype), and 7 were soybean materials with Hap3 genotype (CA genotype).
[0094] Furthermore, the seed protein content of soybeans from different genotypes was compared and significant differences were analyzed. The results are as follows: Figure 2 As shown in the figure. The results showed that the average protein content in soybean seeds of the Hap1 genotype was significantly higher than that of soybeans of the Hap2 and Hap3 genotypes, and the average protein content of soybeans of the Hap2 genotype was significantly higher than that of soybeans of the Hap3 genotype. Specifically, the average protein content of soybeans of the Hap1 genotype was 40.67%, that of soybeans of the Hap2 genotype was 39.71%, and that of soybeans of the Hap3 genotype was 37.61%. Therefore, in practical applications, the protein content of soybeans can be identified by detecting the genotype of the SNP locus to be tested.
[0095] Table 1
[0096]
[0097]
[0098] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.
Claims
1. The application of the substance for detecting the genotype of the soybean SNP locus to be tested in any one of the following (a1)-(a6): (a1) Identify the protein content of the soybean to be tested; (a2) Prepare a product for identifying the protein content of the soybean to be tested; (a3) Screening for soybean varieties with high protein content; (a4) Prepare products from soybean varieties with high protein content; (a5) Soybean breeding; (a6) Preparation of soybean breeding products; The SNP site is located at position 35708082 on chromosome 19 of the soybean genome; The reference genome version number of the soybean genome is Glycine max Wm82.a2.v1; The genotype is Hap1, Hap2, or Hap3. The Hap1 genotype is a homozygous type with C deoxyribonucleic acid at position 35708082 on chromosome 19 of the soybean genome. The Hap2 genotype is a homozygous type with deoxyribonucleic acid A at position 35708082 on chromosome 19 of the soybean genome. The Hap3 genotype is a heterozygous type with C and A deoxyribonucleotides at position 35708082 on chromosome 19 of the soybean genome; The protein content of the soybeans tested with the Hap1 genotype was higher than that of the soybeans tested with the Hap2 genotype; the protein content of the soybeans tested with the Hap2 genotype was higher than that of the soybeans tested with the Hap3 genotype. The purpose of the soybean breeding is to develop soybean varieties with high protein content.
2. The application according to claim 1, characterized in that: The substance used to detect the genotype of the soybean SNP locus is any one of the following (b1)-(b3): (b1) PCR primers for amplifying soybean genomic DNA fragments including the SNP sites; (b2) PCR reagents containing the PCR primers described in (b1); (b3) A kit containing the PCR primers described in (b1) or the PCR reagents described in (b2).
3. The application according to claim 2, characterized in that: The PCR primers consist of single-stranded DNA as shown in Sequence 1 and single-stranded DNA as shown in Sequence 2.
4. A method for identifying or assisting in the identification of the protein content of a soybean to be tested, the method comprising the following steps: detecting whether the genotype of the SNP locus of the soybean to be tested is Hap1, Hap2, or Hap3; and identifying the protein content of the soybean based on the genotype of the SNP locus of the soybean to be tested: the protein content of the soybean to be tested with the Hap1 genotype is higher than that of the soybean to be tested with the Hap2 genotype; the protein content of the soybean to be tested with the Hap2 genotype is higher than that of the soybean to be tested with the Hap3 genotype. The reference genome version number of the soybean genome is Glycine max Wm82.a2.v1; The Hap1 genotype is a homozygous type with C deoxyribonucleic acid at position 35708082 on chromosome 19 of the soybean genome. The Hap2 genotype is a homozygous type with deoxyribonucleic acid A at position 35708082 on chromosome 19 of the soybean genome. The Hap3 genotype is a heterozygous type with C and A deoxyribonucleotides at position 35708082 on chromosome 19 of the soybean genome.
5. A method for screening soybean varieties with high protein content, comprising the following steps: selecting soybean varieties with the Hap1 genotype; wherein the Hap1 genotype is a homozygous type with C as the deoxyribonucleic acid at position 35708082 on chromosome 19 of the soybean genome; and wherein the reference genome version number of the soybean genome is Glycine max Wm82.a2.v1.
6. A method for breeding high-protein soybeans, comprising the following steps: selecting soybean varieties with the Hap1 genotype as parents for breeding; wherein the Hap1 genotype is a homozygous type with deoxyribonucleotide C at position 35708082 on chromosome 19 of the soybean genome; and the reference genome version number of the soybean genome is Glycine max Wm82.a2.v1.