Molecular markers related to protein content in soybeans and their applications

By developing the InDel molecular marker WS205 and the QTL molecular marker WS185 on soybean chromosome 20, the problem of inconsistent QTL positioning of soybean protein content was solved, efficient screening and utilization of high-protein soybean resources were achieved, and the screening success rate of breeding materials was improved.

CN117265168BActive Publication Date: 2025-09-05JILIN ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311384424.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-09-05
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

The existing technology has inconsistent QTL positioning for soybean protein content, making it difficult to efficiently screen and utilize soybean resources with high protein content.

Method used

The InDel molecular marker WS205 and the QTL molecular marker WS185 located on soybean chromosome 20 were developed to predict soybean protein content, and HRM typing technology was used to assist in screening soybean breeding materials with high protein content.

Benefits of technology

The success rate of screening high-protein soybean materials has been improved, especially WS205 has a probability of 62.7%-74.5% to screen out materials with a protein content of more than 45%, and WS185 has a probability of 64.3%-73.0% to screen out materials with a protein content of more than 45%, which significantly improves breeding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117265168B_ABST
    Figure CN117265168B_ABST
Patent Text Reader

Abstract

The present invention provides molecular markers and applications related to soybean protein content, relating to the field of biotechnology. The molecular markers include the InDel molecular marker WS205 located on chromosome 20, and a QTL molecular marker encompassing the region between the InDel molecular markers WS185 and WS205. The insertion / deletion sequence of WS205 is shown as SEQ ID NO. 1, and the insertion / deletion sequence of WS185 is shown as SEQ ID NO. 2. The molecular markers are correlated with soybean protein content and can be used to predict soybean protein content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, in particular to a molecular marker related to protein content in soybeans and its application. Background Art

[0002] Soybean is the food crop with the highest protein content and is an important source of plant protein with high economic value. Therefore, increasing the protein content of soybean is an important research direction in soybean breeding.

[0003] Soybean protein content is a typical quantitative trait, influenced by environmental factors and gene-environment interactions. Therefore, discovering and utilizing QTLs (quantitative trait loci) for high protein content is an important approach to improving the efficiency of wild soybean breeding for high-protein soybeans. Currently, the Soybase database contains over 300 QTLs for soybean protein content, distributed across all 20 soybean chromosomes, of which 16 have been verified. There are also numerous reports on the mapping of QTLs for protein content in China, using parents such as Suinong 76, Shisheng Changye, and Zhonghuang 13, and the mapped QTLs vary. This discrepancy in QTLs indicates that protein content is controlled by multiple genes and reveals the genetic specificity of soybean resources. Therefore, conducting gene mapping research on high-protein wild soybean resources is a prerequisite for their efficient utilization. Providing more QTLs for high protein content will facilitate the development of high-protein soybean resources.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] The present invention aims to provide molecular markers related to soybean protein content, including InDel and QTL markers, which can predict soybean protein content. Based on the functions of these molecular markers, another object of the present invention is to provide applications of these molecular markers or substances for detecting them.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] In a first aspect, an InDel molecular marker related to the protein content in soybean is provided, including WS205, located on chromosome 20; the insertion / deletion sequence of WS205 is shown as SEQ ID NO.1.

[0008] In an optional embodiment, the WS205 is amplified by primers with nucleotide sequences as shown in SEQ ID NOs. 3 and 4, and the length of the amplified product is 189 bp or 174 bp.

[0009] In an optional embodiment, WS205 type A has an insertion of the sequence shown in SEQ ID NO. 1, and type B has a deletion of the sequence shown in SEQ ID NO. 1, and type B is a favorable genotype.

[0010] In the second aspect, a QTL molecular marker related to soybean protein content is provided, including the region between the InDel molecular markers WS185 and WS205; WS185 and WS205 are located on chromosome 20; the insertion / deletion sequence of WS205 is shown as SEQ ID NO.1, and the insertion / deletion sequence of WS185 is shown as SEQ ID NO.2.

[0011] In an optional embodiment, the QTL molecular marker is located at Chr20:20071086-23798747.

[0012] In a third aspect, the use of the InDel molecular marker of the first aspect, or a substance for detecting the InDel molecular marker of the first aspect, or the QTL molecular marker of the second aspect, or a substance for detecting the QTL molecular marker of the second aspect, in any of the following is provided:

[0013] (a) screening or assisting in screening breeding materials of soybeans with high protein content;

[0014] (b) preparing a product for screening or assisting in screening breeding materials of high protein soybeans;

[0015] (c) Used in soybean breeding;

[0016] (d) preparing products for soybean breeding.

[0017] In a fourth aspect, a kit for soybean breeding is provided, which comprises a substance for detecting the InDel molecular marker of the first aspect, or a substance for detecting the QTL molecular marker of the second aspect.

[0018] In a fifth aspect, a method for screening or assisting in screening breeding materials of high-protein soybeans is provided, the method comprising detecting the InDel molecular marker of the first aspect or the QTL molecular marker of the second aspect in the genome of the breeding material to be identified.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The present invention provides InDel molecular markers associated with soybean protein content, including WS205, located on chromosome 20, with an insertion / deletion sequence as shown in SEQ ID NO. 1. This molecular marker is associated with soybean protein content. Using a RIL population derived from "Suinong 14" x ZYD01015 as a sample, favorable genotypes of WS205 have a 62.7% to 74.5% and 30.4% to 47.2% probability of screening for materials with protein contents above 45% and 48%, respectively. Using randomly selected wild soybean resources as a sample, there is a 78.57% probability and a 50% probability of screening for lines with protein contents above 45% and 48%, respectively. The present invention also locates a protein content locus, qPRO20, located between molecular markers WS185 and WS205. These molecular markers can be used to screen or assist in screening for high-protein soybean breeding materials, as well as in soybean breeding, facilitating the development of high-protein soybean resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1a and Figure 1b The distribution of grain protein content of RIL populations in 2020 and 2021, respectively;

[0023] Figure 2a and Figure 2b The numbers of polymorphic variants between BSA samples in 2020 and 2021, respectively;

[0024] Figure 3 Schematic diagram of the associated segments identified by different algorithms;

[0025] Figure 4 This is the QTL map for soybean seed protein content;

[0026] Figure 5a and are the distribution diagrams of the genotype of the molecular marker WS185 of each sample and the corresponding protein content in the sample, where the red scatter points represent the 2021 samples and the blue scatter points represent the 2020 samples;

[0027] Figure 5b and are the distribution diagrams of the genotype of the molecular marker WS205 of each sample and the corresponding protein content in the sample, where the red scatter points represent the 2021 samples and the blue scatter points represent the 2020 samples;

[0028] Figure 6 This is a distribution map of protein content in wild soybean resources. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] The inventors conducted a protein content gene mapping study in wild soybean ZYD01015 using a population of RILs (recombinant inbred lines) (219 advanced families) derived from the cross of Suinong 14 / ZYD01015 (protein content 51.14%). The 2021 and 2022 populations showed significant genetic variation in grain protein content, which conformed to a normal distribution. High- and low-protein pools were constructed for the 2021 and 2022 populations, respectively. Association analysis using BSA-seq technology identified association intervals in soybean Chr20, with common intervals of Chr20: 3,980,000-6,310,000 and Chr20: 25,940,000-30,240,000. Molecular markers were developed based on polymorphic indel marker information from the Chr20:3,980,000-30,240,000 interval, and QTL mapping was performed. A protein content locus, qPRO20, was identified, located between the InDel markers WS185 and WS205, within the Chr20:20071086-23798747 segment. Verification of wild soybean resources revealed significant differences in protein content between alleles of the InDel marker WS205. The insertion sequence deletion is a favorable genotype for the InDel marker WS205, increasing the probability of higher protein content in soybeans and making it suitable for marker-assisted selection.

[0031] Based on the above findings, in a first aspect, a molecular marker indel associated with soybean protein content is provided. The marker includes WS205, located on chromosome 20 of the soybean genome. The insertion / deletion sequence of WS205 is shown in SEQ ID NO. 1. When the sequence shown in SEQ ID NO. 1 is inserted into the soybean genome, it is designated as WS205 type A; when the sequence shown in SEQ ID NO. 1 is deleted from the soybean genome, it is designated as WS205 type B, with type B being the favorable genotype. Using a RIL population derived from "Suinong 14" x ZYD01015 as a sample, the favorable genotypes of WS205 have a 62.7% to 74.5% and a 30.4% to 47.2% probability of screening for materials with protein contents above 45% and 48%, respectively. Using randomly selected wild soybean resources as a sample, the probability of screening for lines with protein contents above 45% and 48% was 78.57% and 50% respectively.

[0032] In an optional embodiment, the WS205 is amplified by primers with nucleotide sequences such as SEQ ID NO. 3 and 4. When the amplified product lacks the sequence shown in SEQ ID NO. 1, the amplified product length is 174 bp; when the amplified product inserts the sequence shown in SEQ ID NO. 2, the amplified product length is 189 bp.

[0033] In a second aspect, a QTL molecular marker related to soybean protein content is provided, which is denoted as qPRO20. The QTL molecular marker includes the region between InDel molecular markers WS185 and WS205.

[0034] WS205 is the InDel molecular marker described in the first aspect.

[0035] WS185 is located on chromosome 20 of the soybean genome. The insertion / deletion sequence of WS185 is shown in SEQ ID NO. 2. When the soybean genome contains a deletion of the sequence shown in SEQ ID NO. 2, it is designated WS185 type A; when the soybean genome contains an insertion of the sequence shown in SEQ ID NO. 2, it is designated WS185 type B. Using a RIL population derived from "Suinong 14" x ZYD01015 as a sample, favorable genotypes of WS185 have a 64.3% to 73.0% probability of screening for materials with protein contents above 45% and 48% with a 31.6% to 48.0% probability, respectively. WS185 can be amplified using primers with nucleotide sequences shown in SEQ ID NOs. 5 and 6. When the amplification product contains a deletion of the sequence shown in SEQ ID NO. 2, the amplification product is 150 bp in length; when the amplification product contains an insertion of the sequence shown in SEQ ID NO. 2, the amplification product is 189 bp in length.

[0036] The QTL molecular marker qPRO20 includes the region between the insertion / deletion site shown in SEQ ID NO. 1 and the insertion / deletion site shown in SEQ ID NO. 2. In an optional embodiment, the QTL molecular marker is located in the Chr20:20071086-23798747 segment, and the reference genome is Willisms82.a4.

[0037] In a third aspect, the InDel molecular marker of the first aspect, or a substance for detecting the InDel molecular marker of the first aspect, or the QTL molecular marker of the second aspect, or a substance for detecting the QTL molecular marker of the second aspect, is provided, for use in any of the following:

[0038] (a) screening or assisting in screening breeding materials of soybeans with high protein content;

[0039] (b) preparing a product for screening or assisting in screening breeding materials of high protein soybeans;

[0040] (c) Used in soybean breeding;

[0041] (d) preparing products for soybean breeding.

[0042] In this article, "auxiliary screening" refers to the process in which the intervention of a certain part of the steps in the entire screening process increases the probability of successful screening. Auxiliary screening does not guarantee that the target trait can be screened out or the target result can be obtained, but the intervention of the auxiliary screening step increases the probability of success. In an optional embodiment, in the method for screening breeding materials of soybeans with high protein content, adding the screening steps of the InDel molecular markers of the first aspect and / or the QTL molecular markers of the second aspect can increase the probability of obtaining breeding materials with protein contents above 45% and above 48%.

[0043] In an optional embodiment, soybean breeding may also optionally include the simultaneous screening of breeding materials with other excellent traits. The target traits include high protein content, and also include but are not limited to one or more combinations of weather resistance, stress resistance, cold resistance, drought resistance, and a specific range of plant height.

[0044] In an optional embodiment, the product for screening or assisting in screening breeding materials of high-protein soybeans and the product for soybean breeding can be, for example, but not limited to, a gene amplification kit, a gene detection chip, a library construction and / or sequencing kit, etc.

[0045] In an optional embodiment, the substance used to detect the InDel molecular marker of the first aspect; or, the substance used to detect the QTL molecular marker of the second aspect can be selected from any substance in the art that can detect gene fragments, including but not limited to amplification primers; probes for binding to regions containing the InDel molecular markers WS185 and / or WS205; reaction reagents for amplification reactions, specifically including but not limited to enzymes, buffer components, solvents, dNTPs, dyes, solid phase carriers, immune grade test papers, antibodies and markers, etc., in combination with one or more of the following.

[0046] In an optional embodiment, the material used to detect molecular markers includes primers for amplifying the target region. By amplifying the region containing the molecular InDel molecular marker or QTL molecular marker in the soybean genome, the type of molecular marker in the soybean genome can be determined, and further, whether the material to be identified has a favorable genotype can be determined. Primers can be designed according to general amplification principles in the art, such as, but not limited to, primers designed based on conventional PCR, primers and / or probes designed based on fluorescent PCR, primers and / or probes designed based on isothermal amplification such as LAMP, RPA, or RCA, etc., and the present invention is not limited thereto.

[0047] In an optional embodiment, the substance for detecting the InDel molecular marker includes a primer for amplifying WS205; and / or, the substance for detecting the QTL molecular marker includes a primer for amplifying WS205 and a primer for amplifying WS185.

[0048] In an optional embodiment, the primers used to amplify WS205 include a primer having a nucleotide sequence as shown in SEQ ID NO.3 and a primer as shown in SEQ ID NO.4.

[0049] In an optional embodiment, the primers used to amplify WS185 include a primer having a nucleotide sequence as shown in SEQ ID NO.5 and a primer as shown in SEQ ID NO.6.

[0050] In a fourth aspect, a kit for soybean breeding is also provided, which contains a substance for detecting the above-mentioned InDel molecular marker, or a substance for detecting the above-mentioned QTL molecular marker.

[0051] In an optional embodiment, the substance used to detect the InDel molecular marker of the first aspect; or the substance used to detect the QTL molecular marker of the second aspect can be selected from any substance in the art that can detect gene fragments, including but not limited to primers for amplifying the target fragment; probes for binding to the area to be tested; reaction reagents for amplification reactions, specifically including but not limited to enzymes, buffer components, solvents, dNTPs, dyes, solid phase carriers, immune grade test paper, antibodies and markers, etc., in combination with one or more of the above.

[0052] In an alternative embodiment, the substances used to detect InDel molecular markers include primers for amplifying WS205; and / or, the substances used to detect QTL molecular markers include primers for amplifying WS205 and primers for amplifying WS185. Primers can be designed according to general amplification principles in the art, such as, but not limited to, primers designed based on conventional PCR, primers and / or probes designed based on fluorescent PCR, or primers and / or probes designed based on isothermal amplification such as LAMP, RPA, or RCA, etc., and the present invention is not limited thereto.

[0053] In an optional embodiment, the primers used to amplify WS205 include a primer having a nucleotide sequence as shown in SEQ ID NO.3 and a primer as shown in SEQ ID NO.4.

[0054] In an optional embodiment, the primers used to amplify WS185 include a primer having a nucleotide sequence as shown in SEQ ID NO.5 and a primer as shown in SEQ ID NO.6.

[0055] In a fifth aspect, a method for screening or assisting in screening breeding materials of high-protein soybeans is also provided, the method comprising detecting the above-mentioned InDel molecular markers or the above-mentioned QTL molecular markers in the genome of the breeding material to be identified.

[0056] In an optional embodiment, the method includes amplifying the InDel molecular marker WS205.

[0057] In an optional embodiment, the method includes screening out breeding materials to be identified whose InDel molecular marker WS205 is type B, where WS205 type A has an insertion of the sequence shown in SEQ ID NO.1, and type B has a deletion of the sequence shown in SEQ ID NO.1;

[0058] In an alternative embodiment, primers with nucleotide sequences as shown in SEQ ID NOs. 3 and 4 are used to amplify WS205.

[0059] In an optional embodiment, the breeding materials in the third aspect and the fifth aspect are derived from the offspring of the hybridization of Suinong 14 and ZYD01015, or from wild soybeans.

[0060] The present invention is further described below by way of specific examples. However, it should be understood that these examples are merely provided for more detailed description and are not to be construed as limiting the present invention in any form.

[0061] Example 1

[0062] (1) Experimental materials:

[0063] The RIL population (219 F 10 The seedlings (100 strains, 100 strains, and 100 strains) were used as test materials and planted in pots at the Gongzhuling Experimental Base of the Jilin Academy of Agricultural Sciences in 2020 and 2021. Each seedling was planted in one pot, with three plants per pot, and individual plants were harvested. Harvested seeds were cleaned of impurities, and mature, plump, disease-free seeds (approximately 10.0 g) were selected. These seeds were pulverized using a tissue grinder, sieved through a 40-mesh sieve, and stored in sealed bags at -20°C for DNA extraction and protein content determination. These were used for BSA-seq analysis to locate protein content QTLs and analyze the efficiency of marker-assisted screening.

[0064] Based on the protein content of the RIL populations planted and harvested in 2020 and 2021, 30 families with the lowest protein content and 30 families with the highest protein content were selected, and the DNA of these families was mixed in equal amounts to construct BSA extreme pools (two BSA extreme pools in 2020 and 2021, respectively, marked as 2020H, 2020L, 2021H, and 2020L). Together with the parents Suinong 14 and ZYD01015, a total of 6 BSA samples (2020H, 2020L, 2021H, and 2020L) were used for BSA-seq analysis. 219 F 10 The family and its parents were used for genotyping of molecular markers WS185 and WS205, and the efficiency of molecular marker-assisted screening was analyzed in combination with their protein content.

[0065] (2) Experimental methods:

[0066] 1. Grain Protein Content Determination: Population crude protein content was determined according to the national standard GB / T 6432-2018 Kjeldahl method. A FOSS8400 fully automatic Kjeldahl nitrogen analyzer was used. Two replicates were used.

[0067] The protein content of grains is expressed as dry basis protein content:

[0068] Dry basis protein content (%) = wet basis protein content / (1-moisture content)×100%.

[0069] 2. Determination of moisture content: Referring to the National Standard GB 5497-85 Grain and Oilseeds - Determination of Moisture, the moisture content of the test soybean seeds was determined by the constant weight method at 105°C with 2 replicates.

[0070] 3. DNA Extraction and DNA Testing: Soy flour genomic DNA was extracted using the Kangwei Plant Tissue DNA Extraction Kit (Model CW0531M). Refer to the kit's instructions for specific procedures. Take 1 μL of the DNA solution and test the DNA concentration and quality using a NanoDrop 2000c.

[0071] 4. BSA-seq and analysis: A total of 6 DNA samples from the extreme pool and parents were sequenced by Beijing Biomark Biotechnology Co., Ltd. using Illumina HiSeq. The extreme pool depth sequencing was set to 30×, and the parent depth sequencing was set to 20×. The original sequencing sequences were quality controlled to obtain high-quality sequences for subsequent alignment, polymorphic site analysis, and association analysis. Using Willisms82.a4 as the reference genome, bwa software was used to locate the sequences, and the sequencing depth, genome coverage, and other information of the samples were statistically analyzed. GATK software was used for variant detection; SnpEff software was used to annotate and predict variants (SNPs and InDels). The Euclidean distance (ED) algorithm and Index method were used for association analysis.

[0072] 5. Marker Development: Based on the InDel (insertion-deletion) marker information, forward and reverse primers were designed on either side of the sequence. Paternal, maternal, and heterozygous (equal amounts of paternal and maternal DNA mixed) samples were used as controls. HRM (High Resolution Melting Analysis) was used to detect the alleles of the samples (PCR amplified fragment lengths ranged from 100-500 bp, with PCR fragment size differences of at least 10 bp between allelic types). Alleles whose melting curves were consistent with the maternal genotype were designated A, those whose melting curves were consistent with the paternal genotype were designated B, and those whose melting curves were consistent with the heterozygous control were designated H. Information on the 12 molecular markers is shown in Table 1.

[0073] Table 1 Information of the 12 developed molecular markers

[0074]

[0075]

[0076] HRM typing was performed on a fluorescence quantitative PCR instrument. A 10-μl PCR reaction system contained 5 μl of 2× AugeGreen Mix for HRM (S2013L), 0.4 μM forward and reverse primers, and 40 ng of DNA template. The reaction procedure included a 2-min incubation at 50°C, a 2-min initial denaturation at 94°C, and 40 cycles of 95°C denaturation for 5 s, 10 s annealing at 58°C, and a 30-s extension at 72°C. Melting curve analysis was then performed using the following steps: denaturation at 95°C for 15 s, annealing at 55°C for 15 s, and then ramping from 55°C to 95°C at a rate of 0.2°C per second. Ten signals were acquired for every 1°C increase, and the collected signals were analyzed using Eco software. Three genotype controls were included in each experimental batch: maternal, paternal, and mixed (equal amounts of paternal and maternal DNA mixed) samples.

[0077] 6. Data Analysis: Genetic variation of protein content was analyzed using SPSS Statistics 17.0. Gene mapping was performed using QTL IciMapping 4.1.

[0078] (3) Results and analysis

[0079] 1. Genetic variation of grain protein content in RIL populations:

[0080] 219 families of the RIL population were planted in pots in Gongzhuling City, Jilin Province in 2020 and 2021. The protein content of their grains was determined by the Kjeldahl method after harvest. The protein content of the parent Suinong 14 was 38.61%, and the protein content of ZYD01015 was 51.57%. The protein content of the 219 families in the RIL population varied greatly ( Figure 1a and Figure 1b ), the mean protein content of the RIL population in 2020 and 2021 was 45.04% and 45.98%, respectively, with a range of 32.58% to 54.18% and 34.93% to 53.05%, respectively, and a standard deviation of 3.79 and 3.54, respectively; the Kolmogorov-Smirnov test showed that the protein content of the population was normally distributed, with P values ​​of 0.916 and 0.776, respectively. The protein content of some families was lower than 38.61% or higher than 51.57%, showing a super-parent separation phenomenon ( Figure 1a and Figure 1b Although the T test found that the mean protein content of the two groups was significantly different (P = 0.000423), there was a significant correlation between the protein content of the groups, with a correlation coefficient of 0.486 (P < 0.05).

[0081] 2.BSA association analysis:

[0082] 2.1 Sequencing analysis:

[0083] Based on the protein content of RIL populations planted and harvested in 2020 and 2021, four extreme protein pools (labeled 2020H, 2020L, 2021H, and 2020L) were constructed. Together with the parents, Suinong 14 (P1) and ZYD01015 (P2), a total of six BSA sequencing samples were obtained (Table 2). A total of 959 gigabytes of clean reads were obtained. Alignment with the soybean reference genome, glyma.Wm82.gnm4, revealed that the genome coverage of the parents was lower, at 98.54% and 97.26%, respectively, while that of the extreme pool was higher, reaching over 99.28%. The sequencing depth of parents P1 and P2 was 21× and 23×, respectively, while the extreme pool had a sequencing depth of over 37×.

[0084] Allelic variation (SNP and InDel) detection of 6 samples revealed that there were 3,470,529 SNPs and 683,079 InDels between the parents, 518,989 SNPs and 127,171 InDels between the two extreme pools 2020H and 2020L, and a total of 382,961 SNPs and 71,878 InDels were detected between the two extreme pools 2021H and 2021L, and 419,600 SNPs and 107,911 InDels were detected between the two extreme pools 2021H and 2021L, and 400,312 SNPs and 75,319 InDels were detected between the two parents. Figure 2a and Figure 2b ).

[0085] Table 2. Data volume and quality of mixed pool and parental sequencing data

[0086] sample Total clean bases Q30(%) depth(%) Coverage (%) 2020H 250,062,892 95.74 37 99.28 2020L 254,385,888 95.79 37 99.52 2021H 273,070,170 95.4 40 99.42 2021L 296,459,240 96.02 44 99.52 P1 146,166,070 95.38 21 98.54 P2 158,090,724 94.71 23 97.26

[0087] 2.2 Correlation Analysis:

[0088] The polymorphic variation data of two groups of BSA samples (2020: 2020H, 2020L and parents; 2021: 2020H, 2020L and parents) were screened for high-quality SNPs and InDels. Then, the ED algorithm and SNP-index (or InDel-index) algorithm were used for association analysis of protein content. Different associated segments were detected, all on chromosome 20 (Chr20) of the soybean genome (Table 2, Figure 3), among which different algorithms for the 2020 group of samples all detected a larger association segment, and the intersection area was Chr20:3100000-33360000. For the 2021 group of samples, the SNP-ED, SNP-index and InDel-ED algorithms respectively detected a larger association area, while the InDel-index algorithm identified 10 association segments, which were distributed in the larger association segments detected by other algorithms, concentrated in the two regions of Chr20:3,980,000-6,310,000 and Chr20:25,940,000-30,240,00.

[0089] From the above, we can see that the protein content BSA extreme pool established using RIL populations planted in different years identified the intersection of the associated segments as Chr20:3,980,000-6,310,000 and Chr20:25,940,000-30,240,00, which contain QTLs controlling soybean protein content.

[0090] Table 2 Protein content associated region information

[0091]

[0092]

[0093] 3. Protein content QTL mapping:

[0094] For the associated segments Chr20:3,980,000-6,310,000 and Chr20:25,940,000-30,240,00, molecular markers based on HRM typing technology were developed based on the InDel information between BSA samples in the Chr20:3,980,000-30,240,00 segment (Table 1). 219 families of the RIL population planted in 2020 and 2021 were used as experimental materials. The grain protein content was measured and gene mapping was performed. Although the positions of the protein content QTLs were slightly different between the two years, they were all located in the WS185-WS205 interval ( Figure 4 ), the QTL contribution rates were 14.00% and 13.26%, the additive effects were -1.42 and -1.31, and the favorable allele for protein content came from the paternal parent ZYD01015.

[0095] 4. Molecular marker-assisted screening of high protein content germplasm:

[0096] Among the 219 RIL lines tested, at the molecular markers WS185 and WS205 at both ends of the QTL, the protein content of the lines with the paternal ZYD01015 allele (B) was significantly higher than that of the lines with the maternal Suinong 14 allele (A) (P < 0.01) (Table 3, Figure 5a and Figure 5b ).

[0097] At the WS185 locus, for the population planted in 2020, there were 102 varieties with the A allele variation, and their average protein content was 43.66%, with a range of 32.58% to 51.77%. There were 98 varieties with the B allele variation, and their average protein content was significantly higher than that of the A allele variation, at 46.46%, with a range of 38.52% to 54.18%. Among them, 63 varieties had a protein content of more than 45%, and 31 varieties had a protein content of more than 48%, that is, there was a 64.3% probability of screening varieties with a protein content of more than 45%, and a 31.6% probability of screening varieties with a protein content of more than 48%. For the 2021 planting population, 108 accessions harboring the A allele had an average protein content of 44.86%, ranging from 34.93% to 51.41%. Meanwhile, 100 accessions harboring the B allele had a significantly higher average protein content of 47.32%, ranging from 38.73% to 53.05%. Of these, 73 accessions had protein contents above 45%, and 48 had protein contents above 48%. This means that there is a 73.0% probability of screening for accessions with protein contents above 45%, and a 48.0% probability of screening for accessions with protein contents above 48%. Therefore, WS185 can be used for molecular marker-assisted screening of high-protein germplasm with the ZYD01015 genotype, with a probability of 64.3% to 73.0% for screening accessions with protein contents above 45%, and a probability of 31.6% to 48.0% for screening accessions with protein contents above 48%.

[0098] At the WS205 locus, for the population planted in 2020, there were 100 varieties with the A allele variation, and their average protein content was 43.78%, with a range of 35.82% to 51.77%. There were 102 varieties with the B allele variation, and their average protein content was significantly higher than that of the A allele variation, at 46.40%, with a range of 38.52% to 54.18%. Among them, 64 varieties had a protein content of more than 45%, and 31 varieties had a protein content of more than 48%, that is, there was a 62.7% probability of screening varieties with a protein content of more than 45%, and a 30.4% probability of screening varieties with a protein content of more than 48%. For the 2021 planting population, 106 accessions harboring the A allele had an average protein content of 44.72%, ranging from 34.93% to 51.41%. There were also 106 accessions harboring the B allele, with a significantly higher average protein content of 47.26%, ranging from 37.25% to 53.05%. Of these, 79 accessions had protein contents above 45%, and 50 had protein contents above 48%. This means there is a 74.5% probability of screening for accessions with protein contents above 45%, and a 47.2% probability of screening for accessions with protein contents above 48%. Therefore, WS205 can be used for molecular marker-assisted screening of high-protein germplasm with the ZYD01015 genotype, with a probability of 62.7% to 74.5% for screening accessions with protein contents above 45%, and a probability of 30.4% to 47.2% for screening accessions with protein contents above 48%.

[0099] Table 3 Distribution of different allelic variants in different protein content ranges

[0100]

[0101] Example 2

[0102] (I) Experimental Materials: 79 accessions of 194 wild soybean samples from Heilongjiang, Jilin, and Liaoning provinces were randomly selected for genotyping at the WS205 locus and analysis of germplasm protein content. These accessions were potted in Gongzhuling in 2022, and individual seeds were harvested for protein content determination and inter-genotype identification.

[0103] (II) Experimental methods: The identification and protein content determination methods and statistical analysis between different genotypes were the same as those in Example 1.

[0104] (3) Results and analysis

[0105] 1. Genetic variation of grain protein content in 79 wild soybean accessions:

[0106] The distribution of protein content among 79 wild soybean germplasms is shown in the figure below. Figure 6The protein content of the germplasms was 38, 45% or less, 18, 45% or less, and 23, 48% or more.

[0107] 2. Differences in grain protein content between germplasms with different genotypes:

[0108] The protein content statistics of germplasms with different alleles at the WS205 locus are shown in Table 4, which have the same distribution trend as the results in Example 1. There are 65 germplasms with the A allele variation, with an average protein content of 44.53% and a range of 34.97% to 53.63%. There are 14 lines with the B allele variation, with an average protein content significantly higher than that of the A allele variation, at 46.76% and a range of 38.86% to 50.47%. Among them, 11 lines have a protein content of more than 45% and 7 have a protein content of more than 48%. In other words, there is a 78.57% probability of screening for lines with a protein content of more than 45% and a 50% probability of screening for lines with a protein content of more than 48%.

[0109] Table 4 Distribution of different allelic variants in different protein content ranges

[0110]

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. Use of an InDel molecular marker, or a substance for detecting the InDel molecular marker, in any of the following: (a) Screening or assisting in the selection of breeding materials for soybeans with high protein content; (b) preparing products for screening or assisting in screening breeding materials of soybeans with high protein content; The InDel molecular marker includes WS205, located on chromosome 20; the insertion / deletion sequence of WS205 is shown in SEQ ID NO.1; WS205 type A has an insertion of the sequence shown in SEQ ID NO.1, and type B has a deletion of the sequence shown in SEQ ID NO.1, and type B is a favorable genotype.

2. The use according to claim 1, characterized in that The substances used to detect the InDel molecular marker include primers for amplifying WS205.

3. The use according to claim 2, characterized in that The primers used to amplify WS205 include a primer having a nucleotide sequence as shown in SEQ ID NO. 3 and a primer having a nucleotide sequence as shown in SEQ ID NO.

4.

4. Use of a QTL molecular marker, or a substance for detecting the QTL molecular marker, in any of the following: (a) Screening or assisting in the selection of breeding materials for soybeans with high protein content; (b) preparing products for screening or assisting in screening breeding materials of soybeans with high protein content; The QTL molecular marker is the region between the InDel molecular markers WS185 and WS205; WS185 and WS205 are located on chromosome 20; the insertion / deletion sequence of WS205 is shown as SEQ ID NO.1, and the insertion / deletion sequence of WS185 is shown as SEQ ID NO.2; the substance used to detect the QTL molecular marker includes primers for amplifying WS205 and primers for amplifying WS185.

5. The use according to claim 4, characterized in that The primers used to amplify WS205 include a primer having a nucleotide sequence as shown in SEQ ID NO. 3 and a primer having a nucleotide sequence as shown in SEQ ID NO.

4.

6. The use according to claim 4, characterized in that The primers used to amplify WS185 include a primer having a nucleotide sequence as shown in SEQ ID NO.5 and a primer as shown in SEQ ID NO.

6.

7. A method for screening or assisting in screening breeding materials of high-protein soybeans, characterized in that: The method comprises detecting the InDel molecular marker described in any one of claims 1 to 3, or the QTL molecular marker described in any one of claims 4 to 6, in the genome of the breeding material to be identified; The method comprises screening out breeding materials to be identified whose InDel molecular marker WS205 is type B, wherein WS205 type A has an insertion of the sequence shown in SEQ ID NO.1, and type B has a deletion of the sequence shown in SEQ ID NO.

1.

8. The method according to claim 7, characterized in that Includes amplification of the InDel molecular marker WS205.

9. The method according to claim 8, wherein WS205 was amplified using primers whose nucleotide sequences are shown in SEQ ID NOs. 3 and 4.