Co-localized quantitative trait loci related to soybean isoflavone and protein content and their applications
By co-localizing quantitative trait loci on soybean chromosomes 6, 8, and 9, the problem of difficult co-localization of soybean isoflavones and protein content was solved, and the synergistic increase of isoflavone and protein content in soybean varieties was achieved to meet the nutritional needs of specific foods.
Patent Information
- Application Number
- CN202210821379.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-07-13
AI Technical Summary
Existing technologies make it difficult to effectively co-locate the quantitative trait loci for soybean isoflavone and protein content, making it difficult to breed soybean varieties that are rich in both isoflavones and protein.
By identifying and co-localizing the quantitative trait loci qISO6.2, qPC6.1, qISO8.1, qPC8, qISO9.1, and qISO9.2 on soybean chromosomes 6, 8, and 9, synergistically associated molecular markers were developed for soybean breeding to improve isoflavone and protein content.
The synergistic increase of isoflavone and protein content in soybean varieties is achieved, and new soybean varieties for preparing specific foods are provided to meet the nutritional needs of different groups.
Smart Images

Figure CN115094159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of molecular biology technology, in particular to a molecular breeding technology related to soybean isoflavone and protein content. Background Art
[0002] Soybean (Glycine max L. Merr.) is a nutritious crop that produces seeds rich in protein and oil. In 2020, protein and oil accounted for approximately 70% of soybean meal and 28% of vegetable oil consumed by humans and livestock. In recent years, with the increasing demand for food and feed, soybean cultivation and production have increased significantly. In 2020, soybean cultivation exceeded 127 million hectares, primarily in Brazil, the United States, Argentina, and China (FAO data), while global soybean production reached 399 million tons, making soybean the fourth most cultivated crop in terms of area and production, after wheat, maize, and rice. Therefore, releasing high-quality soybean varieties rich in protein and / or oil will remain a crucial goal for soybean breeders worldwide for the foreseeable future. In addition to protein and oil, soybeans also produce significant amounts of isoflavones, compounds that contribute to soybean's adaptation to diverse environments and to maintaining human health and well-being. For example, recent studies suggest that isoflavones may reduce the risk of cancer and help prevent the onset of several chronic diseases. Therefore, there is a real and urgent need to breed soybean varieties rich in isoflavones and capable of producing high protein or oil content to meet the needs of different populations. Although soybean seeds may contain high concentrations of isoflavones, there is wide variation in this trait among soybean germplasm. For example, isoflavone concentrations among 1,168 soybean varieties were found to vary from approximately 700 μg g⁻¹ to 5,000 μg g⁻¹, and other studies have also found relatively high heritability of isoflavone concentrations among soybean varieties. Overall, previous results suggest that efforts to breed or genetically engineer soybean varieties that consistently produce high concentrations of isoflavones have a reasonable chance of success.
[0003] Because isoflavones are synthesized via the amino acid (phenylpropanoid) pathway, potential relationships between isoflavone production and protein or oil production have been investigated, although published reports have generated uncertainty due to significant disagreements in the results. One group of researchers found that isoflavone synthesis was negatively correlated with protein production and positively correlated with oil content. Others found that isoflavone content was negatively correlated with protein levels and had no significant correlation with oil production. In other work, decreased protein content was associated with increased isoflavone production in soybean LOX near-isogenic lines, suggesting that the isoflavone and protein synthesis pathways may share common genes or that these pathways include genes with pleiotropic effects. In short, all previous reports have consistently found correlations between protein and isoflavone production and unclear associations between isoflavone and oil synthesis, although the genetic components that may be involved in protein, oil, and isoflavone production remain to be determined.
[0004] Quantitative trait mapping, a method for studying the genetic basis of quantitative traits using linkage maps constructed using molecular markers, is widely used in cereal crops such as rice, maize, and wheat, as well as legumes. Isoflavone, protein, and oil content are complex quantitative traits expected to be controlled by multiple loci that exhibit highly flexible responses to environmental conditions. Previous research has focused on mapping markers for protein and oil production. To date, multiple QTLs associated with protein synthesis and 305 QTLs associated with oil production have been detected in soybean by comparing biparental populations. Studies targeting genetic regions involved in isoflavone production in soybean began later than those investigating protein and oil production, but have nonetheless yielded informative results. The first QTL associated with isoflavone synthesis was reported in 1999. To date, 87 QTLs associated with isoflavone production have been detected in biparental populations of soybean. Many major and stable QTLs associated with isoflavone production have been identified and validated in various soybean populations, and several candidate isoflavone pathway genes have also been characterized. Despite the abundance of relevant results, few studies have investigated the potential genetic colocalization of protein or oil production markers with isoflavone synthesis markers. Some protein and isoflavone synthesis markers have been colocalized, although no shared fragments were identified. Other researchers have localized protein and oil markers, as well as markers for isoflavone production, but no colocalization was detected. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a co-localized quantitative trait locus related to soybean isoflavones and protein content and an application thereof.
[0006] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows.
[0007] Co-localized quantitative trait loci related to soybean isoflavone and protein content, the co-localized quantitative trait genes include three pairs of quantitative trait loci located on soybean chromosomes 6, 8, and 9, and any combination thereof, the three pairs of quantitative trait loci being: qISO6.2 and qPC6.1 co-localized on chromosome 6; qISO8.1 and qPC8 co-localized on chromosome 8; and qISO9.1 and qISO9.2 co-localized on chromosome 9; wherein:
[0008] qISO6.2 is located on soybean chromosome 6, with a genetic distance of 213.609-225.650 cM and a physical locus / interval of 06_18449510-06_21098994;
[0009] qPC6.1 is located on soybean chromosome 6, with a genetic distance of 224.401-224.408 cM and a physical locus / interval of 06_19395795-06_20312314;
[0010] qISO8.1 is located on soybean chromosome 8, with a genetic distance of 95.815-96.096 cM and a physical locus / interval of 08_9020859-08_9054795;
[0011] qPC8 is located on soybean chromosome 8, with a genetic distance of 80.663-103.272 cM and a physical locus / interval of 08_7270752-08_9502316;
[0012] qISO9.1 is located on soybean chromosome 9, with a genetic distance of 31.954 cM and a physical locus / interval of 09_2556374;
[0013] qISO9.2 is located on soybean chromosome 9, with a genetic distance of 46.908-48.224 cM and a physical locus / interval of 09_3204462-09_3672384.
[0014] As a preferred technical solution of the present invention, the qISO6.2, qISO8.1, and qISO9.1 are related to the soybean isoflavone content trait, and the qPC6.1, qPC8, and qISO9.2 are related to the soybean protein content trait.
[0015] The co-localized quantitative trait loci are used to develop molecular markers and primers thereof that are simultaneously correlated with the isoflavone content and protein content of soybeans.
[0016] The co-localized quantitative trait loci are used for soybean protein and isoflavone coordinated molecular marker-assisted breeding.
[0017] The co-localized quantitative trait loci are used for molecular marker-assisted breeding of new soybean varieties with low protein content and high isoflavone content.
[0018] The use of the co-localized quantitative trait loci and the developed new soybean varieties are used to prepare elderly food related to cardiovascular and cerebrovascular diseases.
[0019] The co-localized quantitative trait loci are used for molecular marker-assisted breeding of new soybean varieties with high protein content and low isoflavone content.
[0020] The use of the co-localized quantitative trait loci and the developed new soybean varieties are used for preparing infant and children's food.
[0021] A soybean protein and isoflavone synergistic molecular marker-assisted breeding method based on co-located loci, wherein the soybean molecular marker-assisted breeding is aimed at the isoflavone content trait of soybean, and the cultivated soybean variety JD12 ZDD23040 is selected and crossed with the wild soybean variety Y9 ZYD02739, and a control population consisting of 185 F8-derived recombinant inbred lines (RILs) is constructed using the single seed descent (SSD) method; the parental genotypes are planted in 6 replicates, and the RILs are planted in 3 replicates, and isoflavone, protein, and oil are determined after maturity; then genotyping is performed by sequencing and SNP calling, and DNA is isolated from soybean seedling leaf tissue using the cetyltrimethylammonium bromide (CTAB) method, the collected DNA samples are randomly cut, and a library is constructed and sequenced on the platform for SNP detection, with the minimum SNP length set to 4 bp; then map construction and QTL detection are performed, and SNP data are analyzed in the chi-square test, with the P value threshold set to 0.05 to improve the map quality, and the logarithm of the ratio LOD threshold is set to 2.5 for QTL detection, and A permutation test was performed with a threshold P value set at 0.05 to estimate the LOD significance of 1000 iterations; then, an integrated QTL analysis was performed to construct a composite genetic map based on the QTL, all molecular markers present on the shared genetic map, and all coordinates of the consensus map markers, and significant QTLs associated with the target trait and adjacent markers located on the same position segment were estimated based on the physical distance separating the related characteristics to construct a genetic map and a consensus map; finally, based on statistical analysis, QTLs such as major effect QTLs / clustered QTLs / hotspot QTLs / preferred QTLs that were synergistically related to soybean isoflavone content and protein content were obtained; based on the obtained QTLs, molecular markers and primers were developed for molecular marker-assisted breeding that was synergistically related to soybean protein content and isoflavone content.
[0022] The beneficial effects of adopting the above technical solution are: the research of the present invention uses a high-density genetic linkage map composed of 3943 SNP markers to analyze the genetic basis of soybean isoflavone and / or protein (and / or oil) production, the co-localization of isoflavone markers and / or protein (and / or oil) markers is identified under this marker density, and the verification of the co-localized markers can be applied to the breeding of excellent soybean varieties rich in isoflavones and / or protein (and / or oil). BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Comparison of quality and isoflavone traits between JD12 and Y9. (A) Photographs of JD12 and Y9 seeds; (B) Protein content; (C) Oil content; (D) Content of daidzein isoflavone derivatives (Ds); (E) Content of genistein isoflavone derivatives (Gs); (F) Content of glycidyl isoflavone derivatives (GLs); (G) Total isoflavone content (TIF). Bars represent the mean ± standard deviation of six replicates. Asterisks indicate significant differences between JD12 and J9 at the 0.05 (*), 0.01 (**), and 0.001 (***) significance levels as determined by Student's t-test.
[0024] Figure 2 . Correlation analysis of isoflavone, protein, and oil contents. Histograms of trait distributions are located in the diagonal cells. Correlation coefficients with significance levels are shown above the diagonal, while scatter plots with fitted curves are shown below the diagonal. Red asterisks indicate the significance of RIL differences obtained by Student's t-test at the 0.05 (*), 0.01 (**), and 0.001 (***) significance levels.
[0025] Figure 3 . The isoflavone and protein loci identified in the current population are co-localized with loci identified in previous studies. Colors and bold letters represent the identified loci. Red and black boxes represent loci for isoflavone and protein traits, respectively. The projected areas are highlighted in the corresponding colors. Bold markers of different colors indicate the boundaries of the corresponding loci, and green markers indicate isoflavone and protein traits. The consensus map shows the isoflavone or protein production loci previously identified based on soybase data.
[0026] Figure 4. Effect of combining favorable loci from qISO5 and qISO6.2 on total isoflavone (TIF), protein, and oil content in RIL populations, as described in Materials and Methods. Letters indicate the significance of differences at the 0.01 significance level obtained by LSD-test. Different shapes represent different oil contents, and different colors represent different protein contents. n represents lines in which no favorable loci were detected, while q5 and q6.2 represent lines with favorable loci qISO5 and qISO6.2, respectively. HP and HO indicate that the respective protein and oil contents in the RIL population are above the average. LP and LO indicate that the corresponding protein and oil contents in the RIL population are below the average.
[0027] Figure 5 Correlation analysis of soybean isoflavone, protein, and oil traits in 2020. A histogram with a characteristic fitted curve is placed on the diagonal. Correlation coefficients with significant levels are above the diagonal, and a scatter plot with the fitted curve is below the diagonal. Red asterisks indicate the significance of RIL differences at the 5% (*), 1% (*), and 0.1% (**) levels as determined by Student's t-test.
[0028] Figure 6 Correlation analysis of soybean isoflavone, protein, and oil traits in 2021. Histograms of the trait fitting curves are placed on the diagonal. Correlation coefficients with significant levels are above the diagonal, and scatter plots with fitting curves are below the diagonal. Red asterisks indicate significance of differences between RILs at the 5% (*), 1% (*), and 0.1% (**) levels as determined by Student's t-test.
[0029] Figure 7 Correlation analysis between individual isoflavones and protein / oil traits. Histograms of the trait fitting curves are placed on the diagonal. Above the diagonal are correlation coefficients with significant levels, and below the diagonal are scatter plots with fitting curves. Red asterisks indicate significance of differences between RILs at the 5% (*), 1% (*), and 0.1% (**) levels as determined by Student's t-test.
[0030] Figure 8 Correlation analysis between soy isoflavones and protein / oil traits in 2020. Histograms of the trait fitting curves are placed on the diagonal. Correlation coefficients with significant levels are above the diagonal, and scatter plots with the fitting curves are below the diagonal. Red asterisks indicate significance of differences between RILs at the 5% (*), 1% (*), and 0.1% (**) levels as determined by Student's t-test.
[0031] Figure 9Correlation analysis between soybean isoflavones and quality traits in 2021. Histograms of trait fitting curves are placed on the diagonal. Correlation coefficients with significant levels are above the diagonal, and scatter plots with fitting curves are below the diagonal. Red asterisks indicate significance of differences between RILs at the 5% (*), 1% (*), and 0.1% (**) levels as determined by Student's t-test.
[0032] Figure 10 . 20 linkages of the soybean high-density genetic map. The genetic distance scale is on the right.
[0033] Figure 11 Correlation analysis between isoflavone, protein, and oil traits, eliminating the influence of colocalized fragments on chromosomes 6, 8, and 9 in 2020. The histogram of the trait fitting curves is placed on the diagonal. Above the diagonal are the correlation coefficients with significant levels, and below the diagonal are the scatter plots with the fitting curves. Red asterisks indicate the significance of the differences between RILs at the 5% (*), 1% (*), and 0.1% (**) levels using the Student's t-test. DETAILED DESCRIPTION
[0034] The present invention is described in detail in the following examples. The various raw materials and equipment used in the present invention are conventional commercial products and can be directly purchased from the market.
[0035] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, including, and including, those combinations. References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of the present application include the particular features, structures, or characteristics described in connection with that embodiment. Thus, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in yet other embodiments," etc., appearing in various places in this specification, do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0036] Example 1. Plant materials and growth conditions
[0037] In this study, the cultivated soybean variety JD12 (ZDD23040), which produces relatively low amounts of isoflavones and protein, was crossed with the wild soybean variety Y9 (ZYD02739), which produces relatively high amounts of isoflavones and protein. A control population consisting of 185 F8-derived recombinant inbred lines (RILs) was constructed using the single seed descent (SSD) method. The parents and population were planted in 2020 and 2021 at the Shangdi Experimental Field (E114.48, N38.03) of the Institute of Cereals and Oils Crops, Hebei Academy of Agricultural and Forestry Sciences. The soil at this site is fluvial-type. The average temperature during the soybean growing season is traditionally around 22 degrees Celsius, and the average rainfall is around 500 milliliters. The physicochemical properties of the 25 cm soil surface layer were as follows: pH 8.3; organic matter 19.9 g kg-1; available N, P (Olsen-P), and K 109.2 mg kg-1, 21.6 mg kg-1, and 193.4 mg kg-1, respectively. These field trials used a randomized complete block design. The parental genotypes were planted in six replicates, and the RILs were planted in three replicates. Each plot contained three 2 m long rows spaced 0.5 m apart. Six plants were planted in each row. Seeds were harvested from each plot at maturity and dried at ambient temperature for subsequent isoflavone, protein, and oil determinations.
[0038] Example 2: Extraction and Quantification of Isoflavones
[0039] Approximately 10 g of dry soybean seeds were selected and inspected from each plot to ensure they were free from pests and diseases. Then, they were ground using a cyclone mill (CT 293 Cyclotec, FOSS, Denmark) and sieved through an 80-mesh filter. The method for extracting isoflavones described above was slightly modified. 20 mg (0.01 mg) of soybean powder was dissolved in 1 mL of 70% (v / v) ethanol and 0.1% (v / v) acetic acid in a 1.5 mL plastic tube. The mixture was shaken on an incubation shaker (INNOVA 42, Eppendorf, Germany) at 28 °C and 200 rpm for 12 hours. Then, the mixture in the test tube was centrifuged at 12,000 rpm for 10 minutes, and the supernatant (0.7 mL) was filtered through a 0.45-μm nylon syringe filter (Jinteng, Tianjin, China). Before analyzing the isoflavones, the filtered samples were refrigerated at 4 °C. According to the published method, high-performance liquid chromatography (HPLC) was used to identify and quantify the isoflavones. The chromatography was performed on an Agilent 1260 HPLC system (Agilent 1260 HPLC system, Santa Clara, CA, USA), maintaining a 70-minute linear gradient of 13 - 35% acetonitrile (v / v) at 35 °C, with a solvent flow rate of 1.0 mL·min-1 and an injection volume of 20 μL. The chromatographic column used was a YMC-Pack ODS-AM-303 column (inner diameter 250 mm × 4.6 mm, S-5 μm, 120, YMC Kyoto Co., Ltd.), and isoflavones were detected by ultraviolet absorption at 260 nm with mobile phases A and B.
[0040] Twelve isoflavone standards were provided by the Chinese Academy of Agricultural Sciences. These isoflavones include daidzein (DE), daidzin (D), malonyldaidzein (MD), acetyl daidzein (AD), genistein (GE), genistin (G), malonyl genistein (MG), acetyl genistein (AG), glycitein (GLE), glycitin (GL), malonyl glycitin (MGL), and acetyl glycitin (AGL). Isoflavones were classified according to the aglycones they contained and were defined as follows: Daidzein derivatives (Ds) were defined as the sum of De, D, MD, and AD; Genistein derivatives (Gs) were defined as the sum of GE, G, MG, and AG; Glycitein derivatives (GLs) were defined as the sum of GLE, GL, MGL, and AGL; and total isoflavones (TIF) were defined as the sum of the contents of all individual isoflavones. The concentration of isoflavones was determined based on the retention time and peak area. The peak area data were extracted using the R software package shiny_HPLC. V1.0 (https: / / git hub.com / Zhaoqing Songjia / shiny_HPLC).
[0041] Example 3. Protein and Oil Determination
[0042] Using the constructed model, protein and oil content were measured using a near-infrared spectrometer (NIS). Approximately 10g of disease-free, dry, mature soybean seeds were placed in a measuring cup and scanned using a NIS analyzer (MATRIX-I, BRUKER, Germany). The reflectance spectrum information was converted to logarithmic values and stored. This data was then analyzed using WINISI software (WINISI 1.02) developed by INFROSOFT, ultimately outputting the protein and oil content of the soybean seeds.
[0043] Example 4: Genotyping by sequencing and SNP calling
[0044] DNA was isolated from soybean seedling leaf tissue using the cetyltrimethylammonium bromide (CTAB) method. The collected DNA samples were randomly cut to a length of approximately 350 bp using a Covaris crusher. Libraries were constructed using the TruSeq library construction kit (Novogene, China) and sequenced on the Illumina HiSeq platform. Sequencing data were compared with the reference genome (G.max Wm82.a2) using BWA software using the following parameters: mem-t 4-k 32-M. Format conversion and SNP detection were performed using SAMtools software with a minimum SNP length of 4 bp and a minimum quality value (MQ) of 20.
[0045] Example 5: Map construction and QTL detection
[0046] To improve the quality of the constructed map, SNP data were analyzed using a chi-square test with a p-value threshold of 0.05 to facilitate inclusion of SNPs in subsequent steps. Redundant markers were then removed using the bin function in the QTL IciMapping 4.1 software, run with default parameters. A genetic map was constructed using MSTmap (Linux version). The following parameters were used: a p-value cutoff of 10-20, an on-map distance of 15, a no-map size of 2, and pre-cluster estimation of yes, with all other values remaining at their default values. To detect QTLs associated with protein, oil, and isoflavone content, mapqtl 6.0 (https: / / www.kyazma.nl / index.php / MapQTL / ) was run using the multivariate model (MQM) approach. The logarithm of the odds (LOD) threshold was set to 2.5 to validate the presence of QTLs. A permutation test was performed to estimate LOD significance for 1000 iterations, with a p-value threshold of 0.05.
[0047] Example 6. QTL integration
[0048] The colocalized segments for isoflavone and protein or oil yield identified in this study were compared with previously published QTLs associated with the same traits. A composite genetic map was constructed using the G. max Wm82.a2 genome version available on the soybase website (https: / / soybase.org / ) using all downloaded QTLs, all molecular markers present on the consensus genetic map, and all coordinates of the consensus map markers. Significant QTLs associated with the target traits and neighboring markers located on the same segment were evaluated based on the physical distance separating the associated traits. Genetic and consensus maps were constructed using MapChart 2.32 software.
[0049] Example 7, statistical analysis
[0050] A total of 14 traits were analyzed, including 12 isoflavone traits and protein and oil traits. Since only trace amounts of DE, GE, GLE and AG were detected in the observation population, these compounds were excluded from further analysis. The importance of parents was analyzed using Student's t-test in the basic R package. Population genetic variation was analyzed using the R package psychology. Broad-sense heritability (h2b) was calculated using the R package LME4 (https: / / github.com / lme4) according to the following formula: h2b = VG / (VG+VE), where VG is the variance between RILs and VE is the variance within RILs. Correlation analysis was performed using the R package performance analysis.
[0051] Example 8. Comparison of isoflavones, protein and oil between parents
[0052] The parents used in this study include the representative soybean variety JD12 (ZDD23040) and the wild soybean variety Y9 (ZYD02739). The genotypes of the two parents showed significant differences in seed size, seed color and other seed traits ( Figure 1 ), more importantly, there were significant differences in isoflavone, protein and oil content ( Figure 1 The total isoflavone and protein contents of wild soybean parent Y9 were much higher than those of cultivated parent JD12, but its oil content was lower than that of cultivated parent JD12 ( Figure 1 This suggests that the genetics favored the production of isoflavones and proteins in Y9, while the genetics of JD12 favored oil synthesis. By dividing the aglycones in the isoflavones into three groups, it was possible to analyze Ds, Gs, and GL in the two parents separately. Here, we found that the content of Ds, Gs, and GLs in Y9 was significantly higher than that in JD12, with the concentrations in Y9 being 169%, 161%, and 119% higher than those in JD12, respectively ( Figure 1 ).
[0053] Example 9: Genetic and phenotypic variation results of RIL populations
[0054] To further investigate the genetic basis of isoflavone production in soybean seeds, recombinant inbred lines (RILs) derived from the F8:10 progeny of JD12, Y9, and 185 were evaluated for 12 isoflavone traits, as well as protein and oil content, in a two-year field experiment. Consistent with previous results, no aglycones of DE, GE, or GLE, nor acetylated glycosides of AG, were detected in this study. Furthermore, substantial genetic variation in the contents of other isoflavones, as well as protein and oil, was detected among the parents and RILs (Table 1). The average CVs for isoflavone, protein, and oil content over the two-year observation period were 34.44%, 13.12%, and 6.90%, respectively, indicating that genetic variation in isoflavone production was greater than that in protein or oil production. Widespread heritability was observed for each observed trait, except oil, among the 185 RILs. Population means for most traits fell between the parental values, while the maximum and minimum values exceeded the parental limits, suggesting that both parents may contribute to phenotypic variation (Table 1). Based on the kurtosis and skewness values present in the trait histograms, the contents of isoflavones, protein, and oil detected in the entire population followed a normal and continuous distribution, indicating that these traits were inherited as quantitative traits ( Tables 1 and Figure 1 Furthermore, the broad-sense heritabilities of the traits observed over the two-year experiment ranged from 0.72 to 0.87 (Table 1), indicating that phenotypic variation primarily stems from genetic, rather than environmental, variation. Therefore, localization of trait loci is feasible. The heritabilities for protein and oil content were 0.81 and 0.87, respectively, both higher than those for isoflavones, suggesting that isoflavone production is more sensitive to environmental fluctuations than protein or oil synthesis.
[0055] Example 10. Correlation analysis results of isoflavone, protein and oil contents in soybean seeds
[0056] In order to investigate the potential relationships among the observed traits, correlation analyses were performed, and the correlation coefficients, histograms, and scatter plots for isoflavone, protein, and oil contents are summarized in Figure 2 After two years of experimentation, the protein and oil contents of harvested soybean seeds showed a consistent negative correlation (r = -0.79***), and TIF was positively correlated with Ds, Gs, and GLs ( Figure 2 Total isoflavone content was highly correlated with Ds (r = 0.96***) and Gs (r = 0.98***), which was consistent with the additional observed correlations between TIF and D, MD, G, and MG ( Figure 7 At the same time, the correlation between TIF and GLs was relatively low (r=0.79***), indicating that Ds and Gs may be the main contributors to soybean TIF ( Figure 2Isoflavones were significantly correlated with both protein and oil content, with TIF negatively correlated with protein content (r=-0.36***) and weakly positively correlated with oil content (r=0.067*). Figure 2 When the results for 2020 and 2021 were analyzed separately, similar results were observed, with annual TIF content showing correlations with protein content of -0.42*** and -0.34***, and with oil content of 0.09* and 0.005, respectively. Figure 5-Figure 6 These results suggest that isoflavone and protein traits may be coordinately inherited, whereas oil production appears to be more independent.
[0057] Example 11: Further map construction and verification
[0058] The chromosomal locations of loci affecting isoflavone, protein, and oil content were evaluated using 226,186 SNPs identified between the parents by GBS analysis. Assuming that markers segregate in a 1:1 ratio, 6,072 (2.68%) SNP markers remained after the chi-square test in the QTL software IciMapping 4.1, run with default parameters, and 4,075 (1.8%) SNP markers remained after removing redundant markers that fell into the common bin (Meng et al., 2015). These screened SNP markers generated a high-density linkage map containing 3,943 SNP markers spanning approximately 938 Mb of the 1.1 Gb soybean genome, covering 6,307 cM, with an average distance between adjacent SNP markers of 1.6 cM. The average number of SNPs per linkage group was 197, with the highest number in linkage group 13 (314 SNPs) and the lowest number in linkage group 2 (138 SNPs) ( Figure 10 Table S1). To verify the accuracy of the map, the dominant QTL for seed coat color was located on the constructed map between 8.2 and 8.4 Mb on chromosome 8, with a high LOD value of 39.87, consistent with previously published GWAS data. Therefore, the constructed high-density and high-quality line image map was considered suitable for further research.
[0059] Example 12. QTL Identification Results for Soybean Seed Isoflavone, Protein, and Oil Content
[0060] With the availability of appropriate high-density linkage maps, potential QTLs for isoflavone content were identified annually. A total of 176 QTLs exceeded the LOD and substitution detection thresholds and were subsequently grouped into 25 loci based on genetic and physical distances (Table 2, Table S2). LOD values for the isoflavone loci ranged from 2.52 to 11.34, PVE values ranged from 6.1% to 24.8%, and ADD values were mostly negative, indicating that male-derived alleles promote increased isoflavone content, consistent with the high isoflavone content in Y9. On the other hand, several loci favoring isoflavone production were identified on chromosomes 8, 9, and 20 of JD12. Across both years of the experiment, 14 stable loci were consistently detected, accounting for 56% of the total isoflavone variation (qISO1, qISO5, qISO6.1, qISO6.2, qISO6.3, qISO6.4, qISO8.1, qISO10.3, qISO11, qISO12, qISO14, qISO17, qISO19.2, qISO19.3). These loci are located on chromosomes 1, 5, 6, 8, 9, 10, 11, 12, 14, 17, and 19 (Table 2).
[0061] Two major stable loci with average PEV values exceeding 15% were mapped to narrow genomic segments (Table 4). The qISO5 locus, located on a 1.05 MB physical section (41042159–42098680), was identified as a strong major isoflavone locus for TIF, DS, GS, GLS, D, MD, G, MG, GL, and MGL. Across all assays, the LOD and PVE values for qISO5 ranged from 3.1% to 11.11% and 7.5% to 24.4%, respectively. Furthermore, within the qISO5 locus, the individual QTLs for qDs5, qTIF5, and qGs5 all yielded relatively high LOD values (11.11, 9.94, and 7.12, respectively) and PEV values (24.4%, 22.1%, and 16.4%, respectively), suggesting that these loci may be major controllers of Ds, TIF, and Gs biosynthesis (Table 4). Another locus, qISO6.2, closely associated with TIF, GL, and MGL production, is located within a narrow physical distance of 0.95 Mb (18449150-19395795). LOD and PVE values during the experimental period ranged from 2.94 to 10.15% and 7.1 to 24.8%, respectively. The average LOD and PVE values within qISO6.2 for the combined sections of qGLs6.2 and qMGL6.2 were 10.59% and 23.4%, respectively, indicating that this locus is a major control site for GLs synthesis (Table 4).
[0062] A total of 42 QTLs for protein and oil were detected, including 20 for protein and 22 for oil. LOD values ranged from 2.51 to 10.30, and PVE values ranged from 6.10% to 22.80% (Table 3). Seven stable protein QTLs were significant across both years, including qPC6.1, qPC8, qPC9, qPC15, qPC20.1, qPC20.2, and qPC20.3. With the exception of qPC6.1, the QTL LOD values for each locus were greater than 5. The maximum protein LOD value returned by locus qPC20.2 (22,632,082 bp) was 9.41, with the maximum PVE value of 21.10%. Eight QTLs for oil production were detected in each year, including qOC6.1, qOC8, qOC13.1, qOC15.1, qOC20.1, qOC20.3, qOC20.4, and qOC20.5. With the exception of qOC6.1, qOC13, and qOC20.5, each locus returned an LOD value greater than 5. The qOC20.1 locus (5,834,525 bp-6,101,553 bp) had the highest average LOD value of 10.13 and the highest average PEV of 22.55%. For protein production, all QTLs except the one on chromosome 6 returned negative ADD values, indicating that Y9 contributed to favorable protein inheritance in addition to the allele on chromosome 6. In contrast, JD12 provided favorable alleles for oil production in addition to the allele on chromosome 6. Based on genetic and physical distances, the combined protein and oil loci could be grouped into 15 quality loci. These 10 loci contained both protein QTL and oil QTL (Table 3). This is consistent with previous results, in which protein yield was strongly negatively correlated with oil yield in soybean ( Figure 2 Overall, several stable major QTLs associated with isoflavone, protein, and oil contents were detected in the observed RIL populations.
[0063] Example 13. Co-localization results of soybean isoflavone, protein and oil content sites
[0064] This study found three overlapping regions of colocalized protein and isoflavone QTLs on chromosomes 6, 8, and 9. Among them, qISO6.2 and qPC6.1 were located at the same position, qISO8.1 and qPC8 were located at the same position, and qISO9.1 and qISO9.2 were both located at the same position as qPC9 ( Figure 3 JD12 provides alleles that increase protein and decrease isoflavones on a colocalized segment on chromosome 6, as well as alleles that decrease protein and increase isoflavones on colocalized segments on chromosomes 8 and 9 (Table 2, Table 3). This is consistent with the overall negative correlation between isoflavone and protein content observed in this study ( Figure 2 ).
[0065] Not surprisingly, markers for previously reported protein and isoflavone QTLs were found within the colocalized regions identified in this study. On chromosome 6, the region of the colocalized isoflavone and protein loci identified in this study was also close to the previously identified isoflavone loci Seed isoflavone 1-2, Seed isoflavone 8-1, and the protein loci Seed protein 28-1, Seed protein 29-1, and Protein 35-2. On chromosome 8, the colocalized region identified in this study was very close to the previously reported isoflavone loci Seed isoflavone 7-1, Seed isoflavone 7-7, Seed isoflavone 7-10, and Seed isoflavone 6-7, as well as the protein loci Seed protein 30-4, cqSeed protein 013, cqSeed protein 016, Seed protein 34-4, Seed protein 34-5, and Protein 26-1. On chromosome 9, the colocalization region defined in this study is close to the previously identified isoflavone locus Seed isoflavone 12-6, and the previously identified protein loci Seed protein 36-27, protein 24-3, protein 41-7, and protein 40-3 ( Figure 3 ).
[0066] More interestingly, previously identified isoflavone and protein loci are located in the same region as the major isoflavone and protein loci identified in this study. Near qISO5 lie the previously identified isoflavone loci Seed isoflavone 1-1, Seed isoflavone 6-1, and Seed isoflavone 7-5, as well as the previously identified protein loci Seed protein 34-1, Seed protein 9-1, Protein 2-1, cqSeed protein 011, and Protein 12-1. Meanwhile, the region near qPC15.1 contains the previously identified isoflavone loci Seed isoflavone 6-3, Seed isoflavone 9-2, and Seed isoflavone 11-10, as well as the protein loci Seed protein 30-3, Seed protein 4-5, Seed protein 39-2, Seed protein 3-6, cqSeed protein 001, and Protein 5-1 ( Figure 3 ). All these results clearly indicate that the production of soy isoflavones and protein is coordinated genetically, which may contribute to the effective improvement of the content of soy isoflavones as well as the content of protein or oil.
[0067] Example 14: Effect of isoflavone site combinations on isoflavone and protein / oil content
[0068] Two major favorable loci, qISO5 (primarily targeting Ds and Gs) and qISO6.2 (primarily targeting GL), were selected to further analyze the effects of combined isoflavone loci on isoflavone, protein, and oil content. Progeny without the favorable loci (N group) had significantly lower isoflavone content than those with them. The additive effect of qISO5 (q5 group) was greater than that of qISO6.2 (q6.2 group). Progeny carrying both qISO5 and qISO6.2 (q5q6.2 group) had the highest isoflavone concentrations, indicating a cumulative effect of these two loci on isoflavone content. Members of the N and q5 groups showed a higher number of discrete progenies with high protein content (pink dots), while the q6.2 and q5q6.2 groups showed a higher number of discrete lines with high oil content (round dots). This leads to the speculation that breeding soybeans for high isoflavone and oil content may be simpler than breeding soybeans for high isoflavone and protein content, which is consistent with the rest of the results of this study, in which isoflavone content was negatively correlated with protein content and weakly positively correlated with oil content ( Figure 2 We also found that some protein-rich lines within the q5q6.2 group were scattered among other members, suggesting that it may be possible to breed high-protein, high-isoflavone lines through targeted soybean breeding efforts ( Figure 4 ).
[0069] Example 15: Application Prospects
[0070] Improvements in living standards are driving a growing demand for healthy nutrition, leading to a surge in demand for crop products, such as soybeans, that offer an optimal combination of protein and oil content. Furthermore, soybean seeds may also contain high isoflavone content, which has been linked to health benefits, such as reduced risk of cancer and several chronic diseases. Therefore, in modern soybean breeding programs, there is a pressing need to develop superior soybean varieties that produce both high isoflavone content and high protein or oil content.
[0071] The importance of optimizing protein, oil, and isoflavone content in soybean using quantitative trait loci (QTLs) has long been recognized and studied. However, understanding the relationship between isoflavone, protein, and oil production in soybean remains limited. Here, in two years of field experiments, we investigated the relationship between seed isoflavone, protein, and oil content using QTLs associated with isoflavone, protein, and oil yield. The h2b for isoflavone content in these experiments was 0.81 (Table 1), corresponding to the high heritability previously reported. While genetic variation in isoflavone production exists in soybean, environmental influences are also known to influence seed isoflavone concentrations. Consequently, 11 of 25 loci, including five previously reported loci, were detected in just one year, while six new loci, designated qISO2.2, qISO3, qISO10.1, qISO10.2, qISO19.1, and qISO20, were also identified (Table 2). On the other hand, 14 of the 25 loci, including 10 previously reported loci, were detected in both trials (Table 2). This leaves four significant loci, designated qISO1, qISO6.1, qISO6.3, and qISO6.4, as loci that have not yet been registered in the corresponding composite map at www.soybase.org. Interestingly, we found that the gene Glyma.01g172900 on chromosome 1 (51020047-51025420), located near qISO1 (51600690-54799844) (Table 2), was predicted to be an isoflavone reductase-like protein, potentially a gene at the qISO1 locus controlling isoflavone content. The reconfirmation of QTLs in multiple trials suggests that not only are the regions associated with isoflavone production detected in this study reliable, but also that new stable QTLs may be detected across different populations.
[0072] Previously, a major QTL at the lower end of chromosome 5 was detected in five different populations and narrowed down to a 611.4 kb segment on chromosome 5 between bp positions 38,434,171 and 39,045,620 in the G. maxWm82.a1 reference genome. In this study, we also detected a major QTL marker, qTIF5, within a 373.6 kb segment on chromosome 5 between bp positions 41,042,159 and 41,415,752, explaining 16.1%–22.1% of the variation in isoflavone content observed across the two years of the included field trials. QTLs qD5, qMD5, qDs5, qG5, qMG5, qGs5, qGL5, qMGL5, and qgl5 were also identified near this segment, suggesting that this region may control the production of multiple individual isoflavones (Table 4), consistent with previous reports.
[0073] The low correlation coefficients between glycyrrhizin derivative isoflavones and other isoflavones in this study are consistent with previous reports and suggest that Ds and Gs may be produced in similar metabolic pathways that share metabolic enzymes, while GL may be produced in another, more isolated metabolic pathway that shares few metabolic enzymes with other isoflavone pathways. This suggests that the production of GLs may involve novel genes or loci.
[0074] In this study, we fine-mapped a locus, labeled qISO6.2, located within a 946.3 kb segment on chromosome 6 between positions 18,449,510 and 19,395,795 bp. Within this region, the qGLs6.2 and qML6.2 segments exhibited higher PVE values (average: 23.4% vs. 9.5%) and LOD (average: 10.59% vs. 3.4%) than qTIF6.2, indicating that qISO6.2 primarily controls GL production. Using 480 single-nucleotide polymorphisms (SNPs), we identified a locus associated with GL production near BARC-031337-07051 (physical position: 16,679,945) on chromosome 6, with an associated LOD value of 4.2. The LOD value for the locus detected on chromosome 6 is higher than that of previous studies, facilitating gene cloning and marker-assisted selection (MAS).
[0075] A total of 42 protein and oil QTLs were detected, and 15 of the QTLs detected in the two-year experiment were largely consistent with QTLs identified in previous studies (Table 3). The major locus, qPC20.3, was located between 32.38 and 33.29 Mb on chromosome 20, close to the previously cloned gene Glyma.20g085100, located at a physical location of 31.77 Mb on chromosome 20. Furthermore, the major locus, qOC15.1, was found to be located between 2.32 and 4.37 Mb on chromosome 15, close to GMsweet 39 (Gly ma.15g 049200) at 3.87 Mb on chromosome 15 (mio et al., 2020). These results demonstrate that the high-density map generated in this study accurately localized the relevant genes to the important QTL region (Table 3, Table S3).
[0076] Three of the 14 loci associated with protein content colocalized with the isoflavone loci on three linkage groups (chromosomes 6, 8, and 9) ( Figure 3In these colocalized segments, the additive effects on isoflavone and protein content appear to be opposite. For example, qISO6.2 had a negative additive effect on protein content, with a minimum additive value of -515.976, while the contribution of qPC6.1, located in the same segment, was positive, with a maximum additive value of 1.13 (Tables S2 and S3). This may account for the significant negative correlation between soy isoflavone content and protein content (-0.36***), consistent with previous reports.
[0077] The phenotypic correlations between traits could be due to genetic proximity or pleiotropy within the linkage group. To verify that the observed phenotypic correlations were due to colocalization, the 2020 data were analyzed to understand the effect of eliminating the three colocalized segments. In this permutation, the correlation between protein and isoflavone content was reduced (-0.12) and became non-significant ( Figure 11 ), further demonstrating that the observed correlation is due to colocalization. In addition, previously mapped isoflavone and protein sites were found in all three deleted fragments ( Figure 3 Although only one major isoflavone locus (qISO5) was located on chromosome 5 in this study, isoflavone and protein loci have been found on the same segment in previous studies ( Figure 3 ). The additive effect of this fragment on isoflavone content was found to be opposite to that on protein content. Similarly, although only one major protein locus was detected on chromosome 15 in this study, both protein and isoflavone loci have been previously mapped to the same region ( Figure 3 These results strongly suggest that seed isoflavone and protein content may be coordinately inherited through colocalization of influential loci in the soybean genome.
[0078] Isoflavone production may be negatively correlated with protein content, primarily due to the colocalization of protein and isoflavone loci. However, some loci, such as qPC15.1 and qISO5, are thought to control only protein or isoflavone content, so breeding soybean varieties rich in isoflavones and protein or oil remains a possibility. In this study, no favorable isoflavone loci were found to cluster with high protein loci ( Figure 4 ), despite sufficient segregation to suggest that soybeans with high levels of isoflavones and protein could be cultivated. For example, when qISO5 and qISO6.2 were combined, some progeny lines contained high levels of isoflavones and protein. Conversely, in certain situations, such as infant nutrition, low-isoflavone varieties may be desirable. This study elucidated the genetic basis of protein, oil, and isoflavone production, paving the way for further exploration to produce optimal combinations of isoflavone, protein, and oil content for various applications.
[0079] In summary, the research conducted and presented here identified and mapped stable and major QTLs for soybean isoflavone, protein, and oil yield on a high-density genetic map. Some loci are consistent with previous studies, while others are novel and described for the first time. The consistency with previous results, the significant correlations, and the elimination of the effects of significant loci from the analysis demonstrate the accuracy and novelty of the results. Furthermore, the loci for isoflavone and protein production were mapped to colocalized regions, further demonstrating that the observed correlation between isoflavone and protein content in soybean seeds is likely caused by coordinated inheritance of linked loci rather than pleiotropy. Overall, these results provide a theoretical basis for future soybean breeding efforts to coordinately improve isoflavone and protein quality.
[0080] Example 16, table summary
[0081] Table 1
[0082]
[0083]
[0084] Table 2
[0085]
[0086]
[0087]
[0088]
[0089] Table 3
[0090]
[0091]
[0092] Table 4
[0093]
[0094]
[0095] Table S1
[0096]
[0097] Table s2
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] Table s3
[0104]
[0105]
[0106] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention.
Claims
1. Use of co-localized quantitative trait loci, characterized in that: The co-localized quantitative trait loci include: qISO6.2 and qPC6.1 co-localized on chromosome 6; qISO8.1 and qPC8 co-localized on chromosome 8; qISO9.1 and qISO9.2 co-localized on chromosome 9; wherein: qISO6.2 is located on soybean chromosome 6, with a genetic distance of 213.609-225.650 cM and a physical site / interval of 06_18449510-06_21098994; qPC6.1 is located on soybean chromosome 6, with a genetic distance of 224.401-224.408 cM and a physical site / interval of 06_19395795-06_20312314; qISO8.1 is located on soybean chromosome 8, with a genetic distance of 213.609-225.650 cM and a physical site / interval of 06_18449510-06_21098994; The chromosomes of soybean are qPC8, with a genetic distance of 95.815-96.096 cM and a physical locus / interval of 08_9020859-08_9054795; qPC8 is located on soybean chromosome 8, with a genetic distance of 80.663-103.272 cM and a physical locus / interval of 08_7270752-08_9502316; qISO9.1 is located on soybean chromosome 9, with a genetic distance of 31.954 cM and a physical locus / interval of 09_2556374; qISO9.2 is located on soybean chromosome 9, with a genetic distance of 46.908-48.224 cM and a physical locus / interval of 09_3204462-09_3672384; the version number of the soybean reference genome is G. max Wm82.a2; the use is: to develop molecular markers and primers thereof that are simultaneously correlated with the isoflavone content and protein content of soybeans.
2. Use of the co-localized quantitative trait loci according to claim 1 for synergistic molecular marker-assisted breeding of soybean protein and isoflavones.
3. Use of the co-localized quantitative trait loci according to claim 1 for molecular marker-assisted breeding of new soybean varieties with low protein content and high isoflavone content.
4. Use of the co-localized quantitative trait loci according to claim 1 for molecular marker-assisted breeding of new soybean varieties with high protein content and low isoflavone content.