Haplotype molecular markers related to soybean oil content and their applications
By developing haplotype molecular markers related to soybean oil content traits and using the Glyma.18G027100 and Glyma.03G021800 genes as targets, the problem of low soybean breeding efficiency in existing technologies has been solved, enabling efficient screening of soybean materials with high or low oil content, and improving the accuracy and efficiency of breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YANCHENG INST OF TECH
- Filing Date
- 2024-12-23
- Publication Date
- 2026-07-17
Smart Images

Figure CN119685512B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular marker technology, and particularly relates to haplotype molecular markers related to soybean oil content and their applications. Background Technology
[0002] Soybean (Glycine max (L.) Merr.) is an annual herbaceous plant belonging to the genus Glycine in the family Leguminosae. The oil content of soybeans directly affects their oil yield. Therefore, developing and screening molecular markers closely linked to soybean oil content to improve soybean breeding efficiency is of great significance for soybean variety selection. Summary of the Invention
[0003] The purpose of this invention is to provide haplotype molecular markers related to soybean oil content traits and their applications. The haplotype molecular markers of this invention are closely linked to the oil content of soybeans and can improve soybean breeding efficiency.
[0004] This invention provides the application of the Glyma.18G027100 gene and / or the Glyma.03G021800 gene as targets in regulating or identifying soybean oil content.
[0005] The present invention also provides haplotype molecular markers related to soybean oil content traits, including a first haplotype molecular marker and / or a second haplotype molecular marker;
[0006] The first haplotype molecular marker consists of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8;
[0007] Among them, Glyma.18G027100-SNP1 is located at 2149619 bp of the Glyma.18G027100 gene, and the base is T or C;
[0008] Glyma.18G027100-SNP2 is located at 2149789 bp of the Glyma.18G027100 gene, and the base is A or G.
[0009] Glyma.18G027100-SNP3 is located at 2150035 bp in the Glyma.18G027100 gene, and the base is A or G.
[0010] Glyma.18G027100-SNP4 is located at 2150825 bp in the Glyma.18G027100 gene, with a base of C or T.
[0011] Glyma.18G027100-SNP5 is located at 2150831 bp of the Glyma.18G027100 gene, and the base is either A or G.
[0012] Glyma.18G027100-SNP6 is located at 2155101 bp in the Glyma.18G027100 gene, with a base of C or G.
[0013] Glyma.18G027100-SNP7 is located at 2155366 bp of the Glyma.18G027100 gene, and its base is A.
[0014] Glyma.18G027100-SNP8 is located at 2155379 bp of the Glyma.18G027100 gene, with a base of T or C;
[0015] When the haplotype combination corresponding to Glyma.18G027100-SNP1~Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC.
[0016] The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3 and Glyma.03G021800-SNP4;
[0017] Among them, Glyma.03G021800-SNP1 is located at 2354862 bp of the Glyma.03G021800 gene, and the base is A or C;
[0018] Glyma.03G021800-SNP2 is located at 2354994 bp of the Glyma.03G021800 gene, and the base is either G or A.
[0019] Glyma.03G021800-SNP3 is located at 2357837 bp in the Glyma.03G021800 gene, and the base is either G or A.
[0020] Glyma.03G021800-SNP4 is located at 2358021 bp in the Glyma.03G021800 gene, and the base is A or C.
[0021] When the haplotype combination corresponding to Glyma.03G021800-SNP1~Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC.
[0022] The reference genome for the soybean is the 'Zhonghuang 13' genome.
[0023] The present invention also provides the application of the haplotype molecular markers described above in the identification or auxiliary identification of soybean oil content traits.
[0024] This invention also provides a method for identifying or assisting in the identification of soybean oil content traits using the haplotype molecular markers described above, comprising the following steps:
[0025] Genomic DNA was extracted from the soybean material to be tested;
[0026] By detecting haplotype combinations on the Glyma.18G027100 gene and / or Glyma.03G021800 gene in genomic DNA through gene sequencing, the soybean oil content is higher when the haplotype combination corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 is TAACACAT, and / or when the haplotype combination corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than when the haplotype combinations are AGAC and CAAC.
[0027] The present invention also provides the application of a reagent for detecting the haplotype molecular marker described above in the preparation of a kit for identifying or assisting in the identification of soybean oil content.
[0028] This invention also provides the application of the haplotype molecular markers described above as targets in regulating soybean oil content.
[0029] This invention also provides the application of the haplotype molecular markers described above in soybean molecular breeding.
[0030] Preferably, the soybean molecular breeding includes at least one of identifying the oil content trait of soybean materials, screening soybean materials with high oil content, and screening soybean materials with low oil content.
[0031] This invention also provides a method for molecular breeding of soybeans, comprising the following steps: when the target trait is high oil content, selecting soybean materials with haplotype combinations TAACACAT corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 and / or haplotype combinations AGGA corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4, and eliminating individuals with other genotypes;
[0032] When the target trait is low oil content, soybean materials with haplotype combinations CGGTGGAC corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 and / or haplotype combinations AGAC or CAAC corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4 are selected for subsequent generations, while individuals with other genotypes are eliminated.
[0033] Preferably, the high oil content includes an oil content ≥ 21%; the low oil content includes an oil content ≤ 17%.
[0034] This invention provides the application of the Glyma.18G027100 gene and / or the Glyma.03G021800 gene as targets in regulating or identifying soybean oil content. The haplotypes of the Glyma.18G027100 and Glyma.03G021800 genes exhibit significant differences in oil content (OC) in natural soybean populations. For the Glyma.18G027100 gene, two haplotypes were identified based on SNPs at different positions in the exon regions: Hap1 (TAACACAT) and Hap2 (CGGTGGAC), with Hap1 having the highest frequency (82%). In the population used for haplotype analysis, the average OC of Hap1 was significantly higher than that of Hap2, and Hap1 in Glyma.18G027100 increases OC. Three distinct haplotypes were identified in the analysis of Glyma.03G021800: Hap1 (AGGA), Hap2 (AGAC), and Hap3 (CAAC). Hap1 accounted for 51% of the population, and its average OC was significantly higher than the other two haplotypes, making it a superior haplotype for enhancing OC. In conclusion, Glyma.18G027100 and Glyma.03G021800 are candidate genes associated with soybean OC. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This shows the distribution of important QTNs across the 14 chromosomes; important QTNs located in multiple environments are marked in red. ;
[0037] Figure 2 This is a heatmap showing the expression of 8 differentially expressed genes in 4 groups; where blue represents downregulated genes, red represents upregulated genes, and white represents genes with no significant difference. Gene names displayed in red indicate that the gene is upregulated in all groups, and gene names displayed in blue indicate that the gene is downregulated in all groups. ;
[0038] Figure 3 The graph shows the correlation analysis results between the OC trait and the WGCNA module; where red indicates a positive correlation, blue indicates a negative correlation, * indicates a significant correlation at the 0.05 level, ** indicates a significant correlation at the 0.01 level, and *** indicates a significant correlation at the 0.001 level.
[0039] Figure 4 The correlations between the eight differentially expressed genes and the module are shown; where red indicates a positive correlation, blue indicates a negative correlation, * indicates a significant correlation at the 0.05 level, ** indicates a significant correlation at the 0.01 level, and *** indicates a significant correlation at the 0.001 level.
[0040] Figure 5 The graph shows the haplotype information of Glyma.18G027100 and Glyma.03G021800 and their phenotypic analysis results in natural populations; among them , A shows the haplotype information and corresponding number of individuals for Glyma.18G027100; B shows the box plot of oil content for haplotypes Hap1 and Hap2 of Glyma.18G027100 in 121 materials; C shows the haplotype information and corresponding number of individuals for Glyma.03G021800; D shows the haplotype information and corresponding number of individuals for 141 materials.
[0041] Box plot of oil content of haplotypes Hap1, Hap2 and Hap3 in Glyma.03G021800 material. Detailed Implementation
[0042] This invention provides the application of the Glyma.18G027100 gene and / or the Glyma.03G021800 gene as targets in regulating or identifying soybean oil content.
[0043] In this invention, when the Glyma.18G027100 gene (LOC100784534: XM_003552305.5) and the Glyma.03G021800 gene (LOC100527561: NM_001250951.2) are taken with Williams 82 as the reference genome, they are located at Chr18:2033838..2038697bp reverse and Chr03:2265528..2269179bp forward in the Glycine max Wm82.a2.v1 reference sequence, respectively. When the Glyma.18G027100 and Glyma.03G021800 genes are taken from Zhonghuang 13 as the reference genome, they are located at Chr18:2149555..2155433bp (-strand) and Chr03:2354584..2358026bp (+strand) respectively in the ZH13 reference sequence.
[0044] This invention also provides haplotype molecular markers related to soybean oil content traits, including a first haplotype molecular marker and / or a second haplotype molecular marker; the first haplotype molecular marker is composed of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8; wherein, Glyma.18G0 SNP1 (27100-SNP1) is located at position 2149619 bp of the Glyma.18G027100 gene, with a base of either T or C; SNP2 (2149789 bp) is located at position 2149789 bp of the Glyma.18G027100 gene, with a base of either A or G; SNP3 (2150035 bp) is located at position 2150035 bp of the Glyma.18G027100 gene, with a base of either A or G; SNP4 (2150825 bp) is located at position 2150825 bp of the Glyma.18G027100 gene, with a base of either C or T; SNP5 (27100-SNP5) is located at position 2149619 bp of the Glyma.18G027100 gene, with a base of either C or T. Located at 2150831 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP6 is located at 2155101 bp of the Glyma.18G027100 gene, with a base of C or G; Glyma.18G027100-SNP7 is located at 2155366 bp of the Glyma.18G027100 gene, with a base of A; Glyma.18G027100-SNP8 is located at 2155379 bp of the Glyma.18G027100 gene, with a base of T or C; Glyma.18G027100-SNP1~Glyma.18G02 When the haplotype combination corresponding to 7100-SNP8 is TACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC. The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3, and Glyma.03G021800-SNP4. Among them, Glyma.03G021800-SNP1 is located at 2354862 bp of the Glyma.03G021800 gene, with a base of A or C; Glyma.03G021800-SNP2 is located in Glyma.At position 2354994 bp of the 03G021800 gene, the base is G or A; Glyma.03G021800-SNP3 is located at position 2357837 bp of the Glyma.03G021800 gene, with a base of G or A; Glyma.03G021800-SNP4 is located at position 2358021 bp of the Glyma.03G021800 gene, with a base of A or C; when the haplotype combination corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC; the reference genome of the soybean is the 'Zhonghuang 13' genome.
[0045] In this invention, the first haplotype molecular marker is located in the soybean Glyma.18G027100 gene, specifically at positions 2149619bp to 2155379bp, with a total length of approximately 5760bp.
[0046] In this invention, the second haplotype molecular marker is located in the soybean Glyma.03G021800 gene, specifically at positions 2354862bp to 2358021bp, with a total length of approximately 3159bp.
[0047] The present invention also provides the application of the haplotype molecular markers described above in the identification or auxiliary identification of soybean oil content traits.
[0048] This invention also provides a method for identifying or assisting in the identification of soybean oil content traits using the haplotype molecular markers described above, comprising the following steps: extracting genomic DNA from the soybean material to be tested; detecting haplotype combinations on the Glyma.18G027100 gene and / or the Glyma.03G021800 gene in the genomic DNA by gene sequencing; when the haplotype combination corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC; and / or, when the haplotype combination corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC.
[0049] The present invention also provides the application of a reagent for detecting the haplotype molecular marker described above in the preparation of a kit for identifying or assisting in the identification of soybean oil content.
[0050] This invention also provides the application of the haplotype molecular markers described above as targets in regulating soybean oil content.
[0051] This invention also provides the application of the haplotype molecular markers described above in soybean molecular breeding.
[0052] In the specific implementation of this invention, the soybean molecular breeding includes at least one of the following: identifying the oil content trait of soybean materials, screening soybean materials with high oil content, and screening soybean materials with low oil content.
[0053] This invention also provides a method for molecular breeding of soybeans, comprising the following steps: when the target trait is high oil content, selecting soybean materials with haplotype combinations TAACACAT corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 and / or haplotype combinations AGGA corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4, and eliminating individuals with other genotypes;
[0054] When the target trait is low oil content, soybean materials with haplotype combinations CGGTGGAC corresponding to Glyma.18G027100-SNP1 to Glyma.18G027100-SNP8 and / or haplotype combinations AGAC or CAAC corresponding to Glyma.03G021800-SNP1 to Glyma.03G021800-SNP4 are selected for subsequent generations, while individuals with other genotypes are eliminated.
[0055] In the specific implementation of this invention, the high oil content includes an oil content ≥ 21%; the low oil content includes an oil content ≤ 17%.
[0056] To further illustrate the present invention, the haplotype molecular markers related to soybean oil content and their applications provided by the present invention are described in detail below with reference to the accompanying drawings and embodiments, but these should not be construed as limiting the scope of protection of the present invention.
[0057] Materials: Kenfeng 14, Kenfeng 15, and Kenfeng 19 were bred by the Heilongjiang Provincial Academy of Agricultural Sciences; Heinong 48 was bred by the Heilongjiang Academy of Agricultural Sciences; Kenfeng 14, Kenfeng 15, Kenfeng 19, and Heinong 48 were all provided by the Soybean Research Institute of Northeast Agricultural University;
[0058] The FW-RIL population was obtained through (Kenfeng 14 x Kenfeng 15) x (Heinong 48 x Kenfeng 19). This population was provided by the Soybean Research Institute of Northeast Agricultural University.
[0059] Example 1
[0060] Example 1
[0061] 1. Materials
[0062] In this embodiment, a population of 144 FW-RILs was used, obtained by four-way crosses of the parent lines Kenfeng 14 (OC 20.88%), Kenfeng 15 (OC 21.66%), Kenfeng 19 (OC 18.33%), and Heinong 48 (OC 17.76%).
[0063] 144 FW-RILs and 4 parent lines were planted in 10 environments (hereinafter referred to as E1-E10) at different locations and years (Table 1).
[0064] Table 1 Field Planting Information
[0065] E1 Acheng 126.50°E, 45.27°N 2016 E2 Twin Cities 126.18°E, 45.22°N 2016 E3 Harbin 126.63°E, 45.75°N 2016 E4 Keshan 125.52°E, 48.02°N 2016 E5 Harbin 126.63°E, 45.75°N 2017 E6 Harbin 126.63°E, 45.75°N 2018 E7 Keshan 125.52°E, 48.02°N 2018 E8 Keshan 125.52°E, 48.02°N 2019 E9 Harbin 126.63°E, 45.75°N 2020 E10 Harbin 126.63°E, 45.75°N 2021
[0066] All materials were planted in a completely randomized block design with three replicates under all conditions. Rows were 5m long and 0.7m wide. The management of the experimental plots was consistent with local soybean production practices. At maturity, 10 plants from each line were randomly harvested and threshed. The OC (octane rating) of dry seeds (approximately 10% moisture) of each line was determined using a near-infrared analyzer (Infratec 1241, Foss, Denmark) at the Key Laboratory of Soybean Biology, Ministry of Education, Northeast Agricultural University, China. The calibration regression technique used for this near-infrared analyzer was partial least squares (PLS), which involves combining spectral data with laboratory data (Kjeldahl nitrogen determination) to calculate seed OC as a percentage of seed weight. The calculation results are shown in Table 2. The phenotypic values of each parent and FW-RIL individual used in this example are the average of three replicates.
[0067] Table 2. Descriptive statistics of OC data of parental and FW-RIL populations under 10 environments.
[0068]
[0069] 2. Phenotypic Data Analysis
[0070] Phenotypic data analysis includes determining the mean, standard deviation, range, minimum, maximum, skewness, kurtosis, and coefficient of variation (CV), analysis of variance (ANOVA), and generalized heritability (h). 2 Analysis. All statistical analyses were performed in SAS 9.2 (SAS Institute, Cary, USA). Multi-environmental ANOVA was performed using a combined generalized linear model (GLM) approach, with variance components estimated using a mixed linear model (MLM). h was then calculated using the following formula. 2 :
[0071]
[0072] in and Let r represent the variances of genotype, genotype-environment interaction, and error variance, respectively; r represents the number of replicates in each environment; and n represents the number of environments.
[0073] 3. Genotype, population structure, and linkage disequilibrium
[0074] Genomic DNA was extracted from the young leaves of parental plants and FW-RIL plants using the CTAB method. DNA concentration was determined using a UV752N spectrophotometer (Shanghai Jinko Scientific Instruments Co., Ltd.), and the DNA was diluted to 100 ± 1 ng in deionized water. SNP genotyping was performed at Beijing Boao Biotechnology Co., Ltd. based on the SoySNP660KBeadChip method. After quality screening, with a standard maximum deletion site ratio <10% and a minimum allele frequency (MAF) >5%, 109,676 SNPs were selected for GWAS analysis.
[0075] By analyzing the population structure, the FW-RIL population was divided into two subpopulations, and the Q matrix was obtained and used for multi-site GWAS analysis. The LD region decayed the fastest before 200kb, and then tended to plateau. Therefore, the 200kb regions above and below significant SNPs were identified for searching for potential candidate genes.
[0076] 4. Genome-wide association study (GWAS)
[0077] Five multi-site GWAS methods—mrMLM, FASTmrMLM, FASTmrEMMA, pLARmEB, and ISISEM-BLASSO—were used in mrMLM GUI (version 3.0) software to perform GWAS analysis on the OC trait of soybean grains. The critical p-value parameter for the first stage was set to 0.01 for these methods, except for FASTmrEMMA, whose critical p-value was set to 0.005. For the significant quantitative trait nucleotides (QTN) in the final stage, the critical LOD score was set to 3, and the p-value was set to 0.0002. The parameters were set as follows: probability = "REML", SearchRadius = 100, CriLOD = 3, SelectVariable = 143, Bootstrap = FALSE. The kinship matrix could also be obtained using mrMLM GUI 3.0 software.
[0078] Significant SNPs that meet one of the following conditions are considered important QTNs: (1) the same SNP marker is detected; (2) the decay intervals of different SNP markers intersect (the same QTN detected repeatedly by multiple methods is counted only once).
[0079] The QTN naming convention is as follows: q + trait name - chromosome name - QTN number, where q represents QTN, the trait name is OC, and the QTN number represents the rank of QTN on the corresponding chromosome.
[0080] QTNs detected by multiple GWAS methods or in multiple environments are considered important QTNs.
[0081] Results: This embodiment identified a total of 23 important QTNs (important QTNs are those detected in multiple GWAS methods or in multiple environments) (Table 3 and...). Figure 1 Twenty-one common QTNs were detected in one environment using at least two methods (listed in bold in Table 3); LOD values ranged from 3.24 to 7.78, and PVE percentages in each QTN ranged from 4.22% to 17.68%. Two common QTNs (qOC-7-1 and qOC-18-1) were identified across multiple environments using one or more methods; details of these two QTNs are shown in bold in Table 3; LOD values ranged from 4.83 to 5.20, and PVE percentages in each QTN ranged from 5.52% to 9.20%. For all 23 common QTNs, the direction (positive or negative) of the effect of each QTN was consistent across different methods or environments. No significant QTNs were located on the remaining six chromosomes (chromosomes 06, 12, 13, 15, 19, and 20).
[0082] Table 3. 23 important QTNs related to soybean oil content identified by different methods or in different environments.
[0083]
[0084]
[0085] b ,mrMLM,FASTmrMLM,FASTmrEMMA,pLARmEB andISIS EM-BLASSO represent 1 to 5 respectively.
[0086] 5. Transcriptome sequencing analysis
[0087] The soybean materials used for transcriptome sequencing were Keshan 01 (KS01), Kedou 31 (KD31), Kedou 48 (KD48), and Kedou 57 (KD57). KS01 and KD31 are high-oil varieties with OC contents of 21.82% and 21.26%, respectively; KD48 and KD57 are low-oil varieties with OC contents of 15.78% and 16.44%, respectively. These varieties were provided by the Keshan Branch of the Heilongjiang Academy of Sciences.
[0088] Four uniform seeds of different varieties were selected and potted (5 seeds / pot, 25×18cm pot). After germination, the number of seedlings per pot was reduced to 1. 25 days after flowering (the rapid oil accumulation stage), the seeds were collected, frozen in liquid nitrogen (triple replicate), and stored in an ultra-low temperature freezer at -80℃ until transcriptome sequencing analysis was performed.
[0089] Transcriptome sequencing analysis was performed by Guangzhou Genebio Biotechnology Co., Ltd. (Guangzhou, China) on an Illuminahis q2500 / 4000 system.
[0090] 6. Real-time quantitative PCR verification
[0091] Total RNA extraction, cDNA synthesis, and qRT-PCR analysis of all samples were performed according to the method described by Zhang et al. (see [Zhang, W. et al. A cation diffusion facilitator, GmCDF1, negatively regulates salt tolerance in soybean. PLoS Genet. 15 (2019).]). The internal reference gene was Acint7. Data analysis was performed using 2... -ΔΔCT Method. Primers used for qRT-PCR are shown in Table 4.
[0092] Table 4 qRT-PCR primers
[0093]
[0094]
[0095] 7. Identifying potential candidate genes
[0096] Based on LD decay, a 100kb interval flanking each important QTN was defined as the range for identifying potential candidate genes. All genes within this interval were identified using the Phytozome website (https: / / phytozome.jgi.doe.gov). Genes highly expressed in seeds during lipid formation were then selected from the soybean seed development stage transcriptome dataset (GES42871) downloaded from GEO (Gene Expression Comprehensive Database). Pathway analysis of all highly expressed genes was then performed using the KEGG website (http: / / www.kegg.jp), identifying genes related to lipid synthesis pathways. Finally, differentially expressed genes (DEGs) between high- and low-oil varieties were identified within lipid synthesis pathways based on transcriptome sequencing results. Potential candidate genes were identified by comparing these DEGs with weighted gene association network analysis (WGCNA) results and gene annotation information.
[0097] The results are as follows:
[0098] A total of 441 genes were identified within 23 key QTN regions, of which 196 genes were highly expressed during seed lipid formation. Metabolic pathway analysis of these 196 genes revealed that 132 genes were enriched in 163 pathways, of which only 15 were related to lipid synthesis and metabolism. These included glycolysis / gluconeogenesis, the citric acid cycle (TCA cycle), fructose and mannose metabolism, galactose metabolism, starch and sucrose metabolism, amino sugar and nucleotide sugar metabolism, glyoxylate and dicarboxylate metabolism, propionate metabolism, and inositol phosphate metabolism, all belonging to carbohydrate metabolism pathways; oxidative phosphorylation, photosynthetic antenna proteins, nitrogen metabolism, and sulfur metabolism, belonging to energy metabolism; and glycerol lipid metabolism, belonging to lipid metabolism pathways (Table 5).
[0099] Table 5. Annotation information for the 13 genes and their corresponding metabolic pathways in KEGG.
[0100]
[0101] A total of 13 genes are associated with these 15 pathways. Compared with the DEGs of the KD48-vs-KS01, KD57-vs-KS01, KD48-vs-KD31, and KD57-vs-KD31 groups, there were 8 differentially expressed genes between high-OC and low-OC soybean varieties (Table 5 and 10). Figure 2 These eight DEGs are Glyma.09G173200, Glyma.03G021800, Glyma.14G218800, Glyma.16G213200, Glyma.18G025500, Glyma.18G027100, Glyma.11G093200, and Glyma.18G027200.
[0102] For WGCNA analysis based on transcriptome sequencing data, genes with similar expression patterns were clustered into 20 modules ( Figure 3 Analysis of the relationship between each module and the OC phenotype showed that MM.darkseagreen4 (p = 4e-07), MM.violet (p = 2e-04), MM.darkred (p = 0.01), and MM.floralwhite (p = 0.04) were significantly positively correlated with the OC trait. The module with a significantly negative correlation to the OC trait was MM.brown 4 (p = 0.003). Figure 3Further correlation analysis was performed between the eight DEGs and the five modules mentioned above. Seven DEGs (excluding Glyma.16G13200) were significantly correlated with one or more of the five modules. Glyma.11G093200 and Glyma.18G027100 showed significant or highly significant positive correlations with four modules. Figure 4 (The red text indicates that Glyma.18G027200) showed a positive correlation with the OC trait. Glyma.18G027200 showed a significant positive correlation with MM.darkred and a highly significant positive correlation with MM.floralwhite. Glyma.03G021800, Glyma.09G173200, Glyma.18G025500, and Glyma.14G218800 showed significant or highly significant positive correlations with the MM.brown4 module (the module negatively correlated with the lipid content trait). Therefore, these seven DEGs (except Glyma.16G213200) are considered potential candidate genes for further haplotype analysis.
[0103] 8. Haplotype analysis
[0104] Haplotype analysis was performed on genes from 141 soybean accessions using existing genotype data (https: / / ngdc.cncb.AC.cn / soy omics). Haplotypes were constructed using Haploview software based on SNPs in untranslated regions (UTRs) and exons.
[0105] The results are as follows:
[0106] Haplotypes were classified using natural populations, including 141 soybean accessions with genotype data (https: / / ngdc.cncb.AC.cn / soy omics). Haplotype analysis of the seven DEGs revealed significant differences in oil content (OC) between haplotypes Glyma.18G027100 and Glyma.03G021800 in natural soybean populations. For Glyma.18G027100, two haplotypes were identified based on SNPs at different positions in the exon regions: Hap1 (TAACACAT) and Hap2 (CGGTGGAC). Figure 5 A). Hap1 had the highest frequency (82%). In the population used for haplotype analysis, the mean OC of Hap1 was significantly higher than that of Hap2 (A). Figure 5 (B in the text). These results indicate that Hap1 (TAACACAT) in Glyma.18G027100 increases OC. Three distinct haplotypes were identified from the analysis of Glyma.03G021800: Hap1 (AGGA), Hap2 (AGAC), and Hap3 (CAAC). Figure 5 (C in the group). Hap1 accounts for 51% of the population, and its average OC is significantly higher than the other two haplotypes, making it a superior haplotype for enhancing OC. Figure 5 (D in the text). In summary, Glyma.18G027100 and Glyma.03G021800 can be predicted to be candidate genes associated with soybean OC.
[0107] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, and not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.
Claims
1. The application of a reagent for detecting haplotype molecular markers in the identification of soybean oil content, characterized in that, The haplotype molecular marker is a first haplotype molecular marker and / or a second haplotype molecular marker; The first haplotype molecular marker consists of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8; Among them, Glyma.18G027100-SNP1 is located at 2149619 bp of the Glyma.18G027100 gene, and the base is T or C; Glyma.18G027100-SNP2 is located at 2149789 bp of the Glyma.18G027100 gene, and the base is A or G. Glyma.18G027100-SNP3 is located at 2150035 bp in the Glyma.18G027100 gene, and the base is A or G. Glyma.18G027100-SNP4 is located at 2150825 bp in the Glyma.18G027100 gene, with a base of C or T. Glyma.18G027100-SNP5 is located at 2150831 bp of the Glyma.18G027100 gene, and the base is A or G. Glyma.18G027100-SNP6 is located at 2155101 bp in the Glyma.18G027100 gene, with a base of C or G. Glyma.18G027100-SNP7 is located at 2155366 bp of the Glyma.18G027100 gene, and its base is A. Glyma.18G027100-SNP8 is located at 2155379 bp of the Glyma.18G027100 gene, with a base of T or C; When the haplotype combination corresponding to Glyma.18G027100-SNP1~Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC. The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3 and Glyma.03G021800-SNP4; Among them, Glyma.03G021800-SNP1 is located at 2354862 bp of the Glyma.03G021800 gene, and the base is A or C; Glyma.03G021800-SNP2 is located at 2354994 bp of the Glyma.03G021800 gene, and the base is either G or A. Glyma.03G021800-SNP3 is located at 2357837 bp in the Glyma.03G021800 gene, and the base is either G or A. Glyma.03G021800-SNP4 is located at 2358021 bp in the Glyma.03G021800 gene, and the base is A or C. When the haplotype combination corresponding to Glyma.03G021800-SNP1~Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC. The reference genome for the soybean is the 'Zhonghuang 13' genome.
2. A method for identifying soybean oil content traits using haplotype molecular markers, characterized in that, Includes the following steps: Genomic DNA was extracted from the soybean material to be tested; The haplotype molecular marker is a first haplotype molecular marker; The first haplotype molecular marker consists of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8; Specifically, Glyma.18G027100-SNP1 is located at position 2149619 bp of the Glyma.18G027100 gene, with a base of T or C; Glyma.18G027100-SNP2 is located at position 2149789 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP3 is located at position 2150035 bp of the Glyma.18G027100 gene, with a base of A or G; and Glyma.18G027100-SNP4 is located at position 2150825 bp of the Glyma.18G027100 gene, with a base of C. Or T; Glyma.18G027100-SNP5 is located at position 2150831 bp of the Glyma.18G027100 gene, with base A or G; Glyma.18G027100-SNP6 is located at position 2155101 bp of the Glyma.18G027100 gene, with base C or G; Glyma.18G027100-SNP7 is located at position 2155366 bp of the Glyma.18G027100 gene, with base A; Glyma.18G027100-SNP8 is located at position 2155379 bp of the Glyma.18G027100 gene, with base T or C; By detecting haplotype combinations on the Glyma.18G027100 gene in genomic DNA through gene sequencing, when the haplotype combination corresponding to Glyma.18G027100-SNP1~Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC. The reference genome for the soybean is the 'Zhonghuang 13' genome.
3. A method for identifying soybean oil content traits using haplotype molecular markers, characterized in that, Includes the following steps: Genomic DNA was extracted from the soybean material to be tested; The haplotype molecular marker is a second haplotype molecular marker; The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3 and Glyma.03G021800-SNP4; Specifically, Glyma.03G021800-SNP1 is located at position 2354862 bp of the Glyma.03G021800 gene, with a base of either A or C; Glyma.03G021800-SNP2 is located at position 2354994 bp of the Glyma.03G021800 gene, with a base of either G or A; Glyma.03G021800-SNP3 is located at position 2357837 bp of the Glyma.03G021800 gene, with a base of either G or A; and Glyma.03G021800-SNP4 is located at position 2358021 bp of the Glyma.03G021800 gene, with a base of either A or C. By detecting haplotype combinations on the Glyma.03G021800 gene in genomic DNA through gene sequencing, when the haplotype combination corresponding to Glyma.03G021800-SNP1~Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC. The reference genome for the soybean is the 'Zhonghuang 13' genome.
4. The application of a reagent for detecting haplotype molecular markers in the preparation of a kit for identifying soybean oil content traits; wherein the haplotype molecular marker is a first haplotype molecular marker and / or a second haplotype molecular marker; The first haplotype molecular marker consists of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8; in, Glyma.18G027100-SNP1 is located at position 2149619 bp of the Glyma.18G027100 gene, with a base of T or C; Glyma.18G027100-SNP2 is located at position 2149789 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP3 is located at position 2150035 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP4 is located at position 2150825 bp of the Glyma.18G027100 gene, with a base of C or T. Glyma.18G027100-SNP5 is located at position 2150831 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP6 is located at position 2155101 bp of the Glyma.18G027100 gene, with a base of C or G; Glyma.18G027100-SNP7 is located at position 2155366 bp of the Glyma.18G027100 gene, with a base of A; Glyma.18G027100-SNP8 is located at position 2155379 bp of the Glyma.18G027100 gene, with a base of T or C. When the haplotype combination corresponding to Glyma.18G027100-SNP1~Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC. The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3 and Glyma.03G021800-SNP4; Specifically, Glyma.03G021800-SNP1 is located at position 2354862 bp of the Glyma.03G021800 gene, with a base of either A or C; Glyma.03G021800-SNP2 is located at position 2354994 bp of the Glyma.03G021800 gene, with a base of either G or A; Glyma.03G021800-SNP3 is located at position 2357837 bp of the Glyma.03G021800 gene, with a base of either G or A; and Glyma.03G021800-SNP4 is located at position 2358021 bp of the Glyma.03G021800 gene, with a base of either A or C. When the haplotype combination corresponding to Glyma.03G021800-SNP1~Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC. The reference genome for the soybean is the 'Zhonghuang 13' genome.
5. The application of reagents for detecting haplotype molecular markers in soybean molecular breeding, characterized in that, The soybean molecular breeding includes at least one of the following: identifying the oil content trait of soybean materials, screening soybean materials with high oil content, and screening soybean materials with low oil content. The haplotype molecular marker is a first haplotype molecular marker and / or a second haplotype molecular marker; The first haplotype molecular marker consists of Glyma.18G027100-SNP1, Glyma.18G027100-SNP2, Glyma.18G027100-SNP3, Glyma.18G027100-SNP4, Glyma.18G027100-SNP5, Glyma.18G027100-SNP6, Glyma.18G027100-SNP7 and Glyma.18G027100-SNP8; Specifically, Glyma.18G027100-SNP1 is located at position 2149619 bp of the Glyma.18G027100 gene, with a base of T or C; Glyma.18G027100-SNP2 is located at position 2149789 bp of the Glyma.18G027100 gene, with a base of A or G; Glyma.18G027100-SNP3 is located at position 2150035 bp of the Glyma.18G027100 gene, with a base of A or G; and Glyma.18G027100-SNP4 is located at position 2150825 bp of the Glyma.18G027100 gene, with a base of C. Or T; Glyma.18G027100-SNP5 is located at position 2150831 bp of the Glyma.18G027100 gene, with base A or G; Glyma.18G027100-SNP6 is located at position 2155101 bp of the Glyma.18G027100 gene, with base C or G; Glyma.18G027100-SNP7 is located at position 2155366 bp of the Glyma.18G027100 gene, with base A; Glyma.18G027100-SNP8 is located at position 2155379 bp of the Glyma.18G027100 gene, with base T or C; When the haplotype combination corresponding to Glyma.18G027100-SNP1~Glyma.18G027100-SNP8 is TAACACAT, the soybean oil content is higher than that when the haplotype combination is CGGTGGAC. The second haplotype molecular marker consists of Glyma.03G021800-SNP1, Glyma.03G021800-SNP2, Glyma.03G021800-SNP3 and Glyma.03G021800-SNP4; Specifically, Glyma.03G021800-SNP1 is located at position 2354862 bp of the Glyma.03G021800 gene, with a base of either A or C; Glyma.03G021800-SNP2 is located at position 2354994 bp of the Glyma.03G021800 gene, with a base of either G or A; Glyma.03G021800-SNP3 is located at position 2357837 bp of the Glyma.03G021800 gene, with a base of either G or A; and Glyma.03G021800-SNP4 is located at position 2358021 bp of the Glyma.03G021800 gene, with a base of either A or C. When the haplotype combination corresponding to Glyma.03G021800-SNP1~Glyma.03G021800-SNP4 is AGGA, the soybean oil content is higher than that when the haplotype combinations are AGAC and CAAC. The reference genome for the soybean is the 'Zhonghuang 13' genome.