Molecular marker associated with seed oil content in soybean GmANK gene and application
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-11
AI Technical Summary
然而,现有研究针对大豆籽粒油分的遗传解析仍较为缺乏,且已定位的分子标记在育种实践中的验证和应用尚不充分,难以满足高油高产大豆品种选育的实际需求
本发明通过全基因组关联分析发现与大豆种子含油量性状相关的InDel位点定位到大豆参考基因组Glycine max Wm82.a4.v1版本的17号染色体上,位于GmANK基因的启动子区域。利用该分子标记能够实现高油含量大豆品种的早期快速筛选,可用于在大豆苗期预测大豆种子中油分的含量,为大豆分子标记辅助育种技术体系提供了新的技术支持,有效缩短育种周期,提高育种效率,对加快高油大豆新品种的选育进程具有实际应用价值。
Smart Images

Figure CN122235375B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and in particular to molecular markers in the soybean GmANK gene associated with seed oil content and their applications. Background Technology
[0002] Soybeans Glycine max Soybeans are a common legume and an important economic and oilseed crop. Because their seeds are rich in oil and protein, improving soybean yield and oil extraction rate is a key focus in soybean breeding. Traditional breeding methods are time-consuming, and given the current trends in molecular biology, there is an urgent need for an accurate and efficient molecular breeding method to identify seed oil and protein content, as well as superior soybean varieties based on 100-seed weight.
[0003] Genome-wide association analysis (GWAS) is a method that detects genetic variation polymorphism across the entire genome of multiple individuals, obtains genotypes, and then analyzes the association between genotypes and observable traits to identify genes associated with those traits. In recent years, with the rapid development of genome sequencing technology, GWAS combined with molecular marker technology has been widely applied in crop breeding work, including quantitative trait research, providing a new approach for discovering superior genes in crops. However, current research on the genetic analysis of soybean seed oil content is still relatively lacking, and the validation and application of identified molecular markers in breeding practice are insufficient, making it difficult to meet the actual needs of breeding high-oil and high-yield soybean varieties. Summary of the Invention
[0004] To enable faster and more accurate breeding of high-oil soybean varieties, this invention provides molecular markers in the soybean GmANK gene associated with seed oil content and their applications.
[0005] The present invention adopts the following technical solution: The first objective of this invention is to provide a molecular marker associated with the oil content of soybean seeds, the nucleotide sequence of which is shown in SEQ ID NO.1 and SEQ ID NO.2; the molecular marker is a fragment with the sequence GTAGGTTAT inserted or deleted at position 5332649 on chromosome 17 of the soybean reference genome, version number Glycine max Wm82.a4.v1.
[0006] As a preferred embodiment, the soybean plants with the molecular marker nucleotide sequence shown in SEQ ID NO.1 have a higher seed oil content compared to the soybean plants with the molecular marker nucleotide sequence shown in SEQ ID NO.2.
[0007] A second object of the present invention is to provide primer compositions for identifying molecular markers associated with the oil content of soybean seeds, said primer compositions comprising primers with nucleotide sequences as shown in SEQ ID NO.3 and SEQ ID NO.4.
[0008] As a preferred embodiment, the primer composition comprises an upstream primer with a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer with a nucleotide sequence as shown in SEQ ID NO.4.
[0009] A third object of the present invention is to provide a kit for identifying molecular markers associated with the oil content of soybean seeds, the kit comprising the primer composition described above.
[0010] The fourth objective of this invention is to provide a method for identifying the oil content of soybean seeds, comprising the following steps: using the genomic DNA of the soybean plant to be tested as a template, performing PCR amplification using the primer composition for identifying molecular markers associated with the oil content of soybean seeds; compared with the soybean plant to be tested with a PCR amplification product size of 714 bp, the soybean plant to be tested with a PCR amplification product size of 705 bp has a higher seed oil content.
[0011] As a preferred embodiment, the nucleotide sequence of the PCR amplification product with a size of 705 bp is shown in SEQ ID NO.1, and the nucleotide sequence of the PCR amplification product with a size of 714 bp is shown in SEQ ID NO.2.
[0012] As a preferred approach, genomic DNA from the leaves of the soybean plant to be tested is used as a template.
[0013] As a preferred embodiment, the primer composition comprises an upstream primer with a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer with a nucleotide sequence as shown in SEQ ID NO.4.
[0014] As a more preferred embodiment, the PCR amplification reaction system includes: the upstream primer at a working concentration of 80 nM to 120 nM, and the downstream primer at a working concentration of 80 nM to 120 nM.
[0015] As a specific implementation, the PCR amplification reaction system includes the upstream primer at a working concentration of 80 nM, 90 nM, 100 nM, 110 nM, or 120 nM.
[0016] As a specific implementation, the PCR amplification reaction system includes the downstream primers at working concentrations of 80 nM, 90 nM, 100 nM, 110 nM, or 120 nM.
[0017] As a preferred embodiment, the PCR amplification reaction system comprises: the upstream primer at a working concentration of 100 nM and the downstream primer at a working concentration of 100 nM.
[0018] As a preferred approach, the PCR amplification reaction system includes genomic DNA at a concentration of 5 ng / μL to 10 ng / μL.
[0019] As a specific implementation scheme, the PCR amplification reaction system includes genomic DNA at concentrations of 5 ng / μL, 6 ng / μL, 7 ng / μL, 8 ng / μL, 9 ng / μL, or 10 ng / μL.
[0020] As the preferred method, the PCR amplification reaction system includes genomic DNA at a concentration of 7 ng / μL.
[0021] As a preferred approach, the PCR amplification reaction program includes: pre-denaturation at 94℃~96℃ for 2 min~4 min; 28~35 cycles of amplification reaction, each cycle including denaturation at 94℃~96℃ for 14 s~16 s, annealing at 54℃~56℃ for 14 s~16 s and extension at 71℃~73℃ for 38 s~42 s; and final extension at 71℃~73℃ for 4 min~6 min.
[0022] As a preferred approach, the PCR amplification reaction program includes: pre-denaturation at 95°C for 3 min; 30 cycles of amplification reaction, each cycle including denaturation at 95°C for 15 s, annealing at 55°C for 15 s and extension at 72°C for 40 s; and final extension at 72°C for 5 min.
[0023] As a preferred method, the size of PCR amplification products is detected by agarose gel electrophoresis.
[0024] As a preferred embodiment, the method further includes the following steps: sequencing the PCR amplification products to obtain the nucleotide sequence of the PCR amplification products.
[0025] As a more preferred approach, the following steps are also included: S1. Extract genomic DNA from the soybean plants to be tested; S2. Perform PCR amplification using the primer composition to obtain PCR amplification products; S3. Sequencing the PCR amplification product to obtain the nucleotide sequence of the PCR amplification product; S4. Determine the seed oil content of the soybean plant being tested based on the nucleotide sequence of the PCR amplification product. The determination methods include: Compared with the soybean plant whose nucleotide sequence of PCR amplification product is shown in SEQ ID NO.2, the soybean plant whose nucleotide sequence of PCR amplification product is shown in SEQ ID NO.1 has a higher seed oil content.
[0026] The present invention also provides the application of primer compositions or kits for identifying the molecular markers in molecular marker-assisted breeding of soybean.
[0027] A fifth object of the present invention is to provide the use of primer compositions or kits for identifying the molecular markers in screening and / or breeding soybean varieties with high seed oil content.
[0028] A sixth object of the present invention is to provide the application of primer compositions or kits for identifying the molecular markers in predicting the seed oil content of soybeans.
[0029] The present invention has the following beneficial effects: This invention, through genome-wide association analysis, identified the InDel locus associated with soybean seed oil content as located on chromosome 17 of the soybean reference genome version Glycine max Wm82.a4.v1, within the promoter region of the GmANK gene. This molecular marker enables rapid early screening of high-oil-content soybean varieties and can be used to predict oil content in soybean seeds at the seedling stage. It provides new technical support for soybean marker-assisted breeding technology, effectively shortening the breeding cycle and improving breeding efficiency, thus having practical application value in accelerating the breeding process of new high-oil soybean varieties. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0031] Figure 1 This is an agarose gel electrophoresis image of the PCR product provided in Example 2 of this invention; M represents the DNA molecular weight standard, and the bands from largest to smallest are 8000 bp, 5000 bp, 3000 bp, 2000 bp, 1000 bp, 750 bp, 500 bp, 250 bp and 100 bp; 1 to 10 represent soybean materials from 10 Class I subgroups; 11 to 20 represent soybean materials from 10 Class II subgroups.
[0032] Figure 2This is a scatter box plot of the oil content of soybean seeds from two subgroups of varieties provided in Example 2 of the present invention. The sample size of subgroup I is 473, and the sample size of subgroup II is 43. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0034] Unless otherwise specified, the experimental methods involved in the following embodiments are conventional methods in the art. For example, you can refer to the experimental manual in the art or follow the conditions recommended in the manufacturer's instructions.
[0035] Unless otherwise specified, all experimental materials and reagents used in the following examples are commercially available.
[0036] Example 1: Determination of InDel sites associated with oil content traits This invention uses 521 natural population soybean materials as research objects, the main varieties of which are Dongsheng 112, Dongsheng 137, Dongsheng 124, etc., which are produced in Gongzhuling, Changchun City and are currently preserved at the Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences. The public can obtain these biological materials from the applicant.
[0037] DNA was extracted from the leaves of various soybean materials using the CTAB method and sequenced. The sequencing results were compared with the soybean reference genome version Glycine max Wm82.a4.v1. Genome-wide association analysis was used to identify the InDel loci in seeds associated with oil content traits.
[0038] The oil content of this soybean material, measured in 2024, was analyzed using rMVP software and the GLM model, with a significance threshold set at 0.0000000249.
[0039] Analysis revealed that the InDel site, located at position 5332649 on chromosome 17 of the soybean reference genome Glycine max Wm82.a4.v1, is associated with soybean oil content and has a -log10(p) value greater than 8. This site contains a 9 bp fragment (sequence: GTAGGTTAT) insertion. Soybean seeds without this insertion had higher oil content compared to those with this insertion.
[0040] Based on genome annotation information, the InDel site is located in the promoter region of the GmANK gene.
[0041] The reference sequence of the promoter region of the GmANK gene is as follows: (SEQ ID NO.1).
[0042] The InDel site contains a 9bp fragment (sequence: GTAGGTTAT) with the following inserted sequence: AATCAGTGGTGATACGATCCGACTATATATGTTTTAAATATATATTAGTCGGATAATAATAATTAAAATTATTAAATAACAACAATAACATAACAATGATATATCATTTAATGACAATGTATTTTACCGAAGGAGATATACCCAAATCTCACAAATAT ATTTTATTGTTTAGAGGAAATCTCATAAATATCTGTTTTTACGTTTATAATTATTTTTTAAAATATAACAAGTTTAGACTTTAGATCACTCGAGTTTATCTGAATTGATTAAAACAAAAAGCCATTAATTCAAGTACCACTTTGTCATAGTTTTCCAA GTAGGTTAT GTAGGTTGAAAATCATACTTAGAGGGGAGAGATTAAACTGACCTTTAGGACTTTGATGCTAATGCAGTTTGCTACTATGTGCTACTGATGATTTAGCTATAGACACATGTCATGCTTGGAAAGGTTTCACGCTTCGAGCATAATGGAGGTCGCAATCATTTTGTTTCTTATTGAGTGGACTCTTTCCATCCCCATGC ACCAATTTTTCATGCTTGAGTTGAGTATGTAAAAGTAACTAGGAGTATGTTGTAAGTTGTAAATACTATTCAATCATTATTGGATTGTTGTTGACCAAGTTGTGTCTTAGGAAAAATTTATATATATAAAAACAAAATGTCTTATGGTGTTCATTTAACTTAACTAGAGTCAAATGGGATGCAGCACTGGAGAGAA (SEQ ID NO. 2), the underlined fragment is the 9 bp fragment.
[0043] Example 2: Identification of InDel sites and analysis of oil content in soybeans 1. Identification methods (1) Genomic DNA extraction Leaves from another 516 natural soybean populations were used as test samples. They were placed in 2 mL centrifuge tubes containing steel balls, ground thoroughly, and genomic DNA was extracted using the CTAB method.
[0044] (2) PCR amplification Using genomic DNA as a template, PCR amplification was performed using primer F (sequence: 5'-AATCAGTGGTGATACGATCCG-3' (SEQ ID NO.3)) and primer R (sequence: 5'-TTCTCTCCAGTGCTGCATCC-3' (SEQ ID NO.4)).
[0045] The PCR amplification reaction program was as follows: pre-denaturation at 95℃ for 3 min; 30 cycles of amplification reaction, each cycle including denaturation at 95℃ for 15 s, annealing at 55℃ for 15 s and extension at 72℃ for 40 s; final extension at 72℃ for 5 min; and storage at 4℃.
[0046] The PCR amplification reaction system consisted of 2 μL of genomic DNA at a concentration of 70 ng / μL, 1 μL of primer F (SEQ ID NO.3) at a concentration of 2 μM, 1 μL of primer R (SEQ ID NO.4) at a concentration of 2 μM, 10 μL of 2×Rapid Taq Master Mix, and ddH2O to a final volume of 20 μL.
[0047] (3) Determination of strain subgroups The amplification products were analyzed by agarose gel electrophoresis. The strain subpopulation of the sample was determined based on the size of the electrophoretic bands, using the following method: like Figure 1 As shown, if a 705 bp band (SEQ ID NO.1) is amplified, it indicates that the sample to be tested does not have a 9 bp fragment (sequence: GTAGGTTAT) insertion, and the sample to be tested is determined to be a subpopulation of strain I; if a 714 bp band (SEQ ID NO.2) is amplified, it indicates that the sample to be tested has a 9 bp fragment (sequence: GTAGGTTAT) insertion, and the sample to be tested is determined to be a subpopulation of strain II.
[0048] Analysis revealed that 473 soybean samples belonged to subgroup I, while the remaining 43 samples belonged to subgroup II.
[0049] (4) Analysis of oil content characteristics When the soybean materials in this embodiment reach maturity, the seeds are collected, and the crude fat content of the seeds is determined by Fourier transform near-infrared spectroscopy. The percentage of crude fat content to the dry weight of the seeds is calculated as the oil content (%). The difference in oil content between subgroup I and subgroup II is statistically analyzed by t-test.
[0050] Statistical results are as follows Figure 2 As shown, the oil content of seeds from subgroup I (mean 20.7%) was significantly higher than that of subgroup II (mean 18.6%), with a p-value of 3.9 × 10⁻⁶. -12This indicates that the InDel site provided by this invention can effectively identify the oil content of soybean seeds and can be used for screening and breeding high-oil soybean varieties.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying high and low oil content of soybean seeds, characterized by, Includes the following steps: Using the genomic DNA of the soybean plants to be tested as a template, PCR amplification was performed using a primer combination for identifying molecular markers; compared with the soybean plants to be tested with a PCR amplification product size of 714 bp, the soybean plants to be tested with a PCR amplification product size of 705 bp had higher seed oil content. The nucleotide sequence of the molecular marker is shown in SEQ ID NO.1 and SEQ ID NO.2; the molecular marker is a fragment with the sequence GTAGGTTAT inserted or deleted at position 5332649 on chromosome 17 of the soybean reference genome, and the version number of the soybean reference genome is Glycine max Wm82.a4.v1; The primer composition comprises an upstream primer with a nucleotide sequence as shown in SEQ ID NO.3 and a downstream primer with a nucleotide sequence as shown in SEQ ID NO.4; The soybean plants to be tested were of the varieties Dongsheng 112, Dongsheng 137, or Dongsheng 124.
2. The method of claim 1, wherein, The genomic DNA of the leaves of the soybean plant to be tested was used as a template.
3. The method of claim 1, wherein, The size of PCR amplification products was detected by agarose gel electrophoresis.
4. Use of the primer composition or the kit comprising the primer composition as claimed in claim 1 for screening and / or breeding soybean varieties with high oil content in seeds, characterized in that, The soybean varieties mentioned are Dongsheng 112, Dongsheng 137, or Dongsheng 124.
5. Use of the primer composition or the kit comprising the primer composition described in claim 1 in predicting seed oil content of soybean, characterized in that, The soybean varieties mentioned are Dongsheng 112, Dongsheng 137, or Dongsheng 124.