A molecular marker for identifying high and low soybean oil content and application thereof
Patent Information
- Application Number
- CN202610712359.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-05-22
AI Technical Summary
[0004]目前现有技术中主要集中于开发大豆抗病、抗倒伏等生长性状的分子标记的开发,针对大豆油分含量的全基因组关联分析研究尚为不足,无法为大豆高油品种的高效选育提供充足的技术支撑
本发明在大豆第5号染色体上找到一个与油分含量极显著相关的InDel位点,利用基于该InDel位点开发的分子标记对育种大豆材料进行检测,可在无需种植等待植株成熟、收获种子的情况下,准确高效地预测种子油分含量高低,显著缩短大豆育种周期,提高育种效率。本发明为高油大豆品种的选育提供了重要的分子辅助工具,对提升育种精准度和产业经济效益具有实际应用价值。
Smart Images

Figure CN122256563B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soybean breeding technology, and in particular to a molecular marker for identifying the level of soybean oil content and its application. Background Technology
[0002] Soybeans are annual herbaceous plants belonging to the genus *Glycine* in the legume family. They are characterized by high oil and high protein content, with Northeast my country being a prime soybean-producing region. Oil content directly affects the oil yield of soybeans, making the selection of soybeans with high oil content an important direction in current soybean breeding.
[0003] Currently, with the vigorous development of biotechnology, crop breeding has entered a new stage of design using molecular biology techniques such as molecular marker-assisted selection, molecular design, transgenics, and genome editing. Molecular markers have become an effective tool for crop genetic improvement. With the development and application of high-throughput, low-cost molecular markers and the rapid development of bioinformatics-related technologies, the application of molecular marker-assisted selection has greatly improved the efficiency and accuracy of crop breeding.
[0004] Current technologies mainly focus on developing molecular markers for growth traits such as disease resistance and lodging resistance in soybeans. Research on genome-wide association studies of soybean oil content is still insufficient, which cannot provide sufficient technical support for the efficient breeding of high-oil soybean varieties. Summary of the Invention
[0005] To address the above problems, this invention provides a molecular marker for identifying the level of soybean oil content and its application.
[0006] The technical solution adopted in this invention is as follows: The first objective of this invention is to provide a molecular marker for identifying the level of soybean oil content, the nucleotide sequence of which is shown in SEQ ID NO:1 and SEQ ID NO:2, wherein SEQ ID NO:2 is an insert fragment between 312bp and 313bp of SEQ ID NO:1, and the nucleotide sequence of the insert fragment is TGCTAACAG.
[0007] Compared with soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:1, soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:2 have higher oil content in their seeds.
[0008] In this invention, the InDel site corresponding to the molecular marker is a fragment with the sequence TGCTAACAG inserted or deleted at 2205466 bp on chromosome 5 of the soybean Glycine max Wm82.a4.v1 genome.
[0009] A second objective of this invention is to provide primers for detecting the molecular markers used to identify the level of soybean oil content, the primers comprising nucleotide sequences as shown in SEQ ID NO:3 and SEQ ID NO:4.
[0010] Preferably, the primers include primer F and primer R, wherein primer F includes the nucleotide sequence shown in SEQ ID NO:3 and primer R includes the nucleotide sequence shown in SEQ ID NO:4.
[0011] Specifically, the nucleotide sequence of primer F has at least 95%, 96%, 97%, 98% or 99% similarity to SEQ ID NO:3.
[0012] Specifically, the nucleotide sequence of primer R has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:4.
[0013] More preferably, the nucleotide sequence of primer F is shown in SEQ ID NO:3.
[0014] More preferably, the nucleotide sequence of primer R is shown in SEQ ID NO:4.
[0015] A third objective of this invention is to provide the application of the molecular marker or its detection reagent for identifying the level of soybean oil content in the identification of soybean oil content.
[0016] Preferably, the detection reagent for the molecular marker includes primers, which include nucleotide sequences as shown in SEQ ID NO:3 and SEQ ID NO:4.
[0017] More preferably, the primers include primer F and primer R, wherein primer F includes the nucleotide sequence shown in SEQ ID NO:3 and primer R includes the nucleotide sequence shown in SEQ ID NO:4.
[0018] Specifically, the nucleotide sequence of primer F has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:3.
[0019] Specifically, the nucleotide sequence of primer R has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:4.
[0020] More preferably, the nucleotide sequence of primer F is shown in SEQ ID NO:3.
[0021] More preferably, the nucleotide sequence of primer R is shown in SEQ ID NO:4.
[0022] The fourth objective of this invention is to provide the application of the molecular marker or its detection reagent for identifying the level of soybean oil content in soybean molecular marker-assisted breeding.
[0023] Preferably, the detection reagent for the molecular marker includes primers, which include nucleotide sequences as shown in SEQ ID NO:3 and SEQ ID NO:4.
[0024] More preferably, the primers include primer F and primer R, wherein primer F includes the nucleotide sequence shown in SEQ ID NO:3 and primer R includes the nucleotide sequence shown in SEQ ID NO:4.
[0025] Specifically, the nucleotide sequence of primer F has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:3.
[0026] Specifically, the nucleotide sequence of primer R has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:4.
[0027] More preferably, the nucleotide sequence of primer F is shown in SEQ ID NO:3.
[0028] More preferably, the nucleotide sequence of primer R is shown in SEQ ID NO:4.
[0029] The fifth objective of this invention is to provide the application of the molecular marker or its detection reagent for identifying the level of soybean oil content in the breeding of soybean varieties with high oil content.
[0030] Preferably, the detection reagent for the molecular marker includes primers, which include nucleotide sequences as shown in SEQ ID NO:3 and SEQ ID NO:4.
[0031] More preferably, the primers include primer F and primer R, wherein primer F includes the nucleotide sequence shown in SEQ ID NO:3 and primer R includes the nucleotide sequence shown in SEQ ID NO:4.
[0032] Specifically, the nucleotide sequence of primer F has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:3.
[0033] Specifically, the nucleotide sequence of primer R has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:4.
[0034] More preferably, the nucleotide sequence of primer F is shown in SEQ ID NO:3.
[0035] More preferably, the nucleotide sequence of primer R is shown in SEQ ID NO:4.
[0036] The sixth objective of this invention is to provide a method for identifying the oil content of soybeans, the method comprising: performing PCR amplification on the genomic DNA of a target soybean using primers for detecting the molecular marker used to identify the oil content of soybeans; the target soybean with a PCR amplification product size of 656 bp is a low-oil soybean; the target soybean with a PCR amplification product size of 665 bp is a high-oil soybean.
[0037] Preferably, the genomic DNA of the target soybean leaf or seed is used as a template.
[0038] Preferably, the identification method further includes detecting the size of the PCR amplification product using agarose gel electrophoresis.
[0039] Preferably, the identification method further includes: sequencing the PCR amplification product and determining the type of target soybean based on the sequencing comparison.
[0040] More preferably, the target soybean with the nucleotide sequence of the PCR amplification product as shown in SEQ ID NO:1 is a low-oil soybean; the target soybean with the nucleotide sequence of the PCR amplification product as shown in SEQ ID NO:2 is a high-oil soybean.
[0041] Preferably, the PCR amplification reaction program includes: pre-denaturation at 94℃~96℃ for 2 min~4 min; performing 28~35 cycles of amplification reaction, each cycle including denaturation at 94℃~96℃ for 14 s~16 s, annealing at 54℃~56℃ for 14 s~16 s and extension at 71℃~73℃ for 38 s~42 s; and final extension at 71℃~73℃ for 4 min~6 min.
[0042] More preferably, the PCR amplification reaction program includes: pre-denaturation at 95°C for 3 min; 30 cycles of amplification reaction, each cycle including denaturation at 95°C for 15 s, annealing at 55°C for 15 s and extension at 72°C for 40 s; and final extension at 72°C for 5 min.
[0043] Preferably, the primers include primer F and primer R, wherein primer F includes the nucleotide sequence shown in SEQ ID NO:3 and primer R includes the nucleotide sequence shown in SEQ ID NO:4.
[0044] Specifically, the nucleotide sequence of primer F has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:3.
[0045] Specifically, the nucleotide sequence of primer R has at least 95%, 96%, 97%, 98%, or 99% similarity to SEQ ID NO:4.
[0046] More preferably, the nucleotide sequence of primer F is shown in SEQ ID NO:3.
[0047] More preferably, the nucleotide sequence of primer R is shown in SEQ ID NO:4.
[0048] More preferably, the PCR amplification reaction system includes: primer F with a working concentration of 80 nM to 120 nM, and primer R with a working concentration of 80 nM to 120 nM.
[0049] Specifically, the PCR amplification reaction system includes primer F at working concentrations of 80 nM, 90 nM, 100 nM, 110 nM, or 120 nM.
[0050] Specifically, the PCR amplification reaction system includes primer R at a working concentration of 80 nM, 90 nM, 100 nM, 110 nM, or 120 nM.
[0051] More preferably, the PCR amplification reaction system includes: primer F at a working concentration of 100 nM and primer R at a working concentration of 100 nM.
[0052] More preferably, the PCR amplification reaction system includes genomic DNA at a concentration of 5 ng / μL to 10 ng / μL.
[0053] Specifically, the PCR amplification reaction system includes genomic DNA at concentrations of 5 ng / μL, 6 ng / μL, 7 ng / μL, 8 ng / μL, 9 ng / μL, or 10 ng / μL.
[0054] More preferably, the PCR amplification reaction system comprises genomic DNA at a concentration of 7 ng / μL.
[0055] The present invention has the following beneficial technical effects: This invention identifies an InDel locus on soybean chromosome 5 that is highly correlated with oil content. Molecular markers developed based on this InDel locus are used to detect the oil content of soybean breeding materials. This allows for accurate and efficient prediction of seed oil content without the need for planting, waiting for plant maturity, or harvesting seeds, significantly shortening the soybean breeding cycle and improving breeding efficiency. This invention provides an important molecular-aided tool for the breeding of high-oil soybean varieties and has practical application value for improving breeding accuracy and industrial economic benefits. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0057] Figure 1 This is an agarose gel electrophoresis image of the PCR amplification product provided in Example 3 of the present invention; lane M represents the DNA molecular weight standard, and the bands from largest to smallest are 8000 bp, 5000 bp, 3000 bp, 2000 bp, 1000 bp, 750 bp, 500 bp, 250 bp and 100 bp; lanes 1 to 10 represent 10 low-oil soybeans; lanes 11 to 20 represent 10 high-oil soybeans.
[0058] Figure 2 This is a box plot of seed oil content of low-oil soybeans and high-oil soybeans provided in Embodiment 3 of the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0060] Unless otherwise specified, the experimental methods involved in the following embodiments are conventional methods in the art. For example, you can refer to the experimental manual in the art or follow the conditions recommended in the manufacturer's instructions.
[0061] Unless otherwise specified, all experimental materials and reagents used in the following examples are commercially available.
[0062] The soybean germplasm resources involved in the following examples are all preserved at the Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences, and the public can obtain these biological materials from the applicant.
[0063] Example 1: Acquisition of soybean genomic DNA samples 1. Soybean leaf sample collection Leaf samples were collected from 521 soybean plants at the seed breeding site in Changchun City, Jilin Province. The main varieties included Jiyu 86 and Dongnong 50. Three replicate samples were collected from each soybean plant.
[0064] The leaf samples were placed in 2 mL centrifuge tubes containing steel balls, immediately frozen in liquid nitrogen, thoroughly ground into powder, and then stored at -80°C.
[0065] 2. Extraction of soybean genomic DNA using the CTAB method Take 0.2g of the ground powder sample into a 2mL centrifuge tube, add 600μL of 2×CTAB extraction buffer, mix well and place in a 65℃ water bath. Shake once every 10min during the 30min water bath.
[0066] After the water bath, remove the centrifuge tube, allow it to return to room temperature, add 600 μL of chloroform, mix thoroughly by inverting the tube, and centrifuge at 12000 rpm for 10 min at room temperature. Transfer 500 μL of the supernatant to a new 1.5 mL centrifuge tube, add an equal volume of isopropanol solution, and allow it to settle at -20°C for 1 h or overnight.
[0067] Centrifuge at 12000 rpm at room temperature for 10 min, discard the supernatant; add 500 μL of 75% ethanol solution to wash the precipitate; centrifuge at 12000 rpm for 5 min, discard the supernatant, and repeat the precipitate washing step once.
[0068] Centrifuge at 12000 rpm for 5 min, discard the supernatant, air dry the centrifuge tube, add 100 μL of ddH2O to dissolve the DNA precipitate, and store the dissolved DNA solution at -20℃.
[0069] Example 2: Genome-wide association analysis of soybean seed oil content 1. Genome resequencing The genomic DNA samples of soybean plants obtained in Example 1 were resequencing.
[0070] 2. Genome-wide association analysis The sequencing results were compared and analyzed using the Glycine max Wm82.a4.v1 genome as a reference genome to obtain SNPs / InDel variant maps.
[0071] For the oil content of soybean seeds measured in 2024 from various soybean plants, rMVP software and GLM model were used, and the significance threshold in the GWAS analysis was set at 0.0000000249.
[0072] A highly relevant InDel site was identified at 2205466 bp on chromosome 5, containing an insertion fragment with the sequence TGCTAACAG.
[0073] Gene mapping analysis determined that the InDel site is located in the gene GmRPE3, specifically at the 312th bp position of the nucleotide sequence shown in SEQ ID NO:1.
[0074] A partial sequence of the GmRPE3 gene is shown in SEQ ID NO:1 (underlined bases indicate InDel sites): TCTGGTACCCTTTGTTCAGGCTCTACAATCATCTACAGTGTTTCAATAGAAGCCAGAAAATTACTTAGATTCAATATCAGCTTAATTTAGTATTTGTATAAATAAAAGATGAATACCAGGTGTACATCCAAAGGAAGATCTGTCACAGGGCGCAAT GCATCAACCACAAGAGGTCCAATTGTAATATTTGGAACAAAACGGCCATCCATTACATCAACGTGAATCCAATCACAACCAGCCAACTCCACTGCTTTCACCTAGACAAGATAACACACAAACTACAACAATCATGATCGCTGATAATAGAAAAAC A TGCCTTGAAAACAAAAGGACAGATTTACTGGCAATCAGCAAACGTGATAAAAAAAATTATGAAATATTATATATTGTAGTAGCAAAACAAAACATTGGAAACTGAAATACTGATTACAGAATGTAAAGCATGTTATAGAACCTGCTCTCCCAATTTTGCAAAGTTTGCAGAA AGAATGGATGGAGAAACAATGATATCGCTTTTTGAAAACTTGTCAACACGAGATGTAGCCTTCACTGTGGTTGAAATTTTCTTCCTGAATTCACAGAACATAAAGTAACAAAAGAACATTAATGCAGACTATATTATCCTTTCAAATACTCCAAGGCACAGGCCACATTTAA.
[0075] The partial sequence of the GmRPE3 gene containing the insert fragment is shown in SEQ ID NO:2 (underlined bases indicate the insert fragment): TCTGGTACCCTTTGTTCAGGCTCTACAATCATCTACAGTGTTTCAATAGAAGCCAGAAAATTACTTAGATTCAATATCAGCTTAATTTAGTATTTGTATAAATAAAAGATGAATACCAGGTGTACATCCAAAGGAAGATCTGTCACAGGGCGCAATG CATCAACCACAAGAGGTCCAATTGTAATATTTGGAACAAAACGGCCATCCATTACATCAACGTGAATCCAATCACAACCAGCCAACTCCACTGCTTTCACCTAGACAAGATAACACACAAACTACAACAATCATGATCGCTGATAATAGAAAAACA TGCTAACAG TGCCTTGAAAACAAAAGGACAGATTTACTGGCAATCAGCAAACGTGATAAAAAAAATTATGAAATATTATATATTGTAGTAGCAAAACAAAACATTGGAAACTGAAATACTGATTACAGAATGTAAAGCATGTTATAGAACCTGCTCTCCCAATTTTGCAAAGTTTGCAGAA AGAATGGATGGAGAAACAATGATATCGCTTTTTGAAAACTTGTCAACACGAGATGTAGCCTTCACTGTGGTTGAAATTTTCTTCCTGAATTCACAGAACATAAAGTAACAAAAGAACATTAATGCAGACTATATTATCCTTTCAAATACTCCAAGGCACAGGCCACATTTAA.
[0076] The analysis results showed that soybean plants with the insertion fragment (TGCTAACAG) at the InDel site had higher seed oil content compared to soybean plants without the insertion fragment (TGCTAACAG) at the InDel site.
[0077] Example 3: Reliability evaluation of InDel sites 1. PCR amplification To evaluate the reliability of the InDel site obtained in Example 2, primers F: 5'-TCTGGTACCCTTTGTTCAGGC-3' (SEQ ID NO:3) and primer R: 5'-TTAAATGTGGCCTGTGCCTTG-3' (SEQ ID NO:4) were designed for the sequences in the upstream and downstream regions.
[0078] Following the method in Example 1, genomic DNA was extracted from tissue samples (leaves, seeds, etc.) of the soybean to be tested. Using this DNA as a template, PCR amplification was performed using the reaction system shown in Table 1 and the reaction procedure shown in Table 2.
[0079] The amplified PCR products were detected by agarose gel electrophoresis.
[0080] Table 1 PCR reaction system (20 μL)
[0081] Table 2 PCR amplification reaction procedure
[0082] 2. Interpretation of PCR amplification results The type of soybean to be tested was determined based on the PCR amplification products, and the interpretation criteria are as follows: like Figure 1 As shown, when the PCR amplification product of the soybean to be tested is 656bp, it is determined that the InDel site of the soybean to be tested does not contain the insert fragment (TGCTAACAG), and it belongs to low-oil soybean; when the PCR amplification product of the soybean to be tested is 665bp, it is determined that the InDel site of the soybean to be tested contains the insert fragment (TGCTAACAG), and it belongs to high-oil soybean.
[0083] 3. Interpretation of first-generation sequencing alignment analysis results To further refine the analysis, the PCR products obtained in the previous step were subjected to first-generation sequencing. The sequencing results were then compared with the sequences shown in SEQ ID NO:1 and SEQ ID NO:2 to detect the presence of an insert fragment (TGCTAACAG) between 312bp and 313bp in the PCR product sequence. The soybean species to be tested was determined based on the sequencing alignment results, with the following interpretation criteria (default is homozygous): If the sequence of the PCR amplification product of the soybean to be tested is as shown in SEQ ID NO:1, it indicates that there is no insert fragment (TGCTAACAG), and the soybean to be tested is determined to be a low-oil soybean; if the sequence of the PCR amplification product of the soybean to be tested is as shown in SEQ ID NO:2, it indicates that there is an insert fragment (TGCTAACAG), and the soybean to be tested is determined to be a high-oil soybean.
[0084] 4. Sample population validation Seed samples of 516 newly collected natural population soybean materials (main varieties include Jiyu 86, etc.) from seed breeding sites in Changchun City, Jilin Province were identified according to the method of this embodiment, and the species of each soybean material were determined.
[0085] The results showed that the interpretation of PCR amplification results for all soybean materials was consistent with the interpretation of first-generation sequencing comparison analysis results. Among them, 80 soybean materials were low-oil soybeans and 436 soybean materials were high-oil soybeans.
[0086] The oil content of mature seeds of various soybean materials was determined by Fourier transform near-infrared spectroscopy, and the results were expressed as mass percentages (%). The significance of the differences was analyzed by t-test.
[0087] like Figure 2 As shown, low-oil soybeans have a significantly lower oil content than high-oil soybeans. P <0.05), with mean values of 18.8% and 20.4%, respectively.
[0088] The above results show that the interpretation results of the InDel marker of the present invention are consistent with the determination results of the oil content of soybean seeds. The InDel marker can be used for the prediction and identification of the oil content of soybean seeds, and the results are accurate and reliable.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. The application of a molecular marker detection reagent for identifying the level of soybean oil content in the identification of soybean oil content, wherein the nucleotide sequence of the molecular marker is shown in SEQ ID NO:1 and SEQ ID NO:2, wherein SEQ ID NO:2 is an insert fragment between 312bp and 313bp of SEQ ID NO:1, and the nucleotide sequence of the insert fragment is TGCTAACAG; Compared with soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:1, soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:2 have higher oil content in their seeds.
2. The application of a detection reagent for a molecular marker for identifying the level of soybean oil content in molecular marker-assisted breeding of soybean for traits related to oil content, wherein the nucleotide sequence of the molecular marker is shown in SEQ ID NO:1 and SEQ ID NO:2, wherein SEQ ID NO:2 is an insert fragment between 312bp and 313bp of SEQ ID NO:1, and the nucleotide sequence of the insert fragment is TGCTAACAG; Compared with soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:1, soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:2 have higher oil content in their seeds.
3. The application of a molecular marker detection reagent for identifying the level of soybean oil content in the cultivation of soybean varieties with high oil content, wherein the nucleotide sequence of the molecular marker is shown in SEQ ID NO:1 and SEQ ID NO:2, wherein SEQ ID NO:2 is an insert fragment between 312bp and 313bp of SEQ ID NO:1, and the nucleotide sequence of the insert fragment is TGCTAACAG; Compared with soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:1, soybeans containing molecular markers with nucleotide sequences as shown in SEQ ID NO:2 have higher oil content in their seeds.
4. The application according to any one of claims 1 to 3, characterized in that, The detection reagent for the molecular marker includes primers comprising nucleotide sequences as shown in SEQ ID NO:3 and SEQ ID NO:
4.
5. A method for identifying the content of soybean oil, characterized in that, The identification method includes: The genomic DNA of the target soybean was amplified by PCR using primers, the primers comprising the nucleotide sequences shown in SEQ ID NO:3 and SEQ ID NO:4; The target soybean with the nucleotide sequence of the PCR amplification product as shown in SEQ ID NO:1 is a low-oil soybean; the target soybean with the nucleotide sequence of the PCR amplification product as shown in SEQ ID NO:2 is a high-oil soybean.
6. The identification method according to claim 5, characterized in that, The identification method also includes: detecting the size of PCR amplification products using agarose gel electrophoresis.
7. The identification method according to claim 5 or 6, characterized in that, The identification method further includes: sequencing the PCR amplification product and determining the type of target soybean based on the sequencing comparison.