Soybean protein GmHDA5 / GmHDA17 and application of coding gene of soybean protein GmHDA5 / GmHDA17 in regulation and control of oil content
By regulating the expression and activity of histone deacetylases GmHDA5 and GmHDA17 in soybeans and using gene editing via the CRISPR-Cas system, the unresolved issue of the regulatory mechanism of lipid synthesis was resolved, resulting in a significant increase in the lipid content of soybean seeds.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-19
AI Technical Summary
The regulatory mechanisms of plant oil synthesis have not been fully elucidated in existing technologies, making it difficult to increase the oil yield of plants, especially oil crops.
By regulating the expression and activity of histone deacetylases GmHDA5 and GmHDA17 in soybeans, gene editing was performed using the CRISPR-Cas system to knock out or regulate the expression of their coding genes, thereby increasing the seed oil content.
It significantly increased the oil yield of soybean seeds and enhanced the oil content, especially the content of fatty acid components such as palmitic acid, stearic acid, oleic acid and linoleic acid.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to the application of soybean protein GmHDA5 / GmHDA17 and its encoding gene in regulating oil content. Background Technology
[0002] Seed oil content is one of the main traits in the domestication process from wild soybean to cultivated soybean. To date, we have a considerable understanding of lipid synthesis pathways and have cloned a number of genes involved in lipid synthesis. However, the regulatory mechanisms of lipid synthesis in plants and the related genes are still not fully understood. Therefore, identifying genes that regulate oil accumulation and promoting efficient oil accumulation in plants is of great significance for increasing oil yields, especially in oilseed crops.
[0003] Histone deacetylases (HDA) are a class of proteases involved in important epigenetic modifications. Acetylation neutralizes the positive charge of histones, reducing their binding affinity to negatively charged DNA, loosening chromatin structure, and promoting transcription. The function of histone deacetylases (HDA) is to remove acetyl groups, thus reducing histone acetylation. Histone deacetylases 5 and 17 (HDA5, HDA17) belong to the histone deacetylases family and are enzymes that catalyze histone deacetylation reactions. By removing acetyl groups from lysine residues of histones, they alter the structure and function of chromatin, thereby playing a crucial role in processes such as gene expression regulation. Research on HDA in plants has made some progress. In Arabidopsis, key transcriptional regulatory nodes related to nutrient cycling and senescence have been elucidated, and HDA is recruited through the single-stranded DNA / RNA binding protein WHIRLY1 to participate in leaf senescence and flowering. In 2022, Mastinu A. et al. compared gene expression in drought-tolerant and drought-sensitive plants and found that HDA responds to moderate and severe drought stress. In rice, HDA participates in regulating callus induction and regeneration efficiency, thereby increasing total yield. Functional studies of this protein in soybeans are relatively limited. There are no reports of its involvement in regulating seed oil content. Summary of the Invention
[0004] The main problem this invention aims to solve is to increase the oil yield of soybean seeds for use in soybean breeding.
[0005] To address the aforementioned problems, in a first aspect, the present invention provides the application of a protein, or a substance that knocks out the coding gene of said protein, or a substance that regulates the expression of the coding gene of said protein, or a substance that regulates the activity or content of said protein, wherein the application may be any of the following:
[0006] A1) Application in regulating the oil content of plant seeds; A2) Application in the preparation of products that regulate the oil content of plant seeds; A3) Application in cultivating plants with altered seed oil content; A4) Application in the preparation of products from plants whose seed oil content has been altered; A5) Applications in plant breeding; The protein may be a composition, protein A or protein B, wherein the composition consists of protein A and protein B; The protein A is GmHDA5, and can specifically be any of the following proteins: B1) The amino acid sequence of this protein is SEQ ID NO: 2. B2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 2, but with substitution and / or deletion and / or addition of amino acid residues. Proteins that share more than 80% amino acid sequence identity with B1) and B2) and have the same function, B4) A fusion protein obtained by attaching a tag to the end of any of the proteins defined in B1)-B3); The protein B is GmHDA17, and can specifically be any of the following proteins: C1) The amino acid sequence is that of the protein with sequence SEQ ID NO: 4. C2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 4, but with substitution and / or deletion and / or addition of amino acid residues. Proteins whose amino acid sequences defined by C3) are more than 80% identical to those defined by C1) and C2) and have the same function C4) is a fusion protein obtained by attaching a tag to the end of any of the proteins defined in C1)-C3).
[0007] In the above applications, the protein may be derived from soybeans.
[0008] The proteins mentioned above can be synthesized artificially, or their encoding genes can be synthesized first and then expressed biologically.
[0009] In this invention, the protein tag refers to a polypeptide or protein fused with a target protein using in vitro DNA recombination technology for expression, to facilitate the expression, detection, tracing, and / or purification of the target protein. The tag may be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag, etc.
[0010] The above-mentioned regulation of seed oil content can be achieved by increasing the total oil content in the seed or / and increasing the content of each component of fatty acids in the seed.
[0011] In this invention, the substance that regulates the activity and / or content of the protein can be a substance that regulates gene expression, wherein the gene encodes the proteins GmHDA5 and GmHDA17.
[0012] In the above applications, the substance regulating gene expression can be a substance that performs at least one of the following six types of regulation: 1) regulation at the gene transcription level; 2) post-transcriptional regulation of the gene (i.e., regulation of splicing or processing of the primary transcript of the gene); 3) regulation of RNA transport of the gene (i.e., regulation of mRNA transport of the gene from the nucleus to the cytoplasm); 4) regulation of gene translation; 5) regulation of mRNA degradation of the gene; and 6) post-translational regulation of the gene (i.e., regulation of the activity of the protein translated from the gene).
[0013] In this invention, the identity refers to the identity of amino acid sequences or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, by using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Per residue gap cost, and Lambdaratio to 11, 1, and 0.85 (default values) respectively, a search can be performed to calculate the identity of amino acid sequences, and then the identity value (%) can be obtained.
[0014] In this invention, the 80% or more of identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0015] The SEQ ID NO: 2 may consist of 417 amino acid residues, as detailed below: MHVGSQESVANSSTAIGFDERMLLHAEVEKKSTLHPERPDRLQAIAASLARAGIFPGKCYSIPAREITPEELITVHSLEHIESVEVTSESLSSYFTSDTYANEHSALAARLAAGLCADLASAIVSGRAKNGFALVRPPGHHAGVRQAMGFCLHNNAAVAALAAQAAGARKVLILDWDVHHGNGTQEIFEQNKSVLYISLHRHEGGKFYPGTGAAEEVGSMGAEGFCVNIPWSQGGVGDNDYIFAFQHVVLPIAAEFNPDLTIVSAGFDAARGDPLGCCDITPSGYAHMTNMLNALSGGKLLVILEGGYNLRSISSSATAVIKVLLGESPGCELENSFPSKAGLQTVLEVLKIQMNFWPALGPIFVNLESQWRMYCFERKRKQIKKRRRVLVPMWRWGRKSLLFHFLNGHHNVKSKLC。
[0016] The SEQ ID NO: 4 may consist of 543 amino acid residues, as follows: .
[0017] In the above applications, the substance that knocks out the gene encoding the protein can be any of the following: D1) Nucleic acid molecules that target the genes encoding the proteins described above; D2) Nucleic acid molecules and Cas proteins that target the genes encoding the proteins described above; D3) Nucleic acid molecules encoding genes that target the proteins described above and genes encoding Cas proteins; D4) A biological material containing the nucleic acid molecule described in D1) and the Cas protein encoding gene, wherein the biological material is a vector, expression cassette, or recombinant microorganism; The substance that regulates the expression of the gene encoding the protein or the substance that regulates the activity or content of the protein is a biological material, and the biological material is any one of the following: E1) Nucleic acid molecules that inhibit, reduce, or silence the expression of the genes encoding the proteins described above; E2) expresses the gene encoding the nucleic acid molecule described in E1); E3) contains an expression cassette containing the gene encoding the gene described in E2); E4) A recombinant vector containing the encoding gene described in E2), or a recombinant vector containing the expression cassette described in E3; E5) Recombinant microorganisms containing the encoding gene described in E2), or recombinant microorganisms containing the expression cassette described in E3), or recombinant microorganisms containing the recombinant vector described in E4); E6) A transgenic plant cell line containing the encoding gene described in E2), or a transgenic plant cell line containing the expression cassette described in E3); or a transgenic plant cell line containing the recombinant vector described in E4; E7) Transgenic plant tissue containing the encoding gene described in E2), or transgenic plant tissue containing the expression cassette described in E3), or transgenic plant tissue containing the recombinant vector described in E4; E8) A transgenic plant organ containing the encoding gene described in E2), or a transgenic plant organ containing the expression cassette described in E3), or a transgenic plant organ containing the recombinant vector described in E4).
[0018] In some embodiments of the present invention, the nucleic acid molecule described in D1) may be a gRNA targeting the protein-coding gene. In some embodiments, the gRNA targets a double-stranded DNA whose nucleotide sequence is SEQ ID NO: 5, positions 2128-2150, and / or a double-stranded DNA whose nucleotide sequence is SEQ ID NO: 5, positions 2417-2439, as a target.
[0019] The term "sgRNA (single-guide RNA)" is a component of the CRISPR-Cas system, responsible for guiding Cas proteins to recognize and cleave target nucleic acid molecules. In practical gene editing applications, sgRNA can be synthesized directly or obtained through plasmid expression or in vitro transcription. In this field, "gRNA" and "sgRNA" are often used interchangeably. In this document, "gRNA" and "sgRNA" are also used interchangeably. sgRNA generally refers to a single RNA structure formed by artificially modifying the crRNA / tracrRNA complex (gRNA) with a dual RNA structure, directly (or through a linker) linking the crRNA and tracrRNA. sgRNA is a short RNA containing a recognition region and a framework region.
[0020] The term "recognition region," also known as a guide sequence, is typically an RNA sequence (referred to herein as the "guide sequence") that is identical to or complementary to the target sequence or target site within the target RNA (sgRNA or gRNA). The guide sequence is generally sufficiently complementary to the target sequence to hybridize with the target site and guide the CRISPR / Cas complex to bind specifically to the target. Perfect complementarity between the guide sequence and the target sequence is preferred, but some mismatch (e.g., 1-6 nucleotide mismatch) is permissible as long as it still results in gene knockout. The complementarity between the guide sequence and its corresponding target sequence is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Methods for determining the complementarity of two nucleic acid sequences are within the capabilities of those skilled in the art.
[0021] The term "scaffold" generally refers to the structural or scaffold RNA sequence that guides the binding or interaction of RNA with RNA-directed nucleases and / or other RNA molecules (e.g., tracrRNA) into RNA (sgRNA or gRNA), and can also be called the backbone sequence of sgRNA. The scaffold can be conventionally selected by those skilled in the art; for example, it can be the backbone sequence of the sgRNA corresponding to Cas9, or it can be a mutant constructed based on this sequence that still retains the function of binding the corresponding Cas9.
[0022] In the above applications, the gene described in D3) or the encoding gene described in E2) is introduced into the recipient plant, specifically by transforming plant cells or tissues using conventional biological methods such as Ti plasmid, Ri plasmid, plant virus vector, direct DNA transformation, microinjection, electrocoagulation, or Agrobacterium-mediated transformation, and then cultivating the transformed plant tissues into plants. The transformed cells, tissues, or plants are understood to include not only the final products of the transformation process but also the materials obtained through asexual reproduction and transgenic progeny.
[0023] In one embodiment of the present invention, the recombinant vector may be a plant gene editing vector. The plant gene editing vector may be a pCBSG015-sgRNA vector.
[0024] As a specific embodiment, the recombinant vector is the recombinant vector pCBSG015-sgRNA. The recombinant vector is obtained by replacing the fragments between 5'-ggcaccgagtcggtgc-3' and 5'-gttgaacaacggaaac-3' of the pCBSG015 vector with DNA fragments having nucleotide sequences of 5'-CTTGCTGCTAGGTTGGCAGCTGG-3' and 5'-GGAAGGTGCTAATTTTGGATTGG-3', while keeping other sequences of the pCBSG015 vector unchanged. This recombinant plasmid is named the recombinant vector pCBSG015-sgRNA.
[0025] Those skilled in the art can readily mutate the nucleotide sequences encoding the proteins GmHDA5 and GmHDA17 of this invention using known methods, such as directed evolution or point mutation. Artificially modified nucleotides that possess 75% or more of the nucleotide sequence identity with the proteins GmHDA5 and GmHDA17 isolated in this invention, provided they encode and function as proteins GmHDA5 and GmHDA17, are derived from and equivalent to the sequences of this invention.
[0026] The aforementioned 75% or higher degree of identity can be 75%, 80%, 85%, 90%, or 95% or higher degree of identity.
[0027] In this article, identity refers to the similarity of amino acid or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the procedure, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, a search can be performed to calculate the identity of amino acid sequences, and then the identity value (%) can be obtained.
[0028] In this document, the 80% or more of identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0029] The vectors described herein are known to those skilled in the art and include, but are not limited to: plasmids, bacteriophages (such as λ phage or M13 filamentous phage), granules (i.e., Cos plasmids), Ti plasmids, or viral vectors.
[0030] Recombinant expression vectors containing the GmHDA5 and GmHDA17 genes can be constructed using existing plant expression vectors. These plant expression vectors include, but are not limited to, binary Agrobacterium vectors and vectors suitable for plant microbombardment. The plant expression vectors may also contain the 3' untranslated region of the exogenous gene, i.e., containing a polyadenylate signal and any other DNA fragment involved in mRNA processing or gene expression. The polyadenylate signal can guide the addition of polyadenylate to the 3' end of the mRNA precursor; similar functions exist for the untranslated regions transcribed at the 3' end of Agrobacterium crown gall-inducing (Ti) plasmid genes (such as the Nos gene for lipase) and plant genes (such as the soybean storage protein gene).
[0031] When constructing recombinant plant expression vectors using the GmHDA5 and GmHDA17 genes, any enhancing or constitutive promoter can be added before the transcription initiation nucleotide, including but not limited to the cauliflower mosaic virus (CAMV) 35S promoter and the maize ubiquitin promoter. These can be used alone or in combination with other plant promoters. Furthermore, when constructing plant expression vectors using the genes of this invention, enhancers, including translational enhancers or transcriptional enhancers, can also be used. These enhancer regions can be ATG start codons or adjacent region start codons, but they must be identical to the reading frame of the coding sequence to ensure correct translation of the entire sequence. The sources of the translation control signals and start codons are wide-ranging; they can be natural or synthetic. The translation initiation region can originate from the transcription initiation region or structural genes.
[0032] In the above applications, the microorganism described in D4) or E5) can be yeast, bacteria, algae, or fungi. Specifically, it can be Agrobacterium tumefaciens EHA105. The recombinant Agrobacterium is EHA105 / pCBSG015-sgRNA.
[0033] In the above applications, the transgenic plant cell lines, transgenic plant tissues, and transgenic plant organs described in E6) may or may not include propagation material.
[0034] In the above applications, the plant tissue described in E7) may be derived from roots, stems, leaves, flowers, fruits, seeds, pollen, embryos, and anthers.
[0035] In the above applications, the transgenic plant organs described in E8) can be the roots, stems, leaves, flowers, fruits, and seeds of the transgenic plant.
[0036] Furthermore, in the aforementioned applications, the transgenic plant cell lines, transgenic plant tissues, and transgenic plant organs may or may not include propagation material.
[0037] In the above applications, the plant may be any of the following: N1) dicotyledonous plants; N2) legumes; N3) legumes; N4) soybeans; N5) soybean.
[0038] Secondly, the present invention provides a method, which may be M1, M2, M3 or M4: M1. A method for increasing the oil content of soybean seeds, the method comprising knocking out the coding gene of a soybean to be improved containing the protein coding gene described above to obtain gene knockout soybeans, wherein the oil content of the gene knockout soybeans is higher than that of the soybean to be improved. M2. A breeding method for producing soybeans with increased seed oil content, the method comprising knocking out the coding gene of the soybean to be improved containing the protein coding gene described above to obtain soybeans with increased seed oil content, wherein the seed oil content of the soybeans with increased seed oil content is higher than that of the soybean to be improved. M3. A method for increasing the oil content of soybean seeds, the method comprising reducing the activity and / or content of the aforementioned protein in the soybean to be improved, and / or reducing the expression level of the gene encoding the aforementioned protein, to increase the oil content of the soybean seeds to be improved. M4. A breeding method for producing soybeans with increased seed oil content, the method comprising reducing the activity and / or content of the proteins described above in the soybean to be improved, and / or reducing the expression level of the coding genes of the proteins described above, to obtain improved soybeans, wherein the seed oil content of the improved soybeans is higher than that of the soybeans to be improved.
[0039] In the above method, M1 and M2 include the step of introducing the substance described above that knocks out the gene encoding the protein into the soybean to be improved.
[0040] Specifically, in methods M1 and M2, the knockout includes performing any of the following operations on the genome of the soybean to be improved: F1) Delete the nucleotides at positions 2141-2144 of SEQ ID NO: 5 and positions 4787-4815 of SEQ ID NO: 6 in the soybean genome to be improved, and insert a DNA fragment with the following nucleotide sequence between positions 4786-4787 of SEQ ID NO: 6 in the soybean to be improved: 5´-ATTTGCTTTGGTATGA-3´; F2) Delete nucleotides at positions 2145-2147 and 2427-2434 in SEQ ID NO: 5 and 4758-4803 and 5075-5079 in SEQ ID NO: 6 from the genome of the target plant; F3) Delete nucleotides at positions 2431-2432 of SEQ ID NO: 5 and positions 5076-5080 of SEQ ID NO: 6 in the genome of the target plant, and insert nucleotide A between positions 2146-2147 of SEQ ID NO: 5.
[0041] Thirdly, the present invention provides a substance, which may be the protein or biological material described above.
[0042] Fourthly, the present invention provides gene-edited soybeans, wherein the gene-edited soybeans may be soybeans that do not contain the coding genes for the proteins described above.
[0043] The aforementioned gene-edited soybean can be a gene-knockout soybean obtained by knocking out the coding gene of the soybean to be improved, which contains the coding gene of the protein described above.
[0044] The seed oil content of the gene-edited soybean is higher than that of the soybean to be improved.
[0045] In some embodiments, the oil is at least one of palmitic acid, stearic acid, oleic acid, linoleic acid and / or linolenic acid.
[0046] In some specific embodiments, the gene-edited soybean is obtained by knocking out the coding gene of the soybean to be modified, which contains the coding gene of the protein.
[0047] In some specific embodiments, the knockout is achieved via a CRISPR-Cas system. The CRISPR-Cas system includes an RNA-directed nuclease (Cas protein) and sgRNA.
[0048] In this application, Cas protein is an "RNA-directed nuclease," which refers to an RNA-directed DNA endonuclease associated with the CRISPR system. Unrestricted examples of RNA-directed nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, their homologs or modified forms thereof. In one implementation, the RNA-directed nuclease is Cas9 or nCas9 (D10A).
[0049] The gene-edited soybeans are whole soybean plants or parts thereof (cells, tissues and / or organs).
[0050] The cells include original knockout cells (T0 generation), cells regenerated or developed from T0 generation cells, cells from any progeny or descendant of T0, including seed or embryo cells, or cultured cells, callus cells, etc.
[0051] The tissues and / or organs may be meristems, bud organs / structures (e.g., leaves, stems, or nodes), roots, flowers or floral organs / structures (e.g., flowers, bracts, sepals, petals, stamens, carpels, anthers, and ovules), seeds (e.g., embryos, endosperm, and seed coats), fruits (e.g., mature ovaries), propagules, or other plant tissues (e.g., vascular tissue, dermal tissue, ground tissue).
[0052] In some specific embodiments of the present invention, the soybean to be improved may be the soybean variety Tianlong No. 1 (TL).
[0053] This invention modifies soybeans GmHDA5 Genes and GmHDA17 Genes, via CRISPR / Cas9 vector, are used to target endogenous soybean genes. GmHDA5 or / and GmHDA17 Knockout of the gene yields single or double mutants, resulting in gene knockout soybean plants. hda1 , hda2 and hda3The seed oil content of the mutant plants was significantly higher than that of the recipient controls TL and Null plants. Statistical analysis of these traits indicates that GmHDA5 / GmHDA17 negatively regulates soybean seed oil content, reducing it. GmHDA5 / GmHDA17 The expression level of this substance can significantly increase the oil content of soybean seeds. Attached Figure Description
[0054] Figure 1 for GmHDA5 / GmHDA17 Gene-edited mutants hda1 The gene editing mode was shown. GmHDA5 and GmHDA17 Two target sequences (sgRNA1 and sgRNA2) and their role in GmHDA5 / GmHDA17 Specific location on the gene and hda1 Mutant gene editing was validated by sequencing.
[0055] Figure 2 for GmHDA5 / GmHDA17 Gene-edited mutants hda2 The gene editing mode was shown. GmHDA5 and GmHDA17 Two target sequences (sgRNA1 and sgRNA2) and their role in GmHDA5 / GmHDA17 Specific location on the gene and hda2 Mutant gene editing was validated by sequencing.
[0056] Figure 3 for GmHDA5 / GmHDA17 Gene-edited mutants hda3 The gene editing mode was shown. GmHDA5 and GmHDA17 Two target sequences (sgRNA1 and sgRNA2) and their role in GmHDA5 / GmHDA17 Specific location on the gene and hda3 Mutant gene editing was validated by sequencing.
[0057] Figure 4 for GmHDA5 / GmHDA17 Gene mutants hda1 , hda2 and hda3 In addition, molecular detection of gene expression in mid-development seeds of control TL and Null plants was performed.
[0058] Figure 5 for GmHDA5 / GmHDA17 Gene mutants hda1 , hda2 and hda3 In addition, the statistical analysis of fatty acid and oil content in TL and Null seeds was conducted. Detailed Implementation
[0059] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0060] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0061] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.
[0062] The soybean variety Williams 82 (W82) used in the following examples is an American cultivar, described in the following literature: Schmutz et al., Genome sequence of the palaeopolyploid soybean, Nature, 2010, 463(7278): 178-183; provided by the corresponding author of the paper, Scott A Jackson, Purdue University, Department of Agronomy.
[0063] Tianlong No. 1 (TL) was bred by the Oil Crops Research Institute of the Chinese Academy of Agricultural Sciences through hybridization of Zhongdou 32 and Zhongdou 29, and was approved by the State in 2008 (State Approval No. 2008023). It is available to the public from the Institute of Genetics and Developmental Biology of the Chinese Academy of Sciences. This biological material is intended solely for the replication of experiments related to this invention and may not be used for any other purpose.
[0064] The *Agrobacterium tumefaciens* EHA105 used in the following examples is described in: Rodrigues SD, Karimi M, Impens L, Van Lerberge E, Coussens G, Aesaert S, Rombaut D, Holtappels D, Ibrahim HMM, Van Montagu M, Wagemans J, Jacobs TB, De Coninck B, Pauwels L. Efficient CRISPR-mediated base editing in *Agrobacterium spp.* Proc Natl Acad Sci US A. 2021 Jan, 12;118(2): e2013338118. doi: 10.1073 / pnas.2013338118. This biological material is available to the public from the applicant and is intended solely for the purpose of replicating experiments of this invention and may not be used for any other purpose.
[0065] The gene editing was performed by Wimi Biotechnology Co., Ltd. (formerly known as Baige Biotechnology Co., Ltd.), and the pCBSG015 vector was provided by Wimi Biotechnology Co., Ltd., with the catalog number wimi-pCXB053.
[0066] The structure of the pCBSG015-sgRNA vector is described as follows: The pCBSG015 vector is linearized by digestion with Bsa I. The recombinant plasmid is obtained by replacing the fragment between 5'-ggcaccgagtcggtgc-3' and 5'-gttgaacaacggaaac-3' of the pCBSG015 vector with the DNA sequence of SEQ ID No:7, while keeping the other sequences of the pCBSG015 vector unchanged. The recombinant plasmid is named the recombinant vector pCBSG015-sgRNA. In SEQ ID NO:7, positions 1-441 represent the promoter ATU6-26, positions 442-464 represent the sgRNA1 gene, positions 465-540 represent gRNAscaffold, positions 541-558 represent polyA, positions 559-999 represent a repeat of the promoter ATU6-26, positions 1000-1022 represent the sgRNA2 gene, and positions 1023-1098 represent a repeat of gRNAscaffold. Details are as follows: The recombinant vector pCBSG015-sgRNA contains two editing target sites, 5'-CTTGCTGCTAGGTTGGCAGCTGG-3' and 5'-GGAAGGTGCTAATTTTGGATTGG-3', as well as the gene encoding the Cas9 protein on the vector. After being introduced into the receptor, the two transcribed guide RNAs can target the target sequence near the PAM in the receptor genome through base complementarity, i.e., targeting... GmHDA5 Genes and GmHDA17 Genes, Cas9 protein GmHDA5 and [[ID=…]] A double-strand break in DNA at the target site triggers a DNA damage repair response mechanism in the organism. During this repair process, gene mutations occur in the cleaved region, leading to large deletions, frameshift mutations, or premature termination of translation in the coding gene, thereby achieving the desired effect. and Knockout.
[0067] The experimental results of the following examples are expressed as mean ± standard deviation. A one-way ANOVA test was used, and P < 0.05 was considered satisfactory. () indicates a significant difference, P < 0.01. () indicates a highly significant difference.
[0068] The coding sequence (CDS) of the gene in soybean variety Williams 82 (W82) is SEQ ID NO: 1, and the amino acid sequence encoding the GmHDA5 protein is SEQ ID NO: 2. The genomic gene encoding the GmHDA5 protein in the genomic DNA of soybean variety W82 is shown in SEQ ID NO: 5 of the sequence listing.
[0069] The coding sequence (CDS) of the gene in soybean variety W82 is SEQ ID NO: 3, and the amino acid sequence of the GmHDA17 protein is SEQ ID NO: 4. The genomic gene encoding the GmHDA17 protein in the genomic DNA of soybean variety W82 is shown in SEQ ID NO: 6 of the sequence listing.
[0070] Example 1, Soybeans Acquisition of genes Downstream genes of the protein GmFVE, which negatively regulates soybean seed oil content, were screened. This gene encodes histone deacetylase 5 (HDA5), a member of the histone deacetylase family. HDA5 is an enzyme that catalyzes the deacetylation of histones, altering chromatin structure and function by removing acetyl groups from lysine residues of histones, thus playing a crucial role in gene expression regulation. Further analysis revealed that the histone deacetylase family also includes GmHDA17 (Glyma.17g120900), which has 126 more amino acids than GmHDA5, but shows lower full-length similarity (69%). However, the histone deacetylation domains of the two proteins share 97% similarity, suggesting that they may have similar functions. Currently, there are no reports of its involvement in the regulation of lipid accumulation.
[0071] According to the Williams 82 (W82) reference genome sequence (Genome assembly Glycine_max_v4.0, Genbank: GCA_000004515.5, updated March 10, 2021) and Based on the full-length cDNA sequence information, primers were designed, and the primer sequences are as follows: -up: 5'-ATGCATGTTGGTTCACAGGAA-3'; -dp:5'-TCACATAACTTTGATTTTACATTATGG-3'.
[0072] -up: 5'-ATGGTTAGTGGAGCTCTCAAAAC-3'; -dp:5'-TTAAGGTGGCCGTTCAGAAA-3'.
[0073] RNA was extracted from soybean variety W82 seedlings and reverse transcribed into cDNA using Toyobo's ReverTra Ace reverse transcriptase. The cDNA was then used as a template for further processing. -up / -dp and -up / PCR amplification was performed using -dp primers. PCR was applied to amplify total RNA from soybeans. and Gene extraction: Leaf samples were crushed in liquid nitrogen, suspended in 4 mol / L guanidine thiocyanate, and extracted with acidic phenol and chloroform. Anhydrous ethanol was added to the supernatant to precipitate the RNA, which was then dissolved in water to obtain total RNA. 1 µg of total RNA was reverse transcribed using a Thermo Fisher Scientific kit according to the kit's instructions. The resulting cDNA fragment was then used as a template for PCR amplification.
[0074] The 50 μl PCR reaction mixture consisted of: 1 μl single-stranded cDNA (0.05 μg), 1.5 μl of the above primers (10 μM), 25 μl 2× PCR buffer, 10 μl dNTPs (10 mM), and 1 U KOD DNA polymerase, with the volume made up to 50 μl with ultrapure water. The reaction was performed on a PE9600 PCR instrument with a 94-hour program. o Denaturation at C for 5 minutes; then at 98°C o C 1min, 58 o C 1min, 68 o C 1min, 30-32 cycles in total; then 68 o C extended for 10 minutes; 4 o Store at C. A PCR product of approximately 1.3 kb was obtained. The amplified product was detected by 1% agarose gel electrophoresis. A DNA fragment of approximately 1.5 bp in length was recovered using an agarose gel recovery kit (TIANGEN), cloned into the vector pMD-18 (TaKaRa, catalog number 6011), transformed into competent E. coli cells, positive clones were screened, and plasmids were extracted for sequencing.
[0075] Sequencing results show that: -up and The PCR product amplified by the -dp primer is The CDS sequence (SEQ ID NO: 1, 1254bp) is as follows:
[0076] The amino acid sequence of the translated protein GmHDA5 is sequence 2 in the sequence listing (SEQ ID NO: 2, 417aa), as follows: MHVGSQESVANSSTAIGFDERMLLHAEVEKKSTLHPERPDRLQAIAASLARAGIFPGKCYSIPAREITPEELITVHSLEHIESVEVTSESLSSYFTSDTYANEHSALAARLAAGLCADLASAIVSGRAKNGFALVRPPGHHAGVRQAMGFCLHNNAAVAALAAQAAGARKVLILDWDVHHGNGTQEIFEQNKSVLYISLHRHEGGKFYP GTGAAEEVGSMGAEGFCVNIPWSQGGVGDNDYIFAFQHVVLPIAAEFNPDLTIVSAGFDAARGDPLGCCDITPSGYAHMTNMLNALSGGKLLVILEGGYNLRSI SSSATAVIKVLLGESPGCELENSFPSKAGLQTVLEVLKIQMNFWPALGPIFVNLESQWRMYCFERKRKQIKKRRRVLVPMWRWGRKSLLFHFLNGHHNVKSKLC.
[0077] The genome sequence is SEQ ID NO: 5 (5643bp), of which positions 5381-5643 are 3'UTRs, and it also contains 15 exon sequences located at positions 1-4, 328-404, 518-593, 689-753, 1600-1655, 2088-2211, 2314-2439, 2556-2606, 2818-2886, 3465-3573, 4099-4178, 4420-4502, 4589-4634, 4862-5033, and 5265-5380; the rest are introns.
[0078] Sequencing results show that: -up and The PCR product amplified by the -dp primer is The CDS sequence (SEQ ID NO: 3, 1632bp). SEQ ID NO: 3 is as follows:
[0079] The amino acid sequence of the translated protein GmHDA17 is sequence 4 in the sequence listing (SEQ ID NO: 4, 543aa), as follows: .
[0080] The genome sequence is SEQ ID NO: 6 (9110 bp), where positions 1-389 are the 5' UTR; positions 8077-9110 are the 3' UTR; it also includes 17 exon sequences located at positions 390-485, 959-1156, 1377-1488, 2512-2588, 3461-3536, and 3615-3679, respectively. Positions 4170-4225, 4735-4858, 4960-5085, 5195-5245, 5463-5531, 6163-6271, 6782-6861, 7111-7193, 7279-7324, 7547-7718, 7985-8076, and the rest are introns.
[0081] Example 2: / Gene mutation leads to increased oil content in soybean seeds 1. Knock out the GmHDA family gene in soybean Tianlong No. 1 (TL) and .
[0082] 1.1 Construction of pCBSG015 gene editing vector Target sequence selection: The high-throughput CRISPR-Cas9 target design program developed by Wemi Technology was used. The target design principles of this program are as follows: 1) Knockout sites should be located in the coding sequence (CDS) region and preferably at the protein's front end or in an important functional domain; 2) Coverage should be maximized to include a higher proportion of transcripts; 3) No off-target effects or off-target effects should be located in intergenic regions; 4) Targets with higher editing efficiency should be preferred; 5) The sequence should have a relatively balanced GC content and be less prone to secondary structure formation. The successful application of this program in whole-genome target design in rice has proven its feasibility.
[0083] A dual-gene, dual-target knockout approach was adopted. The selected target site T1: 5'-CTTTGCTGCTAGGTTGGCAGCTGG -3' is located at... The sixth exon region (positions 2128-2150 of SEQ ID NO: 5) corresponds to positions 319-341 of SEQ ID NO: 1; located in The eighth exon region (positions 4775-4797 of SEQ ID NO: 6) corresponds to positions 721-743 of SEQ ID NO: 3.
[0084] T2: 5'- GGAAGGTGCTAATTTTGGATTGG -3' located at The seventh exon region (positions 2417-2439 of SEQ ID NO: 5) corresponds to positions 506-528 of SEQ ID NO: 1; located in The ninth exon region (positions 5063-5085 of SEQ ID NO: 6) corresponds to positions 908-930 of SEQ ID NO: 3.
[0085] Promoter selection: AtU6 derived from Arabidopsis thaliana was used to promote the T1 and T2 target sites.
[0086] Preparation of sgRNA expression cassettes containing target sites: pCBSG015 vector was linearized by Bsa I restriction enzyme digestion. T1 and T2 sequences were directly synthesized using primer synthesis method, and 16bp vector sequences were added to both ends as homologous arms (U6-T1, U6-T2). Reverse complementary sequences (Anti-U6-T1, Anti-U6-T2) were synthesized, annealed to form double strands, and homologously recombined with the backbone linear vector.
[0087] The specific steps are as follows: Sense-U6-T1: 5'-ggcaccgagtcggtgcCTTGCTGCTAGGTTGGCAGCTGGgttgaacaacggaaac-3'; Anti-U6-T1: 5'-gtttccgttgttcaacCCAGCTGCCAACCTAGCAGCAAGgcaccgactcggtgcc-3'; Lowercase letters represent homologous arms, and uppercase letters represent T1 sequences.
[0088] Sense-U6-T2: 5'-ggcaccgagtcggtgcGGAAGGTGCTAATTTTGGATTGGgttgaacaacggaaac-3'; Anti-U6-T2: 5'-gtttccgttgttcaacCCAATCCAAAATTAGCACCTTCCgcaccgactcggtgcc-3'; Lowercase letters represent homologous arms, and uppercase letters represent T2 sequences.
[0089] Preparation of annealed AtU6-T1-gRNA and AtU6-T2-gRNA fragments: The synthesized Sense-U6-T1 and Anti-U6-T1, Sense-U6-T2 and Anti-U6-T2 sequences were annealed to form double strands. The reaction system consisted of dissolving the synthesized sequences in 75 mM NaCl solution to a final concentration of 0.2 nM / μl, and mixing equal volumes of the forward and reverse strand solutions (Sense-U6-T1 and Anti-U6-T1, Sense-U6-T2 and Anti-U6-T2). The mixture was heated in a 95°C water bath for 5–10 min, followed by slow cooling to obtain the annealed AtU6-T1-gRNA and AtU6-T2-gRNA fragments.
[0090] Vector linearization by restriction enzyme digestion: 1-2 μg of pCBSG015 plasmid, 5 μl of 10X CutSmart™ buffer (NEB), 1 μl of Bsa I restriction enzyme, and sterile double-distilled water to a final volume of 50 μl. The mixture was incubated at 37°C for 30 min, then purified using the EZ-10 Column DNA Purification Kit (Shanghai Sangon Biotech). The purified DNA was dissolved in an appropriate amount of water to obtain the linearized pCBSG015 vector.
[0091] Homologous recombination of the target sgRNA expression cassette with the pCBSG015 vector was performed using the EasyGeno Rapid Recombinant Cloning Kit (TianGen). The reaction mixture consisted of 5 μL of 2×EasyGeno Assembly mix buffer, 0.5 μL of linearized pCBSG015 vector, and 4.5 μL of annealed AtU6-T1-gRNA and AtU6-T2-gRNA. The reaction conditions were 50 °C for 15 min. The ligation product was transformed into E. coli DH5α competent cells, and plasmids were extracted from positive colonies. After successful sequencing, the recombinant vector pCBSG015-sgRNA was obtained.
[0092] The recombinant vector pCBSG015-sgRNA is described as follows: It is a recombinant plasmid obtained by replacing the fragments between 5'-ggcaccgagtcggtgc-3' and 5'-gttgaacaacggaaac-3' in the pCBSG015 vector with the DNA sequence of SEQ ID No:7, while keeping the other sequences of the pCBSG015 vector unchanged. This recombinant plasmid is named recombinant vector pCBSG015-sgRNA. In SEQ ID NO:7, 1098 bp, positions 1-441 are the promoter ATU6-26, positions 442-464 are the sgRNA1 gene, positions 465-540 are gRNAscaffold, positions 541-558 are polyA, positions 559-999 are a repeat of the promoter ATU6-26, positions 1000-1022 are the sgRNA2 gene, and positions 1023-1098 are a repeat of the gRNAscaffold. Specifically as follows: The recombinant vector pCBSG015-sgRNA contains the genes encoding the editing targets T1, T2, and the Cas9 protein on the vector. After being introduced into the recipient, the transcribed guide RNA can target the target sequence near the PAM in the recipient genome through base complementarity pairing, i.e., targeting... and Genes, Cas9 protein and A double-strand break in DNA at a gene target site triggers a gene mutation in the cut region through the organism's own DNA damage repair response mechanism. This mutation leads to large deletions, frameshift mutations, or premature termination of translation in the coding gene, thereby achieving the desired effect. and Gene knockout.
[0093] The recombinant vector pCBSG015-sgRNA was transformed into Agrobacterium EHA105 competent cells. Agrobacterium EHA105-pCBSG015-sgRNA was obtained.
[0094] 1.2 Genetic transformation of soybean The *Agrobacterium* EHA105-pCBSG015-sgRNA prepared above was used to infect the soybean variety Tianlong No. 1 (TL) to obtain the proposed... / Genetically edited soybean plants were used to harvest T0 generation seeds. T1 generation seeds were obtained by self-pollination of T1 generation plants. T2 generation seedlings were obtained by planting T1 generation seeds and then subjected to the following tests.
[0095] Null plants were isolated from heterozygous gene-edited lines and do not contain the pCBSG015 vector. / Unedited plants, along with the recipient material TL, served as control materials.
[0096] 1.3 / Screening for homozygous gene mutations Using the genomic DNA of the T2 generation seedlings obtained in section 1.2 as a template, screening and identification were performed using mutant target detection primers. Primers were designed using approximately 222 bp upstream of the target sequence T1 and approximately 111 bp downstream of the target sequence T2. Primers for gene target detection. The primers for gene target detection are: -F: 5'-TACTTGTACCCTGTTTGGTCT -3' and -R: 5'-CCCATTTCCATGATGAACATCCTGTA -3' Amplification product is approximately 692bp; a design was created using approximately 316bp upstream of the target sequence T1 and approximately 104bp downstream of the target sequence T2. Primers for gene target detection, The primers for gene target detection are: -F: 5'-GACTTGTACCCTGTTTGGTTC -3' and -R: 5'-CATTTCCGTGATGAACATCCTGTG-3' The amplified product is approximately 776bp.
[0097] After successful sequencing, the CRISPR target editing method was analyzed using the website DSDecode (http: / / dsdecode.scgene.com / ) and compared with the standard gene sequence using manual peak reading. The editing methods of each target sequence and its upstream and downstream sequences were analyzed and passaged.
[0098] Filtered by the above method / homozygous mutant , and The gene editing methods are respectively , 2 And 3.
[0099] , 2 3 and 3 respectively show / Gene-edited mutants , and The construction process of sgRNA1 (T1) and sgRNA2 (T2) is shown in the figure. and The position of the sequence and the sequence spliced by the mutant. The illustration shows the three mutants. and All achieved the desired editing effect.
[0100] Compared to the recipient soybean variety Tianlong No. 1 (TL), / homozygous mutant In the genome of the strain The gene underwent the following mutation: In both homologous chromosomes, positions 2141-2144 of the genome sequence (SEQ ID NO: 5), corresponding to positions 332-335 of SEQ ID NO: 1, contained a 4-nucleotide 5'-TGGC-3' deletion, resulting in… The translation was terminated early, thus... Gene knockout; and At positions 4786-4787 of the genome sequence (SEQ ID NO: 6), corresponding to positions 732-733 of SEQ ID NO: 3, there is a 16 bp insertion of nucleotide 5'-ATTTGCTTTGGTATGA-3'; simultaneously, at positions 4787-4815 of the genome sequence (SEQ ID NO: 6), corresponding to positions 733-761 of SEQ ID NO: 3, there is a 29 nucleotide deletion of nucleotide 5'-TTGGCAGCTGGGTTATGTGCTGATCTAGC-3', resulting in... Large insertions and deletions, ultimately causing the translation to terminate prematurely, thus... Gene knockout.
[0101] Compared to the recipient soybean variety Tianlong No. 1 (TL), / homozygous mutant In the genome of the strain The gene underwent the following mutation: In two homologous chromosomes, at positions 2145-2147 and 2427-2434 of the genomic sequence (SEQ ID NO: 5), corresponding to positions 336-338 and 517-524 of SEQ ID NO: 1, there were deletions of 3 nucleotides 5'-AGC-3' and 8 nucleotides 5'-ATTTTGGA-3', respectively, leading to… The translation was terminated early, thus... Gene knockout; At positions 4758-4803 and 5075-5079 of the gene genome sequence (SEQ ID NO: 6), corresponding to positions 704-749 and 920-924 of SEQ ID NO: 3, there are deletions of 46 nucleotides 5'-CCAATGAGCATTCAGCACTTGCTGCTAGGTTGGCAGCTGGGTTATG-3' and 5 nucleotides 5'-TTTTG-3', respectively, leading to... Large segments are missing, and the translation was terminated prematurely, thus... Gene knockout.
[0102] Compared to the recipient soybean variety Tianlong No. 1 (TL), / homozygous mutant In the genome of the strain The gene underwent the following mutations: In both homologous chromosomes, at positions 2146-2147 of the genome sequence (SEQ ID NO: 5), corresponding to positions 337-338 of SEQ ID NO: 1, there was an insertion of nucleotide A; and at positions 2431-2432 of the genome sequence (SEQ ID NO: 5), corresponding to positions 521-522 of SEQ ID NO: 1, there was a deletion of 2 nucleotides 5'-TG-3', resulting in… The translation was terminated early, thus... Gene knockout; At positions 5076-5080 of the gene genome sequence (SEQ ID NO: 6), corresponding to positions 921-925 of SEQ ID NO: 3, there is a 5-nucleotide 5'-TTTGG-3' deletion, resulting in... The translation was terminated early, thus... Gene knockout.
[0103] mutant , and Seeds from T2 generation plants (T3 generation) were harvested for subsequent experiments.
[0104] 2. / Gene CRISPR materials , and Phenotypic analysis 2.1 , and middle / Detection of gene expression levels Extracting soybean varieties TL, Null and , and Total RNA from mid-seed development was reverse transcribed, and the resulting cDNA was used as a template for Real-Time PCR identification. / Gene expression levels. Primers used: -qF:5'-AGGATGTTGTGATATTACTCCCTCT-3', -qR:5'- AGGACTCTCACCCAATAATACCT -3'; -F:5'-CAGAAGTTGAAAGAAATCCACTCT -3', -R:5'-TGAAATAACTAGAAAGTGATTCGCT-3'.
[0105] Using the soybean Tubulin gene as an internal standard, the internal standard primers are: Primer-TF: 5′-TGGCCGTTACCTGACAGCAT-3′; Primer-TR: 5′-CTCGGAGGGATGTCACACAC-3′.
[0106] The results are as follows As shown: in TL GmHDA5 The relative expression level was 1.13 ± 0.09 in Null. GmHDA5 The relative expression level was 1.08 ± 0.16. hda1 , hda2 and hda3 middle GmHDA5 The expression levels were approximately 0.56±0.19, 0.51±0.26, and 0.41±0.06, respectively, indicating that... hda1 , hda2 and hda3 In mutants GmHDA5 The expression levels of TL, Null, and others decreased significantly. hda1 , hda2 and hda3 middle GmHDA17 The expression levels were 0.39±0.04, 0.37±0.10, 0.26±0.035, 0.17±0.04, and 0.12±0.04, respectively, indicating that... hda1 , hda2 and hda3 middle GmHDA17 The expression level decreased significantly or extremely significantly.
[0107] 2.2 GmHDA5 / GmHDA17 Gene downregulation increases soybean seed oil content The above hda1 , hda2 and hda3 Seeds from the T2 generation plants of the transgenic event, along with control seeds of Tianlong No. 1 (TL) and Null, were sown in greenhouse pots. Greenhouse conditions included 16 hours of light at 11000 Lux and 8 hours of darkness, with daytime temperatures ranging from 30 to 37°C and nighttime temperatures from 25 to 28°C. Growth and development were observed. Seeds were harvested after 130 days. hda1 , hda2 and hda3 No significant differences were observed in the phenotypes of the control TL and Null during the growth and maturity stages.
[0108] After the potted soybean seeds matured, the seeds harvested from individual plants were dried at 37°C for one week, and the receptor control TL, Null, and mutant levels were measured. hda1 , hda2 and hda3 Seed oil content. Fifteen plants were taken from each line, and the biological experiment was repeated three times. Results are expressed as mean ± standard deviation. One-way ANOVA was used, and P < 0.05 was considered acceptable. () indicates a significant difference, P < 0.01. () indicates a highly significant difference.
[0109] 1) The specific steps for determining the total oil content in seeds are as follows: Grind the dried seeds into powder, weigh 100 mg into four parallel centrifuge tubes. Add 500 μl of n-hexane, mix thoroughly, and incubate overnight at 37°C. Centrifuge slowly for 3 minutes, transferring the n-hexane into a newly weighed centrifuge tube. Repeat the soaking process with n-hexane on the remaining powder, then centrifuge and collect the n-hexane into the same centrifuge tube. Place the centrifuge tube in a vacuum pump and evacuate to allow the n-hexane to evaporate completely. Weigh the centrifuge tube again. The change in weight before and after centrifugation represents the weight of extracted lipids, and the total oil content (%) is calculated. The formula for calculating the total oil content (%) is as follows: Total oil content (%) = (Weight of extracted lipids / Total weight of seeds) × 100%.
[0110] Figure 5 Left display, comparing TL, Null, and hda1 , hda2 and hda3 The total oil content of the seeds was statistically analyzed to be 21.8±0.8%, 21.4±0.4%, 23.5±1.2%, 24.1±0.9%, and 25.0±0.7%, respectively. Mutant hda1 , hda2 and hda3 The total oil content in the seeds of the mutants was significantly higher than that of the controls TL and Null. The total oil content in the mutant seeds increased by 8.8%, 11.6%, and 15.7% compared to the controls, respectively.
[0111] 2) Determination of fatty acid content in seeds The determination of various fatty acids was performed according to the method described by Sukhija and Palmquist (PS Sukhija and DLPalmquist (1988) Rapid Method for Determination of Total Fatty Acid Content and Composition of Feedstuffs and Feces. J Agric Food Chem, 36: 1202-1206): After thorough grinding of the seeds, 10 mg was weighed and added to 1 ml of sulfuric acid methanol solution (2.5 ml sulfuric acid / 100 ml methanol) and 10 µl of internal standard solution (100 µg 17-alkyl acid / ml ethyl acetate). The mixture was incubated at 100 °C for 1 h. The supernatant was collected, and 500 µl of n-hexane and 300 µl of 10% (w / w) NaCl aqueous solution were added to each 1 ml of supernatant. The mixture was then dried under vacuum, and 30 µl of ethyl acetate was added. The supernatant was then subjected to gas chromatography (AGILENT 6890). The determination was performed using the GC method, with Sigma-Aldrich standards for palmitic acid, stearic acid, oleic acid, linoleic acid, and linolenic acid as standard controls.
[0112] In the GC analysis, palmitic acid, stearic acid, oleic acid, linoleic acid and linolenic acid from Sigma were used as standards. The retention time of the standards was used for qualitative analysis and the standard curve method (external standard method) was used for quantitative analysis of palmitic acid, stearic acid, oleic acid, linoleic acid and linolenic acid.
[0113] The gas chromatography parameters are as follows: column inner diameter 0.25 mm, length 30 m; carrier is ethyl acetate with a particle size of 0.25 µm; carrier gas is helium; column temperature is 120 °C, injection port temperature is 120 °C, injection volume is 1 µl, and carrier gas flow rate is 2 ml / min.
[0114] Figure 5 The right side displays TL, Null, and hda1 , hda2 and hda3 The palmitic acid (16:0) content in the seeds was approximately 1.8, 1.7, 2.2, 2.0, and 2.3, respectively; the stearic acid (18:0) content was approximately 0.50, 0.50, 0.70, 0.75, and 0.71, respectively. Both of these fatty acids were significantly or highly significantly higher in the mutant seeds compared to the control. TL, Null, and hda1 , hda2 and hda3 The oleic acid (18:1) content in the seeds was approximately 3.3, 3.3, 3.8, 4.1, and 5.3%, with the mutants showing significantly higher levels than the control. TL, Null, and hda1 , hda2 andhda3 The linoleic acid (18:2) content in the seeds was approximately 9.7, 9.1, 10.3, 10.3 and 10.5; the linolenic acid (18:3) content was approximately 1.6, 1.6, 1.9, 1.9 and 1.8. The content of the above two fatty acids in the mutant seeds was significantly or extremely significantly higher than that in the control.
[0115] The above statistics show that three biological replicate experiments on 15 individual plants grown in greenhouse pots indicated that the mutant... hda2 , hda2 and hda3 The total oil content in seeds was significantly higher than that of the receptor TL and Null controls, indicating that GmHDA5 / GmHDA17 negatively regulates soybean seed oil content and reduces the levels of two genes in the GmHDA family. GmHDA5 and GmHDA17 The expression level of this substance can significantly increase the oil content of soybean seeds.
[0116] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. The use of a protein, or a substance that knocks out the gene encoding the protein, or a substance that regulates the expression of the gene encoding the protein, or a substance that regulates the activity or content of the protein, characterized in that, The application is any one of the following: A1) Application in regulating the oil content of plant seeds; A2) Application in the preparation of products that regulate the oil content of plant seeds; A3) Application in cultivating plants with altered seed oil content; A4) Application in the preparation of products from plants whose seed oil content has been altered; A5) Applications in plant breeding; The protein is a composition, protein A, or protein B, wherein the composition consists of protein A and protein B; The protein A is GmHDA5, which is any of the following proteins: B1) The amino acid sequence of this protein is SEQ ID NO:
2. B2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 2, but with substitution and / or deletion and / or addition of amino acid residues. Proteins that share more than 80% amino acid sequence identity with B1) and B2) and have the same function, B4) A fusion protein obtained by attaching a tag to the end of any of the proteins defined in B1)-B3); The protein B is GmHDA17, which is any of the following proteins: C1) The amino acid sequence is that of the protein with sequence SEQ ID NO:
4. C2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 4, but with substitution and / or deletion and / or addition of amino acid residues. Proteins whose amino acid sequences defined by C3) are more than 80% identical to those defined by C1) and C2) and have the same function C4) is a fusion protein obtained by attaching a tag to the end of any of the proteins defined in C1)-C3).
2. The application according to claim 1, characterized in that, The protein is derived from soybeans.
3. The application according to claim 1 or 2, characterized in that, The substance that knocks out the gene encoding the protein is any one of the following: D1) A nucleic acid molecule that targets the gene encoding the protein described in claim 1 or 2; D2) Nucleic acid molecules and Cas proteins that target the gene encoding the protein described in claim 1 or 2; D3) A nucleic acid molecule encoding a gene that targets the protein described in claim 1 or 2 and a gene that encodes the Cas protein; D4) A biological material containing the nucleic acid molecule described in D1) and the Cas protein encoding gene, wherein the biological material is a vector, expression cassette, or recombinant microorganism; The substance that regulates the expression of the gene encoding the protein or the substance that regulates the activity or content of the protein is a biological material, and the biological material is any one of the following: E1) A nucleic acid molecule that inhibits, reduces, or silences the expression of the gene encoding the protein of claim 1 or 2; E2) expresses the gene encoding the nucleic acid molecule described in E1); E3) contains an expression cassette containing the gene encoding the gene described in E2); E4) A recombinant vector containing the encoding gene described in E2), or a recombinant vector containing the expression cassette described in E3; E5) Recombinant microorganisms containing the encoding gene described in E2), or recombinant microorganisms containing the expression cassette described in E3), or recombinant microorganisms containing the recombinant vector described in E4); E6) A transgenic plant cell line containing the encoding gene described in E2), or a transgenic plant cell line containing the expression cassette described in E3); or a transgenic plant cell line containing the recombinant vector described in E4; E7) Transgenic plant tissue containing the encoding gene described in E2), or transgenic plant tissue containing the expression cassette described in E3), or transgenic plant tissue containing the recombinant vector described in E4; E8) A transgenic plant organ containing the encoding gene described in E2), or a transgenic plant organ containing the expression cassette described in E3), or a transgenic plant organ containing the recombinant vector described in E4).
4. The application according to any one of claims 1-3, characterized in that: The plant is any one of the following: N1) Dicotyledons; N2) Leguminosae; N3) Leguminosae (family legumes); N4) Soybean genus plants; N5) soybeans.
5. The method, characterized in that: The method is M1, M2, M3, or M4: M1. A method for increasing the oil content of soybean seeds, the method comprising knocking out the coding gene of a soybean to be improved containing the protein coding gene of claim 1 to obtain gene knockout soybeans, wherein the oil content of the gene knockout soybeans is higher than that of the soybean to be improved. M2. A breeding method for producing soybeans with increased seed oil content, the method comprising knocking out the coding gene of the soybean to be improved containing the protein coding gene of claim 1 to obtain soybeans with increased seed oil content, wherein the seed oil content of the soybeans with increased seed oil content is higher than that of the soybeans to be improved. M3. A method for increasing the oil content of soybean seeds, the method comprising reducing the activity and / or content of the protein described in claim 1 in the soybean to be improved, and / or reducing the expression level of the gene encoding the protein described in claim 1, to increase the oil content of the soybean seeds to be improved; M4. A breeding method for producing soybeans with increased seed oil content, the method comprising reducing the activity and / or content of the protein described in claim 1 in the soybean to be improved, and / or reducing the expression level of the gene encoding the protein described in claim 1, to obtain improved soybeans, wherein the seed oil content of the improved soybeans is higher than that of the soybean to be improved.
6. The method according to claim 5, characterized in that: M1 and M2 include the step of introducing the substance that knocks out the gene encoding the protein as described in claim 3 into the soybean to be improved.
7. A substance, characterized by: The substance is the protein described in claim 1 or 2, or the biological material described in claim 3.
8. Gene-edited soybeans, characterized by: The gene-edited soybean is a soybean that does not contain the gene encoding the protein described in claim 1.
9. The gene-edited soybean according to claim 8, characterized in that: The gene-edited soybean is a gene-knockout soybean obtained by knocking out the coding gene of the soybean to be improved, which contains the coding gene of the protein described in claim 1.
10. The gene-edited soybean according to claim 8 or 9, characterized in that: The seed oil content of the gene-edited soybean is higher than that of the soybean to be improved.