Soybean protein GmFVE and application of coding gene thereof in regulation and control of oil content

By regulating the expression or activity of GmFVE protein in soybeans, gene editing technology was used to increase the oil content in soybean seeds, solving the problem of unclear regulatory mechanisms of plant oil synthesis and achieving a significant increase in the oil content of soybean seeds.

CN121874248APending Publication Date: 2026-04-17INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
Filing Date
2026-01-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The regulatory mechanism of plant oil synthesis in existing technologies has not been fully elucidated, making it difficult to increase the oil content of soybean seeds.

Method used

By regulating the expression or activity of GmFVE protein in soybeans, and using gene editing technologies such as CRISPR/Cas9 to knock out or silence the GmFVE gene, the oil content of seeds can be increased.

Benefits of technology

It significantly increases the oil content in soybean seeds, including the content of fatty acids such as palmitic acid, stearic acid, oleic acid and linoleic acid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses application of soybean protein GmFVE and a coding gene thereof in regulation and control of oil content. The invention belongs to the technical field of biology, and particularly discloses application of protein GmFVE and a coding gene of the protein or a substance for regulating and controlling expression of the protein or a substance for regulating and controlling activity or content of the protein in regulating and controlling grease content of plant seeds. The protein sequence of the soybean protein GmFVE is as shown in SEQ ID NO: 2. Experiments prove that the oil content in the seeds of gene edited homozygous strains fve1, fve2 and fve3 obtained by gene knockout of GmFVE is obviously higher than that of control plant soybean variety Jack and Null plants, the content of various fatty acids is increased, the protein GmFVE negatively regulates the oil content of the soybean seeds, and the gene knockout GmFVE has important theoretical significance for high-quality soybean creation and soybean breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to the application of soybean protein GmFVE and its encoding gene in regulating oil content. Background Technology

[0002] 71% of the fatty acids in the human diet come from plants. Among the world's major oil-producing crops, soybeans account for about 30% of the total oil production, ranking first in the world's vegetable oil production. Therefore, increasing the oil content in soybean seeds has become one of the main goals of soybean breeding.

[0003] Seed oil content is one of the main traits in the domestication process from wild soybean to cultivated soybean. To date, we have a considerable understanding of lipid synthesis pathways and have cloned a number of genes involved in lipid synthesis. However, the regulatory mechanisms of lipid synthesis in plants and the related genes are still not fully understood. Therefore, identifying genes that regulate oil accumulation and promoting efficient oil accumulation in plants is of great significance for increasing oil yields, especially in oilseed crops.

[0004] FVE / MSI4 (Glyma.09G063100) possesses a WD40 repeat domain and exhibits high homology with the mammalian RbAp protein (retinoblastoma-associated protein). It is a component of the histone deacetylase (HDAC) complex and participates in transcriptional repression. In Arabidopsis, FVE / MSI4 modulates key flowering genes... FLC Maintaining low histone H3ac levels and inhibiting FLC The expression of this protein regulates plant flowering. In 2004, Jungmook Kim et al. found that it not only regulates flowering time but is also related to low-temperature response. There are relatively few studies on the function of this protein in soybeans. There are no reports of its involvement in regulating seed oil content. Summary of the Invention

[0005] The main problem this invention aims to solve is to increase the oil yield of soybean seeds for use in soybean breeding.

[0006] To address the aforementioned problems, in a first aspect, the present invention provides the use of a protein, a substance regulating the expression of a gene encoding said protein, or a substance regulating the activity or content of said protein in any of the following:

[0007] A1) Application in regulating the oil content of plant seeds; A2) Application in the preparation of products that regulate the oil content of plant seeds; A3) Application in cultivating plants with altered seed oil content; A4) Application in the preparation of products from plants whose seed oil content has been altered; A5) Applications in plant breeding; The protein is GmFVE, and can be any of the following proteins: B1) The amino acid sequence of the protein is as shown in SEQ ID NO: 2. B2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 2, but with one or more amino acid residues substituted and / or deleted and / or added. Proteins that share more than 80% amino acid sequence identity with B1) or B2) and have the same function, B4) A fusion protein obtained by attaching a tag to the end of any of the proteins defined in B1)-B3).

[0008] In the above applications, the protein may be derived from soybeans.

[0009] In the above applications, the substance regulating gene expression can be a substance that performs at least one of the following six types of regulation: 1) regulation at the gene transcription level; 2) post-transcriptional regulation of the gene (i.e., regulation of splicing or processing of the primary transcript of the gene); 3) regulation of RNA transport of the gene (i.e., regulation of mRNA transport of the gene from the nucleus to the cytoplasm); 4) regulation of gene translation; 5) regulation of mRNA degradation of the gene; and 6) post-translational regulation of the gene (i.e., regulation of the activity of the protein translated from the gene).

[0010] In this invention, the indicators for plant breeding include the oil content of plant seeds, and the purpose of plant breeding includes cultivating plants with altered seed oil content.

[0011] In this invention, the substance that regulates the expression of the GmFVE protein-encoding gene can be a substance that inhibits, reduces, or downregulates the expression of the gene, and the substance that regulates the oil content of plant seeds can be a substance that increases the oil content in plant seeds.

[0012] The proteins mentioned above can be synthesized artificially, or their encoding genes can be synthesized first and then expressed biologically.

[0013] In this invention, the protein tag refers to a polypeptide or protein fused with a target protein using in vitro DNA recombination technology for expression, to facilitate the expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag, etc.

[0014] In this invention, the identity refers to the identity of amino acid sequences or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Per residue gap cost, and Lambdaratio to 11, 1, and 0.85 (default values) respectively, and performing an identity search on a pair of amino acid sequences to calculate the identity value (%), then the identity value can be obtained.

[0015] In this invention, the 80% or more of identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.

[0016] The SEQ ID NO: 2 may consist of 513 amino acid residues, as detailed below: METPPPQQGVVKKKETRGRKPKPKDEHGKGLKEGRKTQQQQQQQQQHHHQQQQQQDQPSVDEKYTQWKSLVPVLYDWLANHNLVWPSLSCRWGPQLEQATYKNRQRLYLSEQTDGSVPNTLVIANCE VVKPRVAAAEHISQFNEEARSPFVKKYKTIIHPGEVNRIRELPQNSKIVATHTDSPDVLVWDVESQPNRHAVLGATNSRPDLILTGHQDNAEFALAMCPTEPYVLSGGKDKTVVLWSIEDHITSAATDS KSGGSIIKQNSKSGEGNDKTADGPTVGPRGIYCGHEDTVEDVAFCPSSAQEFCSVGDDSCLILWDARVGSSPVVKVEKAHNADLHCVDWNPHDDNLILTGSADNSVRMFDRRNLTTNGVGSPIHKFEG HKAAVLCVQWSPDKSSVFGSSAEDGLLNIWDYEKVGKKIERSGKSISSPPGLFFQHAGHRDKVVDFHWNAYDPWTIVSVSDDCESTGGGGTLQIWRMSDLIYRPEDEVLAELEKFKSHVVACASKTEK.

[0017] In this invention, the substance regulating the expression of the protein-coding gene or the substance regulating the activity or content of the protein can be a biological material, and the biological material can be any one of the following C1) to C8): The biomaterial is any one of the following C1) to C8): C1) Nucleic acid molecules that inhibit, reduce, or downregulate the expression of the genes encoding the proteins described above. C2) expresses the gene encoding the nucleic acid molecule described in C1). C3) contains the expression cassette encoding the gene described in C2). C4) A recombinant vector containing the encoding gene described in C2), or a recombinant vector containing the expression cassette described in C3). C5) Recombinant microorganisms containing the encoding gene described in C2), or recombinant microorganisms containing the expression cassette described in C3), or recombinant microorganisms containing the recombinant vector described in C4). C6) A transgenic plant cell line containing the encoding gene described in C2), or a transgenic plant cell line containing the expression cassette described in C3), or a transgenic plant cell line containing the recombinant vector described in C4). C7) Transgenic plant tissue containing the encoding gene described in C2), or transgenic plant tissue containing the expression cassette described in C3), or transgenic plant tissue containing the recombinant vector described in C4). C8) A transgenic plant organ containing the encoding gene described in C2), or a transgenic plant organ containing the expression cassette described in C3), or a transgenic plant organ containing the recombinant vector described in C4).

[0018] In the above applications, the nucleic acid molecule described in C1) can be a DNA molecule expressing gRNA targeting the protein-coding gene described above, or the gRNA of the protein-coding gene described above.

[0019] The gRNA may correspond to a DNA molecule with nucleotide sequence SEQ ID NO: 3, positions 3670-3692, specifically 5′-GATAATTGACCCACCAGATTTGG-3′.

[0020] In the above applications, the gene encoding C2) is introduced into the recipient plant, specifically by transforming plant cells or tissues using conventional biological methods such as Ti plasmid, Ri plasmid, plant virus vector, direct DNA transformation, microinjection, electrocoagulation, or Agrobacterium-mediated transformation, and then culturing the transformed plant tissues into plants. The transformed cells, tissues, or plants are understood to include not only the final products of the transformation process but also the materials obtained through asexual reproduction and transgenic progeny.

[0021] Those skilled in the art can readily mutate the nucleotide sequence encoding the GmFVE protein of this invention using known methods, such as directed evolution and point mutation. Any artificially modified nucleotides that possess 75% or higher identity to the nucleotide sequence of the GmFVE protein isolated according to this invention, provided they encode and function the GmFVE protein, are derived from and equivalent to the nucleotide sequence of this invention.

[0022] In one embodiment of the present invention, the recombinant vector may be a plant gene editing vector, and the plant gene editing vector may be a pCBSG015-sgRNA vector.

[0023] As a specific embodiment, the recombinant vector is the recombinant vector pCBSG015-sgRNA. The recombinant vector is obtained by replacing the fragment between 5'-ggcaccgagtcggtgc-3' in the pCBSG015 vector with a DNA fragment having the nucleotide sequence 5'-GATAATTGACCCACCAGATTTGG-3', while keeping other sequences of the pCBSG015 vector unchanged. This recombinant plasmid is named the recombinant vector pCBSG015-sgRNA.

[0024] Those skilled in the art can readily mutate the nucleotide sequence encoding the protein GmFVE of this invention using known methods, such as directed evolution or point mutation. Any artificially modified nucleotides that possess 75% or more of the nucleotide sequence identity with the protein GmFVE isolated in this invention, provided they encode and function as protein GmFVE, are derived from and equivalent to the nucleotide sequence of this invention.

[0025] The aforementioned 80% or higher degree of identity can be 80%, 85%, 90%, or 95% or higher degree of identity.

[0026] In this article, identity refers to the similarity of amino acid or nucleotide sequences. The identity of amino acid sequences can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the procedure, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, and setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, a search can be performed to calculate the identity of amino acid sequences, and then the identity value (%) can be obtained.

[0027] In this document, the 80% or more of identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the identity.

[0028] The vectors described herein are known to those skilled in the art and include, but are not limited to: plasmids, bacteriophages (such as λ phage or M13 filamentous phage), granules (i.e., Cos plasmids), Ti plasmids, or viral vectors.

[0029] Existing plant expression vectors can be used to construct structures containing the aforementioned... GmFVE Recombinant gene expression vectors. These plant expression vectors include, but are not limited to, binary Agrobacterium vectors and vectors suitable for plant microbombardment. The plant expression vectors may also contain the 3' untranslated region of the exogenous gene, i.e., containing a polyadenylate signal and any other DNA fragment involved in mRNA processing or gene expression. The polyadenylate signal can guide the addition of polyadenylate to the 3' end of the mRNA precursor, such as, but not limited to, Agrobacterium crown gall tumor inducing (Ti) plasmid genes (e.g., carmine synthase). NosThe untranslated regions transcribed at the 3' end of genes (such as soybean storage protein genes) have similar functions.

[0030] use GmFVE When constructing recombinant plant expression vectors, any type of enhancing promoter or constitutive promoter can be added before the transcription initiation nucleotide, including but not limited to the cauliflower mosaic virus (CAMV) 35S promoter and the maize ubiquitin promoter. These can be used alone or in combination with other plant promoters. Furthermore, when constructing plant expression vectors using the genes of this invention, enhancers, including translational enhancers or transcriptional enhancers, can also be used. These enhancer regions can be ATG start codons or adjacent region start codons, but they must be identical to the reading frame of the coding sequence to ensure correct translation of the entire sequence. The sources of the translation control signals and start codons are wide-ranging; they can be natural or synthetic. The translation initiation region can originate from the transcription initiation region or structural genes.

[0031] To facilitate the identification and screening of transgenic plant cells or plants, the plant expression vectors used can be processed, such as by adding genes that encode enzymes or luminescent compounds that can be expressed in plants, including but not limited to those encoding enzymes that produce color changes. GUS These can include genes such as luciferase genes, antibiotic resistance markers (gentamicin markers, kanamycin markers, etc.), or chemical reagent resistance marker genes (such as herbicide resistance genes). From a safety perspective, transgenic plants can be transformed directly through stress screening without adding any selective marker genes.

[0032] In the above applications, the microorganism described in C5) can be yeast, bacteria, algae, or fungi. Specifically, it can be Agrobacterium tumefaciens EHA105. The recombinant Agrobacterium is EHA105 / pCBSG015-sgRNA.

[0033] In the above applications, the transgenic plant cell lines, transgenic plant tissues, and transgenic plant organs described in C6) may or may not include propagation material.

[0034] In the above applications, the plant tissue described in C7) may be derived from roots, stems, leaves, flowers, fruits, seeds, pollen, embryos, and anthers.

[0035] In the above applications, the transgenic plant organs described in C8) can be the roots, stems, leaves, flowers, fruits, and seeds of the transgenic plant.

[0036] Furthermore, in the aforementioned applications, the transgenic plant cell lines, transgenic plant tissues, and transgenic plant organs may or may not include propagation material.

[0037] Secondly, the present invention provides a method for regulating the oil content of plant seeds, the method comprising step M, wherein step M comprises regulating the activity and / or content of the aforementioned protein in the target plant, and / or regulating the expression level of the encoding gene of the aforementioned protein, thereby regulating the oil content of plant seeds; the target plant may be a seed plant.

[0038] Thirdly, the present invention provides a method for cultivating plants with altered seed oil content, the method comprising step M, wherein step M comprises regulating the activity and / or content of the proteins described above in the target plant, and / or regulating the expression level of the encoding genes of the proteins described above, to obtain plants with altered seed oil content; the target plant may be a seed plant.

[0039] In the above method, step M may include inhibiting or reducing or silencing the activity and / or content of the protein described above in the target plant, or / and inhibiting or reducing or silencing the expression level of the gene encoding the protein described above, in order to increase the seed oil content of the target plant; the target plant contains the gene encoding the protein described above; the target plant may be a seed plant.

[0040] In one embodiment of the present invention, the method for cultivating plants with altered seed oil content includes the following steps: 1) Construct a recombinant expression vector containing DNA molecules shown in SEQ ID NO:1 and SEQ ID NO:3 that inhibit, reduce, or silence them; 2) Transform the recombinant expression vector constructed in step 1) into the recipient plant; 3) Transgenic plants with altered seed oil content were obtained through screening and identification.

[0041] The importation refers to the use of recombination methods, including but not limited to Agrobacterium (…). Agrobacterium Introduced methods include mediated transformation, bio-projectile methods, electroporation, in-planta techniques, and more.

[0042] Using any vector capable of guiding the expression of exogenous genes in plants, the encoding gene for knocking out the protein GmFVE provided in this invention can be expressed. GmFVE By introducing gene fragments into plant cells or recipient plants, transgenic cell lines and transgenic plants with altered seed oil content can be obtained. Expression vectors carrying the gene encoding the knockout protein GmFVE can be used to transform plant cells or tissues using conventional biological methods such as Ti plasmids, Ri plasmids, plant virus vectors, direct DNA transformation, microinjection, electrocoagulation, and Agrobacterium-mediated transformation, and the transformed plant tissues can be cultured into plants.

[0043] The microorganisms mentioned in this article may be yeast, bacteria, algae, or fungi. Among them, bacteria may originate from the genus *Escherichia* (…). Escherichia Erwinia ( Erwinia Agrobacterium tumefaciens ( ), Agrobacterium tumefaciens Agrobacterium Flavobacterium ( Flavobacterium) Alkalophytum genus ( Alcaligenes ), Pseudomonas ( Pseudomonas ), Bacillus spp. ( Bacillus (e.g., Agrobacterium tumefaciens EHA105).

[0044] In this invention, the oil content may be palmitic acid content and / or stearic acid content and / or oleic acid content and / or linoleic acid content and / or linolenic acid content.

[0045] In this invention, the inhibition, reduction, or downregulation of the expression of the GmFVE protein-encoding gene can be achieved through gene knockout or gene silencing.

[0046] Gene knockout refers to the phenomenon of inactivating a specific target gene through homologous recombination. Gene knockout inactivates a specific target gene by altering its DNA sequence.

[0047] Gene silencing refers to the phenomenon of preventing or reducing gene expression without damaging the original DNA. Gene silencing presupposes no change in the DNA sequence, resulting in the absence or reduction of gene expression. Gene silencing can occur at two levels: transcriptional silencing due to DNA methylation, heterochromatinization, and position effects; and post-transcriptional gene silencing, which inactivates the gene at the post-transcriptional level through specific inhibition of target RNA. This includes antisense RNA, co-suppression, gene quelling, RNA interference (RNAi), and microRNA (miRNA)-mediated translational repression.

[0048] The gene encoding the protein may be the genomic gene of the protein or the cDNA gene of the protein.

[0049] The cDNA gene is a cDNA molecule that includes the coding sequence (CDS) of the protein. The coding sequence may be SEQ ID NO: 1.

[0050] Specifically, step M may include introducing a gene knockout vector, such as pCBSG015-sgRNA, targeting bases 3670-3692 of SEQ ID NO:3 into the target plant.

[0051] In some embodiments of the present invention, the target plant contains SEQ ID NO: 3, and step M may include performing any of the following operations on the genome of the target plant: D1) Delete the 33 bp nucleotide TGGGTCAATTATCAAACAAAACTCTAAATCTGG at positions 3680-3712 in the genome of the target plant, resulting in GmFVE Partial segment loss and frameshift mutation, thus GmFVE Gene knockout; D2) Deletion of the 5 bp nucleotide CTGGT at positions 3676-3680 in SEQ ID NO: 3 of the target plant genome leads to premature termination of GmFVE translation, thereby... GmFVE Gene knockout; D3) Delete the 3 bp CTG at positions 3676-3678 in SEQ ID NO: 3 of the target plant genome, resulting in GmFVE Small fragments are missing, thus... GmFVE Gene knockout.

[0052] In this invention, the plant mentioned above can be any of the following: N1) dicotyledonous plants; N2) legumes; N3) legumes; N4) soybeans; N5) soybean.

[0053] In some specific embodiments of the present invention, the soybean may be the soybean variety Jack.

[0054] Fourthly, the present invention also provides a substance, which may be the aforementioned protein or biological material.

[0055] This invention modifies soybeans GmFVE Genes, via CRISPR / Cas9 vector, are used to target endogenous soybean genes. GmFVE Knockout of the gene yields single or double mutants, resulting in gene knockout soybean plants. fve1 , fve2 and fve3 The seed oil content of the mutants was significantly higher than that of the recipient controls Jack and Null. Statistical analysis indicates that GmFVE negatively regulates soybean seed oil content, reducing the level of its encoding gene. GmFVE The expression level of this substance can significantly increase the oil content of soybean seeds. Attached Figure Description

[0056] Figure 1 for GmFVE Gene-edited mutants fve1 The preparation of [the substance] is shown. GmFVE The target sequence (sgRNA1) and its role GmFVE Specific location on the gene andfve1 Mutant gene editing was validated by sequencing.

[0057] Figure 2 for GmFVE Gene-edited mutants fve2 The preparation of [the substance] is shown. GmFVE The target sequence (sgRNA1) and its role GmFVE Specific location on the gene and fve2 Mutant gene editing was validated by sequencing.

[0058] Figure 3 for GmFVE Gene-edited mutants fve3 Preparation. Shown. GmFVE The target sequence (sgRNA1) and its role GmFVE Specific location on the gene and fve3 Mutant gene editing was validated by sequencing.

[0059] Figure 4 For comparison, Jack, Null and GmFVE Gene mutants fve1 , fve2 and fve3 The relative expression level was detected.

[0060] Figure 5 For comparison, Jack, Null and GmFVE mutant fve1 , fve2 and fve3 The content of fatty acids and oils in the seeds was detected. A represents the total oil content, and B represents the content of each component of the fatty acids. Detailed Implementation

[0061] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0062] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0063] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.

[0064] The soybean variety Jack used in the following examples is described in: Jin-Song Zhang1, et al., Atranscriptional regulatory module controls lipid accumulation in soybean, NewPhytologist (2021), 231:661-678. This biological material is available to the public from the applicant and is intended solely for the purpose of replicating experiments of this invention; it may not be used for any other purpose.

[0065] The Agrobacterium tumefaciens EHA105 in the following examples has been described by: Rodrigues SD, Karimi M, Impens L, Van Lerberge E, Coussens G, Aesaert S, Rombaut D, Holtappels D, Ibrahim HMM, Van Montagu M, Wagemans J, Jacobs TB, De Coninck B, Pauwels L. Efficient CRISPR-mediated base editing in Agrobacterium spp. Proc Natl Acad Sci US A.2021 Jan, 12;118(2): e2013338118. doi: 10.1073 / pnas.2013338118. The biological material is available to the public from the applicant and is intended solely for the purpose of repeating experiments of the present invention and may not be used for any other purpose.

[0066] The following examples used statistical software to process the data. The experimental results are expressed as mean ± standard deviation. One-way ANOVA was used, and P < 0.05 was considered satisfactory. () indicates a significant difference, P < 0.01. () indicates a highly significant difference.

[0067] Example 1, Soybeans GmFVE Acquisition of genes Within the established soybean yield regulation network, it was found that... Glyma.09G063100 The FVE / MSI4 protein encoded by the gene possesses a WD40 repeat domain and is a component of the histone deacetylase complex, typically involved in transcriptional repression. It is speculated that this gene may be involved in soybean seed oil accumulation and has been named... GmFVE .

[0068] According to the Williams 82 (W82) reference genome sequence (Genome assembly Glycine_max_v4.0, Genbank: GCA_000004515.5, updated March 10, 2021) GmFVE Based on the full-length cDNA sequence information, primers were designed, and the primer sequences are as follows: GmFVE -up:5′- ATGGAGACTCCTCCTCCCCAACAAG-3′; GmFVE -dp:5′-TCATTTTTCAGTCTTTGAAGCACAC-3′.

[0069] RNA was extracted from soybean variety W82 seedlings and reverse transcribed into cDNA using Toyobo's ReverTra Ace reverse transcriptase. The cDNA was then used as a template... GmFVE -up and GmFVE -dp was used as a primer for PCR amplification. PCR was applied to amplify total RNA from soybeans. GmFVE Gene extraction: Leaf samples were crushed in liquid nitrogen, suspended in 4 mol / L guanidine thiocyanate, and extracted with acidic phenol and chloroform. Anhydrous ethanol was added to the supernatant to precipitate the RNA, which was then dissolved in water to obtain total RNA. 1 µg of total RNA was reverse transcribed using a Thermo Fisher Scientific kit according to the kit's instructions. The resulting cDNA fragment was then used as a template for PCR amplification.

[0070] The 50 μl PCR reaction mixture consisted of: 1 μl single-stranded cDNA (0.05 μg), 1.5 μl of the above primers (10 μM), 25 μl 2× PCR buffer, 10 μl dNTPs (10 mM), and 1 U KOD DNA polymerase, with the volume made up to 50 μl with ultrapure water. The reaction was performed on a PE9600 PCR instrument with a 94-hour program. o C denaturation for 5 minutes; then 98 o C 1min, 58 o C 1min, 68 o C 1min, 30-32 cycles in total; then 68 o C extends for 10 minutes; 4 oStored at C. A PCR product of approximately 1.5 kb was obtained. The amplified product was detected by 1% agarose gel electrophoresis. The approximately 1.5 kb DNA fragment was recovered using an agarose gel recovery kit (TIANGEN) and cloned into the pMD-18 vector (TaKaRa, catalog number 6011). This vector was then transformed into competent *E. coli* cells, positive clones were screened, and plasmids were extracted for sequencing. Sequencing results showed that the PCR product amplified by this primer pair was... GmFVE The CDS sequence (SEQ ID NO:1, 1542bp) has an ORF consisting of nucleotides 1-1539 from the 5' end of sequence 1. SEQ ID NO:1 is as follows:

[0071] The amino acid sequence of this protein is sequence 2 in the sequence listing (SEQ ID NO: 2,513aa), as follows: METPPPQQGVVKKKETRGRKPKPKDEHGKGLKEGRKTQQQQQQQQQHHHQQQQQQDQPSVDEKYTQWKSLVPVLYDWLANHNLVWPSLSCRWGPQLEQATYKNRQRLYLSEQTDGSVPNTLVIANCE VVKPRVAAAEHISQFNEEARSPFVKKYKTIIHPGEVNRIRELPQNSKIVATHTDSPDVLVWDVESQPNRHAVLGATNSRPDLILTGHQDNAEFALAMCPTEPYVLSGGKDKTVVLWSIEDHITSAATDS KSGGSIIKQNSKSGEGNDKTADGPTVGPRGIYCGHEDTVEDVAFCPSSAQEFCSVGDDSCLILWDARVGSSPVVKVEKAHNADLHCVDWNPHDDNLILTGSADNSVRMFDRRNLTTNGVGSPIHKFEG HKAAVLCVQWSPDKSSVFGSSAEDGLLNIWDYEKVGKKIERSGKSISSPPGLFFQHAGHRDKVVDFHWNAYDPWTIVSVSDDCESTGGGGTLQIWRMSDLIYRPEDEVLAELEKFKSHVVACASKTEK.

[0072] GmFVE The genome sequence is SEQ ID NO:3 (7586bp), where positions 1-54 are 5' UTRs; positions 7283-7586 are 3' UTRs; it also includes 15 exon sequences, located at positions 55-329, 417-480, 637-723, 992-1054, 1292-1357, 1463-1537, 3445-3520, 3607-3814, 4563-4644, 4955-5028, 5156-5267, 5355-5429, 5525-5601, 6381-6840, and 7175-7282. The remaining sequences are introns, as detailed below:

[0073] Example 2: GmFVE Gene mutation leads to increased oil content in soybean seeds 1. Knock out soybean Jack GmFVE Family genes GmFVE 1.1 Construction of pCBSG015 gene editing vector The gene editing was performed by Wimi Biotechnology Co., Ltd. (formerly known as Baige Biotechnology Co., Ltd.), and the pCBSG015 vector was provided by Wimi Biotechnology Co., Ltd., with the catalog number wimi-pCXB053.

[0074] Target sequence selection: The high-throughput CRISPR-Cas9 target design program developed by Wemi Technology was used. The target design principles of this program are as follows: 1) Knockout sites should be located in the coding sequence (CDS) region and preferably at the protein's front end or in an important functional domain; 2) Coverage should be maximized to include a higher proportion of transcripts; 3) No off-target effects or off-target effects should be located in intergenic regions; 4) Targets with higher editing efficiency should be preferred; 5) The sequence should have a relatively balanced GC content and be less prone to secondary structure formation. The successful application of this program in whole-genome target design in rice, soybean, and other crops has proven its feasibility.

[0075] A single-gene, single-target knockout approach was adopted. The selected target site T1: 5'-GATAATTGACCCACCAGATTTGG-3' is located in the seventh exon region.

[0076] T1: 5'-GATAATTGACCCACCAGATTTGG-3', the target sequence is located at GmFVE -The 770th-792nd bits of the CDS sequence SEQ ID NO:1 and GmFVE Genome sequence SEQ ID NO:3, positions 3670-3692; Promoter selection: AtU6 from Arabidopsis thaliana was used to promote the T1 target site.

[0077] Preparation of sgRNA expression cassettes containing target sites: The pCBSG015 vector was linearized by Bsa I digestion. Using primer synthesis, the T1 sequence was directly synthesized, and 16 bp vector sequences were added to both ends as homologous arms, yielding Sense-U6-T1: 5'-ggcaccgagtcggtgcGATAATTGACCCACCAGATTTGGgttgaacaacggaaac-3'. Then, the inverse complementary sequence Anti-U6-T1 was synthesized: 5'-gtttccgttgttcaacCCAAATCTGGTGGGTCAATTATCgcaccgactcggtgcc-3', where lowercase letters represent homologous arms and uppercase letters represent the T1 sequence. Annealing was performed to form a double strand, followed by homologous recombination with the backbone linear vector. The specific procedures are as follows: Preparation of the annealed AtU6-T1-gRNA fragment: The synthesized Sense-U6-T1 and Anti-U6-T1 sequences were annealed to form double strands. The reaction system was as follows: the synthesized sequence was dissolved in 75 mM NaCl solution to a final concentration of 0.2 nM / μl, and equal volumes of the forward and reverse strand solutions (Sense-U6-T1 and Anti-U6-T1) were mixed. The mixture was heated in a 95℃ water bath for 5-10 min, followed by slow cooling to obtain the annealed AtU6-T1-gRNA fragment.

[0078] Vector linearization by enzyme digestion: 1-2 μg of pCBSG015 plasmid, 10X CutSmart TM 5 μl of buffer (NEB), 1 μl of BsaI restriction enzyme, and sterile double-distilled water were added to a final volume of 50 μl. The mixture was incubated at 37°C for 30 min. The DNA was then purified using the EZ-10 Column DNA Purification Kit (Shanghai Sangon Biotech). The purified DNA was dissolved in an appropriate amount of water to obtain the linearized pCBSG015 vector.

[0079] Homologous recombination of the target sgRNA expression cassette with the pCBSG015 vector was performed using the EasyGeno Rapid Recombinant Cloning Kit (TianGen). The reaction mixture consisted of 5 μL of 2×EasyGeno Assembly mix buffer, 0.5 μL of linearized pCBSG015 vector, and 4.5 μL of annealed AtU6-T1-gRNA. The reaction conditions were 50 °C for 15 min. The ligation product was transformed into E. coli DH5α competent cells, and plasmids were extracted from positive colonies. After successful sequencing, the recombinant vector pCBSG015-sgRNA was obtained.

[0080] The recombinant vector pCBSG015-sgRNA is described as follows: it is a recombinant plasmid obtained by replacing the fragment between 5'-ggcaccgagtcggtgc-3' in the pCBSG015 vector with a DNA fragment with the nucleotide sequence T1: 5'-GATAATTGACCCACCAGATTTGG-3', while keeping the other sequences of the pCBSG015 vector unchanged.

[0081] The recombinant vector pCBSG015-sgRNA contains the gene encoding the editing target T1 and the Cas9 protein on the vector. After being introduced into the receptor, the transcribed guide RNA can target the target sequence near the PAM in the receptor genome through base complementarity, i.e., targeting... GmFVE Genes, Cas9 protein GmFVE A double-strand break in DNA at a gene target site triggers a gene mutation in the cut region through the organism's own DNA damage repair response mechanism. This mutation leads to large deletions, frameshift mutations, or premature termination of translation in the coding gene, thereby achieving the desired effect. GmFVE Gene knockout.

[0082] The recombinant vector pCBSG015-sgRNA was transformed into Agrobacterium EHA105 competent cells to obtain Agrobacterium EHA105-pCBSG015-sgRNA.

[0083] 1.2 Genetic transformation of soybean The soybean variety Jack was infected with the Agrobacterium EHA105-pCBSG015-sgRNA prepared above. GmFVE Gene-edited soybean plants.

[0084] Harvest T0 generation seeds, self-pollinate T1 generation seeds to obtain T1 generation seeds, plant T1 generation seeds to obtain T2 generation seedlings and perform the following tests.

[0085] Null plants were isolated from heterozygous gene-edited lines and do not contain the pCBSG015 vector. GmFVE Unedited plants, along with the recipient material Jack, served as control materials.

[0086] 1.3 GmFVE Screening for homozygous gene mutations Using the genomic DNA of the T2 generation seedlings obtained in section 1.2 as a template, screening and identification were performed using mutant target detection primers. Primers were designed using approximately 93 bp upstream and 308 bp downstream of the target sequence T1. GmFVE Primers for gene target detection 。GmFVE The primers for gene target detection are: GmFVE -F: 5'- GCCCAACTGAACCCTAT -3' and GmFVE-R: 5'-CAACCAAGCACCAAACA -3', the amplified product is approximately 510bp.

[0087] After successful sequencing, the CRISPR target editing method was analyzed using the website DSDecode (http: / / dsdecode.scgene.com / ) and compared with the standard gene sequence using manual peak reading. The editing methods of each target sequence and its upstream and downstream sequences were analyzed. The T2 generation gene-edited positive plants were then self-crossed and passaged to prepare gene-edited homozygous mutant plants.

[0088] Filtered by the above method GmFVE homozygous mutant fve1 , fve2 and fve3 The gene editing methods are shown in Figure 1 , 2 And 3.

[0089] Figure 1 , 2 3 and 3 respectively show GmFVE Gene-edited mutants fve1 , fve2 and fve3 The construction process of sgRNA1(T1) is shown in the figure. GmFVE The position of the sequence and the sequence spliced ​​by the mutant indicate that in the three mutants, GmFVE All achieved the desired editing effect.

[0090] Compared to the recipient soybean variety Jack, GmFVE homozygous mutant fve1 In the genome of the strain GmFVE The gene underwent the following mutation: A 33-nucleotide deletion was found in positions 3680-3712 of the genome sequence (SEQ ID NO:3), corresponding to positions 780-812 of SEQ ID NO:1, specifically a 5'-TGGGTCAATTATCAAACAAAACTCTAAATCTGG - 3' deletion, resulting in... GmFVE Partial segment loss and frameshift mutation, thus GmFVE Gene knockout.

[0091] Compared to the recipient soybean variety Jack, GmFVE homozygous mutant fve2 In the genome of the strain GmFVEThe gene underwent the following mutation: A 5-nucleotide 5'-CTGGT-3' deletion was found at positions 3676-3680 of the genome sequence (SEQ ID NO:3), corresponding to positions 776-780 of SEQ ID NO:1, resulting in… GmFVE The translation was terminated early, thus... GmFVE Gene knockout.

[0092] Compared to the recipient soybean variety Jack, GmFVE homozygous mutant fve3 In the genome of the strain [[ID=1 The gene underwent the following mutation: In both homologous chromosomes, at positions 3676-3678 of the genome sequence (SEQ ID NO:3), corresponding to positions 776-778 of SEQ ID NO:1, there was a deletion of 3 nucleotides, 5'-CTG-3', resulting in… ​ Small fragments are missing, thus... ​ Gene knockout.

[0093] mutant ​ , ​ and ​ Seeds from T2 generation plants (T3 generation) were harvested for subsequent experiments.

[0094] 2. ​ Gene CRISPR materials ​ Phenotypic analysis 2.1 ​ middle ​ Detection of gene expression levels Extracted from soybean varieties Jack, Null and ​ , ​ and ​ Total RNA from mid-seed development was reverse transcribed, and the resulting cDNA was used as a template for Real-Time PCR identification. ​ Gene expression levels. Primers used: FVE-qPCR-F: 5′-GGAAGGGTTTGAAGGAAGGTAGG -3′; FVE-qPCR-R: 5′-TCGTAGAGGACAGGAACAAGGGA -3′.

[0095] The soybean Tublin gene was used as an internal standard, and the internal standard primers were: Primer-TF: 5′-TGGCCGTTACCTGACAGCAT-3′; Primer-TR: 5′-CTCGGAGGGATGTCACACAC-3′.

[0096] The results are as follows ​ As shown: Jack and Null ​ The relative expression levels were 4.71±0.25 and 4.83±0.19, respectively. ​ and ​ middle ​ The expression levels were approximately 3.78±0.04, 3.88±0.07, and 2.59±0.10, respectively, indicating that the mutants contained... ​ The expression level decreased significantly or extremely significantly.

[0097] 2.2 ​ Gene downregulation increases soybean seed oil content The one obtained in 1.3 of Example 2 ​ and ​ Seeds of the gene-edited T3 generation plants, along with control seeds (Jack and Null), were sown in greenhouse pots. Greenhouse conditions included 16 hours of light at 11000 Lux and 8 hours of darkness, with daytime temperatures ranging from 30 to 37°C and nighttime temperatures from 25 to 28°C. Growth and development were observed. Seeds were harvested after 130 days. ​ and ​ No significant differences were observed in the phenotypes of Jack and Null during the growth and maturity stages.

[0098] After the above-mentioned potted soybean seeds matured, the seeds harvested from individual plants were dried at 37℃ for one week, and the recipient control Jack, Null, and mutant were measured. ​ and ​ Seed oil content. Fifteen plants were taken from each line, and the biological experiment was repeated three times. Results are expressed as mean ± standard deviation. One-way ANOVA was used, and P < 0.05 was considered acceptable. () indicates a significant difference, P < 0.01. () indicates a highly significant difference.

[0099] 1) The specific steps for determining the total oil content in seeds are as follows: Grind the dried seeds into powder, weigh 100 mg into four parallel centrifuge tubes. Add 500 μl of n-hexane, mix thoroughly, and incubate overnight at 37°C. Centrifuge slowly for 3 minutes, transferring the n-hexane into a newly weighed centrifuge tube. Repeat the soaking process with n-hexane on the remaining powder, then centrifuge and collect the n-hexane into the same centrifuge tube. Place the centrifuge tube in a vacuum pump and evacuate to allow the n-hexane to evaporate completely. Weigh the centrifuge tube again. The change in weight before and after centrifugation represents the weight of extracted lipids, and the total oil content (%) is calculated. The formula for calculating the total oil content (%) is as follows: Total oil content (%) = (Weight of extracted lipids / Total weight of seeds) × 100%.

[0100] The results are as follows ​ As shown in A, compare Jack, Null, and ​ and ​ The total oil content of the seeds was statistically analyzed to be 21.7±0.2%, 21.7±0.3%, 23.8±0.3%, 23.3±0.4%, and 24.8±0.3%, respectively. Mutant ​ and ​ The total oil content in the seeds was significantly higher than that of the controls Jack and Null. The total oil content in the mutant seeds increased by 9.0%, 7.5%, and 14.1% compared to the controls, respectively.

[0101] 2) Determination of fatty acid content in seeds The determination of various fatty acids was performed according to the method described by Pritam et al. (Pritam S. Sukhija and DLPalmquist (1988) Rapid Method for Determination of Total Fatty Acid Content and Composition of Feedstuffs and Feces J. Agric. Food Chem., 36: 1202-1206): After the seeds were thoroughly ground, 10 mg was weighed and added to 1 ml of sulfuric acid methanol solution (2.5 ml sulfuric acid / 100 ml methanol) and 10 µl of internal standard solution (100 µg 17 alkyl acetic acid / ml ethyl acetate). The mixture was incubated at 100 °C for 1 h. The supernatant was collected, and 500 µl of n-hexane and 300 µl of 10% (w / w) NaCl aqueous solution were added to each 1 ml of supernatant. The mixture was then dried under vacuum, and 30 µl of ethyl acetate was added. The supernatant was then analyzed by gas chromatography (AGILENT 6890 GC), with various fatty acid standards from Sigma as standard controls.

[0102] The gas chromatography parameters are as follows: column inner diameter 0.25 mm, length 30 m; carrier is ethyl acetate with a particle size of 0.25 µm; carrier gas is helium; column temperature is 120 °C, injection port temperature is 120 °C, injection volume is 1 µl, and carrier gas flow rate is 2 ml / min.

[0103] The results are as follows ​ As shown in B, Jack, Null, and ​ and ​ The palmitic acid (16:0) content in the seeds was approximately 2.1, 2.1, 2.5, 2.3, and 2.4; the stearic acid (18:0) content was approximately 0.7, 0.7, 1.2, 1.1, and 1.1; the oleic acid (18:1) content was approximately 4.8, 4.8, 5.8, 5.7, and 5.6; the linoleic acid (18:2) content was approximately 11.6, 11.6, 13.5, 13.4, and 14.3; and the linolenic acid (18:3) content was approximately 1.6, 1.6, 1.8, 1.8, and 1.9. The mutant... ​ and ​ The content of the above five fatty acids in the seeds was significantly or extremely significantly higher than that in the control.

[0104] The above statistics show that three biological replicate experiments on 15 individual plants in a greenhouse potted plant indicated that the mutant... ​ and ​ The total oil content in soybean seeds was significantly higher than that in the receptor Jack and Null controls, indicating that GmFVE negatively regulates soybean seed oil content and reduces the GmFVE encoding gene. ​ The expression of this substance significantly increases the oil content of soybean seeds.

[0105] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. The use of a protein, or a substance regulating the expression of the gene encoding said protein, or a substance regulating the activity or content of said protein, in any of the following: A1) Application in regulating the oil content of plant seeds; A2) Application in the preparation of products that regulate the oil content of plant seeds; A3) Application in cultivating plants with altered seed oil content; A4) Application in the preparation of products from plants whose seed oil content has been altered; A5) Applications in plant breeding; The protein is GmFVE, and is any of the following proteins: B1) The amino acid sequence of the protein is as shown in SEQ ID NO:

2. B2) A protein having the same function as the amino acid sequence shown in SEQ ID NO: 2, but with one or more amino acid residues substituted and / or deleted and / or added. Proteins that share more than 80% amino acid sequence identity with B1) or B2) and have the same function, B4) A fusion protein obtained by attaching a tag to the end of any of the proteins defined in B1)-B3).

2. Use according to claim 1, characterized in that: The protein is derived from soybeans.

3. Use according to claim 1 or 2, characterized in that: The substance that regulates the expression of the gene encoding the protein or the substance that regulates the activity or content of the protein is a biological material, and the biological material is any one of the following: C1) Nucleic acid molecules that inhibit, reduce, or downregulate the expression of the gene encoding the protein described in claim 1 or 2. C2) expresses the gene encoding the nucleic acid molecule described in C1). C3) contains the expression cassette of the gene described in C2). C4) A recombinant vector containing the gene described in C2), or a recombinant vector containing the expression cassette described in C3). C5) Recombinant microorganisms containing the gene described in C2), or recombinant microorganisms containing the expression cassette described in C3), or recombinant microorganisms containing the recombinant vector described in C4). C6) A transgenic plant cell line containing the gene described in C2), or a transgenic plant cell line containing the expression cassette described in C3), or a transgenic plant cell line containing the recombinant vector described in C4). C7) Transgenic plant tissue containing the gene described in C2), or transgenic plant tissue containing the expression cassette described in C3), or transgenic plant tissue containing the recombinant vector described in C4). C8) A transgenic plant organ containing the gene described in C2), or a transgenic plant organ containing the expression cassette described in C3), or a transgenic plant organ containing the recombinant vector described in C4).

4. Use according to claim 3, characterized in that: C1) The nucleic acid molecule is a gRNA that targets the protein-coding gene described in claim 1.

5. A method of modulating oil content in a seed of a plant, comprising: The method includes step M, which includes regulating the activity and / or content of the protein described in claim 1 or 2 in the target plant, and / or regulating the expression level of the gene encoding the protein described in claim 1 or 2, to regulate the oil content of plant seeds; the target plant is a seed plant.

6. A breeding method for cultivating plants with altered seed oil content, the method comprising step M, wherein step M comprises regulating the activity and / or content of the protein described in claim 1 or 2 in the target plant, and / or regulating the expression level of the gene encoding the protein described in claim 1 or 2, to obtain a plant with altered seed oil content; wherein the target plant is a seed plant.

7. The method according to claim 5 or 6, characterized in that: Step M includes introducing a gene knockout vector targeting bases 3670-3692 of SEQ ID NO:3 into the target plant.

8. The method of claim 5 or 6, wherein: Step M includes performing any of the following operations on the genome of the target plant: D1) Delete the 33bp nucleotide TGGGTCAATTATCAAACAAAACTCTAAATCTGG at positions 3680-3712 in the genome of the target plant; D2) Deletion of 5 bp nucleotide CTGGT at positions 3676-3680 in SEQ ID NO: 3 of the target plant genome; D3) Delete the 3bp CTG at positions 3676-3678 in SEQ ID NO:3 of the target plant genome.

9. Use according to any one of claims 1 to 4, or method according to any one of claims 5 to 8, characterized in that: The plant is any one of the following: N1) Dicotyledons N2) Leguminosae; N3) Leguminosae (family legumes); N4) Plants of the genus *Glycine*; N5) soybeans.

10. A substance, wherein the substance is the protein of claim 1 or 2 or the biological material of claim 3.