Application of protein TaeIF4A in increasing grain number and yield of main spike of wheat
By expressing or mutating the TaeIF4A-6A protein or its encoding gene in wheat, the number of grains per spike, plant height, and tiller number were regulated, solving the problem of unclear genetic basis for wheat yield improvement and achieving significant increases in yield and number of grains per spike.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, the genetic basis and molecular mechanisms of wheat ear grain number and yield are unclear, making it difficult to effectively increase yield per plant.
By expressing or mutating the TaeIF4A-6A protein or its encoding gene in wheat, the number of grains per spike, plant height, and tiller number in wheat can be regulated. Recombinant vectors or gene editing technologies can be used to increase the expression level and activity of the TaeIF4A protein.
Significantly increase the number of grains in the main spike, plant height, and number of tillers in wheat, enhance yield per plant, and achieve the goal of high-yield molecular breeding of wheat.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to the application of protein TaeIF4A in increasing the number of grains in the main ear of wheat and its yield. Background Technology
[0002] Given limited arable land, increasing wheat yield primarily depends on improving yield per plant. Wheat yield per plant is mainly influenced by the number of spikes per unit area, the number of grains per spike, and grain weight. Important traits such as spike length, plant height, spikelet density, number of grains per spike, and thousand-grain weight are closely related to yield and are key indicators that breeders focus on during variety selection. Spike length, plant height, and number of grains per spike are typical quantitative traits with high heritability, but their genetic basis and molecular mechanisms are not yet fully understood. Therefore, identifying and cloning functional genes controlling these traits and elucidating their genetic basis and molecular mechanisms are of great significance for achieving high-yield molecular breeding of wheat. Wheat yield per plant consists of the number of effective spikes, the number of grains per spike, and grain weight. Only the number of grains per spike shows significant variation among different varieties; therefore, genetic improvement of the number of grains per spike plays a crucial role in increasing wheat yield per plant. Numerous studies have shown that increasing the number of grains per spike is a major pathway to further increase wheat yield.
[0003] The DEAD box family of RNA helicases are crucial participants in RNA metabolism in most organisms. Despite their high conservation among DEAD box proteins, they participate in a variety of biological processes, including RNA substrate binding, double-stranded RNA unwinding, ATP hydrolysis, and phosphate release, all of which require tight regulation by other proteins and small molecules. Therefore, DEAD box RNA helicases often function as large complexes. In the exon linker complex, eukaryotic initiation factor 4AIII (eIF4AIII) binds to mRNA, and its ATP hydrolysis and phosphate release are regulated by ligand proteins. Present in all eukaryotic cells as well as many bacteria and archaea, these highly conserved enzymes are essential for RNA metabolism from transcription to degradation, thus playing a key role in gene expression. DEAD box proteins uniquely utilize ATP to unwind short double-stranded RNA, remodeling RNA-protein complexes, and also function as ATP-dependent RNA clamps, providing nucleation sites for the formation of larger RNA-protein complexes. Summary of the Invention
[0004] The purpose of this invention is to increase wheat yield.
[0005] This invention first protects the application of the protein TaeIF4A-6A, which can be S1) or S2).
[0006] S1) Increase wheat yield, number of grains per main spike, plant height, number of tillers and / or spike length;
[0007] S2) Develop transgenic wheat with increased yield, number of grains per spike, plant height, number of tillers and / or spike length.
[0008] In the above applications, the protein TaeIF4A-6A can be a1) or a2).
[0009] a1) The amino acid sequence is that of the protein shown in SEQ ID No. 3;
[0010] a2) A fusion protein with the same function is obtained by attaching a tag or signal peptide to the N-terminus and / or C-terminus of a1).
[0011] Of these, SEQ ID No.3 consists of 413 amino acid residues.
[0012] This invention also protects the application of nucleic acid molecules encoding any of the aforementioned proteins TaeIF4A-6A or biological materials containing said nucleic acid molecules, which may be S1) or S2).
[0013] S1) Increase wheat yield, number of grains per main spike, plant height, number of tillers and / or spike length;
[0014] S2) Develop transgenic wheat with increased yield, number of grains per spike, plant height, number of tillers and / or spike length.
[0015] In the above applications, the nucleic acid molecule encoding any of the aforementioned proteins TaeIF4A-6A can be a DNA molecule as shown in d1), d2), d3), or d4):
[0016] d1) The coding region is the DNA molecule shown in SEQ ID No. 2;
[0017] d2) A DNA molecule with a nucleotide sequence as shown in SEQ ID No. 2 or SEQ ID No. 1;
[0018] d3) A DNA molecule that has 90% or more identity with the nucleotide sequence defined by d1) or d2) and is derived from wheat and encodes any of the proteins TaeIF4A-6A described above;
[0019] d4) Hybridizes under stringent conditions to the nucleotide sequence defined by d1) or d2) a DNA molecule derived from wheat that encodes any of the proteins TaeIF4A-6A described above.
[0020] The nucleic acid molecule can be DNA, such as cDNA, genomic DNA, or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA or hnRNA.
[0021] Of these, SEQ ID No. 2 consists of 1242 nucleotides, and SEQ ID No. 1 consists of 1624 nucleotides. The nucleotides shown in SEQ ID No. 1 or SEQ ID No. 2 encode the amino acid sequence shown in SEQ ID No. 3.
[0022] Those skilled in the art can readily mutate the nucleotide sequence encoding the protein TaeIF4A-6A of this invention using known methods, such as directed evolution and point mutation. Any artificially modified nucleotides that have 90% or higher identity with the nucleotide sequence of the protein TaeIF4A-6A isolated in this invention, as long as they encode the protein TaeIF4A-6A, are derived from and equivalent to the nucleotide sequence of this invention.
[0023] The term "identity" as used herein refers to sequence similarity to a natural nucleic acid sequence. "Identity" includes nucleotide sequences that have 90% or higher, or 95% or higher, identity with the nucleotide sequence encoding the amino acid sequence of the protein TaeIF4A-6A as shown in SEQ ID No. 3 of this invention. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0024] In the above applications, the biomaterial may be at least one of D1)-D4):
[0025] D1) An expression cassette containing the nucleic acid molecule;
[0026] D2) A recombinant vector containing the nucleic acid molecule or a recombinant vector containing the expression cassette;
[0027] D3) Recombinant microorganisms containing the nucleic acid molecule, recombinant microorganisms containing the expression cassette, or recombinant microorganisms containing the recombinant vector;
[0028] D4) Wheat containing the nucleic acid molecule, wheat containing the expression cassette, or wheat containing the recombinant vector; wherein the wheat is wheat cell, wheat tissue, and / or wheat organ.
[0029] The expression cassette may include a promoter, a nucleic acid molecule encoding the protein TaeIF4A-6A, and a terminator.
[0030] The recombinant vector can be a recombinant plasmid obtained by inserting a nucleic acid molecule encoding any of the proteins TaeIF4A-6A described above into an expression vector.
[0031] The recombinant microorganism can be obtained by introducing any of the above-described recombinant vectors into the starting microorganism.
[0032] The starting microorganism may be yeast, bacteria, algae, or fungi. The bacteria may be Gram-positive or Gram-negative. The Gram-negative bacteria may be Agrobacterium tumefaciens or Escherichia coli.
[0033] In any of the above applications, the yield may be the yield per plant.
[0034] The yield mentioned above can be reflected in the number of grains per ear.
[0035] The present invention also protects a method for cultivating transgenic wheat A, which may include the following steps: increasing the expression level and / or activity of any of the proteins TaeIF4A-6A described above in the starting wheat to obtain transgenic wheat A; compared with the starting wheat, the yield, number of grains in the main spike, plant height, number of tillers and / or spike length of transgenic wheat A are increased.
[0036] In the above method, the expression level and / or activity of any of the proteins TaeIF4A-6A in the starting wheat can be increased by methods well known in the art, such as transgenic technology, multiple copying, alteration of promoters, and regulatory factors, to achieve the effect of increasing the expression level and / or activity of any of the proteins TaeIF4A-6A in the starting wheat.
[0037] In the above method, the increase in the expression level and / or activity of any of the proteins TaeIF4A-6A in the starting wheat can be achieved by introducing nucleic acid molecules of any of the proteins TaeIF4A-6A into the starting wheat.
[0038] In the above method, the introduction of nucleic acid molecules encoding any of the proteins TaeIF4A-6A into the starting wheat can be achieved by introducing a recombinant vector into the starting wheat; the recombinant vector can be a recombinant plasmid obtained by inserting a nucleic acid molecule encoding any of the proteins TaeIF4A-6A into an expression vector.
[0039] The recombinant vector may specifically be the overexpression vector pUbi::TaeIF4A-6A mentioned in the embodiments.
[0040] The genetically modified wheat A mentioned in Example 2 can specifically be OE-eIF4A#3 and OE-eIF4A#4. The wheat used in this case is specifically the wheat variety Kenong 199 mentioned in the example.
[0041] This invention also protects a method for cultivating transgenic wheat B, comprising the following steps: mutating the gene encoding the protein TaeIF4A in the starting wheat to obtain transgenic wheat B; compared with the starting wheat, the plant height of transgenic wheat B is reduced;
[0042] The protein TaeIF4A may be at least one of the proteins TaeIF4A-6A, TaeIF4A-6B, and TaeIF4A-6D described above.
[0043] The protein TaeIF4A can specifically be any of the proteins TaeIF4A-6A or TaeIF4A-6D mentioned above.
[0044] The protein TaeIF4A may specifically be composed of any of the proteins TaeIF4A-6A and TaeIF4A-6D described above.
[0045] The protein TaeIF4A-6B mentioned above can be b1) or b2).
[0046] b1) The amino acid sequence is that of the protein shown in SEQ ID No. 6;
[0047] b2) A fusion protein with the same function is obtained by attaching a tag or signal peptide to the N-terminus and / or C-terminus of b1).
[0048] Of these, SEQ ID No. 6 consists of 412 amino acid residues.
[0049] The protein TaeIF4A-6D mentioned above can be c1) or c2).
[0050] c1) The amino acid sequence is that of the protein shown in SEQ ID No. 9;
[0051] c2) A fusion protein with the same function is obtained by attaching a tag or signal peptide to the N-terminus and / or C-terminus of c1).
[0052] Of these, SEQ ID No. 9 consists of 413 amino acid residues.
[0053] The gene encoding any of the aforementioned proteins TaeIF4A-6A is the nucleic acid molecule encoding any of the aforementioned proteins TaeIF4A-6A.
[0054] The gene encoding any of the aforementioned proteins TaeIF4A-6B is the nucleic acid molecule encoding any of the aforementioned proteins TaeIF4A-6B.
[0055] The gene encoding any of the aforementioned proteins TaeIF4A-6D encodes the nucleic acid molecule of any of the aforementioned proteins TaeIF4A-6D.
[0056] The nucleic acid molecule encoding any of the proteins TaeIF4A-6B described above can be a DNA molecule as shown in e1), e2), e3), or e4):
[0057] e1) The coding region is the DNA molecule shown in SEQ ID No. 5;
[0058] e2) The nucleotide sequence is a DNA molecule as shown in SEQ ID No. 5 or SEQ ID No. 4;
[0059] e3) has 90% or more identity with the nucleotide sequence defined by e1) or e2), and is derived from wheat and encodes a DNA molecule that encodes any of the proteins TaeIF4A-6B described above;
[0060] e4) hybridization under stringent conditions with the nucleotide sequence defined by e1) or e2) to a DNA molecule derived from wheat that encodes any of the proteins TaeIF4A-6B described above.
[0061] As used herein, the term "identity" refers to sequence similarity to a natural nucleic acid sequence. "Identity" includes nucleotide sequences that have 90% or higher, or 95% or higher, identity with the nucleotide sequence encoding the amino acid sequence of the protein TaeIF4A-6B as shown in SEQ ID No. 6. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0062] The nucleic acid molecule encoding any of the proteins TaeIF4A-6D described above can be a DNA molecule as shown in f1), f2), f3), or f4):
[0063] f1) The coding region is the DNA molecule shown in SEQ ID No. 8;
[0064] f2) The nucleotide sequence is a DNA molecule as shown in SEQ ID No. 8 or SEQ ID No. 7;
[0065] f3) has 90% or more identity with the nucleotide sequence defined by f1) or f2), and is derived from wheat and encodes a DNA molecule that encodes any of the proteins TaeIF4A-6D described above;
[0066] f4) hybridization under stringent conditions with the nucleotide sequence defined by f1) or f2) of a DNA molecule derived from wheat that encodes any of the proteins TaeIF4A-6D described above.
[0067] As used herein, the term "identity" refers to sequence similarity to a natural nucleic acid sequence. "Identity" includes nucleotide sequences that have 90% or higher, or 95% or higher, identity with the nucleotide sequence encoding the amino acid sequence of the protein TaeIF4A-6D as shown in SEQ ID No. 9. Identity can be evaluated visually or using computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0068] The nucleic acid molecule described above can be DNA, such as cDNA, genomic DNA, or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA or hnRNA.
[0069] Those skilled in the art can readily mutate the nucleotide sequence encoding the protein (protein TaeIF4A-6B or protein TaeIF4A-6D) of this invention using known methods, such as directed evolution and point mutation. Any artificially modified nucleotides having 90% or higher identity with the nucleotide sequence of the protein isolated according to this invention, provided they encode the protein, are derived from and equivalent to the nucleotide sequence of this invention.
[0070] In the above method, the mutation is achieved by gene editing, T-DNA insertion, RNA interference, homologous recombination, zinc finger nuclease or transcription activator-like effector nuclease, to induce missense mutations, DNA fragment deletions and / or DNA fragment insertions in the exons of the gene encoding the protein TaeIF4A.
[0071] In the above method, the gene encoding any of the proteins TaeIF4A in wheat from which the mutation originates can be achieved through (A1), (A2), or (A3).
[0072] The (A1) may be the protein TaeIF4A-6A encoded by SEQ ID No. 2. TaeIF4A-6A Gene mutation TaeIF4A-6A / -1bp Gene;
[0073] The TaeIF4A-6A / -1bp The gene is a DNA molecule obtained by deleting nucleotide A at position 349 from the 5' end of SEQ ID No. 2, while keeping the other nucleotide sequences of SEQ ID No. 2 unchanged.
[0074] The (A2) may be the protein TaeIF4A-6D encoded by SEQ ID No. 8. TaeIF4A-6D Gene mutation TaeIF4A-6D / -1bp Gene;
[0075] The TaeIF4A-6D / -1bp The gene is a DNA molecule obtained by deleting nucleotide A at position 349 from the 5' end of SEQ ID No. 8, while keeping the other nucleotide sequences of SEQ ID No. 8 unchanged;
[0076] The (A3) may be the protein TaeIF4A-6A encoded by SEQ ID No. 2. TaeIF4A-6A Gene mutation TaeIF4A-6A / -2bp The gene and the protein TaeIF4A-6D encoded by SEQ ID No. 8. TaeIF4A-6D Gene mutation as described TaeIF4A-6D / -1bp Gene;
[0077] The TaeIF4A-6A / -2bp The gene is a DNA molecule obtained by deleting two nucleotides CA from position 347-348 of SEQ ID No. 2 starting from the 5' end, while keeping the other nucleotide sequences of SEQ ID No. 2 unchanged.
[0078] In the above method, the gene encoding the protein TaeIF4A in the mutant starting wheat is obtained by introducing a gene editing system into the starting wheat.
[0079] The gene editing system may include a gene editing vector. The gene editing vector may contain a DNA molecule expressing gRNA that targets the gene encoding the protein TaeIF4A.
[0080] Preferably, the target sequence of the gRNA may be as shown in SEQ ID No. 2, positions 332-351 from the 5' end.
[0081] In the above applications, the substance that mutates the expression level and / or activity of the protein TaeIF4A can be either B1) or B2).
[0082] B1) Nucleic acid molecules that reduce the expression of the gene encoding the protein TaeIF4A;
[0083] B2) Expression cassettes, recombinant vectors, recombinant microorganisms, or transgenic wheat cell lines containing the nucleic acid molecules described in B1).
[0084] Preferably, the nucleic acid molecule in B1) may be a DNA molecule expressing gRNA that targets the gene encoding the protein TaeIF4A or gRNA that targets the gene encoding the protein TaeIF4A.
[0085] Preferably, the target sequence of the gRNA may be as shown in SEQ ID No. 2, positions 332-351 from the 5' end.
[0086] In the above method, the starting wheat can be the wheat variety Kenong 199.
[0087] Experiments have shown that overexpression of the protein TaeIF4A-6A can significantly increase wheat plant height, tiller number, spike length, number of grains per main spike, and yield per plant; TaeIF4A-6A Genes and / or TaeIF4A-6D Gene editing occurs when two homologous chromosomes... TaeIF4A-6A Genes and / or TaeIF4A-6D When mutations in all genes lead to premature termination of protein translation, it can significantly reduce wheat plant height, tiller number, spike length, number of grains per spike, and yield per plant. Therefore, the protein TaeIF4A can regulate wheat yield, number of grains per spike, plant height, number of tillers, and / or spike length. This invention has significant application value.
[0088] Terminology Definition
[0089] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, to better understand this invention, definitions and explanations of relevant terms are provided below.
[0090] The term "expression cassette" generally refers to a nucleic acid construct containing sufficient nucleic acid elements to express a target gene. A typical expression cassette includes a promoter, a multiple cloning site (MCS), and / or a terminator. Expression cassettes may also include the target gene, marker genes (such as TK, DHFR, CAT, and NEO genes), ribosome recognition and binding sites (SDs), transcription factor binding sites (TFBSs), enhancers, silencers, repressors, introns, poly(A) signal sequences, and / or mRNA splicing signal sequences. Elements within an expression cassette can be directly linked or indirectly linked through adapters.
[0091] The term "tag" includes, but is not limited to: GST (glutathione thioredoxin) tagged protein, Trx (thioredoxin) tagged protein, nitrogen utilization substrate A (NusA) tagged protein, His-tag protein, MBP (maltose-binding protein) tagged protein, Flag tagged protein, SUMO tagged protein, HA (influenza hemagglutinin) tagged protein, Myc tagged protein, LacZ tagged protein, CBD (cellulose-binding domain) tagged protein, phage T7 protein kinase (T7PK) tagged protein, GFP (green fluorescent protein), CFP (cyan fluorescent protein), YFP (yellow-green fluorescent protein), mCherry (monomer red fluorescent protein), or AviTag tagged protein. Those skilled in the art know how to select appropriate tagged proteins according to the desired purpose. The use of tags does not alter the function of the target protein; its purpose is to separate, purify, detect, or trace it. Therefore, the tagged proteins applicable to this invention are not limited to a specific type. Tags can be separated from the target protein by chemical cleavage methods or enzymatic methods (such as introducing protease cleavage sites to remove the tag using TEV protease).
[0092] The term "vector" generally refers to a vector capable of delivering exogenous DNA or a target gene into host cells for amplification and / or expression. This vector can be a cloning vector or an expression vector. Vectors can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material they carry to be amplified and / or expressed within the host cells. Those skilled in the art can select appropriate vectors based on the purpose of genetic engineering and the properties of the recipient cells. The vectors include, but are not limited to: plasmids, phages (such as λ phage or M13 phage), cosmids (i.e., Cosmids), phagemids, shuttle vectors (such as yeast expression vectors), Ti plasmids, artificial chromosomes (such as yeast artificial chromosomes (YAC), bacterial artificial chromosomes (BAC), P1 artificial chromosomes (PAC), or Ti plasmid artificial chromosomes (TAC)), and viral vectors (such as baculovirus vectors, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, poxviruses, papillomaviruses, papillomaviruses (such as SV40), and herpesviruses (such as herpes simplex virus)). A vector may contain multiple elements controlling expression, including but not limited to promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication origin site.
[0093] The term "microorganism" generally includes bacteria, viruses, fungi, actinomycetes, rickettsiae, mycoplasmas, chlamydiae, spirochetes, algae, etc. For example, the bacteria mentioned could be from the genus *Escherichia* (…). Escherichia sp.(such as Escherichia coli), Erwinia spp. Erwinia sp. ), Agrobacterium ( Agrobacterium sp. (such as Agrobacterium tumefaciens), Flavobacterium spp. ( Flavobacterium sp. Alcaligenes ( ) Alcaligenes sp. ), Pseudomonas spp. Pseudomonas sp. ) and Bacillus spp. ( Bacillus sp. (e.g., Bacillus subtilis). The viruses may include rotavirus, baculovirus, retrovirus (e.g., lentivirus), adenovirus, adeno-associated virus, poxvirus, papillomavirus, influenza virus, papillomavirus (e.g., SV40), and herpesvirus (e.g., herpes simplex virus). The fungi may be derived from yeasts (e.g., Bacillus subtilis). Saccharomyces sp. (such as Saccharomyces cerevisiae, Saccharomyces methylbenzene, Pichia pastoris), Fusarium genus ( Fusarium sp. ), Rhizoctonia spp. ( Rhizoctonia sp. Verticillium ( Verticillium sp. ), Penicillium ( Penicillium sp. Aspergillus ( ) Aspergillus sp. ) and Cephalosporin ( Cephalosporium sp. The actinomycetes may originate from the genus Streptomyces (…). Streptomyces sp. (e.g., Streptomyces). The algae may originate from the phylum Cyanophyta (e.g., cyanobacteria), genus Fucus (e.g., *Fucus*). Fucus sp. ), genus *Cyclocarya* ( Achnanthes sp. ), genus *Codonopsis* ( Amphiprora sp. ), genus Dipterocarpa ( Amphora sp. ), Fiber Algae ( Ankistrodesmus sp. ), genus Styracula ( Asteromonas sp. ) and the genus *Golden Color Algae* ( Boekelovia sp. )wait.
[0094] The term "recombinant vector" generally refers to a recombinant DNA molecule constructed by linking a foreign target gene to a vector in vitro. It can be constructed in any suitable way, as long as the constructed recombinant vector can carry the foreign target gene into the recipient cell and provide the foreign target gene with the ability to replicate, integrate, amplify and / or express in the recipient cell.
[0095] The term "linkage" generally refers to the association of two or more molecules. Linkages can be covalent or non-covalent. The linkages described herein can be direct peptide bonds or linkages via linkers (connectors).
[0096] The term "identity" generally refers to the degree to which two (nucleotide or amino acid) sequences have identical residues at the same position in an alignment, and is usually expressed as a percentage. The identity described herein can refer to the identity of an amino acid sequence or a nucleotide sequence. Two copies having completely identical sequences have 100% identity. Those skilled in the art will recognize that the identity of an amino acid sequence or nucleotide sequence can be determined using identity search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, the identity of an amino acid sequence can be calculated by using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Perresidue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values), and performing a search, thus obtaining the identity value (%). Alternatively, sequence analysis software such as CLC Main Workbench and MegAlign can be used. TM The determination can be performed, for example, using a computer program BLAST with default parameters, especially BLASTP or TBLASTN. The 90% or higher identity mentioned herein can be at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher identity.
[0097] The term "comprising" is not intended to be restrictive, but rather inclusive and implies the presence of other elements besides those listed, and can be interpreted as "including but not limited to". The term "comprising" also encompasses the terms "consisting of" and "substantially composed of". In this document, the terms "comprising" and "including" are used interchangeably. Attached Figure Description
[0098] Figure 1 This is a partial structural diagram of the overexpression vector pUbi::TaeIF4A-6A.
[0099] Figure 2 The relative expression level of the TaeIF4A-6A gene was detected by qRT-PCR.
[0100] Figure 3 For the target point in TaeIF4A A diagram showing the location of a gene.
[0101] Figure 4 T2 generation knockout TaeIF4AThe phenotypes of homozygous mutant strains KO#45, KO#80 and KO#112 at maturity, as well as the statistical results of plant height, tiller number, panicle length, number of grains in the main panicle and yield per plant; data are mean ± SE (n=10-20).
[0102] Figure 5 T2 generation overexpression TaeIF4A-6A Phenotypic results of wheat genes OE-eIF4A#3 and OE-eIF4A#4 at maturity, as well as statistical results of plant height, tiller number, spike length, number of grains per main spike, and yield per plant; data are mean ± SE (n=10-20). Detailed Implementation
[0103] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0104] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0105] In the quantitative experiments in the following examples, three replicate experiments were set up, and the average value of the results was taken.
[0106] KN199 is a common commercially available wheat variety. It was approved by the National Variety Approval Committee in 2006, with the approval number being National Approval Wheat 2006017.
[0107] Example 1: Discovery of the protein TaeIF4A and its encoding gene
[0108] Wheat is an allohexaploid, and its genome consists of three distinct subgenomes (A, B, and D), each derived from a different ancestral species. Through extensive experimentation, the inventors of this application discovered the protein encoding TaeIF4A in KN199. TaeIF4A Genes, and all three genomes of wheat contain them. TaeIF4A Genes were named accordingly. TaeIF4A-6A Gene, TaeIF4A-6B Genes and TaeIF4A-6D The gene encodes proteins TaeIF4A-6A, TaeIF4A-6B, and TaeIF4A-6D, respectively.
[0109] In the genomic DNA of KN199, TaeIF4A-6AThe nucleotide sequence of the gene is shown in SEQ ID No. 1.
[0110] In the cDNA of KN199, TaeIF4A-6A The sequence of the gene coding region is shown in SEQ ID No. 2.
[0111] TaeIF4A-6A The gene encodes the protein TaeIF4A-6A, the amino acid sequence of which is shown in SEQ ID No. 3.
[0112] In the genomic DNA of KN199, TaeIF4A-6B The nucleotide sequence of the gene is shown in SEQ ID No. 4.
[0113] In the cDNA of KN199, TaeIF4A-6B The sequence of the gene coding region is shown in SEQ ID No. 5.
[0114] TaeIF4A-6B The gene encodes the protein TaeIF4A-6B, the amino acid sequence of which is shown in SEQ ID No. 6.
[0115] In the genomic DNA of KN199, TaeIF4A-6D The nucleotide sequence of the gene is shown in SEQ ID No. 7.
[0116] In the cDNA of KN199, TaeIF4A-6D The sequence of the gene coding region is shown in SEQ ID No. 8.
[0117] TaeIF4A-6D The gene encodes the protein TaeIF4A-6D, the amino acid sequence of which is shown in SEQ ID No. 9.
[0118] TaeIF4A-6A Gene, TaeIF4A-6B Genes and TaeIF4A-6D The nucleotide sequence similarity of the genes reaches more than 80%.
[0119] The amino acid sequence similarity of proteins TaeIF4A-6A, TaeIF4A-6B, and TaeIF4A-6D also reached over 90%. 。
[0120] Example 2: Application of protein TaeIF4A in regulating wheat plant height, tiller number, spike length, number of grains per main spike, and yield per plant.
[0121] I. Acquisition of TaeIF4A transgenic materials and gene-editing materials
[0122] (I) Obtaining OE-eIF4A#3 and OE-eIF4A#4 wheat overexpressing the TaeIF4A-6A gene in the T2 generation
[0123] 1. Carrier Construction
[0124] Using wild-type wheat Kenong 199 cDNA as a template, the full-length CDS of TaeIF4A-6A was amplified using the upstream primer TaeIF4A-6A-OE-F (containing a BamHI restriction site and protective bases) and the downstream primer TaeIF4A-6A-OE-R (containing a KpnI restriction site and protective bases). The amplified CDS was then ligated into the pEasy-Blunt vector (purchased from Beijing TransGen Biotech Co., Ltd., catalog number CB101-01). After sequencing verification, the fragment containing the full-length CDS of TaeIF4A-6A was ligated into a plant expression vector via double digestion with BamHI and KpnI. pUbi-163 (hereinafter referred to as) pUbi-163 The overexpression vector pUbi::TaeIF4A-6A was obtained by sequencing and enzyme digestion between the Ubi promoter and Nos terminator in the vector.
[0125] The overexpression vector pUbi::TaeIF4A-6A is a DNA molecule whose nucleotide sequence is SEQ ID No. 2. pUbi-163 An overexpression vector that preserves the nucleotide sequence between the BamHI and KpnI restriction sites while maintaining the integrity of other nucleotides in the vector (see partial structural diagram). Figure 1 ).
[0126] pUbi-163 The vector is described in the non-patent literature "Wenjing Li, Xue He, Yi Chen, YanfuJing, Chuncai Shen, Junbo Yang, Wan Teng, Xueqiang Zhao, Weijuan Hu, Mengyun Hu, Hui Li, Anthony J. Miller and Yiping Tong. A wheat transcription factor positively sets seed vigour by regulating the grain nitrate signal. NewPhytologist (2020) 225: 1667–1680", and its name in the literature is [missing information]. pUbi-163 The vector, which is available to the public from the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences, is a biological material intended solely for repeating experiments related to this invention and may not be used for any other purpose.
[0127] 2. Preparation of TaeIF4A-6A overexpressing plants
[0128] The common wheat cultivar Kenong 199 was selected as the transformation recipient. The overexpression vector pUbi:: TaeIF4A-6A constructed in step 1 was transformed into the recipient common wheat cultivar Kenong 199 using the gene gun-mediated transformation method, and T0 generation transgenic lines were obtained.
[0129] 3. Detection of TaeIF4A-6A overexpressing plants
[0130] (1) Using the genomic DNA of the T0 generation transgenic line as a template, positive seedlings were identified by PCR amplification, and T0 generation transgenic TaeIF4A-6A positive seedlings were obtained.
[0131] The method for identifying positive seedlings by PCR amplification is as follows: PCR amplification of the genomic DNA of the wheat to be tested is performed using primers consisting of OE-TaeIF4A-6A-F: 5'-CACCCAAGCTGTCATTTTCTGC-3' (SEQ ID No. 10) and OE-TaeIF4A-6A-R: 5'-TCCTTGAAGTCGATGCCCTT-3' (SEQ ID No. 11), respectively. If a PCR amplification product of approximately 846 bp can be obtained, the wheat to be tested is a positive seedling.
[0132] (2) qRT-PCR detection of the relative expression level of TaeIF4A-6A gene in T2 generation transgenic TaeIF4A-6A gene-positive seedlings
[0133] ① The T0 generation of TaeIF4A-6A gene-positive seedlings were self-crossed twice, and the positive seedlings were identified by PCR amplification in each generation (the method for identifying positive seedlings by PCR amplification is shown in step (1)). Finally, the T2 generation of TaeIF4A-6A gene-positive seedlings were obtained. RNA was extracted from the aboveground part of the T2 generation of TaeIF4A-6A gene-positive seedlings in the field using the Trizol method, and RT-PCR was performed using Thermo Fisher's K1622 reverse transcription kit to reverse transcribe it into cDNA.
[0134] ② Using the cDNA obtained in step ① as a template, the relative expression levels of the TaeIF4A-6A gene were detected by qRT-PCR (with the TaActin gene as an internal reference).
[0135] The primers for detecting the TaeIF4A-6A gene are RT-TaeIF4A-6A-F: 5'-GAACCCCGGGCAGAGTGTGC-3' (SEQ ID No. 12) and RT-TaeIF4A-6A-R: 5'-ACAGCTTGGGTGATGGTAAGTGTATCATAAAGATCACATA-3' (SEQ ID No. 13).
[0136] The primers for detecting the TaActin gene are RT-TaActin-F: 5'-ACCTTCAGTTGCCCAGCAAT-3' (SEQ ID No. 14) and RT-TaActin-R: 5'-CAGAGTCGAGCACAATACCAGTTG-3' (SEQ ID No. 15).
[0137] Some test results can be found Figure 2 (** indicates a significant difference, with the difference reaching a certain threshold.) P (<0.01 level). The results showed that, compared with KN199, the relative expression level of the TaeIF4A-6A gene was significantly increased in two lines of T2 generation TaeIF4A-6A gene-positive seedlings. These two lines were named OE-eIF4A#3 and OE-eIF4A#4, namely, T2 generation wheat overexpressing the TaeIF4A-6A gene OE-eIF4A#3 and OE-eIF4A#4.
[0138] (ii) T2 generation knockout TaeIF4A Obtaining homozygous mutant strains KO#45, KO#80, and KO#112
[0139] 1. Carrier Construction
[0140] The gene editing vector pYLCRISPR / Cas9-TaeIF4A is a recombinant plasmid obtained by inserting the DNA fragment shown in the TaeIF4A-sgRNA expression cassette into the BsaI restriction site of the pYLCRISPR / Cas9 vector, while keeping other nucleotide sequences unchanged. The target sequence targeted by the TaeIF4A-sgRNA expression cassette is 5'-CCGCTGATACTGTCGCCAACTAGG-3' (as shown in SEQ ID No. 2, positions 332-354 from the 5' end). Its location on the TaeIF4A gene is illustrated in the diagram below. Figure 3 .
[0141] The pYLCRISPR / Cas9 vector is described in the non-patent literature “Hong Yu, Tao Lin, Xiangbing Meng, Huilong Du, Jingkun Zhang, Guifu Liu, Mingjiang Chen, Yanhui Jing, LiquanKou, Xiuxiu Li, Qiang Gao, Yan Liang, Xiangdong Liu, Zhilan Fan, YuntaoLiang, Zhukuan Cheng, Mingsheng Chen, Zhixi Tian, Yonghong Wang, ChengcaiChu, Jianru Zuo, Jianmin Wan, Qian Qian, Bin Han, Andrea Zuccolo, Rod A.Wing, Caixia Gao, Chengzhi Liang, Jiayang Li, A route to de novodomestication of wild allotetraploid rice, Cell, Volume 184, Issue 5, 2021, Pages 1156-1170.e14, ISSN”. 0092-8674, https: / / doi.org / 10.1016 / j.cell.2021.01.013., whose name in the literature is "pYLCRISPR / Cas9Pubi-H", is available to the public from the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. This biological material is only for repeating the relevant experiments of this invention and should not be used for other purposes.
[0142] 2. Construction of gene knockout plants
[0143] The common wheat cultivar Kenong 199 was selected as the transformation recipient. The gene editing vector pYLCRISPR / Cas9-TaeIF4A was transformed into the recipient common wheat cultivar Kenong 199 using the gene gun-mediated transformation method with a minimal expression frame (MFF) to obtain the T0 generation knockout. TaeIF4A Gene mutants.
[0144] 3. Detection of gene knockout plants
[0145] (1) Twenty T0 knockout strains were extracted using the CTAB method. TaeIF4A Genomic DNA of the gene mutant was collected, and then PCR was performed using the following primers to identify the mutation status of groups A, B, and D, respectively, to obtain the mutant strains:
[0146] TaeIF4A-6A-F-cr: 5'-TATGTCCTATTATCTCATGAGTCAA-3' (SEQ ID No. 16);
[0147] TaeIF4A-6A-R-cr: 5'-CCTTACATGCATCCAGAAAAAAAAT-3' (SEQ ID No. 17);
[0148] TaeIF4A-6D-F-cr: 5'-CTGCCAGGATCAATGGTTCCG-3' (SEQ ID No. 18);
[0149] TaeIF4A-6D-R-cr: 5'-ACCATATTAATTGTCAAGCTGAAATGTT-3' (SEQ ID No. 19).
[0150] The specific identification method is as follows: Using the genomic DNA of the test strain as a template, PCR amplification was performed using primer pair 1 (TaeIF4A-6A-F-cr and TaeIF4A-6A-R-cr) and primer pair 2 (TaeIF4A-6D-F-cr and TaeIF4A-6D-R-cr), respectively, to obtain PCR amplification product 1 and PCR amplification product 2. PCR amplification product 1 and PCR amplification product 2 were then sequenced. The sequencing results were compared with... TaeIF4A-6A Gene, TaeIF4A-6B Genes and TaeIF4A-6D The gene editing target sequence is compared, and the mutation types are counted.
[0151] (2) The mutant strains identified in step (1) are subjected to two generations of self-crossing, with identification performed in each generation (the identification method is described in step (1)) until homozygous mutant strains are obtained. Homozygous mutation refers to a mutation in the two homologous chromosomes of the wheat. TaeIF4A Gene( TaeIF4A-6A Gene, TaeIF4A-6B Gene or TaeIF4A-6D The same mutation occurred in the gene.
[0152] Three T2 generation knockout strains were finally obtained through screening. TaeIF4A The homozygous mutant strains were named KO#45, KO#80, and KO#112, respectively. Details are as follows:
[0153] KO#45, TaeIF4A-6B Genes and TaeIF4A-6D There is no change in the genes; in genome A, on two homologous chromosomes TaeIF4A-6AThe gene has the same mutation—a deletion of one nucleotide, specifically the deletion of nucleotide A at position 349 from the 5' end in SEQ ID No. 2; this deletion causes a frameshift, leading to premature termination of translation (a stop codon appears at the position encoding amino acid 126).
[0154] KO#80, TaeIF4A-6A Genes and TaeIF4A-6B There are no changes in the genes; in the D genome, on two homologous chromosomes TaeIF4A-6D The gene has the same mutation—a deletion of one nucleotide, specifically the deletion of nucleotide A at position 349 from the 5' end in SEQ ID No. 8; this deletion causes a frameshift, leading to premature termination of translation (a stop codon appears at the position encoding amino acid 126).
[0155] KO#112, TaeIF4A-6B There is no change in the genes; in genome A, on two homologous chromosomes TaeIF4A-6A The gene underwent the same mutation—a two-nucleotide deletion, specifically the deletion of two CA nucleotides at positions 347-348 from the 5' end of SEQ ID No. 2. This deletion caused a frameshift, leading to premature termination of translation (a stop codon appears at the position encoding amino acid 116), and consequently, a decrease in the expression level of the protein TaeIF4A-6A. In the D genome, on two homologous chromosomes... TaeIF4A-6D The gene has undergone the same mutation—a deletion of one nucleotide, specifically the deletion of nucleotide A at position 349 from the 5' end in SEQ ID No. 8; this deletion causes a frameshift, leading to premature termination of translation (a stop codon appears at the position encoding amino acid 126).
[0156] II. Phenotypic Identification
[0157] Wheat varieties Kenong 199, KO#45, KO#80, KO#112, OE-eIF4A#3, and OE-eIF4A#4 were planted. At maturity, agronomic traits such as plant height, number of tillers, ear length, number of grains per main ear, and yield per plant were collected for these wheat varieties.
[0158] For detailed statistical results on the maturity phenotypes of wheat varieties Kenong 199, KO#45, KO#80, KO#112, OE-eIF4A#3, and OE-eIF4A#4, as well as their plant height, tiller number, spike length, number of grains per main spike, and yield per plant, please refer to [link to relevant documentation]. Figure 4 and Figure 5 .use tThe test method is to test for the significance of differences; where ns indicates no significant difference; * and ** indicate significant differences, respectively, representing differences reaching a certain threshold. P <0.05 and P <0.01 level. The results are as follows:
[0159] Compared to KN199, the T2 generation knockout TaeIF4A The plant height of the homozygous mutant strains KO#45, KO#80, and KO#112 was significantly reduced: KO#45 decreased by an average of 11.80%, approximately 8.03 cm, from 68.11 cm to 60.08 cm; KO#80 decreased by an average of 9.48%, approximately 6.46 cm, from 68.11 cm to 61.65 cm; and KO#112 decreased by an average of 21.57%, approximately 14.69 cm, from 68.11 cm to 53.42 cm. Compared with KN199, overexpression... TaeIF4A-6A The plant height of both genetically modified wheat varieties OE-eIF4A#3 and OE-eIF4A#4 was significantly increased: OE-eIF4A#3 increased by an average of 3.14%, or about 2.14 cm, from 68.11 cm to 70.25 cm; OE-eIF4A#4 increased by an average of 3.72%, or about 2.61 cm, from 68.11 cm to 70.72 cm.
[0160] Compared to KN199, the T2 generation knockout TaeIF4A The number of tillers in the homozygous mutant strains KO#45, KO#80, and KO#112 was significantly reduced: KO#45 from 12.24 to 9.2, with an average decrease of 3.04; KO#80 from 12.24 to 9.3, with an average decrease of 2.94; and KO#112 from 12.24 to 6.67, with an average decrease of 5.57. Compared with KN199, overexpression... TaeIF4A-6A The number of tillers in both the OE-eIF4A#3 and OE-eIF4A#4 genetically modified wheat varieties increased significantly: OE-eIF4A#3 increased from 12.24 to 13.67 tillers, with an average increase of 1.43 tillers; OE-eIF4A#4 increased from 12.24 to 13.72 tillers, with an average increase of 1.49 tillers.
[0161] Compared to KN199, the T2 generation knockout TaeIF4AThe spike length of the homozygous mutant strains KO#45, KO#80, and KO#112 was significantly reduced: KO#45 decreased by an average of 15.05%, approximately 1.09 cm, from 7.22 cm to 6.13 cm; KO#80 decreased by an average of 14.01%, approximately 1.01 cm, from 7.22 cm to 6.21 cm; and KO#112 decreased by an average of 22.44%, approximately 1.62 cm, from 7.22 cm to 5.60 cm. Compared with KN199, overexpression... TaeIF4A-6A Both OE-eIF4A#3 and OE-eIF4A#4 wheat spike lengths increased significantly: OE-eIF4A#3 saw an average increase of 5.10%, approximately 0.37 cm, from 7.22 cm to 7.59 cm; OE-eIF4A#4 saw an average increase of 9.55%, approximately 0.69 cm, from 7.22 cm to 7.91 cm.
[0162] Compared to KN199, the T2 generation knockout TaeIF4A The number of grains in the main ear of the homozygous mutant strains KO#45, KO#80, and KO#112 was significantly reduced: KO#45 saw an average reduction of 16.06%, approximately 9 grains, from 56.06 to 47.06; KO#80 saw an average reduction of 18.84%, approximately 10.56 grains, from 56.06 to 45.50; and KO#112 saw an average reduction of 25.37%, approximately 14.22 grains, from 56.06 to 41.83. Compared with KN199, overexpression... TaeIF4A-6A The number of grains in the main spike of genetically modified wheat OE-eIF4A#3 and OE-eIF4A#4 both increased significantly: OE-eIF4A#3 increased by an average of 10.70%, about 6 grains, from 56.06 to 62.06 grains; OE-eIF4A#4 increased by an average of 11.30%, about 6.33 grains, from 56.06 to 62.39 grains.
[0163] Compared to KN199, the T2 generation knockout TaeIF4A The homozygous mutants KO#45, KO#80, and KO#112 all showed significantly reduced yields per plant: KO#45 saw an average decrease of approximately 41.33%, from 23.03 g to 13.51 g; KO#80 saw an average decrease of approximately 41.12%, from 23.03 g to 13.56 g; and KO#112 saw an average decrease of approximately 53.16%, from 23.03 g to 10.79 g. Compared to KN199, overexpression... TaeIF4A-6ABoth genetically modified wheat varieties OE-eIF4A#3 and OE-eIF4A#4 showed significant increases in yield per plant: OE-eIF4A#3 increased by an average of approximately 14.78%, from 23.03 g to 26.44 g; OE-eIF4A#4 increased by an average of approximately 15.95%, from 23.03 g to 26.71 g.
[0164] In conclusion, overexpression of the protein TaeIF4A-6A can significantly increase wheat plant height, tiller number, spike length, number of grains per main spike, and yield per plant. This also has an effect on the gene encoding TaeIF4A. TaeIF4A Gene editing can significantly reduce wheat plant height, tiller number, spike length, number of grains per spike, and yield per plant.
[0165] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. Application of protein TaeIF4A-6A, for S1) or S2): S1) Increase wheat yield, number of grains per main spike, plant height, number of tillers and / or spike length; S2) Develop transgenic wheat with increased yield, number of grains per spike, plant height, number of tillers and / or spike length; The protein TaeIF4A-6A is either a1) or a2). a1) The amino acid sequence is that of the protein shown in SEQ ID No. 3; a2) A fusion protein with the same function is obtained by attaching a tag or signal peptide to the N-terminus and / or C-terminus of a1).
2. The application of a nucleic acid molecule encoding the protein TaeIF4A-6A as described in claim 1, or a biological material containing said nucleic acid molecule, is S1) or S2). S1) Increase wheat yield, number of grains per main spike, plant height, number of tillers and / or spike length; S2) Develop transgenic wheat with increased yield, number of grains per spike, plant height, number of tillers and / or spike length.
3. The application according to claim 2, characterized in that: The biomaterial is at least one of D1)-D4): D1) An expression cassette containing the nucleic acid molecule; D2) A recombinant vector containing the nucleic acid molecule or a recombinant vector containing the expression cassette; D3) Recombinant microorganisms containing the nucleic acid molecule, recombinant microorganisms containing the expression cassette, or recombinant microorganisms containing the recombinant vector; D4) Wheat containing the nucleic acid molecule, wheat containing the expression cassette, or wheat containing the recombinant vector; wherein the wheat is wheat cell, wheat tissue, and / or wheat organ.
4. The application according to claim 1 or 2, characterized in that: The yield mentioned refers to the yield per plant.
5. A method for breeding transgenic wheat, comprising the following steps: increasing the expression level and / or activity of the protein TaeIF4A-6A as described in claim 1 in the starting wheat to obtain transgenic wheat; and compared with the starting wheat, the yield, number of grains in the main spike, plant height, number of tillers and / or spike length of the transgenic wheat are increased.
6. The method according to claim 5, characterized in that: The improvement in the expression level and / or activity of the protein TaeIF4A-6A of claim 1 in the starting wheat is achieved by introducing a nucleic acid molecule encoding the protein TaeIF4A-6A of claim 1 into the starting wheat.