Codon optimization method for increasing protein expression of target gene in plant
By replacing the third base of the codon with C or G in transgenic plants, the coding sequence of the target protein in dicotyledonous and monocotyledonous plants is optimized, solving the problem of insufficient protein expression in existing technologies and achieving efficient protein expression and cost reduction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2026-04-02
AI Technical Summary
Existing codon optimization methods cannot effectively increase the expression level of target proteins in transgenic plants, and the enumeration method for screening is costly and lacks universality and effectiveness.
Replace the third base of the codon in the target protein coding sequence with C or G, optimize the GC content of the target protein coding sequence to be no less than 55%, especially replace the codon corresponding to threonine with ACC, and apply it to dicotyledonous and monocotyledonous plants.
It significantly increased the expression level of the target protein in plants, with the protein expression level in transient expression systems and transgenic plants increasing by several to tens of times, reducing screening costs, and enhancing the target traits of transgenic plants and reducing protein purification costs.
Smart Images

Figure CN2025118217_02042026_PF_FP_ABST
Abstract
Description
A method for improving the protein expression amount of a target gene in plants by codon optimization
[0001] Cross-reference to Related Applications
[0002] This application claims priority to the Chinese patent application No. 202411393933.1, filed on September 30, 2024, and entitled “A method for improving the protein expression amount of a target gene in plants by codon optimization”, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present application relates to a method for improving the protein expression amount of a target gene in plants by codon optimization, and belongs to the technical field of genetic engineering. BACKGROUND
[0004] In organisms, the DNA sequence of a gene determines the sequence of messenger RNA, and the sequence of messenger RNA determines the sequence of amino acids in a protein. A codon refers to the coding order and rule of three adjacent nucleotides in a messenger RNA or its corresponding DNA molecule. During protein synthesis, the codon order on DNA / RNA is translated by ribosomes into a specific amino acid at a specific position on the target protein. All proteins in organisms are composed of 20 different amino acids; most amino acids are encoded by multiple different codons; and codons encoding the same amino acid are called synonymous codons.
[0005] Generally, a codon with a higher frequency of use is referred to as an optimal codon, and a codon with a lower frequency of use is referred to as a rare codon. Different species have different synonymous codon usage frequencies, resulting in corresponding optimal codons for each species. For example, in 2022, the David Alvarez-Ponce laboratory in the United States statistically analyzed the codon usage of more than 15000 species, and labeled the corresponding optimal codons according to the codon usage frequency (see document: Subramanian K., Payne B., Feyertag F., Alvarez-Ponce D. (2022) The Codon Statistics Database: A Database of Codon Usage Bias. Mol Biol Evol 39: msac157).
[0006] At present, it is generally believed that rare codons will reduce protein translation efficiency and protein expression amount, and therefore, in order to improve protein expression amount, the rare codons in the coding sequence of the target protein are usually replaced by optimal codons. However, in actual operation, according to the existing codon optimization method, the ideal protein expression amount cannot be obtained. It can be seen that although it is known that the expression amount of the same protein encoded by different codons in the organism is different, what is the optimal codon that can effectively improve the protein expression amount of the coding sequence of the target protein in the organism is still controversial, and at the present stage, there is still a lack of some truly effective and universal codon optimization method.
[0007] At the same time, since the different arrangement and combination of codons is an astronomical number, when the existing codon optimization method cannot obtain the ideal protein expression amount, considering the time cost and economic cost, it is also impossible to synthesize and detect all possible codon arrangement and combination one by one by enumeration method. For example, assuming that a protein composed of only 100 amino acids, each amino acid appears 5 times, then the DNA coding sequence corresponding to this protein theoretically exists 1 5 ×1 5 ×2 5 ×2 5 ×2 5 ×2 5 ×2 5 ×2 5 ×2 5 ×2 5 ×2 5 ×3 5 ×4 5 ×4 5 ×4 5 ×4 5 ×4 5 ×6 5 ×6 5 ×6 5 = 8358844170240 kinds (i.e. 8.3 trillion kinds) of possibilities.
[0008] Genetically modified plants refer to plants in which specific recombinant DNA fragments are integrated into the plant genome by genetic engineering technology, so as to produce plants expressing recombinant RNA and recombinant protein encoded by the recombinant DNA. The application scenarios of genetically modified plants include changing the agronomic traits of crops, obtaining recombinant medical proteins or recombinant nutritional proteins for human or animals, etc. The trait change of genetically modified plants can come from the cross between different individuals of species / variety, or from the artificial insertion of specific functional genes by genetic engineering (recombinant DNA) technology.
[0009] Generally, the traits of transgenic plants are positively correlated with the expression amount of the recombinant proteins, and the expression amount of the recombinant proteins is often an important factor limiting the application of excellent genes. At present, it is generally believed that the expression amount of the exogenous proteins is determined by the position effect of the inserted recombinant genes in the genome, and therefore, a large number of transgenic materials are screened to obtain transgenic materials with high expression amount of the recombinant proteins. This process requires a large amount of manpower, material resources, financial resources, time and land cost, and even so, many genes cannot obtain transgenic plant materials with high expression amount of the recombinant proteins. If a codon optimization method with high effectiveness and strong universality for plants can be found, the expression amount of the recombinant proteins in plants can be improved, the target traits of the transgenic plants can be enhanced, and the screening cost of the high-expression transgenic plants can be reduced. SUMMARY
[0010] To solve the above problems, the present application provides a codon optimization method for improving the protein expression amount of a target gene in plants, which comprises: replacing the third base of the codon of the target protein coding sequence with C or G to obtain an optimized target protein coding sequence.
[0011] In an embodiment of the present application, the GC content of the optimized target protein coding sequence is not less than 55%.
[0012] In an embodiment of the present application, in the optimized target protein coding sequence, the GC content (GC3) of the third base of all codons is not less than 70%.
[0013] In an embodiment of the present application, in the method, when the codon is replaced, the third base of the codon corresponding to threonine is replaced with C (i.e., the codon corresponding to threonine is replaced with ACC).
[0014] In an embodiment of the present application, in the optimized target protein coding sequence, more than 50% of the segments have undergone replacement of the third base of the codon.
[0015] In an embodiment of the present application, the target protein coding sequence comprises a DNA sequence encoding a target protein, an RNA sequence encoding a target protein, a complementary sequence of a DNA sequence encoding a target protein, and / or a complementary sequence of an RNA sequence encoding a target protein.
[0016] In an embodiment of the present application, the target protein coding sequence comprises a single strand and / or a double strand.
[0017] In an embodiment of the present application, the plant comprises a dicotyledonous plant and / or a monocotyledonous plant.
[0018] In an embodiment of the present application, the dicotyledonous plant comprises Arabidopsis thaliana, tobacco, soybean, cotton, lettuce, poplar, alfalfa, rape and / or peanut.
[0019] In an embodiment of the present application, the monocotyledonous plant comprises corn, rice, sorghum, wheat, sugarcane and / or bamboo.
[0020] In an embodiment of the present application, the protein of interest comprises a functional protein.
[0021] In an embodiment of the present application, the functional protein comprises a functional protein for enhancing a target trait of a transgenic plant and / or a functional protein produced by using a plant as a bioreactor.
[0022] In an embodiment of the present application, the functional protein comprises a CRY1a.105 protein, a CRY2Ab protein, a Bar protein, a CP4 protein, a new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), a new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, an influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), a giant panda obesity-related protein AmFTO, a red source chicken obesity-related protein GgFTO and / or a fish immune protein OprF.
[0023] The present application also provides a nucleic acid molecule for expressing a protein of interest in a plant, wherein the nucleic acid molecule is optimized by using the above-mentioned codon optimization method.
[0024] In an embodiment of the present application, the plant comprises a dicotyledonous plant and / or a monocotyledonous plant.
[0025] In an embodiment of the present application, the dicotyledonous plant comprises Arabidopsis thaliana, tobacco, soybean, cotton, lettuce, poplar, alfalfa, rape and / or peanut.
[0026] In an embodiment of the present application, the monocotyledonous plant comprises corn, rice, sorghum, wheat, sugarcane and / or bamboo.
[0027] In an embodiment of the present application, the protein of interest comprises a functional protein.
[0028] In an embodiment of the present application, the functional protein comprises a functional protein for enhancing a target trait of a transgenic plant and / or a functional protein produced by using a plant as a bioreactor.
[0029] In an embodiment of the present application, the functional protein comprises CRY1a.105 protein, CRY2Ab protein, Bar protein, CP4 protein, new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO, and / or fish immune protein OprF.
[0030] In an embodiment of the present application, when the target protein is the insect-resistant protein CRY1a.105, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 47;
[0031] When the target protein is the insect-resistant protein CRY2Ab, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 49;
[0032] When the target protein is the herbicide-resistant protein Bar, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 51;
[0033] When the target protein is the herbicide-resistant protein CP4, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 53;
[0034] When the target protein is the new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 55;
[0035] When the target protein is the new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 57;
[0036] When the target protein is the influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 59;
[0037] When the target protein is the giant panda obesity-related protein AmFTO, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 61;
[0038] When the target protein is the red source chicken obesity-related protein GgFTO, the nucleotide sequence of the nucleic acid molecule is as shown in SEQ ID NO. 63;
[0039] When the target protein is fish immune protein OprF, the nucleotide sequence of the nucleic acid molecule is shown as SEQ ID NO. 70.
[0040] The present application also provides a recombinant expression system for expressing a target protein in a plant, which expresses the above-mentioned nucleic acid molecule.
[0041] In an embodiment of the present application, the recombinant expression system comprises a recombinant expression vector or a host cell; the recombinant expression vector carries the above-mentioned nucleic acid molecule; and the host cell is transformed with the recombinant expression vector carrying the above-mentioned nucleic acid molecule.
[0042] In an embodiment of the present application, the recombinant expression vector comprises a plant cell expression vector carrying the above-mentioned nucleic acid molecule.
[0043] In an embodiment of the present application, the plant cell expression vector comprises a plant expression vector pCambia1300 and / or a pCambia1300 derived vector.
[0044] In an embodiment of the present application, the host cell is Agrobacterium transformed with the recombinant expression vector carrying the above-mentioned nucleic acid molecule.
[0045] In an embodiment of the present application, the plant comprises a dicotyledonous plant and / or a monocotyledonous plant.
[0046] In an embodiment of the present application, the dicotyledonous plant comprises Arabidopsis thaliana, tobacco, soybean, cotton, lettuce, poplar, alfalfa, rape and / or peanut.
[0047] In an embodiment of the present application, the monocotyledonous plant comprises corn, rice, sorghum, wheat, sugarcane and / or bamboo.
[0048] In an embodiment of the present application, the target protein comprises a functional protein.
[0049] In an embodiment of the present application, the functional protein comprises a functional protein for enhancing the target traits of a transgenic plant and / or a functional protein produced by using a plant as a biological reactor.
[0050] In an embodiment of the present application, the functional protein comprises CRY1a.105 protein, CRY2Ab protein, Bar protein, CP4 protein, new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO, and / or fish immune protein OprF.
[0051] The present application also provides a transgenic plant or a transformed plant cell, which expresses the above-mentioned nucleic acid molecule.
[0052] In an embodiment of the present application, the method for preparing the transgenic plant comprises: using transgenic technology to transfer the above-mentioned nucleic acid molecule into a plant body to express the above-mentioned nucleic acid molecule in the plant, thereby obtaining the transgenic plant.
[0053] In an embodiment of the present application, the method for preparing the transformed plant cell comprises: using transgenic technology to transfer the above-mentioned nucleic acid molecule into a plant cell to express the above-mentioned nucleic acid molecule in the plant cell, thereby obtaining the transformed plant cell.
[0054] In an embodiment of the present application, the transgenic technology comprises microorganism mediation, microinjection, and / or gene gun bombardment.
[0055] In an embodiment of the present application, when the transgenic technology is microorganism mediation, the method for preparing the transgenic plant comprises: using the above-mentioned recombinant expression system to infect a plant body to express the above-mentioned nucleic acid molecule in the plant, thereby obtaining the transgenic plant.
[0056] In an embodiment of the present application, when the transgenic technology is microorganism mediation, the method for preparing the transformed plant cell comprises: using the above-mentioned recombinant expression system to infect a plant cell to express the above-mentioned nucleic acid molecule in the plant cell, thereby obtaining the transformed plant cell.
[0057] In an embodiment of the present application, the recombinant expression system is a host cell; the host cell is Agrobacterium transformed with a recombinant expression vector carrying the above-mentioned nucleic acid molecule.
[0058] In an embodiment of the present application, the recombinant expression system comprises a recombinant expression vector; the recombinant expression vector carries the above-mentioned nucleic acid molecule.
[0059] In an embodiment of the present application, the recombinant expression vector comprises a plant cell expression vector carrying the above-mentioned nucleic acid molecule.
[0060] In an embodiment of the present application, the plant cell expression vector comprises plant expression vector pCambia 1300 and / or pCambia 1300 derived vector.
[0061] In an embodiment of the present application, the plant comprises dicotyledonous plant and / or monocotyledonous plant.
[0062] In an embodiment of the present application, the dicotyledonous plant comprises Arabidopsis thaliana, tobacco, soybean, cotton, lettuce, poplar, alfalfa, rape and / or peanut.
[0063] In an embodiment of the present application, the monocotyledonous plant comprises corn, rice, sorghum, wheat, sugarcane and / or bamboo.
[0064] In an embodiment of the present application, the functional protein comprises functional protein for enhancing target traits of transgenic plants and / or functional protein produced by using plants as bioreactors.
[0065] In an embodiment of the present application, the functional protein comprises CRY1a.105 protein, CRY2Ab protein, Bar protein, CP4 protein, new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO and / or fish immune protein OprF.
[0066] The present application also provides a method for preparing a transgenic plant or a transformed plant cell, which comprises expressing the above-mentioned nucleic acid molecule in a plant to obtain a transgenic plant, or expressing the above-mentioned nucleic acid molecule in a plant cell to obtain a transformed plant cell.
[0067] In an embodiment of the present application, the method for preparing the transgenic plant comprises using transgenic technology to transfer the above-mentioned nucleic acid molecule into a plant to express the above-mentioned nucleic acid molecule in the plant to obtain a transgenic plant.
[0068] In an embodiment of the present application, the method for preparing the transgenic plant comprises using transgenic technology to transfer the above-mentioned nucleic acid molecule into a plant to express the above-mentioned nucleic acid molecule in the plant to obtain a transgenic plant.
[0069] In an embodiment of the present application, the method for preparing the transformed plant cell comprises using transgenic technology to transfer the above-mentioned nucleic acid molecule into a plant cell to express the above-mentioned nucleic acid molecule in the plant cell to obtain a transformed plant cell.
[0070] In an embodiment of the present application, the transgenic technique comprises microorganism mediation, microinjection and / or biolistic bombardment.
[0071] In an embodiment of the present application, when the transgenic technique is microorganism mediation, the method for preparing the transgenic plant comprises: infecting a plant body with the above-mentioned recombinant expression system to express the above-mentioned nucleic acid molecule in the plant, thereby obtaining the transgenic plant.
[0072] In an embodiment of the present application, when the transgenic technique is microorganism mediation, the method for preparing the transformed plant cell comprises: infecting a plant cell with the above-mentioned recombinant expression system to express the above-mentioned nucleic acid molecule in the plant cell, thereby obtaining the transformed plant cell.
[0073] In an embodiment of the present application, the recombinant expression system is a host cell; the host cell is Agrobacterium transformed with a recombinant expression vector carrying the above-mentioned nucleic acid molecule.
[0074] In an embodiment of the present application, the recombinant expression system comprises a recombinant expression vector; the recombinant expression vector carries the above-mentioned nucleic acid molecule.
[0075] In an embodiment of the present application, the recombinant expression vector comprises a plant cell expression vector carrying the above-mentioned nucleic acid molecule.
[0076] In an embodiment of the present application, the plant cell expression vector comprises plant expression vector pCambia 1300 and / or pCambia 1300 derived vector.
[0077] In an embodiment of the present application, the plant comprises dicotyledonous plant and / or monocotyledonous plant.
[0078] In an embodiment of the present application, the dicotyledonous plant comprises Arabidopsis thaliana, tobacco, soybean, cotton, lettuce, poplar, alfalfa, rape and / or peanut.
[0079] In an embodiment of the present application, the monocotyledonous plant comprises corn, rice, sorghum, wheat, sugarcane and / or bamboo.
[0080] In an embodiment of the present application, the functional protein comprises functional protein for enhancing the target traits of the transgenic plant and / or functional protein produced by using the plant as a biological reactor.
[0081] In an embodiment of the present application, the functional protein comprises functional protein for enhancing the target traits of the transgenic plant and / or functional protein produced by using the plant as a biological reactor.
[0082] In an embodiment of the present application, the functional protein comprises CRY1a.105 protein, CRY2Ab protein, Bar protein, CP4 protein, new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO, and / or fish immune protein OprF.
[0083] The present application also provides the use of the above-mentioned codon optimization method, the above-mentioned nucleic acid molecule, the above-mentioned recombinant expression system or the above-mentioned method in the preparation of a transgenic plant or the transformation of a plant cell.
[0084] The technical solution of the present application has the following advantages:
[0085] The present application provides a codon optimization method for improving the protein expression amount of a target gene in a plant, which comprises: replacing the third base of the codon of the target protein coding sequence with C or G to obtain an optimized target protein coding sequence. The existing concept believes that each species has a corresponding optimal codon combination, therefore, the existing codon optimization method is to replace the rare codon with the optimal codon with high frequency according to the codon usage frequency of each species. Contrary to the existing concept, the present application provides a codon optimization method for improving the protein expression amount of a target protein in a plant, which has effectiveness and universality, and only needs to replace the third base of the codon of the target protein coding sequence with C or G. Studies have shown that after the codon optimization by the method, the protein expression amount of the target gene in a plant protein transient expression system and a transgenic plant is increased by several times to several tens of times, and the transcription level of the target gene is increased by several tens to several hundreds of times, therefore, the method can effectively improve the protein expression amount of the target gene in a plant. In the use of a plant as a biological reactor for the production of a functional protein, the method can increase the yield of the target protein, reduce the purification cost of the target protein, and increase the edible / forage effect of the plant; in the preparation of a transgenic plant, the method can increase the expression amount of the target gene, thereby enhancing the target traits of the transgenic plant and reducing the screening cost of a high-expression transgenic plant. BRIEF DESCRIPTION OF DRAWINGS
[0086] Figure 1: Plasmid map of plant expression vector pHEQ22.
[0087] Figure 2: Effects of A / G / C / T stop codon on protein expression in tobacco leaves. In Figure 2, A: Effects of A / G / C / T stop codon on GFP protein expression in tobacco leaves; B: Effects of A / G / C / T stop codon on RFP protein expression in tobacco leaves; C: Effects of A / G / C / T stop codon on nanoLuc protein expression in tobacco leaves.
[0088] Figure 3: Effects of degenerate bases C and G on protein expression in tobacco leaves.
[0089] Figure 4: Effects of codon optimization ratio on protein expression in tobacco leaves. In Figure 4, A: Effects of codon optimization ratio on GFP protein expression in tobacco leaves; B: Effects of codon optimization ratio on nanoLuc protein expression in tobacco leaves; C: Effects of codon optimization ratio on random protein (PR) expression in tobacco leaves.
[0090] Figure 5: Effects of codon optimization on expression of different functional proteins in tobacco leaves.
[0091] Figure 6: Effects of A / G / C / T stop codon on protein expression in lettuce leaves.
[0092] Figure 7: Plasmid map of plant expression vector pQH-GFP.
[0093] Figure 8: Effects of A / G / C / T stop codon on protein expression in Arabidopsis.
[0094] Figure 9: Infection of poplar seedlings.
[0095] Figure 10: Effects of codon optimization on protein expression in cotton.
[0096] Figure 11: Infection of soybean embryo.
[0097] Figure 12: Effects of codon optimization on protein expression in rice.
[0098] Figure 13: Effects of codon optimization on protein expression in maize. DETAILED DESCRIPTION
[0099] The following examples are provided to better enable those skilled in the art to further understand the application, and are not intended to limit the content and protection scope of the application. Any product that is the same or similar to the present application, which is obtained by the disclosure of the present application or by combining the present application with other prior art features, falls within the protection scope of the present application.
[0100] The specific experimental steps or conditions are not indicated in the following examples, and can be performed according to the conventional experimental steps or conditions described in the literature in the art. The reagents or instruments used are not indicated by the manufacturer, and are conventional reagent products that can be obtained commercially.
[0101] Experimental Example 1: Influence of A / G / C / T stop codon on protein expression in tobacco leaves
[0102] Experiment 1: Taking GFP as the target protein
[0103] The target gene encoding GFP was optimized according to the optimal codons of tobacco, and was denoted as GFP-Nb. The optimized target protein coding sequence was SEQ ID NO: 1.
[0104] Based on GFP-Nb, the codon degenerate base at the 3rd position was optimized to A, T, G or C according to the method described in Table 1, and was denoted as GFP-A, GFP-T, GFP-G and GFP-C. The optimized target protein coding sequences were SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4 and SEQ ID NO: 5.
[0105] The optimized target protein coding sequences were synthesized, and each of the optimized target protein coding sequences was constructed into the BamHI / StuI position of the plant expression vector pHEQ22 (SEQ ID NO: 66) to replace the original EGFP fragment of the plant expression vector pHEQ22 by seamless cloning, to obtain different recombinant pHEQ22 vectors (the plasmid map of the plant expression vector pHEQ22 is shown in FIG. 1, the promoter is 35S, and the terminator is 35S-ter in series with NOS-ter); different recombinant Agrobacterium tumefaciens were obtained by transforming the different recombinant pHEQ22 vectors into Agrobacterium tumefaciens C58C1 competent cells by electroporation (electroporation method is described in the literature: den Dulk-Ras, A., & Hooykaas, P.J. (1995). Electroporation of Agrobacterium tumefaciens. Methods in molecular biology (Clifton, N.J.), 55, 63-72.); Nicotiana benthamiana was selected as the infection object, and 1 mL syringe was used to inject the recombinant Agrobacterium tumefaciens suspension with a living bacteria concentration of OD 600 0.2 into the tobacco leaves from the lower epidermis to infect the tobacco; wherein the formulation of the recombinant Agrobacterium tumefaciens suspension was: recombinant Agrobacterium tumefaciens with OD 600 0.2, 10 mM MgCl2 and 10 mM MES buffer (pH 5.6).
[0106] On the third day after injection, 30 mg of different recombinant Agrobacterium-infected leaves were weighed with a balance, the sample was ground, 800 μL of protein non-denaturing extraction buffer was added for protein extraction, and the expression amount of fluorescent protein GFP was detected by TECAN spark enzyme label instrument. The specific detection parameters are: excitation wavelength 485±20 nm, emission wavelength 535±20 nm, and gain value is set to 54; wherein the formula of the protein non-denaturing extraction buffer is: 150 mM NaCl, 1 mM EDTA, 1% (v / v) Triton X-100, 2% (v / v) glycerol, 1×protease inhibitor cocktail (purchased from Roche, model #04693159001) and 50 mM Tris-HCl buffer (pH 7.5). The detection results are shown in Figure 2A. The expression amount of the synonymous codon optimized according to the G or C mode is significantly higher than that of the A or T mode optimized gene, wherein the expression amount of the synonymous codon optimized according to the G mode is 10 times and 5 times higher than that of the A or T mode, respectively, and the expression amount of the synonymous codon optimized according to the C mode is 17 times and 8 times higher than that of the A or T mode, respectively.
[0107] Experiment two: taking RFP as the target protein
[0108] The target gene encoding RFP is optimized according to the optimal codon of tobacco, denoted as RFP-Nb, and the corresponding optimized target protein coding sequence is SEQ ID NO: 6.
[0109] On the basis of RFP-Nb, the codon is optimized according to the mode described in Table 1, and the degenerate base at the 3rd position of the codon is optimized to A, T, G or C, denoted as RFP-A, RFP-T, RFP-G and RFP-C, and the corresponding optimized target protein coding sequence is SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9 and SEQ ID NO: 10.
[0110] Synthesize the optimized target protein coding sequence, and construct each optimized target protein coding sequence into the BamHI / StuI position of the plant expression vector pHEQ22 by seamless cloning to replace the original EGFP fragment of the plant expression vector pHEQ22, to obtain different recombinant pHEQ22 vectors; transform different recombinant pHEQ22 vectors into Agrobacterium C58C1 competent cells by electroporation method, to obtain different recombinant Agrobacterium; select Nicotiana benthamiana as the infection object, and use a 1 mL syringe to inject 1 mL of different recombinant Agrobacterium with a concentration of OD600=0.5, and the injection site is the 3rd leaf of the Nicotiana benthamiana. 600Different recombinant Agrobacterium suspensions with OD 600 = 0.2 were injected into tobacco leaves from the lower epidermis to infect tobacco with different recombinant Agrobacterium; wherein the formulation of the recombinant Agrobacterium suspension was: OD
[0111] On the third day after injection, 30 mg of leaves infected with different recombinant Agrobacterium were weighed on a balance, ground, and 800 μL of protein non-denaturing extraction buffer was added for protein extraction. The expression of fluorescent protein GFP was detected by a TECAN spark enzyme marker. The specific detection parameters were: excitation wavelength 532 ± 20 nm, emission wavelength 588 ± 20 nm, and gain value set to 68. The detection results are shown in Figure 2B, and the expression of synonymous codons optimized in the G or C mode was significantly higher than that of genes optimized in the A or T mode.
[0112] Experiment three: using nanoLuc as the target protein
[0113] The target gene encoding nanoLuc was optimized according to the optimal codon of tobacco, denoted as nanoLuc-Nb, and the optimized target protein coding sequence was SEQ ID NO: 11.
[0114] Based on nanoLuc-Nb, the codon was optimized according to the mode described in Table 1, and the degenerate base at the 3rd position was optimized to A, T, G or C, denoted as nanoLuc-A, nanoLuc-T, nanoLuc-G and nanoLuc-C, and the optimized target protein coding sequences were SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14 and SEQ ID NO: 15.
[0115] The optimized target protein coding sequences were synthesized, and each optimized target protein coding sequence was constructed into the BamHI / StuI position of the plant expression vector pHEQ22 to replace the original EGFP fragment of the plant expression vector pHEQ22 by seamless cloning, obtaining different recombinant pHEQ22 vectors; different recombinant Agrobacterium were obtained by transforming Agrobacterium C58C1 competent cells with different recombinant pHEQ22 vectors by electroporation; Nicotiana benthamiana tobacco was selected as the infection object, and 1 mL of different recombinant Agrobacterium suspensions with OD 600 Different recombinant Agrobacterium suspensions with OD 600= 0.2, 10 mM MgCl2, and 10 mM MES buffer (pH 5.6).
[0116] On the 3rd day after injection, 30 mg of different recombinant Agrobacterium infected leaves were weighed on a balance, ground, and added with 800 μL of protein non-denaturing extraction buffer for protein extraction to obtain a protein extract; 10 μL of the protein extract was diluted 100-fold by adding 990 μL of the protein non-denaturing extraction buffer, mixed well, to obtain a first dilution; 9 μL of the first dilution was diluted 40-fold by adding 351 μL of the protein non-denaturing extraction buffer, mixed well, to obtain a second dilution; 2 μL of 5 mM Furimazine substrate (purchased from Jinzhida (Shanghai) Biotechnology Co., Ltd.) was added to 360 μL of the second dilution, vortexed, and timing was started from the addition of the substrate; after 3 minutes, the expression amount of the chemiluminescent protein nanoLuc was detected by a TECAN spark microplate reader. The specific detection parameters were: the number of photons emitted between wavelengths of 360-700 nm per second (Integration time). The detection results are shown in Figure 2C, and the expression amount of the synonymous codon optimized in the G or C manner was significantly higher than that of the gene optimized in the A or T manner.
[0117] From the results of Figure 2, it can be seen that the expression amount of the gene codon optimized to end with C is higher than that optimized to end with G, which may be due to the fact that there are 15 amino acids with synonymous codons ending with C, while there are only 11 amino acids with synonymous codons ending with G, i.e., the optimization of the third degenerate base to C requires more codons corresponding to the amino acids than the optimization to G. These results show that the most suitable codon of tobacco tends to use G or C ending codons.
[0118] Table 1 Codon optimization methods ending with A / T / G / C
[0119] Experimental Example 2: Effect of degenerate base G or C on protein expression amount in tobacco leaves
[0120] The results of Experimental Example 1 show that the protein expression amount of the G / C ending codon is higher than that of the A / T ending codon. However, from Table 1, it is found that the third degenerate base of alanine A, glycine G, proline P, threonine T, valine V, arginine R, leucine L, and serine S contains both G and C.
[0121] In order to identify whether there is a difference in the regulation of the third degenerate base G and C of the above amino acids on protein expression, the codons corresponding to alanine A, glycine G, and proline P were optimized to GCC, GGC, and CCC based on GFP-G, denoted as GFP-G / C--AGP, and the corresponding optimized protein coding sequence was SEQ ID NO: 16.
[0122] The codons corresponding to threonine T and valine V are optimized to ACC and GUC, denoted as GFP-G / C--TV, and the optimized coding sequence of the target protein is SEQ ID NO: 18;
[0123] The codons corresponding to leucine L and serine S are optimized to CUC and AGC, denoted as GFP-G / C--LS, and the optimized coding sequence of the target protein is SEQ ID NO: 20;
[0124] The codons corresponding to threonine T and serine S are optimized to ACC and AGC, denoted as GFP-G / C--TS, and the optimized coding sequence of the target protein is SEQ ID NO: 22;
[0125] The codons corresponding to valine V and leucine L are optimized to GUC and CUC, denoted as GFP-G / C--VL, and the optimized coding sequence of the target protein is SEQ ID NO: 24;
[0126] The codons corresponding to alanine A and glycine G are optimized to GCC and GGC, denoted as GFP-G / C--AG, and the optimized coding sequence of the target protein is SEQ ID NO: 26;
[0127] The codons corresponding to proline P and arginine R are optimized to CCC and CGC, denoted as GFP-G / C--PR, and the optimized coding sequence of the target protein is SEQ ID NO: 28.
[0128] Meanwhile, on the basis of GFP-C, the codons corresponding to alanine A, glycine G, and proline P are optimized to GCG, GGG, and CCG, denoted as GFP-C / G--AGP, and the optimized coding sequence of the target protein is SEQ ID NO: 17;
[0129] The codons corresponding to threonine T and valine V are optimized to ACG and GUG, denoted as GFP-C / G--TV, and the optimized coding sequence of the target protein is SEQ ID NO: 19;
[0130] The codons corresponding to leucine L and serine S are optimized to CUG and UCG, denoted as GFP-C / G--LS, and the optimized coding sequence of the target protein is SEQ ID NO: 21;
[0131] The codons corresponding to threonine T and serine S are optimized to ACG and UCG, denoted as GFP-C / G--TS, and the optimized coding sequence of the target protein is SEQ ID NO: 23;
[0132] The codons corresponding to valine V and leucine L are optimized to GUG and CUG, denoted as GFP-C / G--VL, and the corresponding optimized coding sequence of the target protein is SEQ ID NO: 25.
[0133] The codons corresponding to alanine A and glycine G are optimized to GCG and GGG, denoted as GFP-C / G--AG, and the corresponding optimized coding sequence of the target protein is SEQ ID NO: 27.
[0134] The codons corresponding to proline P and arginine R are optimized to CCG and AGG, denoted as GFP-C / G--PR, and the corresponding optimized coding sequence of the target protein is SEQ ID NO: 29.
[0135] After synthesizing the optimized coding sequence of the target protein, the expression amount of the fluorescent protein GFP was detected according to the method of Experiment 1 in Experimental Example 1. The detection results are shown in FIG. 3. By comparing the relative expression amounts of GFP-G, GFP-C, GFP-G / C--AGP, GFP-C / G--AGP, GFP-G / C--AG, GFP-C / G--AG, GFP-G / C--PR and GFP-C / G--PR proteins, it was found that the detection of alanine A, glycine G, proline P and arginine R with G or C ending codon had almost the same effect, and all were the optimal codons. By comparing the relative expression amounts of GFP-G, GFP-C, GFP-G / C--TV, GFP-C / G--TV, GFP-G / C--TS, GFP-C / G--TS, GFP-G / C--VL and GFP-C / G--VL proteins, it was found that the detection of valine V, serine S and leucine L with G or C ending codon had almost the same effect, and all were the optimal codons. However, the protein expression amount increased when the threonine T codon ACG was changed to ACC, and vice versa, indicating that the optimal codon of threonine T was ACC.
[0136] Experimental Example 3: Effect of codon optimization ratio on protein expression amount in tobacco leaves
[0137] Experiment 1: Taking GFP as the target protein
[0138] In order to study the relationship between the codon optimization ratio and the protein expression amount, the GFP-Nb was gradiently optimized, and was recorded as GFP-1, GFP-2, GFP-3, GFP-4, GFP-5, and the corresponding optimized protein coding sequence was SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, and SEQ ID NO: 34. The GC% of GFP-Nb, GFP-1, GFP-2, GFP-3, GFP-4, and GFP-5 was 32%, 40%, 45%, 50%, 55%, and 62%, respectively, and the GC3% was 11%, 35%, 50%, 65%, 77%, and 100%, respectively.
[0139] After the synthesis of the optimized protein coding sequence, the expression amount of the fluorescent protein GFP was detected by referring to the method of experiment one in the experimental example 1. The detection result is shown in A of FIG. 4. With the increase of the optimization ratio, the protein expression amount gradually increased. When the GC3 of the gene codon optimization was greater than 50%, the protein expression amount was 4 times, 5 times, and 9 times of that of GFP-Nb, respectively.
[0140] Experiment two: taking nanoLuc as the target protein
[0141] In order to further verify the result of experiment one, the nanoLuc-Nb was gradiently optimized, and was recorded as nanoLuc-1, nanoLuc-2, nanoLuc-3, nanoLuc-4, and nanoLuc-5, and the corresponding optimized protein coding sequence was SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, and SEQ ID NO: 39. The GC% of nanoLuc-Nb, nanoLuc-1, nanoLuc-2, nanoLuc-3, nanoLuc-4, and nanoLuc-5 was 32%, 40%, 45%, 50%, 55%, and 64%, respectively, and the GC3 was 8%, 30%, 43%, 56%, 72%, and 100%, respectively.
[0142] After the synthesis of the optimized protein coding sequence, the expression amount of the chemiluminescent protein nanoLuc was detected by referring to the method of experiment three in the experimental example 1. The detection result is shown in B of FIG. 4. Consistent with GFP, with the increase of the optimization ratio, the nanoLuc protein expression amount gradually increased. When the GC3 of the gene codon optimization was greater than 50%, the protein expression amount was 3 times, 4 times, and 5 times of that of nanoLuc-Nb, respectively.
[0143] Experiment three: taking a random sequence protein as the target protein
[0144] In order to further verify whether the codon optimization method proposed in the present application is applicable to all genes, three random protein sequences of 200 amino acids were designed by computer. NCBI blast found that there was no correlation with existing proteins. The random protein sequences were named PR1 (Protein Random 1), PR2 and PR3.
[0145] PR1, PR2 and PR3 were reverse translated into protein coding sequences according to the optimal codons of tobacco, and were denoted as PR1-Nb, PR2-Nb and PR3-Nb. The corresponding protein coding sequences were SEQ ID NO: 40, SEQ ID NO: 42 and SEQ ID NO: 44, respectively.
[0146] PR1-Nb, PR2-Nb and PR3-Nb were further optimized, and their GC3 were all optimized to 100%, and were denoted as PR1-GC, PR2-GC and PR3-GC. The corresponding optimized protein coding sequences were SEQ ID NO: 41, SEQ ID NO: 43 and SEQ ID NO: 45, respectively.
[0147] After synthesizing the protein coding sequences and the optimized protein coding sequences, each protein coding sequence and each optimized protein coding sequence was constructed into the SalI / BamHI position of the plant expression vector pHEQ22 by seamless cloning. The expression amount of the target protein was detected by the method of experiment one in experimental example 1. The detection results are shown in C of FIG. 4. After the PR codons were completely optimized to G / C ending codons, the protein expression amount was 8 times, 7 times and 9 times that before optimization, respectively.
[0148] According to the results in FIG. 4, the optimal codon of dicotyledonous tobacco ends with G / C, not A / T, and the higher the codon optimization ratio, the higher the protein expression amount. Importantly, the codon optimization method proposed in the present application is not only applicable to reporter genes, but also may be applicable to all genes.
[0149] Experimental Example 4: Effect of codon optimization on expression amount of different functional proteins in tobacco leaves
[0150] In order to further verify that the codon optimization method proposed in the present application is suitable for common transgenic plant functional genes, the following commonly used transgenic proteins in plants were selected: insect-resistant protein CRY1a.105, insect-resistant protein CRY2Ab, herbicide-resistant protein Bar, herbicide-resistant protein CP4, new coronavirus receptor protein ACE2 (Angiotensin-Converting Enzyme 2), new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA (Influenza Virus Hemagglutinin), giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO, and fish immune protein OprF, and the corresponding protein coding sequences are SEQ ID NO: 46, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 60, SEQ ID NO: 62, and SEQ ID NO: 69, respectively.
[0151] The above protein coding sequences were codon-optimized to have a GC3 of 100%, and are denoted as CRY1a.105-GC, CRY2Ab-GC, Bar-GC, CP4-GC, ACE2-GC, SARS-Cov-RBD-GC, HA-GC, AmFTO-GC, GgFTO-GC, and OprF-GC, respectively. The corresponding optimized protein coding sequences are SEQ ID NO: 47, SEQ ID NO: 49, SEQ ID NO: 51, SEQ ID NO: 53, SEQ ID NO: 55, SEQ ID NO: 57, SEQ ID NO: 59, SEQ ID NO: 61, SEQ ID NO: 63, and SEQ ID NO: 70, respectively.
[0152] After the synthesis of the optimized protein coding sequence, each protein coding sequence and each optimized protein coding sequence were constructed into the Sal I / Bam HI position of the plant expression vector pHEQ22 by seamless cloning and fused expression with the eGFP contained in the vector. The expression amount of the protein was detected according to the method of Experiment 1 in Experimental Example 1. The detection results are shown in Figure 5. The expression amount of these functional gene proteins was increased by at least 1 fold, which will help to further improve the expression level of the resistance genes of the transgenic plants and improve the herbicide resistance and pest resistance and other performances. These results further show that the codon optimization method proposed in the present application is suitable for common functional genes of transgenic plants. The expression amount of the functional genes is not as obvious as that of the reporter gene, which may be due to the fusion expression of the functional genes and GFP, each pair of genes containing the same GFP sequence, thereby weakening the influence of codon optimization on the expression of the functional gene proteins.
[0153] Experimental Example 5: Influence of A / G / C / T ending codon on protein expression amount in lettuce
[0154] In order to study whether the optimal codon of the dicotyledon lettuce is also ended with G / C, the Agrobacterium containing GFP-Nb and GFP-5 obtained in Experiment 1 of Experimental Example 3 was used, and the tender leaves of the 3-week-old lettuce (glass lettuce) were selected as the infection object. The recombinant Agrobacterium suspension with a living bacteria concentration of OD 600 The recombinant Agrobacterium was injected into the lettuce leaves from the lower epidermis of the leaves to infect the lettuce, wherein the formula of the recombinant Agrobacterium suspension was: recombinant Agrobacterium with OD 600 = 0.2, 10 mM MgCl2 and 10 mM MES buffer (pH 5.6).
[0155] On the 3rd day after injection, the expression of GFP was observed by hand-held fluorescence. The observation results are shown in Figure 6. The A / T ending optimized GFP-Nb was almost not observed to express, and the GFP-5 optimized to 100% GC3 was observed to have obvious fluorescence signal, indicating that, like tobacco, the optimal codon in lettuce is also ended with G / C.
[0156] Experimental Example 6: Influence of A / G / C / T ending codon on protein expression amount in Arabidopsis thaliana
[0157] In order to study whether the optimal codon of dicotyledon Arabidopsis thaliana is also ended with G / C, the GFP-A, GFP-T, GFP-G and GFP-C obtained in the first experiment of the experimental example 1 were used to construct the coding sequences of the optimized target proteins into the BamHI / StuI position of the plant expression vector pQH-GFP (SEQ ID NO: 67) to replace the original GFP fragment of the plant expression vector pQH-GFP by seamless cloning, so as to obtain different recombinant pQH-GFP vectors (the plasmid map of the plant expression vector pQH-GFP is shown in FIG. 7, the promoter is 35S, the terminator is 35S-ter in series with NOS-ter); the different recombinant Agrobacterium were obtained by transforming the different recombinant pQH-GFP vectors into the Agrobacterium C58C1 competent cells by electroporation; the Arabidopsis thaliana (Columbia) was selected as the infection object, and the different recombinant Agrobacterium were used to infect the Arabidopsis thaliana by using the dip method (see the literature: Clough S.J., Bent A.F. (1998) Floral dip: a simplified method for Agrobacterium-mediated transformation of Arabidopsis thaliana. Plant J 16: 735-743.).
[0158] The T1 generation seeds were harvested, and the cotyledon stage was used to observe the number of T1 generation positive plants (hygromycin resistance) with GFP fluorescence by using a handheld fluorescence observation. The detection results are shown in Table 2, and no single plant with GFP fluorescence signal was found in 41 GFP-A and 96 GFP-T independently transformed Arabidopsis thaliana T1 generation single plants, but about 30% of the single plants had obvious fluorescence signal in the GFP-G and GFP-C Arabidopsis thaliana T1 generation single plants. These results show that the G or C ending codon is the optimal codon of Arabidopsis thaliana, and after optimization of the codon to end with G or C, the protein expression level can be significantly improved, and the screening efficiency of high expression transgenic single plants can be improved.
[0159] In order to further analyze the expression level of the GFP-A, GFP-T, GFP-G and GFP-C transgenic Arabidopsis thaliana, 8 randomly selected Arabidopsis thaliana infected by each recombinant Agrobacterium were used for qPCR analysis by using a fluorescence PCR kit (Takara, RR820A) at 3 days after injection. The qPCR results are shown in FIG. 8, and the transcription levels of all the detection single plants of GFP-A and GFP-T are very low, and the transcription levels of GFP-G and GFP-C are several times or even several hundred times higher than those of GFP-A or GFP-T.
[0160] The results of Table 2 and Figure 8 show that the optimal codon of Arabidopsis is also ended with G or C. After the codon optimization, the transcription level of the target gene is increased by several times or even several hundred times, and it is easier to screen high-expression transgenic materials.
[0161] Table 2: Fluorescence rate of T1 generation transgenic materials of Arabidopsis
[0162] Experimental Example 7: Influence of A / G / C / T ending codon on protein expression amount in poplar
[0163] In order to study whether the optimal codon of poplar is also ended with G / C, ZsGreen with high fluorescence intensity was selected as the target protein, and the protein coding sequence of the target protein was codon optimized, and the 3rd position of the codon was optimized to end with G / C or A / T, denoted as ZsGreen-G / C and ZsGreen-A / T, and the corresponding optimized target protein coding sequences were SEQ ID NO:64 and SEQ ID NO:65.
[0164] After synthesizing the optimized target protein coding sequence, each optimized target protein coding sequence was constructed into the BamHI / StuI position of the plant expression vector pQH-GFP to replace the original GFP fragment of the plant expression vector pQH-GFP by seamless cloning, and different recombinant pQH-GFP vectors were obtained; different recombinant Agrobacterium was obtained by transforming different recombinant pQH-GFP vectors into Agrobacterium C58C1 competent cells by electroporation method; poplar seedlings (Populus alba x Populus berolinensis) were selected as the infection object, and the leaf disc method (see literature: Bruegmann T., Polak O., Deecke K., Nietsch J., Fladung M. (2019) Poplar Transformation. Methods Mol Biol 1864: 165-177.) was used to infect the tender leaves of poplar seedlings with different recombinant Agrobacterium.
[0165] After 1 month of infection, the fluorescence of T1 generation seedlings was observed by hand-held fluorescence. As shown in Figure 9, no fluorescence was observed in ZsGreen-A / T infected seedlings, and obvious fluorescence was observed in ZsGreen-G / C infected seedlings, indicating that the optimal codon of dicotyledonous plants poplar is also ended with G or C, and the protein expression amount is significantly improved after optimization.
[0166] Experimental Example 8: Influence of codon optimization on protein expression amount in cotton
[0167] Using the agrobacterium containing GFP-Nb and GFP-5 obtained in Experiment 1 of Experimental Example 3, cotton seedlings (Zhongmiansuo 29) grown for 2 weeks were selected as the infection object, and 1 mL syringe was used to inject the recombinant agrobacterium suspension with OD 600 The recombinant agrobacterium was injected into the young leaves of the cotton seedlings from the lower epidermis of the leaves to infect the young leaves of the cotton, wherein the formulation of the recombinant agrobacterium suspension was as follows: recombinant agrobacterium with OD 600 = 0.2, 10 mM MgCl2, and 10 mM MES buffer (pH 5.6).
[0168] On the third day after injection, the expression of GFP was observed by hand-held fluorescence. The observation results are shown in FIG. 10. Whether or not the transcriptional suppressor p19 (the transcriptional suppressor p19 can inhibit post-transcriptional silencing of plants and promote expression of exogenous genes, see the literature: Voinnet O., Rivas S., Mestre P., Baulcombe D. (2003) An enhanced transient expression system in plants based on suppression of gene silencing by the p19 protein of tomato bushy stunt virus. Plant J 33:949-956.) was co-expressed, the expression amount of the protein after codon optimization in the G / C mode was significantly higher than that after optimization of the A / T termination codon, indicating that in the dicotyledonous plant cotton, the most suitable codon also ends with G or C, and the expression amount of the protein after optimization is significantly improved.
[0169] Experimental Example 9: Effect of codon optimization on expression amount of protein in soybean
[0170] The hygromycin resistance gene in pQH-GFP was replaced by the glufosinate-ammonium resistance bar gene by enzyme digestion and seamless cloning method, and the vector was named pQB-GFP (SEQ ID NO: 68).
[0171] The ZsGreen-G / C and ZsGreen-A / T obtained in Experimental Example 7 were connected to the BamHI / StuI position of the vector pQB-GFP by seamless cloning to replace the original GFP fragment of the plant expression vector pQB-GFP, obtaining different recombinant pQB-GFP vectors; the different recombinant pQB-GFP vectors were transformed into Agrobacterium C58C1 competent cells by electroporation, obtaining different recombinant Agrobacterium; soybean immature embryos (Tianlong No. 1) were selected as the infection object, and different recombinant Agrobacterium were used to infect soybean immature embryos through cotyledon nodes (infection method, see the literature: Ko T.S., Korban S.S., Somers D.A. (2006) Soybean (Glycine max) transformation using immature cotyledon explants. Methods Mol Biol 343: 397-405.).
[0172] After the infection was completed, the soybean seedlings were dark cultured for 3 days and recovery cultured for 7 days (dark culture and recovery culture, see the literature: Ko T.S., Korban S.S., Somers D.A. (2006) Soybean (Glycine max) transformation using immature cotyledon explants. Methods Mol Biol 343: 397-405.), and the infection was observed by hand-held fluorescence. As shown in Figure 11, the expression amount of ZsGreen-G / C protein was significantly higher than that of ZsGreen-A / T, indicating that the optimal codon of dicotyledonous plants soybean is also ended with G or C, and the expression amount of the optimized protein is significantly improved.
[0173] Experimental Example 10: Effect of codon optimization on protein expression amount in rice
[0174] The Agrobacterium containing ZsGreen-G / C and ZsGreen-A / T obtained in Experimental Example 9 was used to infect rice seedlings (Japanese early-maturing japonica rice Kitaake) as the infection object, and different recombinant Agrobacterium were used to infect the young leaves of rice seedlings (infection method, see the literature: Nishimura A. (2020) Agrobacterium Transformation in the Rice Genome. Methods Mol Biol 2072: 207-216.).
[0175] After the end of the infection, the rice seedlings were subjected to dark culture, dedifferentiation, differentiation, and rooting culture (dark culture, dedifferentiation, differentiation, and rooting culture refer to the literature: Nishimura A. (2020) Agrobacterium Transformation in the Rice Genome. Methods Mol Biol 2072: 207-216.), and the T1 generation of transgenic rice seedlings was obtained. Five independent T1 generation of transgenic rice seedlings were randomly selected, and Western Blot (Western Blot refer to the literature: Hnasko T.S., Hnasko R.M. (2015) The Western Blot. Methods Mol Biol 1318: 87-96.) was used to detect protein expression. As shown in Figure 12, after the codon optimization of the rice transgenic material ended with G / C, the protein expression was significantly improved, indicating that the optimal codon of monocotyledonous plants rice also ends with G or C, and the protein expression is significantly improved after optimization.
[0176] Experimental Example 11: Effect of codon optimization on protein expression in maize
[0177] Using the Agrobacterium containing ZsGreen-G / C and ZsGreen-A / T obtained in Experimental Example 9, maize immature embryos (maize inbred line B104) about 10 days after pollination were selected as the infection object, and different recombinant Agrobacterium were used to infect the maize immature embryos (the method for obtaining maize immature embryos and the method for infection refer to the literature: Ishida Y., Hiei Y., Komari T. (2007) Agrobacterium-mediated transformation of maize. Nat Protoc 2: 1614-1621.).
[0178] After the end of the infection, the maize immature embryos were subjected to dark culture for 5 days and recovery culture for 5 days (dark culture and recovery culture refer to the literature: Ishida Y., Hiei Y., Komari T. (2007) Agrobacterium-mediated transformation of maize. Nat Protoc 2: 1614-1621.), and the expression of exogenous gene protein in the maize immature embryos was observed by hand-held fluorescence. As shown in Figure 13, only the maize immature embryos infected with ZsGreen-G / C showed fluorescence protein expression, and no fluorescence protein expression was observed in more than 100 maize immature embryos infected with ZsGreen-A / T, indicating that the optimal codon of monocotyledonous plants maize also ends with G or C, and the protein expression is significantly improved after optimization.
[0179] Obviously, the above-mentioned embodiments are only examples for clearly illustrating the present application, but not limitation to the embodiments. Based on the above description, other different forms of changes or variations can be made by those skilled in the art. Here, it is not necessary and also impossible to enumerate all the embodiments. The changes or variations derived from the above are still within the protection scope of the present application.
Claims
1. A method for improving the protein expression amount of a target gene in a plant by codon optimization, characterized in that, The method comprises replacing the third base of the codon of the coding sequence of the target protein with C or G to obtain an optimized coding sequence of the target protein.
2. The codon optimization method of claim 1, wherein, The GC content of the optimized coding sequence of the target protein is not less than 55%.
3. The codon optimization method of claim 1 or 2, wherein, In the optimized coding sequence of the target protein, the GC content of the third base of all codons is not less than 70%.
4. The method of codon optimization according to any one of claims 1 to 3, wherein, In the method, when the codon is replaced, the third base of the codon corresponding to threonine is replaced with C.
5. The method of codon optimization according to any one of claims 1 to 4, wherein, In the optimized coding sequence of the target protein, more than 50% of the segments have undergone replacement of the third base of the codon.
6. A nucleic acid molecule for expressing a protein of interest in a plant, characterized in that, The nucleic acid molecule is optimized by the codon optimization method of any one of claims 1-5.
7. The nucleic acid molecule of claim 6, wherein, The target protein comprises a functional protein; the functional protein comprises a functional protein for enhancing the target traits of a transgenic plant and / or a functional protein produced by using a plant as a biological reactor. The functional protein comprises CRY1a.105 protein, CRY2Ab protein, Bar protein, CP4 protein, new coronavirus receptor protein ACE2, new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, influenza virus hemagglutinin protein HA, giant panda obesity-related protein AmFTO, red source chicken obesity-related protein GgFTO, and / or fish immune protein OprF. When the target protein is the insect-resistant protein CRY1a.105, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
47. When the target protein is the insect-resistant protein CRY2Ab, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
49. When the target protein is the herbicide-resistant protein Bar, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
51. When the target protein is the herbicide-resistant protein CP4, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
53. When the target protein is the new coronavirus receptor protein ACE2, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
55. When the target protein is the new coronavirus spike protein receptor binding domain protein SARS-Cov-RBD, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
57. When the target protein is the influenza virus hemagglutinin protein HA, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
59. When the target protein is the giant panda obesity-related protein AmFTO, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
61. When the target protein is the red source chicken obesity-related protein GgFTO, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
63. When the target protein is the fish immune protein OprF, the nucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO.
70.
8. A recombinant expression system for expressing a protein of interest in a plant, characterized in that, The recombinant expression system expresses the nucleic acid molecule of claim 6 or 7.
9. A transgenic plant or transformed plant cell, characterized in that, The transgenic plant or transformed plant cell expresses the nucleic acid molecule of claim 6 or 7.
10. Use of the codon optimization method according to any one of claims 1 to 5 or the nucleic acid molecule according to claim 6 or 7 or the recombinant expression system according to claim 8 for the production of a transgenic plant or a transformed plant cell.
Citation Information
Patent Citations
Codon optimization method for increasing expression level of target gene in host body
CN110423769A
Codon optimization and ribosome profiling for increasing transgene expression in chloroplasts of higher plants
CN110603046A
2019-nCoV double-target antibody detection microsphere complex combination, preparation method, kit and use method of kit
CN111896734A
Codon optimization method for improving protein expression quantity of target gene in plant
CN119220583A
Method and apparatus for transferring or delivering ai models during handover in wireless communication systems
KR1020250024347A