Novel codon optimization method based on precise transcriptome sequencing and analysis thereof

By obtaining transcriptome data at specific periods of the host and tissue sites, and using weighting method to optimize the frequency of codon use, the problem of insufficient gene expression in the heterologous expression system was solved, and the efficient expression of the target gene was achieved.

CN120544685APending Publication Date: 2025-08-26INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510580851.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the differences in gene expression levels in different development stages, tissues and organs in heterologous expression systems, resulting in insufficient expression of the target gene.

Method used

By obtaining transcriptome data at specific periods of the host and tissue sites, the weighting method is used to calculate the relative synonymous codon usage frequency (RSCU), and replacing rare codons as dominant synonyms without affecting the secondary structure, optimizing the codon usage frequency of the target gene.

Benefits of technology

The expression of the target gene in the host was significantly increased, achieving more efficient gene expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544685A_ABST
    Figure CN120544685A_ABST
Patent Text Reader

Abstract

The invention discloses a novel codon optimization method based on precise transcriptome sequencing and analysis thereof, and belongs to the technical field of bioinformatics. In order to improve the expression quantity of a target gene, corresponding accurate transcriptome data is obtained according to the development period of transgenic expression and parts such as tissues and organs, and efficient expression genes are screened to serve as a host genome expression library; the expression abundance of the gene is comprehensively considered when the use frequency of the codon of the host genome is determined, and the gene optimized by the novel codon optimization method increases the stability of the insecticidal protein gene RNA and improves the expression level of the protein. The invention provides better reference for gene modification and transformation, and provides a novel codon optimization method for high-efficiency expression of transgenes in plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and specifically relates to a novel codon optimization method based on precise transcriptome sequencing and analysis. Background Art

[0002] In organisms, 61 codons encode 20 amino acids. Each amino acid can be expressed by one or more different codons, a phenomenon known as codon degeneracy (Li Ying et al., 2016). For example, arginine, serine, and leucine are each encoded by six codons. Different codons encoding the same amino acid are called synonymous codons, and their frequencies of use during translation vary. This phenomenon is known as codon bias (Li Ying et al., 2016). Genes from different species exhibit distinct codon usage preferences. Even within the same species, genes with different functions, or within the same gene at different developmental stages, tissues, and organs, exhibit significant differences in codon usage frequency, which is positively correlated with intracellular aminoacyl-tRNA levels (Dana, A. & Tuller, T., 2014). Codon bias has profound and complex implications for the translation and efficient expression of heterologous genes across species (Gustafsson C, et al., 2004). For example, if Cry proteins from the prokaryotic bacillus thuringiensis and other heterologous proteins are directly transferred into plants without modification, their expression levels will be significantly reduced (Perlak FJ, et al., 1991). Therefore, codon optimization is a very effective method for efficient expression of foreign genes in heterologous expression systems (Quax and Claassens et al., 2015). Current analysis methods generally use the online software Codon Usage Database (http: / / www.kazusa.or.jp / codon / ) to obtain the codon usage frequency of the host genome. Alternatively, a suitable number of highly expressed genes in the host are selected, the coding region (CDS) sequences corresponding to each gene are extracted, and these CDS sequences are concatenated end-to-end to form a large sequence. The codon usage frequency of each gene is then counted using the online website (http: / / www.geneinfinity.org / sms / sms_codonusage.html), the relative frequency of synonymous codon usage (RSCU) is calculated, and the preference for synonymous codon usage is analyzed to prepare a codon usage table for the host genome. Then, referencing this table, optimize the rare codons in the transformed target gene (Pere Puigbò, et al., 2007). It should be noted that the host genome codon usage frequency obtained using the above method is obtained by calculating the ratio of the total number of a particular codon in the host genome (DNA) library (or in a moderate number of highly expressed genes) to the total number of codons in the genome (or in all moderately expressed genes). This calculation does not take into account the impact of gene expression abundance on codon usage frequency and cannot truly reflect codon usage frequency.However, for individual genes, their expression levels vary significantly in different developmental stages, tissues, and organs of the same species (Nakamura A., 2006). Summary of the Invention

[0003] The technical problem to be solved by the present invention is: how to increase the expression level of the target gene.

[0004] To solve this technical problem, the present invention obtains transcriptome data from specific periods and tissue locations of the host, obtains the codon usage frequency (RSCU) and codon table through transcript analysis and weighted method, and optimizes the codons in the target gene based on the obtained codon usage frequency and codon table.

[0005] The present invention provides a method for optimizing target gene codons, the method comprising:

[0006] S1) receiving host transcriptome sequencing data of the expression site of the target gene at a specific expression period of the host, and selecting n highly expressed genes based on the host transcriptome sequencing data, denoted as gene i, where i represents the i-th gene, n is a natural number greater than or equal to 1, and the value of i ranges from 1 to n;

[0007] S2) obtaining the RSCU (relative synonymous codon usage frequency) of each codon in the host;

[0008] S3) Referring to the RSCU of each codon of the host, the rare codons in the target gene are replaced with dominant synonymous codons, wherein the rare codons are codons with an RSCU of less than 0.8, and the dominant synonymous codons are codons that encode the same amino acid as the rare codons and have an RSCU greater than 1.0.

[0009] Furthermore, in the method described above, obtaining the RSCU of each codon in n highly expressed genes of the host described in S2) comprises the following steps:

[0010] s2-1) Obtaining the number N of various codons encoding amino acid p in each of the n highly expressed genes opj , where i represents the i-th gene, p represents the p-th amino acid, j represents the j-th codon encoding amino acid p, and the value of j is a natural number from 1 to n p , n p represents the number of synonymous codons encoding amino acid p, n p The value of is a natural number from 2 to 6;

[0011] s2-2) Multiply the number of the j-th codon encoding amino acid p in gene i by the expression value of gene i, and obtain the weight value W of the j-th codon encoding amino acid p in gene i according to formula 1ipj ,

[0012] Formula 1 is:

[0013] W ipj =N ipj *FPKM i ,

[0014] In formula 1, W ipj represents the weight value of the jth codon encoding amino acid p in gene i of the host, N ipj represents the number of the jth codon encoding amino acid p in gene i, FPKM i represents the expression level of gene i;

[0015] s2-3) According to formula 2, the W of all synonymous codons encoding amino acid p in the n highly expressed genes of the host is converted to ipj Add the sum and the sum is X pj ,

[0016] Formula 2 is:

[0017]

[0018] In formula 2, X pj represents the sum of the weight values ​​of the j-th codon encoding amino acid p in the host among the n highly expressed genes of the host;

[0019] s2-4) Obtain the RSCU of each codon according to formula 3,

[0020] Formula 3 is:

[0021]

[0022] In formula 3, RSCU pj represents the relative synonymous codon usage frequency of the jth codon encoding amino acid p in n highly expressed genes of the host; It represents the sum of the weight values ​​of all synonymous codons encoding amino acid p in n highly expressed genes of the host.

[0023] In the present invention, the expression level of gene i is the FPKM value of gene i obtained based on transcriptome sequencing data.

[0024] In the present invention, the purpose of optimizing the codons of the target gene is to increase the expression level of the target gene. As common knowledge in the art, the above-mentioned codon optimization needs to ensure that mRNA is normally transcribed and translated. Under this premise, the considerations of the codon optimization may include but are not limited to: codon usage (such as relative synonymous codon usage frequency [RSCU], codon adaptation index [CAI], effective number of codons [ENc] and synonymous codon usage order [SCUO]), codon pairs, tRNA usage (such as tRNA adaptation index [tAI]), GC content, ribosome binding site (RBS), hidden stop codons, motif avoidance, restriction site removal, elimination of the mRNA secondary structure (such as mRNA free energy) of the gene, elimination of cis-acting elements and optimization of the hydrophilicity index. It is well known to those skilled in the art that there are currently a variety of tools for secondary structure prediction such as DNAman, Mfold, RNAfold and RNAstructure. In some embodiments of the present invention, the minimum free energy is calculated by DNAman8.0 software; CAI, GC content and cis-acting elements are analyzed by GenScript. Types of cis-acting elements include: Splice (GGTAAG), Splice (GGTGAT), Splice (GTAAAA), Splice (GTAAGT), Splice (GTACGT), PolyA (AATAAA), PolyA (AATGAA), PolyA (AATGGA), PolyA (TATAAA), PolyA (AATAAT), Destabilizing (ATTTA), PolyT (TTTTTT), PolyA (AAAAAAA).

[0025] Furthermore, the replacement described in S3) can be complete without affecting the secondary structure. That is, all rare codons in the target gene can be replaced with dominant synonymous codons without affecting the secondary structure. The term "without affecting the secondary structure" can be specifically understood as avoiding the formation of an overly large or overly stable neck-loop structure that could affect mRNA translation.

[0026] In the present invention, the target gene may be naturally occurring, such as extracted from naturally occurring organisms. The target gene may also be artificially modified or artificially synthesized. The source of the target gene does not limit the above technical solution. In some embodiments of the present invention, the target gene is a truncated, preliminarily codon-optimized Bt protein encoding gene, and its nucleotide sequence is SEQ ID NO: 1. In the present invention, the host may be a plant. In some embodiments of the present invention, the host is rice and / or corn. It is well known to those skilled in the art that the above limitations are merely exemplary of the optional types of hosts and do not constitute limitations on the hosts in the present invention.

[0027] Furthermore, the host can be an organism that stably or transiently expresses the gene encoding the target protein, including but not limited to cells, microorganisms, plants, and animals. In some embodiments of the present invention, the host is a plant, more specifically, rice and / or corn.

[0028] Furthermore, the expression sites mentioned in S1) include but are not limited to: cells, tissues and / or organs.

[0029] In some embodiments of the present invention, the expression site may be a leaf, and more specifically, the expression site may be a flag leaf.

[0030] The expression period can be a certain stage of the host's entire life cycle. The expression period can be based on the purpose of use, and the above method can be used to achieve effective expression of the target gene in a specific period of the host. At this time, the specific period can be defined as the expression period. Exemplarily, in some embodiments of the present invention, the target gene is a Bt protein encoding gene, and the purpose of introducing this gene into rice is to improve the insect resistance of rice. Taking into account factors such as the insect resistance spectrum of Bt protein, the developmental period of pests, the developmental stage of pests invading rice, and the site of pest invasion, the transcriptome sequencing data of the sword leaf part during the filling period is selected, and the RSCU of the sword leaf part during the filling period of rice is calculated, so as to grasp the codon preference of the sword leaf part during the filling period of rice. With reference to the RSCU of the sword leaf part during the filling period of rice, the rare codons in the target gene are replaced with dominant synonymous codons.

[0031] Furthermore, the method further comprises the step of replacing the codons with 0.8≤RSCU<1 in the target gene with dominant synonymous codons.

[0032] The dominant synonymous codon is a codon that encodes the same amino acid as the codon with 0.8≤RSCU<1 and has an RSCU greater than 1.0.

[0033] Furthermore, the replacement can be partial and / or complete. The partial replacement can be to retain or partially retain the codons with 0.8≤RSCU<1 in the midstream and downstream of the target gene. The complete replacement is to replace all the codons with 0.8≤RSCU<1 in the target gene with dominant codons.

[0034] In the present invention, the principles of codon optimization include but are not limited to A1) and / or A2):

[0035] A1) All codons with a relative synonymous codon usage frequency (RSCU) less than 0.8 are replaced with dominant synonymous codons with an RSCU greater than 1, without affecting the secondary structure. If there are multiple candidates for a synonymous codon with an RSCU greater than 1, the replacement ratio of each synonymous codon is calculated based on the RSCU ratio, and the replacement ratio of the dominant codon with a higher RSCU is increased accordingly.

[0036] A2) If the gene contains a large number of codons with a ratio of 0.8≤RSCU<1, the codons with a ratio of 0.8≤RSCU<1 should also be replaced; at the same time, some codons with a ratio of 0.8≤RSCU<1 can be retained and evenly distributed in the latter part of the sequence.

[0037] Illustratively, in some embodiments of the present invention, the rare codon TCT in the target gene (RSCU is 0.56), the replaceable dominant codons include TCG (1.13), AGC (1.57), and TCC (1.88). According to the above principles, 1 of the 15 rare codons TCT in the target gene is replaced by TCG, 3 are replaced by AGC, and 11 are replaced by TCC.

[0038] The target gene is divided into three sections according to its length, namely the upper section (the first 1 / 3 of the total length), the middle section (the middle 1 / 3) and the rear section (the rear 1 / 3 of the total length). For example, in some embodiments of the present invention, the codons with 0.8≤RSCU<1 are evenly distributed in the rear section of the sequence. Figure 9 shown.

[0039] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any of the above methods.

[0040] The present invention also provides a computer program product, comprising a computer program, which implements any of the above methods when executed by a processor.

[0041] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned methods when executed by a processor.

[0042] The present invention also provides a device for optimizing target gene codons, which may include the following modules:

[0043] B1) Host transcriptome sequencing data receiving and analysis module: used to receive host transcriptome sequencing data of the expression site of the target gene to be optimized during a specific expression period of the host, and select n highly expressed genes based on the host transcriptome sequencing data, each denoted as gene i, where i represents the i-th gene, n is a natural number greater than or equal to 1, and the value of i ranges from 1 to n;

[0044] B2) Host codon RSCU acquisition module: used to obtain the RSCU of each codon in n highly expressed genes of the host, including the following submodules:

[0045] B2-1) Host Nij acquisition module: used to obtain the number value N of various codons encoding amino acid p in each of the n highly expressed genes ipj , where i represents the i-th gene, p represents the p-th amino acid, j represents the j-th codon encoding amino acid p, and the value of j is a natural number from 1 to n p , n p represents the number of synonymous codons encoding amino acid p, n p The value of is a natural number from 2 to 6;

[0046] B2-2) Host W ipj Acquisition module: used to multiply the number of the j-th codon encoding amino acid p in gene i by the expression value of gene i, and obtain the weight value W of the j-th codon encoding amino acid p in gene i according to formula 1 ipj ,

[0047] Formula 1 is:

[0048] W ipj =N ipj *FPKM i ,

[0049] In formula 1, W ipj represents the weight value of the jth codon encoding amino acid p in gene i of the host, N jpj represents the number of the jth codon encoding amino acid p in gene i, FPKM i represents the expression level of gene i;

[0050] B2-3) Host X pj Acquisition module: used to convert the W of all synonymous codons encoding amino acid p in the n highly expressed genes of the host according to formula 2 ipj Add the sum and the sum is X pj ,

[0051] Formula 2 is:

[0052]

[0053] In formula 2, X pj represents the sum of the weight values ​​of the j-th codon encoding amino acid p in the host among the n highly expressed genes of the host;

[0054] B2-4) Host RSCU acquisition module: used to obtain the RSCU of each codon according to formula 3,

[0055] Formula 3 is:

[0056]

[0057] In formula 3, RSCU pj represents the relative synonymous codon usage frequency of the jth codon encoding amino acid p in n highly expressed genes of the host; represents the sum of the weight values ​​of all synonymous codons encoding amino acid p in n highly expressed genes of the host;

[0058] B3) Target gene rare codon optimization module: used to replace rare codons in the target gene with dominant synonymous codons with reference to the RSCU of each codon in the host. The rare codons are codons with an RSCU of less than 0.8, and the dominant synonymous codons are codons that encode the same amino acid as the rare codons and have an RSCU greater than 1.0.

[0059] In the present invention, the host includes but is not limited to cells, microorganisms, plants or animals. In some embodiments of the present invention, the host is rice and / or corn.

[0060] The beneficial technical effects achieved by the present invention are as follows:

[0061] The present invention is based on the transcriptome data obtained by the host at a specific period and tissue site, and obtains the codon usage frequency (RSCU) and codon table by transcript analysis and weighted method, and optimizes the codons in the target gene based on the codon usage frequency and codon table obtained. Specifically, the present invention selects the transcriptome sequencing data of the host's specific expression period and / or specific expression site according to the target gene in the host, and performs bioinformatics analysis on the sequencing data. According to different expression gradients, highly expressed genes are selected, and highly expressed genes are selected based on this data. The relative usage frequency of synonymous codons is calculated, and the accurate optimization of the target gene codon preference is achieved on this basis. The test results of the present invention show that the expression of the target gene in the host can be significantly improved by the optimization method provided by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1Schematic diagram of the structure of transient expression vector pUC19-CryV4-Os;

[0063] Figure 2 Schematic diagram of the structure of transient expression vector pUC19-CryV4-Zm;

[0064] Figure 3 Schematic diagram of the structure of the plant expression vector pCNDL-CryV4-Os;

[0065] Figure 4 PCR detection of some T0 generation transgenic rice with CryV4-Os gene;

[0066] Figure 5 To detect mRNA and protein expression levels in transiently transformed protoplasts;

[0067] Figure 6 To detect the expression levels of transgenic rice mRNA and protein;

[0068] Figure 7 Codon analysis of the Cry1Ab gene (SEQ ID NO: 1);

[0069] Figure 8 Codon analysis of the CryV2-Os gene (SEQ ID NO: 2);

[0070] Figure 9 Codon analysis of the CryV4-Os gene (SEQ ID NO: 4);

[0071] Figure 10 This is a workflow diagram for codon optimization of the present invention. DETAILED DESCRIPTION

[0072] 1. Terms used in the present invention:

[0073] Examples of resources describing many of the molecular biology-related terms used herein can be found in Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, Genes IX, Oxford University Press: New York, 2007.

[0074] Any references cited herein, including, for example, all patents, published patent applications, and non-patent publications, are hereby incorporated by reference in their entirety.

[0075] To facilitate understanding of the present invention, several terms and abbreviations used herein are defined as follows:

[0076] "Gene" refers to a nucleic acid fragment that encodes all or part of a specific protein and includes regulatory sequences preceding (5' non-coding region) and following (3' non-coding region) the coding region.

[0077] "Plant" refers to any plant at any stage of development, in particular seed plants.

[0078] "Transformation" refers to the process of introducing exogenous nucleic acid into a host cell or organism, including Agrobacterium-mediated transformation and gene gun techniques.

[0079] "Transgenic" refers to a gene that is stably integrated into a host organism by introducing a homologous gene in the normal host organism or an exogenous artificially synthesized and modified gene through gene transfer. In the present invention, "transgenic" refers to the "target gene".

[0080] As used herein, "plant" includes an explant, plant part, seedling, plantlet or whole plant at any stage of regeneration or development.

[0081] As used herein, "plant part" can refer to any organ or intact tissue of a plant, such as meristem, bud organ / structure (e.g., leaf, stem, or node), root, flower or flower organ / structure (e.g., flower, bract, sepal, petal, stamen, carpel, anther, and ovule), seed (e.g., embryo, endosperm, and seed coat), fruit (e.g., mature ovary), propagule, or other plant tissue (e.g., vascular tissue, dermal tissue, ground tissue, etc.), or any part thereof. The plant part of the present invention can be viable, non-viable, regenerable, and / or non-regenerable. "Propaganda" can include any plant part that can grow into a whole plant.

[0082] A plant cell is a biological cell of a plant that is taken from a plant or derived from a culture obtained by culturing cells taken from a plant. As used herein, a "transgenic plant cell" refers to any plant cell transformed with a stably integrated recombinant DNA molecule, construct, expression cassette, or sequence. Transgenic plant cells can include the original transformed plant cell, a transgenic plant cell regenerated or developed from an R0 generation transgenic plant cell, a transgenic plant cell cultured from another transgenic plant cell, or a transgenic plant cell from any progeny or subsequent generation of a transformed R0 generation plant, including cells from plant seeds or embryos, or cultured plant cells, callus cells, and the like.

[0083] A "codon" is a set of three consecutive bases in the reading frame of a messenger RNA chain that determines the location of an amino acid. This is also known as a triplet code. The fact that the same amino acid can be encoded by different codons is called "codon degeneracy."

[0084] "Codon preference" refers to the fact that, due to the degeneracy of codons, each amino acid is represented by at least one codon and up to six corresponding codons. Genetic codon usage varies significantly between species and organisms. Various organisms appear to prefer certain synonymous triplet codons. Codons that are frequently used in a species are called its "preferred codons," while codons that are less frequently used are called its "rare codons." In the present invention, codons with an RSCU of less than 0.8 are defined as "rare codons."

[0085] The term "host" may refer to an organism that stably or transiently expresses a gene encoding a protein of interest, including but not limited to cells, microorganisms, plants, and animals.

[0086] The term "transcriptome sequencing data" refers to sequencing data of the transcriptome, which is the sum of all RNAs transcribed by a specific tissue or cell at a certain developmental stage or functional state, mainly including mRNA and non-coding RNA (ncRNA).

[0087] FPKM stands for Fragments Per Kilobase Million (Fragments Per Kilobase ofexon model per Million mapped fragments), which means fragments per kilobase of transcript per million mapped reads).

[0088] The RSCU concept refers to the relative probability of a particular codon among synonymous codons encoding the corresponding amino acid, eliminating the influence of amino acid composition on codon usage. If codon usage is unbiased, the RSCU value for that codon is 1. When the RSCU value of a codon is greater than 1, it indicates that the codon is relatively frequently used, and vice versa.

[0089] As used herein, "plant" includes an explant, plant part, seedling, plantlet or whole plant at any stage of regeneration or development.

[0090] As used herein, "plant part" can refer to any organ or intact tissue of a plant, such as meristem, bud organ / structure (e.g., leaf, stem, or node), root, flower or flower organ / structure (e.g., flower, bract, sepal, petal, stamen, carpel, anther, and ovule), seed (e.g., embryo, endosperm, and seed coat), fruit (e.g., mature ovary), propagule, or other plant tissue (e.g., vascular tissue, dermal tissue, ground tissue, etc.), or any part thereof. The plant part of the present invention can be viable, non-viable, regenerable, and / or non-regenerable. "Propaganda" can include any plant part that can grow into a whole plant.

[0091] 2. The technical solution provided by the present invention:

[0092] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0093] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0094] The quantitative tests in the following examples were repeated three times unless otherwise specified, and the results were averaged.

[0095] The data in the following examples were processed using GraphPad Prism statistical software. The experimental results are expressed as mean ± standard deviation and analyzed using one-way ANOVA combined with Tukey's test. P < 0.001 (***) indicates a highly significant difference.

[0096] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.

[0097] Unless otherwise specified, the materials and reagents used in the following examples can be obtained from commercial sources.

[0098] The transient expression vector pUC19 is a product of Beijing Solebow Technology Co., Ltd., with the product number P3519-20ug.

[0099] During the construction of the following examples, plasmid extraction, enzyme digestion, DNA recovery, and purification were performed according to the methods of molecular cloning (Sambrook and Russell 2001).

[0100] The present invention specifically uses a truncated Bt protein, the amino acid sequence of which is shown in GenBank: AAG16877.1 (01-OCT-2000) from position 1 to position 615.

[0101] Example 1: Establishment of a novel codon optimization method

[0102] 1. Establishment of a frequency table of codon usage for highly expressed genes in rice and maize gene banks

[0103] 1) Establishment of the RICE TOP50 and TOP500 gene banks

[0104] Precision transcriptome sequencing was performed on flag leaves of the super hybrid rice restorer line Minghui 86 (MH86) at the grain filling stage, with three biological replicates. The FPKM values ​​of each gene in the three replicate databases were calculated and averaged. Genes were sorted from highest to lowest according to their FPKM mean, and the top 50 and 500 genes, respectively, were selected as reference gene sequences for high-expression genes. Gene numbers are shown in Tables 1 and 2, respectively. Using the MSU IDs of the selected 50th and 500th highly expressed genes, the corresponding CDS sequences were extracted from the MSU7.0 database (http: / / rice.plantbiology.msu.edu / downloads_gad.shtml) to form the Rice Highly Expressed Reference Gene Sequence Libraries (RICE TOP50 and RICE TOP500).

[0105] Table 1 MSU IDs and expression levels of the top 50 highly expressed genes in rice (3 biological replicates and their means)

[0106]

[0107]

[0108] Table 2 MSU IDs and expression levels of the top 500 highly expressed genes in rice (3 biological replicates and their means)

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122] 2) Establishment of the Maize TOP50 Gene Bank

[0123] The latest version of the RNA-seq database for the representative maize genome, B73 (Zm-B73-REFERENCE-NAM-5.0), was retrieved from the Maize Genetics and Genomics Database (MaizeGDB) (https: / / qteller.maizegdb.org / rna_data_sources.php). Maize leaf RNA-seq data were downloaded and sorted by the mean FPKM of each gene. The top 50 genes were selected as the reference gene sequence library for highly expressed genes. The gene numbers are shown in Table 3. Using the IDs of the 50 selected highly expressed genes, the corresponding CDS sequences of each gene were extracted from the MaiZeGDB database (https: / / staging.maizegdb.org / ) to form the maize highly expressed reference gene sequence library (MAIZE TOP50), as shown in Table 3.

[0124] Table 3 IDs and expression levels of the top 50 highly expressed genes in maize

[0125]

[0126] 3) Determination of weighted relative codon usage (RSCU) in the TOP50 and TOP500 gene pools

[0127] The number of codons for each gene in the rice TOP50, TOP500, and maize TOP50 gene libraries was obtained using the Codon Usage website (http: / / www.geneinfinity.org / sms / sms_codonusage.html). This value was then multiplied by the gene's expression level (FPKM) to determine the weighted codon count for that codon. Because the calculation of FPKM takes into account the number of transcripts and gene length, the codon usage frequency calculated using the FPKM-weighted codon count more accurately reflects the actual frequency of codon usage.

[0128] Based on this, the workflow diagram of host codon optimization established in this application is as follows: Figure 10 As shown, the specific steps include:

[0129] S1) receiving host transcriptome sequencing data of the expression site of the target gene at a specific expression period of the host, and selecting n highly expressed genes based on the host transcriptome sequencing data, denoted as gene i, where i represents the i-th gene, n is a natural number greater than or equal to 1, and the value of i ranges from 1 to n;

[0130] S2) obtaining the RSCU (relative synonymous codon usage frequency) of each codon in n highly expressed genes of the host according to a method comprising the following steps:

[0131] s2-1) Obtaining the number N of various codons encoding amino acid p in each of the n highly expressed genes ipj , where i represents the i-th gene, p represents the p-th amino acid, j represents the j-th codon encoding amino acid p, and the value of j is a natural number from 1 to n p , n p represents the number of synonymous codons encoding amino acid p, n p The value of is a natural number from 2 to 6;

[0132] s2-2) Multiply the number of the j-th codon encoding amino acid p in gene i by the expression value of gene i, and obtain the weight value W of the j-th codon encoding amino acid p in gene i according to formula 1 ipj ,

[0133] W ipj =N ipj *FPKM i ——Formula 1,

[0134] In formula 1, W ipj represents the weight value of the jth codon encoding amino acid p in gene i of the host, N ipj represents the number of the jth codon encoding amino acid p in gene i, FPKM i represents the expression level of gene i;

[0135] s2-3) According to formula 2, the W of all synonymous codons encoding amino acid p in the n highly expressed genes of the host is converted to ipj Add the sum and the sum is X pj ,

[0136]

[0137] In formula 2, X pj represents the sum of the weight values ​​of the j-th codon encoding amino acid p in the host among the n highly expressed genes of the host;

[0138] s2-4) Obtain the RSCU of each codon according to formula 3,

[0139]

[0140] In formula 3, RSCU pj represents the relative synonymous codon usage frequency of the jth codon encoding amino acid p in n highly expressed genes of the host; represents the sum of the weight values ​​of all synonymous codons encoding amino acid p in n highly expressed genes of the host;

[0141] S3) Referring to the RSCU of each codon of the host, the rare codons in the target gene are replaced with dominant synonymous codons, wherein the rare codons are codons with an RSCU of less than 0.8, and the dominant synonymous codons are codons that encode the same amino acid as the rare codons and have an RSCU greater than 1.0.

[0142] To facilitate understanding of Formula 3, the RSCU of the codon AGG corresponding to arginine (Arg) in the rice TOP50W library, which contains six synonymous codons, is calculated. The sum of the weight values ​​of the codon AGG calculated according to Formula 2 is 690073.37. The Arg encoded by this codon also contains five other synonymous codons: AGA, CGG, CGA, CGT, and CGC. The sum of their weight values ​​is 173495.47, 412569.20, 126900.01, 348481.50, and 900118.19, respectively. According to the above RSCU calculation formula 3, the RSCU of AGG is calculated as follows:

[0143] RSCU=690073.37 / [1 / 6*(690073.37+173495.47+412569.20+126900.01+348481.50+900118.19]=1.56.

[0144] The results are shown in Table 4 (RICE TOP 50W), Table 5 (RICE TOP 500W), and Table 6 (MAIZE TOP 50W). In addition, the present invention also used the online software Codon Usage Database to obtain the codon usage frequency of rice and maize genomes, and plotted the codon usage frequency tables of the rice and maize gene libraries (Table 7, Table 8). The present invention defines codons with an RSCU < 0.8 as rare codons. As can be seen from Tables 4 and 5, the independently developed rice RICE TOP 50W and RICE TOP 500W codon usage frequency tables each contain 32 rare codons with RSCU < 0.8, while the rice gene pool codon usage frequency table, drawn based on website data, contains only 14 rare codons (Table 7). The independently developed maize MAIZE TOP 50W codon usage frequency table contains 31 rare codons with RSCU < 0.8 (Table 6), while the maize gene pool codon usage frequency table contains only 17 rare codons (Table 8). Therefore, by optimizing the target gene with reference to the weighted codon usage frequency tables of rice and maize developed by the present invention, the number of rare codons counted from the target gene will be greater, and after optimization, the expression level of the target gene in the host will be more potentially improved. In addition, by comparing the codon usage frequency tables of rice RICE TOP 50W and RICE TOP 500W, it was found that the number of rare codons counted using the top 50 and top 500 highly expressed genes in the transcriptome library was 32, with no difference. Therefore, in the subsequent examples, the target gene was optimized with reference to the codon usage frequency tables of RICE TOP 50W and MAIZE TOP 50W.

[0145] Table 4 Rice TOP 50W codon table

[0146]

[0147] Note: There are 32 codons with RSCU < 0.8, and 2 codons with 0.8 ≤ RSCU < 1.

[0148] Table 5 Rice TOP 500W codon table

[0149]

[0150] Note: There are 32 codons with RSCU < 0.8, and 2 codons with 0.8 ≤ RSCU < 1.

[0151] Table 6 Corn MAIZE TOP 50W codon table

[0152]

[0153] Note: There are 31 codons with RSCU < 0.8 and 2 codons with 0.8 ≤ RSCU < 1.

[0154] Table 7 Codon table of rice genome library

[0155]

[0156] Note: There are 14 codons with RSCU < 0.8, and 19 codons with 0.8 ≤ RSCU < 1.

[0157] Table 8 Codon table of maize genome library

[0158]

[0159] Note: There are 17 codons with RSCU < 0.8, and 18 codons with 0.8 ≤ RSCU < 1.

[0160] 2. New codon optimization methods

[0161] Based on the RSCU of weighted codons in the rice TOP50W library and the maize TOP50W library, rare codons and preferred codons in the target gene were determined and optimized according to the following process:

[0162] 1) Send the target gene sequence to the company for general optimization based on the codon usage frequency table of the host genome library (Table 7 and Table 8).

[0163] 2) Referring to the codon usage frequency in the weighted codon table of the host genome, determine the codon usage preference in the host gene, find rare codons in the universal optimized gene that may lead to decreased protein expression, and replace the rare codons with preferred codons according to the following principles formulated by the present invention.

[0164] Optimization principles:

[0165] All codons with a relative synonymous codon usage frequency (RSCU) < 0.8 are replaced with dominant synonymous codons with an RSCU > 1, without affecting the gene's secondary structure. If multiple synonymous codons with an RSCU > 1 are available, the replacement ratio for each synonymous codon is calculated based on the RSCU ratio, with the replacement ratio for dominant codons with a higher RSCU being increased accordingly. If a gene contains a large number of codons with an RSCU of 0.8 ≤ < 1, the codons with an RSCU of 0.8 ≤ < 1 are also replaced, leaving a small number evenly distributed later in the sequence.

[0166] 3) Perform sequence analysis and evaluation on the deeply optimized genes to detect and eliminate secondary structures, cis-acting elements and other sequences that affect their expression.

[0167] Example 2: Codon Optimization of the Insecticidal Crystal Protein Cry1Ab Gene

[0168] The coding gene of the truncated insecticidal crystal protein is collectively referred to as Cry1Ab gene in the present invention, and its nucleotide sequence is SEQ ID NO: 1. The codon preference is annotated according to RSCU in the weighted codon table of the rice RICE TOP 50W library (Table 4), and the codons with RSCU < less than 0.8 and 0.8 ≤ RSCU < 1 are annotated separately. The annotated results are as follows: Figure 7 According to statistics, there are 301 codons in Cry1Ab with RSCU less than 0.8, shown in red, and 0 codons with 0.8≤RSCU<1, shown in green.

[0169] First, we commissioned Nanjing GeneScript Biotechnology Co., Ltd. (hereinafter referred to as "GeneScript") to optimize the sequence based on the codon usage frequency table of the rice genome library (Table 7). The optimization was performed using GeneScript's codon optimization software OptimumGeneTM according to the current method. The optimized sequence removed enzyme cutting sites such as XmaI, KpnI, NotI and XhoI. At the same time, some motifs that affect the stability of mRNA were also removed, including possible mRNA splicing sites, mRNA polyA addition sites, repeated PolyA and PolyT sequences, and some forward or reverse repeated sequences. The optimized sequence was named CryV2-Os (SEQ ID NO: 2). The codon preference was annotated according to the RSCU in the weighted codon table of the rice RICE TOP 50W library (Table 4), and the codons with RSCU < 0.8 and 0.8 ≤ RSCU < 1 were annotated separately. The annotation results Figure 8 According to statistics, there are 128 codons with RSCU < 0.8 in CryV2-Os, shown in red, and 20 codons with 0.8 ≤ RSCU < 1, shown in green.

[0170] Based on the above-mentioned universal optimized version sequence, the present invention further optimized CryV2-Os according to the method provided in Example 1 to obtain the deeply optimized gene CryV4-Os (SEQ ID NO: 4). The codon preference was annotated according to the RSCU in the rice RICE TOP 50W library weighted codon table (Table 4), and the codons with RSCU < 0.8 and 0.8 ≤ RSCU < 1 were annotated separately. The annotation results are as follows: Figure 9. According to statistics, there are 0 codons with RSCU < 0.8 in CryV4-Os, and 6 codons with 0.8 ≤ RSCU < 1, shown in green. All the codons with RSCU < 0.8 in the CryV2-Os gene were optimized to synonymous codons with RSCU > 1.0. Since the number of codons with 0.8 ≤ RSCU < 1 in the CryV2-Os gene is relatively large (20), according to the optimization principle provided in Example 1: if the gene contains a large number of codons with 0.8 ≤ RSCU < 1, the codons with 0.8 ≤ RSCU < 1 should also be replaced, leaving a small part evenly distributed in the middle and rear part of the sequence. We optimized 12 of the 20 codons with 0.8 ≤ RSCU < 1 in the CryV2-Os gene to synonymous codons with RSCU > 1.0, leaving 6 in the middle and rear part of the CryV4-Os gene.

[0171] Based on the gene version CryV2-Os, we optimized it according to the optimization principles in Implementation 1. The resulting gene version, CryV4-Os, had a total of 142 codons modified, including 128 codons with an RSCU < 0.8 and 14 codons with an RSCU <= 1.0. The codon modifications were as follows: 1 of 3 GCAs was modified to GCG, and 2 to GCC. All 17 GGAs were modified to GGC. All 14 GATs were modified to GAC. Two ATTs were modified to ATC. Four of 13 CCTs were modified to CCC, and nine to CCG. One of 12 TCAs was modified to TCG, three to AGC, and eight to TCC. One of 15 TCTs was modified to TCG, three to AGC, and 11 to TCC. All 24 AATs were modified to AAC. Five CATs were modified to CAC. Thirteen ACAs were modified to ACCs. Two of the ten CCAs were modified to CCCs, and eight to CCGs. Of the 20 modified codons with an RSCU < 1.0 of 0.8, seven ACGs were replaced with ACCs (RSCU = 2.24) and seven CGGs were replaced with CGCs (RSCU = 2.04). The resulting ACC:ACG ratio in the CryV4-Os sequence was approximately 13:1, and the CGC:CGG ratio was approximately 20:3. Details of the modifications are shown in Table 9. Evaluation of the final optimized sequence, CryV4-Os, revealed no cis-acting elements or secondary structures that could affect gene expression. Furthermore, Table 9 shows that when replacing all codons with an RSCU < 0.8 with dominant synonymous codons with an RSCU > 1 without affecting gene secondary structure, if multiple synonymous codons with an RSCU > 1 were available, the replacement ratio for each synonymous codon was adjusted based on the RSCU ratio, with the replacement ratio for dominant codons with higher RSCUs being increased accordingly. For example, there are three GCA codons with RSCU=0.43 in the CRYV2-Os gene sequence. There are two candidate synonymous codons for optimization to dominant synonymous codons with RSCU>1, namely GCG (RSCU=1.28) and GCC (RSCU=1.79). Since the RSCU value of GCC is greater than that of GCG, more GCA will be optimized to GCC during optimization.

[0172] Table 9 Statistics of optimized codons for CryV2-Os and CryV4-Os

[0173]

[0174]

[0175] The Cry1Ab gene (SEQ ID NO: 1) was sent to GenScript for optimization according to the codon usage frequency table of the maize genome library (Table 8) according to the above-mentioned rice optimization process, obtaining the universal optimized sequence CryV2-Zm (SEQ ID NO: 3). Based on this sequence, optimization was then performed according to the maize MAIZE TOP 50W codon table (Table 6) and the principles of Example 1 to obtain the deeply optimized sequence CryV4-Zm (SEQ ID NO: 5).

[0176] The parameters of the resulting deeply optimized rice and maize genes, as well as the company's universally optimized gene sequences, are shown in Table 10. The data in this table demonstrates that the optimized sequences exhibit decreased free energy, improved codon suitability, and increased GC content. Furthermore, the absence of cis-acting elements that could affect expression significantly enhances gene stability. Codon statistics for the deeply optimized genes and the company's universally optimized gene sequences are shown in Table 11.

[0177] Table 10 Comparison of parameters between deep optimization sequence and general optimization sequence

[0178]

[0179] Note: Minimum free energy was calculated using DNAman 8.0 software; CAI, GC content, and cis-acting elements were analyzed by GenScript. Cis-acting element types include: Splice (GGTAAG), Splice (GGTGAT), Splice (GTAAAA), Splice (GTAAGT), Splice (GTACGT), PolyA (AATAAA), PolyA (AATGAA), PolyA (AATGGA), PolyA (TATAAA), PolyA (AATAAT), Destabilizing (ATTTA), PolyT (TTTTTT), and PolyA (AAAAAAA).

[0180] Table 11 Codon statistics of deep optimized sequences and universal optimized sequences

[0181]

[0182]

[0183]

[0184] Example 3: Transformation of Rice and Corn with the Optimized Cry1Ab Gene and Determination of Its Expression Level

[0185] 1. Cloning of the optimized Cry1Ab gene, construction of transient expression vector, and genetic transformation

[0186] 1. Cloning of the optimized Cry1Ab gene and construction of transient expression vector

[0187] The optimized CryV2-Os, CryV4-Os, CryV2-Zm, and CryV4-Zm genes were synthesized by GenScript. Using the synthesized plasmids as templates, the genes were amplified using the following primer pairs:

[0188] CryV2-OsF:ggacgatgacgataagttcgaaATGGATAACAATCCTAATATCAACG;

[0189] CryV2-OsR:cgaaagctctgcaggtcgacTCAGTACTCCGCCTC;

[0190] CryV4-OsF:ggacgatgacgataagttcgaaATGGACAACAACCCGAACATC;

[0191] CryV4-OsR:cgaaagctctgcaggtcgacTCAGTACTCCGCCTCG;

[0192] CryV2-ZmF:ggacgatgacgataagttcgaaATGGATAACAACCCCAATATTAACG;

[0193] CryV2-ZmR:cgaaagctctgcaggtcgacTCAGTACTCGGCC;

[0194] CryV4-ZmF:ggacgatgacgataagttcgaaATGGACAACAACCCGAACATC;

[0195] CryV4-ZmR: cgaaagctctgcaggtcgacTCAGTACTCGGCCTC.

[0196] High-fidelity PCR amplification was used in all construction processes Fast Pfu DNA Polymerase was purchased from Beijing Quanshijin Biotechnology Co., Ltd., catalog number AP221. A 50 μL amplification system was used, as follows:

[0197] Table 12 Amplification system

[0198]

[0199] The PCR amplification program was as follows: 95°C for 2 minutes; 95°C for 20 seconds, 58°C for 20 seconds, 72°C for 2 minutes, 30 cycles; and extension at 72°C for 5 minutes. The PCR products of CryV2-Os, CryV4-Os, CryV2-Zm, and CryV4-Zm genes were recovered and homologously recombined into the transient expression vector pUC19 (double digestion with BstBI and Sal) using the Uniclone One Step Seamless Cloning Kit (SC612) from Beijing Jinsha Biotechnology Co., Ltd., to obtain rice transient expression vectors pUC19-CryV2-Os and pUC19-CryV4-Os, and maize transient expression vectors pUC19-CryV2-Zm and pUC19-CryV4-Zm. Since the expression vector maps are similar, Figure 1 、 Figure 2 Only schematic diagrams of the pUC19-CryV4-Os and pUC19-CryV4-Zm expression vector structures are shown. Each CryV2-Os, CryV4-Os, CryV2-Zm, or CryV4-Zm gene is driven by the constitutively expressed CaMV 35S promoter and terminated by the E9 terminator. After sequencing confirmed the vectors to be correct, bacteria containing the vector plasmids were cultured overnight and the plasmids were then scaled up using the GoldHiEndoFree Plasmid Maxi Kit (CW2104M) produced by Beijing Kangwei Century Biotechnology Co., Ltd. for subsequent transient transformation.

[0200] 2. Genetic transformation of the optimized Cry1Ab gene and positive protoplast detection

[0201] 1) One-week-old etiolated seedlings of rice MH86 or corn B73 were cultured in 1 / 2 MS medium.

[0202] 2) Cut etiolated seedlings into 1 mm segments using a double-edged razor blade. Add 20 mL of enzymatic hydrolyzate per bottle of 50 etiolated seedlings and stir until the tissue is completely immersed in the liquid. Wrap in aluminum foil and protect from light.

[0203] 3) Vacuum for 30 minutes.

[0204] 4) Enzymatic hydrolysis at 28°C (40-60 rpm) for 4-6 hours.

[0205] 5) Filter through a layer of miracloth, squeeze, rinse with an equal volume of W5 solution, and squeeze again. Gently invert 10 times to mix. Divide one tube of tissue culture seedlings into two 50mL centrifuge tubes, rinsing each tube with 30mL of W5 solution. (Do not exceed 20mL, as this will cause excessive pressure and breakage. Alternatively, centrifuge after squeezing, remove the supernatant, and then rinse again with 20mL of W5 solution.)

[0206] 6) Centrifuge at 300g for 5 minutes.

[0207] 7) Remove the supernatant by aspiration, add 10 mL of W5 solution to suspend the suspension, and then filter through a layer of miracloth to remove broken cells.

[0208] 8) Incubate on ice for 30 min to allow the cells to enter the competent state.

[0209] 9) Centrifuge at 200g for 3 minutes and remove the supernatant.

[0210] 10) Add the appropriate volume of Mmg solution to resuspend and gently shake to mix. Calculate the volume based on the number of plasmids to be transfected: 200 μL / tube x the number of tubes.

[0211] 11) Take 200 μL of cells, add 10 μg of plasmid, and shake gently (dilute the plasmid to 1 μg / μL in advance and add 10 μL directly to each tube).

[0212] 12) Add 200 μL of 40% PEG and immediately mix gently by hand. Incubate at 28°C for 20 minutes (starting from the first sample).

[0213] 13) Add 4 mL of W5 solution and shake gently.

[0214] 14) Centrifuge at 250 g for 3 min and remove the supernatant.

[0215] 15) Resuspend in 1 mL of W5 solution and incubate at 28°C in the dark for 12-16 hours. The resulting protoplasts are grouped according to the introduced gene into CryV2-Os, CryV2-Zm, CryV4-Os, and CryV4-Zm. Use directly for microscopic examination and further testing.

[0216] Table 13 Enzyme hydrolysate formula

[0217]

[0218]

[0219] Note: After adding Macerozyme (pectinase), heat to 55°C for 10 minutes to dissolve, and then add subsequent reagents.

[0220] Table 14W5 solution formula

[0221] Final concentration 50mL 500mL 100 mM MES (pH 5.7) 2mM 1mL 10mL 5M NaCl 154mM 1.54mL 15.4mL <![CDATA[CaCl2(CaCl2.2H2O)]]> 125mM 0.69g(0.92g) 6.93g(9.2g) 1M KCl 5mM 0.25mL 2.5mL

[0222] Table 15Mmg solution formula

[0223] Final concentration 10mL 50mL 100 mM MES (pH 5.7) 4mM 400 μL 2mL D-Mannitol 0.6M 1.09g 5.45g <![CDATA[100mM MgCl2]]> 15mM 1.5mL 7.5mL <![CDATA[ddH2O]]> to 10mL Up to 50 mL

[0224] Table 16 40% PEG formulation

[0225] Final concentration 40mL 80mL D-Mannitol 0.2M 1.45g 2.91g <![CDATA[1M CaCl2]]> 100mM 4mL 8mL PEG4000 40% 16g 32g <![CDATA[ddH2O]]> Up to 40 mL Up to 80 mL

[0226] 2. Construction of optimized Cry1Ab gene rice expression vector and genetic transformation

[0227] 1. Construction of optimized Cry1Ab gene rice expression vector

[0228] Using the optimized and synthesized plasmid as a template, the following primers were used to amplify the rice universal optimized gene CryV2-Os and the deeply optimized gene CryV4-Os:

[0229] CryV2P-OsF:tatttacaattacagcggccGCCAGATGGATAACAATCCTAATAT;

[0230] CryV2P-OsR:ccttgggtctcacctacttactCAGTCAGTACTCCGCCTCGAATG;

[0231] CryV4P-OsF:tatttacaattacagcggccGCCAGATGGACAACAACCCGAACATCAACG;

[0232] CryV4P-OsR:ccttgggtctcacctacttactCAGTCAGTACTCCGCCTCGAAGG.

[0233] High-fidelity PCR amplification was also used in the construction process Fast Pfu DNA Polymerase was purchased from Beijing Quanshijin Biotechnology Co., Ltd., catalog number AP221. A 50 μL amplification system was used, as shown in Table 12. The PCR amplification program was as follows: 95°C for 2 minutes; 95°C for 20 seconds, 58°C for 20 seconds, and 72°C for 2 minutes, for 30 cycles; and extension at 72°C for 5 minutes. PCR products of the CryV2-Os and CryV4-Os genes were recovered and seamlessly cloned using the Uniclone One Step Seamless Cloning Kit (SC612) from Beijing Jinsha Biotechnology Co., Ltd. into the PvuII restriction site of the inducible deletion vector system pCNDL, maintaining the remaining nucleotide sequences of the pCNDL vector system unchanged. The resulting plant expression vectors were named pCNDL-CryV2-Os and pCNDL-CryV4-Os, respectively. After sequencing verification, these vectors were transformed into competent Agrobacterium EHA105 cells by electroporation for rice transformation. The nucleotide sequence of the inducible deletion vector system pCNDL is shown in Table 17. This vector has a double LB element, which can effectively prevent the read-through phenomenon of transcription. The nucleotide sequence shown in Table 17 contains the Ubiquitin promoter sequence at positions 9384 to 10362, the Ubiquitin first intron as an enhancing element at positions 10362 to 11372, the Ω sequence at positions 11379 to 11445, the PvuII restriction enzyme recognition site at positions 11454 to 11459, and the TgluB5 terminator at positions 11474 to 11970.

[0234] Table 17 Nucleotide sequence of the inducible deletion vector system pCNDL

[0235]

[0236]

[0237]

[0238]

[0239] 2. Agrobacterium-mediated stable genetic transformation of rice and detection of positive plants

[0240] Genetic transformation of rice was carried out according to conventional methods in the art, and the exemplary steps are as follows:

[0241] 1) The constructed plant expression vectors pCNDL-CryV2-Os and pCNDL-CryV4-Os were introduced into Agrobacterium tumefaciens EHA105 by electroporation to obtain recombinant Agrobacterium.

[0242] 2) A recombinant Agrobacterium monoclone was inoculated into 20 mL of YEB liquid medium containing 50 mg / L kanamycin and 50 mg / L rifampicin, and cultured with shaking at 28°C and 220 rpm for 12 to 16 h. The recombinant Agrobacterium monoclone was then inoculated into YEB liquid medium containing 100 μM acetosyringone at a ratio of 2% (volume percentage) and cultured with shaking at 28°C and 220 rpm until the OD600 value reached approximately 0.5. The cells were collected by centrifugation at 5,000 g for 10 min; the supernatant was discarded, and the cells were suspended in approximately 40 mL of AAM (with AS added to a final concentration of 300 μM) to an OD600 value of 0.4 to 0.5. The cells were cultured in a shaker at 28°C and 100 rpm for 30 min to obtain an Agrobacterium infection solution, which was then used to infect rice calli.

[0243] 3) Mature rice Minghui 86 seeds were shelled and threshed, placed in a 100 mL Erlenmeyer flask, and soaked in a 70% (volume percentage) ethanol aqueous solution for 30 seconds. The seeds were then placed in a 25% (volume percentage) sodium hypochlorite aqueous solution and sterilized by shaking at 120 rpm for 30 minutes. The seeds were rinsed three times with sterile water and dried with filter paper. The seeds were then placed embryo-side down on an induction medium to induce callus at 30°C in full light for 7 days. The grown callus was pinched and used for rice transformation.

[0244] 4) After completing step 3, embryogenic calli with good growth status were taken and immersed in the Agrobacterium infection solution obtained in step 2. The calli were gently shaken at 80 rpm at 28°C for 30 minutes. Then, they were placed on a co-cultivation medium covered with a layer of sterile filter paper and incubated in the dark at 25°C for 4 days.

[0245] 5) After completing step 4, the callus tissue was placed in a sterile culture dish, rinsed 2 to 3 times with sterile water containing 400 mg / L Cef, blotted dry on sterile filter paper, and then placed on pre-culture medium and cultured in the dark at 25°C for 3 to 4 days.

[0246] 6) Take the callus obtained in step 5 and place it on the screening medium. After culturing under alternating light for 2 weeks, replace the new screening medium and continue to culture under alternating light for 2 weeks. Transfer the resistant callus grown on the original callus to the screening medium and culture under alternating light for 2 weeks. The alternating light culture is alternating light culture and dark culture under the following conditions: 28°C; 14 hours of light culture / 10 hours of dark culture; the light intensity during light culture is 90 μE / m 2 / s.

[0247] 7) After completing step 6, take the vigorously growing resistant callus and transfer it to the pre-differentiation medium. After 7 to 10 days, when green spots appear on the callus, place it on the differentiation medium and culture it in alternating light (i.e., alternating light culture and dark culture, conditions: 28°C; 14 hours light culture / 10 hours dark culture; light intensity during light culture is 90 μE / m2 / s) for 3 weeks, resistant buds differentiated from the resistant callus.

[0248] 8) After completing step 7, take the resistant buds, place them on a rooting medium, and culture them under alternating light conditions (i.e., alternating light and dark cultivation, conditions: 28°C; 14 hours of light cultivation / 10 hours of dark cultivation; light intensity during light cultivation is 90 μE / m 2 / s) to obtain resistant plants. When the resistant plants reached 6 to 10 cm in height, they were cultured in open water and, after developing new roots, transplanted into a greenhouse. The resulting T0-generation transgenic rice plants were named CryV2-Os and CryV4-Os, respectively, based on the introduced genes. Genomic DNA was then extracted from the leaves of the T0-generation transgenic rice plants using the CTAB method.

[0249] YEB medium: Add 5g tryptone, 1g yeast extract, 5g sucrose, 5g beef extract, and 0.5g MgSO₄ to every 1L of liquid medium. Adjust the pH to 7.0 with NaOH. Sterilize the medium under high temperature and high pressure (121°C, 1.034×10⁵ Pa) for 20 minutes and store at room temperature. For solid medium, add 15g / L agar powder before autoclaving.

[0250] The rice transformation medium formula is shown in Table 18.

[0251] Table 18 Rice transformation medium formula

[0252]

[0253]

[0254] 9) Molecular testing

[0255] The leaves of the T0 generation transgenic regenerated plants were taken and the genomic DNA was extracted using the DNA rapid extraction method for PCR detection. The primers used in PCR included the detection primers (HPT-F / R) HPTF: 5'-CGTGGATATGTCCTGCGGGT-3', HPTR: 5'-GGCGACCTCGTATTGGG-3' for the selection marker gene hygromycin phosphotransferase gene (HPT) and the target gene detection primers 1Abv4 F: 5'-ACAGCCGGACCTACCCGATC-3'; 1Abv4 R: 5'-CGATGTACACCTCGTTGCC-3'. The obtained PCR products were detected by 1% agarose gel electrophoresis. The detection results of some T0 generation transgenic CryV4-Os rice lines are shown in Figure 2. Figure 4 As shown, the 1047 bp band is the amplified product of the target gene CryV4-Os, and the 561 bp band is the amplified product of the hygromycin resistance gene.

[0256] 3. Optimizing gene expression level detection

[0257] 1. Detection of transcriptional expression levels

[0258] Take 50 μL of the overnight cultured protoplasts prepared in step 1 and centrifuge at 200 g for 3 minutes. Remove the supernatant and add 1 mL of TRIzol (Invitrogen) to extract total RNA. For the rice material stably transformed in step 2, randomly select more than 30 positive lines that are transgenic for CryV2-Os or CryV4-Os. Remove leaves from these lines at the jointing stage, grind them, and add 1 mL of TRIzol (Invitrogen) to extract total RNA. Genomic DNA was removed using RQ1 RNase-free DNase (Promega, M6101) and reverse-transcribed to cDNA using the Oligo(dT)15 primer in the GoScript™ Reverse Transcription System (Promega, A5001). Transcription levels of the optimized genes were detected by qPCR using the following primers:

[0259] Universal primer pair for detecting CryV2-Os and CryV4-Os genes in transgenic rice:

[0260] CryV2 / V4rel-OsF:CAACCTCCAGTTCCAACCTCCATC;

[0261] CryV2 / V4rel-OsR:CCTGAAGGAGCCCGACTGCAG.

[0262] Universal primer pair for detecting CryV2-Zm and CryV4-Zm genes in transgenic maize:

[0263] CryV2 / V4rel-ZmF:AACGAGTGCATCCCGTACAAC;

[0264] CryV2 / V4rel-ZmR:TCCCACTGGCTCGGGCC.

[0265] The rice internal reference gene primers are:

[0266] OsActinrealF:GACCCAGATCATGTTTGAGACC;

[0267] OsActinrealR:CATCACCAGAGTCCAACACAATAC.

[0268] The primers for the maize internal reference gene are

[0269] ZmrealF:ATGGTCAAGGCCGGTTTCG;

[0270] ZmrealR:TCAGGATGCCTCTCTTGGCC.

[0271] The results showed that the relative expression levels of the deep optimized genes CryV4-Os and CryV4-Zm mRNA in the transformed protoplasts were 1.93 times and 3.49 times the relative expression levels of the common optimized genes CryV2-Os and CryV2-Zm mRNA, respectively. Figure 5 The q-PCR average value of the deep optimized gene CryV4-Os in the T0 generation of stably transformed rice was 0.90, and the qPCR average value of the universal optimized gene CryV2-Os was 0.19. The relative expression of the deep optimized gene rice mRNA was 4.74 times that of the universal optimized gene rice mRNA ( Figure 6 ).

[0272] 2. Protein expression level detection

[0273] Take 300 μL of overnight cultured protoplast cells prepared in step 1, centrifuge at 200g for 3 minutes, remove the supernatant and add 400 μL of plant protein extract (Lot01386, Kangwei Century) to extract total protein. The material for extracting total plant protein from the stably transformed rice obtained in step 2 is the same as the material for extracting total RNA. After adding the plant protein extract to the above materials, shake vigorously on an oscillator for 60 seconds, quickly place on ice, and let stand for 30 minutes. Centrifuge at 4°C, 12,000rpm for 10 minutes. Transfer the supernatant to a 1.5mL Eppendorf tube and repeat the above operation. The obtained extract is placed on ice for standby use.

[0274] Refer to the instructions of the BCA protein quantification kit (CW2341) of Beijing Kangwei Century Co., Ltd. for total protein quantification. First, dilute the BSA standard protein (2000ng / μL) into different concentration gradients such as 1000ng / μL, 500ng / μL, 250ng / μL, 125ng / μL, 62.5ng / μL, and 31.25ng / μL. Select the three leftmost columns of test wells on a marked 96-well plate for preparing a standard curve. Repeat each standard three times, and add 25μL of the first and last rows of test wells respectively. BSA standard protein and HO were added to the middle wells, with 25 μL of BSA standard protein diluents (1000 ng / μL, 500 ng / μL, 250 ng / μL, 125 ng / μL, 62.5 ng / μL, and 31.25 ng / μL) added sequentially from top to bottom. 25 μL of a 10-fold diluted protein sample was added to each sample well. Each sample was repeated three times. Then, 200 μL of BCA working solution (reagent A and reagent B were mixed at a 50:1 ratio, depending on the number of BSA standard proteins and sample to be tested) was added to each well. After thorough mixing, the 96-well plate was covered and incubated at 37°C for 30 min, followed by cooling to room temperature. Absorbance was measured at 562 nm using an MD SpectraMax 190 full-wavelength microplate reader. The absorbance OD562 was recorded and a standard curve was plotted. The concentration of the protein sample was then calculated based on the standard curve.

[0275] Protein expression level detection was performed using the Cry1Ab / Ac enzyme-linked immunosorbent assay kit (Cat. No. AA0341) from Shanghai Youlong Biotechnology Co., Ltd. The specific steps are as follows:

[0276] 1) Take out the required reagents from the refrigerated environment and place them at room temperature (20-25℃) for more than 30 minutes. Note that each liquid reagent must be shaken before use.

[0277] 2) Take out the required number of microplates, place the unused microplates in a ziplock bag, and store at 2-8°C.

[0278] 3) The washing working fluid also needs to be warmed up before use.

[0279] 4) Numbering: Number the corresponding microwells of the samples and standards in sequence. Make two parallel wells for each sample and standard, and record the positions of the standard wells and sample wells.

[0280] 5) Add standard / sample: Add 100 μL of sample extract (blank control) / standard / sample to the corresponding microwells, gently shake to mix, and react at 25°C in a dark environment for 45 minutes.

[0281] 6) Washing the plate: Shake off the liquid in the wells and wash thoroughly with 250 μL / well of washing solution 4-5 times, with an interval of 10 seconds between each wash. Pat dry with absorbent paper (air bubbles that are not removed after patting dry can be punctured with a clean pipette tip).

[0282] 7) Add enzyme labeling working solution: Add 100 μL / well of enzyme labeling working solution, gently shake to mix, incubate at 25°C in a dark environment for 30 minutes, remove the plate and repeat step 6.

[0283] 8) Color development: Add 100 μL / well of color developer and incubate at 25°C in a dark environment for 15 min.

[0284] 9) Assay: Add 100 μL / well of stop solution, gently shake to mix, set the microplate reader at 450 nm (dual wavelength 450 / 630 nm detection is recommended, and the data should be read within 5 minutes), and measure the OD value of each well.

[0285] 10) Draw a standard curve with the absorbance of the standard as the y-axis and the concentration (ppb) of the Cry1Ab / Ac standard as the x-axis. Substitute the absorbance of the sample into the standard curve and read the corresponding concentration of the sample from the standard curve.

[0286] The test results showed that the relative expression levels of the deeply optimized genes CryV4-Os and CryV4-Zm were 3.06 and 3.29 times that of the commonly optimized genes CryV2-Os and CryV2-Zm, respectively. Figure 5 The mean ELISA value of the deeply optimized rice gene in the T0 generation of stably transformed rice was 1.09, while the mean ELISA value of the universally optimized gene was 0.53; the relative expression level of the rice protein of the deeply optimized gene CryV4-Os was 2.06 times that of the rice protein of the universally optimized gene CryV2-Os ( Figure 6 The above results show that the protein levels of the deeply optimized genes optimized by the novel codon optimization method independently developed by the present invention are significantly higher than those of the commonly optimized genes in rice and corn.

[0287] The present invention has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present invention, and without the need to carry out unnecessary experimental conditions, the present invention can be implemented in a wide range under equivalent parameters, concentrations and conditions. Although the present invention provides specific embodiments, it should be understood that further improvements can be made to the present invention. In short, according to the principles of the present invention, the present invention is intended to include any changes, uses or improvements to the present invention, including changes that depart from the disclosed scope of the present invention and are made using conventional techniques known in the art.

Claims

1. A method for optimizing target gene codons, characterized in that: The method comprises: S1) receiving host transcriptome sequencing data of the expression site of the target gene at a specific expression period of the host, and selecting n highly expressed genes based on the host transcriptome sequencing data, denoted as gene i, where i represents the i-th gene, n is a natural number greater than or equal to 1, and the value of i ranges from 1 to n; S2) obtaining the RSCU of each codon in n highly expressed genes of the host; S3) Referring to the RSCU of each codon of the host, the rare codons in the target gene are replaced with dominant synonymous codons, wherein the rare codons are codons with an RSCU of less than 0.8, and the dominant synonymous codons are codons that encode the same amino acid as the rare codons and have an RSCU greater than 1.

0.

2. The method according to claim 1, characterized in that S2) obtaining the RSCU of each codon in n highly expressed genes of the host comprises the following steps: s2-1) Obtaining the number N of various codons encoding amino acid p in each of the n highly expressed genes ipj , where i represents the i-th gene, p represents the p-th amino acid, j represents the j-th codon encoding amino acid p, and the value of j is a natural number from 1 to n p , n p represents the number of synonymous codons encoding amino acid p, n p The value of is a natural number from 2 to 6; s2-2) Multiply the number of the j-th codon encoding amino acid p in gene i by the expression value of gene i, and obtain the weight value W of the j-th codon encoding amino acid p in gene i according to formula 1 ipj , Formula 1 is: IN ipj =N ipj *FPKM i , In formula 1, W ipj represents the weight value of the jth codon encoding amino acid p in gene i of the host, N ipj represents the number of the jth codon encoding amino acid p in gene i, FPKM i represents the expression level of gene i; s2-3) According to formula 2, the W of all synonymous codons encoding amino acid p in the n highly expressed genes of the host is converted to ipj Add the sum and the sum is X pj , Formula 2 is: In formula 2, X pj represents the sum of the weight values ​​of the j-th codon encoding amino acid p in the host among the n highly expressed genes of the host; s2-4) Obtain the RSCU of each codon according to formula 3, Formula 3 is: In formula 3, RSCU pj represents the relative synonymous codon usage frequency of the jth codon encoding amino acid p in n highly expressed genes of the host; It represents the sum of the weight values ​​of all synonymous codons encoding amino acid p in n highly expressed genes of the host.

3. The method according to claim 1 or 2, characterized in that The expression sites in S1) include but are not limited to cells, tissues and / or organs.

4. The method according to any one of claims 1 to 3, characterized in that The method further includes replacing codons with 0.8≤RSCU<1 in the target gene with dominant synonymous codons.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement any one of the methods of claims 1 to 4.

6. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, any one of the methods described in claims 1 to 4 is implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, any one of the methods described in claims 1 to 4 is implemented.

8. A device for optimizing target gene codons, characterized in that: The device includes the following modules: B1) Host transcriptome sequencing data receiving and analysis module: used to receive host transcriptome sequencing data of the expression site of the target gene to be optimized during a specific expression period of the host, and select n highly expressed genes based on the host transcriptome sequencing data, each denoted as gene i, where i represents the i-th gene, n is a natural number greater than or equal to 1, and the value of i ranges from 1 to n; B2) Host codon RSCU acquisition module: used to obtain the RSCU of each codon in n highly expressed genes of the host, including the following submodules: B2-1) Host Nij acquisition module: used to obtain the number value N of various codons encoding amino acid p in each of the n highly expressed genes ipj , where i represents the i-th gene, p represents the p-th amino acid, j represents the j-th codon encoding amino acid p, and the value of j is a natural number from 1 to n p , n p represents the number of synonymous codons encoding amino acid p, n p The value of is a natural number from 2 to 6; B2-2) Host W ipj Acquisition module: used to multiply the number of the j-th codon encoding amino acid p in gene i by the expression value of gene i, and obtain the weight value W of the j-th codon encoding amino acid p in gene i according to formula 1 ipj , Formula 1 is: IN ipj =N ipj *FPKM i , In formula 1, W ipj represents the weight value of the jth codon encoding amino acid p in gene i of the host, N ipj represents the number of the jth codon encoding amino acid p in gene i, FPKM i represents the expression level of gene i; B2-3) Host X pj Acquisition module: used to convert the W of all synonymous codons encoding amino acid p in the n highly expressed genes of the host according to formula 2 ipj Add the sum and the sum is X pj , Formula 2 is: In formula 2, X pj represents the sum of the weight values ​​of the j-th codon encoding amino acid p in the host among the n highly expressed genes of the host; B2-4) Host RSCU acquisition module: used to obtain the RSCU of each codon according to formula 3, Formula 3 is: In formula 3, RSCU pj represents the relative synonymous codon usage frequency of the jth codon encoding amino acid p in n highly expressed genes of the host; represents the sum of the weight values ​​of all synonymous codons encoding amino acid p in n highly expressed genes of the host; B3) Target gene rare codon optimization module: used to replace rare codons in the target gene with dominant synonymous codons with reference to the RSCU of each codon in the host. The rare codons are codons with an RSCU of less than 0.8, and the dominant synonymous codons are codons that encode the same amino acid as the rare codons and have an RSCU greater than 1.

0.

9. The device according to claim 8, characterized in that The host includes but is not limited to cells, microorganisms, plants or animals.

10. The device according to claim 8 or 9, characterized in that The target gene includes but is not limited to: a natural gene and / or an artificially synthesized gene.