Protein expression system-based foreign protein production method

By optimizing the D-value grouping of CDS and selecting CDS sequences that are slightly mismatched with the host cell tRNA supply, the problem of insufficient exogenous protein production in the prior art was solved, a balance between exogenous protein translation efficiency and host cell growth rate was achieved, and the total exogenous protein production was significantly improved.

CN121884944APending Publication Date: 2026-04-17SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2026-01-12
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, when improving the translation efficiency of exogenous proteins through whole-sequence optimal codon optimization, the negative impact on the growth rate of host cells is ignored, resulting in insufficient total production of exogenous proteins.

Method used

The CDS optimization method was adopted, which grouped CDS by calculating D value and selected CDS sequences with the highest expression level and slight mismatch with host cell tRNA supply, thus balancing translation efficiency and translation cost to optimize exogenous protein production.

Benefits of technology

It significantly increased the total production of exogenous proteins, balanced translation efficiency with host cell growth rate, and avoided excessive tRNA consumption caused by the all-optimal codons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884944A_ABST
    Figure CN121884944A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological medicine, and particularly relates to a CDS optimization method of a protein expression system and a foreign protein production method based on the system. According to the method, the Euclidean distance between tRNA demand and supply corresponding to CDS synonymous codons for coding foreign proteins is calculated, the expression quantity of a single CDS corresponding to a single D value in each group is obtained through experiments, and the CDS sequence of the D value group corresponding to the maximum expression quantity is obtained. A traditional CDS optimization strategy mostly adopts an optimal preference codon, tRNA can be excessively consumed in the mode, expression of other genes is interfered through the trans-regulation effect of the tRNA, high translation cost is generated by a protein expression system, and cell proliferation can be inhibited. The technical defects are overcome, and the obtained optimized CDS can give consideration to the translation efficiency of the foreign protein and the translation cost of a protein expression system, so that the total yield of the foreign protein is remarkably increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biomedical technology, and in particular relates to a method for producing exogenous proteins based on a protein expression system. Background Technology

[0002] The use of gene-engineered protein expression systems to prepare exogenous proteins enables efficient, controllable, and large-scale production of target proteins. This technology is a core support for biopharmaceuticals, enzyme engineering, antibody preparation, vaccine development, and molecular biology research, and is also one of the core technologies of the biopharmaceutical industry. To obtain high-concentration, high-quality exogenous proteins, conventional techniques involve heterologous preparation via protein expression systems: codon-optimized cDNA encoding the exogenous protein is constructed into an expression vector; after transformation, transfection, or transduction, the recombinant vector is introduced into host cells. The host cell then uses its own transcription and translation system to recognize the exogenous gene information, using its own ribosomes, tRNA, and amino acid library as raw materials to complete the synthesis and folding of the exogenous protein. Finally, high-purity target proteins are obtained through separation and purification.

[0003] Of the 20 commonly encoded amino acids in organisms, excluding two, the remaining 18 are encoded by 2–6 synonymous codons. Influenced by differences in tRNA copy number in the genome and the uneven frequency of use of various synonymous codons, there are significant differences in the frequency of synonymous codon usage; this phenomenon is known as codon usage bias (CUB). The most frequently used codons are the optimal preferred codons, and vice versa. Studies have shown that optimal preferred codons correspond to high-abundance tRNAs; the higher their usage percentage, the stronger the gene translation efficiency and the higher the protein expression level. Therefore, increasing the usage of optimal preferred codons through synonymous codon substitution is a classic optimization method to improve exogenous protein yield.

[0004] However, this classic whole-sequence optimal codon optimization method has significant limitations: it replaces all codons in the cDNA encoding the foreign protein with the host's optimal preferred codons, focusing only on improving the translation efficiency of the foreign protein while ignoring the negative impact of trans-regulation on the host cell growth rate. Intracellular tRNA resources involved in translation are already very limited; efficient overexpression of foreign genes will consume a large amount of the host cell's tRNA pool, leading to translational repression of endogenous genes. This translational repression significantly reduces the host cell's proliferation rate. Ultimately, even if the abundance of foreign protein expression in a single cell is increased, the overall total production of foreign protein will still fall short of expectations and remain insufficient due to the reduced proliferation of host cells. Summary of the Invention

[0005] In view of this, this application provides a CDS optimization method for protein expression systems and a method for producing exogenous proteins based on protein expression systems. The CDS obtained by this optimization method can take into account both the translation efficiency of exogenous proteins and the translation cost of protein expression systems, thereby significantly improving the total yield of exogenous proteins.

[0006] This application provides a CDS optimization method for a protein expression system, the method comprising: S1: Based on the amino acid sequence of the exogenous protein, obtain the CDS encoding the exogenous protein corresponding to the D value, wherein the D value is a real number greater than 0 and less than 1; The D value is defined as the D of 18 amino acids that have two or more synonymous codons. i The geometric mean of is calculated using formula I. Formula I; The D i The calculation formula is Equation II; Formula II; In formula II, i Represents any one of the 18 amino acids that have two or more synonymous codons; n i amino acids i The corresponding number of synonymous codons; Y ij The required value for the CDS synonym codon encoding the amino acid sequence, the first i The amino acid corresponds to the first j The proportion of each codon in the open reading frame (ORF) encoding the amino acid sequence is calculated; X ij The tRNA supply value of cells in the protein expression system; S2: Divide the D values ​​into equal groups from low to high according to their numerical values; S3: Extract the single CDS corresponding to a single D value in each group, and obtain the expression level of each CDS through experimental comparison; S4: Obtain the CDS sequence of the D-value group corresponding to the highest expression level.

[0007] Specifically, of the 20 amino acids that make up human proteins, only methionine and tryptophan are encoded by a single codon, while the remaining 18 amino acids are encoded by two or more synonymous codons. Therefore, the D value is the D of these 18 amino acids with two or more synonymous codons. i The geometric mean; the D i The value is the Euclidean distance between the tRNA demand in the CDS synonymous codon encoding the amino acid sequence and the tRNA supply in the cell of the protein expression system.

[0008] It should be noted that, theoretically, the minimum D value provided by this invention is 0 (i.e., Y). ij =X ij The maximum value of D is √2 (for example, the tRNA supply corresponding to four codons is 1, 0, 0, 0; while the tRNA demand is 0, 1, 0, 0 (only one codon); however, in reality, the tRNA supply is not just one type (i.e., not 1, 0, 0, 0), but rather something like 0.5, 0.3, 0.15, 0.05. For example, in E. coli, the limit range of D value is 1.08. Since the foreign gene cannot obtain enough tRNA when the D value is very large (e.g., D value greater than 1), the translation efficiency is very low and does not need to be considered. Therefore, the D value selected in this invention is a real number greater than 0 and less than 1.

[0009] In some embodiments, a smaller D value indicates a smaller difference between the tRNA demand encoding the exogenous protein CDS and the tRNA supply value of the cells in the protein expression system, while a larger D value indicates a larger difference between the tRNA demand encoding the exogenous protein CDS and the tRNA supply value of the cells in the protein expression system.

[0010] It should be noted that a smaller D value in the appropriate grouping indicates a higher translation efficiency of the corresponding CDS in the protein expression system, but also a higher translation cost and a lower total protein expression level. Conversely, a larger D value in the appropriate grouping indicates a lower translation efficiency of the corresponding CDS in the protein expression system, but also a lower translation cost and a higher total protein expression level. The scheme in this application demonstrates that the optimal CUB is not perfectly matched to the tRNA supply of the protein expression system. This application provides a D-value optimization method where the CUB selected for the CDS has a slight mismatch with the tRNA supply of the protein expression system, simultaneously considering both the translation efficiency of the exogenous protein and the translation cost of the protein expression system, ultimately effectively increasing the total yield of exogenous protein.

[0011] It should be noted that in step S2, the D values ​​are grouped at equal intervals from low to high numerical value, which can be divided into 3-10 groups; each group has 0-30 CDS entries. For example, in the embodiment... Figure 3In a, the numerical range of the D value is divided into three intervals, namely 0 < D < 0.33, 0.33 < D < 0.66, and D > 0.66. In the interval of 0 < D < 0.33, the number of CDS is 19. In the interval of 0.33 < D < 0.66, the number of CDS is 16. In the interval of D > 0.66, the number of CDS is 18. Specifically, the D value range beneficial to the foreign protein in the protein expression system (the foreign protein is highly important relative to the host cell) is 0.33 < D < 0.66, and the D value range not beneficial to the foreign protein in the protein expression system (the foreign protein is less important relative to the host cell) is 0 < D < 0.33.

[0012] In some embodiments, in step S3, the D values are equally spaced into multiple (such as 3 - 10) groups from low to high according to the numerical size. One CDS sequence corresponding to a single D value in each group is randomly selected. Through experimental verification, the expression level of the CDS can be obtained. After comparison, the group corresponding to the CDS with the highest expression level is the D value group with the highest expression level; this D value group contains multiple CDS sequences, and the CUB of the CDS sequences in this group has a slight mismatch with the tRNA supply of the protein expression system, which can take into account both the translation efficiency of the foreign protein and the translation cost of the protein expression system.

[0013] In some embodiments, the same D value can correspond to multiple CDS sequences, and the number of CDS encoding the foreign protein corresponding to the D value is 0 - 10. In the preferred embodiment, the number of CDS encoding the foreign protein corresponding to the D value is 0 - 5. Specifically, the number of CDS encoding the foreign protein corresponding to the D value is 0 - 3. For example, in embodiment Figure 3 a, the number of CDS corresponding to the same D value is 0, 1, 2, or 3.

[0014] It should be noted that the optimization method provided by the present invention confirms the D value group corresponding to the highest expression level. The difference between the tRNA demand value of the CDS sequences with the same D value and the tRNA supply value of the cells in the protein expression system is the same; the difference between the CDS sequences with the same D value lies in the synonymous codons of the amino acids at the same position in the amino acid sequence of the foreign protein.

[0015] In some embodiments, the formula for Y ij is formula III; Formula III; where f j is the number of the i th amino acid corresponding to the j th codon in the CDS sequence, and n i is the iThe number of synonymous codons corresponding to each amino acid.

[0016] In some embodiments, the X ij The calculation formula is Equation IV; Formula IV; Among them, W j For the first i The amino acid corresponds to the first j The tRNA supply index of one codon.

[0017] Specifically, the tRNA supply index (W) corresponding to each codon j The W reference is from the existing literature (dos Reis M, Savva R, Wernisch L. Solving the riddle of codon usage preferences: a test for translational selection. Nucleic Acids Res. 2004;32(17):5036-5044. Published 2004 Sep 24. doi:10.1093 / nar / gkh834). i .

[0018] In some embodiments, the W j The calculation formula is given by equation V; Formula V; Among them, S ij Et is the selective constraint coefficient for the codon-anticodon coupling efficiency. RNA This represents the tRNA expression level.

[0019] Specifically, the S ij It can be determined using existing methods. For example, the S ij This can be obtained by referring to existing literature (Dana, Alexandra, and Tamir Tuller. “The effect of tRNA levels on decoding times of mRNA codons.” Nucleic acids research vol. 42,14 (2014): 9171-81. doi:10.1093 / nar / gku646), where S... ij The codon-anticodon is [S I:U S G:C S U:A S C:G S G:U SI:C S I:A S U:G S L:A ], S of prokaryotes ij For [0, 0, 0, 0, 0.41, 0.63, 0.9749, 0.68, 0.95], the S of eukaryotes ij [0, 0, 0, 0, 0.561, 0.28, 0.9999, 0.68, 0.89].

[0020] In some embodiments, the tRNA expression level can be determined using existing methods. For example, the tRNA expression level is the Transcripts Per Million (TPM) of cellular tRNA genes in a protein expression system, determined and calculated by tRNA sequencing (tRNA-Seq); the calculation formula is Equation VI. Formula VI; Among them, Reads_count i To compare to the first i Number of reads for each tRNA gene, length i For the first i The length of each tRNA gene, where n is the total number of tRNA genes.

[0021] Specifically, the expression level of Escherichia coli tRNA in this embodiment of the invention was obtained through the Gene Expression Omnibus (GEO) database, with dataset number GSE128812.

[0022] In some embodiments, the experiment includes: a) Construct each of the CDSs mentioned in step S3 into the expression cassette of the expression vector of the protein expression system; b) Transform, transfect, or transduce the expression vector into cells of the protein expression system, screen for positive cells containing the vector, and culture them under the same conditions; c) Determine the expression level of exogenous proteins in the positive cells; d) Obtain the CDS with the highest expression level of exogenous protein in the positive cells.

[0023] In some embodiments, the expression level of the exogenous protein is determined using RT-qPCR, Western Blot, ELISA, electrochemiluminescence, immunofluorescence, or flow cytometry.

[0024] In some preferred embodiments, the expression level of the exogenous protein can be a relative expression level or an absolute expression level.

[0025] In some embodiments, in step a), the expression cassette of the expression vector is further connected to a reporter protein gene, which is connected downstream of the CDS.

[0026] In some preferred embodiments, flow cytometry can be used to measure the reporter protein in each cell in step c) to assess the expression level of the exogenous protein.

[0027] Specifically, step c) may also include determining the number of positive cells and plotting the growth curve of the positive cells. The translation cost of the CDS can be evaluated by calculating parameters such as the growth rate of each positive cell line and the expression of exogenous proteins in the cell line. The faster the cell proliferation, the lower the translation cost of the CDS; the slower the cell proliferation, the higher the translation cost of the CDS.

[0028] In some preferred embodiments, flow cytometry can be used to determine the expression level of the exogenous protein in the cells; the expression cassette of the CDS is also connected to a reporter protein gene, and the expression level of the exogenous protein is assessed by measuring the expression level of the reporter protein gene.

[0029] In some preferred embodiments, the reporter protein gene includes, but is not limited to, GFP, BFP, GST, YFP, mCherry, CAT, HRP, CFP, HcRed, or DsRed. Specifically, the reporter protein gene is linked downstream of the CDS via a linker peptide; the linker peptide can be a self-cleaving polypeptide P2A, a TEV restriction site, or other polypeptides. Using a linker peptide allows the expression cassette of the CDS to simultaneously express the reporter protein gene.

[0030] In some embodiments, the method further includes: S5: Compare the GC content, repeat sequence length, number of consecutive base pairing regions, and / or free energy of the CDS in the D-value group corresponding to the highest expression level to obtain the optimized CDS sequence.

[0031] It should be noted that step S4 can obtain multiple CDS, and the D values ​​of these CDS all belong to the D value group corresponding to the expression level peak. Experimental verification shows that the expression levels of these CDS are consistent and all are in the high expression range. Step S5 compares multiple core characteristic parameters of these CDS: GC content, repetitive sequences, continuous base pairing regions and free energy. Through comprehensive analysis, accurate screening is achieved, and finally an optimal CDS sequence is obtained.

[0032] In some preferred embodiments, in the CDS of step S4, a CDS sequence with a GC content of 30%-70%, a repeat sequence length of no more than 30 bp, a number of consecutive base pairing regions of 0-10, and a minimum free energy of 80%-100% of the lowest free energy is selected.

[0033] This application also discloses a method for producing exogenous proteins based on a protein expression system, the method comprising: 1) Based on the CDS optimization method provided by the present invention, CDS in the D-value group corresponding to the maximum expression level are synthesized; 2) Construct the CDS into the expression vector of the protein expression system; 3) Transform, transfect, or transduce the expression vector into cells of the protein expression system, and screen for positive cell lines that stably express the exogenous protein; 4) Culture the positive cell lines on a large scale under conditions suitable for expressing the exogenous protein; 5) Isolate and purify the exogenous protein from the culture medium of the cell line.

[0034] Specifically, the CDS in step 1) above can be any one obtained in step S4, or it can be the optimal CDS obtained in step S5.

[0035] In some embodiments, the protein expression system is selected from Escherichia coli expression systems, yeast expression systems, insect cell expression systems, mammalian cell expression systems, or plant expression systems.

[0036] Specifically, the yeast expression system can be a Saccharomyces cerevisiae expression system or a Pichia pastoris expression system.

[0037] Specifically, the insect cell expression system can be an insect baculovirus expression system.

[0038] Specifically, the mammalian cell expression system can be a CHO expression system or a HEK293 expression system.

[0039] Traditional CDS optimization method ( Figure 1 (a) employs an optimal codon bias strategy (Optimal CUB), completely rejecting non-optimal codons. This method assumes that translation selection is unidirectional, with the selection direction being to minimize the mismatch between the gene's optimal codon bias CUB and the tRNA supply value of the cells in the protein expression system (i.e., CDS uses optimal codons) to maximize the translation efficiency of exogenous proteins and the theoretical peak yield. However, this method ignores the trans-regulatory effect during translation: efficient overexpression of exogenous genes will consume a large amount of tRNA with optimal codons in the host, leading to restricted translation of endogenous genes in the host, and thus reducing the host cell growth rate through trans-regulation. This application overcomes the above shortcomings, and the optimized CDS method provided in this application ( Figure 1b) does not rely entirely on the optimal codon preference, but also includes the second and third most frequently used codons. This method assumes that there is a slight mismatch between the gene's optimal codon preference (CUB) and the tRNA supply value of the cell in the protein expression system, thereby avoiding excessive translation costs for the cell. It can balance the translation efficiency of foreign proteins with the translation cost of the host cell, avoid excessive consumption of tRNA corresponding to all optimal codon preferences, ensure the stability of both translation efficiency and cell growth rate, and ultimately significantly increase the total yield of foreign proteins. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0041] Figure 1 The graph shows the relationship between the D value and host cell growth fitness in the traditional CDS optimization method and the CDS optimization method provided in this application (the horizontal axis is the D value, the vertical axis is cell growth fitness, the black curve represents the correspondence between the horizontal axis and the vertical axis assumed in the traditional CDS optimization method and the CDS optimization method of this application; the red dots, gray arrows and horizontal dashed lines represent the optimal preference codon, the direction of translation selection and genetic drift / drift barrier, respectively). Figure 1 In this context, 'a' represents the curve obtained using the traditional CDS optimization method. Figure 1 In this context, 'b' represents the CDS optimization curve provided in this application.

[0042] Figure 2 This is an experimental schematic diagram of the CDS optimization method for the protein expression system provided in the embodiments of this application.

[0043] Figure 3 This paper examines how the two optimized codon usage methods of the GmR gene provided in the embodiments of this application change the relationship between the translation cost, functional gain (GmR protein translation efficiency), and growth rate of the host cell. Detailed Implementation

[0044] This application provides a CDS optimization method for a protein expression system and a method for producing exogenous proteins based on the system, which addresses the technical shortcomings of existing technologies that use a fully optimal codon strategy, resulting in limited translation of endogenous genes in the host.

[0045] The technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0046] In the following examples, all raw materials used were commercially available or self-made. The 53 GmRCDS sequences of the gentamicin resistance gene GmR were obtained through chemical synthesis. Ampicillin (Sangon Biotech, catalog number A100304-0005), gentamicin (Poferay, catalog number 1146GR001), flow cytometer (Attune N×T, Thermo Fisher Scientific), and Epoch 2 microplate spectrophotometer (BioTek) were used.

[0047] Example

[0048] This application uses the gentamicin resistance gene (GmR, gentamicin-3-acetyltransferase) as a model to explore how the codon usage of exogenous genes changes the translation cost, functional benefits (exogenous protein translation efficiency), and growth rate of host cells.

[0049] like Figure 2 As shown, the GmR gene synonymous mutant CDS was constructed into an E. coli expression vector. The mCherry gene (used to detect GmR expression levels) was linked downstream of the GmR CDS. This expression vector was then transformed into E. coli, which were cultured overnight in gentamicin-containing medium. The E. coli were then divided into two groups: one group was cultured in gentamicin-containing medium until the logarithmic growth phase, and GmR expression levels were determined by flow cytometry using mCherry fluorescence; the other group was cultured in gentamicin-containing medium, and its OD was periodically measured using a microplate reader. 600 The growth rate was calculated by measuring the growth curve of Escherichia coli.

[0050] like Figure 3 As shown, the above method specifically includes: Step 1: 52 GmR synonymous mutant CDS and 1 wild-type GmR CDS were synthesized. The guanine and cytosine contents (GC content) of these synonymous mutant CDS were all within a similar range (46%~55%). All synonymous mutation sites in the 52 mutants were located in the 60-531 nucleotide region of the open reading frame (ORF). The D values ​​of these 52 GmR synonymous mutant CDS ranged from 0.078 to 0.967, with 20 CDS sequences having the same D value (1-3 sequences). A D value of 0.078 indicates that the optimal codon of the CDS almost perfectly matches the tRNA supply value of *E. coli* strain MG1655; a D value of 0.967 indicates that the optimal codon of the synonymous mutant CDS does not significantly match the tRNA supply value of *E. coli* strain MG1655.

[0051] The theoretical maximum D value of the CDS of the GmR synonymous mutant is 1.414. The optimal preferred codon CDS does not match the tRNA supply value of the Escherichia coli MG1655 strain at all. Since the expression level of this type of synonymous mutant CDS is very likely to be too low to produce a measurable protein translation efficiency, it is not tested.

[0052] As Figure 3 shown in a of, the numerical range of the D values of these 53 GmR CDSs (0.078 to 0.967) is divided into three intervals, namely 0 < D < 0.33, 0.33 < D < 0.66, and D > 0.66. In the interval of 0 < D < 0.33, the number of GmR synonymous mutant CDSs is 19. In the interval of 0.33 < D < 0.66, the number of GmR synonymous mutant CDSs is 16. In the interval of D > 0.66, the number of GmR synonymous mutant CDSs is 18.

[0053] Step 2: As Figure 3 shown in b of, these CDSs are respectively constructed into the expression vectors of Escherichia coli to obtain 53 recombinant vectors. These recombinant expression vectors contain two expression cassettes. One expression cassette sequentially contains a strongly linked promoter J23119, a strong ribosome binding site (RBS), GmR CDS, mCherry, and a bidirectional transcription terminator tonB-P14 from Escherichia coli. This expression cassette is used to express the GmR-mCherry fusion protein to ensure that the expression of GmR can have a sufficient impact on the overall translation process of the cell. The other expression cassette sequentially contains a weakly linked promoter J23114, a strong ribosome binding site (RBS), yellow fluorescent protein (YFP), and a bidirectional transcription terminator tonB-P14 from Escherichia coli. This expression cassette is used to express the YFP protein, which can measure the overall translation efficiency of the host cell and can inversely estimate the translation cost brought by GmR. The weak promoter J23114 is used to minimize its impact on the tRNA supply value of the host cell.

[0054] Step 3: Using the calcium chloride transformation method, the 53 recombinant expression vectors constructed above are respectively transformed into Escherichia coli MG1655 competent cells, and then positive clone screening is carried out on the LB solid medium supplemented with 50 μg / mL ampicillin. The minimum inhibitory concentration (MIC) of these 53 engineered strains against gentamicin is determined by the agar dilution method, and then these positive clone strains are cultured in the LB medium supplemented with 120 μg / ml gentamicin. This concentration is lower than the minimum inhibitory concentration (MIC) of 48 out of the 53 strains, ensuring that the phenotypes of most strains can be measured.

[0055] Step 4: Determine the expression levels of fluorescent proteins mCherry and YFP in the logarithmic growth phase of the positive clone strain and the growth curve of the positive clone strain.

[0056] a. Fifty-three positive clones were cultured in LB medium containing 50 μg / mL ampicillin, with gradient concentrations of gentamicin added to the medium. The final concentrations of gentamicin were 0, 10, 20, 30, 40, 60, 70, 80, 100, 120, 140, 160, and 180 μg / mL. After the strains entered the logarithmic growth phase, flow cytometry was used to analyze more than 300,000 cells from each strain to determine the expression levels of the GmR-mCherry fusion protein and YFP protein. The fluorescence signal of mCherry was detected using filters at 620 nm and 15 nm, and the fluorescence signal of YFP was detected using filters at 530 nm and 30 nm. All experiments were performed with three biological replicates and three technical replicates. E. coli cells with mCherry and YFP fluorescence signal values ​​10 times higher than the MG1655 negative control strain were selected for subsequent data analysis. The FSC signal, as well as the fluorescence signals of mCherry and YFP, were detected in all qualified cells. The final expression level of the fluorescent protein was quantified by dividing the fluorescence signal value by the FSC signal value, to assess the cis-regulatory effect of GmR's synonymous codon usage on its own expression and its trans-regulatory effect on other endogenous genes in the host cell. The results are shown in Table 1. Figure 3 cd in the middle.

[0057] Table 1

[0058] b. The growth rate of 53 bacterial strains was determined by spectrophotometry. All strains were cultured in LB medium containing 50 μg / mL ampicillin and a series of gentamicin concentrations: 0, 10, 20, 30, 40, 60, 80, 100, 120, 140, 160, and 180 μg / mL. The strains were inoculated into LB medium and cultured overnight at 37℃ and 250 r / min. The bacterial culture was then diluted to an absorbance value (OD) at 600 nm. 600 =0.15-0.2, and the diluted bacterial solution was transferred to 96-well plates. The growth status of the strain was continuously monitored for 12 h at 37℃ using an Epoch 2 microplate spectrophotometer, with the absorbance (OD) of the bacterial solution at 600 nm measured every 10 min. 600 The doubling time (DT) of a strain is obtained by calculating the initial and final OD values ​​within a certain time period and substituting them into the following formula VII.

[0059] Formula VII; In the above Formula VII, OD t1 and OD t2 are the starting and ending OD values within the measurement period t1 to t2, respectively. Considering that for different strains under different gentamicin screening pressures, the OD range corresponding to their entry into the exponential growth phase varies, a unified standardization method is required to determine the exponential growth phase of each strain. The specific operation is as follows: For each sample to be measured, select all OD detection data combinations where the OD value is within the range of 0.2 < OD < 0.6 and at 3 or more consecutive time points. Use the "ln" linear fitting function in R language to perform linear model fitting on each group of data. Screen out the data set with the highest P - value significance and a determination coefficient R value ranking in the top 40% of all fitting combinations. Determine the growth stage corresponding to this data set as the target exponential growth phase. Use the minimum doubling time (DT) calculated within this exponential growth phase as the final doubling time of the sample. The doubling time of each strain is measured with 3 biological replicates and 3 technical replicates. Finally, divide the average doubling time of all strains in the LB medium without gentamicin by the average doubling time of the strain to be measured. The obtained ratio is the relative Wrightian fitness (Relative fitness) of the strain. The results are as shown in Figure 3 e in

[0060] According to the range of D values (0 < D < 0.33, 0.33 < D < 0.66, D > 0.66), these 53 GmR CDS are divided into 3 types. 0 < D < 0.33 indicates that the optimal preferred codons of GmR CDS have a very high matching degree with the tRNA supply value of Escherichia coli MG1655 strain; 0.33 < D < 0.66 indicates that the optimal preferred codons of GmR CDS have a moderate matching degree with the tRNA supply value of Escherichia coli MG1655 strain; D > 0.66 indicates that the optimal preferred codons of GmR CDS have a very low matching degree with the tRNA supply value of Escherichia coli MG1655 strain.

[0061] Table 1 shows the protein abundance results for different D - value groups (0 < D < 0.33, 0.33 < D < 0.66, D > 0.66) as the gentamicin concentration gradient increases. The protein abundance is calculated by multiplying the average fluorescence intensity of mCherry in each cell measured by flow cytometry by the strain growth rate; Figure 3 ​​​​​​​​​​​​In this, c represents the functional benefits of the GmR gene in strains grouped by different D values (0 < D < 0.33, 0.33 < D < 0.66, D > 0.66) as the gentamicin concentration gradient increases. The functional benefits of this gene are the result of cis-regulation. The average fluorescence intensity of mCherry in each strain was measured by flow cytometry; Figure 3 In this, d represents the translation costs of the GmR gene in strains grouped by different D values (0 < D < 0.33, 0.33 < D < 0.66, D > 0.66) as the gentamicin concentration gradient increases. The translation costs are the result of trans-regulation. The average fluorescence intensity of YFP in 53 strains was measured by flow cytometry. The translation cost of the GmR gene is the difference between the maximum average fluorescence intensity of YFP in the strain and the average fluorescence intensity of YFP in each strain (the calculation formula is as follows). The results show that when the gentamicin concentration is low (0, 60, 120 μg / ml, low gene importance), the supply-demand difference of the optimal preferred codons of Escherichia coli tRNA and the GmR gene is small (0 < D < 0.33), and the GmR CDS with 0 < D < 0.33 can obtain a high protein expression level; while when the gentamicin concentration is high (140, 160, 180 μg / ml), the supply-demand difference of the optimal preferred codons of Escherichia coli tRNA and the GmR gene is moderate (0.33 < D < 0.66), and the GmR CDS with 0.33 < D < 0.66 can obtain a high protein expression level. This shows that: as the gentamicin concentration in the culture medium increases, the importance of the GmR gene in Escherichia coli to the host also increases, that is, the higher the expression level of the GmR gene, the easier it is for Escherichia coli to survive in the culture medium. Figure 3 c (cis-regulation, translation efficiency of the GmR gene) in this and Figure 3 The results of d (trans-regulation, translation cost of the GmR gene) in this also illustrate the above conclusion, that is, the greater the supply-demand difference of the optimal preferred codons of Escherichia coli tRNA and the GmR gene, the lower the expression of the GmR gene and the smaller the functional benefits (translation efficiency); the greater the supply-demand difference of the optimal preferred codons of Escherichia coli tRNA and the GmR gene, the smaller the translation cost.

[0062] Functional benefits of the GmR gene (average translation efficiency) = average fluorescence intensity of mCherry in each strain; Translation cost of the GmR gene (average translation cost) = maximum average fluorescence intensity of YFP in the strain - average fluorescence intensity of YFP in each strain.

[0063] Figure 3In this, e represents the change in the relative Wright fitness of strains in different D-value groups (0 < D < 0.33, 0.33 < D < 0.66, D > 0.66) as the gentamicin concentration gradient increases. The relative Wright fitness of a strain is calculated by dividing the mean doubling time DT of all strains by the DT of each individual strain after measuring the growth curve. The results show that when the gentamicin concentration is low (0 - 60 μg / ml), the growth rate of the D > 0.66 group (with a large difference in the supply and demand of the optimal preferred codons of Escherichia coli tRNA and GmR gene) is the highest; when the gentamicin concentration is high (180 μg / ml), the growth rate of the 0 < D < 0.33 group (with a small difference in the supply and demand of the optimal preferred codons of Escherichia coli tRNA and GmR gene) is the highest; when the drug concentration is moderate (120 - 160 μg / ml), the growth rate of the 0.33 < D < 0.66 group (with a difference in the supply and demand of the optimal preferred codons of Escherichia coli tRNA and GmR gene) is the highest.

[0064] ; Among them, DT i is the doubling time of the i nth strain, and n is the total number of strains.

[0065] The above results indicate that the CDS optimization method provided by this application is designed based on the difference between the tRNA supply value of the host cell and the synonymous codon demand value of the CDS. This method abandons the method that completely relies on the optimal preferred codons, and at the same time incorporates the synonymous codons with the second and third highest usage frequencies. It can not only balance the translation efficiency of the foreign protein and the translation cost of the host cell, but also avoid the excessive consumption of tRNA caused by all-optimal preferred codons, ensuring the double stability of translation efficiency and cell growth rate, and ultimately significantly increasing the total yield of foreign proteins.

[0066] The above is only the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for CDS optimization of a protein expression system, characterized in that, The method includes: S1: Based on the amino acid sequence of the exogenous protein, obtain the CDS encoding the exogenous protein corresponding to the D value, wherein the D value is a real number greater than 0 and less than 1; wherein the D value is 18 amino acids D with 2 or more synonymous codons i the geometric mean of the values of D of formula I Formula I; The D i The formula for the calculation is Formula II; Formula II; In formula II, i Represents any one of the 18 amino acids that have two or more synonymous codons; n i For the first i The number of synonymous codons corresponding to each amino acid; Y ij The required value for the CDS synonym codon encoding the amino acid sequence, the first i The amino acid corresponds to the first j The proportion of each codon in the open reading frame (ORF) encoding the amino acid sequence is calculated; X ij The tRNA supply value of cells in the protein expression system; S2: Divide the D values ​​into equal groups from low to high according to their numerical values; S3: Extract the individual CDS corresponding to each D value in each group, and obtain the expression level of each CDS through experiments; S4: Obtain the CDS sequence of the D-value group corresponding to the highest expression level.

2. The CDS optimization method of claim 1, wherein, The Y ij The calculation formula is formula III; Formula III; Among them, f j In the CDS sequence, the first i The amino acid corresponds to the first j The number of codons, n i For the first i The number of synonymous codons corresponding to each amino acid.

3. The CDS optimization method of claim 1, wherein, The X ij The formula for the calculation is formula IV; Formula IV; Among them, W j For the first i The amino acid corresponds to the first j The tRNA supply index of one codon.

4. The CDS optimization method of claim 3, wherein, The W j The formula for the calculation is formula V; Formula V; where S ij is the selective constraint coefficient on codon-anticodon coupling efficiency, Et RNA is the tRNA expression level.

5. The CDS optimization method of claim 4, wherein, The tRNA expression level is the Transcripts Per Million (TPM) of cellular tRNA in the protein expression system.

6. The CDS optimization method of claim 1, wherein, The experiment includes: a) Construct each of the CDS in S3 into the expression cassette of the expression vector of the protein expression system; b) Transform, transfect, or transduce the expression vector into cells of the protein expression system, screen for positive cells containing the vector, and culture them under the same conditions; c) Determine the expression level of exogenous proteins in the positive cells; d) Obtain the CDS with the highest expression level of exogenous protein in the positive cells.

7. The CDS optimization method of claim 6, wherein, The expression level of the exogenous protein was determined by RT-qPCR, Western Blot, ELISA, electrochemiluminescence, immunofluorescence, or flow cytometry.

8. The CDS optimization method of any of claims 1 to 7, wherein, The method further includes: S5: Compare the GC content, repeat sequence length, number of consecutive base pairing regions, and / or free energy of the CDS in the D-value group corresponding to the highest expression level to obtain the optimized CDS sequence.

9. A method for producing a foreign protein based on a protein expression system, characterized by, The method includes: 1) The CDS optimization method according to any one of claims 1 to 8 synthesizes the CDS in the D-value group corresponding to the maximum expression level; 2) Construct the CDS into the expression vector of the protein expression system; 3) Transform, transfect, or transduce the expression vector into cells of the protein expression system, and screen for positive cell lines that stably express the exogenous protein; 4) Culture the positive cell lines on a large scale under conditions suitable for expressing the exogenous protein; 5) The positive cell lines were broken up, separated, and purified to obtain the exogenous protein.

10. The exogenous protein production method according to claim 9, wherein, The protein expression system is selected from Escherichia coli expression system, yeast expression system, insect cell expression system, mammalian cell expression system or plant expression system.