Method for expressing CRM197 protein

By designing specific signal sequences and optimizing nucleotide sequences, the problem of low expression and secretion efficiency of CRM197 protein in E. coli is solved, efficient expression and secretion is achieved, yield and purity are improved, and the feasibility of large-scale production is enhanced.

CN114929727BActive Publication Date: 2025-06-27GENOFOCUS CO LTD +1
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN201980101045.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-03
Filing Date
2019-10-02
Publication Date
2025-06-27
Estimated Expiration
2039-10-02

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently express and secrete CRM197 protein in E. coli, and the expression and secretion efficiency of proteins is low, which affects the feasibility of large-scale production.

Method used

By designing specific signal sequences and optimizing nucleotide sequences, combining codon background with secondary structural design, the expression and secretion of CRM197 protein into the periplasm of E. coli.

Benefits of technology

The efficient expression and secretion of CRM197 protein in E. coli was achieved, which increased the yield and purity of protein, reduced the toxicity to E. coli, and enhanced the feasibility of large-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114929727B_ABST
    Figure CN114929727B_ABST
Patent Text Reader

Abstract

The present invention relates to signal sequences for expressing CRM197 protein in Escherichia coli and secreting the CRM197 protein into the periplasm and uses thereof, and more particularly to: a signal sequence for expressing CRM197 protein; a nucleic acid encoding the signal sequence; a nucleic acid construct or expression vector comprising the nucleic acid and a CRM197 protein gene; a recombinant microorganism into which the nucleic acid construct or expression vector has been introduced; and a method for producing CRM197 protein including the step of culturing the recombinant microorganism. According to the present invention, the CRM197 protein having the same physicochemical / immunological properties as the protein isolated from the parental bacterium can be expressed even in a conventional Escherichia coli without adjusting the redox potential, and the CRM197 protein having a high periplasmic secretion efficiency can be produced even without changing the pH of the culture medium to increase the secretion into the periplasm, and thus the present invention is very useful in the production of CRM197 protein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to signal sequences for expressing CRM197 protein in Escherichia coli and secreting the CRM197 protein into the periplasm and uses thereof, and more particularly to a signal sequence for expressing CRM197 protein, a nucleic acid encoding the signal sequence, a nucleic acid construct or expression vector comprising the nucleic acid and the CRM197 protein gene, a recombinant microorganism into which the nucleic acid construct or expression vector is introduced, and a method for producing the CRM197 protein including culturing the recombinant microorganism. Background Art

[0002] Diphtheria toxin (DT) is a protein exotoxin and is secreted by pathogenic strains of Corynebacterium diphtheriae. Diphtheria toxin is an ADP-ribosylating enzyme, which is secreted as a zymogen consisting of 535 residues and is separated into two fragments (fragment A and B) by treatment with a trypsin-like protease. Fragment A is the catalytically active region and is a NAD-dependent ADP-ribosyltransferase, which specifically targets the protein synthesis factor EF-2, thereby inactivating EF-2 and interrupting protein synthesis in cells.

[0003] CRM197 was discovered by isolating various non-toxic forms and partially toxic immunologically cross-reactive forms (CRM or cross-reacting material) of diphtheria toxin (Uchida et al., Journal of Biological Chemistry 248, 3845-3850, 1973). Preferably, CRM may have a predetermined size and a composition containing all or part of DT.

[0004] CRM197 is a highly enzymatically inactive and non-toxic form of diphtheria toxin, which contains a single amino acid substitution (G52E). This mutation confers an intrinsic flexibility to the active site loop located in front of the NAD binding site, thereby reducing the binding affinity of CRM197 for NAD and removing the toxicity of DT (Malito et al., Proc. Natl. Acad. Sci. USA 109(14):5229-342012). CRM197, like DT, has two disulfide bonds. One disulfide bond connects Cys186 to Cys201, thereby connecting fragment A to fragment B. Another disulfide bond in fragment B connects Cys461 to Cys471. DT and CRM197 have nuclease activity derived from fragment A (Bruce et al., Proc. Natl. Acad. Sci. USA 87, 2995-8, 1990).

[0005] Many antigens have low immunogenicity unless chemically linked to a protein, especially in infants and young children, and these antigens are therefore produced as conjugates or conjugate vaccines. In these conjugate vaccines, the protein component is also referred to as the "carrier protein". CRM197 is commonly used as the carrier protein in protein-carbohydrate conjugates and hapten-protein conjugates. CRM197 as a carrier protein has several advantages over diphtheria toxoid and other toxoid proteins (Shinefield Vaccine, 28:4335, 2010).

[0006] Methods for preparing diphtheria toxin (DT) are well known in the art. For example, DT can be produced by purifying the toxin from a culture of Corynebacterium diphtheriae and then chemically detoxifying it, or by purifying a recombinant or genetically detoxified analogue of the toxin.

[0007] The abundance of the protein makes it impossible to achieve large-scale production of a diphtheria toxin such as CRM197 for use in vaccines. This problem has previously been solved by expressing CRM197 in Escherichia coli (Bishai, et al., J. Bacteriol. 169:5140-5151), and Bishai et al. have reported recombinant fusion proteins containing the toxin (including the tox signal sequence), resulting in degraded proteins.

[0008] Production of bacterial toxins in the periplasm, as compared to cytoplasmic production, is characterized by: i) the protein being produced in its mature form after cleavage of the signal peptide, or ii) the periplasm of E. coli being an oxidative environment that permits the formation of disulfide bonds, which can contribute to the production of properly folded soluble proteins, iii) the periplasm of E. coli containing fewer proteases compared to the cytoplasm, which helps to avoid proteolytic cleavage of the expressed protein, and iv) the periplasm also containing fewer proteins, which allows for the recombinant protein to be obtained in higher purity.

[0009] Generally, the presence of a signal sequence on a protein facilitates the transport (in prokaryotic hosts) or secretion (in eukaryotic hosts) of the protein into the periplasm. In prokaryotic hosts, the signal sequence causes the newly formed protein to lodge in the periplasm across the inner membrane, and then the signal sequence is cleaved. That is, it is important to search for signal sequences that can more efficiently produce commercial basic proteins on a large scale, and it is necessary to develop recombinant microorganisms.

[0010] Accordingly, due to substantial efforts to develop methods for producing CRM197 protein in an efficient and cost-effective manner, the inventors selected specific signal sequences, designed nucleotide sequences by combining codon context with secondary structure to optimize translation in Escherichia coli, optimized the expression of the CRM197 nucleotide sequence encoding CRM197 protein in Escherichia coli, and found that CRM197 was efficiently expressed in Escherichia coli and efficiently secreted into the periplasm without changing the pH when using these, thereby completing the present invention.

[0011] The information disclosed in this background art section is provided only for a better understanding of the background of the present invention, and thus the information may not include information on the prior art that is already known to those skilled in the art. Summary of the Invention

[0012] An object of the present invention is to provide a signal sequence for expressing a CRM197 protein having a specific sequence; a nucleic acid encoding the signal sequence; and a method for producing a CRM197 protein using the nucleic acid, so as to maximize the expression of the CRM197 protein by secreting the CRM197 protein into the periplasm of Escherichia coli while minimizing the toxicity of the CRM197 protein to Escherichia coli.

[0013] To achieve the above object, the present invention provides a signal sequence for expressing a CRM197 protein, which is represented by any one of the amino acid sequences of SEQ ID NO: 13 to SEQ ID NO: 21.

[0014] The present invention also provides a nucleic acid encoding a signal sequence for expressing the CRM197 protein.

[0015] The present invention also provides a nucleic acid construct or expression vector comprising the nucleic acid and the gene of the CRM197 protein, a recombinant microorganism into which the nucleic acid construct or expression vector is introduced, and a method for producing a CRM197 protein including culturing the recombinant microorganism. Brief Description of the Drawings

[0016] Figure 1 It is a schematic diagram showing the expression module TPB1Tv1.3, where Ptrc refers to the trc promoter, RBS refers to the A / U-rich enhancer + SD, λtR2&T7Te and rrnB T1T2 refer to transcription terminators, and BsaI and DraI refer to restriction enzyme cleavage sites for cloning.

[0017] Figure 2shows the Escherichia coli expression plasmid pHex1.3, where Ori(pBR322) refers to the replication origin of pBR322, KanR refers to the kanamycin marker, lacI refers to the lacI gene, and the other symbols are the same as those in Figure 1 the same as those in

[0018] Figure 3 is a schematic diagram showing the generation of the CRM197 expression plasmid, where T1 refers to the λtR2&T7Te transcription terminator, T2 refers to the rrnB T1T2 transcription terminator, the arrows refer to the PCR primers, the letter characters A, B, C, and D refer to the homologous regions of LIC, and the remaining symbols are the same as those in Figure 1 the same as those in

[0019] Figure 4 shows the expression behavior of the L3 and L5 fusions of CRM197 in various Escherichia coli strains, and the results of Coomassie staining (left) and Western blotting (right) after SDS-PAGE of total Escherichia coli cells, where C refers to C2894H, B refers to BL21(DE3), W refers to W3110-1, O refers to OrigamiTM2, S refers to shuffle, C-v represents C2984H containing pHex1.3 used as a negative control, L3 and L5 represent the fusion signal sequences, and CRM197 represents the reference CRM197, and the cells are loaded in each well at a density corresponding to OD 600 0.025 for Coomassie staining and at a density corresponding to OD 600 0.0005 for Western blotting.

[0020] Figure 5 shows the effects of the type and concentration of the inducer on the L5-induced expression of CRM197, and shows the Coomassie staining (left) and Western blotting (right) after SDS-PAGE, where the amount of cells loaded is the same as that in Figure 4 the same as those in, I refers to the insoluble fraction, S refers to the soluble fraction, (A) shows the culture at 25 °C, and (B) shows the culture at 30 °C.

[0021] Figure 6 shows the position of the CRM197 protein induced by the L5 fusion, where Coomassie staining after SDS-PAGE is shown at the top, Western blotting is shown at the bottom, Cr represents the reference CRM197, the arrow indicates the position of the mature CRM197, P1 refers to the supernatant obtained after treatment with the plasma membrane induction buffer, P2 refers to the periplasmic fraction, and Cy refers to the cytoplasmic fraction.

[0022] Figure 7Shows the effects of the type and concentration of the inducer on the L3-induced CRM197 expression, where (A) shows the culture at 25 °C, and (B) shows the culture at 30 °C.

[0023] Figure 8 Shows the position of the CRM197 protein induced by the L3 fusion.

[0024] Figure 9 Shows the changes in pH, temperature, impeller speed, and dissolved oxygen (DO) of the L3 strain, as well as the addition time of the expression inducer, where (A) shows the case where the temperature is maintained at 30 °C, and (B) shows the situation where the temperature is reduced to 25 °C before expression induction.

[0025] Figure 10 Shows the changes in pH, temperature, impeller speed, and dissolved oxygen (DO) of the L5 strain, as well as the addition time of the expression inducer.

[0026] Figure 11 Shows the changes in CRM197 protein expression before and after the addition of the expression inducer during the culture process, where (A) shows the expression of the L3 strain induced at 30 °C, (B) shows the expression of the L3 strain induced at 25 °C, and (C) shows the expression of the L5 strain induced at 25 °C. 200 ng of reference CRM197 was loaded in columns 3 and 10 in (A), columns 4 and 10 in (B), and columns 3 and 9 in (C). Columns 1 and 2 in (A), columns 1, 2, and 3 in (B), and columns 1 and 2 in (C) are the cell culture solutions before the addition of the expression inducer, and columns 16 in (A) and 16 in (C) are the supernatants after the completion of the culture. The subsequent columns are the cell culture solutions sampled every 2 hours after the addition of the inducer. The loaded amounts are the same as those in Figure 4 the same as in.

[0027] Figure 12 Shows the expression and separation behavior of the CRM197 protein in the L3 / L5 strain, where (A) shows the protein separation behavior in the culture using L3, (B) shows the protein separation behavior in the culture using L5. 200 ng of reference CRM197 was loaded in column 1, column 2 is the total cell culture solution after the completion of the culture, column 3 is the supernatant after treatment with the plasma membrane induction buffer, columns 4 and 7 are the periplasmic fractions (the protein loading amount in column 7 is 4 times that in column 4), and columns 5 and 6 represent the cytoplasmic fractions.

[0028] Figure 13Shows the results of SDS-PAGE of samples purified by DEAE chromatography (A) and HA chromatography (B). In (A), lane 1 refers to CRM197 produced by Corynebacterium, lane 2 refers to the proteins present in the recovered periplasmic fraction, lane 3 refers to the sample concentrated two-fold after ultrafiltration before loading onto the DEAE chromatography, lanes 4 and 5 are used to detect impurity proteins other than CRM197 using the flow-through mode and wash solution of the DEAE chromatography, lane 6 is the eluted sample from which impurities have been removed, and lane 7 is the sample eluted by using high-concentration salt to remove all proteins bound to the resin. In (B), HA Elu is CRM197 eluted after removing proteins not bound to the HA resin.

[0029] Figure 14 Shows the results of SEC-HPLC analysis of the finally purified CRM197, with a purity of 99% or higher.

[0030] Figure 15 Shows the results of SDS-PAGE (left) and Western blot (right). Lane 1 is CRM197 produced by Corynebacterium, and lane 2 is CRM197 produced by Escherichia coli (L3). Both CRM197 show bands at the same position.

[0031] Figure 16 Shows the results of intact protein molecular mass analysis using LC / MS, with a molecular weight of 58,409 Da, which corresponds to the theoretical molecular weight.

[0032] Figure 17 Shows the results of circular dichroism (CD) analysis and indicates that there are no differences in the higher-order structures between CRM197 produced by Corynebacterium ( ● ) and CRM197 produced by Escherichia coli pHex-L3 ( X ).

[0033] Figure 18 Shows the results of fluorescence spectroscopy analysis, where CRM197 produced by Corynebacterium ( ● ) has the same maximum emission wavelength of 338 nm as CRM197 produced by Escherichia coli pHex-L3 (solid line). Detailed Description

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which this invention belongs. Generally, the nomenclature used herein is well-known in the art and is commonly used.

[0035] In one embodiment of the present invention, nine signal sequences were fused to the CRM197 protein to induce its expression. For each signal sequence, nucleotide sequences optimized for translation (SEQ ID NOs: 4 to SEQ ID NO: 12) were designed considering codon context and the secondary structure of the mRNA. These constructs were inserted into the expression plasmid pHex1.3, and the expression of CRM197 was observed in five Escherichia coli strains to find the optimal E. coli strain for each construct. In addition, the culture temperature and the type and concentration of the inducer were set in the selected E. coli. The results of transforming various E. coli strains with the constructs were that the CRM197 protein could be expressed in a soluble form and secreted into the periplasm even in strains in which genes related to the redox potential (trxB, gor) were not engineered.

[0036] Thus, in one aspect, the present invention relates to a signal sequence for expressing the CRM197 protein, which is represented by any one of the amino acid sequences of SEQ ID NOs: 13 to SEQ ID NO: 21.

[0037] In another aspect, the present invention relates to a nucleic acid encoding a signal sequence for expressing the CRM197 protein.

[0038] In the present invention, the nucleic acid may be represented by any one of the nucleotide sequences of SEQ ID NOs: 4 to SEQ ID NO: 12, preferably by the nucleotide sequence of SEQ ID NO: 6 or SEQ ID NO: 8, but is not limited thereto.

[0039] As used herein, the term "signal sequence for expressing the CRM197 protein" refers to a signal sequence for expressing the CRM197 protein and secreting the CRM197 protein into the periplasm.

[0040] In one embodiment of the present invention, signal sequences of proteins targeting the outer membrane of E. coli and signal sequences derived from M13 phage (Table 3) were selected to secrete the CRM197 protein into the periplasm. Nucleotide sequences (SEQ ID NOs: 4 to SEQ ID NO: 12) were designed by combining codon context and secondary structure to optimize the translation of the selected signal sequences in E. coli.

[0041] In another aspect, the present invention relates to a nucleic acid construct comprising a nucleic acid encoding a signal sequence for expressing the CRM197 protein and the gene of the CRM197 protein.

[0042] In another aspect, the present invention relates to an expression vector comprising a nucleic acid encoding a signal sequence for expressing the CRM197 protein and the gene of the CRM197 protein.

[0043] In one embodiment of the present invention, the DNA sequence encoding the amino acid sequence of CRM197 protein (SEQ ID NO:3) was also optimized for expression in E. coli (SEQ ID NO:2). The DNA fragment of each designed signal sequence and the optimized CRM197 DNA fragment were inserted into plasmid pHex1.3( Figure 3 ).

[0044] In the present invention, any CRM197 protein gene can be used without limitation as long as it is a gene encoding CRM197 protein. Preferably, the CRM197 protein gene can be represented by the nucleotide sequence of SEQ ID NO:2, but is not limited thereto.

[0045] As used herein, the term "transformation" refers to the introduction of a specific external DNA strand into a cell from outside the cell. The host microorganism containing the introduced DNA strand is called a "transformed microorganism". As used herein, the term "transformation", which means introducing DNA into a host and enabling the DNA to be replicable through extrachromosomal factors or chromosomal integration, indicates that a vector containing a polynucleotide encoding a target protein is introduced into a host cell or the polynucleotide encoding the target protein is integrated into the chromosome of the host cell to express the protein encoded by the polynucleotide in the host cell. The transformed polynucleotide includes both the transformed polynucleotide inserted and located inside the host cell chromosome and the transformed polynucleotide located outside the chromosome, as long as it can be expressed in the host cell.

[0046] As used herein, the term "nucleic acid construct" includes both the nucleic acid construct inserted and located inside the host cell chromosome and the nucleic acid construct located outside the chromosome, as long as it can be expressed in the host cell.

[0047] Furthermore, as used herein, the term "polynucleotide" is used interchangeably with the term "nucleic acid" and includes DNA and RNA encoding a target protein. The polynucleotide can be introduced in any form as long as it can be introduced into a host cell and expressed therein. For example, the polynucleotide can be introduced into a host cell in the form of an expression cassette, which is a gene construct containing all the elements required for self-expression. The expression cassette usually contains a promoter, a transcription termination signal, a ribosome binding site, and a translation termination signal, which are operably linked to the nucleic acid. The expression cassette can take the form of an expression vector allowing self-replication. The polynucleotide can also be introduced into the host cell in its native form and operably linked to the sequences required for expression in the host cell.

[0048] As used herein, the term "vector" refers to a DNA product that contains a DNA sequence operably linked to a suitable regulatory sequence capable of expressing the DNA in a suitable host. The vector can be a plasmid, a bacteriophage particle, or a simple potential genomic insert. When transformed into a suitable host, the vector can replicate or function independently of the host genome, or some vectors can integrate with the genome. Plasmids are currently the most commonly used form of vector. Thus, the terms "plasmid" and "vector" are used interchangeably.

[0049] For the purposes of the present invention, plasmid vectors are preferably used. Typical plasmid vectors useful for achieving the above purposes contain (a) an origin of replication that replicates efficiently so that there are several to hundreds of plasmid vectors in each host cell, (b) an antibiotic resistance gene for screening host cells transformed with the plasmid vector, and (c) a restriction enzyme cleavage site for inserting an exogenous DNA fragment. Even in the absence of a suitable restriction enzyme cleavage site, the vector and the exogenous DNA can be easily ligated according to conventional methods using synthetic oligonucleotide linkers or adaptors.

[0050] In addition, when a gene is aligned with another nucleic acid sequence based on their functional relationship, it is said to be "operably linked" thereto. This can be in such a way that one or more genes and one or more regulatory sequences are linked such that gene expression is achieved when a suitable molecule (e.g., a transcriptional activator protein) binds to the one or more regulatory sequences. For example, when expressed as a preprotein involved in polypeptide secretion, the DNA of the presequence or secretion leader sequence is operably linked to the DNA of the polypeptide; when a promoter or enhancer affects the transcription of a sequence, it is operably linked to the coding sequence; when a ribosome binding site affects the transcription of a sequence, it is operably linked to the coding sequence; or when positioned to facilitate translation, the ribosome binding site is operably linked to the coding sequence.

[0051] Generally, the term "operably linked" means that the linked DNA sequences are in contact with each other, or that the secretion leader sequence is in contact with and in the reading frame. However, an enhancer does not need to be in contact. The ligation of these sequences is carried out by ligation at convenient restriction enzyme cleavage sites. When such sites are absent, synthetic oligonucleotide linkers or adaptors are used according to conventional methods.

[0052] In the present invention, the expression vector may further contain a Trc promoter.

[0053] In the present invention, the expression vector may be pHex1.3, but is not limited thereto.

[0054] It is known that CRM197 has high toxicity to Escherichia coli due to its nuclease activity. Therefore, the expression of CRM197 under undesired conditions can have an adverse effect on the growth of Escherichia coli. The Escherichia coli expression plasmid pHex1.3 has the LacI gene and can inhibit the background expression of the trc promoter, and the expression module TPB1Tv1.3 ( Figure 1 ) inserted into this plasmid is designed to inhibit the expression of CRM197 transcribed from the promoter derived from the plasmid by inserting the tR2 transcriptional terminator derived from λ phage and the Te transcriptional terminator derived from T7 phage upstream of the CRM197 gene to be expressed, and inserting the rrnB T1T2 transcriptional terminator derived from Escherichia coli downstream of said gene.

[0055] On the other hand, the present invention relates to a recombinant microorganism into which the nucleic acid construct or the expression vector has been introduced.

[0056] In the present invention, the recombinant microorganism can be Escherichia coli, but is not limited thereto.

[0057] Generally, a host cell with high DNA introduction efficiency and high expression efficiency of the introduced DNA is used as the recombinant microorganism. All microorganisms, including all bacteria, yeasts, molds, etc., that is, all microorganisms including prokaryotic and eukaryotic cells are available, and Escherichia coli is used in the examples of the present invention, but the present invention is not limited thereto, and any type of microorganism can be used as long as the CRM197 protein can be sufficiently expressed.

[0058] It should be understood that not all vectors function equally when expressing the DNA sequence of the present invention. Similarly, not all hosts act the same on the same expression system. However, those skilled in the art will be able to make appropriate selections from various different vectors, expression regulatory sequences, and hosts without excessive experimental burden and without departing from the scope of the present invention. For example, the vector should be selected considering the host, because the vector should be able to replicate in the host. The replication times of the vector, the ability to control the replication times of the vector, and the expression of other proteins encoded by the corresponding vector (such as the expression of antibiotic markers) should also be considered.

[0059] The transformed recombinant microorganism can be prepared according to any known transformation method.

[0060] In the present invention, the method for inserting a gene into the chromosome of a host cell can be selected from conventionally known genetic manipulation methods, such as methods using retroviral vectors, adenoviral vectors, adeno-associated viral vectors, herpes simplex viral vectors, poxviral vectors, lentiviral vectors, or non-viral vectors.

[0061] In addition to using expression vectors, transformation can also be carried out by directly inserting a nucleic acid construct into the chromosome of a host cell.

[0062] Generally, electroporation, liposome transfection, particle bombardment, virosomes, liposomes, immunoliposomes, polycations or lipid:nucleic acid conjugates, naked DNA, artificial virus-like particles, chemically enhanced DNA uptake, calcium phosphate (CaPO4) precipitation, calcium chloride (CaCl2) precipitation, microinjection, lithium acetate-DMSO method, etc. can be used.

[0063] For example, sonoporation, such as the method using the Sonitron 2000 system (Rich-Mar), can also be used for nucleic acid delivery, and other representative nucleic acid delivery systems include Amaxa Biosystems (Cologne, Germany), Maxcyte, Inc. (Rockville, Maryland), and BTX Molecular System (Holliston, Massachusetts). Liposome transfection methods are disclosed in U.S. Patent No. 5,049,386, U.S. Patent No. 4,946,787, and U.S. Patent No. 4,897,355, and liposome transfection reagents are commercially available, such as TRANSFECTAM TM and LIPOFECTIN TM . Cationic or neutral lipids suitable for efficient receptor recognition liposome transfection of polynucleotides include Felgner's lipids (WO 91 / 17424 and WO 91 / 16024), which can be delivered to cells by ex vivo transduction and targeted to tissues by in vivo transduction. Methods for preparing lipid:nucleic acid complexes (such as immunolipid complexes) containing target liposomes are well known in the art (Crystal, Science., 270:404-410, 1995; Blaese et al., Cancer Gene Ther., 2:291-297, 1995; Behr et al., Bioconjugate Chem., 5:382-389, 1994; Remy et al., Bioconjugate Chem., 5:647-654, 1994; Gao et al., Gene Therapy., 2:710-722, 1995; Ahmad et al., Cancer Res., 52:4817-4820, 1992; U.S. Patent No. 4,186,183; U.S. Patent No. 4,217,344; U.S. Patent No. 4,235,871; U.S. Patent No. 4,261,975; U.S. Patent No. 4,485,054; U.S. Patent No. 4,501,728; U.S. Patent No. 4,774,085; U.S. Patent No. 4,837,028; U.S. Patent No. 4,946,787).

[0064] In one embodiment of the present invention, the results of transforming various Escherichia coli strains (Table 4) with the prepared plasmid vectors were that for the L3 fusion, CRM197 was expressed in all strains, while for the L5 fusion, CRM197 was expressed in strains other than Origami TM 2 ( Figure 4 ). Even in strains in which genes related to the redox potential (trxB, gor) were not engineered, both the L3 and L5 fusions were able to express the CRM197 protein in soluble form, and the protein had the same physicochemical / immunological properties as the protein isolated from the parental strain. In addition, in the case of the L3 fusion, the CRM197 protein could be expressed in soluble form at 25 °C as well as 30 °C ( Figure 7 and Figure 12 A).

[0065] Previous reports have shown that the secretion of CRM197 into the periplasm was improved when cultured at a pH of 6.5 to 6.8 and then shifted to pH 7.5 during the induction process. In contrast, it was found that the strains prepared in the present invention efficiently secreted CRM197 into the periplasm without changing the pH of the medium ( Figure 12 ). In the case of the L5 fusion cultured at high concentration without changing the pH, the productivity of CRM197 was 3.7 g / L, and 2 g / L or more of CRM197 was secreted into the periplasm.

[0066] On the other hand, the present invention relates to a method for producing the CRM197 protein, which comprises (a) culturing a recombinant microorganism introduced with the nucleic acid construct or the expression vector to produce the CRM197 protein, and (b) recovering the produced CRM197 protein.

[0067] In the present invention, step (b) may include recovering the CRM197 protein secreted into the periplasm.

[0068] Hereinafter, the present invention will be described in more detail with reference to the examples. However, it will be apparent to those skilled in the art that these examples are provided only for illustrating the present invention and should not be construed as limiting the scope of the present invention.

[0069] Example 1: Preparation of a CRM197 overexpression plasmid

[0070] Example 1.1: Construction of plasmid pHex1.3

[0071] Prepare Escherichia coli expressing plasmid pHex1.3 as follows. After double digestion of ptrc99a (Amann et al., Gene. 69, 301-15, 1988) with SspI and DraI, a DNA fragment of approximately 3.2 kb was purified using agarose electrophoresis. The kanamycin resistance gene was amplified by PCR. The template used herein was plasmid pCR2.1, and the primers used herein were KF2 and KR (Table 1).

[0072] [Table 1]

[0073] PCR primers

[0074]

[0075]

[0076]

[0077] A PCR reaction solution was prepared using 2.5 mM of each dNTP, 10 pmol of each primer, 200 to 500 ng of template DNA, 1.25 U of PrimeSTAR HS DNA polymerase (Takara Bio Inc., Japan), and a reaction volume of 50 μl, and PCR was performed for 30 cycles, each cycle consisting of three steps: 98°C for 10 seconds, 60°C for 5 seconds, and 72°C for min / kb. After double digestion of the approximately 0.8 kb DNA fragment generated under PCR conditions with BamHI and HindIII, the DNA fragment was filled in with Klenow fragment to form blunt ends, and then ligated to the 3.2 kb DNA fragment prepared in the previous process using T4 DNA ligase. The reaction solution was transformed into Escherichia coli C2984H to prepare pHex1.1, in which the selectable marker Amp was replaced with Km.

[0078] An expression module TPB1Bv1.3 containing a promoter, RBS, and transcription terminator was synthesized in Bioneer ( Figure 1)(SEQ ID NO:1). The process of loading the expression module TPB1Bv1.3 onto pHex1.3 is as follows. The DNA fragment (1255 bp) obtained by PCR amplification using the expression module TPB1Bv1.3 as a template and TPB_F and TPB_R as primers (Table 1) was assembled with the DNA fragment (4561 bp) obtained by PCR amplification using pHex1.1 as a template and pHex_F and pHex_R as primers under the conditions shown in Table 2 via ligation-independent cloning (LIC; Jeong et al., Appl Environ Microbiol. 78, 5440 - 3, 2012), and the resulting product was transformed into Escherichia coli C2984H to prepare pHex1.3( Figure 2 ).

[0079] [Table 2]

[0080] LIC reaction solution

[0081] Material concentration Linearized vector 100 ng Insert 1 40 ng Insert 2 (if necessary) 40 ng T4 DNA polymerase (NEB) 1U <![CDATA[H2O]]> Up to 10 μL

[0082] LIC reaction conditions: Digest the vector with restriction enzymes or generate linear DNA fragments by PCR. Prepare the insert by PCR. Mix the vector with the insert as shown in the table above, and then react at room temperature for 2 minutes and 30 seconds.

[0083] Example 1.2: Construction of the CRM197 gene

[0084] The nucleotide sequence of CRM197 that is optimally expressed in Escherichia coli (CRM197ec) (SEQ ID NO:2) was synthesized in GenScript. The amino acid sequence encoded by CRM197ec is SEQ ID NO:3.

[0085] Example 1.3: Construction of the signal sequence gene

[0086] The signal sequences for secreting the CRM197 protein into the periplasm of Escherichia coli are shown in Table 3 below (SEQ ID NO:13 to 21). To optimize expression in Escherichia coli, DNA of SEQ ID NO:4 to 12 was synthesized considering codon context and secondary structure (Table 3).

[0087] [Table 3]

[0088] Signal sequences for the present invention

[0089]

[0090]

[0091] Example 1.4: Construction of Plasmids Overexpressing CRM197

[0092] As Figure 3 shown, plasmids were prepared for overexpressing CRM197 in Escherichia coli and secreting CRM197 into the periplasm. After double digestion of pHex1.3 with BsaI and DraI, DNA fragments approximately 5 kb in length were separated using agarose electrophoresis. The signal sequences L1 to L9 were amplified by PCR using the primers and templates shown in Table 1. The CRM197 DNA fragments in which the CRM197 gene was fused with each signal sequence using the LIC method were amplified by PCR using the primers and templates shown in Table 1. In vitro, the pHex1.3 digested with BsaI and DraI, each signal sequence fragment, and the compatible CRM197 fragment were assembled together using the LIC conditions described above, and the resulting product was then transformed into Escherichia coli C2984H to prepare plasmids pHex-L1-CRM, pHex-L2-CRM, pHex-L3-CRM, pHex-L4-CRM, pHex-L5-CRM, pHex-L6-CRM, pHex-L7-CRM, pHex-L8-CRM, and pHex-L9-CRM containing CRM197 fused with the corresponding signal sequence.

[0093] Example 2: Expression of CRM197 Protein in Escherichia coli

[0094] Considering the expression level and the degree of cell growth, pHex-L3-CRM and pHex-L5-CRM were selected from the plasmids prepared in Example 1. After transforming pHex-L3-CRM and pHex-L5-CRM into the Escherichia coli strains shown in Table 4 below, the expression and expression location of the CRM197 protein were evaluated.

[0095] [Table 4]

[0096] Escherichia coli Strains Used in the Present Invention

[0097]

[0098] The culture method was as follows. Colonies grown on solid medium (10 g / L soy peptone, 5 g / L yeast extract, 10 g / L NaCl, 15 g / L agar) were shaken in LB liquid medium containing 100 mM potassium phosphate (pH 7.5), km 50 μg / ml, and 0.2% lactose. All cells were subjected to SDS-PAGE, and then Coomassie staining and Western blot analysis were used to analyze CRM197 expression ( Figure 4)。For the L3 fusion, the CRM197 protein was expressed in all strains used, while for the L5 fusion, growth in liquid medium was not possible when transformed into Origami TM 2. Growth was observed in other strains with CRM197 protein expression.

[0099] Example 2.1: Expression of CRM197 via L5

[0100] BL21(DE3) containing pHex-L5-CRM was cultured in LB liquid medium (50 mL / 500 mL baffled flask) containing 100 mM potassium phosphate (pH 7.5) and 50 μg / ml kanamycin until the OD 600 reached 0.4 to 0.6. Then, 0.2%, 0.4%, or 0.6% lactose or 0.02 mM, 0.2 mM, or 2 mM IPTG (isopropyl β-D-1-thiogalactopyranoside) was added as an inducer to induce expression. The culture temperature was 25 °C or 30 °C. After culturing, the cells were harvested, suspended in 50 mM potassium phosphate (pH 7.0), and then disrupted by sonication. After disruption, centrifugation was performed to separate the supernatant (soluble fraction) from the precipitate (insoluble fraction). After SDS-PAGE of each sample, Coomassie staining and Western blotting were used to analyze the expression of the CRM197 protein ( Figure 5 ). When the inducer (lactose, IPTG) was added and then cultured at 25 °C, most of the expressed CRM197 protein was present in soluble form and had the same molecular weight as the reference CRM197, indicating that the expressed protein was the mature CRM197 with the L5 signal sequence removed ( Figure 5 A). On the other hand, when lactose was used as an inducer and cultured at 30 °C, approximately 50% of the CRM197 protein was found in the insoluble fraction, and when IPTG was added as an inducer, most of the CRM197 protein was found in the insoluble fraction ( Figure 5 B).

[0101] Osmotic shock was used to recover the periplasmic fraction to detect the location of CRM197 expression. The procedure was as follows. BL21(DE3) containing pHex-L5-CRM was cultured at 25 °C and then centrifuged to harvest the cells. The cells were resuspended in periplasm induction buffer [30 mM Tris-HCl (pH 8.0), 20% sucrose, 1 or 10 mM EDTA, 1 mM PMSF (phenylmethylsulfonyl fluoride)] until the cell concentration reached OD 600was 10, and the mixture was stirred at room temperature for 0.5 to 1 hour. Then, the cells were collected by centrifugation at 4,000 x g for 15 minutes, and the same amount of 30 mM cold (4 °C or lower) Tris-HCl (pH 8.0) was added, followed by stirring at room temperature for 0.5 to 1 hour. Then, the cells were centrifuged at 4,000 x g for 15 minutes to obtain the supernatant (periplasmic fraction, P2). After treatment with the plasma membrane induction buffer, the supernatant (P1), periplasmic fraction (P2), and cytoplasmic fraction were developed using SDS-PAGE, and then the expression of CRM197 and the location of the expressed CRM197 were evaluated using Coomassie staining and Western blotting ( Figure 5 ). It was found that L5 could successfully secrete CRM197 into the periplasm ( Figure 6 ). EDTA as the plasma membrane induction buffer was more effective at 10 mM than at 1 mM.

[0102] Example 2.2: Expression of CRM197 by L3

[0103] BL21(DE3) containing pHex-L3-CRM was expressed under the same conditions as in Example 2.1. It was found that, different from L5, CRM197 induced by L3 existed in a soluble form under all conditions ( Figure 7 ), and was secreted into the periplasm ( Figure 8 ).

[0104] Example 3: Cultivation of Escherichia coli BL21(DE3)

[0105] The BL21(DE3) strain containing pHex-L3-CRM or pHex-L5-CRM was cultivated using the following method. During the main cultivation period, feeding was carried out using the constant pH method, and the pH was maintained at 7.3 using the feeding solution (600 g / L glucose, 30 g / L yeast extract) and the alkali solution (14%-15% ammonia). The compositions of the solution and medium used for cultivation are shown in Tables 5 and 6.

[0106] [Table 5]

[0107] Compositions of the solution and medium used for cultivation in the present invention

[0108]

[0109] [Table 6]

[0110] Composition of trace metals used for cultivation in the present invention

[0111] Trace metals (100x, / L) EDTA 840 mg <![CDATA[CoCl2·6H2O]]> 250 mg <![CDATA[MnCl2·4H2O]]> 1.5g <![CDATA[CuCl2·2H2O]]> 150 mg <![CDATA[H3BO3]]> 300 mg <![CDATA[Na2MoO4·2H2O]]> 250 mg <![CDATA[Zn(CH3COO)2·2H2O]]> 1.3g Ferric citrate (III) 10g

[0112] A single colony formed on modified LB agar medium [modified Luria - Bertani (LB) agar: 10 g / L soy peptone, 5 g / L yeast extract, 10 g / L sodium chloride, 15 g / L agar, 50 mg / L kanamycin] was inoculated into the seed medium and then incubated at 30 °C for 18 hours. The obtained seed culture was again inoculated into the main medium (3 L / 5 L fermenter) at a ratio of 1% (v / v) and cultured at 30 °C. The main medium was obtained by adding 0.1% of a sterilized antifoaming agent to the seed medium. After the absorbance of the culture solution reached 30 - 40, the temperature was lowered to 25 °C. Then, 10 mM IPTG was added, and the culture was terminated after the absorbance of the culture solution reached 100 to 120.

[0113] The cultivation behavior of Escherichia coli BL21(DE3) containing pHex - L3 - CRM in a 5 L fermenter is shown in Figure 9 and the cultivation behavior of Escherichia coli BL21(DE3) containing pHex - L5 - CRM is shown in Figure 10 The expression behavior of CRM197 during fermentation in a 5 L fermenter is shown in Figure 11 For the L3 fusion, the expression yield was 1.1 to 1.2 g / L, and for the L5 fusion, the expression yield was 3.0 to 3.7 g / L. After SDS - PAGE / Coomassie staining, the amount of CRM197 was measured by comparing the relative amounts with reference CRM197 using a densitometer (GS - 900 TM , Bio - Rad laboratories Ins., Hercules, California).

[0114] Example 4: Protein purification

[0115] Example 4.1: Generation of the periplasmic fraction from cell culture

[0116] The procedure for recovering the periplasmic fraction from the cells cultured in a 5 L fermenter is as follows. The cell culture medium was centrifuged at 4 °C and 4,000 x g for 15 minutes to precipitate the cells. The cell pellet was resuspended in the modified protein periplasmic induction buffer (Table 7) based on an absorbance of 100 and stirred at room temperature for 0.5 to 1 hour.

[0117] [Table 7]

[0118] Composition of the buffer solution for preparing the periplasmic fraction of the present invention

[0119]

[0120] Then, the cells were collected by centrifugation at 4,000 x g for 15 minutes and the same amount of 30 mM cold (at 4 °C or lower) Tris-HCl (pH 8.0) was added, followed by stirring at room temperature for 0.5 to 1 hour. Then, the cells were centrifuged at 4,000 x g for 15 minutes to obtain the supernatant, and the impurities were removed using MF. The SDS-PAGE analysis of the periplasmic fraction recovered from Escherichia coli BL21(DE3) containing pHex-L3-CRM by the above process is shown in Figure 12 (A), and the SDS-PAGE analysis of the periplasmic fraction recovered from Escherichia coli BL21(DE3) containing pHex-L5-CRM is shown in Figure 12 (B). The amount of CRM197 protein present in the periplasmic fraction was found to be 1.2 g / L for the L3 strain and 2.3 g / L for the L5 strain (Table 8).

[0121] [Table 8]

[0122] The amount of CRM197 protein obtained by the cultivation of the present invention

[0123]

[0124] Example 4.2: Purification of CRM197 protein

[0125] The periplasmic fraction of the pHex-L3-CRM medium was concentrated twice using a TFF system with a 10 kDa cut-off membrane and ultrafiltered using ten volumes of 10 mM sodium phosphate solvent (pH 7.2). Purification was completed by a two-column process using an AKTA pure (GE Healthcare) system. The first column process was anion exchange chromatography (diethylaminoethyl agarose fast flow resin, DEAE) and was used to remove nucleic acids and impurity proteins. The DEAE resin is negatively charged (-) and binds to positively charged (+) proteins. The unbound proteins and impurities were mainly extracted and removed from the ultrafiltered sample by DEAE chromatography, and then the impure proteins with low binding ability other than CRM197 were removed by a subsequent washing process based on salt concentration. Then, CRM197 was eluted only by increasing the salt concentration. SDS-PAGE analysis of the sample was performed during the purification process using DEAE chromatography, and the results are shown in Figure 13 (A). First, the unbound impurities and proteins were mainly removed from the sample subjected to DEAE chromatography in a flow-through manner using hydroxyapatite (HA) chromatography, and CRM197 was eluted with a 100 mM potassium phosphate and 100 mM NaCl solvent. The results of SDS-PAGE of the eluted CRM197 are shown in Figure 13 (B).

[0126] Example 5: Comparison with CRM197 Produced Using Corynebacterium

[0127] The quality and characteristics of the finally purified CRM197 were analyzed. As a result of SEC-HPLC analysis, a purity of more than 99% was found ( Figure 14 ). SDS-PAGE and Western blot analysis showed that bands appeared at the same positions as CRM197 produced using Corynebacterium ( Figure 15 ).

[0128] Furthermore, it was found that the entire sequence of 535 amino acids constituting CRM197 was 100% identical, and the result of molecular weight measurement was that a main peak of 58,409 Da was identified, which corresponded to the theoretical molecular weight ( Figure 16 ). The higher-order structure was identified by circular dichroism (CD) analysis, and no difference was found from the higher-order structure of CRM197 produced using Corynebacterium ( Figure 17 ). Fluorescence spectrum analysis showed that the maximum emission wavelength was 338 nm, the same as that of CRM197 produced using Corynebacterium ( Figure 18 ).

[0129] The above results showed that CRM197 produced using Escherichia coli pHex-L3 strain was the same as CRM197 produced using Corynebacterium in terms of physicochemistry and immunology.

[0130] Utility

[0131] The present invention is very useful for the production of CRM197 protein because a CRM197 protein having the same physicochemical / immunological properties as the protein isolated from the parent bacterium can be expressed in ordinary Escherichia coli without regulating the redox potential, and a CRM197 protein having a high secretion efficiency into the periplasm can be produced without changing the pH of the culture medium to increase the secretion into the periplasm.

[0132] Although the specific configuration of the present invention has been described in detail, those skilled in the art should understand that this specification is provided for illustrative purposes to explain the preferred embodiments and should not be construed as limiting the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.

[0133] Sequence Listing Free Text

[0134] An electronic file is attached. <110> GenoFox Co., Ltd. EUBIO Co., Ltd. <120> Method for Expressing CRM197 Protein <130> PP-B2284 <150> KR 10-2019-0108892 <151> 2019-09-03 <160> 56 <170> KoPatentIn 3.0 <210> 1 <211> 1226 <212> DNA <213> Artificial Sequence <220> <223> TPB1Tv1.3 <400> 1 tactgagcta ataacaggcc tgctggtaat cgcaggcctt tttatttctg gctcaccttc 60 gggtgggcct ttctgcgttt acggttctgg caaatattct gaaatgagct gttgacaatt 120 aatcatccgg ctcgtataat gtgtggaatt gtgagcggat aacaatttac tgctctttaa 180 caatttatca gatccaattg gaggaacaat atgggagacc acgcgcgcga ggtctcatct 240 gcggcggcag cagggggatc gatgagcaaa ggtgaagagc tgtttaccgg tgtggtgccg 300 attctggttg aactggatgg cgacgttaac ggccataaat tcagcgtcag cggcgagggc 360 gaaggggatg ccacctacgg taaactgacc ctgaagttta tttgcaccac cggcaaatta 420 ccggttccgt ggccgacgct ggtgacgacc tttagctatg gcgtgcagtg cttttcccgt 480 tatccggacc atatgaaaca gcatgatttt tttaaaagcg cgatgccgga aggctatgtt 540 caggaacgta ccattttctt taaggatgac ggcaattaca aaacccgcgc ggaagtgaaa 600 tttgaaggtg ataccctggt caaccgcatt gaactgaaag gcattgattt caaagaagat 660 ggtaatattc tcggtcataa gctggaatat aactacaaca gccataacgt ttatatcatg 720 gcggataaac aaaaaaacgg tattaaagtg aactttaaaa ttcgccataa tatcgaagat 780 ggtagcgtgc aactggcgga tcattatcag cagaacaccc caattggcga tggcccggtg 840 ttgctgccgg ataaccacta tctgagcacc cagtcggcgc tgagtaaaga tccgaacgaa 900 aaacgcgatc acatggtgct gctggagttt gtcaccgccg ccggtatcac ccacggcatg 960 gatgaactgt ataaataatg atttaaagcg gatatcaaat aaaacgaaag gctcagtcga 1020 aagactgggc ctttcgtttt atctgttgtt tgtcggtgaa cgctctcctg agtaggacaa 1080 atccgccggg agcggatttg aacgttgcga agcaacggcc cggagggtgg cgggcaggac 1140 gcccgccata aactgccagg catcaaatta agcagaaggc catcctgacg gatggccttt 1200 ttgcgtttct acaaactctt tttgtt 1226 <210> 2 <211> 1625 <212> DNA <213> Artificial Sequence <220> <223> CRM197ec <400> 2 gaagacacgg cgctgacgat gttgttgaca gcagcaagag ctttgttatg gagaatttca 60 gcagctatca cggcaccaaa ccgggctatg ttgacagcat tcagaaaggc atccaaaagc 120 cgaaaagcgg tacccagggc aactacgacg atgactggaa agagttttat agcaccgata 180 acaagtacga cgcggcgggc tatagcgttg ataacgaaaa cccgctgagc ggtaaagcgg 240 gtggcgtggt taaggtgacc tacccgggtc tgaccaaagt gctggcgctg aaggttgaca 300 acgcggaaac catcaagaaa gagctgggcc tgagcctgac cgaaccgctg atggagcagg 360 ttggtaccga ggaattcatt aagcgttttg gtgatggtgc gagccgtgtg gttctgagcc 420 tgccgttcgc ggaaggtagc agcagcgtgg agtacatcaa caactgggaa caagcgaaag 480 cgctgagcgt ggagctggaa attaacttcg aaacccgtgg caagcgtggc caggatgcga 540 tgtacgaata tatggcgcaa gcgtgcgcgg gtaaccgtgt gcgtcgtagc gttggtagca 600 gcctgagctg cattaacctg gattgggacg ttatccgtga caagaccaaa accaagatcg 660 aaagcctgaa agagcacggt ccgattaaaa acaagatgag cgagagcccg aacaagaccg 720 tgagcgagga aaaagcgaag cagtacctgg aggaatttca ccaaaccgcg ctggagcacc 780 cggaactgag cgagctgaaa accgtgaccg gcaccaaccc ggttttcgcg ggtgcgaact 840 atgcggcgtg ggcggtgaac gttgcgcagg tgatcgatag cgaaaccgcg gacaacctgg 900 aaaagaccac cgcggcgctg agcatcctgc cgggtattgg cagcgtgatg ggcatcgcgg 960 atggtgcggt tcaccacaac accgaggaaa tcgtggcgca gagcattgcg ctgagcagcc 1020 tgatggttgc gcaagcgatc ccgctggttg gtgagctggt tgacattggt ttcgcggcgt 1080 acaactttgt ggaaagcatc attaacctgt tccaggtggt tcacaacagc tacaaccgtc 1140 cggcgtatag cccgggccac aaaacccaac cgtttctgca cgatggttat gcggtgagct 1200 ggaacaccgt tgaggacagc atcattcgta ccggtttcca gggcgagagc ggtcacgata 1260 tcaagattac cgcggaaaac accccgctgc cgattgcggg cgttctgctg ccgaccattc 1320 cgggtaaact ggacgttaac aaaagcaaga cccacattag cgtgaacggc cgtaagatcc 1380 gtatgcgttg ccgtgcgatt gatggtgacg tgaccttttg ccgtccgaaa agcccggtgt 1440 acgttggtaa cggcgtgcac gcgaacctgc acgttgcgtt ccaccgtagc agcagcgaga 1500 agatccacag caacgaaatt agcagcgaca gcatcggcgt tctgggttat caaaaaaccg 1560 tggatcatac caaagtgaat agcaagctga gcctgttctt cgagattaaa agctctgagg 1620 tcttc 1625 <210> 3 <211> 535 <212> PRT <213> Artificial Sequence <220> <223> CRM197 <400> 3 Gly Ala Asp Asp Val Val Asp Ser Ser Lys Ser Phe Val Met Glu Asn 1 5 10 15 Phe Ser Ser Tyr His Gly Thr Lys Pro Gly Tyr Val Asp Ser Ile Gln 20 25 30 Lys Gly Ile Gln Lys Pro Lys Ser Gly Thr Gln Gly Asn Tyr Asp Asp 35 40 45 Asp Trp Lys Glu Phe Tyr Ser Thr Asp Asn Lys Tyr Asp Ala Ala Gly 50 55 60 Tyr Ser Val Asp Asn Glu Asn Pro Leu Ser Gly Lys Ala Gly Gly Val 65 70 75 80 Val Lys Val Thr Tyr Pro Gly Leu Thr Lys Val Leu Ala Leu Lys Val 85 90 95 Asp Asn Ala Glu Thr Ile Lys Lys Glu Leu Gly Leu Ser Leu Thr Glu 100 105 110 Pro Leu Met Glu Gln Val Gly Thr Glu Glu Phe Ile Lys Arg Phe Gly 115 120 125 Asp Gly Ala Ser Arg Val Val Leu Ser Leu Pro Phe Ala Glu Gly Ser 130 135 140 Ser Ser Val Glu Tyr Ile Asn Asn Trp Glu Gln Ala Lys Ala Leu Ser 145 150 155 160 Val Glu Leu Glu Ile Asn Phe Glu Thr Arg Gly Lys Arg Gly Gln Asp 165 170 175 Ala Met Tyr Glu Tyr Met Ala Gln Ala Cys Ala Gly Asn Arg Val Arg 180 185 190 Arg Ser Val Gly Ser Ser Leu Ser Cys Ile Asn Leu Asp Trp Asp Val 195 200 205 Ile Arg Asp Lys Thr Lys Thr Lys Ile Glu Ser Leu Lys Glu His Gly 210 215 220 Pro Ile Lys Asn Lys Met Ser Glu Ser Pro Asn Lys Thr Val Ser Glu 225 230 235 240 Glu Lys Ala Lys Gln Tyr Leu Glu Glu Phe His Gln Thr Ala Leu Glu 245 250 255 His Pro Glu Leu Ser Glu Leu Lys Thr Val Thr Gly Thr Asn Pro Val 260 265 270 Phe Ala Gly Ala Asn Tyr Ala Ala Trp Ala Val Asn Val Ala Gln Val 275 280 285 Ile Asp Ser Glu Thr Ala Asp Asn Leu Glu Lys Thr Thr Ala Ala Leu 290 295 300 Ser Ile Leu Pro Gly Ile Gly Ser Val Met Gly Ile Ala Asp Gly Ala 305 310 315 320 Val His His Asn Thr Glu Glu Ile Val Ala Gln Ser Ile Ala Leu Ser 325 330 335 Ser Leu Met Val Ala Gln Ala Ile Pro Leu Val Gly Glu Leu Val Asp 340 345 350 Ile Gly Phe Ala Ala Tyr Asn Phe Val Glu Ser Ile Ile Asn Leu Phe 355 360 365 Gln Val Val His Asn Ser Tyr Asn Arg Pro Ala Tyr Ser Pro Gly His 370 375 380 Lys Thr Gln Pro Phe Leu His Asp Gly Tyr Ala Val Ser Trp Asn Thr 385 390 395 400 Val Glu Asp Ser Ile Ile Arg Thr Gly Phe Gln Gly Glu Ser Gly His 405 410 415 Asp Ile Lys Ile Thr Ala Glu Asn Thr Pro Leu Pro Ile Ala Gly Val 420 425 430 Leu Leu Pro Thr Ile Pro Gly Lys Leu Asp Val Asn Lys Ser Lys Thr 435 440 445 His Ile Ser Val Asn Gly Arg Lys Ile Arg Met Arg Cys Arg Ala Ile 450 455 460 Asp Gly Asp Val Thr Phe Cys Arg Pro Lys Ser Pro Val Tyr Val Gly 465 470 475 480 Asn Gly Val His Ala Asn Leu His Val Ala Phe His Arg Ser Ser Ser 485 490 495 Glu Lys Ile His Ser Asn Glu Ile Ser Ser Asp Ser Ile Gly Val Leu 500 505 510 Gly Tyr Gln Lys Thr Val Asp His Thr Lys Val Asn Ser Lys Leu Ser 515 520 525 Leu Phe Phe Glu Ile Lys Ser 530 535 <210> 4 <211> 85 <212> DNA <213> Artificial Sequence <220> <223> PelB (L1) <400> 4 ggtctcatat gaaatatctg ttaccgaccg ccgctgccgg actgctgtta ctggcggcgc 60 agccggcgat ggcgggcgag agacc 85 <210> 5 <211> 88 <212> DNA <213> Artificial Sequence <220> <223> G8 (L2) <400> 5 ggtctcatat gaaaaaaagc ctggttctga aagcgtctgt tgcggtggcg acgctggtgc 60 cgatgctgtc gtttgccggc gagagacc 88 <210> 6 <211> 73 <212> DNA <213> Artificial Sequence <220> <223> Wp (L3) <400> 6 ggtctcatat gcgttctgtg attgttgcct tcctgtttgc ctgtagcttt tgcgtgagcg 60 ccggcgagag acc 73 <210> 7 <211> 79 <212> DNA <213> Artificial Sequence <220> <223> OmpT (L4) <400> 7 ggtctcatat gcgtgcgaaa ctgctcggca ttgttctgac caccccgatt gccatttcca 60 gctttgccgg cgagagacc 79 <210> 8 <211> 82 <212> DNA <213> Artificial Sequence <220> <223> OmpA (L5) <400> 8 ggtctcatat gaaaaaaacc gccatcgcca ttgccgttgc cctcgctggc tttgccaccg 60 tggcgcaggc gggcgagaga cc 82 <210> 9 <211> 82 <212> DNA <213> Artificial Sequence <220> <223> G4 (L6) <400> 9 ggtctcatat gaaactgctg aacgtgatca actttgtttt cctgatgttt gtcagcagca 60 gtagttttgc cggcgagaga cc 82 <210> 10 <211> 73 <212> DNA <213> Artificial Sequence <220> <223> G3 (L7) <400> 10 ggtctcatat gaaaaaactg ctgtttgcca ttccgctggt tgtaccgttt tacagccaca 60 gcggcgagag acc 73 <210> 11 <211> 79 <212> DNA <213> Artificial Sequence <220> <223> Lpp (L8) <400> 11 ggtctcatat gaaagcgacg aaactggtgc tgggtgctgt gattctgggc agcacgctgc 60 tggcgggcgg cgagagacc 79 <210> 12 <211> 88 <212> DNA <213> Artificial Sequence <220> <223> Gsp (L9) <400> 12 ggtctcatat gaaaggtctg aataaaatta cctgctgttt actggcggcg ctgctgatgc 60 cgtgcgcggg tcatgcgggc gagagacc 88 <210> 13 <211> 22 <212> PRT <213> Artificial Sequence <220> <223> PelB (L1) <400> 13 Met Lys Tyr Leu Leu Pro Thr Ala Ala Ala Gly Leu Leu Leu Leu Ala 1 5 10 15 Ala Gln Pro Ala Met Ala 20 <210> 14 <211> 23 <212> PRT <213> Artificial Sequence <220> <223> G8 (L2) <400> 14 Met Lys Lys Ser Leu Val Leu Lys Ala Ser Val Ala Val Ala Thr Leu 1 5 10 15 Val Pro Met Leu Ser Phe Ala 20 <210> 15 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> Wp (L3) <400> 15 Met Arg Ser Val Ile Val Ala Phe Leu Phe Ala Cys Ser Phe Cys Val 1 5 10 15 Ser Ala <210> 16 <211> 20 <212> PRT <213> Artificial Sequence <220> <223> OmpT (L4) <400> 16 Met Arg Ala Lys Leu Leu Gly Ile Val Leu Thr Thr Pro Ile Ala Ile 1 5 10 15 Ser Ser Phe Ala 20 <210> 17 <211> 21 <212> PRT <213> Artificial Sequence <220> <223> OmpA (L5) <400> 17 Met Lys Lys Thr Ala Ile Ala Ile Ala Val Ala Leu Ala Gly Phe Ala 1 5 10 15 Thr Val Ala Gln Ala 20 <210> 18 <211> 21 <212> PRT <213> Artificial Sequence <220> <223> G4 (L6) <400> 18 Met Lys Leu Leu Asn Val Ile Asn Phe Val Phe Leu Met Phe Val Ser 1 5 10 15 Ser Ser Ser Phe Ala 20 <210> 19 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> G3 (L7) <400> 19 Met Lys Lys Leu Leu Phe Ala Ile Pro Leu Val Val Pro Phe Tyr Ser 1 5 10 15 His Ser <210> 20 <211> 20 <212> PRT <213> Artificial Sequence <220> <223> Lpp (L8) <400> 20 Met Lys Ala Thr Lys Leu Val Leu Gly Ala Val Ile Leu Gly Ser Thr 1 5 10 15 Leu Leu Ala Gly 20 <210> 21 <211> 23 <212> PRT <213> Artificial Sequence <220> <223> Gsp (L9) <400> 21 Met Lys Gly Leu Asn Lys Ile Thr Cys Cys Leu Leu Ala Ala Leu Leu 1 5 10 15 Met Pro Cys Ala Gly His Ala 20 <210> 22 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 22 gcggatccaa gagacaggat gaggatcgtt tcgc 34 <210> 23 <211> 40 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 23 cggatatcaa gcttggaaat gttgaatact catactcttc 40 <210> 24 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 24 gagatccgga gcttatactg agctaataac 30 <210> 25 <211> 29 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 25 gaaaaataaa caaaaacaaa aagagtttg 29 <210> 26 <211> 31 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 26 tacaaactct ttttgttttt gtttattttt c 31 <210> 27 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 27 cctgttatta gctcagtata agctccggat ctcg 34 <210> 28 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 28 aattggagga acaatatgaa atatct 26 <210> 29 <211> 29 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 29 aacatcgtca gcgcccgcca tcgccggct 29 <210> 30 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 30 ttggaggaac aatatgaaaa aaagcct 27 <210> 31 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 31 aacatcgtca gcgccggcaa acgacagcat 30 <210> 32 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 32 ttggaggaac aatatgcgtt ctgtga 26 <210> 33 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 33 aacatcgtca gcgccggcgc tcacgcaa 28 <210> 34 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 34 ttggaggaac aatatgcgtg cgaaact 27 <210> 35 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 35 aacatcgtca gcgccggcaa agctggaaat 30 <210> 36 <211> 25 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 36 ttggaggaac aatatgaaaa aaacc 25 <210> 37 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 37 aacatcgtca gcgcccgcct gcgccacggt 30 <210> 38 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 38 ttggaggaac aatatgaaac tgctga 26 <210> 39 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 39 aacatcgtca gcgccggcaa aactactgct 30 <210> 40 <211> 25 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 40 ttggaggaac aatatgaaaa aactg 25 <210> 41 <211> 29 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 41 acatcgtcag cgccgctgtg gctgtaaaa 29 <210> 42 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 42 ttggaggaac aatatgaaag cgacgaaa 28 <210> 43 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 43 aacatcgtca gcgccgcccg ccagcagcgt 30 <210> 44 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 44 ttggaggaac aatatgaaag gtctgaa 27 <210> 45 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 45 aacatcgtca gcgcccgcat gacccgcgca 30 <210> 46 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 46 agccggcgat ggcgggcgct gacgatg 27 <210> 47 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 47 atgctgtcgt ttgccggcgc tgacgatg 28 <210> 48 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 48 ttgcgtgagc gccggcgctg acgatg 26 <210> 49 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 49 tttccagctt tgccggcgct gacgatg 27 <210> 50 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 50 accgtggcgc aggcgggcgc tgacgatg 28 <210> 51 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 51 agcagtagtt ttgccggcgc tgacgatg 28 <210> 52 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 52 ttttacagcc acagcggcgc tgacgatg 28 <210> 53 <211> 24 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 53 tgctggcggg cggcgctgac gatg 24 <210> 54 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 54 cgcgggtcat gcgggcgctg acgatg 26 <210> 55 <211> 36 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 55 gatatccgct tttcattagc ttttaatctc gaagaa 36 <210> 56 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> primer <400> 56 ggcgcaagcg tgcgcgggta accgtgtgcg 30

Claims

1. A nucleic acid encoding a signal peptide for expressing CRM197 protein, wherein the nucleic acid is represented by the nucleotide sequence of SEQ ID NO:

8.

2. A nucleic acid construct comprising the nucleic acid according to claim 1 and the gene of the CRM197 protein.

3. The nucleic acid construct according to claim 2, wherein the gene of the CRM197 protein is represented by the nucleotide sequence of SEQ ID NO:

2.

4. An expression vector comprising the nucleic acid according to claim 1 and the gene of the CRM197 protein.

5. The expression vector according to claim 4, wherein the gene of the CRM197 protein is represented by the nucleotide sequence of SEQ ID NO:

2.

6. The expression vector according to claim 4, further comprising a Trc promoter.

7. A recombinant microorganism into which the nucleic acid construct according to claim 2 or the expression vector according to claim 4 is introduced.

8. The recombinant microorganism according to claim 7, wherein the recombinant microorganism is Escherichia coli.

9. A method for producing CRM197 protein, the method comprising: (a) culturing the recombinant microorganism according to claim 7 to produce CRM197 protein; and (b) recovering the produced CRM197 protein.

10. The method according to claim 9, wherein the step (b) comprises recovering the CRM197 protein secreted into the periplasm.

Citation Information

Patent Citations

  • Method for preparing poly(3-hydroxypropionate-b-lactate) block copolymer by using microorganisms

    KR1020190108892A

  • Liposome carriers in chemotherapy of leishmaniasis

    US4186183A

  • Compositions containing aqueous dispersions of lipid spheres

    US4217344A

  • Method of encapsulating biologically active materials in lipid vesicles

    US4235871A

  • Viral liposome particle

    US4261975A