Method for secretory production of unnatural-amino-acid-containing protein
Patent Information
- Application Number
- JP2023533180
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2022-07-07
- Filing Date
- 2022-07-07
- Publication Date
- 2025-06-13
AI Technical Summary
There is no known technology for producing proteins containing non-natural amino acids using coryneform bacteria, which limits the functional modification of proteins through the introduction of unnatural amino acids.
A method for secretory production of proteins containing non-natural amino acids is developed by culturing modified coryneform bacteria expressing an orthogonal pair of tRNA corresponding to non-natural amino acids and aminoacyl-tRNA synthetase, using a gene construct with a promoter sequence, signal peptide, and nucleic acid sequences encoding the protein, allowing for the secretion and collection of unnatural amino acid-containing proteins.
This method enables the efficient secretion and production of proteins with non-natural amino acids, such as tyrosine derivatives and lysine derivatives, into the culture medium or on the bacterial surface, facilitating their recovery and use in various applications.
Abstract
Description
Secretory production of proteins containing unnatural amino acids
[0001] The present invention relates to a method for secretory production of proteins containing non-natural amino acids (noncanonical amino acids; ncAAs).
[0002] Technologies have been developed to introduce non-natural amino acids (ncAAs) into the amino acid sequence of proteins, other than the 20 natural amino acids, for the purpose of modifying protein function, etc. For example, techniques have been reported for expressing proteins containing ncAAs in their amino acid sequence (ncAA-containing proteins) in hosts such as Escherichia coli, S. cerevisiae, or mammalian cells, using orthogonal pairs of tRNA and aminoacyl-tRNA synthetase (aaRS) corresponding to the ncAA (Non-Patent Documents 1 to 3).
[0003] However, no technology is known for producing ncAA-containing proteins using coryneform bacteria.
[0004] Liu CC, Schultz PG. Adding new chemistries to the genetic code. Annu Rev Biochem. 2010;79:413-44.Jason W Chin et. al., Addition of p-azido-L-phenylalanine to the genetic code of Escherichia coli, J Am Chem Soc. 2002 Aug 7;124(31):9026-7.Wei Wan et. al., Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochim Biophys Acta. 2014 Jun;1844(6):1059-70.
[0005] An objective of the present invention is to provide a method for secretory production of proteins containing unnatural amino acids (ncAAs).
[0006] As a result of intensive research to solve the above-mentioned problems, the present inventors discovered that ncAA-containing proteins can be secreted and produced by using coryneform bacteria modified to express an orthogonal pair of tRNA and aminoacyl-tRNA synthetase (aaRS) corresponding to an unnatural amino acid (ncAA), thereby completing the present invention.
[0007] That is, the present invention can be exemplified as follows: [1] A method for producing an unnatural amino acid-containing protein, comprising culturing a coryneform bacterium having a gene construct for secretory expression of the unnatural amino acid-containing protein in a medium containing the unnatural amino acid, and recovering the secreted and produced unnatural amino acid-containing protein, wherein the coryneform bacterium has been modified to express an orthogonal pair of a tRNA and an aminoacyl-tRNA synthetase corresponding to the unnatural amino acid. [2] The method, wherein the gene construct comprises, from 5' to 3', a promoter sequence that functions in coryneform bacteria, a nucleic acid sequence encoding a signal peptide that functions in coryneform bacteria, and a nucleic acid sequence encoding an unnatural amino acid-containing protein, and the unnatural amino acid-containing protein is expressed as a fusion protein with the signal peptide. [3] The method, wherein the unnatural amino acid is encoded by a stop codon or a four-residue codon. [4] The method, wherein the stop codon is UAG or UGA. [5] The method, wherein the unnatural amino acid is a tyrosine derivative or a lysine derivative. [6] The unnatural amino acid is p-azido-L-phenylalanine, 3-azido-L-tyrosine, 3-chloro-L-tyrosine, 3-nitro-L-tyrosine, O-sulfo-L-tyrosine, L-pyrrolidine, or N δ[7] The method described above, wherein the tRNA is tRNA(Tyr) or tRNA(Pyl). [8] The method described above, wherein the tRNA(Tyr) is RNA described in the following (a), (b), or (c): (a) RNA comprising the nucleotide sequence set forth in SEQ ID NO: 42 or 44; (b) RNA comprising the nucleotide sequence set forth in SEQ ID NO: 40, 42, or 44 in which the anticodon has been modified; (c) RNA comprising a nucleotide sequence having 90% or more identity to the nucleotide sequence of the RNA described in (a) or (b), and which functions as a tRNA corresponding to the unnatural amino acid. [9] The method, wherein the tRNA(Pyl) is an RNA described in the following (a), (b), or (c): (a) an RNA comprising the nucleotide sequence set forth in SEQ ID NO: 46, 119, or 121; (b) an RNA comprising the nucleotide sequence set forth in SEQ ID NO: 46, 119, or 121 in which the anticodon has been modified; or (c) an RNA comprising a nucleotide sequence having 90% or more identity to the nucleotide sequence of the RNA described in (a) or (b), and which functions as a tRNA corresponding to the unnatural amino acid.
[10] The method, wherein the aminoacyl-tRNA synthetase is tyrosyl-tRNA synthetase or pyrrolysyl-tRNA synthetase.
[11] The method as described above, wherein the tyrosyl-tRNA synthetase is a protein described in the following (a), (b), (c), or (d): (a) a protein comprising the amino acid sequence set forth in SEQ ID NO: 50 or 52; (b) a protein comprising an amino acid sequence set forth in SEQ ID NO: 48, 50, or 52, which has a mutation that alters substrate specificity, and which has aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (c) a protein comprising the amino acid sequence of the protein described in (a) or (b) above, but which comprises substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues; (d) a protein comprising an amino acid sequence that is 90% or more identical to the amino acid sequence of the protein described in (a) or (b).
[12] The method as mentioned above, wherein the mutation that alters substrate specificity comprises a mutation in one or more amino acid residues selected from the following: Y32, H70, E107, D158, I159, L162, and D286.
[13] The method as described above, wherein the pyrrolysyl-tRNA synthetase is a protein described in the following (a), (b), (c), or (d): (a) a protein comprising the amino acid sequence set forth in SEQ ID NO: 54 or 115; (b) a protein comprising an amino acid sequence having a mutation in the amino acid sequence set forth in SEQ ID NO: 54 or 115 that alters substrate specificity, and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (c) a protein comprising an amino acid sequence of the protein described in (a) or (b) that includes a substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues, and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (d) a protein comprising an amino acid sequence that is 90% or more identical to the amino acid sequence of the protein described in (a) or (b), and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid.
[14] The method described above, wherein the mutation that alters substrate specificity comprises a mutation in one or more amino acid residues selected from the following: M241, L266, A267, L270, Y271, L274, N311, C313, M315, Y349, V367, and W383.
[15] The method described above, wherein the coryneform bacterium has been further modified to harbor a phoS gene encoding a mutant PhoS protein.
[16] The method described above, wherein the mutation is a substitution of the amino acid residue corresponding to the tryptophan residue at position 302 of SEQ ID NO: 2 in the wild-type PhoS protein with an amino acid residue other than an aromatic amino acid or histidine.
[17] The method described above, wherein the amino acid residue other than an aromatic amino acid or histidine is a lysine residue, an alanine residue, a valine residue, a serine residue, a cysteine residue, a methionine residue, an aspartic acid residue, or an asparagine residue.
[18] The method, wherein the wild-type PhoS protein is a protein described in (a), (b), or (c) below: (a) a protein comprising the amino acid sequence set forth in any one of SEQ ID NOs: 2 to 7; (b) a protein comprising the amino acid sequence set forth in any one of SEQ ID NOs: 2 to 7, with substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues, and having a function as a sensor kinase of the PhoRS system; (c) a protein comprising an amino acid sequence having 90% or more identity to the amino acid sequence set forth in any one of SEQ ID NOs: 2 to 7, and having a function as a sensor kinase of the PhoRS system.
[19] The method, wherein the signal peptide is a Tat system-dependent signal peptide.
[20] The method, wherein the Tat system-dependent signal peptide is any one signal peptide selected from the group consisting of TorA signal peptide, SufI signal peptide, PhoD signal peptide, LipA signal peptide, and IMD signal peptide.
[21] The method described above, wherein the coryneform bacterium has been further modified to increase expression of one or more genes selected from genes encoding the Tat system secretion apparatus compared to an unmodified strain.
[22] The method described above, wherein the genes encoding the Tat system secretion apparatus consist of the tatA gene, the tatB gene, the tatC gene, and the tatE gene.
[23] The method described above, wherein the signal peptide is a Sec system-dependent signal peptide.
[24] The method described above, wherein the Sec system-dependent signal peptide is any one signal peptide selected from the group consisting of the PS1 signal peptide, the PS2 signal peptide, and the SlpA signal peptide.
[25] The method described above, wherein the gene construct further comprises, between the nucleic acid sequence encoding the signal peptide functional in coryneform bacteria and the nucleic acid sequence encoding the unnatural amino acid-containing protein, a nucleic acid sequence encoding an amino acid sequence comprising Gln-Glu-Thr.
[26] The method described above, wherein the gene construct further comprises, between the nucleic acid sequence encoding the amino acid sequence comprising Gln-Glu-Thr and the nucleic acid sequence encoding the unnatural amino acid-containing protein, a nucleic acid sequence encoding an amino acid sequence used for enzymatic cleavage.
[27] The method, wherein the coryneform bacterium is a bacterium of the genus Corynebacterium.
[28] The method, wherein the coryneform bacterium is Corynebacterium glutamicum.
[29] The method, wherein the coryneform bacterium is a modified strain derived from Corynebacterium glutamicum AJ12036 (FERM BP-734) or a modified strain derived from Corynebacterium glutamicum ATCC 13869.
[30] The method, wherein the coryneform bacterium is a coryneform bacterium in which the number of cell surface protein molecules per cell is reduced compared to an unmodified strain.
[31] The method, wherein the coryneform bacterium has a first expression vector carrying the gene construct and a second expression vector carrying a gene encoding the tRNA and a gene encoding the aminoacyl-tRNA synthetase.
[32] The method, wherein the first expression vector further carries a gene encoding the tRNA and / or a gene encoding the aminoacyl-tRNA synthetase.
[33] The method described above, wherein the first expression vector is a pPK-based vector and the second expression vector is a pVC-based vector.
[34] The method described above, wherein the first expression vector is pPK4 or pPK5 and the second expression vector is pVC7 or pVC7N.
[35] The method described above, wherein the coryneform bacterium has a single expression vector carrying the gene construct, the gene encoding the tRNA, and the gene encoding the aminoacyl-tRNA synthetase.
[36] The method described above, wherein the expression vector is a pPK-based vector.
[37] The method described above, wherein the expression vector is pPK4 or pPK5.
[38] The method described above, wherein the unnatural amino acid-containing protein is an antibody-related molecule, an antibody mimetic, or a physiologically active protein.
[39] The method described above, wherein the unnatural amino acid-containing protein is a VHH fragment, the Z domain of protein A, a fluorescent protein, or a growth factor.
[0008] Figure showing an example of an ncAA. Figure showing the nucleotide sequence (SEQ ID NO: 57) of the gene encoding the anti-epidermal growth factor receptor (EGFR) VHH antibody 9g8 and the amino acid sequence of 9g8 (SEQ ID NO: 58). The ncAA insertion site and the corresponding triplet are indicated in bold italics. Figure (photograph) showing the structure of the expression construct for the AzF-introduced 9g8 mutant in C. glutamicum as an expression host and the results of SDS PAGE of the culture supernatant. Panel (A) shows the results obtained from culture in the absence of AzF, and panel (B) shows the results obtained from culture in the presence of 0.3 [mM] AzF. Lane 1, molecular weight marker; Lane 2, control strain PC (strain transfected with wild-type 9g8 vector); Lane 3, control strain WT (strain transfected with AzFN3 vector and wild-type 9g8 + AzFRS vector); Lanes 4-14, mutant 9g8-expressing strains (numbers (32-116) indicate the site of modification). The arrow indicates the position corresponding to the full length of 9g8. Diagram showing the nucleotide sequence (SEQ ID NO: 83) of the gene encoding the anti-Human Epidermal Growth Factor Receptor 2 (HER2) antibody ZHER2 affibody and the amino acid sequence (SEQ ID NO: 84) of the ZHER2 affibody. The ncAA insertion site and the corresponding triplet are indicated in bold italics. Diagram (photo) showing the structure of the expression construct for the AzF-introduced ZHER2 affibody mutant in C. glutamicum as an expression host and the results of SDS-PAGE of the culture supernatant. Panel (A) shows the results obtained when cultured in the absence of AzF, and panel (B) shows the results obtained when cultured in the presence of 0.3 mM AzF. Lane 1, molecular weight marker; Lane 2, control strain PC (strain transfected with wild-type ZHER2 affibody vector); Lane 3, control strain WT (strain transfected with AzFN3 vector and wild-type ZHER2 affibody + AzFRS vector); Lanes 4-11, mutant ZHER2 affibody-expressing strains (F7, W16, P22, and Y37 indicate the sites of modification). Arrows indicate the positions corresponding to the full-length ZHER2 affibody.A diagram showing the nucleotide sequence (SEQ ID NO: 94) of the gene encoding monomeric red fluorescent protein (mRFP) and the amino acid sequence of mRFP (SEQ ID NO: 95). The ncAA insertion site and the corresponding triplet are shown in bold italics. A diagram (photograph) showing the structure of the expression construct for the AzF-introduced mRFP mutant using C. glutamicum as the expression host and the results of SDS PAGE of the culture supernatant. "No AzF" indicates the results obtained from culture in the absence of AzF, and "0.3 [mM] AzF" indicates the results obtained from culture in the presence of 0.3 [mM] AzF. Lanes 1 and 6, control strain WT + tRNA (AzFN3 vector and wild-type mRFP + tRNA). CTA Vector-transfected strains); Lanes 2 and 7, AzFN3 vector and mutant mRFP + tRNA CTAStrains transfected with the vector; Lanes 3 and 8, control strain WT+AzFRS (strain transfected with the AzFN3 vector and wild-type mRFP + AzFRS vector); Lanes 4 and 9, strains transfected with the AzFN3 vector and mutant mRFP + AzFRS vector; Lane 5, molecular weight marker. The arrow indicates the position corresponding to the full-length mRFP. Figure (photo) shows the structure of the expression construct for the AzF-introduced mRFP mutant in C. glutamicum as an expression host, and the results of SDS PAGE of the culture supernatant. "No AzF" indicates the results obtained from culture in the absence of AzF, and "0.3 [mM] AzF" indicates the results obtained from culture in the presence of 0.3 [mM] AzF. Lanes 1-3 and 7-9: WT control strain (strain transfected with AzFN3 vector and wild-type mRFP vector); Lanes 4-6 and 10-12: mutant mRFP-expressing strain; Lane 13: molecular weight marker. The arrow indicates the position corresponding to the full-length mRFP. Figure (photo) shows the structure of the expression construct for the AzF-introduced mRFP mutant in C. glutamicum as an expression host, and the results of SDS PAGE of the culture supernatant. "No AzF" indicates the results obtained from culture in the absence of AzF, and "0.3 [mM] AzF" indicates the results obtained from culture in the presence of 0.3 [mM] AzF. Lanes 1 and 4, control strain WT (strain introduced with pVC7T7pol1 vector and wild-type mRFP + AzFN3 vector); Lanes 2-3 and 5-6, mutant mRFP-expressing strains (numbers (36 or 80) indicate the modification sites); Lane 7, molecular weight marker. The arrow indicates the position corresponding to the full-length mRFP. Diagram showing the nucleotide sequence (SEQ ID NO: 103) of the gene encoding the anti-Izumo protein 1 N-terminal extracellular domain (NDOM) VHH antibody N15 and the amino acid sequence of N15 (SEQ ID NO: 104). The ncAA insertion site and the corresponding triplet are indicated in bold italics. Diagram (photograph) showing the structure of the expression construct for the ClY-introduced N15 mutant using C. glutamicum as an expression host and the results of SDS PAGE of the culture supernatant."No ClY" indicates results obtained in culture without ClY, and "1 [mM] ClY" indicates results obtained in culture with 1 [mM] ClY. Lanes 1 and 7, control strain PC (strain transfected with wild-type N15 vector); Lanes 2 and 8, control strain WT (strain transfected with pVC7T7pol1 vector and wild-type N15 + IYN3 vector); Lanes 3-5 and 9-11, mutant N15 expression strains (numbers (60, 81, or 96) indicate the site of modification); Lane 6, molecular weight marker. The arrow indicates the position corresponding to the full-length N15. Diagram (photo) showing the structure of the expression construct for the AzY-introduced N15 mutant in C. glutamicum as an expression host and the results of SDS PAGE of the culture supernatant. "No AzY" indicates results obtained in culture without AzY, and "0.3 [mM] AzY" indicates results obtained in culture with 0.3 [mM] AzY. Lanes 1 and 7, control strain WT (strain transfected with pVC7T7pol1 vector and wild-type N15 + IYN3 vector); Lanes 2-4 and 7-9, mutant N15-expressing strains (numbers (60, 81, or 96) indicate the site of modification); Lane 5, molecular weight marker. The arrow indicates the position corresponding to the full-length N15. Figure (photo) shows the structure of the expression construct for the AllocLys-introduced 9g8 mutant in C. glutamicum as an expression host, and the results of SDS-PAGE of the culture supernatant. "-" indicates results obtained in culture without AllocLys, and "+" indicates results obtained in culture with 1 [mM] AllocLys. Lanes 1 and 2, control strain PC (strain transfected with wild-type 9g8 vector); Lanes 3 and 12, molecular weight markers; Lanes 4, 5, 10, and 11, control strain WT (strain transfected with PylRS + wild-type 9g8 + tRNA(Pyl)CTA vector); Lanes 6 and 7, mutant 9g8-expressing strain (9g8 (Y32TAG) and tRNA. CTA combination); Lanes 8 and 9, mutant 9g8 expression strain (9g8 (Y107TGA) and tRNA TCALanes 13 and 14, mutant 9g8 expressing strains (9g8 (Y32TAG) and tRNA TCA (A combination of E. coli and wild-type 9g8 vectors). The arrow indicates the position corresponding to the full-length 9g8. Figure (photograph) shows the structure of the expression construct for the AzF-introduced 9g8 mutant in E. coli as an expression host, as well as the results of SDS PAGE and Western blotting of the culture supernatant. Panel (A) shows the results of SDS PAGE, and panel (B) shows the results of Western blotting. Lane 1, molecular weight marker; Lanes 2, 4, 6, and 8, control strain WT (a strain transfected with the E. coli AzFN3 vector and the E. coli wild-type 9g8 vector); Lanes 3, 5, 7, and 9, mutant 9g8 expression strain (a strain transfected with the E. coli AzFN3 vector and the E. coli mutant 9g8 vector). The arrow indicates the position corresponding to the full-length 9g8. Figure (photograph) shows the structure of the expression construct for the AzF-introduced 9g8 mutant in C. glutamicum as an expression host, as well as the results of SDS PAGE of the culture supernatant. Lane 1, molecular weight marker; Lanes 2, 4, 6, and 8, control strain WT (AzFRS + wild-type 9g8 + tRNA CTA Vector-transfected strains); Lanes 3, 5, 7, and 9, mutant 9g8 expression strain (AzFRS + mutant 9g8 + tRNA CTA(Strains transfected with the vector). The arrow indicates the position corresponding to the full-length of 9g8. Diagram showing the nucleotide sequences of the wild-type and mutant aMD4dY-PA22 genes and the alignment of wild-type and mutant aMD4dY-PA22. The triplet encoding AzF is shown in bold italics. Diagram (photograph) showing the structure of the expression construct for the AzF-introduced aMD4dY-PA22 mutant in C. glutamicum and the results of SDS PAGE of the culture supernatant. The arrow indicates the position corresponding to the full-length of aMD4dY-PA22. Diagram (photograph) showing the nucleotide sequences of the wild-type and mutant EPO-PA22 genes and the alignment of wild-type and mutant EPO-PA22. The triplet encoding AzF is shown in bold italics. Diagram (photograph) showing the structure of the expression construct for the AzF-introduced EPO-PA22 mutant in C. glutamicum and the results of SDS PAGE of the culture supernatant. The arrow indicates the position corresponding to the full-length of EPO-PA22. Figure showing the results of PEG modification of wild-type and mutant aMD4dY-PA22 by strain-promoted alkyne azide cycloaddition (SPAAC) reaction. Figure showing the results of PEG modification of wild-type and mutant EPO-PA22 by strain-promoted alkyne azide cycloaddition (SPAAC) reaction. Figure showing the results of reporter assay of PEG-modified and non-PEG-modified wild-type and mutant aMD4dY-PA22. Figure showing the results of reporter assay of PEG-modified and non-PEG-modified wild-type and mutant EPO-PA22.
[0009] The present invention will be described in detail below.
[0010] The method of the present invention is a method for secretory production of proteins containing unnatural amino acids (ncAAs) using coryneform bacteria.
[0011] Specifically, the method of the present invention may be a method for producing an unnatural amino acid-containing protein, comprising: culturing a coryneform bacterium having a gene construct for secretory expression of the unnatural amino acid-containing protein; and recovering the secreted and produced unnatural amino acid-containing protein, wherein the coryneform bacterium has been modified to express an orthogonal pair of tRNA corresponding to the unnatural amino acid and an aminoacyl-tRNA synthetase.
[0012] The coryneform bacterium (i.e., the coryneform bacterium used in the method of the present invention) is also referred to as the "bacterium of the present invention" or the "coryneform bacterium of the present invention." The bacterium of the present invention or a parent strain used to construct the bacterium is also referred to as the "host."
[0013] The above gene construct (i.e., a gene construct for secretory expression of an ncAA-containing protein) is also referred to as a "gene construct for secretory expression."
[0014] The above pair (i.e., an orthogonal pair of a tRNA corresponding to an ncAA and an aminoacyl-tRNA synthetase (aaRS)) is also called an "orthogonal tRNA(ncAA) / ncAA-aaRS pair."
[0015] <1> Coryneform Bacteria of the Present Invention The coryneform bacterium of the present invention is a coryneform bacterium having a gene construct for secretory expression of an ncAA-containing protein (i.e., a gene construct for secretory expression), and is a coryneform bacterium modified to express an orthogonal pair of tRNA and aaRS corresponding to the ncAA (i.e., an orthogonal tRNA(ncAA) / ncAA-aaRS pair).
[0016] <1-1> Coryneform Bacteria Capable of Secreting and Producing ncAA-Containing Proteins The coryneform bacteria of the present invention are capable of secreting and producing ncAA-containing proteins. The coryneform bacteria of the present invention are capable of secreting and producing ncAA-containing proteins based on at least a combination of having a gene construct for secretory expression and expressing an orthogonal tRNA(ncAA) / ncAA-aaRS pair. Specifically, the coryneform bacteria of the present invention may be capable of secreting and producing ncAA-containing proteins through a combination of having a gene construct for secretory expression and expressing an orthogonal tRNA(ncAA) / ncAA-aaRS pair, or through a combination of having a gene construct for secretory expression, expressing an orthogonal tRNA(ncAA) / ncAA-aaRS pair, and other properties.
[0017] "Secretion" of a protein refers to the transfer of the protein outside the bacterial cell (extracellularly). Examples of "extracellular" include the medium and the bacterial cell surface. That is, the secreted protein molecule may be present, for example, in the medium, on the bacterial cell surface, or both in the medium and on the bacterial cell surface. That is, "secretion" of a protein is not limited to the case where all of the protein molecules are ultimately placed in a completely free state in the medium, but also includes, for example, the case where all of the protein molecules are present on the bacterial cell surface, or the case where some of the protein molecules are present in the medium and the remaining molecules are present on the bacterial cell surface.
[0018] Specifically, the "ability to secrete and produce an ncAA-containing protein" refers to the ability of the bacterium of the present invention to secrete the ncAA-containing protein into the medium and / or onto the cell surface when cultured in a medium, and to accumulate the ncAA-containing protein to an extent that it can be recovered from the medium and / or the cell surface. The amount of accumulation in the medium may be, for example, preferably 10 μg / L or more, more preferably 1 mg / L or more, particularly preferably 100 mg / L or more, and even more preferably 1 g / L or more. Furthermore, the amount of accumulation on the cell surface may be, for example, such that when the ncAA-containing protein from the cell surface is recovered and suspended in an equal volume of liquid as the medium, the concentration of the ncAA-containing protein in the suspension is preferably 10 μg / L or more, more preferably 1 mg / L or more, and particularly preferably 100 mg / L or more.
[0019] Coryneform bacteria are aerobic, gram-positive rod-shaped bacteria. Examples of coryneform bacteria include bacteria of the genera Corynebacterium, Brevibacterium, and Microbacterium. The advantages of using coryneform bacteria include the fact that, compared with filamentous fungi, yeast, Bacillus bacteria, and the like that have traditionally been used for the secretory production of proteins, very little protein is secreted outside the bacterial cells, which is expected to simplify or eliminate the purification process when proteins are secreted and produced. Furthermore, they grow well in simple media containing sugars, ammonia, inorganic salts, and the like, and are therefore excellent in terms of medium cost, culture method, and culture productivity.
[0020] Specific examples of coryneform bacteria include the following species: Corynebacterium acetoacidophilum Corynebacterium acetoglutamicum Corynebacterium alkanolyticum Corynebacterium callunae Corynebacterium crenatum Corynebacterium glutamicum Corynebacterium lilium Corynebacterium melassecola Corynebacterium thermoaminogenes (Corynebacterium efficiens) Corynebacterium herculis Brevibacterium divaricatum (Corynebacterium glutamicum) Brevibacterium flavum (Corynebacterium glutamicum) Brevibacterium immariophilum Brevibacterium lactofermentum (Corynebacterium glutamicum) Brevibacterium roseum Brevibacterium saccharolyticum Brevibacterium thiogenitalisCorynebacterium ammoniagenes (Corynebacterium stationis) Brevibacterium album Brevibacterium cerinum Microbacterium ammoniaphilum
[0021] Specific examples of coryneform bacteria include the following strains: Corynebacterium acetoacidophilum ATCC 13870, Corynebacterium acetoglutamicum ATCC 15806, Corynebacterium alkanolyticum ATCC 21511, Corynebacterium callunae ATCC 15991, and Corynebacterium crenatum AS1.542 Corynebacterium glutamicum ATCC 13020, ATCC 13032, ATCC 13060, ATCC 13869, FERM BP-734 Corynebacterium lilium ATCC 15990 Corynebacterium melassecola ATCC 17965 Corynebacterium efficiens (Corynebacterium thermoaminogenes) AJ12340 (FERM BP-1539) Corynebacterium herculis ATCC 13868 Brevibacterium divaricatum (Corynebacterium glutamicum) ATCC 14020 Brevibacterium flavum (Corynebacterium glutamicum) ATCC 13826, ATCC 14067, AJ12418 (FERM BP-2205) Brevibacterium immariophilum ATCC 14068 Brevibacterium lactofermentum (Corynebacterium glutamicum) ATCC 13869 Brevibacterium roseum ATCC 13825 Brevibacterium saccharolyticum ATCC 14066 Brevibacterium thiogenitalis ATCC 19240 Corynebacterium ammoniagenes (Corynebacterium stationis) ATCC 6871, ATCC 6872 Brevibacterium album ATCC 15111 Brevibacterium cerinum ATCC 15112 Microbacterium ammoniaphilum ATCC 15354.
[0022] The genus Corynebacterium also includes bacteria that were previously classified as Brevibacterium but have now been integrated into the genus Corynebacterium (Int. J. Syst. Bacteriol., 41, 255(1991)). Corynebacterium stationis also includes bacteria that were previously classified as Corynebacterium ammoniagenes but have been reclassified as Corynebacterium stationis based on 16S rRNA sequence analysis and other factors (Int. J. Syst. Evol. Microbiol., 60, 874-879(2010)).
[0023] These strains can be obtained, for example, from the American Type Culture Collection (address: 12301 Parklawn Drive, Rockville, Maryland 20852, PO Box 1549, Manassas, VA 20108, United States of America). Each strain is assigned a corresponding accession number, and can be obtained using this accession number (see http: / / www.atcc.org / ). The accession numbers corresponding to each strain are listed in the catalog of the American Type Culture Collection. These strains can also be obtained, for example, from the depository institution where they were deposited.
[0024] In particular, C. glutamicum AJ12036 (FERM BP-734), a streptomycin (Sm)-resistant mutant isolated from wild-type C. glutamicum ATCC 13869, is predicted to have a mutation in a gene controlling protein secretion functions compared to its parent (wild-type) strain. This mutation is believed to have significantly higher protein secretion and production capacity, with the amount of protein accumulated under optimal culture conditions being approximately two to three times higher than that of the wild-type strain, making it a suitable host bacterium (WO 2002 / 081694). AJ12036 was originally deposited as an international deposit with the Fermentation Research Institute, Agency of Industrial Science and Technology (now the National Institute of Technology and Evaluation, Patent Organism Depositary Center, Room 120, 2-5-8 Kazusa Kamatari, Kisarazu City, Chiba Prefecture, Japan, postal code: 292-0818) on March 26, 1984, and has been assigned the accession number FERM BP-734.
[0025] Corynebacterium thermoaminogenes AJ12340 (FERM BP-1539) was originally deposited as an international deposit at the Fermentation Research Institute, Agency of Industrial Science and Technology (currently the Patent Organism Deposit Center, National Institute of Technology and Evaluation, Japan; postal code: 292-0818; address: Room 120, 2-5-8 Kazusa Kamatari, Kisarazu City, Chiba Prefecture, Japan) on March 13, 1987, and has been assigned the accession number FERM BP-1539. Brevibacterium flavum AJ12418 (FERM BP-2205) was originally deposited as an international deposit at the Fermentation Research Institute, Agency of Industrial Science and Technology (currently the Patent Organism Deposit Center, National Institute of Technology and Evaluation, Independent Administrative Institution, Postal Code: 292-0818, Address: Room 120, 2-5-8 Kazusa Kamatari, Kisarazu City, Chiba Prefecture, Japan) on December 24, 1988, and has been assigned the accession number FERM BP-2205.
[0026] Alternatively, a strain with enhanced protein secretory production ability may be selected using mutation or genetic recombination techniques from the above-described coryneform bacteria as a parent strain and used as a host. For example, strains with enhanced protein secretory production ability can be selected after treatment with ultraviolet light or a chemical mutagen such as N-methyl-N'-nitrosoguanidine.
[0027] Furthermore, using such a strain modified to not produce cell surface proteins as a host is particularly preferred, as it facilitates the purification of ncAA-containing proteins secreted into the medium or onto the bacterial cell surface. Such modifications can be achieved by introducing mutations into the coding region of the cell surface protein or its expression regulatory region on the chromosome using mutagenesis or genetic recombination techniques. An example of a coryneform bacterium modified to not produce cell surface proteins is the C. glutamicum YDK010 strain (WO 2002 / 081694), which is a strain of C. glutamicum AJ12036 (FERM BP-734) that lacks the cell surface protein PS2.
[0028] The bacterium of the present invention can be obtained by appropriately modifying the coryneform bacterium described above (e.g., by introducing a gene construct for secretory expression, by introducing a gene encoding an orthogonal tRNA(ncAA) / ncAA-aaRS pair, and optionally by introducing other modifications). That is, the bacterium of the present invention may be, for example, a modified strain derived from the coryneform bacterium described above. Specifically, the bacterium of the present invention may be, for example, a modified strain derived from C. glutamicum AJ12036 (FERM BP-734) or a modified strain derived from C. glutamicum ATCC 13869. Note that a modified strain derived from C. glutamicum AJ12036 (FERM BP-734) also corresponds to a modified strain derived from C. glutamicum ATCC 13869. The modifications for constructing the bacterium of the present invention can be performed in any order.
[0029] <1-2> Gene construct for secretory expression of ncAA-containing protein and introduction thereof The coryneform bacterium of the present invention has a gene construct for secretory expression of an ncAA-containing protein (i.e., a gene construct for secretory expression).
[0030] Secretory proteins are generally translated as preproteins (also called prepeptides) or preproproteins (also called prepropeptides), which are then processed to mature proteins. Specifically, secretory proteins are generally translated as preproteins or preproproteins, and then the signal peptide (preprotein) is cleaved by a protease (commonly called a signal peptidase) to convert them into mature proteins or proproteins. The proprotein is then further cleaved by a protease to convert it into a mature protein. Therefore, in the present invention, signal peptides are used for the secretory production of ncAA-containing proteins. The preproteins and preproproteins of secretory proteins are sometimes collectively referred to as "secretory protein precursors." A "signal peptide" (also called a "signal sequence") refers to an amino acid sequence present at the N-terminus of a secretory protein precursor that is not typically present in the native mature protein.
[0031] The gene construct for secretory expression comprises, from 5' to 3', a promoter sequence functional in coryneform bacteria, a nucleic acid sequence encoding a signal peptide functional in coryneform bacteria, and a nucleic acid sequence encoding an ncAA-containing protein. The nucleic acid sequence encoding the signal peptide may be linked downstream of the promoter sequence so that the signal peptide is expressed under the control of the promoter. The nucleic acid sequence encoding the ncAA-containing protein may be linked downstream of the nucleic acid sequence encoding the signal peptide so that the ncAA-containing protein is expressed as a fusion protein with the signal peptide. Such a fusion protein is also referred to as the "fusion protein of the present invention." Note that in the fusion protein of the present invention, the signal peptide and the ncAA-containing protein may or may not be adjacent to each other. In other words, "the ncAA-containing protein is expressed as a fusion protein with a signal peptide" does not necessarily mean that the ncAA-containing protein is expressed adjacent to the signal peptide as a fusion protein with the signal peptide, but also includes cases where the ncAA-containing protein is expressed as a fusion protein with the signal peptide via another amino acid sequence. For example, as described below, the fusion protein of the present invention may contain an inserted sequence between the signal peptide and the ncAA-containing protein, such as an amino acid sequence containing Gln-Glu-Thr or an amino acid sequence used for enzymatic cleavage. Furthermore, as described below, the final ncAA-containing protein does not need to have a signal peptide. In other words, "the ncAA-containing protein is expressed as a fusion protein with a signal peptide" means that the ncAA-containing protein forms a fusion protein with the signal peptide upon expression; the final ncAA-containing protein does not necessarily form a fusion protein with the signal peptide. The term "nucleic acid sequence" may also be interpreted as "gene." For example, a nucleic acid sequence encoding an ncAA-containing protein is also referred to as an "ncAA-containing protein-encoding gene" or an "ncAA-containing protein gene." Examples of nucleic acid sequences include DNA.Furthermore, the gene construct for secretory expression may have control sequences (such as an operator, SD sequence, or terminator) that are effective for expressing the fusion protein of the present invention in coryneform bacteria, at appropriate positions so that they can function.
[0032] The promoter used in the present invention is not particularly limited as long as it is a promoter that functions in coryneform bacteria. A "promoter that functions in coryneform bacteria" refers to a promoter that has promoter activity (i.e., gene transcription activity) in coryneform bacteria. Examples of promoters that function in coryneform bacteria include those described below in the section "Methods for increasing protein activity."
[0033] The signal peptide used in the present invention is not particularly limited as long as it functions in coryneform bacteria. The signal peptide may be a signal peptide derived from coryneform bacteria (e.g., derived from a host) or from a heterologous species. The signal peptide may be a signal peptide inherent to the ncAA-containing protein or a signal peptide from another protein. A "signal peptide that functions in coryneform bacteria" refers to a peptide that, when linked to the N-terminus of a protein of interest, enables coryneform bacteria to secrete the protein. Whether a signal peptide functions in coryneform bacteria can be confirmed, for example, by fusing the protein of interest with the signal peptide, expressing the protein, and confirming whether the protein is secreted.
[0034] Examples of signal peptides include Tat-dependent signal peptides and Sec-dependent signal peptides.
[0035] The term "Tat system-dependent signal peptide" refers to a signal peptide recognized by the Tat system. Specifically, the "Tat system-dependent signal peptide" may be a peptide that, when linked to the N-terminus of a protein of interest, causes the protein to be secreted by the Tat system secretion apparatus.
[0036] Examples of Tat-dependent signal peptides include the signal peptide of the E. coli TorA protein (trimethylamine-N-oxide reductase), the signal peptide of the E. coli SufI protein (suppressor of ftsI), the signal peptide of the Bacillus subtilis PhoD protein (phosphodiesterase), the signal peptide of the Bacillus subtilis LipA protein (lipoic acid synthase), and the signal peptide of the Arthrobacter globiformis IMD protein (isomaltodextranase). The amino acid sequences of these signal peptides are as follows: TorA signal peptide: MNNNDLFQASRRRFLAQLGGLTVAGMLGPSLLTPRRATA (SEQ ID NO: 18) SufI signal peptide: MSLSRRQFIQASGIALCAGAVPLKASA (SEQ ID NO: 19) PhoD signal peptide: MAYDSRFDEWVQKLKEESFQNNTFDRRKFIQGAGKIAGLSLGLTIAQS (SEQ ID NO: 20) LipA signal peptide: MKFVKRRTTALVTTLMLSVTSLFALQPSAKAAEH (SEQ ID NO: 21) IMD signal peptide: MMNLSRRTLLTTGSAATLAYALGMAGSAQA (SEQ ID NO: 22)
[0037] Tat-dependent signal peptides have a twin-arginine motif, such as S / TRRXFLK (SEQ ID NO: 23) and RRX-#-# (X: naturally occurring amino acid residue, #: hydrophobic amino acid residue).
[0038] The term "Sec system-dependent signal peptide" refers to a signal peptide recognized by the Sec system. Specifically, the term "Sec system-dependent signal peptide" may be a peptide that, when linked to the N-terminus of a protein of interest, causes the protein to be secreted by the Sec system secretion apparatus.
[0039] Examples of Sec system-dependent signal peptides include signal peptides of cell surface proteins of coryneform bacteria. Cell surface proteins of coryneform bacteria are as described above. Examples of cell surface proteins of coryneform bacteria include PS1 and PS2 (CspB) derived from C. glutamicum (JP Patent Publication No. 6-502548) and SlpA (CspA) derived from C. stationis (JP Patent Publication No. 10-108675). The amino acid sequence of the signal peptide of PS1 from C. glutamicum (PS1 signal peptide) is shown in SEQ ID NO: 25, the amino acid sequence of the signal peptide of PS2 (CspB) from C. glutamicum (PS2 signal peptide) is shown in SEQ ID NO: 26, and the amino acid sequence of the signal peptide of SlpA (CspA) from C. stationis (SlpA signal peptide) is shown in SEQ ID NO: 27.
[0040] The Tat-dependent signal peptide may be a variant of the above-exemplified Tat-dependent signal peptide, so long as it has a twin-arginine motif and maintains its original function. The Sec-dependent signal peptide may be a variant of the above-exemplified Sec-dependent signal peptide, so long as it maintains its original function. The description of conservative variants of ncAA-aaRS and ncAA-aaRS genes described below applies mutatis mutandis to variants of signal peptides and genes encoding them. For example, the signal peptide may be a peptide having an amino acid sequence in which one or several amino acids are substituted, deleted, inserted, and / or added at one or several positions in the amino acid sequence of the above-exemplified signal peptide. The term "one or several" in the context of a signal peptide variant specifically refers to preferably 1 to 7, more preferably 1 to 5, even more preferably 1 to 3, and particularly preferably 1 to 2. The terms "TorA signal peptide," "SufI signal peptide," "PhoD signal peptide," "LipA signal peptide," "IMD signal peptide," "PS1 signal peptide," "PS2 signal peptide," and "SlpA signal peptide" are intended to encompass the peptides set forth in SEQ ID NOs: 18 to 22 and 25 to 27, respectively, as well as conservative variants thereof.
[0041] With regard to a Tat-dependent signal peptide, "maintaining its original function" means that it is recognized by the Tat system, and specifically, that it has the function of secreting the protein of interest via the Tat-dependent secretion apparatus when linked to the N-terminus of the protein. Whether a peptide functions as a Tat-dependent signal peptide can be confirmed, for example, by confirming that the secretory production of a protein to which the peptide has been added at its N-terminus is increased by enhancing the Tat-dependent secretion apparatus, or by confirming that the secretory production of a protein to which the peptide has been added at its N-terminus is decreased by deficiency of the Tat-dependent secretion apparatus.
[0042] With regard to a Sec system-dependent signal peptide, "maintaining its original function" means that it is recognized by the Sec system, and more specifically, it may have the function of secreting a protein of interest by the Sec system secretion apparatus when linked to the N-terminus of the protein. Whether a peptide functions as a Sec system-dependent signal peptide can be confirmed, for example, by confirming that the secretory production of a protein to which the peptide has been added at its N-terminus is increased by enhancing the Sec system secretion apparatus, or by confirming that the secretory production of a protein to which the peptide has been added at its N-terminus is decreased by deficiency of the Sec system secretion apparatus.
[0043] Signal peptides are generally cleaved by signal peptidase when the translation product is secreted outside the bacterial cell. In other words, the final ncAA-containing protein does not need to have a signal peptide. The gene encoding the signal peptide can be used in its native form, but it can also be modified to have optimal codons depending on the codon usage frequency of the host used.
[0044] In the gene construct for secretory expression, a nucleic acid sequence encoding an amino acid sequence containing Gln-Glu-Thr may be inserted between the nucleic acid sequence encoding the signal peptide and the nucleic acid sequence encoding the ncAA-containing protein (WO2013 / 062029). The "amino acid sequence containing Gln-Glu-Thr" is also referred to as the "insertion sequence used in the present invention." Examples of the insertion sequence used in the present invention include the amino acid sequence containing Gln-Glu-Thr described in WO2013 / 062029. The insertion sequence used in the present invention is particularly suitable for use in combination with a Sec-dependent signal peptide.
[0045] The insertion sequence used in the present invention is preferably a sequence consisting of three or more amino acid residues from the N-terminus of the mature protein of the cell surface protein CspB of coryneform bacteria (hereinafter also referred to as "mature CspB" or "mature CspB protein"). The "sequence consisting of three or more amino acid residues from the N-terminus" refers to the amino acid sequence from the first amino acid residue at the N-terminus to the third or more amino acid residues.
[0046] The cell surface protein CspB of coryneform bacteria is described below. Specific examples of CspB include CspB from C. glutamicum ATCC13869, CspB from the 28 strains of C. glutamicum described below, and variants thereof. In the amino acid sequence of C. glutamicum ATCC13869 CspB shown in SEQ ID NO: 11, amino acid residues 1 to 30 correspond to the signal peptide, and amino acid residues 31 to 499 correspond to the mature CspB protein. The amino acid sequence of the mature CspB protein from C. glutamicum ATCC13869 excluding the 30 amino acid residues in the signal peptide portion is shown in SEQ ID NO: 28. In the mature CspB from C. glutamicum ATCC13869, amino acid residues 1 to 3 at the N-terminus correspond to Gln-Glu-Thr.
[0047] The insertion sequence used in the present invention is preferably an amino acid sequence extending from the amino acid residue at position 1 to any one of amino acid residues 3 to 50 of mature CspB. The insertion sequence used in the present invention is more preferably an amino acid sequence extending from the amino acid residue at position 1 to any one of amino acid residues 3 to 8, 17, or 50 of mature CspB. The insertion sequence used in the present invention is particularly preferably an amino acid sequence extending from the amino acid residue at position 1 to any one of amino acid residues 4, 6, 17, or 50 of mature CspB.
[0048] The insertion sequence used in the present invention is preferably an amino acid sequence selected from the group consisting of the amino acid sequences A to H below. (A) Gln-Glu-Thr (B) Gln-Glu-Thr-Xaa1 (C) Gln-Glu-Thr-Xaa1-Xaa2 (D) Gln-Glu-Thr-Xaa1-Xaa2-Xaa3 (E) An amino acid sequence in which amino acid residues at positions 4 to 7 of mature CspB are added to Gln-Glu-Thr (F) An amino acid sequence in which amino acid residues at positions 4 to 8 of mature CspB are added to Gln-Glu-Thr (G) An amino acid sequence in which amino acid residues at positions 4 to 17 of mature CspB are added to Gln-Glu-Thr (H) An amino acid sequence in which amino acid residues at positions 4 to 50 of mature CspB are added to Gln-Glu-Thr In the amino acid sequences A to H, Xaa1 is Asn, Gly, Thr, Pro, or Ala, Xaa2 is Pro, Thr, or Val, and Xaa3 is Thr or Tyr. Furthermore, in the amino acid sequences A to H, "amino acid residues at positions 4 to X of mature CspB are added to Gln-Glu-Thr" means that amino acid residues at positions 4 to X of the N-terminus of mature CspB are added to Thr in Gln-Glu-Thr. Typically, the first to third amino acid residues at the N-terminus of mature CspB are Gln-Glu-Thr, and in that case, "an amino acid sequence in which amino acid residues at positions 4 to X of mature CspB are added to Gln-Glu-Thr" is synonymous with the amino acid sequence consisting of amino acid residues at positions 1 to X of mature CspB.
[0049] Specifically, the insertion sequence used in the present invention is preferably an amino acid sequence selected from the group consisting of Gln-Glu-Thr-Asn-Pro-Thr (SEQ ID NO: 32), Gln-Glu-Thr-Gly-Thr-Tyr (SEQ ID NO: 33), Gln-Glu-Thr-Thr-Val-Thr (SEQ ID NO: 34), Gln-Glu-Thr-Pro-Val-Thr (SEQ ID NO: 35), and Gln-Glu-Thr-Ala-Val-Thr (SEQ ID NO: 36).
[0050] The "amino acid residue at position X in mature CspB" refers to the amino acid residue corresponding to the amino acid residue at position X in SEQ ID NO: 28. In the amino acid sequence of any mature CspB, which amino acid residue is "the amino acid residue corresponding to the amino acid residue at position X in SEQ ID NO: 28" can be determined by aligning the amino acid sequence of any mature CspB with the amino acid sequence of SEQ ID NO: 28.
[0051] "Unnatural amino acid-containing protein (ncAA-containing protein)" refers to a protein that contains an ncAA in its amino acid sequence. When a protein contains an ncAA, it is also referred to as "the protein contains an ncAA residue."
[0052] "Unnatural amino acid (ncAA)" refers to an amino acid other than a natural amino acid. An ncAA may be an L-amino acid or a D-amino acid. An ncAA may specifically be an L-amino acid.
[0053] "Natural amino acids" refers to the following 20 types of amino acids: K (Lys), R (Arg), H (His), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), F (Phe), W (Trp), Y (Tyr), C (Cys), M (Met), D (Asp), E (Glu), N (Asn), and Q (Gln). All natural amino acids (except Gly) are L-amino acids.
[0054] Examples of ncAAs include derivatives of natural amino acids. Examples of derivatives of natural amino acids include tyrosine derivatives and lysine derivatives. The term "derivative of a natural amino acid" refers to a compound having a partially modified structure of a natural amino acid. Examples of structural modifications include substituting a component of a natural amino acid with another component. Examples of such components include atoms and functional groups. For example, in the case of a tyrosine derivative, the components of tyrosine to be substituted include a phenolic hydroxyl group, a phenolic hydrogen atom, and a hydrogen atom on the benzene ring. The components introduced by substitution are not particularly limited as long as they enable the secretion and expression of an ncAA-containing protein. The components introduced by substitution can be appropriately selected depending on various conditions, such as the intended use of the ncAA. Examples of components introduced by substitution include a halogen atom, an azide group, a nitro group, a sulfo group, a hydroxyl group, an alkyl group, an aryl group, an alkoxy group, and an acyl group. Examples of halogen atoms include a fluorine atom, a chlorine atom, a bromine atom, and an iodine atom. Examples of structural modifications include the conversion of an L-amino acid to a D-amino acid. That is, derivatives of natural amino acids also include D-forms of natural amino acids.
[0055] Specifically, ncAAs include p-Azido-L-phenylalanine (AzF), 3-Azido-L-tyrosine (AzY), 3-chloro-L-tyrosine (ClY), 3-nitro-L-tyrosine (NOY), O-sulfo-L-tyrosine (SfY), L-Pyrrolysine (Pyl), N δExamples of compounds that can be used include α-Alloc-L-lysine (AllocLys), a compound described in Liu CC, Schultz PG. Adding new chemistries to the genetic code. Annu Rev Biochem. 2010;79:413-44, and a compound described in Wei Wan et al., Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochim Biophys Acta. 2014 Jun;1844(6):1059-70. The compounds described in Liu CC, Schultz PG. Adding new chemistries to the genetic code. Annu Rev Biochem. 2010;79:413-44 are shown in Figure 1. Figure 1 is a reference to Figure 1 from the same publication. AzF is the same as compound No. 7 in Figure 1. Pyl is the same as compound No. 59 in Figure 1. AllocLys is the same as compound No. 67 in Figure 1. ncAAs include, in particular, AzF, AzY, ClY, NOY, and SfY. ncAAs also include, in particular, AllocLys.
[0056] For example, AzF, AzY, ClY, NOY, SfY, and compounds Nos. 1 to 27, 31, 32, 34 to 36, 41 to 44, 46, and 48 to 50 in FIG. 1 may all be tyrosine derivatives.
[0057] For example, Pyl, AllocLys, and compounds Nos. 28, 33, 51, and 53 to 71 in FIG. 1 may all be lysine derivatives.
[0058] An ncAA-containing protein may contain one type of ncAA residue, or two or more types of ncAA residues. An ncAA-containing protein may contain an ncAA residue at one site, or two or more sites. When an ncAA-containing protein contains ncAA residues at two or more sites, the types of ncAA residues contained at each site may or may not be the same.
[0059] The ncAA-containing protein is not particularly limited other than containing an ncAA. The ncAA-containing protein may be a host-derived protein or a heterologous protein. A "heterologous protein" refers to a protein that is exogenous to the host (i.e., the bacterium of the present invention) that produces the protein. The ncAA-containing protein may be, for example, a microbial, plant, animal, or viral protein, or even a protein with an artificially designed amino acid sequence. The ncAA-containing protein may be, in particular, a human-derived protein. The ncAA-containing protein may be a monomeric or multimeric protein. A multimeric protein is a protein that can exist as a multimer consisting of two or more subunits. In a multimer, the subunits may be linked by covalent bonds such as disulfide bonds, non-covalent bonds such as hydrogen bonds or hydrophobic interactions, or a combination thereof. The multimer preferably contains one or more intermolecular disulfide bonds. The multimer may be a homomultimer consisting of a single type of subunit, or a heteromultimer consisting of two or more types of subunits. When the multimeric protein is a heteromultimer, at least one of the subunits constituting the multimer may be an ncAA-containing protein. That is, all subunits may contain ncAA, or only some of the subunits may contain ncAA. The ncAA-containing protein may be a naturally secreted protein or a naturally non-secreted protein, but is preferably a naturally secreted protein. Furthermore, the ncAA-containing protein may be a naturally secreted protein dependent on the Tat system or a naturally secreted protein dependent on the Sec system.
[0060] The secretory production of an ncAA-containing protein may be limited to only one type, or may be limited to two or more types. Furthermore, when an ncAA-containing protein is a heteromultimer, it is sufficient that only the subunits containing the ncAA are secreted. That is, when an ncAA-containing protein is a heteromultimer, only one type of subunit may be secreted, or two or more types of subunits may be secreted. In other words, "secretory production of an ncAA-containing protein" does not necessarily mean secretory production of all subunits constituting the ncAA-containing protein, but also encompasses secretory production of only some of the subunits.
[0061] Examples of ncAA-containing proteins include enzymes, physiologically active proteins, receptor proteins, antigenic proteins, and any other proteins that contain ncAAs. Note that the term "protein" may also include so-called peptides, such as oligopeptides and polypeptides.
[0062] Examples of the enzyme include cellulase, xylanase, transglutaminase, protein glutaminase, protein asparaginase, isomaltodextranase, protease, endopeptidase, exopeptidase, aminopeptidase, carboxypeptidase, collagenase, chitinase, γ-glutamylvaline synthetase, glutamic acid-cysteine ligase, and glutathione synthetase. Examples of the transglutaminase include secretory transglutaminases from actinomycetes such as Streptoverticillium mobaraense IFO 13819 (WO 01 / 23591), Streptoverticillium cinnamoneum IFO 12852, Streptoverticillium griseocarneum IFO 12776, and Streptomyces lydicus (WO 9606931), and filamentous fungi such as Oomycetes (WO 9622366). Examples of protein glutaminases include protein glutaminase from Chryseobacterium proteolyticum (WO2005 / 103278), and examples of isomaltodextranase include isomaltodextranase from Arthrobacter globiformis (WO2005 / 103278).
[0063] Physiologically active proteins include growth factors, hormones, cytokines, antibody-related molecules, and antibody mimetics.
[0064] Growth factors include epidermal growth factor (EGF), insulin-like growth factor-1 (IGF-1), transforming growth factor (TGF), nerve growth factor (NGF), brain-derived neurotrophic factor (BDNF), vascular endothelial growth factor (VEGF), granulocyte-colony stimulating factor (G-CSF), granulocyte-macrophage-colony stimulating factor (GM-CSF), platelet-derived growth factor (PDGF), erythropoietin (EPO), thrombopoietin (TPO), and acidic fibroblast growth factor (AGF). These include fibroblast growth factor (aFGF or FGF1), basic fibroblast growth factor (bFGF or FGF2), keratinocyte growth factor (KGF-1 or FGF7, KGF-2 or FGF10), hepatocyte growth factor (HGF), stem cell factor (SCF), activin, and peptides that mimic their functions. Activins include activin A, C, and E. Peptides that mimic the functions of the growth factors listed above include aMD4dY-PA22 and EPO-PA22 (WO2021 / 112249).
[0065] Examples of hormones include insulin, glucagon, somatostatin, human growth hormone (hGH), parathyroid hormone (PTH), calcitonin, and exenatide.
[0066] Cytokines include interleukins, interferons, and tumor necrosis factors (TNFs).
[0067] It is not necessary to strictly distinguish between growth factors, hormones, and cytokines. For example, a physiologically active protein may belong to any one group selected from growth factors, hormones, and cytokines, or may belong to multiple groups selected from these groups.
[0068] Furthermore, the physiologically active protein may be the entire protein or a portion thereof. Examples of the portion of the protein include a physiologically active portion. Specific examples of the physiologically active portion include teriparatide, a physiologically active peptide consisting of the N-terminal 34 amino acid residues of the mature form of parathyroid hormone (PTH).
[0069] The term "antibody-related molecule" refers to a protein containing a molecular species consisting of a single domain selected from the domains constituting a complete antibody or a combination of two or more domains. The domains constituting a complete antibody include the heavy chain domains VH, CH1, CH2, and CH3, and the light chain domains VL and CL. An antibody-related molecule may be a monomeric or multimeric protein, as long as it contains the above-mentioned molecular species. When the antibody-related molecule is a multimeric protein, it may be a homomultimer consisting of a single type of subunit, or a heteromultimer consisting of two or more types of subunits. Specific examples of antibody-related molecules include complete antibodies, Fab, F(ab'), F(ab')2, Fc, a dimer consisting of a heavy chain (H chain) and a light chain (L chain), an Fc fusion protein, a heavy chain (H chain), a light chain (L chain), a single-chain Fv (scFv), sc(Fv)2, a disulfide-linked Fv (sdFv), a diabody, and a VHH fragment (Nanobody®). Antibody-related molecules include, in particular, VHH fragments (Nanobody (registered trademark)). More specific examples of antibody-related molecules include trastuzumab, adalimumab, nivolumab, VHH antibody N15, and VHH antibody 9g8.
[0070] "Antibody mimetic" may refer to an organic compound that can specifically bind to an antigen but is not structurally related to an antibody. Specific examples of antibody mimetics include the Z domain of protein A (Affibody). More specific examples of antibody mimetics include ZHER2 affibody.
[0071] Receptor proteins include receptor proteins for physiologically active proteins and other physiologically active substances. Other physiologically active substances include neurotransmitters such as dopamine. Receptor proteins may also be orphan receptors for which the corresponding ligand is unknown.
[0072] The antigen protein is not particularly limited as long as it can induce an immune response. The antigen protein can be appropriately selected depending on, for example, the target of the expected immune response. The antigen protein can be used, for example, as a vaccine.
[0073] Other proteins include liver-type fatty acid-binding protein (LFABP), fluorescent proteins, immunoglobulin-binding proteins, albumin, fibroin-like proteins, and extracellular proteins. Fluorescent proteins include green fluorescent protein (GFP) and monomeric red fluorescent protein (mRFP). Immunoglobulin-binding proteins include protein A, protein G, and protein L. Albumins include human serum albumin. Fibroin-like proteins include those disclosed in WO2017 / 090665 and WO2017 / 171001.
[0074] Extracellular proteins include fibronectin, vitronectin, collagen, osteopontin, laminin, and their partial sequences. Laminin is a heterotrimeric protein consisting of an α chain, a β chain, and a γ chain. Examples of laminin include mammalian laminins. Mammals include primates such as humans, monkeys, and chimpanzees; rodents such as mice, rats, hamsters, and guinea pigs; and various other mammals such as rabbits, horses, cows, sheep, goats, pigs, dogs, and cats. Mammals, in particular, include humans. Laminin subunit chains (i.e., α chains, β chains, and γ chains) include five α chains (α1-α5), three β chains (β1-β3), and three γ chains (γ1-γ3). Laminin forms various isoforms based on the combination of these subunit chains. Specific examples of laminins include laminin 111, laminin 121, laminin 211, laminin 213, laminin 221, laminin 311, laminin 321, laminin 332, laminin 411, laminin 421, laminin 423, laminin 511, laminin 521, and laminin 523. Laminin subsequences include laminin E8, the E8 fragment of laminin. Laminin E8 is a heterotrimeric protein consisting of an α-chain E8 fragment (α-chain E8), a β-chain E8 fragment (β-chain E8), and a γ-chain E8 fragment (γ-chain E8). The subunit chains of laminin E8 (i.e., α-chain E8, β-chain E8, and γ-chain E8) are collectively referred to as the "E8 subunit chain." Examples of E8 subunit chains include E8 fragments of the laminin subunit chains listed above. Laminin E8 is composed of various isoforms depending on the combination of these E8 subunit chains. Specific examples of laminin E8 include laminin 111E8, laminin 121E8, laminin 211E8, laminin 221E8, laminin 332E8, laminin 421E8, laminin 411E8, laminin 511E8, and laminin 521E8.
[0075] The ncAA-containing protein gene is not particularly limited as long as it encodes the above-mentioned ncAA-containing protein.
[0076] The ncAA-containing protein gene contains a codon encoding the ncAA. The codon encoding the ncAA is also referred to as the "ncAA codon." The ncAA codon is not particularly limited as long as it allows the ncAA-containing protein to be secreted and expressed. Examples of the ncAA codon include stop codons and unnatural codons. Examples of the ncAA codon include stop codons in particular. Examples of stop codons include UAA (ochre), UAG (amber), and UGA (opal). Examples of stop codons in particular include UAG (amber) and UGA (opal). Examples of stop codons in particular include UAG (amber). Note that "U" and "T" in the base sequence should be interpreted as appropriate depending on the type of nucleic acid. Examples of unnatural codons include codons of four or more residues in length. Examples of four-residue codons include CGGG and GGGU (Takahiro Hosaka, Development and application of artificial protein synthesis systems using an expanded genetic code, Seibutsubutsu 47(2), 124-128 (2007)).
[0077] The ncAA-containing protein gene may be, for example, a gene having the known or naturally occurring nucleotide sequence of a gene encoding such a protein, except for the inclusion of an ncAA codon. Similarly, the ncAA-containing protein may be, for example, a protein having the known or naturally occurring amino acid sequence of such a protein, except for the inclusion of an ncAA. Furthermore, the ncAA-containing protein gene may be, for example, a variant of a gene having the known or naturally occurring nucleotide sequence of a gene encoding such a protein, except for the inclusion of an ncAA codon. Similarly, the ncAA-containing protein may be, for example, a variant of a protein having the known or naturally occurring amino acid sequence of such a protein, except for the inclusion of an ncAA. The description of conservative variants of ncAA-aaRS and ncAA-aaRS genes described below can be applied mutatis mutandis to variants of ncAA-containing proteins and ncAA-containing protein genes. For example, the ncAA-containing protein gene may be a gene encoding a protein having the amino acid sequence of the known or naturally occurring amino acid sequence of such a protein, except for the inclusion of an ncAA, in which one or several amino acids are substituted, deleted, inserted, and / or added at one or several positions. Furthermore, for example, an ncAA-containing protein gene may be a gene encoding a protein having an amino acid sequence that, apart from containing an ncAA, is 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire known or naturally occurring amino acid sequence of such a protein. Proteins identified by their originating biological species are not limited to proteins found in the biological species themselves, but also include proteins having the amino acid sequence of proteins found in the biological species and their variants. These variants may or may not be found in the biological species. That is, for example, a "human-derived protein" is not limited to proteins found in humans, but also includes proteins having the amino acid sequence of proteins found in humans and their variants.The gene encoding the ncAA-containing protein may be modified to contain an ncAA codon. Alternatively, the gene encoding the ncAA-containing protein may be modified by substituting an arbitrary codon with an equivalent codon. For example, the gene encoding the ncAA-containing protein may be modified to have an optimal codon depending on the codon usage frequency of the host used.
[0078] The gene construct of the present invention may further contain a nucleic acid sequence encoding an amino acid sequence used for enzymatic cleavage between the nucleic acid sequence encoding the amino acid sequence containing Gln-Glu-Thr and the nucleic acid sequence encoding the ncAA-containing protein. By inserting the amino acid sequence used for enzymatic cleavage into the fusion protein of the present invention, the expressed fusion protein can be enzymatically cleaved to obtain the ncAA-containing protein.
[0079] The amino acid sequence used for enzymatic cleavage is not particularly limited as long as it is recognized and cleaved by an enzyme that hydrolyzes peptide bonds, and a usable sequence may be appropriately selected depending on the amino acid sequence of the ncAA-containing protein. The nucleic acid sequence encoding the amino acid sequence used for enzymatic cleavage can be appropriately designed based on the amino acid sequence. For example, the nucleic acid sequence encoding the amino acid sequence used for enzymatic cleavage can be designed to have optimal codons depending on the codon usage frequency of the host.
[0080] The amino acid sequence used for enzymatic cleavage is preferably a recognition sequence for a protease with high substrate specificity. Specific examples of such amino acid sequences include the recognition sequences for Factor Xa protease and proTEV protease. Factor Xa protease recognizes the amino acid sequence Ile-Glu-Gly-Arg (IEGR) (SEQ ID NO: 37) in proteins, and proTEV protease recognizes the amino acid sequence Glu-Asn-Leu-Tyr-Phe-Gln (ENLYFQ) (SEQ ID NO: 38) in proteins, and specifically cleaves the C-terminus of each sequence.
[0081] The N-terminal region of the ncAA-containing protein finally obtained by the method of the present invention may or may not be identical to that of the naturally occurring protein. For example, the N-terminal region of the finally obtained ncAA-containing protein may have one or several extra amino acids added or deleted compared to that of the naturally occurring protein. Note that the term "one or several" varies depending on the full length and structure of the ncAA-containing protein, but specifically preferably means 1 to 20, more preferably 1 to 10, even more preferably 1 to 5, and particularly preferably 1 to 3 amino acids.
[0082] Furthermore, the ncAA-containing protein to be secreted and produced may be a protein with a pro-structure (proprotein). When the ncAA-containing protein to be secreted and produced is a proprotein, the final ncAA-containing protein may or may not be a proprotein. That is, the proprotein may be cleaved to form a mature protein by cleaving the pro-structure. Cleavage can be performed, for example, with a protease. When using a protease, from the perspective of the activity of the final protein, it is generally preferable for the proprotein to be cleaved at approximately the same position as the native protein, and more preferably, to cleave the proprotein at exactly the same position as the native protein to produce a mature protein identical to the native protein. Therefore, in general, specific proteases that cleave the proprotein at a position that produces a protein identical to the naturally occurring mature protein are most preferred. However, as mentioned above, the N-terminal region of the final ncAA-containing protein does not have to be identical to that of the native protein. For example, depending on the type of ncAA-containing protein to be produced and its intended use, a protein with an N-terminus that is one to several amino acids longer or shorter than the native protein may have more appropriate activity. Proteases that can be used in the present invention include commercially available proteases such as Dispase (Boehringer Mannheim) as well as those obtained from microbial culture media, such as culture media of actinomycetes. Such proteases can be used in an unpurified state, or may be purified to an appropriate purity as needed. When a mature protein is obtained by cleaving the pro-structure, the inserted amino acid sequence containing Gln-Glu-Thr is cleaved and removed together with the pro-structure, so that the target protein can be obtained without locating an amino acid sequence used for enzymatic cleavage after the amino acid sequence containing Gln-Glu-Thr.
[0083] The method for introducing a gene construct for secretory expression into a coryneform bacterium is not particularly limited. "Introduction of a gene construct for secretory expression" refers to having the gene construct retained in the host. "Introduction of a gene construct for secretory expression" is not limited to the case where a pre-constructed gene construct is introduced into the host all at once, but also includes the case where at least an ncAA-containing protein gene is introduced into the host and the gene construct is constructed within the host. In the bacterium of the present invention, the gene construct for secretory expression may be present on a vector that autonomously replicates outside the chromosome, such as a plasmid, or may be integrated into the chromosome. Introduction of a gene construct for secretory expression can be carried out, for example, in the same manner as the introduction of genes in the "method for increasing protein activity" described below.
[0084] The gene construct for secretory expression can be introduced into a host, for example, using a vector containing the gene construct. For example, the gene construct for secretory expression can be ligated to a vector to construct an expression vector for the gene construct, and the host can be transformed with the expression vector to introduce the gene construct into the host. Furthermore, for example, when a vector has a promoter that functions in coryneform bacteria, an expression vector for the gene construct for secretory expression can also be constructed by ligating a nucleotide sequence encoding the fusion protein of the present invention downstream of the promoter. There are no particular limitations on the vector, as long as it is capable of autonomous replication in coryneform bacteria. Vectors that can be used in coryneform bacteria are as described above.
[0085] The secretory expression gene construct can also be introduced into the host chromosome using a transposon, such as an artificial transposon. When a transposon is used, the secretory expression gene construct is introduced into the chromosome by homologous recombination or its own transposition ability. The secretory expression gene construct can also be introduced into the host chromosome by other transfer methods utilizing homologous recombination. Examples of transfer methods utilizing homologous recombination include methods using linear DNA, a plasmid containing a temperature-sensitive replication origin, a conjugatively transferable plasmid, or a suicide vector lacking a replication origin functional in the host. Alternatively, at least the ncAA-containing protein gene may be introduced into the chromosome to construct the secretory expression gene construct on the chromosome. In this case, some or all of the components of the secretory expression gene construct, other than the ncAA-containing protein gene, may be originally present on the host chromosome. Specifically, for example, a gene construct for secretory expression can be constructed on a chromosome, and the bacterium of the present invention can be constructed by simply using a promoter sequence originally present on the host chromosome and a nucleic acid sequence encoding a signal peptide connected downstream of the promoter sequence, and replacing only the gene connected downstream of the nucleic acid sequence encoding the signal peptide with an ncAA-containing protein gene. Introduction of a portion of the gene construct for secretory expression, such as an ncAA-containing protein gene, into a chromosome can be carried out in the same manner as introduction of the gene construct for secretory expression into a chromosome.
[0086] Gene constructs for secretory expression and their components (promoter sequence, nucleic acid sequence encoding a signal peptide, nucleic acid sequence encoding an ncAA-containing protein, etc.) can be obtained, for example, by cloning. Specifically, for example, an ncAA-containing protein gene is obtained by cloning from an organism that has the ncAA-containing protein, and then modifications such as the introduction of an ncAA codon, a nucleotide sequence encoding a signal peptide, and a promoter sequence are performed to obtain a gene construct for secretory expression. Gene constructs for secretory expression and their components can also be obtained by chemical synthesis (Gene, 60(1), 115-127 (1987)). The obtained gene constructs and their components can be used as is or with appropriate modifications.
[0087] When expressing two or more types of proteins, the gene constructs for secretory expression of each protein need only be retained in the bacterium of the present invention so that the secretory expression of the ncAA-containing protein can be achieved. Specifically, for example, all of the gene constructs for secretory expression of each protein may be retained on a single expression vector, or all may be retained on a chromosome. Alternatively, the gene constructs for secretory expression of each protein may be retained separately on multiple expression vectors, or may be retained separately on a single or multiple expression vectors and on a chromosome. "Expressing two or more types of proteins" refers, for example, to the secretion and production of two or more types of ncAA-containing proteins or the secretion and production of a heteromultimeric protein.
[0088] The method for introducing a gene construct for secretory expression into a coryneform bacterium is not particularly limited, and commonly used methods such as the protoplast method (Gene, 39, 281-286 (1985)), electroporation (Bio / Technology, 7, 1067-1070 (1989)), and electric pulse method (Japanese Patent Laid-Open Publication No. 2-207791) can be used.
[0089] <1-3> Expression of an orthogonal pair of tRNA and aaRS corresponding to an ncAA The coryneform bacterium of the present invention has been modified to express an orthogonal pair of tRNA and aaRS corresponding to an ncAA (i.e., an orthogonal tRNA(ncAA) / ncAA-aaRS pair). "Expression of an orthogonal tRNA(ncAA) / ncAA-aaRS pair" refers to the expression of the tRNA and aaRS that constitute the orthogonal tRNA(ncAA) / ncAA-aaRS pair. The tRNA and aaRS that constitute the orthogonal tRNA(ncAA) / ncAA-aaRS pair are expressed from the genes that encode them, respectively. In other words, the coryneform bacterium of the present invention has been modified to have genes that encode the tRNA and aaRS that constitute the orthogonal tRNA(ncAA) / ncAA-aaRS pair, respectively.
[0090] The tRNA that constitutes an orthogonal tRNA (ncAA) / ncAA-aaRS pair is the tRNA that corresponds to the ncAA. The tRNA that corresponds to the ncAA is referred to as "tRNA (ncAA)" or "tRNA ncAA The gene encoding tRNA(ncAA) is also called "tRNA gene corresponding to ncAA," "tRNA(ncAA) gene," or "tRNA ncAA Also called "genes."
[0091] The aaRS that constitutes an orthogonal tRNA (ncAA) / ncAA-aaRS pair is the aaRS that corresponds to the ncAA. The aaRS that corresponds to the ncAA is also called "ncAA-aaRS." The gene that encodes the ncAA-aaRS is also called "the aaRS gene that corresponds to the ncAA" or "ncAA-aaRS gene."
[0092] The term "orthogonal" in the context of an orthogonal tRNA (ncAA) / ncAA-aaRS pair means that the tRNA (ncAA) and the ncAA-aaRS interact exclusively with each other. "The tRNA (ncAA) and the ncAA-aaRS interact exclusively with each other" may mean that the tRNA (ncAA) is recognized as a substrate by the ncAA-aaRS but is not substantially recognized as a substrate by any of the host's endogenous aaRSs, and that the ncAA-aaRS recognizes the tRNA (ncAA) as a substrate but is not substantially recognized as a substrate by any of the host's endogenous tRNAs. Examples of the host's endogenous aaRS include the host's endogenous aaRSs corresponding to the 20 naturally occurring amino acids. Examples of the host's endogenous tRNAs include the host's endogenous tRNAs corresponding to the 20 naturally occurring amino acids. "A tRNA(ncAA) is not substantially recognized as a substrate by any endogenous aaRS of the host" may mean, for example, that for each endogenous aaRS of the host, the Km of the aaRS for tRNA(ncAA) is at least 10-fold, at least 100-fold, or at least 1000-fold greater than the Km of the ncAA-aaRS for tRNA(ncAA), or that the tRNA(ncAA) is not recognized as a substrate by the aaRS at all. "A ncAA-aaRS is not substantially recognized as a substrate by any endogenous tRNA of the host" may mean, for example, that for each endogenous tRNA of the host, the Km of the ncAA-aaRS for the tRNA is at least 10-fold, at least 100-fold, or at least 1000-fold greater than the Km of the ncAA-aaRS for tRNA(ncAA), or that the ncAA-aaRS does not recognize the tRNA at all as a substrate.
[0093] "tRNA corresponding to ncAA" (tRNA(ncAA)) refers to an RNA that functions as an adapter molecule to transfer ncAA to a peptide chain during translation. The function as an adapter molecule is also referred to as "the function of tRNA(ncAA)." Functioning as an adapter molecule is also referred to as "having the function of tRNA(ncAA)." Specifically, "having the function of tRNA(ncAA)" may mean that the RNA is aminoacylated with ncAA to become an aminoacyl-tRNA of ncAA, and functions as an adapter molecule to transfer ncAA to a peptide chain during translation. The aminoacyl-tRNA of ncAA (i.e., tRNA aminoacylated with ncAA) is also referred to as "ncAA-tRNA."
[0094] tRNA(ncAA) has an anticodon corresponding to the ncAA codon. The anticodon corresponding to the ncAA codon is also referred to as an "ncAA anticodon." tRNA(ncAA) may inherently have an ncAA anticodon, or may be modified to have an ncAA anticodon. When tRNA(ncAA) has an ncAA anticodon, it is also referred to as "the tRNA(ncAA) gene has an ncAA anticodon."
[0095] Examples of tRNA(ncAA) include tRNA(Tyr) and tRNA(Pyl).
[0096] "tRNA(Tyr)" means a tRNA corresponding to tyrosine. A gene encoding tRNA(Tyr) is also called a "tRNA(Tyr) gene."
[0097] Examples of tRNA(Tyr) genes and tRNA(Tyr) include those from various organisms other than the host. Specific examples of tRNA(Tyr) genes and tRNA(Tyr) include those from archaea such as methanogenic archaea, and those from bacteria such as Escherichia and Geobacillus. Examples of methanogenic archaea include Methanocaldococcus archaea such as Methanocaldococcus jannaschii (formerly known as Methanococcus jannaschii). Examples of Escherichia bacteria include Escherichia coli. Examples of Geobacillus bacteria include Geobacillus stearothermophilus (formerly known as Bacillus stearothermophilus). Examples of tRNA(Tyr) genes and tRNA(Tyr) include those from archaea. More particularly, examples of tRNA(Tyr) genes and tRNA(Tyr) include those of archaea of the genus Methanocaldococcus, such as Methanocaldococcus jannaschii. The nucleotide sequences of tRNA(Tyr) genes and tRNA(Tyr) derived from various organisms can be obtained, for example, from public databases such as NCBI and technical documents such as patent documents. The nucleotide sequence of the tRNA(Tyr) gene of Methanocaldococcus jannaschii and the nucleotide sequence of the tRNA(Tyr) encoded by the gene are shown in SEQ ID NOs: 39 and 40, respectively. Note that the nucleotide sequence of the tRNA gene becomes the nucleotide sequence of the tRNA if "T" is replaced with "U." Furthermore, the nucleotide sequence of the tRNA becomes the nucleotide sequence of the tRNA gene if "U" is replaced with "T."
[0098] The tRNA(Tyr) gene and tRNA(Tyr) exemplified above may be used as a tRNA(ncAA) gene and tRNA(ncAA) either as is or after appropriate modification. For example, when at least the anticodon of tRNA(Tyr) is not an ncAA anticodon, the anticodon of the tRNA(Tyr) is modified to an ncAA anticodon before use. In the nucleotide sequences shown in SEQ ID NOs: 39 and 40, the nucleotide sequence at positions 35 to 37 corresponds to the anticodon. The nucleotide sequences of an example of a modified tRNA(Tyr) gene from Methanocaldococcus jannaschii having an anticodon corresponding to UAG (amber), and the nucleotide sequence of the modified tRNA(Tyr) encoded by the gene are shown in SEQ ID NOs: 41 and 42, respectively. The nucleotide sequence of another example of a modified tRNA(Tyr) gene from Methanocaldococcus jannaschii having an anticodon corresponding to UAG (amber), and the nucleotide sequence of the modified tRNA(Tyr) encoded by the gene are shown in SEQ ID NOs: 43 and 44, respectively. Furthermore, for example, the tRNA(Tyr) may be modified to enhance its affinity for a host elongation factor (Jiantao Guo et al., "Evolution of amber suppressor tRNAs for efficient bacterial production of proteins containing nonnatural amino acids," Angew Chem Int Ed Engl, 2009;48(48):9148-51). Any of the above-exemplified modified tRNA(Tyr)s may be modified to enhance its affinity for a host elongation factor. The above-exemplified modified tRNA(Tyr)s may also be further modified before use. For example, the anticodon of the above-exemplified modified tRNA(Tyr) may be further modified.
[0099] "tRNA(Pyl)" means a tRNA corresponding to pyrrolysine. A gene encoding tRNA(Pyl) is also called a "tRNA(Pyl) gene."
[0100] Examples of tRNA(Pyl) genes and tRNA(Pyl) include those from various organisms other than the host. Specific examples of tRNA(Pyl) genes and tRNA(Pyl) include those from archaea such as methanogenic archaea, and those from bacteria such as bacteria of the genus Desulfitobacterium. Examples of methanogenic archaea include archaea of the genus Methanosarcina such as Methanosarcina barkerii and Methanosarcina mazei, archaea of the genus Methanomethylophilus such as Methanomethylophilus alvus, and unclassified archaea such as methanogenic archaeon ISO4-G1. Examples of bacteria of the genus Desulfitobacterium include Desulfitobacterium hafniense. Examples of tRNA(Pyl) genes and tRNA(Pyl) include those from archaea in particular. More particularly, tRNA(Pyl) genes and tRNA(Pyl)s are those of archaea of the genus Methanosarcina, such as Methanosarcina barkerii and Methanosarcina mazei. The nucleotide sequences of tRNA(Pyl) genes and tRNA(Pyl)s derived from various organisms can be obtained from public databases such as NCBI and technical documents such as patent documents. The nucleotide sequence of the tRNA(Pyl) gene of Methanosarcina barkerii and the nucleotide sequence of the tRNA(Pyl) encoded by the gene are shown in SEQ ID NOs: 45 and 46, respectively.
[0101] The tRNA(Pyl) gene and tRNA(Pyl) exemplified above may be used as is or after appropriate modification as the tRNA(Pyl) gene and tRNA(Pyl). For example, if at least the anticodon of the tRNA(Pyl) is not an ncAA anticodon, the tRNA(Pyl) is used after modifying the anticodon to an ncAA anticodon. In the nucleotide sequences shown in SEQ ID NOs: 45 and 46, the nucleotide sequence at positions 31 to 33 corresponds to the anticodon. Note that the anticodon in the nucleotide sequences shown in SEQ ID NOs: 45 and 46 is originally an anticodon corresponding to UAG (amber). The nucleotide sequences of the modified tRNA(Pyl) gene of Methanosarcina mazei having an anticodon corresponding to UAG (amber) and the nucleotide sequence of the modified tRNA(Pyl) encoded by the same gene are shown in SEQ ID NOs: 118 and 119, respectively. The nucleotide sequence of the modified tRNA(Pyl) gene of Methanosarcina mazei having an anticodon corresponding to UGA (opal), and the nucleotide sequence of the modified tRNA(Pyl) encoded by the gene are shown in SEQ ID NOs: 120 and 121, respectively.
[0102] "An aaRS corresponding to an ncAA (ncAA-aaRS)" refers to a protein having the activity of catalyzing a reaction in which a tRNA (ncAA) is aminoacylated with an ncAA to generate an aminoacyl-tRNA (ncAA-tRNA). This activity is also referred to as "an aaRS activity corresponding to an ncAA" or "ncAA-aaRS activity." Specifically, the ncAA-aaRS activity may be the activity of catalyzing a reaction in which a tRNA (ncAA) is aminoacylated with an ncAA in the presence of ATP to generate an aminoacyl-tRNA (ncAA-tRNA).
[0103] Examples of ncAA-aaRS include Tyr-RS and Pyl-RS.
[0104] "Tyr-RS" refers to tyrosyl-tRNA synthetase. "Tyrosyl-tRNA synthetase" refers to an aaRS corresponding to tyrosine. The gene encoding Tyr-RS is also referred to as the "Tyr-RS gene." ncAAs that can serve as substrates for Tyr-RS include tyrosine derivatives. Specific examples of ncAAs that can serve as substrates for Tyr-RS include AzF, AzY, ClY, NOY, SfY, and compounds 1 to 15, 17 to 26, 31, 32, 34 to 36, 41 to 44, 46, and 48 to 50 in Figure 1.
[0105] Examples of Tyr-RS genes and Tyr-RS include those from various organisms other than the host. Specific examples of Tyr-RS genes and Tyr-RS include those from the organisms exemplified as the origins of tRNA(Tyr) genes and tRNA(Tyr). Examples of Tyr-RS genes and Tyr-RS include those from archaea, in particular. Examples of Tyr-RS genes and Tyr-RS include those from archaea of the genus Methanocaldococcus, such as Methanocaldococcus jannaschii. The nucleotide sequences of Tyr-RS genes and the amino acid sequences of Tyr-RS from various organisms can be obtained, for example, from public databases such as NCBI and technical literature such as patent documents. The nucleotide sequence of the Tyr-RS gene from Methanocaldococcus jannaschii and the amino acid sequence of the Tyr-RS encoded by the gene are shown in SEQ ID NOs: 47 and 48, respectively.
[0106] The Tyr-RS gene and Tyr-RS exemplified above may be used as an ncAA-aaRS gene or ncAA-aaRS, either as is or after appropriate modification. For example, at least when a Tyr-RS does not recognize an ncAA as a substrate, the Tyr-RS is used after modifying its substrate specificity so that it recognizes an ncAA as a substrate (i.e., has ncAA-aaRS activity). That is, the Tyr-RS may have a mutation that modifies its substrate specificity.
[0107] A Tyr-RS having a mutation that alters substrate specificity is also referred to as a "mutant Tyr-RS" or "modified Tyr-RS." A gene encoding a mutant Tyr-RS is also referred to as a "mutant Tyr-RS gene" or "modified Tyr-RS gene." A Tyr-RS without a mutation that alters substrate specificity is also referred to as a "wild-type Tyr-RS." A gene encoding a wild-type Tyr-RS is also referred to as a "wild-type Tyr-RS gene." Note that the term "wild-type" here is a convenient description to distinguish "wild-type" Tyr-RS from "mutant" Tyr-RS, and is not limited to naturally occurring Tyr-RSs as long as they do not have a mutation that alters substrate specificity. For example, the wild-type Tyr-RS may be a variant of the wild-type Tyr-RS exemplified above (e.g., a protein having the amino acid sequence set forth in SEQ ID NO: 48), as long as the mutant Tyr-RS has ncAA-aaRS activity. The description of the ncAA-aaRS variants described below applies mutatis mutandis to the wild-type Tyr-RS variants. "The Tyr-RS does not have a mutation that alters substrate specificity" may mean that the Tyr-RS does not have a mutation selected to alter substrate specificity. The wild-type Tyr-RS may or may not have a mutation that was not selected to alter substrate specificity, as long as it does not have a mutation selected to alter substrate specificity. The nucleotide sequence of the modified Tyr-RS gene from Methanocaldococcus jannaschii that recognizes AzF as a substrate and the amino acid sequence of the modified Tyr-RS encoded by the gene are shown in SEQ ID NOs: 49 and 50, respectively. The nucleotide sequence of the modified Tyr-RS gene from Methanocaldococcus jannaschii that recognizes halogenated tyrosine and AzY as substrates and the amino acid sequence of the modified Tyr-RS encoded by the gene are shown in SEQ ID NOs: 51 and 52, respectively. The modified Tyr-RSs exemplified above may be further modified before use. For example, the substrate specificity of the modified Tyr-RS exemplified above may be further modified.
[0108] Mutations that alter the substrate specificity of Tyr-RS include mutations in the following amino acid residues (Jason W Chin et al., Addition of p-azido-L-phenylalanine to the genetic code of Escherichia coli, J Am Chem Soc. 2002 Aug 7;124(31):9026-7; Biochem. Biophys. Res. Commun. 411, 757-761 (2011); and WO2004 / 070024): Y32, H70, E107, D158, I159, L162, and D286.
[0109] Mutations at these amino acid residues can be effective for modifying the Tyr-RS to recognize, for example, AzF, AzY, and / or halogenated tyrosine (e.g., ClY) as a substrate. Specifically, mutations at Y32, E107, D158, I159, and L162 can be effective for modifying the Tyr-RS to recognize, for example, AzF as a substrate. Furthermore, mutations at H70, D158, and I159 can be effective for modifying the Tyr-RS to recognize, for example, AzY and / or halogenated tyrosine as a substrate. Furthermore, mutations at D286 can be effective for enhancing substrate specificity for, for example, amber suppressor tRNAs.
[0110] A mutation that alters the substrate specificity of Tyr-RS may be a mutation in a single amino acid residue or a combination of mutations in two or more amino acid residues. That is, a mutation that alters the substrate specificity of Tyr-RS may include, for example, a mutation in one or more amino acid residues selected from these amino acid residues. A mutation that alters the substrate specificity of Tyr-RS may be, for example, a mutation in a single amino acid residue selected from these amino acid residues or a combination of mutations in two or more amino acid residues selected from these amino acid residues. A mutation that alters the substrate specificity of Tyr-RS may be, for example, a combination of a mutation in one or more amino acid residues selected from Y32, H70, E107, D158, I159, and L162 and a mutation in D286. A mutation that alters the substrate specificity of Tyr-RS may specifically be, for example, a combination of a mutation in one or more amino acid residues selected from Y32, E107, D158, I159, and L162 and a mutation in D286. Specifically, the mutation that alters the substrate specificity of Tyr-RS may be, for example, a combination of a mutation at one or more amino acid residues selected from H70, D158, and I159 and a mutation at D286.
[0111] In the above notation for identifying amino acid residues, the numbers indicate positions in the amino acid sequence shown in SEQ ID NO: 48, and the letters to the left of the numbers indicate the amino acid residues at each position (i.e., the amino acid residues at each position before modification) in the amino acid sequence shown in SEQ ID NO: 48. That is, for example, "Y32" indicates the Y (Tyr) residue at position 32 in the amino acid sequence shown in SEQ ID NO: 48.
[0112] In any Tyr-RS, these amino acid residues each represent "the amino acid residue corresponding to the amino acid residue in the amino acid sequence shown in SEQ ID NO: 48." That is, for example, "Y32" in any Tyr-RS represents the amino acid residue corresponding to the Y (Tyr) residue at position 32 in the amino acid sequence shown in SEQ ID NO: 48.
[0113] Each of the above mutations may be a substitution of an amino acid residue. In each of the above mutations, the modified amino acid residue may be any amino acid residue other than the amino acid residue before modification, as long as the desired substrate specificity is obtained. In other words, the modified amino acid residue may be selected so as to obtain the desired substrate specificity. Specific examples of modified amino acid residues include amino acid residues selected from K (Lys), R (Arg), H (His), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), F (Phe), W (Trp), Y (Tyr), C (Cys), M (Met), D (Asp), E (Glu), N (Asn), and Q (Gln), other than the amino acid residue before modification.
[0114] Specific examples of mutations that alter the substrate specificity of Tyr-RS include the following (Jason W Chin et al., Addition of p-azido-L-phenylalanine to the genetic code of Escherichia coli, J Am Chem Soc. 2002 Aug 7;124(31):9026-7; Biochem. Biophys. Res. Commun. 411, 757-761 (2011); and WO2004 / 070024): Y32 (T, L, A, G), H70A, E107 (N, S, T, R, P), D158 (P, V, T, Q), I159 (L, S, V, I, Y), L162 (Q, D, L, S), and D286 (A, R, Y).
[0115] That is, mutations that alter the substrate specificity of Tyr-RS may include, for example, one or more mutations selected from these mutations. Mutations that alter the substrate specificity of Tyr-RS may be, for example, a single mutation selected from these mutations, or a combination of two or more mutations selected from these mutations. These mutations may be effective, for example, for modifying Tyr-RS to recognize AzF, AzY, and / or halogenated tyrosine (e.g., ClY) as a substrate. Specifically, Y32 (T, L, A, G), E107 (N, S, T, R, P), D158 (P, V, T, Q), I159 (L, S, V, I, Y), and L162 (Q, D, L, S) may be effective, for example, for modifying Tyr-RS to recognize AzF as a substrate. Specifically, H70A, D158 (P, V, T, Q), and I159 (L, S, V, I, Y) can be effective for modifying Tyr-RS to recognize, for example, AzY and / or halogenated tyrosine as a substrate. Specifically, D286 (A, R, Y) can be effective for enhancing substrate specificity for, for example, amber suppressor tRNA.
[0116] In the above notation for specifying a mutation, the meaning of the number and the letter to the left of it are the same as above. In the above notation for specifying a mutation, the letter to the right of the number indicates the amino acid residue after modification at each position. That is, for example, "Y32(T, L, A, G)" indicates a mutation in which the Y (Tyr) residue at position 32 in the amino acid sequence shown in SEQ ID NO: 48 is substituted with a T (Thr) residue, an L (Leu) residue, an A (Ala) residue, or a G (Gly) residue.
[0117] In any Tyr-RS, these mutations each represent "a mutation corresponding to the mutation in the amino acid sequence shown in SEQ ID NO: 48." In any Tyr-RS, "a mutation corresponding to a mutation in which the amino acid residue at position X in the amino acid sequence shown in SEQ ID NO: 48 is substituted with a certain amino acid residue" is to be read as "a mutation in which the amino acid residue corresponding to the amino acid residue at position X in the amino acid sequence shown in SEQ ID NO: 48 is substituted with a certain amino acid residue." That is, for example, in any Tyr-RS, "Y32T" represents a mutation in which the amino acid residue corresponding to the Y (Tyr) residue at position 32 in the amino acid sequence shown in SEQ ID NO: 48 is substituted with a T (Thr) residue.
[0118] The combination of mutations is not particularly limited. Examples of the combination of mutations include the following combinations (Jason W Chin et al., Addition of p-azido-L-phenylalanine to the genetic code of Escherichia coli, J Am Chem Soc. 2002 Aug 7;124(31):9026-7; and Biochem. Biophys. Res. Commun. 411, 757-761 (2011)): Y32T / E107N / D158P / I159L / L162Q Y32T / E107S / D158P / I159S / L162Q Y32T / E107S / D158P / I159L / L162Q Y32L / E107T / D158P / I159V / L162Q Y32A / E107A / D158V / I159I / L162A Y32G / E107T / D158T / I159Y / L162L Y32L / E107P / D158Q / I159I / L162S H70A / D158T / I159S / D286R.
[0119] That is, mutations that alter the substrate specificity of Tyr-RS may include, for example, any combination of these. Mutations that alter the substrate specificity of Tyr-RS may include, for example, any combination of these. These combinations may be effective, for example, to alter Tyr-RS to recognize AzF, AzY, and / or halogenated tyrosine (such as ClY) as a substrate. Specifically, Y32T / E107N / D158P / I159L / L162Q, Y32T / E107S / D158P / I159S / L162Q, Y32T / E107S / D158P / I159L / L162Q, Y32L / E107T / D158P / I159V / L162Q, Y32A / E107A / D158V / I159I / L162A, Y32G / E107T / D158T / I159Y / L162L, and Y32L / E107P / D158Q / I159I / L162S can be effective for modifying Tyr-RS to recognize, for example, AzF as a substrate. Specifically, H70A / D158T / I159S / D286R may be effective for modifying Tyr-RS to recognize, for example, AzY and / or halogenated tyrosine as a substrate.
[0120] In the above notation for specifying a combination, the meanings of the numbers and the letters to the left and right of them are the same as those described above. In the above notation for specifying a combination, the combination of two or more mutations separated by " / " indicates a double mutation or multiple mutations. For example, "Y32T / E107N / D158P / I159L / L162Q" indicates a quintuple mutation of Y32T, E107N, D158P, I159L, and L162Q.
[0121] With regard to the positions of the amino acid residues referred to in each of the above mutations in any Tyr-RS, the explanation of the position of the "amino acid residue at position X of the wild-type PhoS protein" described below can be applied mutatis mutandis, except that the amino acid sequence shown in SEQ ID NO: 48 is used as the reference sequence.
[0122] "Pyl-RS" refers to pyrrolysyl-tRNA synthetase. "Pyrrolysyl-tRNA synthetase" refers to an aaRS corresponding to pyrrolysine. The gene encoding Pyl-RS is also referred to as the "Pyl-RS gene." Examples of ncAAs that can serve as substrates for Pyl-RS include lysine derivatives. Specific examples of ncAAs that can serve as substrates for Pyl-RS include Pyl, AllocLys, compounds 40, 51, and 59-71 in Figure 1, and the compounds described in Wei Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biochim Biophys Acta. 2014 Jun;1844(6):1059-70.
[0123] Pyl-RS genes and Pyl-RSs can be found in various organisms other than the host. Specific examples of Pyl-RS genes and Pyl-RSs include those from the organisms exemplified as the origins of tRNA(Pyl) genes and tRNA(Pyl). Pyl-RS genes and Pyl-RSs can particularly be found in archaea. More particularly, Pyl-RS genes and Pyl-RSs can be found in archaea of the genus Methanosarcina, such as Methanosarcina barkerii and Methanosarcina mazei. The nucleotide sequences of Pyl-RS genes and Pyl-RSs from various organisms can be obtained, for example, from public databases such as NCBI and technical literature such as patent documents. The nucleotide sequence of the Pyl-RS gene from Methanosarcina barkerii and the amino acid sequence of the Pyl-RS encoded by this gene are shown in SEQ ID NOs: 53 and 54, respectively. The nucleotide sequence of the Pyl-RS gene from Methanosarcina mazei and the amino acid sequence of the Pyl-RS encoded by this gene are shown in SEQ ID NOs: 114 and 115, respectively.
[0124] The Pyl-RS gene and Pyl-RS exemplified above may be used as an ncAA-aaRS gene or ncAA-aaRS, either as is or after appropriate modification. For example, when a Pyl-RS does not recognize an ncAA as a substrate, the Pyl-RS is used after modifying its substrate specificity so that it recognizes an ncAA as a substrate (i.e., has ncAA-aaRS activity). That is, the Pyl-RS may have a mutation that modifies its substrate specificity.
[0125] A Pyl-RS having a mutation that alters substrate specificity is also referred to as a "mutant Pyl-RS" or "modified Pyl-RS." A gene encoding a mutant Pyl-RS is also referred to as a "mutant Pyl-RS gene" or "modified Pyl-RS gene." A Pyl-RS without a mutation that alters substrate specificity is also referred to as a "wild-type Pyl-RS." A gene encoding a wild-type Pyl-RS is also referred to as a "wild-type Pyl-RS gene." Note that the term "wild-type" here is used for convenience to distinguish "wild-type" Pyl-RS from "mutant" Pyl-RS, and is not limited to naturally occurring Pyl-RSs as long as they do not have a mutation that alters substrate specificity. For example, a wild-type Pyl-RS may be a variant of the wild-type Pyl-RS exemplified above (e.g., a protein having the amino acid sequence set forth in SEQ ID NO: 54 or 115), as long as the mutant Pyl-RS has ncAA-aaRS activity. The description of wild-type Pyl-RS variants described below for ncAA-aaRS variants can be applied mutatis mutandis. "Pyl-RS does not have a mutation that alters substrate specificity" may mean that Pyl-RS does not have a mutation selected as a mutation that alters substrate specificity. As long as the wild-type Pyl-RS does not have a mutation selected as a mutation that alters substrate specificity, it may or may not have a mutation that was not selected as a mutation that alters substrate specificity.
[0126] Mutations that alter the substrate specificity of Pyl-RS include mutations in the following amino acid residues (Wei Wan et al., Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochim Biophys Acta. 2014 Jun;1844(6):1059-70.): M241, L266, A267, L270, Y271, L274, N311, C313, M315, Y349, V367, and W383.
[0127] A mutation that alters the substrate specificity of Pyl-RS may be a mutation in a single amino acid residue, or a combination of mutations in two or more amino acid residues. That is, a mutation that alters the substrate specificity of Pyl-RS may include, for example, a mutation in one or more amino acid residues selected from these amino acid residues. A mutation that alters the substrate specificity of Pyl-RS may be, for example, a mutation in a single amino acid residue selected from these amino acid residues, or a combination of mutations in two or more amino acid residues selected from these amino acid residues. Mutations in these amino acid residues may be effective, for example, to modify Pyl-RS to recognize the ncAA described in Wei Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biochim Biophys Acta. 2014 Jun;1844(6):1059-70, as a substrate.
[0128] Each of the above mutations may be a substitution of an amino acid residue. In each of the above mutations, the modified amino acid residue may be any amino acid residue other than the amino acid residue before modification, as long as the desired substrate specificity is obtained. In other words, the modified amino acid residue may be selected so as to obtain the desired substrate specificity. Specific examples of modified amino acid residues include amino acid residues selected from K (Lys), R (Arg), H (His), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), F (Phe), W (Trp), Y (Pyl), C (Cys), M (Met), D (Asp), E (Glu), N (Asn), and Q (Gln), other than the amino acid residue before modification.
[0129] Specific mutations that alter the substrate specificity of Pyl-RS include the following (Wei Wan et al., Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochim Biophys Acta. 2014 Jun;1844(6):1059-70): M241F, L266(M, V, L), A267(S, L, F, T), L270(I, M, F), Y271(A, G, M, L, I, C, F), L274(A, G, M, P, L, S), N311(A, S, T, V, G), C313(A, V, S, F, C, T, K, L, W, G), M315F, Y349(F, W, L), V367(L, I), and W383(Y).
[0130] That is, mutations that alter the substrate specificity of Pyl-RS may include, for example, one or more mutations selected from these mutations. Mutations that alter the substrate specificity of Pyl-RS may be, for example, a single mutation selected from these mutations, or a combination of two or more mutations selected from these mutations. These mutations may be effective, for example, for modifying Pyl-RS to recognize the ncAA described in Wei Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biochim Biophys Acta. 2014 Jun;1844(6):1059-70, as a substrate.
[0131] The combination of mutations is not particularly limited. Examples of the combination of mutations include the following (Wei Wan et al., Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochim Biophys Acta. 2014 Jun;1844(6):1059-70).<h2 style=";text-align:left;direction:ltr">): L274A / C313A / Y349F / 、L274A / C313V / Y349F / 、A267S / C313V / M315F / 、C313V / 、Y271A / Y349F / 、L274A / C313V / 、Y271A / Y349F / 、Y271G / Y 349F / Y271M / L274G / C313A / Y271A / L274M / C313A / L274A / C313A / Y349F / L274A / C313S / Y349F / L266M / Y271L / L274A / C313F / Y349W / L274M / C313A / Y349F / Y271M / L274A / C313A / Y349F / Y271I / L274A / C313A / Y349F / A267S / Y271C / L274M / C313C / M241F / Y349F / Y271M / L274A / C313T / 、Y271M / L274A / C313C / 、Y271M / L274P / C313C / 、Y271I / L274M / C313A / 、Y349W / 、Y271M / L274G / C313A / 、Y271M / L27 4G / C313A / Y349W / 、L266M / Y271L / L274A / C313F / 、L266M / Y271L / L2 74L / C313S / 、L266V / L270I / Y271F / L274A / C313F / 、L266L / L270I / Y 271L / L274A / C313F / 、L266M / L270I / Y271F / L274A / C313F / 、Y271A / L274M / Y349F / 、Y349W / 、L274A / C313F / Y349F / 、N311A / C313L / 、N311 A / C313K / 、A267L / Y271M / N311S / C313L / Y349L / 、A267F / Y271L / N311T / C313F / Y349L / 、L270M / Y271L / L274S / N311S / C313M / 、N311A / C313A / 、A267T / N311V / C313W / Y349F / V367L / 、L270F / Y271M / N311G / C313G / 、A267T / N311T / C313T / 、A267T / N311G / C313T / V367I / W383Y。.<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">
[0132] That is, mutations that alter the substrate specificity of Pyl-RS may include, for example, any combination of these. Mutations that alter the substrate specificity of Pyl-RS may include, for example, any combination of these. These combinations may be effective for modifying Pyl-RS to recognize, as a substrate, the ncAA described in, for example, Wei Wan et al., "Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool," Biochim Biophys Acta. 2014 Jun;1844(6):1059-70.
[0133] Regarding the notation of mutations that alter the substrate specificity of Pyl-RS, the explanation for the notation of mutations that alter the substrate specificity of Tyr-RS can be applied mutatis mutandis, except that the amino acid sequence shown in SEQ ID NO: 54 is used as the reference sequence.
[0134] The tRNA(ncAA) gene may be, for example, a gene having the nucleotide sequence of the above-exemplified tRNA(ncAA) gene (e.g., a nucleotide sequence in which the anticodon has been modified in the nucleotide sequence shown in SEQ ID NO: 39, 41, 43, 45, 118, or 120, or a nucleotide sequence shown in SEQ ID NO: 41, 43, 45, 118, or 120). The tRNA(ncAA) may be, for example, an RNA having the nucleotide sequence of the above-exemplified tRNA(ncAA) (e.g., a nucleotide sequence in which the anticodon has been modified in the nucleotide sequence shown in SEQ ID NO: 40, 42, 44, 46, 119, or 121, or a nucleotide sequence shown in SEQ ID NO: 42, 44, 46, 119, or 121). The ncAA-aaRS gene may be, for example, a gene having the nucleotide sequence of the above-exemplified ncAA-aaRS gene (e.g., a nucleotide sequence having a mutation that alters substrate specificity in the nucleotide sequence shown in SEQ ID NO: 47, 49, 51, 53, or 114, or a nucleotide sequence shown in SEQ ID NO: 49, 51, 53, or 114). The ncAA-aaRS may be, for example, a protein having the amino acid sequence of the above-exemplified ncAA-aaRS (e.g., an amino acid sequence having a mutation that alters substrate specificity in the amino acid sequence shown in SEQ ID NO: 48, 50, 52, 54, or 115, or an amino acid sequence shown in SEQ ID NO: 50, 52, 54, or 115). The expression "a gene or RNA has a nucleotide sequence" may mean that the gene or RNA contains the nucleotide sequence, and may also include cases where the gene or RNA consists of the nucleotide sequence, unless otherwise specified. Furthermore, unless otherwise specified, the expression "a protein has an amino acid sequence" may mean that the protein contains the amino acid sequence, and may also include the case where the protein consists of the amino acid sequence.
[0135] The tRNA(ncAA) gene may be a variant of the above-exemplified tRNA(ncAA) gene (e.g., a gene having a nucleotide sequence in which the anticodon is modified in the nucleotide sequence shown in SEQ ID NO: 39, 41, 43, 45, 118, or 120, or a gene having a nucleotide sequence shown in SEQ ID NO: 41, 43, 45, 118, or 120), so long as the original function is maintained. The tRNA(ncAA) may be a variant of the above-exemplified tRNA(ncAA) (e.g., an RNA having a nucleotide sequence in which the anticodon is modified in the nucleotide sequence shown in SEQ ID NO: 40, 42, 44, 46, 119, or 121, or an RNA having a nucleotide sequence shown in SEQ ID NO: 42, 44, 46, 119, or 121), so long as the original function is maintained. The ncAA-aaRS gene may be a variant of the above-exemplified ncAA-aaRS genes (e.g., a gene having a nucleotide sequence with a mutation that alters substrate specificity in the nucleotide sequence set forth in SEQ ID NO: 47, 49, 51, 53, or 114, or a gene having a nucleotide sequence set forth in SEQ ID NO: 49, 51, 53, or 114), so long as the original function is maintained. Similarly, the ncAA-aaRS may be a variant of the above-exemplified ncAA-aaRS (e.g., a protein having an amino acid sequence with a mutation that alters substrate specificity in the amino acid sequence set forth in SEQ ID NO: 48, 50, 52, 54, or 115, or a protein having an amino acid sequence set forth in SEQ ID NO: 50, 52, 54, or 115), so long as the original function is maintained. Such variants that maintain the original function are sometimes referred to as "conservative variants." The term "tRNA(ncAA) gene" encompasses not only the above-exemplified tRNA(ncAA) genes, but also their conservative variants. Similarly, the term "tRNA(ncAA)" encompasses the above-exemplified tRNA(ncAA) as well as their conservative variants. The term "ncAA-aaRS gene" encompasses the above-exemplified ncAA-aaRS gene as well as their conservative variants. Similarly, the term "ncAA-aaRS" encompasses the above-exemplified ncAA-aaRS as well as their conservative variants.Conservative variants include, for example, homologs and artificially modified forms of the above-mentioned tRNA(ncAA) genes, tRNA(ncAA), ncAA-aaRS genes, and ncAA-aaRS.
[0136] "Maintaining the original function" means that a gene or protein variant has a function (e.g., activity or property) corresponding to the function (e.g., activity or property) of the original gene or protein. That is, "maintaining the original function" with respect to a tRNA(ncAA) gene may mean that a gene variant encodes a tRNA(ncAA). "Maintaining the original function" with respect to a tRNA(ncAA) may mean that an RNA variant has the function of a tRNA(ncAA). "Maintaining the original function" with respect to an ncAA-aaRS gene may mean that a gene variant encodes an ncAA-aaRS. Furthermore, "maintaining the original function" with respect to an ncAA-aaRS may mean that a protein variant has ncAA-aaRS activity.
[0137] ncAA-aaRS activity can be measured, for example, by incubating the enzyme with substrates (ncAA and tRNA(ncAA)) in the presence of ATP and measuring the production of enzyme- and substrate-dependent products (AMP or ncAA-tRNA).
[0138] Examples of conservative variants are shown below.
[0139] Homologs of tRNA(ncAA) genes, ncAA-aaRS genes, or ncAA-aaRS genes can be easily obtained from public databases by, for example, BLAST or FASTA searches using the nucleotide sequences of the exemplified tRNA(ncAA) genes, ncAA-aaRS genes, or ncAA-aaRS amino acid sequences as query sequences. Homologs of tRNA(ncAA) genes or ncAA-aaRS genes can also be obtained by PCR using, for example, the chromosomes of various organisms as templates and oligonucleotides prepared based on the nucleotide sequences of these known tRNA(ncAA) genes or ncAA-aaRS genes as primers.
[0140] As long as the original function is maintained, the ncAA-aaRS gene may be a gene encoding a protein having an amino acid sequence in which one or several amino acids have been substituted, deleted, inserted, and / or added at one or several positions in the amino acid sequence (e.g., the amino acid sequence shown in SEQ ID NO: 48, 50, 52, 54, or 115, which has a mutation that alters substrate specificity, or the amino acid sequence shown in SEQ ID NO: 50, 52, 54, or 115). For example, the encoded protein may have its N-terminus and / or C-terminus extended or shortened. Note that the term "one or several" varies depending on the position and type of amino acid residue in the three-dimensional structure of the protein, but specifically means, for example, 1 to 50, 1 to 40, 1 to 30, preferably 1 to 20, more preferably 1 to 10, even more preferably 1 to 5, and particularly preferably 1 to 3.
[0141] The above-mentioned substitution, deletion, insertion, and / or addition of one or several amino acids is a conservative mutation that maintains normal protein function. A typical conservative mutation is a conservative substitution. A conservative substitution is a mutation in which Phe, Trp, and Tyr are substituted for each other when the substitution site is an aromatic amino acid; Leu, Ile, and Val are substituted for each other when the substitution site is a hydrophobic amino acid; Gln and Asn are substituted for each other when the substitution site is a polar amino acid; Lys, Arg, and His are substituted for each other when the substitution site is a basic amino acid; Asp and Glu are substituted for each other when the substitution site is an acidic amino acid; and Ser and Thr are substituted for each other when the substitution site is an amino acid having a hydroxyl group. Specific examples of substitutions that are considered to be conservative substitutions include substitution of Ala with Ser or Thr, substitution of Arg with Gln, His, or Lys, substitution of Asn with Glu, Gln, Lys, His, or Asp, substitution of Asp with Asn, Glu, or Gln, substitution of Cys with Ser or Ala, substitution of Gln with Asn, Glu, Lys, His, Asp, or Arg, substitution of Glu with Gly, Asn, Gln, Lys, or Asp, substitution of Gly with Pro, substitution of His with Asn, Lys, Gln, Arg, or Tyr, substitution of Il Examples of substitutions include substitutions of Lys with Leu, Met, Val, or Phe, substitutions of Leu with Ile, Met, Val, or Phe, substitutions of Lys with Asn, Glu, Gln, His, or Arg, substitutions of Met with Ile, Leu, Val, or Phe, substitutions of Phe with Trp, Tyr, Met, Ile, or Leu, substitutions of Ser with Thr or Ala, substitutions of Thr with Ser or Ala, substitutions of Trp with Phe or Tyr, substitutions of Tyr with His, Phe, or Trp, and substitutions of Val with Met, Ile, or Leu. The above-mentioned amino acid substitutions, deletions, insertions, or additions also include those resulting from naturally occurring mutations (mutants or variants) based on individual differences or differences in species of the organism from which the gene is derived.
[0142] Furthermore, the ncAA-aaRS gene may be a gene encoding a protein having an amino acid sequence that is, for example, 50% or more, 65% or more, 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire amino acid sequence described above, so long as the original function is maintained.
[0143] Furthermore, as long as the original function is maintained, the tRNA(ncAA) gene or ncAA-aaRS gene may be a gene (e.g., DNA) that hybridizes under stringent conditions with a probe that can be prepared from the above-mentioned nucleotide sequence (e.g., a nucleotide sequence in which the anticodon is modified in the nucleotide sequence shown in SEQ ID NO: 39, 41, 43, 45, 118, or 120 for the tRNA(ncAA) gene, or a nucleotide sequence shown in SEQ ID NO: 41, 43, 45, 118, or 120 for the ncAA-aaRS gene; a nucleotide sequence in which a mutation that modifies substrate specificity is present in the nucleotide sequence shown in SEQ ID NO: 47, 49, 51, 53, or 114 for the ncAA-aaRS gene, or a nucleotide sequence shown in SEQ ID NO: 49, 51, 53, or 114), for example, a sequence complementary to all or part of the above-mentioned nucleotide sequence. As long as the original function is maintained, the tRNA(ncAA) may be an RNA that hybridizes under stringent conditions with a probe that can be prepared from the above-mentioned nucleotide sequence (e.g., a nucleotide sequence in which the anticodon in the nucleotide sequence shown in SEQ ID NO: 40, 42, 44, 46, 119, or 121 has been modified, or a nucleotide sequence shown in SEQ ID NO: 42, 44, 46, 119, or 121), such as a sequence complementary to all or part of the above-mentioned nucleotide sequence. "Stringent conditions" refer to conditions under which a so-called specific hybrid is formed and a non-specific hybrid is not formed. One example of such conditions is a condition under which DNAs with high identity, for example, DNAs with an identity of 50% or more, 65% or more, 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more, hybridize with each other, while DNAs with lower identity do not hybridize with each other; or a condition in which washing is performed once, preferably two to three times, at a salt concentration and temperature equivalent to the washing conditions for conventional Southern hybridization, namely, 60°C, 1×SSC, 0.1% SDS, preferably 60°C, 0.1×SSC, 0.1% SDS, more preferably 68°C, 0.1×SSC, 0.1% SDS.
[0144] As mentioned above, the probe used in the hybridization may be a portion of the complementary sequence of the gene. Such a probe can be prepared by PCR using oligonucleotides prepared based on a known gene sequence as primers and a DNA fragment containing the gene as a template. For example, a DNA fragment of about 300 bp in length can be used as the probe. When a DNA fragment of about 300 bp in length is used as the probe, washing conditions for the hybridization include 50°C, 2×SSC, and 0.1% SDS.
[0145] Furthermore, since codon degeneracy differs depending on the host, the ncAA-aaRS gene may be one in which any codon has been replaced with an equivalent codon. That is, the ncAA-aaRS gene may be a variant of the above-exemplified ncAA-aaRS gene due to the degeneracy of the genetic code. For example, the ncAA-aaRS gene may be modified to have optimal codons depending on the codon usage frequency of the host used.
[0146] The "identity" between amino acid sequences refers to the identity between amino acid sequences calculated by blastp using default scoring parameters (Matrix: BLOSUM62; Gap Costs: Existence = 11, Extension = 1; Compositional Adjustments: Conditional compositional score matrix adjustment). The "identity" between nucleotide sequences refers to the identity between nucleotide sequences calculated by blastn using default scoring parameters (Match / Mismatch Scores = 1, -2; Gap Costs = Linear).
[0147] The above descriptions regarding conservative variants of genes and proteins can be applied mutatis mutandis to any genes and proteins.
[0148] <1-4> Other Properties The bacterium of the present invention may have any desired properties as long as it is capable of secreting and producing an ncAA-containing protein. For example, the bacterium of the present invention may have reduced activity of a cell surface protein (WO2013 / 065869, WO2013 / 065772, WO2013 / 118544, WO2013 / 062029). The bacterium of the present invention may also be modified to reduce the activity of a penicillin-binding protein (WO2013 / 065869). The bacterium of the present invention may also be modified to increase the expression of a gene encoding a metallopeptidase (WO2013 / 065772). The bacterium of the present invention may also be modified to have a mutant ribosomal protein S1 gene (mutant rpsA gene) (WO2013 / 118544). The bacterium of the present invention may also be modified to have a mutant phoS gene (WO2016 / 171224). The bacterium of the present invention may also be modified to reduce the activity of the RegX3 protein (WO2018 / 074578). The bacterium of the present invention may also be modified to reduce the activity of the HrrSA system (WO2018 / 074579). The bacterium of the present invention may also be modified to increase the activity of the Tat secretion apparatus. These properties or modifications can be used alone or in appropriate combination.
[0149] <1-4-1> Introduction of a mutant phoS gene The bacterium of the present invention may be modified to harbor a mutant phoS gene. "Harboring a mutant phoS gene" is also referred to as "having a mutant phoS gene" or "having a mutation in the phoS gene." Furthermore, "harboring a mutant phoS gene" is also referred to as "having a mutant PhoS protein" or "having a mutation in the PhoS protein."
[0150] The phoS gene and the PhoS protein are described below. The phoS gene encodes the PhoS protein, a sensor kinase in the PhoRS system. The PhoRS system is a two-component regulatory system that initiates a response to environmental phosphate deficiency. The PhoRS system consists of the sensor kinase PhoS, encoded by the phoS gene, and the response regulator PhoR, encoded by the phoR gene.
[0151] A PhoS protein having a "specific mutation" is also called a "mutant PhoS protein," and the gene encoding it is also called a "mutant phoS gene." In other words, a "mutant phoS gene" is a phoS gene having a "specific mutation." A PhoS protein that does not have a "specific mutation" is also called a "wild-type PhoS protein," and the gene encoding it is also called a "wild-type phoS gene." In other words, a "wild-type phoS gene" is a phoS gene that does not have a "specific mutation." Note that the term "wild-type" used here is a convenient description to distinguish it from a "mutant," and is not limited to naturally occurring ones, as long as they do not have a "specific mutation." A "specific mutation" will be described later.
[0152] Examples of wild-type phoS genes include the phoS genes of coryneform bacteria. Specific examples of phoS genes of coryneform bacteria include the phoS genes of C. glutamicum YDK010 strain, C. glutamicum ATCC13032 strain, C. glutamicum ATCC14067 strain, C. callunae, C. crenatum, and C. efficiens. The nucleotide sequence of the phoS gene of the C. glutamicum YDK010 strain is shown in SEQ ID NO: 1. The amino acid sequences of the wild-type PhoS proteins encoded by these phoS genes are shown in SEQ ID NOs: 2 to 7, respectively.
[0153] The wild-type phoS gene may be a variant of the wild-type phoS gene exemplified above, so long as it does not have the "specific mutation" and maintains its original function. Similarly, the wild-type PhoS protein may be a variant of the wild-type PhoS protein exemplified above, so long as it does not have the "specific mutation" and maintains its original function. In other words, the term "wild-type phoS gene" is not limited to the wild-type phoS gene exemplified above, but also includes its conservative variant that does not have the "specific mutation." Similarly, the term "wild-type PhoS protein" is not limited to the wild-type PhoS protein exemplified above, but also includes its conservative variant that does not have the "specific mutation." The above descriptions regarding the ncAA-aaRS gene and conservative variants of ncAA-aaRS apply mutatis mutandis to variants of the wild-type PhoS protein and wild-type phoS gene. For example, the wild-type phoS gene may be a gene that encodes a protein having an amino acid sequence in which one or more amino acids at one or more positions in the above amino acid sequence have been substituted, deleted, inserted, and / or added, so long as it does not have a "specific mutation" and the original function is maintained. Furthermore, for example, the wild-type phoS gene may be a gene that encodes a protein having an amino acid sequence that is 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire above amino acid sequence, so long as it does not have a "specific mutation" and the original function is maintained.
[0154] Note that "maintaining the original function" may mean that, in the case of a wild-type PhoS protein, a protein variant retains the function of a PhoS protein (e.g., the function of a protein consisting of the amino acid sequences set forth in SEQ ID NOS: 2 to 7). Furthermore, "maintaining the original function" may mean that, in the case of a wild-type PhoS protein, a protein variant retains the function of a sensor kinase in the PhoRS system. That is, "the function of a PhoS protein" may specifically mean the function of a sensor kinase in the PhoRS system. "The function of a sensor kinase in the PhoRS system" may specifically mean the function of conjugating with the PhoR protein, which is a response regulator, to elicit a response to environmental phosphate deficiency. More specifically, "the function of a sensor kinase in the PhoRS system" may mean the function of sensing environmental phosphate deficiency, undergoing autophosphorylation, and activating the PhoR protein by phosphoryl transfer.
[0155] Whether a PhoS protein variant functions as a sensor kinase of the PhoRS system can be confirmed, for example, by introducing a gene encoding the variant into a phoS gene-deficient strain of a coryneform bacterium and determining whether the responsiveness to phosphate deficiency is complemented. Complementation of the responsiveness to phosphate deficiency can be detected, for example, by improved growth under phosphate-deficient conditions or by induction of expression of a gene known to be induced under phosphate-deficient conditions (J. Bacteriol., 188, 724-732 (2006)). Examples of phoS gene-deficient strains of coryneform bacteria that can be used include the phoS gene-deficient strain of C. glutamicum YDK010 and the phoS gene-deficient strain of C. glutamicum ATCC13032.
[0156] In the wild-type PhoS protein, the histidine residue that undergoes autophosphorylation is preferably conserved. That is, the conservative mutation preferably occurs at an amino acid residue other than the histidine residue that undergoes autophosphorylation. The "histidine residue that undergoes autophosphorylation" refers to the histidine residue at position 276 in the wild-type PhoS protein. Furthermore, the wild-type PhoS protein preferably has a conserved sequence of, for example, the wild-type PhoS protein exemplified above. That is, the conservative mutation preferably occurs at an amino acid residue that is not conserved in, for example, the wild-type PhoS protein exemplified above.
[0157] The mutant PhoS protein has a "specific mutation" in the amino acid sequence of the wild-type PhoS protein as described above.
[0158] In other words, the mutant PhoS protein may be identical to the wild-type PhoS protein or a conservative variant thereof exemplified above, except for the "specific mutation." Specifically, for example, the mutant PhoS protein may be a protein having the amino acid sequence set forth in SEQ ID NOs: 2 to 7, except for the "specific mutation." Furthermore, specifically, for example, the mutant PhoS protein may be a protein having an amino acid sequence containing one or more amino acid substitutions, deletions, insertions, and / or additions in the amino acid sequence set forth in SEQ ID NOs: 2 to 7, except for the "specific mutation." Furthermore, specifically, for example, the mutant PhoS protein may be a protein having an amino acid sequence that is 80% or more, preferably 90% or more, more preferably 95% or more, more preferably 97% or more, and particularly preferably 99% or more identical to the amino acid sequence set forth in SEQ ID NOs: 2 to 7, except for the "specific mutation."
[0159] In other words, the mutant PhoS protein may be a variant of the wild-type PhoS protein exemplified above that has a "specific mutation" and further contains conservative mutations at positions other than the "specific mutation." Specifically, for example, the mutant PhoS protein may be a protein having an amino acid sequence shown in SEQ ID NOs: 2 to 7 that has a "specific mutation" and further contains one or more amino acid substitutions, deletions, insertions, and / or additions at positions other than the "specific mutation."
[0160] The mutant phoS gene is not particularly limited as long as it encodes the above-mentioned mutant PhoS protein.
[0161] The "specific mutation" contained in the mutant PhoS protein will be explained below.
[0162] The "specific mutation" is not particularly limited as long as it changes the amino acid sequence of the wild-type PhoS protein as described above and is effective for the secretory production of an ncAA-containing protein.
[0163] The "specific mutation" is preferably a mutation that improves the secretory production of the ncAA-containing protein. "Improving the secretory production of the ncAA-containing protein" means that a coryneform bacterium (modified strain) modified to have a mutant phoS gene can secrete and produce a greater amount of ncAA-containing protein than an unmodified strain. The "unmodified strain" refers to a control strain that does not have a mutation in the phoS gene, i.e., a control strain that does not have a mutant phoS gene, and may be, for example, a wild-type strain or a parent strain. The term "secreting and producing a greater amount of ncAA-containing protein than an unmodified strain" is not particularly limited as long as the secretory production of the ncAA-containing protein is increased compared to an unmodified strain. For example, it may mean secreting and producing an amount of ncAA-containing protein that is accumulated in the medium and / or on the bacterial cell surface that is preferably 1.1-fold or more, more preferably 1.2-fold or more, even more preferably 1.3-fold or more, even more preferably 2-fold or more, and particularly preferably 5-fold or more, of the unmodified strain. Furthermore, "secreting and producing a greater amount of ncAA-containing protein than a non-modified strain" may mean that when the unconcentrated culture supernatant of the unmodified strain is subjected to SDS-PAGE and stained with CBB, the ncAA-containing protein cannot be detected, but when the unconcentrated culture supernatant of the modified strain is subjected to SDS-PAGE and stained with CBB, the ncAA-containing protein can be detected. Note that "improving the secretory production of ncAA-containing proteins" does not necessarily mean improving the secretory production of all ncAA-containing proteins; it is sufficient to improve the secretory production of the ncAA-containing protein set as the target for secretory production. Specifically, "improving the secretory production of ncAA-containing proteins" may mean, for example, improving the secretory production of the ncAA-containing proteins described in the Examples.
[0164] Whether a certain mutation improves the secretory production of an ncAA-containing protein can be confirmed, for example, by creating a strain based on a strain belonging to the coryneform bacteria that has been modified to have a gene encoding a PhoS protein with the mutation, quantifying the amount of ncAA-containing protein secreted when the modified strain is cultured in a medium, and comparing this with the amount of ncAA-containing protein secreted when the unmodified strain (unmodified strain) is cultured in a medium.
[0165] The amino acid sequence change is preferably an amino acid residue substitution. That is, the "specific mutation" is preferably a substitution of any amino acid residue in the wild-type PhoS protein with another amino acid residue. The amino acid residue substituted by the "specific mutation" may be a single residue, or a combination of two or more residues. The amino acid residue substituted by the "specific mutation" may preferably be an amino acid residue other than an autophosphorylated histidine residue. The amino acid residue substituted by the "specific mutation" may more preferably be an amino acid residue in the HisKA domain other than an autophosphorylated histidine residue. The "autophosphorylated histidine residue" refers to the histidine residue at position 276 of the wild-type PhoS protein. The "HisKA domain" refers to the region consisting of amino acid residues 266 to 330 of the wild-type PhoS protein. The amino acid residue substituted by the "specific mutation" may particularly preferably be the tryptophan residue at position 302 (W302) of the wild-type PhoS protein.
[0166] In the above mutations, the substituted amino acid residues include those other than the original amino acid residues among K (Lys), R (Arg), H (His), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), F (Phe), W (Trp), Y (Tyr), C (Cys), M (Met), D (Asp), E (Glu), N (Asn), and Q (Gln). The substituted amino acid residue can be selected, for example, to improve the secretory production of the ncAA-containing protein.
[0167] When W302 is substituted, the amino acid residue after substitution can be an amino acid residue other than aromatic amino acids and histidine. Specific examples of "amino acid residues other than aromatic amino acids and histidine" include K (Lys), R (Arg), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), C (Cys), M (Met), D (Asp), E (Glu), N (Asn), and Q (Gln). More specific examples of "amino acid residues other than aromatic amino acids and histidine" include K (Lys), A (Ala), V (Val), S (Ser), C (Cys), M (Met), D (Asp), and N (Asn).
[0168] The "specific mutation" in the phoS gene refers to a mutation in the nucleotide sequence that causes the above-mentioned "specific mutation" in the encoded PhoS protein.
[0169] The "amino acid residue at position X of the wild-type PhoS protein" refers to the amino acid residue corresponding to the amino acid residue at position X in SEQ ID NO: 2. For example, "W302" refers to the amino acid residue corresponding to the tryptophan residue at position 302 in SEQ ID NO: 2. The above amino acid residue positions indicate relative positions, and their absolute positions may change due to amino acid deletion, insertion, addition, etc. For example, if a wild-type PhoS protein consisting of the amino acid sequence set forth in SEQ ID NO: 2 is deleted or inserted at a position N-terminal to position X, the original amino acid residue at position X becomes the X-1 or X+1 amino acid residue, respectively, counting from the N-terminus, and is still considered to be the "amino acid residue at position X of the wild-type PhoS protein." Specifically, in the amino acid sequences of the wild-type PhoS proteins set forth in SEQ ID NOs: 2 to 7, "W302" refers to the tryptophan residues at positions 302, 302, 302, 321, 275, and 286, respectively. In the amino acid sequences of the wild-type PhoS proteins shown in SEQ ID NOs: 2 to 7, the "histidine residue at position 276 of the wild-type PhoS protein (the histidine residue that is autophosphorylated)" refers to the histidine residues at positions 276, 276, 276, 295, 249, and 260, respectively. In the amino acid sequences of the wild-type PhoS proteins shown in SEQ ID NOs: 2 to 7, the "region consisting of amino acid residues at positions 266 to 330 of the wild-type PhoS protein (HisKA domain)" refers to the regions consisting of amino acid residues at positions 266 to 330, 266 to 330, 266 to 330, 285 to 349, 239 to 303, and 250 to 314, respectively.
[0170] Note that "W302" as used herein is typically a tryptophan residue, but does not necessarily have to be a tryptophan residue. That is, when the wild-type PhoS protein has an amino acid sequence other than those set forth in SEQ ID NOS: 2 to 7, "W302" may not be a tryptophan residue. Therefore, for example, "a mutation in which W302 is substituted with a cysteine residue" is not limited to a mutation in which "W302" is a tryptophan residue and the tryptophan residue is substituted with a cysteine residue. It also encompasses a mutation in which "W302" is K (Lys), R (Arg), H (His), A (Ala), V (Val), L (Leu), I (Ile), G (Gly), S (Ser), T (Thr), P (Pro), F (Phe), Y (Tyr), M (Met), D (Asp), E (Glu), N (Asn), or Q (Gln) and the amino acid residue is substituted with a cysteine residue. The same applies to other mutations.
[0171] In the amino acid sequence of any PhoS protein, which amino acid residue corresponds to the amino acid residue at position X in SEQ ID NO: 2 can be determined by aligning the amino acid sequence of the PhoS protein with the amino acid sequence of SEQ ID NO: 2. Alignment can be performed using, for example, known genetic analysis software. Specific examples of such software include DNASIS manufactured by Hitachi Solutions and GENETYX manufactured by Genetyx (Elizabeth C. Tyler et al., Computers and Biomedical Research, 24(1), 72-96, 1991; Barton GJ et al., Journal of molecular biology, 198(2), 327-37, 1987).
[0172] A mutant phoS gene can be obtained, for example, by modifying a wild-type phoS gene so that the encoded PhoS protein has the above-mentioned "specific mutation." The wild-type phoS gene that is the source of the modification can be obtained, for example, by cloning from an organism having a wild-type phoS gene or by chemical synthesis. A mutant phoS gene can also be obtained without using a wild-type phoS gene. For example, a mutant phoS gene can be obtained directly by chemical synthesis. The obtained mutant phoS gene can be further modified and used.
[0173] Genetic modification can be performed using known techniques. For example, a desired mutation can be introduced into a desired site in DNA by site-directed mutagenesis. Examples of site-directed mutagenesis include PCR-based methods (Higuchi, R., 61, in PCR technology, Erlich, HA Eds., Stockton Press (1989); Carter, P., Meth. in Enzymol., 154, 382 (1987)) and phage-based methods (Kramer, W. and Frits, HJ, Meth. in Enzymol., 154, 350 (1987); Kunkel, TA et al., Meth. in Enzymol., 154, 367 (1987)).
[0174] Hereinafter, a method for modifying a coryneform bacterium so that it has a mutant phoS gene will be described.
[0175] Modification of a coryneform bacterium to have a mutant phoS gene can be achieved by introducing the mutant phoS gene into the coryneform bacterium. Modification of a coryneform bacterium to have a mutant phoS gene can also be achieved by introducing the above-mentioned "specific mutation" into the phoS gene on the chromosome of the coryneform bacterium. Introduction of a mutation into a gene on the chromosome can be achieved by natural mutation, mutagen treatment, or genetic engineering techniques.
[0176] The method for introducing a mutant phoS gene into a coryneform bacterium is not particularly limited. In the bacterium of the present invention, the mutant phoS gene may be maintained in an expressible state under the control of a promoter that functions in the coryneform bacterium. The promoter may be a promoter derived from the host or a heterologous promoter. The promoter may be a promoter native to the phoS gene or a promoter for another gene. In the bacterium of the present invention, the mutant phoS gene may be present on an extrachromosomally autonomously replicating vector such as a plasmid, or may be integrated into the chromosome. The bacterium of the present invention may have only one copy of the mutant phoS gene, or may have two or more copies. The bacterium of the present invention may have only one type of mutant phoS gene, or may have two or more types of mutant phoS genes. Introduction of the mutant phoS gene can be carried out, for example, in the same manner as the introduction of genes in the methods for increasing gene expression or the introduction of a gene construct for secretory expression, as described below.
[0177] The bacterium of the present invention may or may not have a wild-type phoS gene, but preferably does not have one.
[0178] A coryneform bacterium lacking the wild-type phoS gene can be obtained by disrupting the wild-type phoS gene on its chromosome. Disruption of the wild-type phoS gene can be carried out by known techniques. Specifically, for example, the wild-type phoS gene can be disrupted by deleting part or all of the promoter region and / or coding region of the wild-type phoS gene.
[0179] Furthermore, by replacing the wild-type phoS gene on the chromosome with a mutant phoS gene, it is possible to obtain a coryneform bacterium that has been modified to have the mutant phoS gene but not the wild-type phoS gene. Examples of methods for performing such gene replacement include methods using linear DNA, such as "Red-driven integration" (Datsenko, K. A., and Wanner, BL Proc. Natl. Acad. Sci. USA 97:6640-6645 (2000)), a method combining Red-driven integration with an excision system derived from λ phage (Cho, EH, Gumport, RI, Gardner, JFJ Bacteriol. 184:5200-5203 (2002)) (see WO2005 / 010175), methods using a plasmid containing a temperature-sensitive replication origin, methods using a conjugatively transferable plasmid, and methods using a suicide vector that does not have a replication origin that functions in the host (U.S. Pat. No. 6,303,383, JP 05-007491 A).
[0180] The PhoS protein functions in conjunction with the PhoR protein, a response regulator, i.e., it triggers a response to environmental phosphate deficiency. Therefore, the bacterium of the present invention has a phoR gene so that the mutant PhoS protein can function. The phoR gene encodes the PhoR protein, which is a response regulator of the PhoRS system. "Having the phoR gene" is also referred to as "having the PhoR protein." Generally, it is sufficient for the PhoR protein inherent in the bacterium of the present invention to function in conjunction with the mutant PhoS protein. Alternatively, an appropriate phoR gene may be introduced into the bacterium of the present invention in addition to or instead of the phoR gene inherent in the bacterium of the present invention. The phoR gene to be introduced is not particularly limited, as long as it encodes a PhoR protein that functions in conjunction with the mutant PhoS protein.
[0181] Examples of the phoR gene include the phoR gene of coryneform bacteria. Specific examples of the phoR gene of coryneform bacteria include the phoR genes of C. glutamicum YDK010 strain, C. glutamicum ATCC13032 strain, C. glutamicum ATCC14067 strain, C. callunae, C. crenatum, and C. efficiens. The nucleotide sequence of the phoR gene and the amino acid sequence of the PhoR protein of the C. glutamicum ATCC13032 strain are shown in SEQ ID NOs: 8 and 9, respectively.
[0182] The phoR gene may be a variant of the exemplified phoR gene, as long as the original function is maintained. Similarly, the PhoR protein may be a variant of the exemplified PhoR protein, as long as the original function is maintained. In other words, the term "phoR gene" encompasses not only the exemplified phoR genes but also their conservative variants. Similarly, the term "PhoR protein" encompasses not only the exemplified PhoR proteins but also their conservative variants. The above descriptions regarding the ncAA-aaRS gene and conservative variants of ncAA-aaRS can be applied mutatis mutandis to variants of the PhoR protein and phoR gene. For example, the phoR gene may be a gene encoding a protein having an amino acid sequence in which one or several amino acids at one or several positions in the above amino acid sequence have been substituted, deleted, inserted, and / or added, as long as the original function is maintained. For example, the phoR gene may be a gene encoding a protein having an amino acid sequence that is 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire amino acid sequence, so long as the original function is maintained. Note that, in the case of PhoR protein, "maintaining the original function" may mean that the protein variant retains the function of the PhoR protein (e.g., the function of a protein consisting of the amino acid sequence set forth in SEQ ID NO: 9). Also, in the case of PhoR protein, "maintaining the original function" may mean that the protein variant retains the function as a response regulator of the PhoRS system. Specifically, "the function as a PhoR protein" may mean the function as a response regulator of the PhoRS system. Specifically, "the function as a response regulator of the PhoRS system" may mean the function of conjugating with the sensor kinase PhoS protein to elicit a response to environmental phosphate deficiency.More specifically, the "function of the PhoRS system as a response regulator" may be a function that is activated by phosphoryl transfer from the autophosphorylated PhoS protein upon sensing phosphate deficiency in the environment, thereby controlling the expression of genes that respond to phosphate deficiency in the environment.
[0183] Whether a PhoR protein variant functions as a response regulator of the PhoRS system can be confirmed, for example, by introducing a gene encoding the variant into a phoR gene-deficient strain of a coryneform bacterium and determining whether the responsiveness to phosphate deficiency is complemented. Complementation of the responsiveness to phosphate deficiency can be detected, for example, by improved growth under phosphate deficiency conditions or by induction of expression of a gene known to be induced under phosphate deficiency conditions (J. Bacteriol., 188, 724-732 (2006)). Examples of phoR gene-deficient strains of coryneform bacteria that can be used include the phoR gene-deficient strain of C. glutamicum YDK010 and the phoR gene-deficient strain of C. glutamicum ATCC13032.
[0184] <1-4-2> Decreased Activity of Cell Surface Protein The bacterium of the present invention may have reduced activity of a cell surface protein. Specifically, the bacterium of the present invention may have reduced activity of a cell surface protein compared to a non-modified strain. "Decreased activity of a cell surface protein" may particularly mean a reduction in the number of molecules of the cell surface protein per cell. Cell surface proteins and the genes encoding them are described below.
[0185] Cell surface proteins are proteins that make up the cell surface (S-layer) of bacteria and archaea. Examples of cell surface proteins of coryneform bacteria include PS1 and PS2 (CspB) of C. glutamicum (JP Patent Publication No. 6-502548) and SlpA (CspA) of C. stationis (JP Patent Publication No. 10-108675). Of these, it is preferable to reduce the activity of the PS2 protein.
[0186] The nucleotide sequence of the cspB gene of C. glutamicum ATCC13869 and the amino acid sequence of the PS2 protein (CspB protein) encoded by the gene are shown in SEQ ID NOs: 10 and 11, respectively.
[0187] For example, the amino acid sequences of CspB homologs from 28 strains of C. glutamicum have been reported (J. Biotechnol., 112, 177-193 (2004)). The GenBank accession numbers of these 28 C. glutamicum strains and their cspB gene homologs in the NCBI database are shown below (the numbers in parentheses indicate the GenBank accession numbers).C. glutamicum ATCC13058(AY524990) C. glutamicum ATCC13744(AY524991) C. glutamicum ATCC13745(AY524992) C. glutamicum ATCC14017(AY524993) C. glutamicum ATCC14020(AY525009) C. glutamicum ATCC14067(AY524994) C. glutamicum ATCC14068(AY525010) C. glutamicum ATCC14747(AY525011) C. glutamicum ATCC14751(AY524995) C. glutamicum ATCC14752(AY524996) C. glutamicum ATCC14915(AY524997) C. glutamicum ATCC15243(AY524998) C. glutamicum ATCC15354(AY524999) C. glutamicum ATCC17965(AY525000) C. glutamicum ATCC17966(AY525001) C. glutamicum ATCC19223(AY525002) C. glutamicum ATCC19240(AY525012) C. glutamicum ATCC21341(AY525003) C. glutamicum ATCC21645(AY525004) C. glutamicum ATCC31808(AY525013) C. glutamicum ATCC31830(AY525007) C. glutamicum ATCC31832(AY525008) C. glutamicum LP-6(AY525014) C. glutamicum DSM20137(AY525015) C. glutamicum DSM20598(AY525016) C. glutamicum DSM46307(AY525017) C. glutamicum 22220(AY525005) C. glutamicum 22243(AY525006)。
[0188] Because the nucleotide sequences of genes encoding cell surface proteins may differ depending on the species or strain of coryneform bacteria, the genes encoding the cell surface proteins may be variants of the genes encoding the above-exemplified cell surface proteins, as long as the original function is maintained. Similarly, the cell surface proteins may be variants of the above-exemplified cell surface proteins, as long as the original function is maintained. That is, for example, the term "cspB gene" encompasses not only the above-exemplified cspB genes but also their conservative variants. Similarly, the term "CspB protein" encompasses not only the above-exemplified CspB proteins but also their conservative variants. The above descriptions regarding the ncAA-aaRS genes and conservative variants of ncAA-aaRS can be applied mutatis mutandis to variants of cell surface proteins and genes encoding them. For example, the genes encoding the cell surface proteins may be genes encoding proteins having an amino acid sequence in which one or several amino acids are substituted, deleted, inserted, and / or added at one or several positions in the above-exemplified amino acid sequence, as long as the original function is maintained. Furthermore, for example, the gene encoding the cell surface protein may be a gene encoding a protein having an amino acid sequence that is 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire amino acid sequence, so long as the original function is maintained. Note that, in the case of a cell surface protein, "maintaining the original function" may mean, for example, that the cell surface protein has the property of increasing the secretory production of the ncAA-containing protein when its activity is reduced in a coryneform bacterium compared to a non-modified strain.
[0189] The term "the property of increasing the secretory production of an ncAA-containing protein when activity is reduced in a coryneform bacterium compared to a non-modified strain" refers to the property of conferring to the coryneform bacterium the ability to secrete and produce a greater amount of ncAA-containing protein than a non-modified strain when activity is reduced in the coryneform bacterium. A "non-modified strain" refers to a control strain in which the activity of the cell surface protein is not reduced, and may be, for example, a wild-type strain or a parent strain. The term "secreting and producing a greater amount of ncAA-containing protein than a non-modified strain" is not particularly limited as long as the amount of ncAA-containing protein secreted and produced is increased compared to a non-modified strain. For example, the term may refer to the secretory production of an ncAA-containing protein in a medium and / or on the bacterial cell surface, with the amount accumulated in the medium and / or on the bacterial cell surface being preferably at least 1.1 times, more preferably at least 1.2 times, even more preferably at least 1.3 times, and particularly preferably at least 2 times that of a non-modified strain. Furthermore, "secreting and producing a greater amount of ncAA-containing protein than a non-modified strain" may also mean that when the unconcentrated culture supernatant of the non-modified strain is subjected to SDS-PAGE and stained with CBB, the ncAA-containing protein cannot be detected, but when the unconcentrated culture supernatant of the modified strain is subjected to SDS-PAGE and stained with CBB, the ncAA-containing protein can be detected.
[0190] Whether a protein has the property of increasing the secretory production of an ncAA-containing protein when its activity is reduced in a coryneform bacterium compared to an unmodified strain can be confirmed by creating a strain based on a coryneform bacterium that has been modified to reduce the activity of the protein, quantifying the amount of ncAA-containing protein secreted when the modified strain is cultured in a medium, and comparing this with the amount of ncAA-containing protein secreted when the unmodified strain (unmodified strain) is cultured in a medium.
[0191] "Decreased cell surface protein activity" includes cases where a coryneform bacterium has been modified to reduce the activity of the cell surface protein, as well as cases where the activity of the cell surface protein is originally reduced in the coryneform bacterium. "Decreased cell surface protein activity in the coryneform bacterium" also includes cases where the coryneform bacterium does not originally have a cell surface protein. That is, an example of a coryneform bacterium with reduced cell surface protein activity is a coryneform bacterium that does not originally have a cell surface protein. An example of a coryneform bacterium that does not originally have a cell surface protein is a coryneform bacterium that does not originally have a gene encoding a cell surface protein. Note that "coryneform bacterium does not originally have a cell surface protein" may mean that the coryneform bacterium does not originally have one or more proteins selected from cell surface proteins found in other strains of the species to which the coryneform bacterium belongs. For example, "C. glutamicum does not originally have cell surface proteins" may mean that the C. glutamicum strain does not originally have one or more proteins selected from cell surface proteins found in other C. glutamicum strains, i.e., PS1 and / or PS2 (CspB). An example of a coryneform bacterium that does not originally have cell surface proteins is C. glutamicum ATCC 13032, which does not originally have the cspB gene.
[0192] <1-4-3> Protein Secretion System The bacterium of the present invention has a protein secretion system. The bacterium of the present invention may inherently have a protein secretion system. The protein secretion system is not particularly limited, as long as it is capable of secreting an ncAA-containing protein. Examples of protein secretion systems include the Sec system (Sec system secretion apparatus) and the Tat system (Tat system secretion apparatus). The bacterium of the present invention may be modified to increase the activity of the protein secretion system (e.g., the Tat system secretion apparatus). Specifically, the bacterium of the present invention may be modified to increase the activity of the protein secretion system (e.g., the Tat system secretion apparatus) compared to an unmodified strain. The activity of the Tat system secretion apparatus can be increased, for example, by increasing the expression of one or more genes selected from genes encoding the Tat system secretion apparatus. More specifically, the bacterium of the present invention may be modified to increase the expression of one or more genes selected from genes encoding the Tat system secretion apparatus. Increased activity of the Tat system secretion apparatus is particularly advantageous when secreting and producing ncAA-containing proteins using a Tat system-dependent signal peptide. A method for increasing the expression of genes encoding the Tat-based secretion system is described in Japanese Patent No. 4730302.
[0193] Genes encoding the Tat system secretion apparatus include the tatA gene, the tatB gene, the tatC gene, and the tatE gene.
[0194] Specific examples of genes encoding the Tat system secretion apparatus include the tatA gene, tatB gene, and tatC gene of C. glutamicum. The tatA gene, tatB gene, and tatC gene of C. glutamicum ATCC 13032 correspond to the complementary sequence of the sequence from positions 1571065 to 1571382, the sequence from positions 1167110 to 1167580, and the sequence from positions 1569929 to 1570873, respectively, in the genome sequence registered in the NCBI database as GenBank accession NC_003450 (VERSION NC_003450.3 GI:58036263). The TatA, TatB, and TatC proteins of C. glutamicum ATCC 13032 have been registered as GenBank accession numbers NP_600707 (version NP_600707.1 GI:19552705, locus_tag="NCgl1434"), NP_600350 (version NP_600350.1 GI:19552348, locus_tag="NCgl1077"), and NP_600706 (version NP_600706.1 GI:19552704, locus_tag="NCgl1433"), respectively. The nucleotide sequences of the tatA gene, tatB gene, and tatC gene of C. glutamicum ATCC 13032 and the amino acid sequences of the TatA protein, TatB protein, and TatC protein are shown in SEQ ID NOs: 12 to 17.
[0195] Specific examples of genes encoding the Tat system secretion apparatus include the tatA, tatB, tatC, and tatE genes of E. coli. The tatA, tatB, tatC, and tatE genes of E. coli K-12 MG1655 correspond to the sequences of positions 4019968 to 4020237, 4020241 to 4020756, 4020759 to 4021535, and 658170 to 658373, respectively, in the genome sequence registered in the NCBI database as GenBank accession NC_000913 (VERSION NC_000913.2 GI:49175990). The TatA, TatB, TatC, and TatE proteins of E. coli K-12 MG1655 were identified in GenBank accession numbers NP_418280 (version NP_418280.4 GI:90111653, locus_tag="b3836"), YP_026270 (version YP_026270.1 GI:49176428, locus_tag="b3838"), NP_418282 (version NP_418282.1 GI:16131687, locus_tag="b3839"), and NP_415160 (version NP_415160.1 GI:16131687), respectively. It is registered as GI:16128610, locus_tag="b0627").
[0196] The gene encoding the Tat system secretion apparatus may be a variant of the gene encoding the Tat system secretion apparatus exemplified above, so long as the original function is maintained. Similarly, the Tat system secretion apparatus may be a variant of the Tat system secretion apparatus exemplified above, so long as the original function is maintained. That is, for example, the terms "tatA gene," "tatB gene," "tatC gene," and "tatE gene" encompass the tatA gene, tatB gene, tatC gene, and tatE gene exemplified above, respectively, as well as conservative variants thereof. Similarly, the terms "TatA protein," "TatB protein," "TatC protein," and "TatE protein" encompass the TatA protein, TatB protein, TatC protein, and TatE protein exemplified above, as well as conservative variants thereof, respectively. The above descriptions regarding the ncAA-aaRS gene and conservative variants of ncAA-aaRS apply mutatis mutandis to variants of the Tat system secretion apparatus and the gene encoding it. For example, a gene encoding a Tat system secretion system may be a gene encoding a protein having an amino acid sequence in which one or more amino acids at one or more positions have been substituted, deleted, inserted, and / or added, as long as the original function is maintained. Furthermore, for example, a gene encoding a Tat system secretion system may be a gene encoding a protein having an amino acid sequence that is 80% or more, preferably 90% or more, more preferably 95% or more, even more preferably 97% or more, and particularly preferably 99% or more identical to the entire amino acid sequence, as long as the original function is maintained. Note that, in the case of a Tat system secretion system, "maintaining the original function" may mean that the protein having a Tat system-dependent signal peptide added to its N-terminus is capable of being secreted extracellularly.
[0197] Increased activity of the Tat secretion apparatus can be confirmed, for example, by confirming an increase in the secretory production of a protein having a Tat-dependent signal peptide added to its N-terminus. The secretory production of a protein having a Tat-dependent signal peptide added to its N-terminus may be increased, for example, by 1.5-fold or more, 2-fold or more, or 3-fold or more compared to that of an unmodified strain.
[0198] <1-5> Methods for Increasing Protein Activity Methods for increasing protein activity (including methods for increasing gene expression) are described below.
[0199] "Increased protein activity" means that the activity of the protein is increased compared to that of an unmodified strain. Specifically, "increased protein activity" means that the activity of the protein per cell is increased compared to that of an unmodified strain. Here, "unmodified strain" refers to a control strain that has not been modified to reduce the activity of the target protein. Examples of unmodified strains include wild-type strains and parent strains. Specific examples of unmodified strains include the type strains of each bacterial species. Specific examples of unmodified strains also include the strains exemplified in the description of coryneform bacteria. That is, in one embodiment, the activity of the protein may be increased compared to that of the type strain (i.e., the type strain of the species to which the bacterium of the present invention belongs). In another embodiment, the activity of the protein may be increased compared to that of the C. glutamicum ATCC 13869 strain. In another embodiment, the activity of the protein may be increased compared to that of the C. glutamicum ATCC 13032 strain. In another embodiment, the activity of the protein may be increased compared to that of C. glutamicum AJ12036 (FERM BP-734). In another embodiment, the activity of the protein may be increased compared to that of the C. glutamicum YDK010 strain. "Increased protein activity" is also referred to as "enhanced protein activity." More specifically, "increased protein activity" may mean an increase in the number of molecules of the protein per cell and / or an increase in the function of the protein per molecule compared to a non-modified strain. That is, the "activity" in "increased protein activity" is not limited to the catalytic activity of the protein, but may also refer to the transcription amount (mRNA amount) or translation amount (protein amount) of the gene encoding the protein. The "number of protein molecules per cell" may mean the average number of molecules of the protein per cell. Furthermore, "increasing the activity of a protein" includes not only increasing the activity of a target protein in a strain that originally has the activity of that protein, but also imparting the activity of that protein to a strain that does not originally have the activity of that protein.Furthermore, as long as the resulting protein activity is increased, the activity of a suitable target protein may be imparted after reducing or eliminating the activity of a target protein that the host naturally possesses.
[0200] The degree of increase in protein activity is not particularly limited as long as the protein activity is increased compared to that of an unmodified strain. The protein activity may be increased, for example, by 1.5 times or more, 2 times or more, or 3 times or more compared to that of an unmodified strain. Furthermore, if the unmodified strain does not have the activity of the target protein, the protein may be produced by introducing a gene encoding the protein, and for example, the protein may be produced to an extent that its activity can be measured.
[0201] Modifications that increase the activity of a protein can be achieved, for example, by increasing the expression of the gene encoding the protein. "Increased gene expression" means that the expression of the gene is increased compared to an unmodified strain such as a wild-type strain or a parent strain. "Increased gene expression" specifically means that the expression level of the gene per cell is increased compared to an unmodified strain. "Expression level of the gene per cell" may refer to the average expression level of the gene per cell. "Increased gene expression" may more specifically mean that the transcription level (mRNA level) of the gene is increased and / or the translation level (protein level) of the gene is increased. "Increased gene expression" is also referred to as "enhanced gene expression." Gene expression may be increased, for example, by 1.5-fold or more, 2-fold or more, or 3-fold or more compared to an unmodified strain. "Increased gene expression" includes not only increasing the expression level of the target gene in a strain in which the target gene is originally expressed, but also expressing the gene in a strain in which the target gene is not originally expressed. That is, "gene expression is increased" may mean, for example, introducing a target gene into a strain that does not harbor the gene, thereby expressing the gene.
[0202] Increased gene expression can be achieved, for example, by increasing the copy number of the gene.
[0203] Increasing the copy number of a gene can be achieved by introducing the gene into a host chromosome. Introduction of a gene into a chromosome can be achieved, for example, by homologous recombination (Miller, JH, Experiments in Molecular Genetics, 1972, Cold Spring Harbor Laboratory). Gene introduction methods that utilize homologous recombination include, for example, methods using linear DNA such as Red-driven integration (Datsenko, K. A., and Wanner, BL, Proc. Natl. Acad. Sci. USA 97:6640-6645 (2000)), methods using plasmids containing a temperature-sensitive replication origin, methods using conjugatively transferable plasmids, methods using suicide vectors lacking a replication origin that functions in the host, and transduction methods using phages. Specifically, a host can be transformed with recombinant DNA containing a target gene, and the gene can be introduced into the host chromosome by homologous recombination with the target site on the host chromosome. The structure of the recombinant DNA used for homologous recombination is not particularly limited as long as it allows homologous recombination to occur in the desired manner. For example, a host can be transformed with linear DNA containing a target gene, with base sequences at both ends of the gene that are homologous to the target site on the chromosome, respectively, and homologous recombination can occur upstream and downstream of the target site, thereby replacing the target site with the gene. The recombinant DNA used for homologous recombination may contain a marker gene for selecting transformants. Only one copy of the gene may be introduced, or two or more copies may be introduced. For example, multiple copies of a gene can be introduced into a chromosome by performing homologous recombination targeting a base sequence that exists in multiple copies on a chromosome. Examples of sequences that exist in multiple copies on a chromosome include repetitive DNA sequences and inverted repeats at both ends of transposons. Homologous recombination may also be performed targeting an appropriate sequence on a chromosome, such as a gene unnecessary for the production of a target substance.Genes can also be randomly introduced into chromosomes using transposons or Mini-Mu (Japanese Patent Laid-Open No. 2-109985, US Pat. No. 5,882,888, EP805867B1). Such chromosome modification techniques using homologous recombination are not limited to the introduction of target genes, but can also be used for any chromosome modification, such as modification of expression regulatory sequences.
[0204] The introduction of the target gene into the chromosome can be confirmed by Southern hybridization using a probe having a sequence complementary to all or part of the gene, or by PCR using primers prepared based on the sequence of the gene.
[0205] The copy number of a gene can also be increased by introducing a vector containing the gene into a host. For example, a DNA fragment containing a target gene can be ligated to a vector that functions in the host to construct an expression vector for the gene, and the host can be transformed with the expression vector to increase the copy number of the gene. A DNA fragment containing a target gene can be obtained, for example, by PCR using the genomic DNA of a microorganism containing the target gene as a template. A vector capable of autonomous replication within host cells can be used. The vector is preferably a multicopy vector. Furthermore, the vector preferably contains a marker such as an antibiotic resistance gene for the selection of transformants. The vector may also contain a promoter or terminator for expressing the inserted gene. The vector may be, for example, a bacterial plasmid-derived vector, a yeast plasmid-derived vector, a bacteriophage-derived vector, a cosmid, or a phagemid. Specific examples of vectors capable of autonomous replication in coryneform bacteria include pHM1519 (Agric. Biol. Chem., 48, 2901-2903 (1984)); pAM330 (Agric. Biol. Chem., 48, 2901-2903 (1984)); plasmids having drug resistance genes improved from these; pCRY30 (JP-A-3-210184); pCRY21, pCRY2KE, pCRY2KX, pCRY31, pCRY3KE, and pCRY3KX (JP-A-2-72876, U.S. Pat. No. 5,185,262); pCRY2 and pCRY3 (JP-A-1-191686); pAJ655, pAJ611, and pAJ1844 (JP-A-1983) -192900); pCG1 (JP 57-134500); pCG2 (JP 58-35197); pCG4 and pCG11 (JP 57-183799); pVK7 (JP 10-215883); pVK9 (US2006-0141588); pVC7 (JP 9-070291); pVS7 (WO2013 / 069634); pPK4 (JP 9-322774); pPK5 (WO2018 / 074579).Specific examples of vectors capable of autonomous replication in coryneform bacteria include pVC7N, a variant of pVC7 (Shuhei Hashiro et al., High copy number mutants derived from Corynebacterium glutamicum cryptic plasmid pAM330 and copy number control, J Biosci Bioeng, 2019 May;127(5):529-538.), and pVC7H1, pVC7H2, pVC7H3, pVC7H4, pVC7H5, pVC7H6, and pVC7H7 (WO2018 / 179834). Specific examples of vectors capable of autonomous replication in coryneform bacteria include pPK4H1, pPK4H2, pPK4H3, pPK4H4, pPK4H5, and pPK4H6 (WO2018 / 179834). Examples of vectors include pVC-based vectors and pPK-based vectors. pVC-based vectors include pVC7 and its variants, vectors in which the antibiotic resistance gene has been replaced with another antibiotic resistance gene, and vectors with 90% or more, 95% or more, 97% or more, or 99% or more nucleotide sequence identity therewith. pPK-based vectors include pPK4, pPK5, and their variants, vectors in which the antibiotic resistance gene has been replaced with another antibiotic resistance gene, and vectors with 90% or more, 95% or more, 97% or more, or 99% or more nucleotide sequence identity therewith. More particularly, examples of vectors include pVC7, pVC7N, pPK4, and pPK5.
[0206] When a gene is introduced, it is sufficient that the gene is retained in the host in an expressible manner. Specifically, it is sufficient that the gene is retained so that it is expressed under the control of a promoter that functions in the host. The promoter is not particularly limited as long as it functions in the host. A "promoter that functions in the host" refers to a promoter that has promoter activity in the host. The promoter may be a promoter derived from the host or a heterologous promoter. The promoter may be a promoter native to the gene to be introduced or a promoter of another gene. Furthermore, the promoter may be inducible or constitutive for gene expression.
[0207] Examples of promoters that can be used in coryneform bacteria include promoters of genes involved in the glycolysis pathway, the pentose phosphate pathway, the TCA cycle, the amino acid biosynthesis pathway, and cell surface proteins. Specific examples of promoters for amino acid biosynthesis genes include the glutamate dehydrogenase gene for glutamate biosynthesis, the glutamine synthetase gene for glutamine synthesis, the aspartokinase gene for lysine biosynthesis, the homoserine dehydrogenase gene for threonine biosynthesis, the acetohydroxyacid synthase gene for isoleucine and valine biosynthesis, the 2-isopropylmalate synthase gene for leucine biosynthesis, the glutamate kinase gene for proline and arginine biosynthesis, the phosphoribosyl-ATP pyrophosphorylase gene for histidine biosynthesis, the deoxyarabinoheptulosonate phosphate (DAHP) synthase gene for aromatic amino acid biosynthesis such as tryptophan, tyrosine, and phenylalanine, and the phosphoribosylpyrophosphate (PRPP) amidotransferase gene, inosinate dehydrogenase gene, and guanylate synthase gene for nucleic acid biosynthesis such as inosinate and guanylate. Furthermore, promoters that can be used in coryneform bacteria also include the stronger promoters described below.
[0208] A terminator for terminating transcription can be placed downstream of the gene. The terminator is not particularly limited as long as it functions in the host. The terminator may be a terminator derived from the host or a heterologous terminator. The terminator may be a terminator inherent to the gene to be introduced or a terminator of another gene. Specific examples of terminators include the terminator of bacteriophage BFK20.
[0209] Vectors, promoters, and terminators that can be used in various microorganisms are described in detail in, for example, "Basic Microbiology Lectures 8: Genetic Engineering, Kyoritsu Shuppan, 1987," and they can be used.
[0210] Furthermore, when two or more genes are introduced, it is sufficient that each gene is maintained in an expressible state in the host. For example, two or more genes may all be maintained on a single expression vector, or all may be maintained on a chromosome. Two or more genes may be maintained separately on multiple expression vectors, or may be maintained separately on a single or multiple expression vectors and on a chromosome. Two or more genes may constitute an operon and be introduced. For example, the tRNA(ncAA) gene and the ncAA-aaRS gene may or may not be maintained on a single expression vector. For example, the tRNA(ncAA) gene and the ncAA-containing protein gene may or may not be maintained on a single expression vector. For example, the ncAA-aaRS gene and the ncAA-containing protein gene may or may not be maintained on a single expression vector. For example, the tRNA(ncAA) gene, the ncAA-aaRS gene, and the ncAA-containing protein gene may or may not be maintained on a single expression vector. For example, multiple copies of a tRNA(ncAA) gene may be carried on multiple expression vectors, respectively. For example, multiple copies of an ncAA-aaRS gene may be carried on multiple expression vectors, respectively. For example, multiple copies of an ncAA-containing protein gene may be carried on multiple expression vectors, respectively.
[0211] That is, the bacterium of the present invention may have a single expression vector carrying, for example, an ncAA-containing protein gene (specifically, a gene construct for secretory expression), a tRNA(ncAA) gene, and an ncAA-aaRS gene, and the single expression vector may be, for example, a pPK-based vector such as pPK4 or pPK5.
[0212] Furthermore, the bacterium of the present invention may have, for example, a first expression vector carrying an ncAA-containing protein gene (specifically, a gene construct for secretory expression) and a second expression vector carrying a tRNA(ncAA) gene and an ncAA-aaRS gene. The first expression vector may further carry a tRNA(ncAA) gene and / or an ncAA-aaRS gene. The first expression vector may be, for example, a pPK-based vector such as pPK4 or pPK5. The second expression vector may be, for example, a pVC-based vector such as pVC7 or pVC7N.
[0213] The gene to be introduced is not particularly limited as long as it encodes a protein that functions in the host. The gene to be introduced may be a gene derived from the host or a gene derived from a heterologous species. The gene to be introduced can be obtained, for example, by PCR using primers designed based on the nucleotide sequence of the gene and the genomic DNA of an organism carrying the gene or a plasmid carrying the gene as a template. Alternatively, the gene to be introduced may be totally synthesized based on the nucleotide sequence of the gene (Gene, 60(1), 115-127 (1987)). The obtained gene can be used as is or after appropriate modification. In other words, by modifying the gene, its variants can be obtained. Gene modification can be performed using known techniques. For example, site-directed mutagenesis can be used to introduce a desired mutation into a target site in DNA. For example, site-directed mutagenesis can be used to modify the coding region of a gene so that the encoded protein contains substitutions, deletions, insertions, and / or additions of amino acid residues at specific sites. Site-directed mutagenesis methods include PCR-based methods (Higuchi, R., 61, in PCR Technology, Erlich, HA Eds., Stockton Press (1989); Carter, P., Meth. in Enzymol., 154, 382 (1987)) and phage-based methods (Kramer, W. and Frits, HJ, Meth. in Enzymol., 154, 350 (1987); Kunkel, TA et al., Meth. in Enzymol., 154, 367 (1987)). Alternatively, gene variants may be totally synthesized.
[0214] When a protein functions as a complex consisting of multiple subunits, all or only a portion of the multiple subunits may be modified, as long as the resulting protein activity is increased. That is, for example, when increasing protein activity by increasing gene expression, the expression of all or only a portion of the multiple genes encoding the subunits may be enhanced. It is usually preferable to enhance the expression of all of the multiple genes encoding the subunits. Furthermore, each subunit constituting the complex may be derived from a single organism, or from two or more different organisms, as long as the complex has the function of the target protein. That is, for example, genes encoding multiple subunits derived from the same organism may be introduced into a host, or genes derived from different organisms may be introduced into a host.
[0215] Increased gene expression can also be achieved by improving gene transcription efficiency. Increased gene expression can also be achieved by improving gene translation efficiency. Gene transcription efficiency and translation efficiency can be improved, for example, by modifying expression regulatory sequences. "Expression regulatory sequence" is a general term for sites that affect gene expression. Examples of expression regulatory sequences include promoters, Shine-Dalgarno (SD) sequences (also known as ribosome binding sites (RBS)), and spacer regions between the RBS and the start codon. Expression regulatory sequences can be determined using promoter search vectors or genetic analysis software such as GENETYX. These expression regulatory sequences can be modified, for example, by a method using a temperature-sensitive vector or the Red-driven integration method (WO2005 / 010175).
[0216] The efficiency of gene transcription can be improved, for example, by replacing the promoter of a gene on a chromosome with a stronger promoter. A "stronger promoter" refers to a promoter that enhances gene transcription compared to the wild-type promoter that is originally present. Examples of stronger promoters that can be used in coryneform bacteria include the artificially engineered P54-6 promoter (Appl. Microbiol. Biotechnol., 53, 674-679(2000)), pta, aceA, aceB, adh, and amyE promoters that can be induced by acetate, ethanol, pyruvate, etc., and other strong promoters such as cspB, SOD, and tuf (EF-Tu) promoters (Journal of Biotechnology 104 (2003) 311-323, Appl. Environ Microbiol. 2005 Dec;71(12):8587-96), lac promoter, tac promoter, trc promoter, F1 promoter, T7 promoter, T5 promoter, T3 promoter, and SP6 promoter. Furthermore, stronger promoters can be obtained by using various reporter genes to obtain highly active versions of existing promoters. For example, promoter activity can be increased by adjusting the -35 and -10 regions of the promoter region to resemble consensus sequences (WO 00 / 18935). Examples of highly active promoters include various tac-like promoters (Katashkina JI et al., Russian Federation Patent Application 2006134574). Methods for evaluating promoter strength and examples of strong promoters are described in Goldstein et al. (Prokaryotic promoters in biotechnology. Biotechnol. Annu. Rev., 1, 105-128 (1995)).
[0217] The translation efficiency of a gene can be improved, for example, by replacing the Shine-Dalgarno (SD) sequence (also known as the ribosome binding site (RBS)) of a gene on a chromosome with a stronger SD sequence. A "stronger SD sequence" refers to an SD sequence that improves mRNA translation compared to the native wild-type SD sequence. An example of a stronger SD sequence is the RBS of gene 10 from phage T7 (Olins PO et al., Gene, 1988, 73, 227-235). Furthermore, it is known that substitution, insertion, or deletion of several nucleotides in the spacer region between the RBS and the start codon, particularly in the sequence immediately upstream of the start codon (5'-UTR), significantly affects mRNA stability and translation efficiency. Gene translation efficiency can also be improved by modifying these sequences.
[0218] The translation efficiency of a gene can also be improved by, for example, codon modification. For example, the translation efficiency of a gene can be improved by replacing rare codons present in the gene with synonymous codons that are used more frequently. That is, the gene to be introduced may be modified to have optimal codons depending on the codon usage frequency of the host used. Codon substitution can be performed, for example, by site-directed mutagenesis. Alternatively, a gene fragment with substituted codons may be totally synthesized. The codon usage frequencies in various organisms are disclosed in the "Codon Usage Database" (http: / / www.kazusa.or.jp / codon; Nakamura, Y. et al., Nucl. Acids Res., 28, 292 (2000)).
[0219] Furthermore, increasing gene expression can also be achieved by amplifying regulators that increase gene expression, or by deleting or weakening regulators that decrease gene expression.
[0220] The above-mentioned methods for increasing gene expression may be used alone or in any combination.
[0221] Modifications that increase protein activity can also be achieved by, for example, enhancing the specific activity of the protein. Proteins with enhanced specific activity can be obtained, for example, by searching for and obtaining proteins from various organisms. Highly active proteins can also be obtained by introducing mutations into existing proteins. The mutations introduced may be, for example, substitutions, deletions, insertions, and / or additions of one or several amino acids at one or several positions in the protein. Mutations can be introduced, for example, by site-directed mutagenesis, as described above. Mutations can also be introduced by, for example, mutagenesis. Examples of mutagenesis include X-ray irradiation, ultraviolet irradiation, and treatment with mutagens such as N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), ethyl methanesulfonate (EMS), and methyl methanesulfonate (MMS). Alternatively, random mutations can be induced by directly treating DNA with hydroxylamine in vitro. Specific activity enhancement can be used alone or in any combination with the above-mentioned methods for enhancing gene expression.
[0222] The transformation method is not particularly limited, and conventionally known methods can be used, such as a method reported for Escherichia coli K-12 in which recipient cells are treated with calcium chloride to increase DNA permeability (Mandel, M. and Higa, A., J. Mol. Biol. 1970, 53, 159-162), or a method reported for Bacillus subtilis in which DNA is introduced into competent cells prepared from cells in the growth stage (Duncan, C.H., Wilson, G.A. and Young, F.E., 1977, Gene 1: 153-167). Alternatively, a method known for Bacillus subtilis, actinomycetes, and yeasts can be used in which the recipient cells are transformed into protoplasts or spheroplasts, which readily incorporate recombinant DNA, and then the recombinant DNA is introduced into the recipient cells (Chang, S. and Choen, SN, 1979, Mol. Gen. Genet. 168: 111-115; Bibb, MJ, Ward, JM, and Hopwood, OA, 1978, Nature 274: 398-400; Hinnen, A., Hicks, JB, and Fink, GR, 1978, Proc. Natl. Acad. Sci. USA 75: 1929-1933). Alternatively, an electric pulse method, as reported for coryneform bacteria (JP 2-207791), can be used.
[0223] The increase in protein activity can be confirmed by measuring the activity of the protein.
[0224] Increased protein activity can also be confirmed by confirming increased expression of the gene encoding the protein, which can be confirmed by confirming increased transcription of the gene or increased amount of protein expressed from the gene.
[0225] Increased gene transcription levels can be confirmed by comparing the amount of mRNA transcribed from the gene with that of a wild-type strain or a non-modified strain such as the parent strain. Methods for assessing mRNA levels include Northern hybridization, RT-PCR, microarrays, RNA-seq, etc. (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual / Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (USA), 2001). The amount of mRNA (e.g., number of molecules per cell) may be increased, for example, by 1.5-fold or more, 2-fold or more, or 3-fold or more compared to that of a non-modified strain.
[0226] The increase in the amount of the protein can be confirmed by Western blotting using an antibody (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual / Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (USA), 2001). The amount of the protein (e.g., the number of molecules per cell) may be increased by, for example, 1.5-fold or more, 2-fold or more, or 3-fold or more compared to that of an unmodified strain.
[0227] The above-mentioned methods for increasing protein activity can be used to enhance the activity of any protein or enhance the expression of any gene.
[0228] <1-6> Methods for reducing protein activity Methods for reducing protein activity are described below. Note that the methods for reducing protein activity described below can also be used to disrupt wild-type PhoS protein.
[0229] "Decreased protein activity" means that the activity of the protein is reduced compared to that of a non-modified strain. Specifically, "decreased protein activity" means that the activity of the protein per cell is reduced compared to that of a non-modified strain. Here, "non-modified strain" refers to a control strain that has not been modified to reduce the activity of the target protein. Examples of non-modified strains include wild-type strains and parent strains. Specific examples of non-modified strains include type strains of each bacterial species. Specific examples of non-modified strains also include the strains exemplified in the description of coryneform bacteria. That is, in one embodiment, the activity of the protein may be reduced compared to that of a type strain (i.e., the type strain of the species to which the bacterium of the present invention belongs). In another embodiment, the activity of the protein may be reduced compared to that of C. glutamicum ATCC 13032. In another embodiment, the activity of the protein may be reduced compared to that of C. glutamicum ATCC 13869. In another embodiment, the activity of the protein may be reduced compared to that of C. glutamicum AJ12036 (FERM BP-734). In another embodiment, the activity of the protein may be reduced compared to that of the C. glutamicum YDK010 strain. Note that "reduced protein activity" also encompasses cases where the activity of the protein is completely lost. More specifically, "reduced protein activity" may mean a reduction in the number of molecules of the protein per cell and / or a reduction in the function of the protein per molecule compared to a non-modified strain. That is, the "activity" in "reduced protein activity" is not limited to the catalytic activity of the protein, but may also refer to the transcription amount (mRNA amount) or translation amount (protein amount) of the gene encoding the protein. The "number of molecules of protein per cell" may refer to the average number of molecules of the protein per cell. Note that "reduced number of molecules of protein per cell" also encompasses cases where the protein is completely absent. Furthermore, "the function per molecule of a protein is reduced" also includes a case where the function per molecule of the protein is completely lost.The degree of reduction in protein activity is not particularly limited, as long as the protein activity is reduced compared to that of an unmodified strain. For example, the protein activity may be reduced to 50% or less, 20% or less, 10% or less, 5% or less, or 0% of that of an unmodified strain.
[0230] Modifications that reduce the activity of a protein can be achieved, for example, by reducing the expression of the gene encoding the protein. "Reduced gene expression" means that the expression of the gene is reduced compared to an unmodified strain. "Reduced gene expression" specifically means that the expression level of the gene per cell is reduced compared to an unmodified strain. "Expression level of the gene per cell" may refer to the average expression level of the gene per cell. "Reduced gene expression" may more specifically mean a reduction in the transcription level (mRNA level) of the gene and / or a reduction in the translation level (protein level) of the gene. "Reduced gene expression" also includes cases where the gene is not expressed at all. "Reduced gene expression" is also referred to as "attenuated gene expression." Gene expression may be reduced to, for example, 50% or less, 20% or less, 10% or less, 5% or less, or 0% of that of an unmodified strain.
[0231] Decreased gene expression may be due to, for example, decreased transcription efficiency, decreased translation efficiency, or a combination thereof. Decreased gene expression can be achieved, for example, by modifying the expression regulatory sequence of the gene. "Expression regulatory sequence" is a general term for sites that affect gene expression, such as promoters, Shine-Dalgarno (SD) sequences (also known as ribosome binding sites (RBS)), and spacer regions between the RBS and the start codon. Expression regulatory sequences can be determined, for example, using promoter search vectors or gene analysis software such as GENETYX. When modifying an expression regulatory sequence, preferably one or more bases, more preferably two or more bases, and particularly preferably three or more bases are modified in the expression regulatory sequence. Decreased gene transcription efficiency can be achieved, for example, by replacing the promoter of a gene on a chromosome with a weaker promoter. A "weaker promoter" refers to a promoter that weakens gene transcription compared to the native wild-type promoter. An example of a weaker promoter is an inducible promoter. That is, an inducible promoter can function as a weaker promoter under non-inducing conditions (e.g., in the absence of an inducer). Part or all of the expression regulatory sequence may also be deleted (deleted). Reduced gene expression can also be achieved, for example, by manipulating factors involved in expression control. Factors involved in expression control include small molecules (inducers, inhibitors, etc.), proteins (transcription factors, etc.), and nucleic acids (siRNA, etc.) involved in transcription and translation control. Reduced gene expression can also be achieved, for example, by introducing a mutation into the coding region of the gene that reduces gene expression. For example, gene expression can be reduced by replacing codons in the coding region of the gene with synonymous codons that are used less frequently in the host. Furthermore, gene expression itself can be reduced, for example, by gene disruption as described below.
[0232] Furthermore, a modification that reduces the activity of a protein can be achieved, for example, by disrupting the gene encoding the protein. "Disrupting a gene" means that the gene is modified so that it does not produce a protein that functions normally. "Not producing a protein that functions normally" includes cases where no protein is produced from the gene at all, and cases where the gene produces a protein with reduced or lost function per molecule (e.g., activity or properties).
[0233] Gene disruption can be achieved, for example, by deleting (deleting) the gene on a chromosome. "Gene deletion" refers to the deletion of part or all of the coding region of a gene. Furthermore, the entire gene may be deleted, including the sequences before and after the coding region of the gene on the chromosome. The sequences before and after the coding region of the gene may include, for example, a gene expression regulatory sequence. As long as a reduction in protein activity can be achieved, the region to be deleted may be any region, such as the N-terminal region (the region encoding the N-terminal side of the protein), an internal region, or a C-terminal region (the region encoding the C-terminal side of the protein). Generally, the longer the region to be deleted, the more reliably the gene can be inactivated. The region to be deleted may be, for example, 10% or more, 20% or more, 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, or 95% or more of the entire length of the coding region of the gene. Furthermore, it is preferable that the reading frames of the sequences before and after the region to be deleted do not match. Reading frame mismatches can result in frameshifts downstream of the region to be deleted.
[0234] Gene disruption can also be achieved by, for example, introducing an amino acid substitution (missense mutation) into the coding region of a gene on a chromosome, introducing a stop codon (nonsense mutation), or adding or deleting one or two bases (frameshift mutation) (Journal of Biological Chemistry 272:8611-8617(1997), Proceedings of the National Academy of Sciences, USA 95 5511-5515(1998), Journal of Biological Chemistry 26 116, 20833-20839(1991)).
[0235] Gene disruption can also be achieved, for example, by inserting another base sequence into the coding region of the gene on the chromosome. The insertion site may be anywhere in the gene, but the longer the inserted base sequence, the more reliably the gene can be inactivated. Furthermore, it is preferable that the reading frames of the sequences before and after the insertion site do not match. A mismatch in the reading frame can cause a frameshift downstream of the insertion site. The other base sequence is not particularly limited as long as it reduces or eliminates the activity of the encoded protein, and examples include marker genes such as antibiotic resistance genes and genes useful for producing target substances.
[0236] Gene disruption may be carried out, particularly to delete (delete) the amino acid sequence of the encoded protein. In other words, modification that reduces the activity of a protein can be achieved, for example, by deleting the amino acid sequence of the protein (partial or entire region of the amino acid sequence), specifically by modifying the gene to encode a protein from which the amino acid sequence (partial or entire region of the amino acid sequence) has been deleted. The term "deletion of the amino acid sequence of a protein" refers to the deletion of part or entire region of the amino acid sequence of a protein. The term "deletion of the amino acid sequence of a protein" refers to the absence of the original amino acid sequence in the protein, and also encompasses cases in which the original amino acid sequence is changed to a different amino acid sequence. For example, a region that has been changed to a different amino acid sequence due to frameshifting may be considered a deleted region. While deletion of the amino acid sequence of a protein typically shortens the overall length of the protein, it may also be possible for the overall length of the protein to remain unchanged or to be extended. For example, deletion of part or entire region of the coding region of a gene can delete the region encoded by the deleted region in the amino acid sequence of the encoded protein. For example, by introducing a stop codon into the coding region of a gene, the region coded for by the region downstream of the introduction site in the amino acid sequence of the encoded protein can be deleted. For example, by causing a frameshift in the coding region of a gene, the region coded for by the frameshift site can be deleted. The position and length of the region to be deleted in the deletion of an amino acid sequence can be determined mutatis mutandis from the explanation of the position and length of the region to be deleted in the deletion of a gene.
[0237] The above-described modification of a gene on a chromosome can be achieved, for example, by creating a disrupted gene modified so that it does not produce a normally functioning protein, transforming a host with recombinant DNA containing the disrupted gene, and inducing homologous recombination between the disrupted gene and the wild-type gene on the chromosome, thereby replacing the wild-type gene on the chromosome with the disrupted gene. In this case, incorporating a marker gene into the recombinant DNA according to the host's traits, such as its nutritional requirements, facilitates manipulation. Examples of disrupted genes include genes lacking part or all of the coding region of a gene, genes with missense mutations, genes with nonsense mutations, genes with frameshift mutations, and genes with insertion sequences such as transposons or marker genes. Even if a protein encoded by a disrupted gene is produced, it will have a different three-dimensional structure from the wild-type protein, resulting in reduced or lost function. The structure of the recombinant DNA used for homologous recombination is not particularly limited, as long as it allows homologous recombination to occur in the desired manner. For example, a host can be transformed with linear DNA containing a disrupted gene, the linear DNA having upstream and downstream sequences of a wild-type gene on a chromosome at both ends, and homologous recombination can occur upstream and downstream of the wild-type gene, thereby replacing the wild-type gene with the disrupted gene.Such gene disruption by gene replacement using homologous recombination has already been established, and includes methods using linear DNA, such as a method called "Red-driven integration" (Datsenko, K. A., and Wanner, BL Proc. Natl. Acad. Sci. USA 97:6640-6645 (2000)), a method combining the Red-driven integration method with an excision system derived from λ phage (Cho, EH, Gumport, RI, Gardner, JFJ Bacteriol. 184: 5200-5203 (2002)) (see WO2005 / 010175), methods using a plasmid containing a temperature-sensitive replication origin, methods using a conjugatively transferable plasmid, and methods using a suicide vector that does not have a replication origin that functions in the host (U.S. Pat. No. 6,303,383, JP 05-007491 A). Such a method of modifying a chromosome using homologous recombination is not limited to disrupting a target gene, but can be used for any modification of a chromosome, such as modification of an expression regulatory sequence.
[0238] Modifications that reduce the activity of a protein may also be performed by, for example, mutation treatments, such as X-ray irradiation, ultraviolet irradiation, and treatment with mutagens such as N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), ethyl methanesulfonate (EMS), and methyl methanesulfonate (MMS).
[0239] The above-mentioned methods for reducing protein activity may be used alone or in any combination.
[0240] The decrease in the activity of the protein can be confirmed by measuring the activity of the protein.
[0241] A decrease in protein activity can also be confirmed by confirming a decrease in expression of the gene encoding the protein. A decrease in gene expression can be confirmed by confirming a decrease in the transcription level of the gene or a decrease in the amount of protein expressed from the gene.
[0242] The reduction in the transcription level of a gene can be confirmed by comparing the amount of mRNA transcribed from the gene with that of a non-modified strain. Methods for assessing the amount of mRNA include Northern hybridization, RT-PCR, microarray, RNA-Seq, etc. (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual / Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (USA), 2001). The amount of mRNA may be reduced to, for example, 50% or less, 20% or less, 10% or less, 5% or less, or 0% of that of a non-modified strain.
[0243] The reduction in protein amount can be confirmed by performing SDS-PAGE and checking the intensity of the separated protein bands. The reduction in protein amount can also be confirmed by Western blotting using an antibody (Sambrook, J., et al., Molecular Cloning: A Laboratory Manual / Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (USA), 2001). The protein amount (e.g., number of molecules per cell) may be reduced to, for example, 50% or less, 20% or less, 10% or less, 5% or less, or 0% of that of an unmodified strain.
[0244] Gene disruption can be confirmed by determining the nucleotide sequence, restriction enzyme map, or full length of a part or all of the gene, depending on the means used for disruption.
[0245] The above-mentioned methods for reducing protein activity can be used to reduce the activity of any protein or the expression of any gene.
[0246] <2> Method for producing ncAA-containing protein The bacterium of the present invention obtained as described above is cultured in a medium containing ncAA to express the ncAA-containing protein, thereby obtaining a large amount of the ncAA-containing protein secreted outside the bacterial cells.
[0247] The medium used is not particularly limited, as long as it contains ncAA, allows the bacterium of the present invention to grow, and produces an ncAA-containing protein. For example, a medium prepared by adding ncAA to a conventional medium used for culturing bacteria such as coryneform bacteria can be used. Specific examples of the medium include conventional media containing, in addition to ncAA, a carbon source, a nitrogen source, and inorganic ions. Organic micronutrients such as vitamins and amino acids can also be added to the medium as needed. The culture conditions are not particularly limited, as long as the bacterium of the present invention can grow and produce an ncAA-containing protein. Culture can be carried out under, for example, conventional conditions used for culturing bacteria such as coryneform bacteria. The types and concentrations of medium components and the culture conditions can be appropriately determined depending on various conditions, such as the type of coryneform bacteria and the type of ncAA-containing protein.
[0248] Carbon sources include carbohydrates such as glucose and sucrose, organic acids such as acetic acid, alcohols, or other suitable carbon sources. Nitrogen sources include ammonia gas, aqueous ammonia, ammonium salts, or other suitable nitrogen sources. Inorganic ions include calcium ions, magnesium ions, phosphate ions, potassium ions, and iron ions, as needed. Culture can be performed aerobicly for approximately 1 to 7 days at a suitable pH range of 5.0 to 8.5 and 15 to 37°C. Culture conditions for L-amino acid production by coryneform bacteria and other protein production methods using Sec-dependent or Tat-dependent signal peptides can also be used (see WO01 / 23591 and WO2005 / 103278). When an inducible promoter is used to express an ncAA-containing protein and / or an orthogonal tRNA(ncAA) / ncAA-aaRS pair, their expression may or may not be induced. By culturing the bacterium of the present invention under these conditions, the ncAA-containing protein is produced in large amounts within the bacterium and efficiently secreted outside the bacterium. Furthermore, since the produced ncAA-containing protein is secreted outside the bacterium, proteins such as transglutaminase, which are generally lethal when accumulated in large amounts within microbial cells, can be continuously produced without any lethal effects.
[0249] Cultivation can be carried out by batch culture, fed-batch culture, continuous culture, or a combination of these. The medium at the start of cultivation is also called the "initial medium." The medium supplied to the cultivation system (fermenter) in fed-batch culture or continuous culture is also called the "fed-batch medium." Supplying a fed-batch medium to the cultivation system in fed-batch culture or continuous culture is also called "fed-batch." Cultivation may be carried out in two stages: seed culture and main culture.
[0250] The ncAA concentration in the medium is not particularly limited, as long as an ncAA-containing protein is produced. For example, the ncAA may be contained in the medium at a concentration of 0.1 mM or more, 0.2 mM or more, 0.3 mM or more, 0.4 mM or more, 0.5 mM or more, 0.7 mM or more, or 1 mM or more, or at a concentration of 20 mM or less, 10 mM or less, 5 mM or less, 3 mM or less, 2 mM or less, 1 mM or less, 0.7 mM or less, or 0.5 mM or less, or any compatible combination thereof. The ncAA may be contained in the initial medium within the concentration range exemplified above, and / or may be added during culture to achieve the concentration range exemplified above.
[0251] The ncAA may or may not be present in the medium throughout the entire culture period. For example, the ncAA may or may not be present in the medium at a predetermined concentration range, such as the concentration ranges exemplified above, throughout the entire culture period. That is, the ncAA may be present in the medium at a concentration outside the predetermined concentration range, such as the concentration ranges exemplified above, during a portion of the culture period. The ncAA may be in short supply during a portion of the culture period. "Insufficient" means that the required amount is not met, and may mean, for example, that the concentration in the medium is zero. For example, the ncAA may or may not be present in the medium from the start of culture. If the ncAA is not present in the medium at the start of culture, the ncAA is supplied to the medium after the start of culture. The timing of supply can be appropriately determined depending on various conditions, such as the culture time. The ncAA may be supplied to the medium after the bacterium of the present invention has grown sufficiently. Furthermore, for example, the ncAA may be consumed during culture, and its concentration in the medium may become zero. A "partial period" may refer to, for example, 1% or less, 5% or less, 10% or less, 20% or less, 30% or less, or 50% or less of the total culture period. When the culture is divided into a seed culture and a main culture, "the entire culture period" may refer to the entire main culture period. Thus, even if ncAA is insufficient for a partial period, as long as there is a culture period in a medium containing ncAA, this is included in "culturing bacteria in a medium containing ncAA." Typically, ncAA may be contained in the medium at least throughout the desired period for expression of the ncAA-containing protein.
[0252] The ncAA-containing protein secreted into the medium by the method of the present invention can be isolated and purified from the culture medium after cultivation using methods well known to those skilled in the art. For example, after removing the bacterial cells by centrifugation or other methods, the ncAA-containing protein can be isolated and purified using known appropriate methods such as salting out, ethanol precipitation, ultrafiltration, gel filtration chromatography, ion exchange column chromatography, affinity chromatography, medium- to high-pressure liquid chromatography, reversed-phase chromatography, and hydrophobic chromatography, or a combination of these. In some cases, fractions containing the ncAA-containing protein, such as the culture or culture supernatant, can be used directly as the ncAA-containing protein. The ncAA-containing protein secreted onto the bacterial cell surface by the method of the present invention can also be solubilized by methods well known to those skilled in the art, such as increasing the salt concentration or using a surfactant, and then isolated and purified in the same manner as if it were secreted into the medium. In some cases, the ncAA-containing protein secreted onto the bacterial cell surface can be used, for example, as an immobilized enzyme, without solubilization.
[0253] The secretion and production of ncAA-containing proteins can be confirmed by performing SDS-PAGE using the culture supernatant and / or fractions containing the bacterial cell surface as samples and determining the molecular weight of the separated protein bands. The secretion and production of ncAA-containing proteins can also be confirmed by Western blotting using antibodies with the culture supernatant and / or fractions containing the bacterial cell surface as samples (Molecular Cloning (Cold Spring Harbor Laboratory Press, Cold Spring Harbor (USA), 2001)). The secretion and production of ncAA-containing proteins can also be confirmed by detecting the N-terminal amino acid sequence of the secreted protein using a protein sequencer. The secretion and production of ncAA-containing proteins can also be confirmed by determining the mass of the secreted protein using a mass spectrometer. If the ncAA-containing protein is an enzyme or has some measurable physiological activity, the secretion and production of the ncAA-containing protein can be confirmed by measuring the enzymatic or physiological activity of the ncAA-containing protein using the culture supernatant and / or fractions containing the bacterial cell surface as samples.
[0254] The produced ncAA-containing protein may be modified before use. That is, the method of the present invention may further include a step of modifying the produced ncAA-containing protein. The type of modification can be selected appropriately depending on various conditions, such as the intended use of the ncAA-containing protein. Modifications include PEGylation, fluorescent labeling, the addition of fatty acid molecules, and the addition of functional molecules (e.g., antibodies, antibody fragments, peptides, proteins, nucleic acids, organic compounds, inorganic compounds, sugar chains, lipids, polymers, metals (e.g., gold), or chelators). Modifications may be site-specific or not. For example, the azide group of an azide-containing ncAA residue such as AzF can be site-specifically modified by strain-promoted alkyne azide cycloaddition (SPAAC) reaction.
[0255] The present invention will now be described in more detail with reference to the following non-limiting examples.
[0256] In this example, "azidophenylalanine" (AzF) refers to p-Azido-L-phenylalanine, "3-chlorotyrosine" (ClY) refers to 3-chloro-L-tyrosine, "3-azidotyrosine" (AzY) refers to 3-Azido-L-tyrosine, "pyrrolidine" (Pyl) refers to L-Pyrrolysine, and "alloclysine" (AllocLys) refers to N- δ -Alloc-L-lysine, respectively.
[0257] Example 1: Expression of a mutant anti-epidermal growth factor receptor (EGFR) VHH antibody 9g8 with site-specific substitution to azidophenylalanine (AzF) 1. Preparation of a plasmid vector (AzFN3 vector) containing an AzFN3 DNA cassette The AzFN3 DNA cassette (SEQ ID NO: 55) was obtained by total synthesis. The AzFN3 DNA cassette contains a T7 promoter and an azidophenylalanyl-tRNA synthetase (AzFRS) gene linked downstream thereof, and an F1 promoter and a suppressor tRNA (tRNA) linked downstream thereof that can translate a UAG codon into azidophenylalanine (AzF). CTA AzFRS contains the tRNA gene. AzFRS is a modified tyrosyl-tRNA synthetase derived from the archaeon Methanococcus jannaschii, which has the mutations Y32T / E107N / D158P / I159L / L162Q and has been modified to use azidophenylalanine as a substrate (J. AM. CHEM. SOC. 2002, 124, 9026-9027). The nucleotide sequence of the AzFRS gene and the amino acid sequence of AzFRS are shown in SEQ ID NOs: 49 and 50, respectively. tRNA CTA is a modified tyrosyl-tRNA from the archaeon Methanococcus jannaschii that has been engineered to have an anticodon (CTA) corresponding to UAG (amber). CTAThe base sequence of the gene and the base sequence of tRNA(Tyr) encoded by the gene are shown in SEQ ID NOs: 41 and 42, respectively. When this AzFN3 DNA cassette is introduced into coryneform bacteria and cultured in a medium containing azidophenylalanine, azidophenylalanine is taken up by the bacteria and converted to tRNA(Tyr) by AzFRS. CTA The azidophenylalanyl-tRNA binds to the end of the codon, UAG, in the mRNA during protein translation, and the codon is translated as azidophenylalanine, resulting in the synthesis of a protein with azidophenylalanine introduced at the UAG codon position.
[0258] The AzFN3 DNA cassette was cloned into the pVC7T7pol1 vector (Shuhei Hashiro et al., Efficient production of long double-stranded RNAs applicable to agricultural pest control by Corynebacterium glutamicum equipped with a coliphage T7-expression system, Appl Microbiol Biotechnol, 2021 Jun 7.) to obtain the AzFN3 vector. The pVC7T7pol1 vector is a pVC7N vector (chloramphenicol drug resistance) incorporating the lacI gene and T7 RNA polymerase gene.
[0259] 2. Construction of a Plasmid Vector (9g8+AzFRS Vector) Containing a Wild-Type or Mutant 9g8 DNA Cassette and an AzFRS DNA Cassette. The wild-type 9g8 DNA cassette (SEQ ID NO: 56) was obtained by total synthesis. This wild-type 9g8 DNA cassette contains a cspB promoter and a gene encoding wild-type 9g8 linked downstream thereof with a CspA signal sequence at the N-terminus and a His-tag sequence at the C-terminus. 9g8 is a VHH antibody against the epidermal growth factor receptor (EGFR). The nucleotide sequence of the gene encoding wild-type 9g8 (excluding the additional sequence) and the amino acid sequence of wild-type 9g8 (excluding the additional sequence) are shown in SEQ ID NOs: 57 and 58, respectively. The wild-type 9g8 DNA cassette was cloned into a pPK4 vector (Japanese Patent Laid-Open No. 9-322774; kanamycin drug resistance) to obtain a pPK4 vector (wild-type 9g8 vector) containing the wild-type 9g8 DNA cassette.
[0260] Next, the wild-type 9g8 vector was modified to express 11 types of mutant 9g8 in which 11 amino acid residues at positions 32, 59, 60, 80, 94, 95, 100, 107, 112, 114, and 116 of wild-type 9g8 were substituted with azidophenylalanine. Specifically, using the wild-type 9g8 vector as a template, PCR was performed with a primer pair for introducing mutations to change the triplets encoding the 11 amino acid residues in the wild-type 9g8 gene to TAG triplets corresponding to the UAG codon, a stop codon, to obtain pPK4 vectors (mutant 9g8 vectors) containing 11 types of mutant 9g8 DNA cassettes. The primer pairs used were the pair of SEQ ID NOs: 59 and 60, the pair of SEQ ID NOs: 61 and 62, the pair of SEQ ID NOs: 63 and 64, the pair of SEQ ID NOs: 65 and 66, the pair of SEQ ID NOs: 67 and 68, the pair of SEQ ID NOs: 69 and 70, the pair of SEQ ID NOs: 71 and 72, the pair of SEQ ID NOs: 73 and 74, the pair of SEQ ID NOs: 75 and 76, the pair of SEQ ID NOs: 77 and 78, and the pair of SEQ ID NOs: 79 and 80, for positions 32, 59, 60, 80, 94, 95, 100, 107, 112, 114, and 116 of wild-type 9g8, respectively. The modified parts are shown in Figure 2.
[0261] The synthesized AzFRS DNA cassette (SEQ ID NO: 81) was cloned downstream of the wild-type and mutant 9g8 genes in the wild-type and mutant 9g8 vectors to obtain wild-type and mutant 9g8 + AzFRS vectors. The AzFRS DNA cassette contains a T7 promoter and the AzFRS gene linked downstream of it.
[0262] 3. Expression of Mutant 9g8 The C. glutamicum YDK0107 strain was transformed with the AzFN3 vector and each mutant 9g8 + AzFRS vector to obtain 11 mutant 9g8-expressing strains containing one of the 11 mutant 9g8 genes and the AzFN3 DNA cassette. As a control, the C. glutamicum YDK0107 strain was transformed with the wild-type 9g8 vector or with the AzFN3 vector and the wild-type 9g8 + AzFRS vector to obtain wild-type 9g8-expressing strains (also referred to as control strain PC and control strain WT, respectively). Strain YDK0107 is a spontaneous mutant of strain YDK010 (WO2002 / 081694), a mutant lacking the cell surface protein CspB of C. glutamicum AJ12036 (FERM BP-734), in which the PhoS(W302C) mutation was introduced into the chromosomal phoS gene (WO2016 / 171224). Each 9g8-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, each strain's culture medium was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, a band was observed at a position corresponding to the full-length of 9g8 in the presence of azidophenylalanine in 11 mutant 9g8-expressing strains, confirming the expression of mutant 9g8 into which azidophenylalanine had been introduced (Figure 3).
[0263] Example 2: Expression of mutant anti-epidermal growth factor receptor (Human Epidermal Growth Factor Receptor 2: HER2) antibody ZHER2 affibody with site-specific azidophenylalanine (AzF) substitution 1. Preparation of a plasmid vector (AzFN3 vector) containing an AzFN3 DNA cassette The AzFN3 vector was prepared as described in Example 1-1.
[0264] 2. Construction of a Plasmid Vector Containing a Wild-Type or Mutant ZHER2 Affibody DNA Cassette and an AzFRS DNA Cassette (ZHER2 Affibody + AzFRS Vector) The wild-type ZHER2 affibody DNA cassette (SEQ ID NO: 82) was obtained by total synthesis. This wild-type ZHER2 affibody DNA cassette contains a cspB promoter and, linked downstream thereof, a gene encoding the wild-type ZHER2 affibody, which is sequentially enriched at its N-terminus with a CspB signal sequence, the N-terminal six amino acid residues of the mature CspB protein, a TEV protease recognition sequence, and a His tag sequence. The ZHER2 affibody is an affibody for the epidermal growth factor receptor 2 (HER2). The nucleotide sequence of the gene encoding the wild-type ZHER2 affibody (excluding the additional sequence) and the amino acid sequence of the wild-type ZHER2 affibody (excluding the additional sequence) are shown in SEQ ID NOs: 83 and 84, respectively. The wild-type ZHER2 affibody DNA cassette was cloned into the pPK5 vector (WO2018 / 074579; kanamycin drug resistance) to obtain a pPK5 vector containing the wild-type ZHER2 affibody DNA cassette (wild-type ZHER2 affibody vector).
[0265] Next, the wild-type ZHER2 affibody vector was modified to express four mutant ZHER2 affibodies in which the amino acid residues at positions 7, 16, 22, and 37 of the wild-type ZHER2 affibody were substituted with azidophenylalanine. Specifically, using the wild-type ZHER2 affibody vector as a template, PCR was performed with a primer pair for introducing mutations to change the triplets encoding the four amino acid residues in the wild-type ZHER2 affibody gene to TAG triplets corresponding to the UAG codon, a stop codon, to obtain pPK4 vectors (mutant ZHER2 affibody vectors) containing four mutant ZHER2 affibody DNA cassettes. The primer pairs used were the pair of SEQ ID NOs: 85 and 86, the pair of SEQ ID NOs: 87 and 88, the pair of SEQ ID NOs: 89 and 90, and the pair of SEQ ID NOs: 91 and 92 for positions 7, 16, 22, and 37 of the wild-type ZHER2 affibody, respectively. The modifications are shown in Figure 4.
[0266] The fully synthesized AzFRS DNA cassette (Example 1-2) was cloned into the downstream region of the wild-type and mutant-type ZHER2 affibody genes in the wild-type and mutant-type ZHER2 affibody vectors to obtain wild-type and mutant-type ZHER2 affibody+AzFRS vectors.
[0267] 3. Expression of Mutant ZHER2 Affibodies. C. glutamicum YDK0107 was transformed with the AzFN3 vector and each mutant ZHER2 affibody + AzFRS vector to obtain four mutant ZHER2 affibody-expressing strains containing one of the four mutant ZHER2 affibody genes and the AzFN3 DNA cassette. As a control, C. glutamicum YDK0107 was transformed with the wild-type ZHER2 affibody vector or with the AzFN3 vector and the wild-type ZHER2 affibody + AzFRS vector to obtain wild-type ZHER2 affibody-expressing strains (also referred to as control strain PC and control strain WT, respectively). Each ZHER2 affibody-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM azidophenylalanine. After incubation, 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). In the presence of azidophenylalanine, bands corresponding to the full-length ZHER2 affibody were observed in the four mutant ZHER2 affibody-expressing strains, confirming the expression of the azidophenylalanine-introduced mutant ZHER2 affibody (Fig. 5).
[0268] Example 3: Expression of a site-specifically substituted azidophenylalanine (AzF) mutant of a monomeric red fluorescent protein (mRFP) 1. Preparation of a plasmid vector (AzFN3 vector) containing an AzFN3 DNA cassette The AzFN3 vector was prepared as described in Example 1, section 1.
[0269] 2. A plasmid vector containing a wild-type or mutant mRFP DNA cassette and an AzFRS DNA cassette (mRFP+AzFRS vector), and a wild-type or mutant mRFP DNA cassette and tRNA CTA Plasmid vector containing DNA cassette (mRFP + tRNA CTAA wild-type mRFP DNA cassette (SEQ ID NO: 93) was obtained by total synthesis. This wild-type mRFP DNA cassette contains a cspB promoter and a gene encoding wild-type mRFP linked downstream thereof, to which the CspB signal sequence and the N-terminal 6 amino acid residues of the mature CspB protein have been added, in that order, to the N-terminus. The nucleotide sequence of the gene encoding wild-type mRFP (excluding the additional sequence) and the amino acid sequence of wild-type mRFP (excluding the additional sequence) are shown in SEQ ID NOs: 94 and 95, respectively. The wild-type mRFP DNA cassette was cloned into the pPK4 vector (JP 9-322774; kanamycin drug resistance) to obtain a pPK4 vector (wild-type mRFP vector) containing the wild-type mRFP DNA cassette.
[0270] Next, the wild-type mRFP vector was modified to express mutant mRFP, in which a single amino acid residue at position 36 of wild-type mRFP was substituted with azidophenylalanine. Specifically, using the wild-type mRFP vector as a template, PCR was performed using a pair of mutation-introducing primers (pair of SEQ ID NOs: 96 and 97) to change the triplet encoding the single amino acid residue in the wild-type mRFP gene to a TAG triplet corresponding to the UAG codon, a stop codon, to obtain a pPK4 vector (mutant mRFP vector) containing a mutant mRFP DNA cassette. The modifications (including the modification at position 80 described in Example 5 below) are shown in Figure 6.
[0271] The fully synthesized AzFRS DNA cassette (Example 1-2) was cloned into the downstream region of the wild-type and mutant mRFP genes in the wild-type and mutant mRFP vectors to obtain wild-type and mutant mRFP + AzFRS vectors. CTA The DNA cassette (SEQ ID NO: 100) was cloned into the downstream region of the wild-type and mutant mRFP genes in the wild-type and mutant mRFP vectors, and wild-type and mutant mRFP + tRNA CTA Vector was obtained. tRNA CTA The DNA cassette contains the F1 promoter and a tRNA linked downstream of it. CTA Contains genes.
[0272] 3. Expression of mutant mRFP AzFN3 vector and mutant mRFP + AzFRS vector or mutant mRFP + tRNA CTA The C. glutamicum YDK0107 strain was transformed with the vector to obtain a mutant mRFP-expressing strain containing both the mutant mRFP gene and the AzFN3 DNA cassette. As controls, the AzFN3 vector and wild-type mRFP + AzFRS vector or wild-type mRFP + tRNA were used. CTA C. glutamicum YDK0107 was transformed with the vector to obtain wild-type mRFP-expressing strains (also referred to as the control strain WT+AzFRS or control strain WT+tRNA, respectively). Each mRFP-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, each strain's culture was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). In the presence of azidophenylalanine, a band corresponding to the full-length mRFP was observed in the mutant mRFP-expressing strain, confirming the expression of the azidophenylalanine-introduced mutant mRFP (Figure 7).
[0273] Example 4: Expression of a site-specifically substituted azidophenylalanine (AzF) mutant of monomeric red fluorescent protein (mRFP) 2> 1. Preparation of a plasmid vector containing an AzFN3 DNA cassette (AzFN3 vector) The AzFN3 vector was prepared as described in Example 1, section 1.
[0274] 2. Preparation of Plasmid Vectors Containing Wild-Type or Mutant-Type mRFP Genes Wild-type and mutant-type mRFP vectors were prepared as described in Example 3, section 2.
[0275] 3. Expression of Mutant mRFP. C. glutamicum YDK0107 was transformed with the AzFN3 vector and mutant mRFP vector to obtain a mutant mRFP-expressing strain containing both the mutant mRFP gene and the AzFN3 DNA cassette. As a control, C. glutamicum YDK0107 was transformed with the AzFN3 vector and wild-type mRFP vector to obtain a wild-type mRFP-expressing strain (also referred to as the control strain WT). Each mRFP-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, each culture was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, in the mutant mRFP-expressing strain, a band was observed at a position corresponding to the full-length mRFP in the presence of azidophenylalanine, confirming the expression of mutant mRFP into which azidophenylalanine had been introduced (FIG. 8).
[0276] Example 5: Expression of mutant monomeric red fluorescent protein (mRFP) with site-specific substitution with azidophenylalanine (AzF) 3 1. Preparation of plasmid vector containing wild-type or mutant mRFP DNA cassette and AzFN3 DNA cassette (mRFP+AzFN3 vector) A wild-type mRFP vector was prepared as described in Example 3, section 2.
[0277] Next, the wild-type mRFP vector was modified to express two mutant mRFPs in which the amino acid residues at positions 36 and 80 of wild-type mRFP were substituted with azidophenylalanine. Specifically, using the wild-type mRFP vector as a template, PCR was performed using a primer pair for mutation introduction to change the triplets encoding these two amino acid residues in the wild-type mRFP gene to TAG triplets, thereby obtaining pPK4 vectors (mutant mRFP vectors) containing two mutant mRFP DNA cassettes. The primer pairs used for positions 36 and 80 of wild-type mRFP were the pair of SEQ ID NOs: 96 and 97, and the pair of SEQ ID NOs: 98 and 99, respectively.
[0278] The AzFN3 DNA cassette (1 in Example 1) was cloned into the downstream region of the wild-type and mutant-type mRFP genes in the wild-type and mutant-type mRFP vectors to obtain wild-type and mutant-type mRFP+AzFN3 vectors.
[0279] 2. Expression of Mutant mRFP. C. glutamicum YDK0107 was transformed with the pVC7T7pol1 vector and the mutant mRFP + AzFN3 vector to obtain two mutant mRFP-expressing strains, each containing one of the two mutant mRFP genes and the AzFN3 DNA cassette. As a control, C. glutamicum YDK0107 was transformed with the pVC7T7pol1 vector and the wild-type mRFP + AzFN3 vector to obtain a wild-type mRFP-expressing strain (also referred to as the control strain WT). Each mRFP-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, the culture medium was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, in the mutant mRFP-expressing strain, a band was observed at a position corresponding to the full-length mRFP in the presence of azidophenylalanine, confirming the expression of mutant mRFP into which azidophenylalanine had been introduced (FIG. 9).
[0280] Example 6: Expression of mutants of anti-Izumo protein 1 N-terminal extracellular domain (NDOM) VHH antibody N15 site-specifically substituted with 3-chlorotyrosine (ClY) 1. Preparation of plasmid vectors (N15+IYN3 vectors) containing wild-type or mutant N15 DNA cassettes and IYN3 DNA cassettes The IYN3 DNA cassette (SEQ ID NO: 101) was obtained by total synthesis. The IYN3 DNA cassette contains a T7 promoter and a halogenated tyrosyl-tRNA synthetase (IYRS) gene linked downstream thereof, and a cspB promoter and a suppressor tRNA (tRNA) capable of translating a UAG codon into halogenated tyrosine linked downstream thereof. CTA IYRS is a modified tyrosyl-tRNA synthetase derived from the archaeon Methanococcus jannaschii, which has the mutations H70A, D158T, I159S, and D286R and has been modified to use halogenated tyrosine as a substrate (Biochem. Biophys. Res. Commun. 411, 757-761 (2011)). The nucleotide sequence of the IYRS gene and the amino acid sequence of IYRS are shown in SEQ ID NOs: 51 and 52, respectively. tRNA CTA is a modified tyrosyl-tRNA from the archaeon Methanococcus jannaschii that has been engineered to have an anticodon (CTA) corresponding to UAG (amber). CTA The nucleotide sequence of the gene and the nucleotide sequence of tRNA(Tyr) encoded by the gene are shown in SEQ ID NOs: 43 and 44, respectively. When this IYN3 DNA cassette is introduced into a coryneform bacterium and cultured in a medium containing halogenated tyrosine, the halogenated tyrosine is taken up by the bacterium and converted to tRNA(Tyr) by IYRS. CTA The halogenated tyrosyl-tRNA binds to the end of the codon, UAG, in mRNA during protein translation, and the codon is translated as a halogenated tyrosine, resulting in the synthesis of a protein with a halogenated tyrosine at the UAG codon position.
[0281] The wild-type N15 DNA cassette (SEQ ID NO: 102) was obtained by total synthesis. The wild-type N15 DNA cassette contains a cspB promoter and a gene encoding wild-type N15 linked downstream thereof, to which the CspB signal sequence and the N-terminal six amino acid residues of the mature CspB protein are added, in order, at the N-terminus. N15 is an anti-NDOM VHH antibody. The nucleotide sequence of the gene encoding wild-type N15 (excluding the additional sequence) and the amino acid sequence of wild-type N15 (excluding the additional sequence) are shown in SEQ ID NOs: 103 and 104, respectively. The wild-type N15 DNA cassette was cloned into the pPK4 vector (JP 9-322774; kanamycin drug resistance) to obtain the pPK4 vector (wild-type N15 vector) containing the wild-type N15 DNA cassette.
[0282] Next, the wild-type N15 vector was modified to express three mutant N15s in which the amino acid residues at positions 60, 81, and 96 of wild-type N15 were each substituted with 3-chlorotyrosine. Specifically, using the wild-type N15 vector as a template, PCR was performed with a pair of mutation-introducing primers to change the triplets encoding these three amino acid residues in the wild-type N15 gene to TAG triplets corresponding to the UAG codon, a stop codon, to obtain pPK4 vectors (mutant N15 vectors) containing three mutant N15 DNA cassettes. The primer pairs used for positions 60, 81, and 96 of wild-type N15 were the pair of SEQ ID NOs: 105 and 106, the pair of SEQ ID NOs: 107 and 108, and the pair of SEQ ID NOs: 109 and 110, respectively. The modifications are shown in Figure 10.
[0283] The IYN3 DNA cassette was cloned into the downstream region of the wild-type and mutant N15 genes in the wild-type and mutant N15 vectors to obtain the wild-type and mutant N15 + IYN3 vectors.
[0284] 2. Expression of Mutant N15. C. glutamicum YDK0107 was transformed with the pVC7T7pol1 vector and the mutant N15 + IYN3 vector to obtain three mutant N15-expressing strains containing one of the three mutant N15 genes and an IYN3 DNA cassette. As a control, C. glutamicum YDK0107 was transformed with the wild-type N15 vector or the pVC7T7pol1 vector and the wild-type N15 + IYN3 vector to obtain wild-type N15-expressing strains (also referred to as control strains PC and WT, respectively). Each N15-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 1 mM 3-chlorotyrosine. After the culture was completed, the culture medium was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, a band was observed at a position corresponding to the full-length N15 in the presence of 3-chlorotyrosine in the mutant N15-expressing strain, confirming the expression of mutant N15 into which 3-chlorotyrosine had been introduced (Figure 11).
[0285] Example 7: Expression of mutant anti-Izumo protein 1 N-terminal extracellular domain (NDOM) VHH antibodies site-specifically substituted with 3-azidotyrosine (AzY) Halogenated tyrosyl-tRNA synthetase (IYRS) is known to also use 3-azidotyrosine (AzY) as a substrate (Nucleic Acids Research 2010, Vol. 38, 3682-3691). Therefore, when the IYN3 DNA cassette (Example 6-1) is introduced into coryneform bacteria and cultured in a medium containing 3-azidotyrosine, 3-azidotyrosine is taken up by the bacteria and synthesized into tRNA by IYRS. CTA During protein translation, 3-azidotyrosyl tRNA binds to the UAG codon in mRNA, which is translated as 3-azidotyrosine, resulting in the synthesis of a protein with 3-azidotyrosine introduced at the UAG codon position.
[0286] Each N15-expressing strain obtained in Example 6-2 was cultured at 30°C for 72 hours in medium supplemented with 0.3 mM 3-azidotyrosine. After the culture was completed, the culture medium for each strain was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, a band was observed at the position corresponding to the full-length N15 in the presence of 3-azidotyrosine for the mutant N15-expressing strain, confirming the expression of mutant N15 into which 3-azidotyrosine had been introduced (Figure 12).
[0287] Example 8: Expression of mutant anti-epidermal growth factor receptor (EGFR) VHH antibody 9g8 with site-specific substitution of alloclysine (AllocLys) 1. Construction of a plasmid vector (PylRS+9g8+tRNA(Pyl)CTA / tRNA(Pyl)TCA vector) containing a wild-type or mutant 9g8 DNA cassette, a PylRS DNA cassette, and a tRNA(Pyl)CTA or tRNA(Pyl)TCA DNA cassette. The wild-type 9g8 DNA cassette (SEQ ID NO: 56) was obtained by total synthesis. This wild-type 9g8 DNA cassette contains a gene encoding wild-type 9g8, linked downstream of the cspB promoter and with a CspA signal sequence at the N-terminus and a His-tag sequence at the C-terminus. 9g8 is a VHH antibody against the epidermal growth factor receptor (EGFR). The nucleotide sequence of the gene encoding wild-type 9g8 (excluding the additional sequence) and the amino acid sequence of wild-type 9g8 (excluding the additional sequence) are shown in SEQ ID NOs: 57 and 58, respectively. The wild-type 9g8 DNA cassette was cloned into the pPK4 vector (JP 9-322774 A; kanamycin drug resistance) to obtain the pPK4 vector (wild-type 9g8 vector) containing the wild-type 9g8 DNA cassette.
[0288] Next, the wild-type 9g8 vector was modified to express two types of mutant 9g8, in which the amino acid residues at positions 32 and 107 of wild-type 9g8 were substituted with alloclysines, respectively. Specifically, using the wild-type 9g8 vector as a template, PCR was performed using a pair of mutation-introducing primers to either change the triplet encoding the amino acid residue at position 32 in the wild-type 9g8 gene to a TAG triplet corresponding to the UAG codon (amber), a stop codon (this mutant 9g8 is also referred to as 9g8 (Y32TAG)), or change the triplet encoding the amino acid residue at position 107 in the wild-type 9g8 gene to a TGA triplet corresponding to the UGA codon (opal), a stop codon (this mutant 9g8 is also referred to as 9g8 (Y107TGA)). These mutant 9g8 vectors contained two types of mutant 9g8 DNA cassettes. The primer pairs used were the pair of SEQ ID NOs: 59 and 60 and the pair of SEQ ID NOs: 111 and 112 for positions 32 and 107 of wild-type 9g8, respectively. The modified sites are shown in Figure 2.
[0289] The PylRS DNA cassette (SEQ ID NO: 113) was synthesized. The PylRS DNA cassette contains the lac and T7 promoters and the pyrrolysyl-tRNA synthetase (PylRS) gene linked downstream. PylRS is a pyrrolysyl-tRNA synthetase derived from the archaebacterium Methanosarcina mazei (Chemistry & Biology 2008 15, 1187-1197). The nucleotide sequence of the PylRS gene and the amino acid sequence of PylRS are shown in SEQ ID NOs: 114 and 115, respectively.
[0290] The tRNA(Pyl)CTA DNA cassette (SEQ ID NO: 116) and the tRNA(Pyl)TCA DNA cassette (SEQ ID NO: 117) were obtained by total synthesis. Both the tRNA(Pyl)CTA DNA cassette and the tRNA(Pyl)TCA DNA cassette contain an F1+44 promoter and a suppressor tRNA gene capable of translating a stop codon downstream of the promoter into a pyrrolysine (Pyl) analogue, and a rrnC terminator downstream of the promoter. The tRNA(Pyl)CTA is a modified pyrrolysyl-tRNA derived from the archaebacterium Methanosarcina mazei that has been modified to have an anticodon (CTA) corresponding to UAG (amber). The nucleotide sequences of the tRNA(Pyl)CTA gene and the tRNA(Pyl)CTA encoded by the gene are shown in SEQ ID NOs: 118 and 119, respectively. tRNA(Pyl)TCA is a modified pyrrolysyl-tRNA derived from the archaea Methanosarcina mazei that has been modified to have an anticodon (TCA) corresponding to UGA (opal). The nucleotide sequences of the tRNA(Pyl)TCA gene and the tRNA(Pyl)TCA encoded by the gene are shown in SEQ ID NOs: 120 and 121, respectively.
[0291] When a PylRS DNA cassette and a tRNA(Pyl)CTA DNA cassette or a tRNA(Pyl)TCA DNA cassette are introduced into a coryneform bacterium and cultured in a medium containing a pyrrolysine analogue, the pyrrolysine analogue is taken up by the bacterium and bound to the end of tRNA(Pyl)CTA or tRNA(Pyl)TCA by PylRS, producing pyrrolysyl-tRNA. During protein translation, the pyrrolysyl-tRNA binds to the UAG or UGA codon in mRNA, and each codon is translated as a pyrrolysine analogue, resulting in the synthesis of a protein with the pyrrolysine analogue introduced at the UAG or UGA codon.
[0292] The PylRS DNA cassette was cloned upstream of the wild-type and mutant 9g8 genes in the wild-type and mutant 9g8 vectors, and the tRNA(Pyl)CTA DNA cassette or the tRNA(Pyl)TCA DNA cassette was cloned downstream of the wild-type and mutant 9g8 genes in the wild-type and mutant 9g8 vectors. These vectors were then cloned to obtain the PylRS+wild-type and mutant 9g8+tRNA(Pyl)CTA vectors and the PylRS+wild-type and mutant 9g8+tRNA(Pyl)TCA vectors.
[0293] 2. Expression of Mutant 9g8. C. glutamicum YDK0107 was transformed with either the PylRS + mutant 9g8 + tRNA(Pyl)CTA vector or the PylRS + mutant 9g8 + tRNA(Pyl)TCA vector to obtain four mutant 9g8-expressing strains, each containing one of the two mutant 9g8 genes and a PylRS DNA cassette and either a tRNA(Pyl)CTA DNA cassette or a tRNA(Pyl)TCA DNA cassette. As a control, C. glutamicum YDK0107 was transformed with either the wild-type 9g8 vector or the PylRS + wild-type 9g8 + tRNA(Pyl)CTA vector to obtain wild-type 9g8-expressing strains (also referred to as control strain PC or control strain WT, respectively). Each 9g8-expressing strain was cultured at 30°C for 72 hours in medium supplemented with 1 mM alloclysine. After the culture was completed, the culture medium of each strain was centrifuged, and 10 μl of the resulting culture supernatant was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). As a result, in the two mutant 9g8-expressing strains in which the codon combination on the mutant 9g8 gene and the anticodon on tRNA(Pyl) were correct, a band was observed at the position corresponding to the full-length of 9g8 in the presence of alloclysine, confirming the expression of alloclysine-introduced mutant 9g8 (lanes 7 and 9 in panel (A) of Figure 13).
[0294] Example 9: Expression of mutant anti-epidermal growth factor receptor (EGFR) VHH antibody 9g8 with site-specific azidophenylalanine (AzF) substitution in Escherichia coli 1. Preparation of a plasmid vector containing an AzFN3 DNA cassette (AzFN3 vector for E. coli) As with coryneform bacteria, when the AzFN3 DNA cassette (1 of Example 1) is introduced into Escherichia coli, a protein with azidophenylalanine introduced at the UAG codon position is synthesized.
[0295] The AzFN3 DNA cassette was cloned downstream of the lacI gene of a plasmid in which the drug resistance of the pCDF-1b vector (Novagen) was changed to kanamycin resistance, to obtain the AzFN3 vector for E. coli.
[0296] 2. Construction of Plasmid Vectors Containing Wild-Type or Mutant 9g8 DNA Cassettes (9g8 Vectors for E. coli) The wild-type 9g8 DNA cassette for E. coli (SEQ ID NO: 122) was obtained by total synthesis. This wild-type 9g8 DNA cassette for E. coli contains a T7 promoter and a gene encoding wild-type 9g8, linked downstream of the promoter and equipped with a pelB leader sequence at the N-terminus and a His tag sequence at the C-terminus. The wild-type 9g8 DNA cassette for E. coli was cloned into a pET26b vector (pET26b(+) vector (Novagen) in which the kanamycin drug resistance was changed to ampicillin drug resistance) to obtain a pET26b vector (wild-type 9g8 vector for E. coli) containing the wild-type 9g8 DNA cassette for E. coli.
[0297] Next, using the wild-type 9g8 vector for E. coli as a template, the triplet encoding the amino acid residue at position 32 of wild-type 9g8 was changed to a TAG triplet (this mutant 9g8 is also referred to as 9g8 (Y32TAG)) in the same manner as in Example 1-2, to obtain a pET26b vector (mutant 9g8 vector for E. coli) containing a mutant 9g8 DNA cassette for E. coli.
[0298] 3. Expression of Mutant 9g8: E. coli BL21(DE3) was transformed with the E. coli AzFN3 vector and the E. coli mutant 9g8 vector to obtain mutant 9g8-expressing strains containing both the mutant 9g8 gene and the AzFN3 DNA cassette. As a control, E. coli BL21(DE3) was transformed with the E. coli AzFN3 vector and the wild-type 9g8 vector to obtain wild-type 9g8-expressing strains. Each 9g8-expressing strain was cultured overnight at 37°C in medium supplemented with 0, 0.1, 0.3, or 1.5 mM azidophenylalanine to express wild-type 9g8 or mutant 9g8 containing azidophenylalanine. The cells were then harvested, 100 μl of TE buffer (2.5% SDS) was added, and the mixture was heated at 95°C for 5 minutes to obtain protein extracts containing wild-type or mutant 9g8. Ten microliters of each extract was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). The SDS-PAGE results were analyzed using the ChemiDoc (Bio-Rad) automated image detection system to calculate the ratio of mutant 9g8 expression to wild-type 9g8 expression (Amb / wt) at each AzF concentration (Figure 14, panel (A)). The expression of mutant 9g8 in the absence of AzF is thought to be due to nonspecific incorporation of tyrosine. Proteins from the SDS-PAGE gel were transferred to a 0.2 μm PVDF membrane (Bio-Rad), conjugated with an alkaline phosphatase-fused anti-C-terminal His tag antibody (Invitrogen), and detected by color development using an alkaline phosphatase colorimetric reagent (Bio-Rad) (Figure 14, panel (B)).
[0299] 4. Comparative Evaluation of Azidophenylalanine Incorporation Efficiency in C. glutamicum and E. coli The AzFRS DNA cassette, wild-type 9g8 or mutant 9g8 DNA cassette, and tRNA were introduced into C. glutamicum and E. coli in the same manner as in Example 1-2 and Example 8-1. CTA pPK4 vector containing the DNA cassette (AzFRS + wild-type 9g8 + tRNA CTA Vector or AzFRS + mutant 9g8 + tRNACTA A vector (Amb / wt) was constructed. This mutant 9g8 DNA cassette encodes 9g8 (Y32TAG). Following the same procedure as in Example 1, section 3, C. glutamicum was cultured in medium supplemented with 0, 0.1, 0.3, or 1.5 mM azidophenylalanine to express wild-type and mutant 9g8. After the culture was completed, 10 μl of each culture medium was subjected to reducing SDS-PAGE and stained with Coomassie Brilliant Blue (CBB) R-250 (Bio-Rad). The SDS-PAGE results were used to calculate the expression ratio (Amb / wt) of mutant 9g8 to wild-type 9g8 at each AzF concentration using the automated image detection system ChemiDoc (Bio-Rad) (Figure 15). The expression of mutant 9g8 without AzF is thought to be due to nonspecific tyrosine uptake.
[0300] At all AzF concentrations (0.1-1.5 mM), C. glutamicum exhibited a higher expression ratio (Amb / wt) than E. coli, indicating a higher azidophenylalanine incorporation efficiency than E. coli (Figure 14, panel (A) and Figure 15). Furthermore, without AzF, C. glutamicum exhibited a lower expression ratio (Amb / wt) than E. coli, indicating that nonspecific tyrosine uptake was more suppressed than in E. coli (Figure 14, panel (A) and Figure 15). These results demonstrate that coryneform bacteria can produce ncAA-containing proteins, such as azidophenylalanine-containing proteins, more efficiently and / or specifically than E. coli.
[0301] Example 10: Expression of aMD4dY-PA22 mutants containing site-specific azidophenylalanine (AzF) insertion 1. Construction of vectors (pPK14c and pPK14d) carrying an AzFRS DNA cassette in the pPK5 vector. The fully synthesized AzFRS DNA cassettes (SEQ ID NOs: 123 and 124) were integrated into the KpnI-XbaI site of the pPK5 vector described in WO2016 / 171224 by infusion reaction to construct the basic vectors pPK14c and pPK14d incorporating the AzFRS DNA cassette. The AzFRS DNA cassette (SEQ ID NO: 123) in pPK14c contains a T7 promoter and the AzFRS gene linked downstream of it, a KpnI site at the 5'-end, and ApaI and XbaI sites at the 3'-end. The AzFRS DNA cassette (SEQ ID NO: 124) in pPK14d contains the rrnC terminator and the T7 promoter and AzFRS gene linked downstream of it. It contains KpnI and BamHI sites at the 5'-end and ApaI and XbaI sites at the 3'-end. The infusion reaction was performed using the In-Fusion® HD Cloning Kit (Takara Bio) according to the manufacturer's recommended protocol. Nucleotide sequencing of the inserts confirmed that the AzFRS DNA cassette-containing vectors were constructed as designed. Nucleotide sequencing was performed using the BigDye® Terminator v3.1 Cycle Sequencing Kit (Applied Biosystems) and a 3500xL Genetic Analyzer (Applied Biosystems).
[0302] Both pPK14c and pPK14d vectors can be used as base vectors for expressing AzF transduced proteins. When subcloning an AzF transduced protein expression cassette into pPK14c, the order of insertion of the AzFRS DNA cassette and the AzF transduced protein expression cassette can be adjusted by using the KpnI or ApaI site. Similarly, when subcloning an AzF transduced protein expression cassette into pPK14d, the order of insertion of the AzFRS DNA cassette and the AzF transduced protein expression cassette can be adjusted by using the BamHI or ApaI site.
[0303] 2. Construction of a plasmid vector (aMD4dY-PA22+AzFRS vector) containing a wild-type or mutant aMD4dY-PA22 DNA cassette and an AzFRS DNA cassette. The peptide used to introduce azidophenylalanine was the HGF-PAS dimer peptide aMD4dY-PA22 described in WO2021 / 112249. aMD4dY-PA22 is a peptide with HGF-like activity.
[0304] A wild-type aMD4dY-PA22 DNA cassette (SEQ ID NO: 125) was obtained by total synthesis. Three types of mutant aMD4dY-PA22 DNA cassettes were also obtained by total synthesis for expressing mutant aMD4dY-PA22: one in which AzF was inserted at the N-terminus (between positions 3 and 4) of wild-type aMD4dY-PA22; one in which AzF was inserted within the PA22 linker (between positions 32 and 33) of wild-type aMD4dY-PA22; and one in which Pro (Pro at position 31) within the PA22 linker of wild-type aMD4dY-PA22 was substituted with AzF. The nucleotide sequence of each mutant aMD4dY-PA22 DNA cassette is identical to that of the wild-type aMD4dY-PA22 DNA cassette, except that it contains each mutant aMD4dY-PA22 gene instead of the wild-type aMD4dY-PA22 gene. Specifically, the mutant aMD4dY-PA22 gene contains an insertion or substitution of a TAG triplet corresponding to the UAG codon (amber), a stop codon, in the nucleotide sequence of the wild-type aMD4dY-PA22 gene, and the TAG triplet can be translated into AzF. Wild-type aMD4dY-PA22 is referred to as WT-aMD4dY-PA22, and the three mutant aMD4dY-PA22 types are referred to as N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, and L2-AzF-aMD4dY-PA22, respectively. The wild-type or mutant aMD4dY-PA22 DNA cassette contains a cspB promoter and a wild-type or mutant aMD4dY-PA22 gene linked downstream thereof to a CspB signal sequence added to the N-terminus. The nucleotide and amino acid sequences of the gene encoding WT-aMD4dY-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 126 and 127, respectively. The nucleotide and amino acid sequences of the gene encoding N-AzF-aMD4dY-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 128 and 129, respectively. The nucleotide and amino acid sequences of the gene encoding L1-aMD4dY-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 130 and 131, respectively. The nucleotide and amino acid sequences of the gene encoding L2-aMD4dY-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 132 and 133, respectively.
[0305] The nucleotide sequences of the wild-type and mutant aMD4dY-PA22 genes and the alignment of the wild-type and mutant aMD4dY-PA22 are shown in FIG.
[0306] Each of the synthesized aMD4dY-PA22 DNA cassettes was inserted into the ApaI site of the pPK14d vector constructed in Example 10-1 by infusion reaction to obtain wild-type or mutant aMD4dY-PA22 expression vectors, pPK14d(A)_WT-aMD4dY-PA22, pPK14d(A)_N-AzF-aMD4dY-PA22, pPK14d(A)_L1-AzF-aMD4dY-PA22, and pPK14d(A)_L2-AzF-aMD4dY-PA22. Nucleotide sequencing of the inserts confirmed that the wild-type or mutant aMD4dY-PA22 + AzFRS-carrying vectors had been constructed as designed.
[0307] 3. Expression of Mutant aMD4dY-PA22 The C. glutamicum YDK0107 strain was transformed with the AzFN3 vector described in Example 1 and the mutant aMD4dY-PA22 + AzFRS expression vectors pPK14d(A)_N-AzF-aMD4dY-PA22, pPK14d(A)_L1-AzF-aMD4dY-PA22, or pPK14d(A)_L2-AzF-aMD4dY-PA22 constructed in Example 10-2, to obtain three mutant aMD4dY-PA22-expressing strains into which one of the three mutant aMD4dY-PA22 genes, an AzFRS DNA cassette, and an AzFN3 DNA cassette were introduced. As a control, the C. glutamicum YDK0107 strain was transformed with the AzFN3 vector and the wild-type aMD4dY-PA22 + AzFRS expression vector pPK14d(A)_WT-aMD4dY-PA22 to obtain a wild-type aMD4dY-PA22-expressing strain.
[0308] Each aMD4dY-PA22-expressing strain was cultured for 96 hours at 30°C in a medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, each culture was centrifuged, and 6.5 μL of the resulting culture supernatant was subjected to reducing SDS-PAGE using NuPAGE® 12% Bis-Tirs Gel (Thermo Fisher Scientific) and then stained with Quick-CBB (Wako).
[0309] As a result, in the culture supernatants of the L1-AzF-aMD4dY-PA22-expressing strain and the L2-AzF-aMD4dY-PA22-expressing strain, no band was observed at the position corresponding to the full-length of aMD4dY-PA22 in the absence of azidophenylalanine, whereas a band was observed at the position corresponding to the full-length of aMD4dY-PA22 in the presence of azidophenylalanine, suggesting that mutant aMD4dY-PA22 into which azidophenylalanine had been introduced was expressed (lanes 6-9 in Figure 17).
[0310] In the culture supernatant of the N-AzF-aMD4dY-PA22-expressing strain, a band corresponding to the full-length aMD4dY-PA22 was observed even in the absence of azidophenylalanine (lane 4 in Figure 17). Analysis of the N-terminal amino acid sequence of this band confirmed the sequence AETY..., confirming the incorporation of tyrosine at the UAG codon. It has been previously reported that tyrosine misincorporation occurs at the UAG codon (amber) in the absence of unnatural amino acids because the codons for tyrosine are UAC or UAU (Nucleic Acids Research, 2002, 30(21), 4692-4699). This finding confirmed that tyrosine misincorporation occurs at the UAG codon (amber) in the absence of azidophenylalanine during secretion of N-AzF-aMD4dY-PA22 in C. glutamicum.
[0311] Similarly, analysis of the N-terminal amino acid sequence of the N-AzF-aMD4dY-PA22 expression band (lane 5 in Figure 17) in the culture supernatant obtained in the presence of azidophenylalanine confirmed the sequence AETF... That is, tyrosine was not detected at the azidophenylalanine incorporation site of the UAG codon, but phenylalanine was. These results demonstrate that in the secretion and expression of N-AzF-aMD4dY-PA22 by C. glutamicum, incorporation of azidophenylalanine is prioritized over misincorporation of tyrosine at the UAG codon (amber) in the presence of azidophenylalanine, enabling highly efficient secretion and expression of azidophenylalanine-containing N-AzF-aMD4dY-PA22.
[0312] Example 11: Expression of EPO-PA22 mutants with site-specific azidophenylalanine (AzF) insertion 1. Construction of a plasmid vector (EPO-PA22 + AzFRS vector) containing a wild-type or mutant EPO-PA22 DNA cassette and an AzFRS DNA cassette EPO-PA22, an EPO-PAS dimer peptide described in WO2021 / 112249, was used as the peptide into which azidophenylalanine was inserted. EPO-PA22 is a peptide with erythropoietin-like activity.
[0313] A wild-type EPO-PA22 DNA cassette (SEQ ID NO: 134) was obtained by total synthesis. Three types of mutant EPO-PA22 DNA cassettes were also obtained by total synthesis for expressing mutant EPO-PA22: one in which AzF was inserted at the N-terminus (between positions 3 and 4) of wild-type EPO-PA22; one in which AzF was inserted within the PA22 linker (between positions 35 and 36) of wild-type EPO-PA22; and one in which Pro (Pro at position 34) in the PA22 linker of wild-type EPO-PA22 was substituted with AzF. The nucleotide sequence of each mutant EPO-PA22 DNA cassette was identical to that of the wild-type EPO-PA22 DNA cassette, except that the mutant EPO-PA22 gene was replaced with the wild-type EPO-PA22 gene. Specifically, the mutant EPO-PA22 gene contains an insertion or substitution of a TAG triplet corresponding to the UAG (amber) codon, one of the stop codons in the nucleotide sequence of the wild-type EPO-PA22 gene, and this TAG triplet can be translated into AzF. Wild-type EPO-PA22 is designated WT-EPO-PA22, and the three mutant EPO-PA22 types are designated N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22, respectively. The wild-type or mutant EPO-PA22 DNA cassette contains a wild-type or mutant EPO-PA22 gene with a CspB signal sequence added to its N-terminus, linked downstream of the cspB promoter. The nucleotide and amino acid sequences of the gene encoding WT-EPO-PA22 (excluding the additional sequence) are set forth in SEQ ID NOs: 135 and 136, respectively. The nucleotide sequence and amino acid sequence of the gene encoding N-AzF-EPO-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 137 and 138, respectively. The nucleotide sequence and amino acid sequence of the gene encoding L1-AzF-EPO-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 139 and 140, respectively. The nucleotide sequence and amino acid sequence of the gene encoding L2-AzF-EPO-PA22 (excluding the additional sequence) are shown in SEQ ID NOs: 141 and 142, respectively.
[0314] The nucleotide sequences of the wild-type and mutant EPO-PA22 genes and the alignment of wild-type and mutant EPO-PA22 are shown in FIG.
[0315] Each of the fully synthesized EPO-PA22 DNA cassettes was inserted into the ApaI site of the pPK14c vector constructed in Example 10-1 by infusion reaction to obtain wild-type or mutant EPO-PA22 expression vectors, pPK14c(A)_WT-EPO-PA22, pPK14c(A)_N-AzF-EPO-PA22, pPK14c(A)_L1-AzF-EPO-PA22, and pPK14c(A)_L2-AzF-WPO-PA22. Nucleotide sequencing of the inserts confirmed that the wild-type or mutant EPO-PA22 + AzFRS-containing vectors had been constructed as designed.
[0316] 2. Expression of Mutant EPO-PA22 The C. glutamicum YDK0107 strain was transformed with the AzFN3 vector described in Example 1 and the mutant EPO-PA22+AzFRS expression vectors pPK14c(A)_N-AzF-EPO-PA22, pPK14c(A)_L1-AzF-EPO-PA22, and pPK14c(A)_L2-AzF-EPO-PA22 constructed in Example 11-1, to obtain three mutant EPO-PA22-expressing strains into which one of the three mutant EPO-PA22 genes, an AzFRS DNA cassette, and an AzFN3 DNA cassette were introduced. As a control, the C. glutamicum YDK0107 strain was transformed with the AzFN3 vector and the wild-type EPO-PA22 + AzFRS vector pPK14c(A)_WT-EPO-PA22 to obtain a wild-type EPO-PA22-expressing strain.
[0317] Each EPO-PA22-expressing strain was cultured for 96 hours at 30°C in medium supplemented with 0.3 mM azidophenylalanine. After the culture was completed, each culture was centrifuged, and 6.5 μL of the resulting culture supernatant was subjected to reducing SDS-PAGE using NuPAGE® 12% Bis-Tirs Gel (Thermo Fisher Scientific) and then stained with Quick-CBB (Wako).
[0318] As a result, in the culture supernatants of the three mutant EPO-PA22-expressing strains, almost no band was observed at the position corresponding to the full-length EPO-PA22 in the absence of azidophenylalanine, whereas a band was observed at the position corresponding to the full-length EPO-PA22 in the presence of azidophenylalanine, suggesting that mutant EPO-PA22 into which azidophenylalanine had been introduced was expressed (lanes 4-9 in Figure 19).
[0319] Example 12: Molecular modification of azidophenylalanine-introduced polypeptide 1. Molecular modification of azidophenylalanine-introduced polypeptide An example of site-selective molecular modification of an azidophenylalanine-introduced polypeptide using the strain-promoted alkyne azide cycloaddition (SPAAC) reaction is shown below.
[0320] Culture supernatants containing azidophenylalanine-introduced aMD4dY-PA22 (N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, and L2-AzF-aMD4dY-PA22), azidophenylalanine-introduced EPO-PA22 (N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22), azidophenylalanine-free aMD4dY-PA22 (WT-aMD4dY-PA22), and azidophenylalanine-free EPO-PA22 (WT-EPO-PA22) (Examples 10 and 11) were each dialyzed against 1 mL of PBS using an Amicon Ultra-0.5 3 kDa (Merck Millipore) and then concentrated to a 2.5-fold concentration. 3 mM m-dPEG (registered trademark) was added to the concentrated culture supernatant. 24 -DBCO (Quanta Biodesign) was mixed and incubated at 4°C for 72 hours.
[0321] 2. Confirmation of azidophenylalanine introduction and molecular modification using MALDI-TOF-MS Introduction of azidophenylalanine into aMD4dY-PA22 and EPO-PA22 and m-dPEG (registered trademark) 24 Modification with -DBCO was confirmed by matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF-MS) analysis.
[0322] Culture supernatants containing WT-aMD4dY-PA22, N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, L2-AzF-aMD4dY-PA22, WT-EPO-PA22, N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22, and the above-mentioned m-dPEG (registered trademark) 24The samples that underwent the -DBCO modification reaction were purified using ZipTip C18 (Merck Millipore), mixed with α-cyano-4-hydroxycinnamic acid (CHCA), and analyzed by MALDI-TOF-MS (Shimadzu, AXIMA-TOF 2 ) was used to measure the molecular weight.
[0323] The results for aMD4dY-PA22 are shown in Figure 20. WT-aMD4dY-PA22 was obtained by the same method as m-dPEG (registered trademark). 24 The molecular weight remained unchanged after the -DBCO modification, indicating that the PA22 and PA22 were not modified by the SPAAC reaction. The molecular weights of N-AzF-aMD4dY-PA22 and L1-AzF-aMD4dY-PA22 increased by approximately 162 compared to WT-aMD4dY-PA22, which corresponds to the molecular weight increase that would occur if 4-amino-L-phenylalanine were inserted into WT-aMD4dY-PA22. Azidophenylalanine is known to be converted to 4-amino-L-phenylalanine in MALDI-TOF-MS analysis. Therefore, these results indicate that azidophenylalanine is inserted into N-AzF-aMD4dY-PA22 and L1-AzF-aMD4dY-PA22. The molecular weight of L2-AzF-aMD4dY-PA22 is approximately 65 times larger than that of WT-aMD4dY-PA22, which is consistent with the increase in molecular weight that occurs when L-proline in WT-aMD4dY-PA22 is replaced with 4-amino-L-phenylalanine. This result indicates that L-proline is replaced with azidophenylalanine in N-AzF-aMD4dY-PA22 and L1-AzF-aMD4dY-PA22. Furthermore, N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, and L2-AzF-aMD4dY-PA22 are substituted with dPEG (registered trademark). 24 After the -DBCO modification reaction, the molecular weight increased by approximately 1400, which is the same as the molecular weight of 4-amino-L-phenylalanine and m-dPEG (registered trademark). 24These results indicate that azidophenylalanine was introduced into N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, and L2-AzF-aMD4dY-PA22, and that the azidophenylalanine was substituted with dPEG (registered trademark). 24 -DBCO modification was confirmed.
[0324] The results for EPO-PA22 are shown in Figure 21. WT-EPO-PA22 was m-dPEG (registered trademark). 24 The molecular weights remained unchanged after the -DBCO modification, indicating that they were not modified by the SPAAC reaction. The molecular weights of N-AzF-EPO-PA22 and L1-AzF-EPO-PA22 increased by approximately 162 compared to WT-EPO-PA22, which corresponds to the molecular weight increase that would occur if 4-amino-L-phenylalanine were inserted into WT-EPO-PA22. This result indicates that azidophenylalanine was inserted in N-AzF-EPO-PA22 and L1-AzF-EPO-PA22. The molecular weight of L2-AzF-EPO-PA22 increased by approximately 65 compared to WT-EPO-PA22, which corresponds to the molecular weight increase that would occur if L-proline in WT-EPO-PA22 were replaced with 4-amino-L-phenylalanine. This result indicates that L-proline was replaced with azidophenylalanine in L2-AzF-EPO-PA22. In addition, N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22 were dPEG (registered trademark) 24 After the -DBCO modification reaction, the molecular weight increased by approximately 1400, which is the same as the molecular weight of 4-amino-L-phenylalanine and m-dPEG (registered trademark). 24 These results indicate that azidophenylalanine was introduced into N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22, and that the azidophenylalanine was substituted for dPEG (registered trademark). 24 -DBCO modification was confirmed.
[0325] 3. Confirmation of the cellular activity of PEG-modified azidophenylalanine-introduced aMD4dY-PA22 using a reporter assay. The introduction of unnatural amino acids or molecular modifications into a polypeptide can result in the loss of activity of the original polypeptide. It has been reported that aMD4dY-PA22 induces Met-Erk-SRE signaling, and EPO-PA22 induces EPOR-JAK signaling (WO2021 / 112249; Communications Biology, 5, 56 (2022)). Therefore, we investigated the cellular activity of the azidophenylalanine-introduced aMD4dY-PA22 and EPO-PA22, as well as the azidophenylalanine-introduced m-dPEG (registered trademark). 24 The cellular activity of -DBCO-modified aMD4dY-PA22 and EPO-PA22 was evaluated using a reporter assay to confirm whether the polypeptide activity was maintained after the introduction of unnatural amino acids and molecular modifications.
[0326] The activity of aMD4dY-PA22 was evaluated by the following reporter assay. 0.6 μL of Attractene Transfection Reagent (QIAGEN) was added to 25 μL of Opti-MEM medium (Thermo Fisher Scientific) and incubated at room temperature for 5 minutes. A mixture of 25 μL of Opti-MEM and 1 μL of SRE reporter vector (QIAGEN) was added to this mixture and incubated at room temperature for 20 minutes. This mixture was added to a 96-well cell culture plate, and HEK293 cells were seeded on top at a density of 40,000 cells / well. The cells were then cultured overnight at 37°C and 5% CO2 for transfection. After transfection, all culture supernatant was removed, and 100 μL of evaluation medium (Opti-MEM medium containing 0.5% FBS (Thermo Fisher Scientific), 1% non-essential amino acid solution (Thermo Fisher Scientific), and penicillin-streptomycin (Nacalai Tesque)) was added. The cells were then cultured at 37°C for 4 hours to starve the cells. WT-aMD4dY-PA22 and azidophenylalanine-conjugated aMD4dY-PA22 (N-AzF-aMD4dY-PA22, L1-AzF-aMD4dY-PA22, and L2-AzF-aMD4dY-PA22) were treated with m-dPEG (registered trademark). 24The PEG-modified sample (PEG+) and the unmodified sample (PEG-) were each diluted 2500-fold with the assay medium, and 100 μL of this mixture was added to the cells. Cells were stimulated by overnight incubation at 37°C in a 5% CO2 incubator. Signal intensity was measured using the Dual-Luciferase Reporter Assay System (Promega). 100 μL of culture supernatant was removed, and 50 μL of Glo reagent was added. Cells were lysed at room temperature for 10 minutes. Erk-SRE pathway activation was quantified by detecting Firefly luciferase luminescence from the cell lysate using a plate reader. Next, 50 μL of Glo & Stop reagent was added, and after 10 minutes, the luminescence of the internal standard, Renilla luciferase, was detected and cell number was quantified. For each well, the signal activity value was calculated as follows: (Firefly luciferase luminescence intensity) / (Renilla luciferase luminescence intensity). For each sample-added group, the relative signal activity value to the non-sample-added group (Mock) was determined and used as the reporter activity.
[0327] The results of activity evaluation are shown in Figure 22. In all sample-added groups, reporter activity was 10 times or more higher than in the sample-free group (Mock). 24 It was revealed that the activity of the original aMD4dY-PA22 molecule was not lost by the -DBCO molecular modification.
[0328] The activity of EPO-PA22 was evaluated by a reporter assay using the PathHunter® eXpress EpoR-JAK2 Functional Assay Kit (DiscoverX). Cells for EPO activity evaluation included in the PathHunter® eXpress EpoR-JAK2 Functional Assay Kit were seeded into a 384-well plate and cultured at 37°C and 5% CO2 for 24 hours. WT-EPO-PA22 and azidophenylalanine-conjugated EPO-PA22 (N-AzF-EPO-PA22, L1-AzF-EPO-PA22, and L2-AzF-EPO-PA22) were transfected with m-dPEG®. 24 A sample modified with PEG (PEG+) and an unmodified sample (PEG-) with -DBCO were diluted 1000-fold and added to cells. After 3 hours of stimulation at room temperature, the prepared Substrate Reagent was added and incubated for 60 minutes. Chemiluminescence intensity was quantified using a Nivo plate reader (Perkin Elmer). The relative chemiluminescence intensity of each sample-added group compared to the non-sample-added group (Mock) was calculated and used as reporter activity.
[0329] The activity evaluation results are shown in Figure 23. In all sample-added groups, reporter activity was 1.5 times or more higher than in the sample-free group (Mock). This indicates that the introduction of unnatural amino acids and dPEG (registered trademark) increased the reporter activity. 24 It was revealed that the activity of EPO-PA22 was not lost by the -DBCO molecular modification.
[0330] According to the present invention, proteins containing unnatural amino acids (ncAAs) can be efficiently secreted and produced.
[0331] [Explanation of the sequence listing] SEQ ID NO: 1: Nucleotide sequence of the phoS gene of C. glutamicum YDK010 SEQ ID NO: 2: Amino acid sequence of the PhoS protein of C. glutamicum YDK010 SEQ ID NO: 3: Amino acid sequence of the PhoS protein of C. glutamicum ATCC 13032 SEQ ID NO: 4: Amino acid sequence of the PhoS protein of C. glutamicum ATCC 14067 SEQ ID NO: 5: Amino acid sequence of the PhoS protein of C. callunae SEQ ID NO: 6: Amino acid sequence of the PhoS protein of C. crenatum SEQ ID NO: 7: Amino acid sequence of the PhoS protein of C. efficiens SEQ ID NO: 8: Nucleotide sequence of the phoR gene of C. glutamicum ATCC 13032 SEQ ID NO: 9: Amino acid sequence of the PhoR protein of C. glutamicum ATCC 13032 SEQ ID NO: 10: Nucleotide sequence of the cspB gene of C. glutamicum ATCC 13869 SEQ ID NO: 11: Amino acid sequence of the PhoS protein of C. glutamicum ATCC 13869 SEQ ID NO: 12: Nucleotide sequence of the tatA gene of C. glutamicum ATCC 13032 SEQ ID NO: 13: Amino acid sequence of the TatA protein of C. glutamicum ATCC 13032 SEQ ID NO: 14: Nucleotide sequence of the tatB gene of C. glutamicum ATCC 13032 SEQ ID NO: 15: Amino acid sequence of the TatB protein of C. glutamicum ATCC 13032 SEQ ID NO: 16: Nucleotide sequence of the tatC gene of C. glutamicum ATCC 13032 SEQ ID NO: 17: Amino acid sequence of the TatC protein of C. glutamicum ATCC 13032 SEQ ID NO: 18: Amino acid sequence of the TorA signal peptide SEQ ID NO: 19: Amino acid sequence of the SufI signal peptide SEQ ID NO: 20: Amino acid sequence of the PhoD signal peptide SEQ ID NO: 21: Amino acid sequence of the LipA signal peptide SEQ ID NO: 22: Amino acid sequence of the IMD signal peptide SEQ ID NO: 23: Amino acid sequence of the twin arginine motif SEQ ID NO: 24: Skipped sequence SEQ ID NO: 25: Amino acid sequence of the PS1 signal peptideSEQ ID NO: 26: Amino acid sequence of PS2 signal peptide SEQ ID NO: 27: Amino acid sequence of SlpA signal peptide SEQ ID NO: 28: Amino acid sequence of CspB mature protein of C. glutamicum ATCC 13869 SEQ ID NOs: 29 to 31: Skipped sequences SEQ ID NOs: 32 to 36: Amino acid sequences of one embodiment of the insertion sequence used in the present invention SEQ ID NO: 37: Recognition sequence for Factor Xa protease SEQ ID NO: 38: Recognition sequence for ProTEV protease SEQ ID NO: 39: Nucleotide sequence of tRNA(Tyr) gene of Methanocaldococcus jannaschii SEQ ID NO: 40: Nucleotide sequence of tRNA(Tyr) of Methanocaldococcus jannaschii SEQ ID NO: 41: Nucleotide sequence of modified tRNA(Tyr) gene of Methanocaldococcus jannaschii SEQ ID NO: 42: Nucleotide sequence of modified tRNA(Tyr) of Methanocaldococcus jannaschii SEQ ID NO: 43: Methanocaldococcus SEQ ID NO: 44: Nucleotide sequence of modified tRNA(Tyr) gene of Methanocaldococcus jannaschii SEQ ID NO: 45: Nucleotide sequence of tRNA(Pyl) gene of Methanosarcina barkeri SEQ ID NO: 46: Nucleotide sequence of tRNA(Pyl) of Methanosarcina barkeri SEQ ID NO: 47: Nucleotide sequence of Tyr-RS gene of Methanocaldococcus jannaschii SEQ ID NO: 48: Amino acid sequence of Tyr-RS of Methanocaldococcus jannaschii SEQ ID NO: 49: Nucleotide sequence of modified Tyr-RS gene of Methanocaldococcus jannaschii SEQ ID NO: 50: Amino acid sequence of modified Tyr-RS of Methanocaldococcus jannaschii SEQ ID NO: 51: Nucleotide sequence of modified Tyr-RS gene of Methanocaldococcus jannaschii SEQ ID NO: 52: Methanocaldococcus Amino acid sequence of modified Tyr-RS of Methanosarcina jannaschii SEQ ID NO: 53barkeri Pyl-RS gene SEQ ID NO:54: Amino acid sequence of Pyl-RS of Methanosarcina barkeri SEQ ID NO:55: Nucleotide sequence of AzFN3 DNA cassette SEQ ID NO:56: Nucleotide sequence of wild-type 9g8 DNA cassette SEQ ID NO:57: Nucleotide sequence of wild-type 9g8 gene SEQ ID NO:58: Amino acid sequence of wild-type 9g8 SEQ ID NOs:59-80: Primers SEQ ID NO:81: Nucleotide sequence of AzFRS DNA cassette SEQ ID NO:82: Nucleotide sequence of wild-type ZHER2 affibody DNA cassette SEQ ID NO:83: Nucleotide sequence of wild-type ZHER2 affibody gene SEQ ID NO:84: Amino acid sequence of wild-type ZHER2 affibody SEQ ID NOs:85-92: Primers SEQ ID NO:93: Nucleotide sequence of wild-type mRFP DNA cassette SEQ ID NO:94: Nucleotide sequence of wild-type mRFP gene SEQ ID NO:95: Amino acid sequence of wild-type mRFP SEQ ID NOs:96-99: Primers SEQ ID NO:100: tRNACTA Nucleotide sequences of DNA cassettes SEQ ID NO: 101: Nucleotide sequence of IYN3 DNA cassette SEQ ID NO: 102: Nucleotide sequence of wild-type N15 DNA cassette SEQ ID NO: 103: Nucleotide sequence of wild-type N15 gene SEQ ID NO: 104: Amino acid sequence of wild-type N15 SEQ ID NOs: 105 to 112: Primers SEQ ID NO: 113: Nucleotide sequence of PylRS DNA cassette SEQ ID NO: 114: Nucleotide sequence of Pyl-RS gene of Methanosarcina mazei SEQ ID NO: 115: Amino acid sequence of Pyl-RS of Methanosarcina mazei SEQ ID NO: 116: Nucleotide sequence of tRNA(Pyl)CTA DNA cassette SEQ ID NO: 117: Nucleotide sequence of tRNA(Pyl)TCA DNA cassette SEQ ID NO: 118: Nucleotide sequence of modified tRNA(Pyl) gene of Methanosarcina mazei SEQ ID NO: 119: Nucleotide sequence of modified tRNA(Pyl) of Methanosarcina mazei SEQ ID NO: 120: Methanosarcina SEQ ID NO: 121: Nucleotide sequence of modified tRNA(Pyl) gene from Methanosarcina mazei SEQ ID NO: 122: Nucleotide sequence of wild-type 9g8 DNA cassette for E. coli SEQ ID NO: 123: Nucleotide sequence of AzFRS DNA cassetteSEQ ID NO: 124: Nucleotide sequence of AzFRS DNA cassette SEQ ID NO: 125: Nucleotide sequence of wild-type aMD4dY-PA22 DNA cassette SEQ ID NO: 126: Nucleotide sequence of WT-aMD4dY-PA22 gene SEQ ID NO: 127: Amino acid sequence of WT-aMD4dY-PA22 SEQ ID NO: 128: Nucleotide sequence of N-AzF-aMD4dY-PA22 gene SEQ ID NO: 129: Amino acid sequence of N-AzF-aMD4dY-PA22 SEQ ID NO: 130: Nucleotide sequence of L1-AzF-aMD4dY-PA22 gene SEQ ID NO: 131: Amino acid sequence of L1-AzF-aMD4dY-PA22 SEQ ID NO: 132: Nucleotide sequence of L2-AzF-aMD4dY-PA22 gene SEQ ID NO: 133: Amino acid sequence of L2-AzF-aMD4dY-PA22 SEQ ID NO: 134: Wild-type EPO-PA22 Nucleotide sequences of DNA cassettes SEQ ID NO: 135: Nucleotide sequence of WT-EPO-PA22 gene SEQ ID NO: 136: Amino acid sequence of WT-EPO-PA22 SEQ ID NO: 137: Nucleotide sequence of N-AzF-EPO-PA22 gene SEQ ID NO: 138: Amino acid sequence of N-AzF-EPO-PA22 SEQ ID NO: 139: Nucleotide sequence of L1-AzF-EPO-PA22 gene SEQ ID NO: 140: Amino acid sequence of L1-AzF-EPO-PA22 SEQ ID NO: 141: Nucleotide sequence of L2-AzF-EPO-PA22 gene SEQ ID NO: 142: Amino acid sequence of L2-AzF-EPO-PA22
Claims
1. A method for producing a protein containing a non-natural amino acid, comprising: culturing a coryneform bacterium having a gene construct for secretory expression of the protein containing the non-natural amino acid in a medium containing the non-natural amino acid; and recovering the secreted and produced protein containing the non-natural amino acid, wherein the coryneform bacterium is modified to express an orthogonal pair of a tRNA corresponding to the non-natural amino acid and an aminoacyl-tRNA synthetase.
2. The gene construct contains, in the 5' to 3' direction, a promoter sequence functional in a coryneform bacterium, a nucleic acid sequence encoding a signal peptide functional in a coryneform bacterium, and a nucleic acid sequence encoding a protein containing a non-natural amino acid, and the protein containing the non-natural amino acid is expressed as a fusion protein with the signal peptide, according to the method of Claim 1.
3. The method according to Claim 1 or 2, wherein the non-natural amino acid is encoded by a stop codon or a four-residue codon.
4. The method according to Claim 3, wherein the stop codon is UAG or UGA.
5. The method according to Claim 1 or 2, wherein the non-natural amino acid is a tyrosine derivative or a lysine derivative.
6. The unnatural amino acid is p-azido-L-phenylalanine, 3-azido-L-tyrosine, 3-chloro-L-tyrosine, 3-nitro-L-tyrosine, O-sulfo-L-tyrosine, L-pyrrolidine, or N δ -alloc-L-lysine, and the method according to claim 1 or 2.
7. The method according to Claim 1 or 2, wherein the tRNA is tRNA(Tyr) or tRNA(Pyl).
8. The method according to Claim 7, wherein the tRNA(Tyr) is an RNA described in the following (a), (b), or (c): (a) an RNA containing the nucleotide sequence shown in SEQ ID NO: 42 or 44; (b) an RNA containing a nucleotide sequence in which the anticodon is modified in the nucleotide sequence shown in SEQ ID NO: 40, 42, or 44; (c) an RNA having a nucleotide sequence having 90% or more identity to the nucleotide sequence of the RNA described in (a) or (b) above and having a function as a tRNA corresponding to the non-natural amino acid.
9. The method according to Claim 7, wherein the tRNA(Pyl) is an RNA described in the following (a), (b), or (c): (a) an RNA containing the nucleotide sequence shown in SEQ ID NO: 46, 119, or 121; (b) an RNA containing a nucleotide sequence in which the anticodon is modified in the nucleotide sequence shown in SEQ ID NO: 46, 119, or 121; (c) an RNA having a nucleotide sequence having 90% or more identity to the nucleotide sequence of the RNA described in (a) or (b) above RNA that contains the nucleotide sequence to be used and has a function as a tRNA corresponding to the unnatural amino acid.
10. The method according to claim 1 or 2, wherein the aminoacyl-tRNA synthetase is tyrosyl-tRNA synthetase or pyrrolidyl-tRNA synthetase.
11. The method according to claim 10, wherein the tyrosyl-tRNA synthetase is a protein described in any one of the following (a), (b), (c), or (d): (a) A protein containing the amino acid sequence shown in SEQ ID NO: 50 or 52; (b) A protein containing an amino acid sequence having a mutation that modifies substrate specificity in the amino acid sequence shown in SEQ ID NO: 48, 50, or 52 and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (c) A protein containing an amino acid sequence including substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues in the amino acid sequence of the protein described in (a) or (b) above; (d) A protein containing an amino acid sequence having 90% or more identity to the amino acid sequence of the protein described in (a) or (b) above.
12. The method according to claim 11, wherein the mutation that modifies the substrate specificity includes a mutation at one or more amino acid residues selected from the following: Y32, H70, E107, D158, I159, L162, D286.
13. The method according to claim 10, wherein the pyrrolidyl-tRNA synthetase is a protein described in any one of the following (a), (b), (c), or (d): (a) A protein containing the amino acid sequence shown in SEQ ID NO: 54 or 115; (b) A protein containing an amino acid sequence having a mutation that modifies substrate specificity in the amino acid sequence shown in SEQ ID NO: 54 or 115 and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (c) A protein containing an amino acid sequence including substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues in the amino acid sequence of the protein described in (a) or (b) above and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid; (d) A protein containing an amino acid sequence having 90% or more identity to the amino acid sequence of the protein described in (a) or (b) above and having aminoacyl-tRNA synthetase activity corresponding to the unnatural amino acid.
14. The method according to claim 13, wherein the mutation for modifying the substrate specificity comprises a mutation at one or more amino acid residues selected from the following: M241, L266, A267, L270, Y271, L274, N311, C313, M315, Y349, V367, W383.
15. The method according to claim 1 or 2, wherein the coryneform bacterium is further modified to harbor a phoS gene encoding a mutant PhoS protein.
16. The method according to claim 15, wherein the mutation is a mutation in which the amino acid residue corresponding to the tryptophan residue at position 302 of SEQ ID NO: 2 in the wild-type PhoS protein is substituted with an amino acid residue other than an aromatic amino acid and histidine.
17. The method according to claim 16, wherein the amino acid residue other than the aromatic amino acid and histidine is a lysine residue, an alanine residue, a valine residue, a serine residue, a cysteine residue, a methionine residue, an aspartic acid residue, or an asparagine residue.
18. The method according to claim 16, wherein the wild-type PhoS protein is a protein described in the following (a), (b), or (c): (a) a protein comprising the amino acid sequence shown in any one of SEQ ID NOs: 2 to 7; (b) a protein comprising an amino acid sequence containing substitution, deletion, insertion, and / or addition of 1 to 10 amino acid residues in the amino acid sequence shown in any one of SEQ ID NOs: 2 to 7, and having a function as a sensor kinase of the PhoRS system ; (c) a protein comprising an amino acid sequence having 90% or more identity to the amino acid sequence shown in any one of SEQ ID NOs: 2 to 7, and having a function as a sensor kinase of the PhoRS system .
19. The method according to claim 2, wherein the signal peptide is a Tat-dependent signal peptide.
20. The method according to claim 19, wherein the Tat-dependent signal peptide is any one signal peptide selected from the group consisting of TorA signal peptide, SufI signal peptide, PhoD signal peptide, LipA signal peptide, and IMD signal peptide.
21. The method according to claim 19, wherein the coryneform bacterium is further modified such that the expression of one or more genes selected from the genes encoding the Tat-dependent secretion apparatus is increased as compared with the unmodified strain.
22. The method according to claim 21, wherein the gene encoding the Tat secretion apparatus consists of the tatA gene, the tatB gene, the tatC gene, and the tatE gene.
23. The method according to claim 2, wherein the signal peptide is a Sec system-dependent signal peptide.
24. The method according to claim 23, wherein the Sec system-dependent signal peptide is any one signal peptide selected from the group consisting of the PS1 signal peptide, the PS2 signal peptide, and the SlpA signal peptide.
25. The method according to claim 2, wherein the gene construct further comprises a nucleic acid sequence encoding an amino acid sequence containing Gln-Glu-Thr between a nucleic acid sequence encoding a signal peptide that functions in coryneform bacteria and a nucleic acid sequence encoding a non-natural amino acid-containing protein.
26. The method according to claim 25, wherein the gene construct further comprises a nucleic acid sequence encoding an amino acid sequence used for enzymatic cleavage between a nucleic acid sequence encoding an amino acid sequence containing Gln-Glu-Thr and a nucleic acid sequence encoding a non-natural amino acid-containing protein.
27. The method according to claim 1 or 2, wherein the coryneform bacterium is a bacterium belonging to the genus Corynebacterium.
28. The method according to claim 27, wherein the coryneform bacterium is Corynebacterium glutamicum.
29. The method according to claim 28, wherein the coryneform bacterium is a modified strain derived from Corynebacterium glutamicum AJ12036 (FERM BP-734) or a modified strain derived from Corynebacterium glutamicum ATCC 13869.
30. The method according to claim 1 or 2, wherein the coryneform bacterium is a coryneform bacterium in which the number of molecules of the cell surface protein per cell is reduced as compared with the unmodified strain.
31. The method according to claim 1 or 2, wherein the coryneform bacterium has a first expression vector carrying the gene construct and a second expression vector carrying the gene encoding the tRNA and the gene encoding the aminoacyl-tRNA synthetase.
32. The method according to claim 31, wherein the first expression vector further carries the gene encoding the tRNA and / or the gene encoding the aminoacyl-tRNA synthetase.
33. The method according to claim 31, wherein the first expression vector is a pPK-based vector and the second expression vector is a pVC-based vector.
34. The method according to claim 31, wherein the first expression vector is pPK4 or pPK5 and the second expression vector is pVC7 or pVC7N.
35. The method according to claim 1 or 2, wherein the coryneform bacterium has a single expression vector carrying the gene construct, the gene encoding the tRNA, and the gene encoding the aminoacyl-tRNA synthetase.
36. The method according to claim 35, wherein the expression vector is a pPK-based vector.
37. The method according to claim 35, wherein the expression vector is pPK4 or pPK5.
38. The method according to claim 1 or 2, wherein the non-natural amino acid-containing protein is an antibody-related molecule, an antibody mimetic, or a bioactive protein.
39. The non-natural amino acid-containing protein is a VHH fragment, the Z domain of protein A, a fluorescent protein, or a growth factor. The method according to claim 1 or 2.