Aminoacyl-tRNA synthetase set

By developing new classes of pyrrolysyl-tRNA synthetases without the N-terminal domain and modifying their specificity, the challenge of discovering orthogonal pairs is addressed, allowing for the efficient incorporation of multiple non-canonical amino acids and synthesis of diverse polymers in cells.

JP2026518309APending Publication Date: 2026-06-04UNITED KINGDOM RESEARCH AND INNOVATION

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
UNITED KINGDOM RESEARCH AND INNOVATION
Filing Date
2024-05-29
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Despite advances in incorporating non-canonical amino acids into proteins and synthesizing polymers, there is a lack of effective criteria for discovering mutually orthogonal pyrrolysyl-tRNA synthetase/tRNA pairs, limiting the ability to incorporate multiple non-canonical amino acids and synthesize diverse polymers in cells.

Method used

The development of a new class of pyrrolysyl-tRNA synthetases (PylRS) lacking the N-terminal domain, combined with modified derivatives, allows for the creation of triple orthogonal pairs, enabling the incorporation of several non-canonical amino acids and synthesis of polymers by identifying and modifying PylRS enzymes into different classes (A, B, and C) to achieve mutual orthogonality.

Benefits of technology

This approach enables the site-specific incorporation of multiple non-canonical amino acids and synthesis of non-canonical polymers and macrocyclic molecules in cells, expanding the genetic code's versatility and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026518309000001_ABST
    Figure 2026518309000001_ABST
Patent Text Reader

Abstract

The present invention relates to classes of aminoacyl-tRNA synthetases (aRS), pyrrolidyl-tRNA synthetases (PylRS), sets of aRS, and cells expressing said aRS and / or sets of aRS. The present invention also relates to methods for generating sets of aRS and methods for generating cells containing exogenous sets of aRS. Furthermore, the present invention relates to methods for generating polymers and the use of cells for generating polymers.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a class of acyl-tRNA synthetases (aRS, acyl-tRNA synthetase), pyrrolysyl-tRNA synthetases (PylRS, pyrrolysyl-tRNA synthetase), a set of aRS, and cells expressing said aRS and / or a set of aRS. The present invention also relates to a method for generating a set of aRS, and a method for generating cells containing an exogenous set of aRS. Furthermore, the present invention relates to a method for generating polymers, and the use of cells for generating polymers. [Background technology]

[0002] By reprogramming the genetic code of living cells, it is possible to site-specifically incorporate non-canonical amino acids (ncAAs) and hydroxy acids into proteins, and to synthesize encoded non-canonical polymers, macrocyclic peptides, and depsipeptides. (Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53 (2017). https: / / doi.org:10.1038 / nature24031, De La Torre, D. & Chin, JW Reprogramming the genetic code. Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7, Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057-1062 (2021). https: / / doi.org:10.1126 / science.abg3029, Spinck, M. et al. Genetically programmed cell-based synthesis of non-natural peptide and depsipeptide macrocycles. Nature Chemistry) These advances are supported by the discovery of aminoacyl-tRNA synthetases (aaRS) and tRNAs whose aminoacylation specificity is orthogonal to the host organism's synthetase and tRNA, and mutually orthogonal to each other. One such pair is Methanococcus janaschii (Mj) tyrosyl-tRNA synthetase (TyrRS) / MjtRNA. TyrArchaeoglobus fulgidus (Af)TyrRS / AftRNA Tyr Methanococcus maripaludis (Mmp) phosphoseryl-tRNA synthetase (SepRS) / MjtRNA Sep Saccharomyces cerevisiae (Sc) tryptophanyl-tRNA synthetase (TrpRS) / SctRNA Trp ), Methanosarcina mazei (Mm) or Methanosarcina barkeri (Mb) pyrrolidyl-tRNA synthetase (PylRS) / MmtRNA Pyl or MbtRNA Pyl , as well as Candidatus Methanomethylophilus species 1R26 (1R26)PylRS / Candidatus Methanomethylophilus alvus (Alv) tRNA Pyl-8 and Methanomassiliicoccus luminyensis 1 (Lum1)PylRS / Candidatus Methanomassiliicoccus intestinalis (Int)tRNA Pyl-17C10These include modified, mutually orthogonal PylRS / tRNAPyl pairs (Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol., 1-11 (2020). https: / / doi.org:10.1038 / s41587-020-0479-2, Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5, Srinivasan, G., James, CM & Krzycki, JA Pyrrolysine Encoded by UAG in Archaea: Charging of a UAG-Decoding Specialized tRNA. Science 296, 1459--1462 (2002). https: / / doi.org:10.1126 / science.1069588, Krzycki, JA The direct genetic encoding of pyrrolysine. Curr. Opin. Microbiol. 8, 706--712 (2005). https: / / doi.org:10.1016 / j.mib.2005.10.009、Neumann, H., Peak-Chew, S. Y. & Chin, J. W. Genetically encoding Nε-acetyllysine in recombinant proteins. Nat. Chem. Biol. 4, 232 (2008). https: / / doi.org:10.1038 / nchembio.73、Wang, L., Brock, A., Herberich, B. & Schultz, P. G. Expanding the Genetic Code of Escherichia coli. Science 292, 498--500 (2001). https: / / doi.org:10.1126 / science.1060077、Borrel, G. et al. Unique Characteristics of the Pyrrolysine System in the 7th Order of Methanogens: Implications for the Evolution of a Genetic Code Expansion Cassette. Archaea 2014, 11 (2014). https: / / doi.org:10.1155 / 2014 / 374146、Park, H.-S. et al. Expanding the Genetic Code of Escherichia coli with Phosphoserine. Science 333, 1151--1154 (2011). https: / / doi.org:10.1126 / science.1207203、Rogerson, D. T. et al. Efficient genetic encoding of phosphoserine and its nonhydrolyzable analog. Nat. Chem. Biol. 11, 496 (2015). https: / / doi.org:10.1038 / nchembio.1823、Hughes, R. A. & Ellington, A. D.Rational design of an orthogonal tryptophanyl nonsense suppressor tRNA. Nucleic Acids Research 38, 6813-6830 (2010). https: / / doi.org:10.1093 / nar / gkq521, Chatterjee, A., Sun, SB, Furman, JL, Xiao, H. & Schultz, PG A Versatile Platform for Single- and Multiple-Unnatural Amino Acid Mutagenesis in Escherichia coli. American Chemical Society (2013). https: / / doi.org:10.1021 / bi4000244), these have been modified to recognize different amino acids. In early studies, ncAA was incorporated in response to the amber codon, but recent studies have shown the use of additional stop codons (Italia, JS et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. (2019). https: / / doi.org:10.1021 / jacs.8b12954) and quadruplet codons (Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441 (2010). https: / / doi.org:10.1038 / nature08817, Wang, K. et al.).Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393 (2014). https: / / doi.org:10.1038 / nchem.1919, Anderson, J. C. et al. An expanded genetic code with a functional quadruplet codon. Proceedings of the National Academy of Sciences 101, 7566-7571 (2004). https: / / doi.org:10.1073 / pnas.0401517101, Dunkelmann, D. L., Oehm, S. B., Beattie, A. T. & Chin, J. W. A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA design. Nature Chemistry 13, 1110-1117 (2021). https: / / doi.org:10.1038 / s41557-021-00764-5), codons containing non-standard bases (Malyshev, D. A. et al. A semi-synthetic organism with an expanded genetic alphabet. Nature 509, 385-388 (2014). https: / / doi.org:10.1038 / nature13314, Fischer, E. C. et al. New codons for efficient production of unnatural proteins in a semisynthetic organism. Nature Chemical Biology 16, 570-576 (2020). https: / / doi.org:10.1038 / s41589-020-0507-z, Zhang, Y. et al. A semi-synthetic organism that stores and retrieves increased genetic information. Nature 551, 644 (2017). https: / / doi.org:10.1038 / nature24659), and sense codons in organisms with genome coding compression and tRNA deletion (Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057-1062 (2021). https: / / doi.org:10.1126 / science.abg3029, Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514-518 (2019). Other codons are also used, including those mentioned in https: / / doi.org:10.1038 / s41586-019-1192-5 and Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59 (2016). https: / / doi.org:10.1038 / nature20124. Mutually orthogonal pairs are the basis for the incorporation of ncAA combinations and the synthesis of encoded polymers in cells, but despite recent advances, the discovery of such pairs remains an unresolved challenge. (Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53 (2017). https: / / doi.org:10.1038 / nature24031, De La Torre, D. & Chin, JW Reprogramming the genetic code.Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7、Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol., 1-11 (2020). https: / / doi.org:10.1038 / s41587-020-0479-2、Dunkelmann, D. L., Willis, J. C. W., Beattie, A. T. & Chin, J. W. Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x、Italia, J. S. et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. (2019). https: / / doi.org:10.1021 / jacs.8b12954、Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, J. W. Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441 (2010). https: / / doi.org:10.1038 / nature08817、Neumann, H., Slusarczyk, A. L. & Chin, J. W. De Novo Generation of Mutually Orthogonal Aminoacyl-tRNA Synthetase / tRNA Pairs. American Chemical Society (2010). https: / / doi.org:10.1021 / ja9068722、Beranek, V., Willis, J. C. W. & Chin, J. W. An Evolved Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase / tRNA Pair Is Highly Active and Orthogonal in Mammalian Cells. Biochemistry 58, 387-390 (2019). https: / / doi.org:10.1021 / acs.biochem.8b00808、Chatterjee, A., Xiao, H. & Schultz, P. G. Evolution of multiple, mutually orthogonal prolyl-tRNA synthetase / tRNA pairs for unnatural amino acid mutagenesis in Escherichia coli. Proc. Natl. Acad. Sci. U.S.A. 109, 14841--14846 (2012). https: / / doi.org:10.1073 / pnas.1212454109、Italia, J. S. et al. An orthogonalized platform for genetic code expansion in both bacteria and eukaryotes. Nature Chemical Biology 13, 446-450 (2017). https: / / doi.org:10.1038 / nchembio.2312).

[0003] ピロリジル-tRNAシンテターゼPylRS / tRNA PylPairs are the most widely used system for extending the genetic code. (De La Torre, D. & Chin, JW Reprogramming the genetic code. Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7) These pairs enable site-specific incorporation of ncAAs in all domains of life (Chin, JW Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu. Rev. Biochem. 83, 379--408 (2014). https: / / doi.org:10.1146 / annurev-biochem-060713-035737), and the anticodon of the tested pyl tRNA is not a recognition element of the PylRS enzyme (Suzuki, T. et al. Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 13, 1261 (2017). https: / / doi.org:10.1038 / nchembio.2497, and mutations can be induced to decode diverse codons (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Wang, K. et al.Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393 (2014). https: / / doi.org:10.1038 / nchem.1919, Dunkelmann, DL, Oehm, SB, Beattie, AT & Chin, JW A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA design. Nature Chemistry 13, 1110-1117 (2021). https: / / doi.org:10.1038 / s41557-021-00764-5, Ambrogelly, A. et al. Pyrrolysine is not hardwired for cotranslational insertion at UAG codons. Proc. Natl. Acad. Sci. USA 104, 3141--3146 (2007). https: / / doi.org:10.1073 / pnas.0611634104, Elliott, TS et al. Proteome labeling and protein identification in specific tissues and at specific developmental stages in an animal. Nat. Biotechnol. 32, 465 (2014). https: / / doi.org:10.1038 / nbt.2860), The active site of PylRS does not recognize standard amino acids and can accept a variety of ncAAs and hydroxy acids, or can be evolved to accept them. (Spinck, M. et al.)Genetically programmed cell-based synthesis of non-natural peptide and depsipeptidemacrocycles. Nature Chemistry、Neumann, H., Peak-Chew, S. Y. & Chin, J. W. Genetically encoding Nε-acetyllysine in recombinant proteins. Nat. Chem. Biol. 4, 232 (2008). https: / / doi.org:10.1038 / nchembio.73、Kobayashi, T., Yanagisawa, T., Sakamoto, K. & Yokoyama, S. Recognition of Non-α-amino Substrates by Pyrrolysyl-tRNA Synthetase. J. Mol. Biol. 385, 1352-1360 (2009). https: / / doi.org:10.1016 / j.jmb.2008.11.059、Polycarpo, C. R. et al. Pyrrolysine analogues as substrates for pyrrolysyl-tRNA synthetase. FEBS Letters 580, 6695-6700 (2006). https: / / doi.org:10.1016 / j.febslet.2006.11.028、Bindman, N. A., Bobeica, S. C., Liu, W. R. & Van Der Donk, W. A. Facile Removal of Leader Peptides from Lanthipeptides by Incorporation of a Hydroxy Acid. J. Am. Chem. Soc. 137, 6975-6978 (2015). https: / / doi.org:10.1021 / jacs.5b04681、Li, Y.-M. et al. Ligation of Expressed Protein α-Hydrazides . viaGenetic Incorporation of an α-Hydroxy Acid. ACS Chemical Biology 7, 1015-1022 (2012). https: / / doi.org:10.1021 / cb300020s, Ohtake, K. et al. Engineering an Automaturing Transglutaminase with Enhanced Thermostability by Genetic Code Expansion with Two Codon Reassignments. ACS Synthetic Biology 7, 2170-2176 (2018). https: / / doi.org:10.1021 / acssynbio.8b00157, Polycarpo, C. et al. An aminoacyl-tRNA synthetase that specifically activates pyrrolysine. Proc. Natl. Acad. Sci. USA 101, 12450-12454 (2004). (https: / / doi.org:10.1073 / pnas.0405362101)

[0004] Most studies on genetic coding extension using the pyrrolidil system have focused on MmPylRS / MmtRNA. Pyl CUA Paired and closely related MbPylRS / MbtRNAs Pyl CUAThe focus has been on pairs. (Chin, JW Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu. Rev. Biochem. 83, 379--408 (2014). https: / / doi.org:10.1146 / annurev-biochem-060713-035737) These paired PylRS enzymes consist of two domains: an amino (N)-terminal domain and a carboxy (C)-terminal domain. The C-terminal domain binds to an amino acid substrate and catalyzes the aminoacylation of the congeneral tRNAPyl, while the N-terminal domain binds to the tRNA PylIt contacts the variable loop and T-loop to improve binding affinity and specificity. (Suzuki, T. et al. Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 13, 1261 (2017). https: / / doi.org:10.1038 / nchembio.2497, Nozawa, K. et al. Pyrrolysyl-tRNA synthetase-tRNAPyl structure reveals the molecular basis of orthogonality. Nature 457, 1163 (2008). https: / / doi.org:10.1038 / nature07611) It was widely believed that both domains are necessary to produce a functional MmPylRS / MmtRNAPyl pair in Escherichia coli (E. coli), and that all PylRS systems require both domains for activity.(Herring, S. et al. The amino-terminal domain of pyrrolysyl-tRNA synthetase is dispensable in vitro but required for in vivo activity. FEBS Lett. 581, 3197--3203 (2007). https: / / doi.org:10.1016 / j.febslet.2007.06.004, Jiang, R. & Krzycki, JA PylSn and the homologous N-terminal domain of pyrrolysyl-tRNA synthetase bind the tRNA that is essential for the genetic encoding of pyrrolysine. J. Biol. Chem., jbc.M112.396754 (2012). https: / / doi.org:10.1074 / jbc.M112.396754) The inventors have found that ΔN lacks the N-terminal domain (in the same polypeptide or in trans). We demonstrated that a new group of PylRS enzymes, known as PylRS (Borrel, G. et al. Unique Characteristics of the Pyrrolysine System in the 7th Order of Methanogens: Implications for the Evolution of a Genetic Code Expansion Cassette. Archaea 2014, 11 (2014). https: / / doi.org:10.1155 / 2014 / 374146), is both active and orthogonal.(Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5) By combining these pairs and their modified derivatives with pairs derived from the standard +N group, it became possible to create mutually orthogonal Pyl systems. The inventors have identified ΔN group PylRS and tRNA. PylThe inventors further demonstrated that the sequences are clustered into two classes, A and B, based on their sequence identity, and created a triple orthogonal pair consisting of a pair derived from the +N group, a class A pair, and a class B pair. (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x) The discovery of a new mutually orthogonal Pyl system, combined with strategies for providing codons to encode non-canonical monomers, makes it possible to incorporate several different non-canonical amino acids into proteins, as well as to synthesize non-canonical polymers and macrocyclic molecules encoded in cells. (Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057-1062 (2021). https: / / doi.org:10.1126 / science.abg3029, Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5, Dunkelmann, DL, Oehm, SB, Beattie, AT & Chin, JW.A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA design. Nature Chemistry 13, 1110-1117 (2021). https: / / doi.org:10.1038 / s41557-021-00764-5、Beranek, V., Willis, J. C. W. & Chin, J. W. An Evolved Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase / tRNA Pair Is Highly Active and Orthogonal in Mammalian Cells. Biochemistry 58, 387-390 (2019). https: / / doi.org:10.1021 / acs.biochem.8b00808、Meineke, B., Heimgartner, J., Eirich, J., Landreh, M. & Elsasser, S. J. Site-Specific Incorporation of Two ncAAs for Two-Color Bioorthogonal Labeling and Crosslinking of Proteins on Live Mammalian Cells. Cell Reports 31, 107811 (2020). https: / / doi.org:10.1016 / j.celrep.2020.107811、Meineke, B., Heimgartner, J., Lafranchi, L. & Elsasser, S. J. Methanomethylophilus alvus Mx1201 Provides Basis for Mutual Orthogonal Pyrrolysyl tRNA / Aminoacyl-tRNA Synthetase Pairs in Mammalian Cells. ACS Chemical Biology 13, 3087-3096 (2018). https: / / doi.org:10.1021 / acschembio.8b00571, Zhang, H. et al. The tRNA discriminator base defines the mutual orthogonality of two distinct pyrrolysyl-tRNA synthetase / tRNAPyl pairs in the same organism. Nucleic Acids Research 50, 4601-4615 (2022). https: / / doi.org:10.1093 / nar / gkac271, Fischer, JT, Soll, D. & Tharp, JM Directed Evolution of Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase Generates a Hyperactive and Highly Selective Variant. Front. Mol. Biosci. 0 (2022). https: / / doi.org:10.3389 / fmolb.2022.850613, Tharp, J.M., Vargas-Rodriguez, O., (Schepartz, A. & Soll, D. Genetic Encoding of Three Distinct Noncanonical Amino Acids Using Reprogrammed Initiator and Nonsense Codons. ACS Chemical Biology 16, 766-774 (2021). https: / / doi.org:10.1021 / acschembio.1c00120) Despite these advances, there are no effective criteria for searching genomic data for mutually orthogonal Pyl systems, and the inventors hypothesized that many orthogonal and mutually orthogonal systems have yet to be discovered. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Chin, JW Expanding and reprogramming the genetic code. Nature 550, 53 (2017). https: / / doi.org:10.1038 / nature24031 [Non-Patent Document 2] De La Torre, D. & Chin, JW Reprogramming the genetic code. Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7 [Non-Patent Document 3] Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057-1062 (2021). https: / / doi.org:10.1126 / science.abg3029 [Non-Patent Document 4] Spinck, M. et al. Genetically programmed cell-based synthesis of non-natural peptide and depsipeptide macrocycles. Nature Chemistry [Non-Patent Document 5] Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol., 1-11 (2020). https: / / doi.org:10.1038 / s41587-020-0479-2 [Non-Patent Document 6] Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x [Non-Patent Document 7] Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5 [Non-Patent Document 8] Srinivasan, G., James, CM & Krzycki, JA Pyrrolysine Encoded by UAG in Archaea: Charging of a UAG-Decoding Specialized tRNA. Science 296, 1459--1462 (2002). https: / / doi.org:10.1126 / science.1069588 [Non-Patent Document 9] Krzycki, JA The direct genetic encoding of pyrrolysine. Curr. Opin. Microbiol. 8, 706--712 (2005). https: / / doi.org:10.1016 / j.mib.2005.10.009 [Non-Patent Document 10] Neumann, H., Peak-Chew, SY & Chin, JW Genetically encoding Nε-acetyllysine in recombinant proteins. Nat. Chem. Biol. 4, 232 (2008). https: / / doi.org:10.1038 / nchembio.73 [Non-Patent Document 11] Wang, L., Brock, A., Herberich, B. & Schultz, PG Expanding the Genetic Code of Escherichia coli. Science 292, 498--500 (2001). https: / / doi.org:10.1126 / science.1060077 [Non-Patent Document 12] Borrel, G. et al. Unique Characteristics of the Pyrrolysine System in the 7th Order of Methanogens: Implications for the Evolution of a Genetic Code Expansion Cassette. Archaea 2014, 11 (2014). https: / / doi.org:10.1155 / 2014 / 374146 [Non-Patent Document 13] Park, H.-S. et al. Expanding the Genetic Code of Escherichia coli with Phosphoserine. Science 333, 1151--1154 (2011). https: / / doi.org:10.1126 / science.1207203 [Non-Patent Document 14] Rogerson, D. T. et al. Efficient genetic encoding of phosphoserine and its nonhydrolyzable analog. Nat. Chem. Biol. 11, 496 (2015). https: / / doi.org:10.1038 / nchembio.1823

Non-Patent Document 15

Non-Patent Document 16

Non-Patent Document 17

Non-Patent Document 18

Direct Account 28

Direct Environment 29

Draw 30 pages

Non-Patent Document 39

Non-Patent Document 40

Non-Patent Document 41

Non-Patent Document 42

Direct Environment 43

Direct Environment 44

Direct Environment 45

[0006] In one embodiment, cells are provided that contain an exogenous class C acyl-tRNA synthetase (aRS) and include one, two, three, or four members from the group consisting of exogenous class A aRS, exogenous class B aRS, exogenous class N aRS, and exogenous class S aRS, wherein each aRS is a pyrrolidyl-tRNA synthetase (PylRS) or a variant of PylRS modified to alter its acylation specificity.

[0007] In one embodiment, a cell is provided which contains an exogenous class S aRS and one, two, three, or four members from the group of exogenous class A aRS, exogenous class B aRS, exogenous class C aRS, and exogenous class N aRS, wherein each aRS is either a PylRS or a variant modified to alter the acylation specificity of a PylRS.

[0008] In one embodiment, a method for producing cells containing at least two exogenous acyl-tRNA synthetases, i) Screening one or more PylRS to identify a first PylRS belonging to class C or S, ii) Screening one or more PylRS to identify a second PylRS belonging to class A, B, C, N, or S, Optionally, iii) modifying the first PylRS and / or the second PylRS to change the acylation specificity, iv) To generate cells expressing a first PylRS and a second PylRS, wherein the first PylRS and the second PylRS are not of the same class. A method is provided that includes this.

[0009] In one embodiment, cells obtained or obtainable by any of the methods disclosed herein are provided.

[0010] In one embodiment, cells are provided that include nucleic acid sequences encoding an exogenous protein having an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 29 or SEQ ID NO: 30, and a protein having an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 62.

[0011] In one embodiment, cells are provided that include nucleic acid sequences encoding an exogenous protein having an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 31 or SEQ ID NO: 32, and a protein having an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 63.

[0012] In one embodiment, the use of any of the cells disclosed herein is provided for the production of polymers comprising at least one non-natural amino acid or a monomer that is not an alpha-amino acid.

[0013] In one embodiment, a method for producing a polymer containing at least one non-natural amino acid or a monomer that is not an alpha-amino acid, To culture any of the cells disclosed herein, To supply cells with genes that code for polymers, and Obtaining a polymer A method is provided that includes this. [Brief explanation of the drawing]

[0014] [Figure 1-1] ~ [Figure 1-4]This figure shows the selection of candidate PylRS and tRNAPylCUA sequences, and the division of the Pyl system into five different sequence-defining classes. 'a' plots the activity of each combination of ΔN PylRSi and ΔN tRNAPylj against sequence identity between ΔN PylRSi and ΔN PylRSj, measured by the production of GFP(150TAG)His6 from cells carrying the GFP(150TAG)His6 gene in the presence of 4 mM AllocK 1, where ΔN PylRSj is a synthetase of the same biological origin as ΔN tRNAPylj. ΔN PylRS proteins with sequence identity higher than 55% (gray dashed line) primarily exhibit activity with each other's pyl tRNAs (88% case). ΔN PylRS proteins with sequence identity lower than 55% may or may not exhibit activity with each other's pyl tRNAs. b plots the activity of each combination of ΔN PylRSi and ΔN tRNAPylj against the sequence identity between ΔN tRNAPyli and ΔN tRNAPylj, where ΔN tRNAPyli is a tRNAPyl from the same organism as ΔN PylRSi. ΔN pyl tRNAs with sequence identity higher than 75% (gray dashed line) primarily show activity with each other's synthetases (93% of cases). ΔN pyl tRNAs with sequence identity lower than 75% may or may not show activity with each other's synthetases. c is a clustergram of the amino acid sequences of the C-terminal domains of 351 PylRS proteins read from a BLAST search for the AΔ-AlvPylRS protein sequence, with three groups highlighted in red (+N), blue (ΔN), and green (sN) on the dendrogram. Using a clustering threshold of 55%, 37 clusters were obtained. This data was generated using the scikit-learn package (version 1.0.1) in the Python programming language (version 3.9.7). Sequence identity scores were converted to a Euclidean distance index, and unweighted average concatenation clustering was performed. The heatmap represents the sequence identity percentage scores.d is a dendrogram showing 37 clusters generated from aggregated hierarchical clustering of amino acid sequences of the C-terminal domains of 351 PylRS. 37 PylRS sequences selected as cluster representatives are indicated. The radial coordinates represent sequence identity percentage (logarithmic scale), and the gray contour lines correspond to 20% intervals. The red contour lines represent the clustering threshold of 55% sequence identity. e is a clustergram of 35 identified tRNAPyl sequences from the same organism as the representative PylRS in each cluster, with five Pyl system classes highlighted in red (N), purple (A), light blue (B), dark blue (C), and green (S) on the dendrogram. f is a dendrogram showing 8 clusters generated from aggregated hierarchical clustering of 35 identified tRNAPyl sequences. The colored markings correspond to 16 tRNAPyl sequences selected for experimental characterization along with their corresponding PylRS enzymes. The radial coordinates represent sequence identity percentages (logarithmic scale), and the gray contour lines correspond to 20% intervals. The red contour lines represent the clustering threshold of 75% sequence identity. g shows representative tRNAPyl for each class. Significant structural differences from standard N-MmtRNAPyl are highlighted with blue nucleotides. h is a schematic diagram of the three Pyl system groups and their classification into five classes. The names of the Pyl systems selected for characterization are noted below each class. For classes A and B, the pairs AΔ-1R26PylRS / A-AlvtRNAPyl and BΔ-Lum1PylRS / B-InttRNAPyl were used. For all other classes, the PylRS / tRNAPyl pairs were induced from the same organism. i is a schematic diagram of the interaction network between all five classes. Classes N and S are known to interact (red arrows). Interactions between classes N, A, and B can be neutralized by tRNA modification (gray arrows). All other interclass interactions have not been explored (pink arrow). [Figure 2-1] ~ [Figure 2-3]This figure shows the activity mapping of candidate PylRS enzymes and pyl tRNAs, as well as the discovery of novel triple orthogonal PylRS / tRNAPyl pairs. Figure a shows the structure of N6-((allyloxy)carbonyl)-L-lysine (AllocK)1, the amino acid used in this study. Figure b is a heatmap showing the activity of all selected pyl tRNA and PylRS enzyme combinations, measured by the production of GFP(150AllocK)His6 from cells carrying the GFP(150TAG)His6 gene in the presence of 4 mM AllocK1. The values ​​are shown as a percentage relative to wild-type GFP expression. Only PylRS enzymes and pyl tRNAs with activity higher than 30% in combination with at least one tRNAPyl or PylRS enzyme are shown. It is noted that most active pairs consist of heterogeneous combinations of PylRS enzymes and pyl tRNAs. Each heatmap value represents the average of three biological replicate experiments. c is an activity heatmap of representative sets in each family of double orthogonal PylRS / tRNAPyl pairs obtained from activity screening. The orthogonality coefficient (oc), defined as the quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity, is shown in gray. The set with the highest oc in each family is displayed. Each family shares the same combination of PylRS enzymes but has different combinations of pyl tRNAs. For them to be considered orthogonal to each other, the intrapair activity must be higher than 40% of the wtGFP control, and the interpair cross-reactivity must be lower than 20% of the wtGFP control. Furthermore, the oc of each pair must be at least 2.5. d is an activity heatmap of the set with the highest oc in each family of triple orthogonal PylRS / tRNAPyl pairs obtained from activity screening. e concerns the generation of SΔPylRS variants by deleting the N-terminal domain from class S PylRS enzymes. The inventors considered the SΔPylRS variant to be a modified member of the ΔN group because the diversity of their activity profiles is too high to consider them as separate classes.f is the interaction network between all five PylRS classes, based on the activity between the characterized PylRS enzymes and pyl tRNAs. Using PylRS enzymes of classes A and B, A and S, B and S, C and N, and C and S, mutually orthogonal pairs can be found (gray double arrows). Thus, five of the ten possible mutually orthogonal combinations were discovered. The other five each showed one type of unwanted cross-reactivity (red single arrow). These orthogonal interactions were identified without modifications to adjust the PylRS:tRNAPyl interactions. For any two PylRS classes, not all combinations exhibited bidirectional cross-reactivity. [Figure 3-1] ~ [Figure 3-3]Screening of modified pyl tRNAs allows control over 18 of 20 possible cross-reactivity patterns between specific members of five Pyl classes. Figure a is a schematic diagram of the interactions between five specific PylRS / tRNAPyl pairs (one from each class), representing a logical starting point for developing quintuple orthogonal pairs through the tRNAPyl modification strategy. For classes N, A, and B, the inventors selected active pairs whose interclass cross-reactivity (highlighted in blue arrows) had previously been controlled by tRNAPyl modification. For class C, the inventors selected the CΔ-NitraPylRS / C-Therm1tRNAPyl pair, where tRNAPyl is naturally orthogonal to all other PylRS classes. Finally, for class S, the inventors selected the most active PylRS / tRNAPyl pair within the class. Figure b is an activity heatmap of the PylRS enzyme and pyl tRNA sets selected as the basis for developing quintuple orthogonal pairs. Pyl tRNAs that require modification or replacement to control unwanted cross-reactivity are marked in red, while pyl tRNAs that already meet all required orthogonality requirements are marked in green. Green box: Innate orthogonality of C-Therm1tRNAPyl. Blue box: Class interactions previously orthogonalized by tRNAPyl modification and screening. c is a screening of previously reported class N pyl tRNAs against the major active PylRS enzymes (and SΔPylRS variants) of each class. Several class N pyl tRNAs are highly specific to N+-MmPylRS. Each heatmap value represents the average of three biological replicate experiments. d is an updated activity heatmap of b based on the screening results of class N tRNAPyls. N-MettRNAPyl, the tRNA with the highest orthogonality in the screening of class N tRNAPyls against the selected PylRS enzymes, is paired with N+-MmPylRS. N-MettRNAPyl satisfies all orthogonality requirements (green box).e is a screening of previously reported A-AlvtRNAPyl mutants against the major active PylRS enzymes (and SΔPylRS variants) of each class. Several mutants are highly specific to AΔ-1R26PylRS. Each heatmap value represents the average of three biological replicate experiments. f is an updated activity heatmap from d based on the A-AlvtRNAPyl screening results. A-AlvtRNAPyl-21, the tRNAPyl with the highest orthogonality in the screening of class A tRNAPyls, is paired with AΔ-1R26PylRS. A-AlvtRNAPyl-21 satisfies all orthogonality requirements (green box). g is a screening of previously reported B-InttRNAPyl mutants against the major active PylRS enzymes (and SΔPylRS variants) of each class. The activity of BΔ-Lum1PylRS with class B pyl tRNAs is well-matched to SΔ-ClosPylRS, SΔ-I2PylRS, and, most remarkably, CΔ-NitraPylRS. Each heatmap value represents the average of three biological replicate experiments. h is an updated activity heatmap of f based on the screening results for class B tRNAPyls. B-InttRNAPyl-17C10 was the most orthogonal tRNAPyl in the screening of class B tRNAPyls, but it does not meet all the necessary orthogonality requirements, solely due to cross-reactivity with CΔ-NitraPylRS (red box). i is a schematic diagram summarizing the main results of the N, A, and B tRNAPyl screening. By screening native and modified pyl tRNAs, we were able to control 18 of the 20 interactions that needed to be orthogonalized to generate quintuple orthogonal PylRS / tRNAPyl pairs (left figure, blue highlighted arrows indicate interactions that were successfully controlled). Three completely orthogonal pyl-tRNAs (shown in green along with their corresponding PylRS enzymes) were identified, while for each of the remaining two pyl-tRNAs (shown in red), only one cross-reaction remained unregulated. [Figure 4-1] ~ [Figure 4-3]The unique activity pattern of the SΔPylRS enzyme enables the development of quadruple orthogonal pairs. a is a schematic representation of a strategy to replace a class B or C PylRS enzyme with an SΔPylRS variant to resolve undesirable cross-reactivity between class C PylRS and class B tRNAPyl. The diverse activity of the SΔPylRS variants means that it may be possible to substitute class B PylRS enzymes with different variants (e.g., BΔ-Lum1PylRS with SΔ-ClosPylRS as indicated) or class C PylRS enzymes (e.g., CΔ-NitraPylRS with SΔ-I2PylRS as indicated). b is an activity heatmap of the set with the highest oc in each family of triple orthogonal PylRS / tRNAPyl pairs obtained from tRNAPyl screening results for N, A, and B. Substituting B or C class PylRS with different SΔPylRS variants (indicated in two different shades of blue) allows for the generation of many new families. Gold frame: Representative sets in the previously reported triple orthogonal N+-MmPylRS, AΔ-1R26PylRS, and BΔ-Lum1PylRS families. Silver frame: Representative sets in the only triple orthogonal family found before tRNAPyl screening for classes N, A, and B. The orthogonality coefficient oc is shown in gray. Each family shares the same combination of PylRS enzymes but has different combinations of pyl tRNAs. For each pair to be considered orthogonal, the activity within each pair must be higher than 40% of the wtGFP control, and the cross-reactivity between each pair must be lower than 20% of the wtGFP control. Furthermore, the oc of each pair must be at least 2.5. c is the activity heatmap of the set with the highest oc in each family of quadruple orthogonal PylRS / tRNAPyl pairs obtained from tRNAPyl screening results for classes N, A, and B. By substituting PylRS of class B or C with different SΔPylRS variants (indicated in two different shades of blue), all generation of such families becomes possible.d is an activity heatmap of two quadruplet families in which a single SΔPylRS variant replaces a class B or class C PylRS, shown together with the fifth pair with the highest orthogonality in the final class (class N). Pyl tRNAs that require modification or replacement to neutralize unwanted cross-reactivity are marked in red, while pyl tRNAs that already meet all the necessary orthogonality requirements are marked in green. e is a schematic diagram of the interaction between the quadruple orthogonality set with the highest oc (PylRS of class B is replaced by SΔB, oc3.9) and the fifth pair with the highest orthogonality in the final class (class N). Cross-reactivity between class C PylRS and class B tRNAPyl is neutralized by replacing the class B pair with the pair formed with the SΔPylRS variant (green arrow). However, two other types of cross-reactivity must be neutralized to obtain a quintuple orthogonality pair (red arrow). f is a schematic diagram of the interaction between the lowest quadruple orthogonality set (PylRS of class C is substituted with SΔC, oc 2.5) and the fifth pair with the highest orthogonality in the final class (class N). Cross-reactivity between PylRS of class C and tRNAPyl of class B is neutralized by substituting the class C pair with the pair formed with the SΔPylRS variant (green arrow). However, two other types of cross-reactivity must be neutralized to obtain a quintuple orthogonality pair (red arrow). g is an activity heatmap of two (overlapping) quadruplet families in which both PylRS enzymes of class B and class C are substituted with SΔPylRS variants (marked in two different shades of blue), shown together with the fifth pair with the highest orthogonality in the final class (class N or class S). tRNAPyl that require modification or substitution to neutralize unwanted cross-reactivity are marked in red, and pyl tRNA that already meets all the necessary orthogonality requirements are marked in green. h is a schematic diagram of the interaction in the quadruple preset of g (all oc2.9).In all cases, cross-reactivity between class N and class S must be eliminated to obtain quintuple orthogonal pairs (red arrow). Furthermore, the difficult-to-resolve cross-reactivity between class C PylRS and class B tRNAPyl is mitigated by substituting both class B and C pairs with pairs formed with the SΔPylRS variant, but the remaining cross-reactivity limits the oc of these sets (yellow arrow). [Figure 5-1] ~ [Figure 5-5]This diagram illustrates the quintuple orthogonal PylRS / tRNAPyl pairs via directed evolution. a is an activity heatmap of the quadruple pret with the highest oc values, combined with the fifth pair with the highest orthogonality in the final class. Cross-reactivity is indicated by a red box. Pyl tRNAs to be replaced by the evolved quintuple orthogonal variants are marked in red. b is a schematic diagram showing the library used to evolve quintuple orthogonal pyl tRNAs from the S-I2tRNAPyl scaffold. The cloverleaf structure of S-I2tRNAPyl is shown, and randomized nucleotides are represented by blue circles. c is a heatmap showing the activity of hits obtained by sequential negative screening using PylRS enzymes of classes N, A, C, and S after positive selection of the library using SΔB-ClosPylRS. Each heatmap value represents the average of three biological replicate experiments. d is the activity heatmap from a, updated based on the results of directed evolution from c. S-I2tRNAPyl-B32 is the most orthogonal tRNA obtained from directed evolution and satisfies all the necessary orthogonality requirements (green box). As a result, the quintuple orthogonal SΔB-ClosPylRS / S-I2tRNAPyl-B32 pair effectively substitutes class B and requires no further modification. e is a heatmap showing the activity of S-I2tRNAPyl-S52, a tRNAPyl hit obtained by sequential negative screening with other classes of PylRS enzymes after positive selection of the library with S+-DebPylRS. This tRNAPyl is active only with S+-DebPylRS. Each heatmap value represents the average of three biological replicate experiments. f is the activity heatmap of d updated based on the results of directed evolution of e. S-I2tRNAPyl-S52 satisfies all the necessary orthogonality requirements (green box). As a result, the S+-DebPylRS / S-I2tRNAPyl-S52 pair completes the quintuple orthogonality set of the pair without requiring further modification (oc4.0). g is a schematic diagram of the overall tRNAPyl evolutionary strategy and the resulting pair.By sequentially disrupting cross-reactivity between class N and class B (or its SΔ equivalent), and between class N and class S, a set of five pairs is obtained in which all 20 possible cross-reactivity is minimized. h is an activity heatmap of each family of quadruple orthogonal PylRS / tRNAPyl pairs obtained under the tRNAPyl-oriented evolutionary strategy, showing the quadruple pret with the highest oc. Dark gray box: Quadruple orthogonal family discovered in Figure 4. Light gray box: Quadruple orthogonal family discovered in Figure 4, but with higher oc when a newly evolved tRNAPyl is incorporated. The orthogonality coefficient oc is shown in gray. i is an activity heatmap of two families of quintuple orthogonal pairs incorporating evolved pyl tRNA, showing the quintoplet with the highest oc. j is a schematic diagram of the interactions in the quintuple orthogonal set of pairs showing the highest oc (5.4), formed using one PylRS of each class. k is a schematic diagram illustrating the success of classifying pyrrolidine systems into five mutually orthogonal functional classes: N, A, B or SΔB, C or SΔC, and S. [Figure 6-1] ~ [Figure 6-2] This is the alignment of the C-terminal domain of PylRS. The sequences in Figure 6 are sequence numbers 2, 4, 6, 9, 12, 15, 18, 21, 24, 27, 29, 31, and 33. This figure shows the alignment of the C-terminal domain performed in mean concatenation clustering. [Modes for carrying out the invention]

[0015] The present inventors provide herein a method for classifying pyrrolidine tRNA synthetases (PylRS) into one of five categories: class A, class B, class C, class N, and class S. Classes A, B, and N are based on our previous research (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW: Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x and Willis, JCW & Chin, JW: Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018)). The present inventors have found that these classes of PylRS have a remarkable degree of natural orthogonality with respect to one another. Therefore, by selecting one PylRS from each of these classes, it is possible to use them to form a system of mutual orthogonality. PylRS may be modified or altered to alter their acylation specificity. For example, PylRS may be modified or altered so that the resulting acyl-tRNA synthetase (aRS) can charge the tRNA with an alternative standard amino acid or a non-standard monomer. For example, the non-standard monomer may be a non-natural alpha-amino acid, or a monomer that is not an alpha-amino acid (referred to herein as “non-alpha-amino acid”). This monomer may be, for example, a hydroxy acid or a beta-amino acid.

[0016] The inventors demonstrate that pairing five PylRS classes with appropriate tRNAs enables the formation of quintuple orthogonal PylRS / tRNA pairs in host cells.

[0017] Accordingly, in one embodiment, there is a cell comprising an exogenous class C aRS and one, two, three, or four members from the group of exogenous class A aRS, exogenous class B aRS, exogenous class N aRS, and exogenous class S aRS, wherein each aRS is a PylRS, or a variant modified to alter the acylation specificity of PylRS.

[0018] In certain embodiments, cells are provided that contain an exogenous class C aRS and one, two, or three members of the group consisting of exogenous class A aRS, exogenous class B aRS, and exogenous class N aRS, wherein each aRS is a PylRS, or a variant modified to alter the acylation specificity of a PylRS.

[0019] In another embodiment, there is a cell comprising an exogenous class S aRS and one, two, three, or four members from the group of exogenous class A aRS, exogenous class B aRS, exogenous class C aRS, and exogenous class N aRS, wherein each aRS is a PylRS or a variant modified to alter the acylation specificity of PylRS.

[0020] As used herein, “exogenous” aRS refers to aRS present in cells that do not naturally express the aforementioned aRS. For example, exogenous aRS may be encoded in an episomal replicon such as a plasmid. Alternatively, a sequence encoding exogenous aRS may be inserted into the genome of a host cell.

[0021] In one embodiment, PylRS of class A, B, C, N, or S may be defined by comparing the amino acid sequence of the C-terminal domain of a candidate PylRS with the amino acid sequence of a representative PylRS enzyme known to belong to all five classes. The sequences can be compared using unweighted mean concatenation clustering. For example, using python (version 3.9.7), an identity percentage matrix can be calculated from the multiple sequence alignment of the C-terminal domain for all pairs of PylRS sequences in the database. The database may be Table 1 (Example 11) herein. This matrix can then be used to perform unweighted mean concatenation clustering (UPGMA) of aligned PylRS C-terminal domain sequences, using the python libraries biopython (version 1.79) and scikit-learn (version 1.0.1), with a clustering threshold of 55% sequence identity.

[0022] Therefore, in one embodiment, a class A, class B, or class C aRS is a PylRS that does not have an N-terminal domain, or a PylRS in which the N-terminal domain is not expressed in the host cell, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 and the aligned sequences of the C-terminal domains, with a lower threshold of sequence identity set to 55%, forms a cluster with the class A, class B, or class C (each) PylRS in Table 1. This definition is referred to herein as the “sequence-dependent” definition.

[0023] In certain embodiments, aRS of class A, class B, or class C is a PylRS that does not have an N-terminal domain, or a PylRS in which the N-terminal domain is not expressed in the host cell, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of the C-terminal domains, with a lower threshold of sequence identity set to 55%, forms a cluster with PylRS of class A, class B, or class C (each) in Figure 6, or a PylRS referred to as the cluster representative of each class in Table 1.

[0024] Class A aRS may cluster with sequence number 35 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of those C-terminal domains, with a lower threshold of 55% for sequence identity. Class B aRS may cluster with one of sequence numbers 36, 375, 392, 399, or 400 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of those C-terminal domains, with a lower threshold of 55% for sequence identity. Class C aRS may cluster with one of sequence numbers 8, 11, 14, 17, 20, or 23 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of those C-terminal domains, with a lower threshold of 55% for sequence identity. Class N aRS may cluster with sequence number 331 after mean concatenation clustering of the aligned C-terminal domain sequences of aRS sequences in Table 1 with the aligned sequences of those C-terminal domains, with a lower threshold of sequence identity of 55%. Class S aRS may cluster with one of sequence numbers 26, 29, 31, 33, 149, 174, 187, 198, 200, 272, 277, 278, 285, 288, 291, 294, 376, 377, 381, 382, ​​384, 388, 393, 394, or 397 after mean concatenation clustering of the aligned C-terminal domain sequences of aRS sequences in Table 1 with the lower threshold of sequence identity of 55%.

[0025] Similarly, in one embodiment, a class N aRS is a PylRS that forms a cluster with a class N PylRS in Table 1 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 and the aligned sequences of those C-terminal domains, with a lower threshold of sequence identity set to 55%. In a particular embodiment, a class N aRS is a PylRS that forms a cluster with a class N PylRS in Figure 6, or a PylRS referred to as the cluster representative of class N in Table 1, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 and the aligned sequences of those C-terminal domains, with a lower threshold of sequence identity set to 55%.

[0026] Furthermore, in one embodiment, the aRS of class S is a PylRS that forms a cluster with the PylRS of class S in Table 1 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 and the aligned sequences of those C-terminal domains, with a lower threshold of sequence identity set to 55%. In a particular embodiment, the aRS of class S is a PylRS that forms a cluster with the PylRS of class S in Figure 6, or the PylRS referred to as the cluster representative of class S in Table 1, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 and the aligned sequences of those C-terminal domains, with a lower threshold of sequence identity set to 55%.

[0027] The C-terminal domain sequence of PylRS is defined as starting from the amino acid that aligns with S186 of MmPylRS (SEQ ID NO: 1) in multiplex sequence alignment. Multiplex sequence alignment can be performed using Clustal Omega with default parameters (Dealign Input Sequences: No, MBED-Like Clustering Guide-Tree: Yes, MBED-Like Clustering Iteration: Yes, Number of Combined Iterations: default (0), Max Guide Tree Iterations: default, Distance Matrix: No, Guide Tree: yes, Order: aligned). The identity percentage between two sequences is defined as the quotient of the number of identical positions and the total number of aligned positions, expressed as a percentage.

[0028] Classes A, B, and C aRS all lack an N-terminal domain; therefore, host cells neither express the associated N-terminal domain as part of the protein containing the C-terminal domain of the aRS, nor separately. Classes A, B, and C aRS may originate from archaeal species.

[0029] Alternatively, aRS of class A, B, or C may be derived from bacterial species in which the naturally occurring N-terminal domain is not expressed by the host cell. In such circumstances, aRS is defined according to the functionality-dependent definition rather than the sequence-dependent definition disclosed herein. The functionality-dependent definition is further described below.

[0030] Class A, B, or C aRS may have a sequence identical to the wild-type sequence of naturally occurring PylRS. Alternatively, class A, B, or C aRS may be a modified variant of naturally occurring PylRS. For example, the active site of PylRS may be modified to recognize a non-pyrrolidine substrate.

[0031] In addition to, or as an alternative to, sequence-dependent and / or functionality-dependent definitions, a class N aRS may be defined as a PylRS derived from an archaeal species, wherein the N-terminal domain of PylRS is included as part of the same polypeptide as the C-terminal domain of PylRS. A class N aRS may have a sequence identical to the wild-type sequence of a naturally occurring PylRS. Alternatively, this aRS may be a modified variant of a naturally occurring PylRS. For example, the active site of PylRS may be modified to recognize a substrate that is not pyrrolidine. This definition is referred to herein as the “origin-dependent” definition.

[0032] In addition to, or as an alternative to, sequence-dependent and / or functionality-dependent definitions, a class S aRS may be defined as a PylRS derived from a bacterial species, wherein the N-terminal domain of the PylRS is separately encoded. This N-terminal domain is also expressed as an exogenous protein in host cells. A class S aRS may have a sequence identical to the wild-type sequence of a naturally occurring PylRS. Alternatively, this aRS may be a modified variant of a naturally occurring PylRS. For example, the active site of the PylRS may be modified to recognize a substrate that is not pyrrolidine. This definition is referred to herein as the “origin-dependent” definition.

[0033] In certain embodiments, aRS of class A, class B, or class C is a PylRS derived from an archaeal species, lacking an N-terminal domain, which forms a cluster with the PylRS of class A, class B, or class C (each) in Figure 6 after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of the C-terminal domains, with a lower threshold of sequence identity of 55%; aRS of class S is a PylRS derived from a bacterial species, in which the N-terminal domain of the PylRS is separately encoded and expressed as an exogenous protein in host cells; and aRS of class N is a PylRS derived from an archaeal species, in which the N-terminal domain of the PylRS is included as part of the same polypeptide as the C-terminal domain of the PylRS.

[0034] In another embodiment, aRS of class A, B, C, N, or S may be identified by functional comparison. Such comparison tests a candidate aRS to ensure that it is capable of charging the same tRNA as a particular class of PylRS enzyme, and also tests that it is not capable of charging the same tRNA as any of the other four classes of PylRS enzyme.

[0035] A suitable assay involves modifying test cells to express candidate aRS and test tRNA. These cells also contain a marker gene that can only be decoded when the test tRNA is charged. For example, this marker gene may contain a stop codon specific to the test tRNA, and this stop codon may be positioned within the gene such that termination of translation at the stop codon triggers marker expression. The level of marker expression can be compared to the level of a control gene. The control gene is identical to the marker gene except that it can be decoded by endogenous tRNA in the test cells.

[0036] An example of such an assay is described in the Examples section and the cited references. In this assay, the amber stop codon is located within the green fluorescent protein (GFP) gene. Termination at the amber stop codon does not result in the expression of the fluorescent protein, but skipping the amber stop codon results in the expression of functional GFP (marker). The test tRNA contains an anticodon that can recognize the amber stop codon. If the candidate aRS can charge the test tRNA with the appropriate amino acids, the stop codon is read as a sense codon, and functional GFP is expressed. Candidate aRS enzymes that cannot charge the test tRNA do not result in GFP expression. The control gene is the wild-type GFP gene, which is identical to the marker gene except that it has a wild-type codon instead of an experimental stop codon.

[0037] Candidate aRSs can be functionally classified and tested using the assays described above. The DNA sequences encoding the test tRNAs used in such assays are listed below. The test tRNA used in the Class A functionality test is: GGGGGACGGTCCGGCGACCAGCGGGTCTCTAAAACCTAGCataagCGGGGTTCGACcCCCCGGTCTCTCGCCA SEQ ID NO: 37 It is coded by. The test tRNA used in the Class B functionality test is: GGGGTGTTGATCGGATTGATCGCGTGGACTCTAAATCCGCGGTAGACGGGTGAAACTCCCGTACACCTCACCA Sequence No. 38 It is coded by. The test tRNA used in the functional testing of Class C is: GGGGGGCTGGTCGGGTGGCCAAGGGGGCTCTAAACCCTCGGTTGCCGGTTCAACTCCCGGGCTCCCCACCA SEQ ID NO: 39 It is coded by. The test tRNA used in the functional testing of Class N is: GGAGACTTGATCATGTAGATCGAACGGACTCTAAATCCGTTCAGCCGGGTTAGATTCCCGGAGTTTCCGCCA Sequence No. 40 It is coded by. The test tRNA used in the functional testing of class S is: GGGGCGTTGATCGGATTGATCGCGTGGACTCTAAATCCGCGGCCGACGGGTGAAACTCCCGTACACCTCTCCA Sequence ID 41 It is coded by.

[0038] In such embodiments, an aRS is defined as a class A aRS if, when paired with the tRNA encoded by SEQ ID NO: 37, it induces the expression of a marker gene at a level higher than or equal to 40% of the control gene's expression, and when individually paired with each of the tRNAs encoded by SEQ ID NO: 38, 39, 40, and 41, it induces the expression of a marker gene at a level lower than 20% of the control gene's expression. Thus, in such embodiments, a class A aRS has activity above the threshold with respect to the tRNA that defines class A, but activity below the threshold with respect to all four tRNAs that define each of the other classes. This is referred to herein as the "functionality-dependent" definition.

[0039] Classes B, C, N, and S can be functionally classified in a similar manner, with aRS of class B having above-threshold activity with the tRNA encoded by SEQ ID NO: 38, aRS of class C having above-threshold activity with the tRNA encoded by SEQ ID NO: 39, aRS of class N having above-threshold activity with the tRNA encoded by SEQ ID NO: 40, aRS of class S having above-threshold activity with the tRNA encoded by SEQ ID NO: 41, and each aRS having below-threshold activity with the tRNA that defines each of the other classes.

[0040] In some embodiments, the cell includes an aRS defined according to the sequence-dependent definition disclosed herein. In some embodiments, the cell includes an aRS defined according to the sequence-dependent definition and / or the origin-dependent definition. In some embodiments, the cell includes an aRS defined according to the functionality-dependent definition disclosed herein. In other embodiments, the cell may include at least one aRS defined according to the sequence-dependent definition and at least one aRS defined according to the functionality-dependent definition disclosed herein. In further embodiments, the cell includes at least one aRS defined according to the sequence-dependent definition and / or at least one aRS defined according to the functionality-dependent definition disclosed herein, as well as aRS of class N and / or class S defined according to the origin-dependent definition. In yet another embodiment, the cell includes an aRS that satisfies both the sequence-dependent definition and the functionality-dependent definition for each class.

[0041] The cells may include aRS of class A and class C derived from archaeal species that lack an N-terminal domain, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of the C-terminal domains, with a lower threshold of sequence identity set at 55%, form clusters with class A or class C (respectively) PylRS in Figure 6; class B aRS that satisfy the functional definition; and class N aRS derived from archaeal species that include the N-terminal domain of PylRS as part of the same polypeptide as the C-terminal domain of PylRS.

[0042] The cells may include aRS of class A and class B derived from archaeal species that lack an N-terminal domain, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of the aRS sequences in Table 1 with the aligned sequences of the C-terminal domains, with a lower threshold of sequence identity set at 55%, form clusters with class A or class B (respectively) PylRS in Figure 6; class C aRS that satisfy the functional definition; and class N aRS derived from archaeal species that include the N-terminal domain of PylRS as part of the same polypeptide as the C-terminal domain of PylRS.

[0043] A cell may contain a class C aRS satisfying the sequence-dependent definition or the functionality-dependent definition, optionally a class A aRS satisfying the sequence-dependent definition, optionally a class B aRS satisfying the sequence-dependent definition, and optionally a class N aRS satisfying the origin-dependent definition.

[0044] The aRS expressed within a cell is an aRS that is active in the host cell. An active aRS is one that can be expressed in the host cell without rendering it unviable and can acylate tRNA. In one example, the cell is an E. coli cell, and each exogenous aRS expressed by the E. coli cell is compatible with the E. coli cell and is active in the E. coli cell.

[0045] An example of a class N PylRS derived from Metanosarkina mazei (Mm) is provided below. MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSVARA PKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKD LQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRARSESYYNGISTNL(Sequence ID 1)

[0046] The C-terminal domain of the above sequence is as follows: SAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRARSESYYNGISTNL(Sequence ID 2)

[0047] An example of a modified class N PylRS derived from Mm is provided below. MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSVARA PKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKD LQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPXLXPNXXNYXRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLXFXQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVXGDTLDVMHGDLELSSAXVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRARSESYYNGISTNL(Sequence ID 3)

[0048] An example of a class A PylRS derived from Candidatus metanomethylophilus alvus (Alv) is provided below. MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKIKGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWIDEKRALRPMLAPNLYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLNLVDMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPHYLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN (Sequence ID 4)

[0049] The C-terminal domain of the above sequence is as described above.

[0050] An example of a modified Class A PylRS derived from Alv is provided below. MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKIKGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWIDEKRALRPXLXPNXXSVXRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLXLXDMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVXKETIDVEINGQEVCSAXVGPHYLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN (Sequence ID 5)

[0051] An example of a class B PylRS derived from Candidatus metanomassillicoccus intestinalis (Int) is provided below. MPVEWTASQKQRLKELGIPAEADRIFNDTKEREEVFKDITSEHLSKVRKDIKHMLDYPERHQLSQIESILAQALVDNGFIEVKTPSIISRSALEKMGIDRSHPLHEQVFWLDEKRCLRPMLAPNLYFMMRHMYRYSKGPLRLFEIGSCFRKESKGSNHLEEFTMLNLVEMAPDNDPADQLLVHIKTIMDALGLEYSLVECESDVYVKTLDVEIDGVEVASGAVGPHKLDPAHGITQSWAGVGFGLERLSMMKYGMDNIKKSGRSLIYLRGVRLDI (Sequence ID 6)

[0052] The C-terminal domain of the above sequence is as described above.

[0053] An example of a PylRS for a modified class B derived from Int is provided below. MPVEWTASQKQRLKELGIPAEADRIFNDTKEREEVFKDITSEHLSKVRKDIKHMLDYPERHQLSQIESILAQALVDNGFIEVKTPSIISRSALEKMGIDRSHPLHEQVFWLDEKRCLRPXLXPNXXFMXRHMYRYSKGPLRLFEIGSCFRKESKGSNHLEEFTMLXLXEMAPDNDPADQLLVHIKTIMDALGLEYSLVECESDVXVKTLDVEIDGVEVASGXVGPHKLDPAHGITQSWAGVGFGLERLSMMKYGMDNIKKSGRSLIYLRGVRLDI (Sequence ID 7)

[0054] An example of class C PylRS from the archaeon Candidatus Bathyarchaeota (Bathy) is provided below. MGNNGLQKLPRSRMENQKILEGKKSSMKDCQFTLSQKRRLKELGADSYVNLTFKNEKERDNAFDNLATILERKHKEALLNLLTLTKRPLIRQLESKLIEALTSAGFVEVNTPFIIPRKFIECMGIRETHKLWKQIHWLRNGRCLRPMLAPNLYHIMRLLRKFTKPVSIFEIGPCFRKESKGREHVEEFTMLNVVELAPNQDPFERLKEIINIVTKTVELPHYRLRNVKSEIYGETMDVIVDEHEIASAVVGPHSLDKNWGIFEAWAGVGFGIERIAMVKKDIKRIRHVARSLTYLDGASLDVQ(Sequence ID 8)

[0055] The C-terminal domain of the above sequence is as follows: KDCQFTLSQKRRLKELGADSYVNLTFKNEKERDNAFDNLATILERKHKEALLNLLTLTKRPLIRQLESKLIEALTSAGFVEVNTPFIIPRKFIECMGIRETHKLWKQIHWLRNGRCLRPMLAPNLYHIMRLLRKFTKPVSIFEIGPCFRKESKGREHVEEFTMLNVVELAPNQDPFERLKEIINIVTKTVELPHYRLRNVKSEIYGETMDVIVDEHEIASAVVGPHSLDKNWGIFEAWAGVGFGIERIAMVKKDIKRIRHVARSLTYLDGASLDVQ(Sequence ID 9)

[0056] An example of a modified Class C PylRS derived from Bathy is provided below. MGNNGLQKLPRSRMENQKILEGKKSSMKDCQFTLSQKRRLKELGADSYVNLTFKNEKERDNAFDNLATILERKHKEALLNLLTLTKRPLIRQLESKLIEALTSAGFVEVNTPFIIPRKFIECMGIRETHKLWKQIHWLRNGRCLRPXLXPNXXHIXRLLRKFTKPVSIFEIGPCFRKESKGREHVEEFTMLXVXELAPNQDPFERLKEIINIVTKTVELPHYRLRNVKSEIXGETMDVIVDEHEIASAXVGPHSLDKNWGIFEAWAGVGFGIERIAMVKKDIKRIRHVARSLTYLDGASLDVQ(Sequence ID 10)

[0057] An example of a class C PylRS from the Nitrososphaeria archaeon (Nitra) is provided below. MSKIRFTRGQIHRLIELGAEPTELERDFETEAERDKEFNKIAENLARKNLKNIKDFLEQRRKPLVRVIEEKLRTTALRLGFSEVVTPIIIPRLFIKRMGIDEGDPLWKQVMLIDDKRALRPMLAPNLYVLMAKLSNIVRPVKIFEIGPCFRRETGGRYHLEEFTMFNMVELAPEGDPKERLLDYIDTIMRDIGLNYTISVEPSNVYGETLDVVVNGIEVASAAIGPKPIDANWGVREPWIGVGFGVERLAMLVGGYNSIARIAKSLSYLDGSTLSVIKLRW (Sequence ID 11)

[0058] The C-terminal domain of the above sequence is as follows: SKIRFTRGQIHRLIELGAEPTELERDFETEAERDKEFNKIAENLARKNLKNIKDFLEQRRKPLVRVIEEKLRTTALRLGFSEVVTPIIIPRLFIKRMGIDEGDPLWKQVMLIDDKRALRPMLAPNLYVLMAKLSNIVRPVKIFEIGPCFRRETGGRYHLEEFTMFNMVELAPEGDPKERLLDYIDTIMRDIGLNYTISVEPSNVYGETLDVVVNGIEVASAAIGPKPIDANWGVREPWIGVGFGVERLAMLVGGYNSIARIAKSLSYLDGSTLSVIKLRW (Sequence ID 12)

[0059] An example of a modified class C PylRS derived from Nitra is provided below. MSKIRFTRGQIHRLIELGAEPTELERDFETEAERDKEFNKIAENLARKNLKNIKDFLEQRRKPLVRVIEEKLRTTALRLGFSEVVTPIIIPRLFIKRMGIDEGDPLWKQVMLIDDKRALRPXLXPNXXVLXAKLSNIVRPVKIFEIGPCFRRETGGRYHLEEFTMFXMXELAPEGDPKERLLDYIDTIMRDIGLNYTISVEPSNVXGETLDVVVNGIEVASAXIGPKPIDANWGVREPWIGVGFGVERLAMLVGGYNSIARIAKSLSYLDGSTLSVIKLRW (Sequence ID 13)

[0060] An example of a class C PylRS from the candidate phylum MSBL1 archaeon SCGC-AAA382A20 (SCGC) is provided below. MNLTSSQKQRLRELGWDGSIPDFDNKKERDQFFNKTATKLKNRNKERFLKLLENKVPSWRRVERKLRNIFYELGFVEVQTPSIISPSLLEKMDIGEESKLYNQIYQIKGEKKSLRPMLAPNLYRELRYFSRISDEEVIRLFELGSCFRKENGGERHLNEFKMLNAVEMGNIKDTKKRLDELISNVFSPFANYKVEKEKSTVYEETVDVNIKNTEVASCVIGPHFLDSNWHIDEPWVGLGIGVERLTRVIEGEPSVKPFGKSYVYQDGIRLDIE(Sequence ID 14)

[0061] The C-terminal domain of the above sequence is as follows: MNLTSSQKQRLRELGWDGSIPDFDNKKERDQFFNKTATKLKNRNKERFLKLLENKVPSWRRVERKLRNIFYELGFVEVQTPSIISPSLLEKMDIGEESKLYNQIYQIKGEKKSLRPMLAPNLYRELRYFSRISDEEVIRLFELGSCFRKENGGERHLNEFKMLNAVEMGNIKDTKKRLDELISNVFSPFANYKVEKEKSTVYEETVDVNIKNTEVASCVI (Sequence ID 15)

[0062] An example of a modified Class C PylRS derived from SCGC is provided below. MNLTSSQKQRLRELGWDGSIPDFDNKKERDQFFNKTATKLKNRNKERFLKLLENKVPSWRRVERKLRNIFYELGFVEVQTPSIISPSLLEKMDIGEESKLYNQIYQIKGEKKSLRPXLXPNXXREXRYFSRISDEEVIRLFELGSCFRKENGGERHLNEFKMLXAXEMGNIKDTKKRLDELISNVFSPFANYKVEKEKSTVXEETVDVNIKNTEVASCXIGPHFLDSNWHIDEPWVGLGIGVERLTRVIEGEPSVKPFGKSYVYQDGIRLDIE(Sequence ID 16)

[0063] An example of a class C PylRS derived from Candidatus Methanohalarchaeum thermophilum 1 (Therm1) is provided below. MELTRSQSQRLRELGYQGEAPTFEDQEERDEFFERKETELQKKNRNKFKKLQRINEPDWKKTEQKLRKNLYESDFTEVQTPHIISMSVLKNKMNISEESNIYNQIYKLDEGNKCLRPMLAPNLYRQMKHFLRISKKDVVKLFELGTCFRKEQGKNHVREFKMLNAVEVGEIKDKEKRTREMIDEIIGNLVDYKIEEEKSTVYGKTLDIEVNGLEIASSVIGPHPLDANFSINKPWIGIGIGVERLIQTKNEGNSIKSYARSLSYQDGIRLEIN(Sequence ID 17)

[0064] The C-terminal domain of the above sequence is as follows: MELTRSQSQRLRELGYQGEAPTFEDQEERDEFFERKETELQKKNRNKFKKLQRINEPDWKKTEQKLRKNLYESDFTEVQTPHIISMSVLKNKMNISEESNIYNQIYKLDEGNKCLRPMLAPNLYRQMKHFLRISKKDVVKLFELGTCFRKEQGKNHVREFKMLNAVEVGEIKDKEKRTREMIDEIIGNLVDYKIEEEKSTVYGKTLDIEVNGLEIASSVI (Sequence ID 18)

[0065] An example of a modified Class C PylRS derived from Therm1 is provided below. MELTRSQSQRLRELGYQGEAPTFEDQEERDEFFERKETELQKKNRNKFKKLQRINEPDWKKTEQKLRKNLYESDFTEVQTPHIISMSVLKNKMNISEESNIYNQIYKLDEGNKCLRPXLXPNXXRQXKHFLRISKKDVVKLFELGTCFRKEQGKNHVREFKMLXAXEVGEIKDKEKRTREMIDEIIGNLVDYKIEEEKSTVXGKTLDIEVNGLEIASSXIGPHPLDANFSINKPWIGIGIGVERLIQTKNEGNSIKSYARSLSYQDGIRLEIN (Sequence ID 19)

[0066] An example of a class C PylRS derived from Candidatus metanohalarcaeum thermopyrum.9096 (Therm2) is provided below. MEFTETQKQRLRELGYKGEFPELDTKEEVNEAYSQLEKKLRKKHRKKLNDLFESKKPTWKNTVENIRQNLQDLGFIEVQTPLIISKNLLKKMKIDQKSDLMNQVYRINDNKVLRPMLAQNLYKELENFSKLSNRDTIQLFEIGTCFRKEKGGKDHLNEFKMLNAVELGNFKDKEKRLKEVISTLFKDFDEYVLEKEKSTVYGETYDVLVNGTELASCAIGPHQLDEKWDINRPWIGIGIGIERFTRELNNSDSTVKAYGRSFVYQDGIRLDIK(Sequence ID 20)

[0067] The C-terminal domain of the above sequence is as follows: MEFTETQKQRLRELGYKGEFPELDTKEEVNEAYSQLEKKLRKKHRKKLNDLFESKKPTWKNTVENIRQNLQDLGFIEVQTPLIISKNLLKKMKIDQKSDLMNQVYRINDNKVLRPMLAQNLYKELENFSKLSNRDTIQLFEIGTCFRKEKGGKDHLNEFKMLNAVELGNFKDKEKRLKEVISTLFKDFDEYVLEKEKSTVYGETYDVLVNGTELASCAI (Sequence ID 21)

[0068] An example of a modified Class C PylRS derived from Therm2 is provided below. MEFTETQKQRLRELGYKGEFPELDTKEEVNEAYSQLEKKLRKKHRKKLNDLFESKKPTWKNTVENIRQNLQDLGFIEVQTPLIISKNLLKKMKIDQKSDLMNQVYRINDNKVLRPXLXQNXXKEXENFSKLSNRDTIQLFEIGTCFRKEKGGKDHLNEFKMLXAXELGNFKDKEKRLKEVISTLFKDFDEYVLEKEKSTVXGETYDVLVNGTELASCXIGPHQLDEKWDINRPWIGIGIGIERFTRELNNSDSTVKAYGRSFVYQDGIRLDIK(Sequence ID 22)

[0069] An example of a class C PylRS from the archaeon Methanonatronarchaeia archaeon (Tron) is provided below. MEFTVTQKQRLQELGFEGVFPSDFEDVDERNRFFEELVGRLRDRNRKRFERLVGNKIPFWRKVSSDLRNRFYELGFVEVRTPEIISYSLLEKMEISDDLREQVYWLEEDNRCLRPMLAPNLYNELRHFNRISNQSKVRIFEIGTCFRREKSSSEHLNEFTMLNAVEMGDIGDTEERLDRLIEEVFGEFTDYKKVGEESSLYGKTVDVLVDGVEVASCIAGPHPLDSNWSIDQPWVGIGLGVERLAMLLDDGSTAKAYGNSYIYQDGVRLDIK (Sequence ID 23)

[0070] The C-terminal domain of the above sequence is as follows: MEFTVTQKQRLQELGFEGVFPSDFEDVDERNRFFEELVGRLRDRNRKRFERLVGNKIPFWRKVSSDLRNRFYELGFVEVRTPEIISYSLLEKMEISDDLREQVYWLEEDNRCLRPMLAPNLYNELRHFNRISNQSKVRIFEIGTCFRREKSSSEHLNEFTMLNAVEMGDIGDTEERLDRLIEEVFGEFTDYKKVGEESSLYGKTVDVLVDGVEVASCIA (Sequence ID 24)

[0071] An example of a modified Class C PylRS derived from Tron is provided below. MEFTVTQKQRLQELGFEGVFPSDFEDVDERNRFFEELVGRLRDRNRKRFERLVGNKIPFWRKVSSDLRNRFYELGFVEVRTPEIISYSLLEKMEISDDLREQVYWLEEDNRCLRPXLXPNXXNEXRHFNRISNQSKVRIFEIGTCFRREKSSSEHLNEFTMLXAXEMGDIGDTEERLDRLIEEVFGEFTDYKKVGEESSLXGKTVDVLVDGVEVASCXAGPHPLDSNWSIDQPWVGIGLGVERLAMLLDDGSTAKAYGNSYIYQDGVRLDIK (Sequence ID 25)

[0072] An example of class S PylRS from Clostridiales bacteria (Clos) is provided below. MENFTITQTERLKQLNCENDVLELEFEDSEARNSKFREIEIGRVKKGKENIKNLLKEKHITISDEVGNKLSDWLMSKDYTKVLTPTIISKDQLKAMTIDEENHLFSQVFWIDNNKCLRPMLAPNLYIVMRELKRITNEPVKIFEIGSCFRKESQGARHMNEFTMLNMVELASVEDGKQLDTLKALAHEAMESLGVESYELVIEESAVYGSTLDIEIDGIEVASGSYGPHELDANWDIFDTWVGIGFGIERLAMAINGGSTIKKYGRSINFIDGETMKL (Sequence ID 26)

[0073] The C-terminal domain of the above sequence is as follows: MENFTITQTERLKQLNCENDVLELEFEDSEARNSKFREIEIGRVKKGKENIKNLLKEKHITISDEVGNKLSDWLMSKDYTKVLTPTIISKDQLKAMTIDEENHLFSQVFWIDNNKCLRPMLAPNLYIVMRELKRITNEPVKIFEIGSCFRKESQGARHMNEFTMLNMVELASVEDGKQLDTLKALAHEAMESLGVESYELVIEESAVYGSTLDIEIDGIEVASGSY (Sequence ID 27)

[0074] An example of a modified PylRS for class S derived from Clos is provided below. MENFTITQTERLKQLNCENDVLELEFEDSEARNSKFREIEIGRVKKGKENIKNLLKEKHITISDEVGNKLSDWLMSKDYTKVLTPTIISKDQLKAMTIDEENHLFSQVFWIDNNKCLRPXLXPNXXIVXRELKRITNEPVKIFEIGSCFRKESQGARHMNEFTMLXMXELASVEDGKQLDTLKALAHEAMESLGVESYELVIEESAVXGSTLDIEIDGIEVASGXYGPHELDANWDIFDTWVGIGFGIERLAMAINGGSTIKKYGRSINFIDGETMKL (Sequence ID 28)

[0075] An example of the C-terminal domain of PylRS from class S bacteria (Deb) of the class Deltaproteobacteria is provided below. MNSSWTEVQRHRLKELNGAEKDLETAFGDDLQRNRAFQKLEKQLVYQERKRLDRLLDTRFRPLRCELESLLIDALKCEGFTRVETPTIISQNDLERMSIDRSHPFNDQVYRVDSKHCLRPMLAPGLYRLMKDLARIRSGKPVRIFEIGPCFRKETSGARHAGEFTMLNLVEMRIEKGSRRFRIETLAKRIMHAAGIDTYDLVDEPSEVYNTTLDIVCGSDPLEVASCAMGPHPLDAAWGIIDTWVGLGFGLERLLMARENSPGIGKWCKSVSYLDGIRLTL(Sequence No. 29)

[0076] An example of a modified C-terminal domain of class S PylRS derived from Deb is provided below. MNSSWTEVQRHRLKELNGAEKDLETAFGDDLQRNRAFQKLEKQLVYQERKRLDRLLDTRFRPLRCELESLLIDALKCEGFTRVETPTIISQNDLERMSIDRSHPFNDQVYRVDSKHCLRPXLXPGXXRLXKDLARIRSGKPVRIFEIGPCFRKETSGARHAGEFTMLXLXEMRIEKGSRRFRIETLAKRIMHAAGIDTYDLVDEPSEVXNTTLDIVCGSDPLEVASCXMGPHPLDAAWGIIDTWVGLGFGLERLLMARENSPGIGKWCKSVSYLDGIRLTL(Sequence ID 30)

[0077] When expressed intracellularly, the N-terminal domain of Deb may also be expressed. The N-terminal domain sequence is as follows: MKETKPAAKRFYRKRVELFRLIDKIKIWPSRTGVLHGIRSVDKRGDIAIITTHCNETFTVRNSRNSRAARWLRNKWFKSVCPACRVPDWKLEKYASTRFKRHFGSDLSRRAD(Sequence ID 62)

[0078] An example of the C-terminal domain of PylRS from class S bacteria (Gem) of the phylum Gemmatimonadetes is provided below. MGITWSKTQKDRLRALRADGARLADSFEGRPQRDQAFQDLEGALAKARRKELEDLRAGHGRPGLCRLQTTLEGTLVGAGFVQVATPTIMSRGLLAKMGVTKNHDLFEQVFWLDRDRCLRPMLAPHLYYVIKDLLRLWEKPLGIFEVGSCFRKDSQGARHSNEFTMLNLCEFGLPEEDRGGRLREMAEVVTRAAGVHEYELEESASTVYGGTLDVVSVDGLELGSGAMGPHPLDHAWRITDTWVGIGFGLERLLMTVNRETSIGKMGRSLAYLDGIPLSI (Sequence ID 31)

[0079] An example of a modified C-terminal domain of class S PylRS derived from a gem is provided below. MGITWSKTQKDRLRALRADGARLADSFEGRPQRDQAFQDLEGALAKARRKELEDLRAGHGRPGLCRLQTTLEGTLVGAGFVQVATPTIMSRGLLAKMGVTKNHDLFEQVFWLDRDRCLRPXLXPHXXYVXKDLLRLWEKPLGIFEVGSCFRKDSQGARHSNEFTMLXLXEFGLPEEDRGGRLREMAEVVTRAAGVHEYELEESASTVXGGTLDVVSVDGLELGSGXMGPHPLDHAWRITDTWVGIGFGLERLLMTVNRETSIGKMGRSLAYLDGIPLSI (Sequence ID 32)

[0080] When expressed intracellularly, the N-terminal domain of Gem may also be expressed. The N-terminal domain sequence is as follows: MSENKKAERERYYRKRVELFRLIDKIKIWPSRKGLLHGIRTTDKMGDVARVTTHCNKTFMVNNSRNSRAARWLRNKWFTGVCPECRIPDWKLAKYSSTHFRRHHGSDLQEGGPPRVDRQPAAPKGAAEVGSGSQEGL (Sequence ID 63)

[0081] An example of the C-terminal domain of class S PylRS derived from Desulfosporosinus sp. I2 (accession number WP_045576271.1) is provided below. MGIIWTPIQKQRLQELNASEAQREMCFESQQARDRAFQEQEHSLVVEGKRRLMELRDIKRRPSLSVLEQQLVEALTQQGFVQVVTPTIISKTSLAKMSVSDDHPLFSQVFWLDSKRCLRPMLAPNLYTLWKDLLRLWEKPIRIFEIGTCYRKESKGSLHLNEFTMLNLTELGLPEDQRHQRLEELASLVMETVGIADYEMELTTSVVYGDTLDVVKGIELGSSAMGPHPLDDQWGIIDPWVGIGFGLERLLMIKEGSQNVQSMGRSLTYLNGVRLNI (Sequence ID 33)

[0082] An example of a modified C-terminal domain of class S PylRS derived from I2 is provided below. MGIIWTPIQKQRLQELNASEAQREMCFESQQARDRAFQEQEHSLVVEGKRRLMELRDIKRRPSLSVLEQQLVEALTQQGFVQVVTPTIISKTSLAKMSVSDDHPLFSQVFWLDSKRCLRPXLXPNXXTLXKDLLRLWEKPIRIFEIGTCYRKESKGSLHLNEFTMLXLXELGLPEDQRHQRLEELASLVMETVGIADYEMELTTSVVXGDTLDVVKGIELGSSXMGPHPLDDQWGIIDPWVGIGFGLERLLMIKEGSQNVQSMGRSLTYLNGVRLNI (Sequence ID 34)

[0083] An example of class A PylRS derived from Candidatus methamethylophilus species 1R26 (1R26) is provided below. MAEHFTDAQIQRLREYGNGTYKDMEFADVSAREKAFTKLMSDASRDNESALKGMIAHPARQGLSRLMNDIADALVADGFIEVRTPIIISKDALAKMTITPDKPLFKQVFWIDDKRALRPMLAPSLYTVMRSLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLNLVDMGPAGDATESLKKYIGIVMKAAGLPDYQLVHEESDVYKETIDVEINGQEVCSAAVGPHYLDAAHDVHEPWAGAGFGLERLLTIRQGYSTVMKGGASTTYLNGAKMD(Sequence ID 35)

[0084] 1R26PylRS may be as described in International Publication No. 2022248061 (incorporated herein by reference). For example, a modified variant may include a mutation that is one or a combination of substitutions at residues L121, L125, Y126, M129, N166, V168, Y206, and A223, such that a variant with altered selectivity is produced.

[0085] An example of a class B PylRS derived from Metanomassillicoccus luminuensis 1 (Lum1) is provided below. MDTRLTPAQAQRIREMGGTVDPSLAFSSEAERESAFQRISADLQGANLAKIRRCAEAPERHPIGSLENTLACALAAKGFIEVKTPMMIPADGLVKMGIDESHPLWNQVFWVGPKKALRPMLAPNLYFLMRHLRRSVPAPLLLFEIGPCFRKESRGSNHLEEFTMLNLVELAPQADATERLKEHIATVMNAVGLPYELVVEGSEVYGTTIDVEVDGVELASGAVGPLPMDKPHGITEPWAGVGFGLERIALMRTKEQNIKKVGRSLVYVNGARIDI (Sequence ID 36)

[0086] Lum1PylRS can be as described in International Publication No. 2022 / 248061. For example, the modified variant can include a mutation that is any one or combination of substitutions at residues L121, L125, Y126, M129, and V168 such that a variant with altered selectivity is generated.

[0087] "X" in the above sequences (i.e., any of SEQ ID NOs: 3, 5, 7, 10, 13, 16, 19, 22, 25, 28, 30, 32, and 34) can be any amino acid, any naturally occurring amino acid, or any of the 20 standard amino acids. The X residue is positioned at a common site that enables modification of PylRS, for example, for alteration of substrate specificity.

[0088] In certain embodiments, MmPylRS, 1R26PylRS, Lum1PylRS or S Δ -ClosPylRS, NitraPylRS or S Δ -I2PylRS, and DebPylS cells are provided that include 2, 3, 4, or all 5 members of the group.

[0089] In certain embodiments, MmPylRS, 1R26PylRS, Lum1PylRS or S Δ -ClosPylRS, and NitraPylRS or S Δ -I2PylRS cells are provided that include.

[0090] In certain embodiments, MmPylRS, 1R26PylRS, Lum1PylRS, NitraPylRS or S Δ -I2PylRS, and DebPylS Cells containing 2, 3, 4, or all 5 members of the group are provided.

[0091] In one embodiment, MmPylRS, 1R26PylRS, S Δ -ClosPylRS, NitraPylRS or S Δ -I2PylRS, and DebPylS Cells containing 2, 3, 4, or all 5 members of the group are provided.

[0092] In one embodiment, MmPylRS, 1R26PylRS, Lum1PylRS, NitraPylRS, and DebPylS Cells containing 2, 3, 4, or all 5 members of the group are provided.

[0093] In one embodiment, MmPylRS, 1R26PylRS, Lum1PylRS, and NitraPsteinRS Cells containing the following are provided.

[0094] Therefore, in one embodiment, exogenous NitraPylRS or S Δ -Includes I2PylRS and, Exogenous MmPylRS, Exogenous 1R26PylRS, Exogenous Lum1PylRS or S Δ -ClosPylRS, and Exogenous DebPylS Cells containing one, two, three, or four members of the group are provided.

[0095] Therefore, in one embodiment, it includes exogenous DebPylS, and Exogenous MmPylRS, Exogenous 1R26PylRS, Exogenous Lum1PylRS or S Δ -ClosPylRS, and Exogenous NitraPylRS or S Δ -I2PylRS Cells containing one, two, three, or four members of the group are provided.

[0096] In one embodiment, MmPylRS and the tRNA encoded by Sequence ID No. 40, 1R26PylRS, and the tRNA encoded by Sequence ID No. 37, Lum1PylRS and the tRNA encoded by Sequence ID No. 59, NitraPylRS, and the tRNA encoded by SEQ ID NO: 39, DebPylS and the tRNA encoded by Sequence ID No. 41 Cells containing 2, 3, 4, or all 5 members of the group are provided.

[0097] In the above embodiment, MmPylRS, 1R26PylRS, Lum1PylRS, S Δ -ClosPylRS, NitraPylRS, S Δ -I2PylRS and DebPylS may be modified to recognize substrates other than pyrrolidine, for example. The anticodon of the tRNA may be modified to recognize a different codon, such as another triplet codon or quadruplet codon.

[0098] In the embodiments described above, MmPylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 1 or 3. 1R26PylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 35. 1R26PylRS may be modified, for example, by mutations in any one or any combination of residues disclosed herein. Lum1PylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 36. Lum1PylRS may be modified, for example, by mutations in any one or any combination of the residues disclosed herein. Δ -ClosPylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of sequence numbers 26, 27, or 28. NitraPylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of sequence numbers 11, 12, or 13. Δ -I2PylRS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 33 or 34. DebPylS may have an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 29 or 30.

[0099] tRNA can be encoded by a sequence having 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity with any one of sequence numbers 37, 39, 40, 41, or 59.

[0100] Exemplary class C and class S PylRS The inventors provide herein PylRS enzymes and demonstrate the functionality of these enzymes when exogenously expressed in host cells.

[0101] In another aspect, a cell is provided that comprises a nucleic acid sequence encoding an exogenous protein having an amino acid sequence that has at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 30 or 32 or SEQ ID NO: 29 or 31, and a protein having an amino acid sequence that has at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identity to SEQ ID NO: 62 or 63. The proteins according to SEQ ID NO: 29 and 30 can pair with SEQ ID NO: 62. The proteins according to SEQ ID NO: 31 and 32 can pair with SEQ ID NO: 63. One or both proteins may be modified to alter acylation specificity. For example, they may be modified to be able to charge a tRNA with an alternative natural amino acid, a non-natural alpha amino acid, or a monomer that is not an alpha amino acid. These proteins may have lost specificity for pyrrolidine.

[0102] The cells described above may have any other features or properties disclosed herein. For example, a cell may express any other acyl-tRNA synthetase disclosed herein or any combination of acyl-tRNA synthetases disclosed herein. A cell may express any tRNA disclosed herein, any combination of tRNAs disclosed herein, or any acyl-tRNA synthetase-tRNA pair disclosed herein. A cell may express two, three, four, five, or six or more mutually orthogonal acyl-tRNA synthetase-tRNA pairs. A cell may be of any type disclosed herein and may be modified in any way.

[0103] Design method The inventors herein provide a classification of PylRS that enables the generation of cells containing a set of aRS that have innate orthogonality with respect to each other. As described herein, cells may express 2, 3, 4, or 5 aRS, each belonging to a different class. The classes as defined herein are classes A, B, C, N, and S.

[0104] Therefore, in one embodiment, a method for producing cells containing at least two exogenous acyl-tRNA synthetases, i) Screening one or more PylRS to identify a first PylRS belonging to class C or S, ii) Screening one or more PylRS to identify a second PylRS belonging to class A, B, C, N, or S, Optionally, iii) modifying the first PylRS and / or the second PylRS to change the acylation specificity, iv) To generate cells expressing a first PylRS and a second PylRS, wherein the first PylRS and the second PylRS are not of the same class. A method is provided that includes this.

[0105] Classes A, B, C, and N are as described in relation to the first aspect of the present invention. Therefore, screening may be performed to determine whether PylRS satisfies any of the definitions described herein.

[0106] Screening for one or more PylRS may involve comparing the amino acid sequences of candidate PylRS with the amino acid sequences of representative PylRS enzymes known to belong to all five classes. The sequences can be compared using unweighted average concatenation clustering. Thus, PylRS of class A, class B, and / or class C can be identified using the sequence-dependent definitions disclosed herein. Additionally, PylRS of class N and / or class S can be identified using the sequence-dependent definitions disclosed herein.

[0107] As described herein, aRSs of classes A, B, and C all lack an N-terminal domain, and therefore the resulting cells neither express the associated N-terminal domain as part of the protein containing the C-terminal domain of the aRS, nor separately. PylRSs of classes A, B, and C may originate from archaeal species.

[0108] Alternatively, PylRS of classes A, B, and C may be derived from bacterial species in which the naturally occurring N-terminal domain is not expressed by the cells that produce them. The classes of such PylRS are defined by a functionality-dependent definition rather than a sequence-dependent definition.

[0109] Alternatively, or further, a PylRS of class N may be defined as a PylRS derived from an archaeal species, wherein the N-terminal domain of the PylRS is included as part of the same polypeptide as the C-terminal domain of the PylRS.

[0110] Alternatively, or further, a PylRS of class S may be defined as a PylRS derived from a bacterial species, where the generated cell expresses a separately encoded N-terminal domain associated with the PylRS of class S.

[0111] Alternatively, or furthermore, screening for one or more PylRS may involve performing assays to determine the class of candidate PylRS. Thus, PylRS of class A, class B, class C, class N, and / or class S may be identified using the functionality-dependent definitions disclosed herein.

[0112] In one embodiment, screening may classify a PylRS as a Class A, Class B, Class C, Class N, or Class S PylRS if it satisfies both sequence-dependent / origin-dependent and functionality-dependent definitions.

[0113] In some embodiments, the method identifies PylRS defined according to the sequence-dependent definitions disclosed herein. In some embodiments, the method identifies PylRS defined according to the sequence-dependent definitions and / or origin-dependent definitions. In some embodiments, the method identifies PylRS defined according to the functionality-dependent definitions disclosed herein. In some embodiments, the method identifies at least one PylRS defined according to the sequence-dependent definitions and at least one PylRS defined according to the functionality-dependent definitions disclosed herein. In some embodiments, the method identifies at least one PylRS defined according to the sequence-dependent definitions and / or at least one PylRS defined according to the functionality-dependent definitions disclosed herein, as well as PylRS of class N and / or class S defined according to the origin-dependent definitions. In some embodiments, the method identifies PylRS that satisfy both the sequence-dependent definitions and the functionality-dependent definitions for each class.

[0114] This method can identify aRS of class C that satisfies the sequence-dependent definition or the functionality-dependent definition, optionally aRS of class A that satisfies the sequence-dependent definition, optionally aRS of class B that satisfies the sequence-dependent definition, and optionally aRS of class N that satisfies the origin-dependent definition.

[0115] This method may include a step of determining whether the PylRS is active in the desired host cell after it has been identified. For example, this method may include a step of determining whether the candidate can be expressed and acylate tRNA in viable E. coli cells. Inactive PylRS are rejected, and a reasonable screening step is repeated.

[0116] The identified PylRS can be modified to alter its acylation specificity. Modified PylRS are described further herein, and the modifications may be for generating such modified PylRS. For example, mutations known to alter the active site of PylRS may be introduced into the backbone of the identified PylRS.

[0117] This method may include further screening one or more PylRS to identify a third PylRS, optionally a fourth PylRS, and optionally a fifth PylRS. Each PylRS may be modified to alter its acylation specificity, as described herein. Cells expressing the first PylRS, the second PylRS, the third PylRS, an optional fourth PylRS, and an optional fifth PylRS may then be generated, wherein each PylRS belongs to a different class.

[0118] The cells produced may be any cells disclosed herein, for example, bacterial cells such as E. coli cells.

[0119] Methods for producing polymers Any of the cells disclosed herein may be used to produce polymers. In a particular embodiment, the use of any of the cells disclosed herein is provided for the production of polymers comprising at least one non-natural amino acid or non-alpha-amino acid.

[0120] In another embodiment, a method is provided for producing a polymer comprising at least one non-natural amino acid or non-alpha-amino acid, comprising culturing any of the cells disclosed herein, supplying the cells with a gene encoding the polymer, and obtaining the polymer.

[0121] Polymers are produced, at least partially, by the genetic integration of monomers. Cells express exogenous aRS as part of an aRS-tRNA pair, and this aRS is capable of charging its tRNA with non-native or non-alpha-amino acids for integration into polymers.

[0122] The polymer may contain standard amino acids. The polymer may contain at least one non-natural amino acid, such as a non-natural alpha-amino acid. The polymer may contain only non-natural amino acids. The polymer may be a macrocyclic molecule. The polymer may contain at least one non-alpha-amino acid. The polymer may contain at least one beta-amino acid. The polymer may contain at least one hydroxy acid, such as at least one alpha-hydroxy acid.

[0123] At least one of the monomers can be supplied exogenously. At least one of the monomers can be synthesized by the cell.

[0124] Non-natural amino acids, non-alpha-amino acids, and alpha-hydroxy acids may be any of those disclosed herein.

[0125] The cells used for polymer production may be any of those disclosed herein. For example, prokaryotic cells such as E. coli cells may be used.

[0126] Modified aRS The aRS enzymes disclosed herein may be wild-type or genetically modified PylRS enzymes. Disclosures of the modified PylRS enzyme and methods for producing the enzyme are found in: (Neumann, H., Peak-Chew, SY & Chin, JW Genetically encoding Nε-acetyllysine in recombinant proteins. Nat. Chem. Biol. 4, 232 (2008). https: / / doi.org:10.1038 / nchembio.73, Yanagisawa et al. (Chem Biol 2008, 15:1187), De La Torre, D. & Chin, JW Reprogramming the genetic code. Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7, Wan et al. (Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics, Volume 1844, Issue 6, June) This includes International Publication No. 2009 / 056803, Pamphlet A1 (2014, Pages 1059–1070), and International Publication No. 2013 / 171485 (each incorporated herein by reference). Modified PylRS is also referred to as mutated PylRS.

[0127] Sequence IDs 3, 5, 7, 10, 13, 16, 19, 22, 25, 28, 30, 32, and 34 indicate residues that can be mutated to alter the acylation specificity of PylRS. These mutations can also be transferred to alternative aRS skeletons to alter the acylation specificity. As those skilled in the art will see, the corresponding residues for mutation can be identified by aligning the sequences disclosed herein with the sequences of another PylRS. In some examples, only the catalytic region of PylRS is aligned to simplify the analysis.

[0128] PylRS may be mutated to enable the tRNA to be charged with specific native or standard amino acids. PylRS may also be mutated to enable the tRNA to be charged with non-native amino acids. "Non-native amino acids" are amino acids other than L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan, L-tyrosine, L-pyrrolidine, or L-selenocysteine.

[0129] PylRS may be mutated to allow tRNA to be charged with non-standard amino acids. "Non-standard amino acids" are amino acids other than L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L-phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L-asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L-tryptophan, or L-tyrosine.

[0130] Suitable non-natural amino acids include those disclosed in Neumann, H., 2012. FEBS letters, 586(15), pp.2057-2064, and Liu, CC and Schultz, PG, 2010. Annual review of biochemistry, 79, pp.413-444. For example, non-natural amino acids include p-acetylphenylalanine, m-acetylphenylalanine, O-allyl tyrosine, phenylselenocysteine, p-propargyloxyphenylalanine, p-azidophenylalanine, p-boronophenylalanine, O-methyltyrosine, p-aminophenylalanine, p-cyanophenylalanine, m-cyanophenylalanine, p-fluorophenylalanine, p-iodophenylalanine, p-bromophenylalanine, p-nitrophenylalanine, L-DOPA, 3-aminotyrosine, 3-iodotyrosine, p-isopropylphenylalanine, 3-(2-naphthyl)alanine, biphenylalanine, and homoglutamine. One or more of the following may be selected: D-tyrosine, p-hydroxyphenyllactic acid, 2-aminocaprylic acid, bipyridylalanine, HQ-alanine, p-benzoylphenylalanine, o-nitrobenzylcysteine, o-nitrobenzylserine, 4,5-dimethoxy-2-nitrobenzylserine, o-nitrobenzyllysine, o-nitrobenzyltyrosine, 2-nitrophenylalanine, dansylalanine, p-carboxymethylphenylalanine, 3-nitrotyrosine, sulfotyrosine, acetyllysine, methylhistidine, 2-aminononanoic acid, 2-aminodecanoic acid, pyrrolidine, Cbz-lysine, Boc-lysine, and allyloxycarbonyllysine.

[0131] PylRS may be mutated to allow the tRNA to be charged with non-standard monomers. Non-standard monomers are, for example, monomers that are not alpha-amino acids (see, incorporated herein by reference, (Spinck, M. et al. Genetically programmed cell-based synthesis of non-natural peptide and depsipeptide macrocycles. Nature Chemistry volume 15, pages 61-69 (2023))). In some embodiments, the non-alpha-amino acid is an alpha-hydroxy acid or a beta-amino acid.

[0132] Non-alpha-amino acids are monomers that lack the RCH(NH2)COOH structure of an amino acid, and in particular, the alpha(NH2) group. Alpha-hydroxy acids are exemplary non-alpha-amino acids. Some alpha-hydroxy acids are hydroxy variants of standard amino acids. Alternatively, they may be non-standard acids. Exemplary alpha-hydroxy acids include hydroxy acids having aromatic side chains, which may be selected from F-OH, pIF-OH, and NapA-OH, as well as alpha-hydroxy acids having aliphatic side chains, which may be selected from BocK-OH, PenK-OH, AllocK-OH, NorK-OH, AlkynK-OH, CbzK-OH, ButK-OH, and AcK-OH. Hydroxy acid analogs of O4BBy, O2beY, pCaaF, pVsaf, and pAaF are also considered. See (Iannuzzelli and Fasan, Chem. Sci., 2020, 11, 6202).

[0133] Generally, alpha-hydroxy acids can be derived from standard or non-standard amino acids. For example, alpha-hydroxy acids include, but are not limited to, p-hydroxy-L-phenyllactic acid (an alpha-hydroxy analog of tyrosine), leucic acid (an alpha-hydroxy analog of leucine), lactic acid (an alpha-hydroxy analog of alanine), 2-hydroxy-3-methylbutyrate (an alpha-hydroxy analog of valine), 2-hydroxy-3-phenylpropionic acid, hydroxy derivatives (F-OH) of phenylalanine, and alpha-hydroxy analogs of other natural and non-natural amino acids. Examples of non-standard amino acid derivatives include the hydroxy derivatives of Nε-Alloc-L-lysine (AllocK-OH), p-iodophenylalanine (pIF-OH), Nε-((propa-2-in-1-yloxy)carbonyl)-L-lysine (AlkynK-OH), p-azidophenylalanine (pAzF-OH), L-3-(2-naphthyl)alanine (NapA-OH), Nε-tert-butoxycarbonyl-L-lysine (BocK-OH), and N6-carbobenzyloxy-L-lysine (CbzK-OH), as well as other lysine derivatives including PenK-OH, NorK-OH, ButK-OH, and AcK-OH.

[0134] tRNA The PylRS enzymes disclosed herein can be used in conjunction with exogenous tRNAs to form exogenous PylRS-tRNA pairs. Each exogenous PylRS-tRNA pair may be capable of incorporating monomers at the codon. For example, the monomers may be native amino acids, non-native amino acids, or non-alpha-amino acids. PylRS may be mutated to allow for the charging of specific monomers, such as alternative native amino acids, non-native amino acids, or certain hydroxy acids, into the tRNA (see the description herein of modified PylRS enzymes). The tRNA may have anticodons that have been modified for specific sense codons or stop codons.

[0135] In embodiments using 2, 3, 4, or 5 exogenous PylRS-tRNA pairs, they may be directed to incorporate 2, 3, 4, or 5 different monomers. Monomers may include native amino acids, unnatural alpha-amino acids, and non-alpha-amino acids (such as beta-amino acids or hydroxy acids). In embodiments using 2, 3, 4, or 5 exogenous PylRS-tRNA pairs, they may be directed to incorporate 2, 3, 4, or 5 monomers (each).

[0136] The tRNA may be a tRNA naturally associated with a PylRS. The tRNA may be a modified version of a tRNA naturally associated with a PylRS. Alternatively, the tRNA, or a modified version thereof, may not be naturally associated with a PylRS. The tRNA may also be one that is normally associated with the same or a different class of PylRS compared to the PylRS paired with the tRNA.

[0137] The tRNA may be derived from, for example, an archaeal species of the genus Methanosarkina. The tRNA may also be derived from a bacterial species, for example, Desulfitobacterium.

[0138] In some embodiments, two or more of the exogenous tRNAs expressed by the cell have less than 75% sequence identity with respect to each other. In other embodiments, all of the exogenous tRNAs expressed by the cell have less than 75% sequence identity with respect to each other. Alternatively, or further, two or more of the intracellular exogenous tRNAs are different due to exogenous features. Examples of exogenous features include 6-8 base pair D-loops, long variable loops, unusual (e.g., adenine or uracil) nucleic acid bases at the identification base position, bulges in the anticodon stem, and loops in the anticodon stem. Long variable loops may have a nucleotide length greater than 3. For example, a long variable loop may have a nucleotide length of 4.

[0139] Examples of tRNA include those encoded by the following DNA sequences: Bathy-tRNA: gggggtttggccgaggcggtcgcgagggtactaagctctcgcagccgggttcaactcccgggacccccgcca (SEQ ID NO: 42) Nitra-tRNA: gggggctcggccgaggcggccacaggggctctatacccctgcagccgggttcaactcccggagcccccgcca (SEQ ID NO: 43) SCGC-tRNA: ggggggcaggccgaggatggccagtggggctctaaaccctgcgtctaccgggttcaattcccgggccccccacca (SEQ ID NO: 44) Therm1-tRNA: SEQ ID NO: 39 Therm2-tRNA:ggggggttggtcgggttgaccaaaggaggctctaaaccttctcaagggttcaggcaaatcctgggcctttaccgggttcgactctcgggccccccgcca(SEQ ID NO: 45) Tron-tRNA:ggggggctggtcggggtgaccacggaggccctatacctcccttagccgggttcaactcccgggtccctcgcca (SEQ ID NO: 46) Gem-tRNA:ggggagtggatcgatacaagatcgtgtgggctctaaacccacatagctcggtgtgactccggggctccccacca (SEQ ID NO: 47) Bacc-tRNA: ggagggtgaatcggagtgattatgcggactctaaatccgtacagccgggtgcaactcccggactcttcgcca (SEQ ID NO: 48) I2-tRNA:ggggggtagatcggattgatcgcgtggactctaaatccgcgtagacgggtgaaactcccgtactcctcgcca(SEQ ID NO: 49) Clos-tRNA:ggggaacggatcggatggatcacatagactctaaatctatgtagccgagtgaaactctcgggttcctcgcca (SEQ ID NO: 50) Deb-tRNA:ggggagtggatcggatgagatcgcatgggctctaaacccatgtagccgggtgcgactcccgggcttccctcca (SEQ ID NO: 51) Spi-tRNA: ggggggcggatcgcaagcgagatcgcgcggactctaaatccgcgcagccgggtgcaactcccgggctcctcgcca (SEQ ID NO: 52) Mic-tRNA:ggagtgttgatcgacaacgggatcacatggactctaaatccatgcagccgggtgcaactcccggacactccgcca (SEQ ID NO: 53) Met-tRNA: SEQ ID NO: 40 Bur-tRNA:GGAGACTTGATCATGTAGATCGAACGGACTCTAAATCCTTTCAGCCGGGTTAGATTCCCGGAGTTTCCGCCA (Sequence ID 54) Pro-tRNA:GGAAATCAGATCATGTTGATCGAATGGACTCTAAATCCGTTCAGTCGGGTTAAATTCCCGAGGTTTCCGCCA(Sequence ID 55) Therm-tRNA (intron removed): SEQ ID NO: 39 Tron-tRNA (intron removed and anticodon loop modified): GGGGGGCTGGTCGGGGTGACCACGGAGGCTCTAAACCTCCCTTAGCCGGGTTCAACTCCCGGGTCCCTCGCCA (Sequence ID 56) I2CUAG-tRNA:GGGGGGTAGATCGGATTGATCGCGTGGACTCTAAATCCGCGCTAGACGGGTGAAACTCCCGTACTCCTCGCCA (Sequence ID 57) I2b8-tRNA:GGGGCGTCGATCGGATTGATCGCGTGGACTCTAAATCCGCGCGCAACGGGTGAAACTCCCGTACGCCTCGCCA (Sequence ID 58) I2b32-tRNA: SEQ ID NO: 38 I2b72-tRNA:GGGGTGTAGATCGGATTGATCGCGTGGACTCTAAATCCGCGAACAACGGGTGAAACTCCCGTACACCTCGCCA(Sequence ID 59) I2H52-tRNA: SEQ ID NO: 41 Alv17-tRNA:GGGGGACGGTCCGGCGACCAGCGGGTCTCTAAAACCTAGCgtaagCGGGGTTCGACaCCCCGGTCTCTCGCCA (Sequence ID 60) Alv21-tRNA: SEQ ID NO: 37 Alv22-tRNA:GGGGGACGGTCCGGCGACCAGCGGGTCTCTAAAACCTAGCtcaaggCGGGGTTCGACtCCCCGGTCTCTCGCCA (Sequence ID 61)

[0140] Therefore, cells may express 1, 2, 3, 4, 5, or 6 or more exogenous tRNAs. Cells may express 1, 2, 3, 4, 5, or 6 or more exogenous PylRS-tRNA pairs.

[0141] As described herein, cells may express one, two, three, four, or five tRNAs encoded by any combination of SEQ ID NO: 37, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, or SEQ ID NO: 59, or their variants, in particular variants in which the anticodon is altered.

[0142] The tRNA may be selected to exhibit activity with one of the exogenous PylRS enzymes and to have reduced activity with other aRS expressed by the host cell. Thus, a PylRS-tRNA pair may be orthogonal with respect to any other PylRS-tRNA pair in the host cell. In some embodiments, the host cell may contain one, two, three, four, five, or six or more mutually orthogonal exogenous PylRS-tRNA pairs.

[0143] Orthogonal aRS-tRNA pairs induce the expression of a marker gene at or above 40% of the control gene expression level in test cells, whereas pairing aRS with any other tRNA in the cell, or pairing aRS with any other aRS in the cell, induces the expression of a marker gene at or below 20% of the control gene expression level. Appropriate assays are disclosed herein (for example, the marker gene is a GPF gene with a stop codon located internally, and the control gene is wild-type GFP).

[0144] An orthogonal set of aRS-tRNA pairs is defined as a set where each aRS-tRNA pair satisfies the above criteria (i.e., each pair achieves expression of a marker gene at a level higher than 40%, each cross-reactivity causes expression of a marker gene at a level lower than 20%), and the quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity is greater than 2.5.

[0145] host cell The cells of this disclosure, including all embodiments of cells disclosed herein, may be prokaryotic cells. The cells may be bacteria. The bacterial cells may be of any species, as long as they are suitable for the production of heterologous proteins, in particular polypeptides. Suitable bacterial host cells include Escherichia species (e.g., Escherichia coli), Caulobacter species (e.g., Caulobacter crescentus), phototrophic bacteria (e.g., Rhodobacter sphaeroides), cold-adapted bacteria (e.g., Pseudoalteromonas haloplanktis, Shewanella strain Ac10), Pseudomonas bacteria (e.g., Pseudomonas fluorescens, Pseudomonas putida, Pseudomonas aeruginosa), and halophilic bacteria (e.g., Halomonas elongata). elongate), Chromohalobacter salexigens, Streptomyces (e.g., Streptomyces lividans, Streptomyces griseus), Nocardia (e.g., Nocardia lactamdurans), Mycobacteria (e.g., Mycobacterium smegmatis), Corynebacteria (e.g., Corynebacterium glutamicum, Corynebacterium ammoniagenes, Brevibacterium lactofermentum) Bacillus lactofermentum), rod-shaped bacteria (for example, Bacillus subtilis, Bacillus brevis)Examples include Bacillus brevis, Bacillus megaterium, Bacillus licheniformis, Bacillus amyloliquefaciens, Vibrio bacteria (e.g., Vibrio cholera, Vibrio natriegens), and lactic acid bacteria (e.g., Lactococcus lactis, Lactobacillus plantarum, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri). In some cases, the bacteria are Gram-negative.

[0146] In certain cases, the bacteria are Escherichia coli, Salmonella enterica, or Shigella dysenteriae. More preferably, the bacteria are Escherichia coli. Suitable E. coli cells include K-12, MG1655, BL21, BL21(DE3), AD494, Origami, HMS174, BLR(DE3), HMS174(DE3), Tuner(DE3), Origami2(DE3), Rosetta2(DE3), Lemo21(DE3), NiCo21(DE3), T7 Express, SHuffle Express, C41(DE3), C43(DE3), and m15 pREP4, or their derivatives (Rosano, GL and Ceccarelli, EA, 2014. Frontiers in microbiology, 5, p.172). In particular, the cells may be MG1655 or BL21, or derivatives thereof. MG1655 is considered to be the wild-type strain of E. coli. The GenBank ID for the genome sequence of this strain is U00096. BL21 is widely available commercially. For example, it can be purchased from New England BioLabs under catalog number C2530H.

[0147] Host cells may contain a recoded genome in which one or more sense codons are recoded. Examples of such organisms (Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019). https: / / doi.org:10.1038 / s41586-019-1192-5, Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057–1062 (2021). https: / / doi.org:10.1126 / science.abg3029) are provided in International Publication No. 2020 / 229592 A1 and International Publication No. 2022 / 248061, each of which is incorporated herein by reference in whole.

[0148] As used herein, “re-coding” means replacing the occurrence of a codon with a different codon so that the occurrence of a certain type of codon disappears from the genome. A re-coded sense codon may be replaced with a synonymous codon so as to result in a different codon usage frequency without altering the encoded polypeptide. Alternatively, a sense codon may be replaced with a non-synonymous codon, for example, if the alteration of the encoded polypeptide sequence does not affect viability. The number of occurrences of a type of sense codon being re-coded may be sufficient to allow the removal of the corresponding tRNA of the same family while maintaining the viability of the cell. Stop codons can also be re-coded, for example, by replacing the occurrence of a first type of stop codon with a synonymous codon. The number of occurrences of a type of stop codon being re-coded may be sufficient to allow the removal of the corresponding termination factor of the same family while maintaining the viability of the cell.

[0149] In certain cases, the host cell genome is recoded such that sense codon TCG is replaced with AGC, sense codon TCA is replaced with AGT, and stop codon TAG is replaced with TAA, with a sufficient number of such codons being recoded without the need for two congeneral tRNAs and congeneral termination factors.

[0150] The host cell may be a bacterium, such as E. coli, whose genome has been recoded. For example, the cell may be a bacterial cell whose genome has been recoded with respect to codons TCA and TCG. The bacterial cell has tRNA Ser UGA and tRNA Ser CGA The cells may lack , and may also lack RF-1. The cells may be Syn61, a strain derived from Syn61, or re-encoded in the same way as Syn61. The cells may be Syn61Δ3, a strain derived from Syn61Δ3, or modified in the same way as Syn61Δ3.

[0151] The host cell can be recoded to add four types of sense codons, such as a new quadruplet codon, and to recode one type of stop codon.

[0152] Cells disclosed herein, including cells expressing one or more aRS, are viable cells. Viable cells are capable of being metabolically active. In certain embodiments, viable cells may be able to grow when cultured in a suitable medium and under suitable conditions for a particular species or strain. Such cells may be referred to as culturable. For example, if the cells are bacterial cells such as Escherichia coli, the viability may be evaluated by culturing the bacteria at 37°C in a medium containing LB medium or agar containing LB agar. The medium or agar may be supplemented with 2% glucose. Bacterial growth is OD 600 Monitoring can be performed using standard methods such as the measurement of [specific parameters]. Alternative methods, or methods adapted for specific cells, bacterial strains, bacterial species, or to account for the inclusion of marker genes, are known to those skilled in the art.

[0153] Sequence comparison can be performed using readily available sequence comparison programs. These publicly and commercially available computer programs can calculate sequence identity between two or more sequences.

[0154] An experienced technician will know how to calculate the percentage of identity between two nucleic acid sequences. To calculate the percentage of identity between two nucleic acid sequences, it is first necessary to create an alignment of the two sequences and then calculate the value of sequence identity. The percentage of identity between two sequences may take different values ​​depending on (i) the method used to align the sequences, e.g., the Needleman-Wunsch algorithm (e.g., applied by Needle(EMBOSS) or Stretcher(EMBOSS)), the Smith-Waterman algorithm (e.g., applied by Water(EMBOSS)), or the LALIGN application (e.g., applied by Matcher(EMBOSS)), and (ii) the parameters used by the alignment method, e.g., whether it is a local or global alignment, the matrix used, and the parameters applied to the gaps. In certain embodiments, the sequence identity disclosed herein may be calculated based on a global alignment of reasonable features, e.g., a comparison of the C-terminal domain of one PylRS with the C-terminal domain of another PylRS.

[0155] After alignment, there are many different ways to calculate the percentage of identity between two sequences. For example, the number of identities can be divided by (i) the length of the shortest sequence, (ii) the length of the alignment, (iii) the average length of the sequences, (iv) the number of non-gap positions, or (iv) the number of equivalent positions excluding overhangs. Furthermore, it will be found that the percentage of identity is strongly dependent on length. Therefore, the shorter the pair of sequences, the higher the sequence identity that can be expected to occur by chance.

[0156] Next, the percentage of identity between two nucleic acid sequences can be calculated from such alignments as (N / T)*100, where N is the number of positions in which the sequences share identical residues, and T is the total number of positions to compare, including gaps but excluding overhangs.

[0157] Sequence alignment can be pairwise sequence alignment. Suitable services include Needle(EMBOSS), Stretcher(EMBOSS), Water(EMBOSS), Matcher(EMBOSS), LALIGN, or GeneWise. In one example, the identity between two amino acid sequences can be calculated using the Needle(EMBOSS) service with default parameters set to, for example, matrix (BLOSUM62), gap open (10), gap extend (0.5), end gap penalty (false), end gap open (10), and end gap extend (0.5). In another example, the identity between two amino acid sequences can be calculated using the Matcher(EMBOSS) service with default parameters set to, for example, matrix (BLOSUM62), gap open (14), gap extend (4), and alternative matches (1). In one example, the identity between two nucleic acid sequences can be calculated using the service Needle (EMBOSS) with default parameters set to, for example, matrix (DNAfull), gap open (10), gap extend (0.5), end gap penalty (false), end gap open (10), and end gap extend (0.5). In another example, the identity between two nucleic acid sequences can be calculated using the service Matcher (EMBOSS) with default parameters set to, for example, matrix (DNAfull), gap open (16), gap extend (4), and alternative matches (1).

[0158] All of the features described herein (including any accompanying claims, abstracts, and drawings), and / or all of the steps of any methods or processes disclosed herein, can be combined with any of the above embodiments in any combination, except for any combination in which at least some of such features and / or steps are mutually exclusive.

[0159] All published documents, patent applications, patents, and other references referenced herein are incorporated herein by reference as a whole. References cited herein are not considered prior art relating to the claimed disclosures. In case of any conflict, this specification shall prevail, including definitions.

[0160] Examples are provided below to aid in understanding the present invention and to illustrate how its embodiments may be carried out. The examples are not intended to limit the present invention in any way. [Examples]

[0161] Abstract The discovery of mutually orthogonal aminoacyl-tRNA synthetase (aaRS) / tRNA pairs provides a basis for incorporating non-standard amino acid (ncAA) combinations into proteins within cells, as well as for synthesizing encoded non-standard polymers and macrocyclic molecules. Here, we define a cluster of pyrrolidine tRNA-synthetase (PylRS) sequences that exceed an empirically determined threshold for mutual orthogonality. 84% of the resulting cluster belongs to a PylRS class that has not been explored in the search for orthogonal pairs. We determine that 95% of the members of the PylRS cluster are derived from the same organism as tRNA. Pyl Sequence identification and clustering are performed, and they are filtered using empirically sequence-based mutual orthogonality thresholds and the presence of novel structural features. The inventors thereby perform mutual orthogonality searches on PylRS / tRNA Pyl Define a group of pairs. The inventors define PylRS and tRNAPyl Two novel classes of sequences were identified, and the majority of the novel PylRS enzymes and pyl tRNAs are active and orthogonal. The inventors define five classes of PylRS systems that form the basis of quintuple orthogonal pairs. Notably, the inventors' computational method is used to identify quintuple orthogonal PylRS / tRNAs. Pyl The inventors directly address 20 of the 25 aminoacyl-tRNA synthetase / tRNA pairwise specificities required to create pairs. The remaining 5 specificities are addressed by tRNA Pyl This is controlled by modification and oriented evolution. In general, the inventors have identified 924 mutually orthogonal PylRS / tRNAs. Pyl We generate pairs, 1324 triple orthogonal pairs, 128 quadruple orthogonal pairs, and 8 quintuple orthogonal pairs. These advances will provide an important foundation for the synthesis of encoded polymers.

[0162] Introduction Here, the inventors of the present invention have developed an experimental PylRS / tRNA Pyl Using cross-reactivity data, the inventors empirically define sequence identity thresholds for mutually orthogonal PylRS enzymes and pyl tRNAs. Next, they perform aggregate clustering of 351 PylRS sequences to define clusters of sequences exceeding the empirical threshold. 84% of the resulting clusters belong to PylRS classes not explored in the orthogonal pair search. The inventors identify tRNAs of the same biological origin for 95% of the members of the PylRS sequence clusters. Pyl Sequence identification and clustering are performed. By using both an empirical orthogonality threshold and the presence of exogenous structural features that can result in orthogonality, the inventors select a set of pyl tRNAs that serve as a starting point for experimentally searching for mutually orthogonal pairs in conjunction with the PylRS enzyme of the same organism.

[0163] The inventors of this invention have identified PylRS and tRNA PylWe identify two new classes of sequences (which we name Class C and Class S), and we demonstrate that the majority of our new PylRS enzymes and pyl tRNAs are active and orthogonal in E. coli. We explore the specificity of Class S and Class C systems against each other and against previously characterized Class N, A, and B PylRS systems. Notably, our sequence-based method enables quintuple orthogonal PylRS / tRNA without additional modification. Pyl This allows control over 20 of the 25 aminoacyl-tRNA synthetase / tRNA pairwise specificities required to create a pair set. The inventors have controlled the remaining 5 specificities of tRNA Pyl This is controlled by modification and oriented evolution. In general, the inventors have identified 924 mutually orthogonal PylRS / tRNAs. Pyl We generate pairs, 1324 triple orthogonal pairs, 128 quadruple orthogonal pairs, and 8 quintuple orthogonal pairs. [Examples]

[0164] PylRS / tRNA Pyl Pair cross-reactivity and sequence identity The inventors have identified three different classes of PylRS / tRNA (N, A, and B). Pyl Having previously defined the cross-reactivity profiles of pairs, we have shown that certain non-homogeneous pairs derived from ΔN PylRS classes A and B exhibit a surprising degree of innate orthogonality with respect to each other. We hypothesize that this mutual orthogonality may be related to sequence identity between pairs, and therefore, we have identified ΔN PylRS / tRNA PylWe decided to quantify this relationship by utilizing activity data (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x).

[0165] Notably, the inventors have found that when two ΔN PylRS enzymes have more than 55% sequence identity, one ΔN PylRS enzyme will naturally pair with the other ΔN PylRS enzyme's tRNA. Pyl We found that high activity was observed in approximately 90% of cases (Figure 1a). With less than 55% sequence identity, the ΔN PylRS enzyme exhibited broad activity with the pyl tRNAs of other ΔN PylRS enzymes. Similarly, two tRNAs Pyl When genes share more than 75% sequence identity, one of the tRNAs Pyl is the other tRNA Pyl Together with synthetase (and vice versa), it showed high activity in approximately 90% of cases (Figure 1b). With less than 75% sequence identity, pyl tRNA exhibited broad activity with the ΔN PylRS enzyme of other pyl tRNAs.

[0166] This analysis suggests that the development of new multiple orthogonal pairs involves PylRS / tRNAs where the sequence identity of the synthetase and tRNA is less than 55% and less than 75%, respectively. Pyl It was suggested that emphasis should be placed on pairs. [Examples]

[0167] PylRS sequence identification and clustering Candidate PylRS / tRNAs for the Development of Multiple Orthogonal Pairs Pyl To identify the combinations, the inventors first constructed a database of PylRS sequences. AlvΔN PylRS sequences (Class A; hereafter, A Δ By performing a BLAST search for sequence similarity with (referred to as -AlvPylRS), the inventors read 351 PylRS protein sequences. Of these, 79 belonged to the Archaeon +N group, 66 belonged to the Archaeon ΔN group, and 204 belonged to the Bacteria (sN) group. Furthermore, two PylRS genes, despite being classified as Archaeon, possessed separately encoded N-terminal domains. The inventors named these the Archaeon sN group.

[0168] The inventors performed aggregated hierarchical clustering to visualize sequence diversity among PylRS catalytic domains (Figure 1c). They observed two main groups: dense clusters consisting of PylRS sequences from class N, and loose clusters consisting of PylRS sequences from bacteria and other archaea. The latter clusters themselves contained several relatively dense subclusters, including those corresponding to sequences from known archaeal classes A and B.

[0169] To discover mutually orthogonal systems, the inventors focused on identifying PylRS sequences with a pairwise sequence identity of less than 55% (Figure 1a). To achieve this, the inventors set a linkage distance threshold for aggregated clustering such that two clusters were merged only if the average identity percentage of each PylRS in the two clusters was greater than 55%. This yielded 37 clusters (Figure 1d). Three of the clusters corresponded to known PylRS classes N, A, and B. In contrast, there were 25 bacterial sN group clusters and 9 additional archaeal clusters (7 ΔN group and 2 archaeal sN group). This analysis demonstrated that substantial sequence diversity between Pyl systems has not yet been explored. [Examples]

[0170] tRNA Pyl Sequence identification and clustering The inventors then selected representative PylRS enzymes from each of the 37 clusters and aimed to identify pyl tRNAs originating from the same organism. They obtained DNA sequences or metagenomic reads of the host genome containing each PylRS gene and used the tRNA detection program ARAGORN (Laslett, D. ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences. Nucleic Acids Research 32, 11-16 (2004). https: / / doi.org:10.1093 / nar / gkh152) to detect tRNAs. Pyl The presence or absence of genes was tested. In total, the inventors found tRNA for 35 out of 37 types of PylRS enzymes. Pyl We obtained the genes.

[0171] The inventors have identified 35 types of tRNA Pyl Agglomerative hierarchical clustering of sequences was performed (Figure 1e). The inventors observed tighter grouping than with the PylRS sequences. Of the nine new archaeal pyl tRNAs, three formed a group with class B, and six formed a distinct new group, though more loosely related, which the inventors named class C. On the other hand, the pyl tRNAs of bacteria sN formed a strongly aggregated group, which the inventors assigned to a new class S. Interestingly, the pyl tRNAs of the PylRS enzymes from two archaeal sNs were fairly weakly related and classified into classes B and C, suggesting that they may not share a common origin.

[0172] By setting the linkage distance threshold for aggregated clustering to 75% sequence identity (Figure 1b), the inventors have found that tRNA PylEight gene clusters were generated (Figure 1f). Of these clusters, three corresponded to known PylRS classes N, A, and B. There were two bacterial clusters, one of which contained a single tRNA. Pyl It contained only [this element]. The remaining three clusters were of class C, which matched the relatively loose interrelationships among members of this class.

[0173] Some of the newly discovered pyl tRNAs are standard class N tRNAs. Pyl It contains exogenous structural features not observed in tRNA (Figure 1g). For example, certain class S pyl tRNAs contain a 6-8 base pair D loop, while some class C pyl tRNAs contain a long variable loop. Previous studies have shown that structural elements are present in tRNA. Pyl It has been shown that tRNAPyl can strongly influence PylRS interactions (Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5, Fan, C., Xiong, H., Reynolds, NM & Soll, D. Rationally evolving tRNAPyl for efficient incorporation of noncanonical amino acids. Nucleic Acids Research 43, e156-e156 (2015). https: / / doi.org:10.1093 / nar / gkv800), and in fact, PylRS / tRNA interactions with endogenous aminoacyl-tRNA synthetases in various host organisms have been shown. Pyl The orthogonality of the system is tRNA PylThis is attributed to the compact structure of the main body. (Nozawa, K. et al. Pyrrolysyl-tRNA synthetase-tRNAPyl structure reveals the molecular basis of orthogonality. Nature 457, 1163 (2008). https: / / doi.org:10.1038 / nature07611) In particular, tRNA PylThe extension of the variable loop has been used to maintain activity with congeneral PylRS classes while reducing cross-reactivity with non-congeneral PylRS classes. (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5, Meineke, B., Heimgartner, J., Eirich, J., Landreh, M. & Elsasser, SJ Site-Specific Incorporation of Two ncAAs for Two-Color Bioorthogonal Labeling and Crosslinking of Proteins on Live Mammalian Cells. Cell Reports 31, 107811 (2020). https: / / doi.org:10.1016 / j.celrep.2020.107811, Meineke, B., Heimgartner, J., Lafranchi, L. & Elsasser, S.J.Methanomethylophilus alvus Mx1201 Provides Basis for Mutual Orthogonal Pyrrolysyl tRNA / Aminoacyl-tRNA Synthetase Pairs in Mammalian Cells. ACS Chemical Biology 13, 3087-3096 (2018). https: / / doi.org:10.1021 / acschembio.8b00571).

[0174] Furthermore, multiple pyl-tRNAs of both classes were found at the discriminator base position, which is a known identity element for previously characterized PylRS proteins (Ambrogelly, A. et al. Pyrrolysine is not hardwired for cotranslational insertion at UAG codons. Proc. Natl. Acad. Sci. USA 104, 3141–-3146 (2007). https: / / doi.org:10.1073 / pnas.0611634104, Zhang, H. et al. The tRNA discriminator base defines the mutual orthogonality of two distinct pyrrolysyl-tRNA synthetase / tRNAPyl pairs in the same organism. Nucleic Acids Research 50, 4601–4615 (2022).). (https: / / doi.org:10.1093 / nar / gkac271), it contains unusual (adenine or uracil) nucleic acid bases. Therefore, when selecting pyl tRNA for further characterization, each tRNA PylIn addition to ensuring that at least one member is selected from a cluster, the inventors also selected additional pyl tRNAs from a particular cluster based on such non-standard characteristics. The inventors hypothesized that, despite their relatively high sequence identity (over 75%) with the other selected pyl tRNAs, their structural differences could lead to orthogonal interactions with each other.

[0175] The inventors selected a total of 16 pyl-tRNAs for further investigation (Figure 1h). This group included 13 newly selected pyl-tRNAs (6 archaeal class C and 7 bacterial class S) that had not been previously characterized in E. coli, along with three previously characterized pyl-tRNAs from classes N, A, and B. PylRS / tRNA for experimental characterization Pyl To determine a representative set of pairs, the inventors used the previously reported PylRS enzyme A Δ -1R26PylRS and B Δ A-AlvtRNAs that form highly active heterologous pairs with Lum1PylRS, respectively. Pyl and B-InttRNA Pyl With the exception of the selected tRNA Pyl This was combined with a synthetase of the same organism (Figure 1h). (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x) Of the 10 interclass relationships, it was found that only 4 had been partially characterized in E. coli (Figure 1i). +-MbPylRS is S-Desulfitobacterium hafniense (Dh)-tRNA Pyl It is known to interact with (Herring, S. et al. The amino-terminal domain of pyrrolysyl-tRNA synthetase is dispensable in vitro but required for in vivo activity. FEBS Lett. 581, 3197--3203 (2007). https: / / doi.org:10.1016 / j.febslet.2007.06.004), and a set of modified pairs of classes N, A, and B (N + -MmPylRS / N-Methanosarcina spelaei (Spe)tRNA Pyl , A Δ -1R26PylRS / A-AlvtRNA Pyl-8 , and B Δ -Lum1PylRS / B-InttRNA Pyl-17C10These (etc.) are known to be triple orthogonal to each other. (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5)

[0176] In preparing a new system for characterization, the inventors observed that several class S PylRS genes were difficult to clone, and that even those successfully cloned resulted in reduced growth when expressed in E. coli cells. The inventors hypothesized that these problems might be related to the separate expression of their N-terminal domain proteins (PylSn), and prepared variants by removing the PylSn gene from each class S PylRS system, thereby neutralizing the toxic effect. These variants (S Δ (to be named) wild-type enzyme (S + The characteristics were evaluated in conjunction with (or instead of) naming it as follows. [Examples]

[0177] Novel active PylRS enzyme and pyl tRNA In the presence of the non-standard amino acid N6-((allyloxy)carbonyl)-L-lysine (AllocK1, Fig. 2a), a known substrate of the previously characterized PylRS enzyme, the inventors produced GFP from a gene encoding green fluorescent protein (GFP) containing an amber codon at position 150 for each selected tRNA Pyl and measured the activity of each selected PylRS enzyme with it (Fig. 2b). (Katayama, H., Nozawa, K., Nureki, O., Nakahara, Y. & Hojo, H. Pyrrolysine Analogs as Substrates for Bacterial Pyrrolysyl-tRNA Synthetase in Vitro and in Vivo. Bioscience, Biotechnology, and Biochemistry 76, 205-208 (2012). https: / / doi.org:10.1271 / bbb.110653)

[0178] Fifteen out of the 16 types of pyl tRNAs caused PylRS-dependent GFP production (at least 30% of the level produced from the control GFP gene without an amber stop codon, "wtGFP control") in the presence of at least one PylRS enzyme. This included all of the class C pyl tRNAs and all but one of the class S tRNAs Pyl . Furthermore, 13 out of the 20 types of PylRS enzymes caused GFP production at least 30% of the level of the wtGFP control in the presence of at least one tRNA Pyl . Additionally, six new PylRS enzymes (C Δ -nitrososphaera phylum archaea (Nitra) PylRS, C<tmp Δ -methanonatronarchaeum phylum archaea (Tron) PylRS, S Δ -desulfosporosinus species I2 (I2) PylRS, S Δ -clostridiales bacterium (Clos) PylRS, S Δ -delta-proteobacteria phylum bacterium (Deb) PylRS, and S Δ- Spirochaeta sp. PylRS caused GFP production at levels of at least 50% of the wtGFP control in the presence of a suitable new class C or S tRNA Pyl and.

[0179] Interestingly, the active class C PylRS enzyme showed greater specificity for a particular class C pyl tRNA than for class S pyl tRNA. In particular, C-TronPylRS had high activity with C-TrontRNA Pyl (76% of the wtGFP control), but activity with all but one of the class S tRNAs tested Pyl was less than 10%. Most of the active class S PylRS enzymes aminoacylated both class C and class S pyl tRNAs. However, C-Candidatus Methanohalalkalicoccus thermophilus 1 (Therm1) tRNA Pyl (and, to a lesser extent, C-candidate phylum MSBL1 archaeon SCGC-AAA382A20 (SCGC) tRNA Pyl ) was not well recognized by most of the active class S PylRS enzymes, but formed a high-activity pair with C Δ -NitraPylRS.

[0180] Regarding previously characterized pyl-tRNAs, see (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW: Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Willis, JCW & Chin, JW: Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018)). (https: / / doi.org:10.1038 / s41557-018-0052-5), most class C and class S PylRS enzymes use B-InttRNA Pyl It was proven to have high activity with N-MmtRNA, and in some cases produced GFP levels exceeding 80% of the wtGFP control. Furthermore, several class C and class S PylRS enzymes were found to have high activity with N-MmtRNA. Pyl and A Δ -AlvtRNA Pyl Both showed moderate to strong activity. Among the PylRS enzymes previously characterized, N + -MmPylRS exhibited the most promiscuous substrate specificity, producing GFP levels exceeding 50% of the wild type in the presence of 11 out of 16 pyl tRNAs (including all pyl tRNAs of class S except two). These included the most active class S PylRS enzyme and S-ClostRNA, which showed only moderate activity. Pyl and S-DebtRNA PylIt was included. To a relatively low degree, class A and class B PylRS enzymes also cross-reacted with certain class S and class C pyl tRNAs. Nevertheless, C-SCGCtRNA Pyl and C-Therm1 tRNA Pyl is N + -MmPylRS, a Δ -1R26PylRS, and B Δ It was pleasing to observe orthogonality with -Lum1PylRS. This confirms that the naturally occurring tRNA is orthogonal to all other PylRS enzymes obtained from other classes. Pyl It has been demonstrated that it can be found.

[0181] Two wild-type class S PylRS enzymes, S + - Gemmatimonas bacteria (Gem) PylRS and S + -DebPylRS was expressed and showed conclusive activity. + -GemPylRS is its S Δ tRNA similar to the variant Pyl It exhibited specificity. However, the most active S was evaluated. + System S + -DebPylRS is S Δ - It exhibits a significantly different activity profile from DebPylRS, for example, A-AlvtRNA Pyl Activity was significantly higher with C-TrontRNA (72% vs. 2% respectively compared to wtGFP control), but C-TrontRNA activity was also significantly higher. Pyl The activity was significantly lower (10% vs. 64%). This is because the PylSn protein was not tRNA PylThis is consistent with reports that it regulates specificity (Suzuki, T. et al. Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 13, 1261 (2017). https: / / doi.org:10.1038 / nchembio.2497, Herring, S. et al. The amino-terminal domain of pyrrolysyl-tRNA synthetase is dispensable in vitro but required for in vivo activity. FEBS Lett. 581, 3197--3203 (2007). https: / / doi.org:10.1016 / j.febslet.2007.06.004). [Examples]

[0182] Mutually orthogonal PylRS / tRNA Pyl pair The inventors then investigated whether any of the pyrrolidine systems they discovered naturally form sets of mutual orthogonality. First, they defined a criterion for mutual orthogonality pairs by referring to the interactions between them. The network of interactions between multiple aaRS / tRNA pairs can be represented as a matrix where each element xij is the activity of the aaRS protein in column j and the tRNA in row i (in this case, measured by the production of GFP(150AllocK)His6 in the presence of aaRSj and tRNAi). If aaRSi represents a congener of tRNAi, all diagonal elements xii represent the activity of the pair that is to be maximized. All off-diagonal elements represent the cross-reactivity between non-congeneral aaRS and tRNA that should be minimized. Thus, a diagonal interaction matrix of order N represents a set of N pairs that are perfectly orthogonal.

[0183] The inventors have determined that, in order to exclude sets of pairs with unacceptably low activity or unacceptably high cross-reactivity, the activity of a congeneral pair must be higher than 40% of wild-type GFP production, but each cross-reactivity (tRNA belonging to a different pair) is important. Pyl It was determined that the ratio between PylRS enzyme and tRNA should be lower than 20% of wild-type GFP production. Pyl Since pairs composed of these may have activity equal to or greater than that of the corresponding homologous pair, the inventors have identified heterologous PylRS / tRNA as possible homologous pairs in their search. Pyl This included combinations of the following. Furthermore, the inventors defined a new metric as the quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity. Hereafter referred to as the "orthogonality coefficient" (oc). This metric is a quantitative indicator of the mutual orthogonality between sets of aaRS / tRNA pairs. The previously characterized triple orthogonal pairs (N) used for the incorporation of three different non-standard amino acids. + -MmPylRS / N-SpetRNA Pyl , A Δ -1R26PylRS / A-AlvtRNA Pyl-8 B Δ -Lum1PylRS / B-InttRNA Pyl-17C10 While the formula has an oc of approximately 5.0 (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x), the inventors believed that a lower cutoff (oc greater than 2.5) would be more useful in initial screening, as mutual orthogonality can be improved through further modification.

[0184] In our initial search, we considered an interaction matrix to be sufficiently orthogonal if (i) all diagonal elements are greater than 40% of the wtGFP control, (ii) all off-diagonal elements are less than 20% of the wtGFP control, and (iii) the quotient obtained by dividing the smallest diagonal element by the largest off-diagonal element is greater than 2.5.

[0185] Forty-six double orthogonal pairs were identified (hereinafter referred to as "doublets"). The highest oc value (oc) of the doublet was 15.7. Since many doublets contain the same two PylRS enzymes and differ only in the pyl tRNAs they use, the inventors grouped these doublets into families that use the same two PylRS enzymes. In this way, the inventors obtained 15 doublet families (Figure 2c). All but one family contain a new class C or class S PylRS enzyme. Similarly, the inventors obtained two triple orthogonal pairs (i.e., "triplets") from the same family. The highest oc value of the triplet was 2.9 (Figure 2d).

[0186] These families include PylRS / tRNA Pyl This provides important insights into the activity profile. As expected, C-Therm1tRNA is a highly orthogonal class C pyl tRNA. Pyl and C-SCGCtRNA Pyl is any of the tRNAs Pyl to C Δ -When paired with NitraPylRS, it appears as a doublet. However, these pyl tRNAs are each S Δ -I2PylRS and S Δ -When paired with ClosPylRS, it also forms an unexpected interclass doublet family. This doublet is N + -MmPylRS pair (e.g., N + -MmPylRS / S-SpitRNA Pyl ) forms part of a triplet family.+ or S Δ Contains PylRS enzyme (e.g., S + -DebPylRS and S Δ -DebPylRS, or S Δ -DebPylRS and S Δ -Including I2PylRS) Further doublet families include S + PylRS enzyme and their S Δ Not only the differences between variants, but also different S Δ The differences between PylRS variants are also shown. Therefore, the inventors of S Δ We do not consider the PylRS enzyme as a separate class, but rather as a synthetically induced PylRS variant that expands the ΔN group (Figure 2e).

[0187] The relationships between the five PylRS classes can be explained based on the double orthogonality pairs formed by a representative PylRS enzyme of each class and an appropriate pyl-tRNA (which may or may not belong to the same class). In five of these ten possible interclass relationships, the inventors have found that representative PylRS / RNAs are mutually orthogonal. Pyl A pair was obtained (Figure 2f). For the remaining 5 combinations, the PylRS / tRNA satisfies the criterion of mutual orthogonality. Pyl There were no pairs. Two cases where there is no mutual orthogonality between the common PylRS classes R1 and R2 are (1) "bidirectional cross-reactivity," i.e., R1-PylRS / Ti-tRNA Pyl and R2-PylRS / Tj-tRNA Pyl For two pairs of morphologies (where Ti and Tj are arbitrary tRNA classes), R1-PylRS / Tj-tRNA Pyl and R2-PylRS / Ti-tRNA Pyl (1) Both cross-reactivity is too high (i.e., off-diagonal elements in the interaction matrix are greater than 20% of the wtGFP control or result in an oc of less than 2.5), (2) "one-way cross-reactivity", i.e., R1-PylRS / Tj-tRNA Pyl or R2-PylRS / Ti-tRNA PylPairs of R1-PylRS / Ti-tRNAs where only one of the cross-reactivity levels is too high (i.e., only one off-diagonal element in the interaction matrix results in an oc greater than 20% or less than 2.5 compared to the wtGFP control). Pyl and R2-PylRS / Tj-tRNA Pyl It can be defined that such a relationship exists. Notably, all five non-orthogonal interclass relationships fall into the second class. Therefore, of the 25 pairwise singularities required to construct a set of quintuple orthogonal pairs (using one PylRS from each of the five classes), our computational method was able to resolve 20 of them, i.e., all five congeneral interactions and 15 of the 20 non-congeneral interactions. [Examples]

[0188] Elimination of interclass cross-reactivity As a starting point for generating quintuple orthogonal pairs, the inventors have identified five specific PylRS / tRNAs. Pyl The interclass interaction matrices of pairs were examined (Figures 3a-b). For classes N, A, and B, N + -MmPylRS / N-MmtRNA Pyl , A Δ -1R26PylRS / A-AlvtRNA Pyl , and B Δ -Lum1PylRS / B-InttRNA PylSince the pair was the starting point for the previously reported triplet (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x), the inventors used these. For class C, the natural starting point is C Δ -NitraPylRS / C-Therm1tRNA Pyl This is the most active class C pair, and its tRNA Pyl This is because it is orthogonal to all other PylRS classes. For class S, N + -The diverse PylRS activity profiles of MmPylRS and their high cross-reactivity with class S pyl tRNA make it more difficult to define a good starting point. The inventors simply paired the most active wild-type class S PylRS enzyme with S + -DebRS / S-SpitRNA Pyl The following were selected. Of the 20 possible interclass synthetase / tRNA interactions (off-diagonal elements) in this matrix, only 9 were low enough to satisfy the initial criterion for cross-reactivity (less than 20% of wild-type GFP production levels). To make the other 11 interactions below this threshold, the inventors attempted to replace the tRNAs involved in unwanted cross-reactivity (off-diagonal interactions) with more orthogonal variants, i.e., to rearrange the rows of the interaction matrix so that off-diagonal elements are progressively eliminated.

[0189] tRNA of class N that is orthogonal to all other classes PylTo find it, the inventors screened seven previously reported pyl tRNAs (Figures 3c-d) derived from homologous class N Pyl systems (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x). Fortunately, N-methanococcoides methylutens (Met)tRNA Pyl and N-Methanococcoides burtonii (Bur) tRNA Pyl This results in activity of approximately 10% or less compared to PylRS enzymes of classes A, B, C, and S, and N-MmtRNA Pyl and N + - It retained over 85% of the activity with MmPylRS.

[0190] Class A tRNAs that are orthogonal to all other classes Pyl To find this, the inventors of the present invention have identified 10 previously reported A-AlvtRNAs Pyl Modified variants (Figures 3e-f) (Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5) were screened. Fortunately, two pyl tRNAs (A-AlvtRNA) were found. Pyl-17 and A-AlvtRNA Pyl-21) produces less than 10% of wild-type GFP levels in the presence of PylRS enzymes of classes N, B, C, and S, and A Δ It retained over 70% of the activity compared to -1R26PylRS.

[0191] Class B tRNAs that are orthogonal to all other classes Pyl To find this, the inventors of the present invention have identified seven previously reported B-InttRNAs Pyl Variants (Figure 3g~h) (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x) were screened. However, while orthogonality to the PylRS protein of classes N, A, and S was obtained, all pyl tRNAs tested were C Δ -In the presence of NitraPylRS, significant levels of GFP production (over 40%) were induced. B-InttRNA Pyl The variant has parental tRNA in both the acceptor stem and the variable loop. Pyl Because it has a wide range of variations in origin, the inventors of the present invention have found that C Δ -NitraPylRS, B-InttRNA Pyl It is possible to recognize multiple identity elements in B, and therefore, Δ We hypothesized that unless recognition by -Lum1PylRS is also disrupted, it would be difficult to easily neutralize that interaction.

[0192] Overall, this screening further examines nine different interclass PylRS / tRNA PylThis led to the inactivation of the interaction (Figure 3i). This resulted in the interaction between class C PylRS and class B tRNA. Pyl The interaction between, and the PylRS of class N and the tRNA of class S Pyl Only two undesirable interactions remained: the interaction between the two. In particular, three of the five tRNAs now satisfied all orthogonality requirements. Our results demonstrate that evolutionary differences in PylRS sequences can be utilized to a considerable extent to generate orthogonal interactions. [Examples]

[0193] Quadruple orthogonal PylRS / tRNA Pyl pair The inventors have identified a fourth PylRS / tRNA Pyl We investigated additional techniques to eliminate interclass (off-diagonal) interactions between pairs and form mutually orthogonal quadruplets.

[0194] As described above, the modified S Δ The PylRS variant belongs to the extended ΔN group (Figure 2e). Different S Δ The variant's activity is similar to the activity of different ΔN classes. For example, S Δ -I2PylRS is C-Therm1 tRNA Pyl They also exhibit activity (Figure 2b; C Δ - (very similar to NitraPylRS), S Δ -ClosPylRS is a modified B-InttRNA Pyl It shows activity along with the variant (Figure 3g; B) Δ -Lum1PylRS is very similar). Therefore, the inventors have identified a pair of class B or C that yields the desired B-like or C-like activity. Δ We hypothesized that the problem of cross-reactivity between class C PylRS enzymes and class B pyl tRNAs could be resolved by substituting with pairs containing PylRS (Figure 4a). These substituted PylRS variants were then assigned to S ΔB or S ΔCThis is called [a specific term]. From the perspective of the interaction matrix, this required a permutation of columns such that off-diagonal interactions were eliminated.

[0195] Class B or C PylRS enzyme S Δ After replacing with the PylRS variant, a total of 946 doublets were obtained in 25 families, 1425 triplets in 16 families, and, importantly, 96 quadruplets in 4 families (Figures 4b-c). Notably, the highest oc value of the triplet was 24.5, compared to the previously reported triplet (N) used for incorporating three different non-standard amino acids. + -MmPylRS / N-SpetRNA Pyl , A Δ -1R26PylRS / A-AlvtRNA Pyl-8 B Δ -Lum1PylRS / B-InttRNA Pyl-17C10 ) is approximately 5 times higher than previously reported N + -MmPylRS, a Δ -1R26PylRS, B Δ -The highest oc triplet from the Lum1PylRS family (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x) is approximately twice as high. Most triplet families and all quadruplet families have a pair of class B and / or class C, S Δ Includes substitution with pairs containing the PylRS variant.

[0196] S ΔTo understand how much the PylRS substitution strategy has advanced the development of quintuple orthogonal pairs, we examined a novel quadruplet in terms of an interclass interaction network (Figures 4d-h). Δ [A, S] are formed by substituting class B or class C PylRS in PylRS variants, respectively. ΔB , C, S](oc3.9) and [A, B, S ΔC In the quadruplet family of [S](oc2.5), cross-reactivity between class B and C (Figure 3h) is observed between class N and S ΔB Alternatively, it was effectively replaced by cross-reactivity with B (Figures 4d-f). As a result, a diagonal submatrix of order 4 in the overall interaction matrix, and consequently a quadruplet, is generated. However, class N(N + -MmPylRS) and S(S + - Any tRNA paired with DebPylRS Pyl Including cross-reactivity between ), two types of cross-reactivity remained uneliminated. On the other hand, when both class B and class C PylRS enzymes were used, Δ [N, A, S] formed by substitution with PylRS variants. ΔB S ΔC ] and [A, S ΔB S ΔC In the quadruple presets of [S] (all oc2.9), there is one major cross-reactivity (N + -MmPylRS and S + - Any tRNA paired with DebPylRS Pyl (between) remained (Figure 4g~h). However, Class S ΔB and S ΔC Among the 22 possible pairs, there were still some unresolved cross-reactivity issues that would limit the potential oc of the quintoplets. The inventors believe that a general solution to all of these cross-reactivity issues is Class B / S ΔB We hypothesized that this was a further modification of pyl tRNA paired with the class S PylRS enzyme. [Examples]

[0197] Quintuple orthogonal PylRS / tRNA Pyl pair The inventors of this invention have identified (1) a class B PylRS enzyme (or S ΔB tRNA that functions with PylRS variants but is orthogonal to all other PylRS classes. Pyl (2) tRNAs that function with PylRS of class S but are orthogonal to all other PylRS classes. Pyl To discover (3) the quadruplet with the highest oc, S ΔB -ClosPylRS and S + The aim was to replace the pyl tRNA paired with DebPylRS with a new pyl tRNA. The inventors predicted that this would reduce unwanted cross-reactivity with the fifth pair, and as a result enable the generation of a quintuple orthogonal set of pairs (Figure 5a).

[0198] C Δ -NitraPylRS is a modified B-tRNA Pyl Since it was observed that the variant exhibited activity together (Figure 3g), the inventors identified different parental tRNAs, such as those derived from the bacterial class S system. Pyl However, due to directional evolution, C Δ -Class B specific or S-type enzymes that are orthogonal to NitraPylRS (and other PylRS enzymes as needed) ΔB Specific tRNA Pyl We assumed this could be a better starting point for discovering variants.

[0199] The present inventors have identified bacterial tRNA Pyl S-I2tRNA Pyl is, S ΔB -It has excellent activity with ClosPylRS (producing GFP levels of 93% of the wild type), C Δ -Because cross-reactivity with NitraPylRS is quite low (23% GFP level of wild type), this was selected as a starting point for oriented evolution. In fact, S-I2tRNA Pylis N + -It showed the highest cross-reactivity with MmPylRS (68% of the wild-type GFP level), and therefore, the inventors believe that N + - S-I2tRNA by MmPylRS Pyl We focused our efforts on invalidating that perception.

[0200] From previous efforts, N + -MmPylRS is more efficient than the ΔN PylRS enzymes of class A7 and B6 in tRNA Pyl It has been demonstrated that it can be significantly more sensitive to elongation in short variable loops, which is consistent with structural and biochemical studies. (Suzuki, T. et al. Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 13, 1261 (2017). https: / / doi.org:10.1038 / nchembio.2497, Herring, S. et al. The amino-terminal domain of pyrrolysyl-tRNA synthetase is dispensable in vitro but required for in vivo activity. FEBS Lett. 581, 3197--3203 (2007). https: / / doi.org:10.1016 / j.febslet.2007.06.004) In fact, (unlike the parental pyl tRNA) N + A-AlvtRNA orthogonal to MmPylRS Pyl and B-InttRNA PylThe variants (Figures 3e and 3g) were obtained from a library of mutants in which the variable loop was extended (from 3 bases) to 4, 5, or 6 randomized nucleotides. (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x, Willis, JCW & Chin, JW Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5) For example, insertion of a single uracil in a variable loop enables the parental tRNA Pyl Different B-InttRNA Pyl-B03 This is parental tRNA Pyl Compared to N + - Its activity with MmPylRS is less than 1 / 30th.

[0201] The inventors of the present invention, S ΔB -ClosPylRS is an artificial ΔN class PylRS, therefore S-I2tRNA Pyl Based on the hypothesis that insertion into the variable loop might be permitted, we extended the variable loop to 4 nucleotides in S-I2tRNA. Pyl A library of mutants (Figure 5b) was synthesized. Positions 8 and 22, which can form important tertiary contact points with the variable loop, were also randomized. Furthermore, to increase the diversity of activity profiles within the library, S-I2tRNA was randomized. PylThree bases generally considered important for recognition were also randomized: position 69 (a key identity element of PylRS), and positions 5 and 64. The latter two are S-I2tRNA. Pyl They form fluctuating base pairs in the acceptor stem. Such pairs are known to control aminoacyl-tRNA synthetase recognition by distorting the RNA helix. (Varani, G. & McClain, WH The G·U wobble base pair. EMBO reports 1, 18-23 (2000). https: / / doi.org:10.1093 / embo-reports / kvd001)

[0202] The inventors have enabled the production of a protein from a gene encoding chloramphenicol acetyltransferase containing an amber codon at position 111, ΔB Cells expressing -ClosPylRS were grown in 100 μg mL-1 of chloramphenicol in the presence of AllocK 1, and S-I2tRNA was also observed. Pyl A mutant was selected. The inventors then selected the S-I2tRNA. Pyl Perform sequential negative screening on the variant, N + -MmPylRS, a Δ -1R26PylRS, C Δ -NitraPylRS, and S + - We identified pyl tRNAs with minimal cross-reactivity with DebPylRS: GFP(150TAG)His6 and S-I2tRNA. Pyl Variant, and N + -MmPylRS, a Δ -1R26PylRS, C Δ -NitraPylRS, or S + AllocK was supplied to cells containing one of the -DebPylRS strains, and the absence of GFP expression was screened. These screenings revealed the absence of S-I2tRNA Pyl-B8 S-I2tRNA Pyl-B32 , and S-I2tRNAPyl-B72 The three library members, S ΔB -Activity with ClosPylRS is maintained at a high level, but N + -MmPylRS, a Δ -1R26PylRS, C Δ -NitraPylRS, and S + -It was revealed that there is almost no cross-reactivity with DebPylRS (Figure 5c-d). In particular, S-I2tRNA Pyl-B32 and S-I2tRNA Pyl-B72 In all cases, their parental S-I2tRNA Pyl and S ΔB - It retains more than 75% of the activity with ClosPylRS, but N + -Activity with MmPylRS is less than 1 / 20th compared to their parental tRNAs, C Δ -The activity with NitraPylRS was about one-third that of their parents.

[0203] S-I2tRNA Pyl When negative screening was performed on mutants, the inventors identified several S-I2tRNAs. Pyl The mutant, despite those variable loops being stretched, S + We observed increased activity with -DebPylRS. The inventors found that the divided S + -The N-terminal domain protein in the DebPylRS enzyme is N + We hypothesized that it might recognize variable loop nucleotides differently from PylRS. Therefore, we paired the tRNA of class S with PylRS, which is orthogonal to all other PylRS classes used in the quintoplet. Pyl In order to discover S + S-I2tRNA with an extended variable loop that is selectively aminoacylated by DebPylRS. Pyl I selected a member of the Mutant Library.

[0204] The inventors of the present invention, S +-S-I2tRNA that exhibits activity together with DebPylRS Pyl After performing a positive selection for mutants, N + -MmPylRS, a Δ -1R26PylRS, S ΔB -ClosPylRS, and C Δ Negative screening was performed to minimize cross-reactivity with NitraPylRS. These screenings were conducted to identify S-I2tRNA Pyl-S52 One library member was identified as S. + -DebPylRS shows activity (43% of wild-type GFP production), S Δ -ClosPylRS, N + -MmPylRS, a Δ -1R26PylRS, and C Δ -NitraPylRS has very low activity (all less than 2% of wild-type GFP production) (Figure 5e-f). Wild-type parental tRNA is S-I2tRNA. Pyl In comparison, S-I2tRNA Pyl-S52 is, S + -DebPylRS shows more than 3 times the activity, S ΔB -ClosPylRS exhibits 1 / 120th the activity.

[0205] S-I2tRNA evolved by the inventors Pyl Replace the mutant and select the quadruplet ([A, S ΔB By incorporating it into the [C, S] family, the inventors have found that N + - Minimized cross-reactivity with MmPylRS. This allowed the updated quadruplet to use a fifth orthogonal pair of class N (e.g., N + -MmPylRS / N-BurtRNA Pyl By combining this with the above, it became possible to generate a family of mutually orthogonal quintoplets with oc values ​​up to 4.0 (Figure 5f). In this way, the inventors eliminated the last two types of cross-reactivity by mutation, selection, and screening from a single tRNA scaffold (Figure 5g). S-I2tRNAPyl-B32 and S-I2tRNA Pyl-S52 Despite their vastly different activities, these molecules differ by only four nucleotides. This demonstrates the influence of synthetically modifying identity elements to control tRNA recognition by aminoacyl-tRNA synthetase.

[0206] S-I2tRNA Pyl-S52 is B Δ -Lum1PylRS and S ΔC Since it is orthogonal to both -I2PylRS, the inventors have found that S ΔB -ClosPylRS to B Δ Replace with -Lum1PylRS, or C Δ -NitraPylRS to S ΔC Substitution with -I2PylRS yielded two more quintoplet families. Overall, using our original oc threshold of 2.5, we obtained 1136 doublets in 27 families (maximum oc 72.0), 2359 triplets in 26 families (maximum oc 39.6), 919 quadruplets in 14 families (maximum oc 7.2), and 90 quintoplets in 3 families (maximum oc 5.4). Using an oc threshold of 5.0 (PylRS / tRNA previously used to incorporate three different non-standard amino acids) Pyl By increasing the oc value (equivalent to that of a pair), the inventors obtained 924 doublets in 22 families, 1324 triplets in 18 families, 128 quadruplets in 7 families, and 8 quintoplets in 1 family (Figures 5h-i). The quintoplet with the highest oc value has the following composition: N + -MmPylRS and (non-homogeneous) natural class N tRNA Pyl , A Δ A-AlvtRNA modified to -1R26PylRS Pyl Variant, B Δ - Lum1PylRS modified S-I2tRNA Pyl Variant, CΔ -NitraPylRS and (non-homogeneous) natural class C tRNA Pyl , and S + -DebPylRS and modified S-I2tRNA Pyl Variant (Figure 5j). The inventors produced GFP150AllocKHis6 and Ub11AllocKHis6 from GFP150TAGHis6 and Ub11TAGHis6, respectively, and measured protein titer and MS spectrum to determine the quintuple orthogonal PylRS / tRNA in the most orthogonal quintoplet. Pyl The efficiency and precision of amber suppression in each pair were characterized. Our results show that homologous PylRS / tRNA Pyl This study demonstrates, for the first time in history, the successful classification of systems into five classes (Figure 5k) that are mutually orthogonal in terms of aminoacylation specificity. [Examples]

[0207] Consideration The inventors of this invention have identified PylRS and tRNA PylWe defined sequence identity threshold criteria to effectively search for orthogonality from genomic data. By applying these thresholds to generate sequence clusters, we computationally searched PylRS sequences for multiple orthogonal systems. Using this method, we identified orthogonal systems from hundreds of PylRS systems spanning all PylRS groups. By combining our computational method with directed evolution and modification, we generated the first quadruple orthogonal and quintuple orthogonal PylRS systems. These advances, along with strategies for generating codons that can be used to encode non-standard monomers, will be central to the synthesis of proteins containing an increased number of ncAAs and the synthesis of more diverse encoded polymers and macrocyclic molecules in cells. (De La Torre, D. & Chin, JW Reprogramming the genetic code. Nature Reviews Genetics 22, 169-184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7, Robertson, WE et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057-1062 (2021). https: / / doi.org:10.1126 / science.abg3029, Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, JW Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441 (2010). https: / / doi.org:10.1038 / nature08817, Wang, K. Et al.Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393 (2014). https: / / doi.org:10.1038 / nchem.1919、Dunkelmann, D. L., Oehm, S. B., Beattie, A. T. & Chin, J. W. A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA design. Nature Chemistry 13, 1110-1117 (2021). https: / / doi.org:10.1038 / s41557-021-00764-5、Malyshev, D. A. et al. A semi-synthetic organism with an expanded genetic alphabet. Nature 509, 385-388 (2014). https: / / doi.org:10.1038 / nature13314、Fischer, E. C. et al. New codons for efficient production of unnatural proteins in a semisynthetic organism. Nature Chemical Biology 16, 570-576 (2020). https: / / doi.org:10.1038 / s41589-020-0507-z、Zhang, Y. et al. A semi-synthetic organism that stores and retrieves increased genetic information. Nature 551, 644 (2017). https: / / doi.org:10.1038 / nature24659、Fredens, J. et al.(Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514–518 (2019). https: / / doi.org:10.1038 / s41586-019-1192-5; Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59 (2016). https: / / doi.org:10.1038 / nature20124) Furthermore, mining the data generated through our method may provide further insights into the sequence requirements of mutually orthogonal systems, potentially enabling the creation of more sophisticated rules for predicting orthogonality. [Examples]

[0208] method Identification of the PylRS sequence The inventors performed a BLAST search against the NCBI non-redundant protein sequence database, and A Δ - Using the AlvPylRS protein as the query sequence, 1 × 10 -30 PylRS sequences were identified by filtering for expected values ​​less than . After manually checking and removing partial protein sequences and synthetic construct sequences, the inventors ultimately obtained 351 PylRS sequences. The inventors aggregated these sequences into a database and assigned each PylRS sequence a unique identifier based on the biological name and NCBI entrusted ID. The inventors aligned the obtained PylRS sequences using Clustal Omega (Sievers, F. et al. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Molecular Systems Biology 7, 539 (2011). https: / / doi.org:10.1038 / msb.2011.75).+ -By referring to known annotations of the C-terminal domain (CTD) in MmPylRS6, sequences corresponding to the CTD of the PylRS enzyme were extracted.

[0209] Identification of tRNAPyl sequences Using the NCBI nucleotide database, the inventors obtained metagenomic reads containing either all nucleotide sequences available in the host genome or PylRS sequences identified by the inventors. The inventors ran the tRNA detection program ARAGORN (Laslett, D. ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences. Nucleic Acids Research 32, 11-16 (2004). https: / / doi.org:10.1093 / nar / gkh152) (version 1.2.38) on each nucleotide sequence, allowing introns consisting of up to 100 nucleic acid bases, and setting the scoring threshold to the default level of 90%. The discovered pyl tRNAs were added to the PylRS sequence database. The PylRS sequence (isolated from the archaeon SCGC-AAA382A20, candidate phylum MSBL1, hereafter referred to as C) was labeled with identifier Marc.6481. ΔRegarding the sequence referred to as -SCGCPylRS, ARAGORN was unable to find a corresponding tRNA, but since the putative sequence had been previously reported (Guan, Y., Haroon, MF, Alam, I., Ferry, JG & Stingl, U. Single-cell genomics reveals pyrrolysine-encoding potential in members of uncultivated archaeal candidate division MSBL1. Environmental Microbiology Reports 9, 404-410 (2017). https: / / doi.org:10.1111 / 1758-2229.12545), it was added to the database. On the other hand, the PylRS sequences with identifiers Mthe.7552 and Mthe.9096 were found to originate from the same organism (Candidatus metanohalarcaeum thermopilum), but ARAGORN was unable to find a corresponding tRNA for one of these PylRS sequences (Mthe.7552, hereafter C). Δ -Therm1PylRS (a corresponding tRNA) Pyl Only was found. By manually searching the nucleotide sequence, the other PylRS sequence (Mthe.9096, hereafter C) was found. Δ A second tRNA that is close to (called -Therm2PylRS) PylThis was revealed. This was also added to the database. During the preparation of this paper, these two pyl tRNAs were reported independently. (Zhang, H. et al. The tRNA discriminator base defines the mutual orthogonality of two distinct pyrrolysyl-tRNA synthetase / tRNAPyl pairs in the same organism. Nucleic Acids Research 50, 4601-4615 (2022). https: / / doi.org:10.1093 / nar / gkac271) After removing pseudogenes by manual curation, the inventors obtained pyl tRNAs for 284 different PylRS genes.

[0210] Previously characterized PylRS / tRNA Pyl Pair analysis PylRS and tRNA Pyl From a sequence database, the inventors identified the sequences and tRNAs of class A and B PylRS that they had previously experimentally characterized (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x). PylThe sequence was obtained. Then, using Python (version 3.9.7) (Cock, PJA et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422-1423 (2009). https: / / doi.org:10.1093 / bioinformatics / btp163), all pairs of PylRS sequences and tRNAs were obtained. Pyl A matrix of pairwise sequence identity percentages was calculated for all pairs. For PylRS sequences, the identity percentage was calculated from the multiple sequence alignment of the C-terminal domain. tRNA Pyl For the sequences, the identity percentage was calculated by manually performing multiple sequence alignment based on secondary structure prediction from ARAGORN and RNAfold (Gruber, AR, Lorenz, R., Bernhart, SH, Neubock, R. & Hofacker, IL The Vienna RNA Websuite. Nucleic Acids Research 36, W70-W74 (2008). https: / / doi.org:10.1093 / nar / gkn188, Lorenz, R. et al. ViennaRNA Package 2.0. Algorithms for Molecular Biology 6, 26 (2011). https: / / doi.org:10.1186 / 1748-7188-6-26).

[0211] The inventors have previously reported experimental activity data for these sequences, namely, PylRS enzyme and tRNA Pyl CUA Each combination, and 8mM N ε GFP production level from a GFP gene (GFP150TAGHis6) containing an in-frame amber codon at position 150, obtained in the presence of -Boc-L-lysine (tRNA by endogenous aaRS) PylCUA Due to aminoacylation in the background, the GFP gene and tRNA Pyl CUA We considered the GFP production levels (subtracted from the levels in the presence of only PylRS enzyme and tRNA). Pyl CUA For each combination, the inventors have determined that this activity is related to the PylRS sequence and tRNA Pyl CUA The sequence identity percentage was plotted against the sequence of PylRS from the same organism. Similarly, the inventors plotted the sequence identity of the PylRS enzyme against tRNA. Pyl CUA The activity in each combination is determined by tRNA Pyl CUA And, tRNA of the same biological origin as the PylRS enzyme Pyl CUA The sequence was plotted against the percentage of identity with the given sequence.

[0212] Clustering of the C-terminal domain sequence of PylRS Using Python (version 3.9.7), the inventors calculated an identity percentage matrix from the multiple sequence alignment of the C-terminal domain for all pairs of PylRS sequences in the database. Next, using this matrix, we performed unweighted average concatenation clustering (UPGMA) of aligned PylRS CTD sequences using the Python libraries biopython (version 1.79) and scikit-learn (version 1.0.1) (Cock, PJA et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422-1423 (2009). https: / / doi.org:10.1093 / bioinformatics / btp163, Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12, 2825-2830 (2011)), with a clustering threshold of 55% sequence identity.

[0213] tRNA Pyl Sequence alignment and clustering For each PylRS cluster, the inventors have identified the corresponding tRNA Pyl We selected representative PylRS sequences from which we were able to find sequences. Of the 37 PylRS clusters, two were identified as tRNA Pyl No sequences were found. These were excluded from further analysis. For the 35 pyl tRNAs obtained, the inventors performed manual multiplex sequence alignment based on secondary structure prediction from ARAGORN and RNAfold. Then, using this multiplex sequence alignment, the selected tRNAs were analyzed using the biopython (version 1.79) and scikit-learn (version 1.0.1) Python libraries. PylA matrix of identity percentages was calculated for all pairs. The inventors then used this matrix to select aligned tRNAs with a cluster integration threshold of 75% sequence identity. Pyl Agglomerative hierarchical clustering (UPGMA) was performed on the array using the unweighted average concatenation method.

[0214] DNA construct PylRS and tRNA PylThe genes were synthesized as gBlock double-stranded DNA fragments by IDT. The inventors cloned all novel pyl tRNAs into a minimal pMB1 backbone under the lpp promoter. Previously reported pyl tRNAs were used in the same format. Certain tRNAs differed from the standard sequence in the anticodon loop. These positions were mutated to consensus bases found in E. coli to improve the efficiency of tRNA translation in E. coli, as previously reported (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x). PylRS sequences were cloned into the p15A backbone under the glnS promoter. For each novel PylRS sequence, a 5' untranslated region was generated using the online tool De Novo DNA (Salis, HM, Mirsky, EA & Voigt, CA Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009). https: / / doi.org:10.1038 / nbt.1568), which is predicted to maximize translation initiation efficiency, and inserted between the +1 site of the glnS promoter and the gene's start codon. For class S PylRS enzymes, a polycistronic operon consisting of separately expressed N-terminal and C-terminal domains was constructed using the intergenetic region predicted by De Novo DNA. The optimal arrangement of the two domains was selected by maximizing the predicted translation initiation rate.+ MmPylRS, A Δ -AlvPylRS, and B Δ -Lum1PylRS was used in a similar p15A construct containing a C-terminal tag, as previously described (Dunkelmann, DL, Willis, JCW, Beattie, AT & Chin, JW Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x). The p15A vector also encoded the chloramphenicol acetyltransferase (CAT) gene with an amber codon at position 111 under a constitutive cat promoter, and the GFP gene with an amber codon at position 150 under an L-arabinose-inducible pBAD promoter.

[0215] PylRS / tRNA Pyl Measurement of CUA pair activity and specificity PylRS / tRNA Pyl To measure the activity of the CUA pair, the inventors added tRNA to 4-10 μL of E. coli DH10B chemically competent cells carrying the p15A plasmid encoding the PylRS gene, CAT111TAG gene, and GFP150TAGHis6 gene. Pyl0.4 μL of the pMB1 plasmid encoding the CUA gene was transduced. The inventors recovered the transduced cells in 180 μL of SOC medium (Super Optimal Broth containing Catabolite Repression Medium) in a 96-well Costar microtiter plate format at 37°C and 750 r.pm for approximately 1 hour. Next, the inventors seeded 40 μL of the rescued cells in 760 μL of selective 2xYT-st medium (2xYT medium containing 75 μg mL-1 of spectinomycin and 12.5 μg mL-1 of tetracycline) in a 1.2 mL 96-well plate format, and grew the culture overnight at 37°C and 750 r.pm. After a minimum of 16 hours, 40 μL of the overnight culture was seeded in 760 μL of 2xYT-st medium containing 0.05% L-arabinose and 4 mM Nε-Alloc-L-lysine (AllocK) in a 1.2 mL 96-well plate format. Cells were grown at 37°C and 750 r.pm for 18–24 hours. Finally, 100 μL of each culture was transferred to a 96-well flat-bottom Costar plate, and fluorescence and optical density (OD) were measured using a PHERAstar FS plate reader. Measured GFP OD 600 -1 The value is the GFP OD600 of cells expressing GFP from the GFP150AsnHis6 gene (referred to as "wtGFP control"). -1 The values ​​were normalized.

[0216] Mutually orthogonal PylRS / tRNA Pyl Identifying pairs Using Python (version 3.9.7), the inventors determined mutually orthogonal PylRS / tRNA based on GFP activity data. Pyl The paired sets were identified. PylFor pairs, the orthogonality coefficient oc was defined as the quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity. A pair was considered mutually orthogonal if the lowest intrapair activity was higher than 40% of the wtGFP control, the highest interpair cross-reactivity was lower than 20% of the wtGFP control, and oc was higher than 2.5. The inventors grouped mutually orthogonal pairs into families if they contained the same PylRS enzyme.

[0217] S-I2tRNA Pyl Generating a CUA library S-I2tRNA containing randomized nucleotides Pyl The CUA library was constructed by Golden Gate cloning into a pMB1 vector using PCR primers, Q5 DNA polymerase, Bbs1-HF restriction enzyme, and T4 DNA ligase (all enzymes purchased from New England Biolabs (NEB)). This library had a transduction efficiency of 1 × 10⁻¹⁶. 8 Transduction was performed into electrocompetent Escherichia coli DH10B cells, which have a higher colony-forming unit level.

[0218] Orthogonal S-I2tRNA Pyl Selection and screening to identify CUA hits CAT(111TAG), GFP(150TAG)His6, and S Δ -Clos Pyl RS or S + -Electrocompetent E. coli DH10B cells carrying a p15A plasmid encoding one of the DebPylRS were subjected to S-I2tRNA PylThe CUA library was transduced. Cells were recovered for 1 hour in 1 mL of SOC with AllocK added at 37°C and 220 rpm, and seeded onto LB agar plates containing 4 mM AllocK, 75 μg mL-1 spectinomycin, 12.5 μg mL-1 tetracycline, and 100 μg mL-1 chloramphenicol. The plates were incubated at 37°C for 18-24 hours. After incubation, a total of 576 colonies (S) were selected from a combination of one of the selected groups. Δ -ClosPylRS or S + Colonies grown in the presence of DebPylRS were picked onto 500 μL of 2xYT-st and grown overnight at 750 rpm and 37°C. Then, 40 μL of the overnight culture was transferred to 760 μL of 2xYT-st containing 0.05% L-arabinose, both in the presence and absence of 4 mM AllocK. Plasmids of clones that selectively fluoresced in the presence of AllocK were extracted (using Qiagen DNA miniprep), digested with NcoI restriction enzyme and T5 exonuclease (both from NEB), and CAT (111 TAG), GFP (150 TAG), His6, and N + -MmPylRS, S Δ -ClosPylRS, C Δ -NitraPylRS, A Δ -1R26PylRS, or S + Chemically competent cells carrying the p15A plasmid encoding one of the PylRS genes from DebPylRS were transduced again. Plasmids of clones that met the orthogonality requirements were isolated and sequenced.

[0219] Quintuple orthogonal PylRS / tRNA Pyl Quantification of GFP150AllocKHis6 and Ub11AllocKHis6 protein production using CUA pairs. Quintuple orthogonal PylRS / tRNA in the set with the highest oc Pyl To measure the yield of proteins incorporating a single ncAA using a CUA pair, the inventors used tRNAPyl 0.8 μL of the pMB1 plasmid encoding the CUA gene and 0.8 μL of the p15A plasmid encoding the PylRS gene, CAT111TAG gene, and GFP150TAGHis6 or Ub11TAGHis6 gene were co-transduced into competent Escherichia coli DH10B by electroporation. As a control, the inventors also introduced GFPHis or Ub11TCAHis6 into AlvtRNA. Pyl-21 CUA Co-transduction was performed using the same plasmid configuration as with MmPylRS.

[0220] The inventors restored transduced cells in 600 μL of SOC medium (Super Optimal Broth containing catabolite inhibitory medium) at 37°C and 220 r.pm for approximately 1 hour. Using 160 μL of the rescued cells, the inventors administered 5 mL of selective 2xYT-st medium (75 μg mL) in a 50 mL glass tube. -1 Spectinomycin and 12.5 μg mL -1 Cells were seeded in 2xYT medium containing tetracycline, and the cultures were grown overnight in a shaking incubator at 37°C and 220 r.pm. After a minimum of 16 hours, 140 μL of the overnight culture was used to seed 5 mL of 2xYT-st medium containing 0.05% L-arabinose and 4 mM Nε-Alloc-L-lysine (AllocK) into a 50 mL glass tube. The cells were grown at 37°C and 220 r.pm for 16–18 hours. The cells were centrifuged, aspirated, and the cell pellet was frozen at -20°C for a minimum of 1 hour.

[0221] The pellet was resuspended in 800 μL of BugBuster® Protein Extraction Reagent containing cOmplete® protease inhibitor and lysed for 1 hour using an up-and-down rotation motion. The lysed cells were centrifuged, and the supernatant was incubated with 160 μL of NiNTA agarose beads at 4°C for 1–16 hours. The beads were washed five times with 800 μL of pH 8.5 PBS containing 25 mM imidazole, and the proteins were eluted five times with 160 μL of pH 8.5 PBS containing 250 mM imidazole (for GFP samples) or five times with 100 μL of pH 8.5 PBS containing 250 mM imidazole (for Ub samples). The protein concentration of GFP was measured by quantifying the absorbance at 280 nm. The protein concentration of ubiquitin was measured using the Pierce® BCA Protein Assay Kit from Thermo Fisher, according to the manufacturer's protocol.

[0222] Electrospray ionization mass spectrometry Denatured protein samples (approximately 10 μM) were subjected to liquid chromatography-mass spectrometry analysis. Briefly, proteins were separated using a modified nanoAcquity (Waters) on a 1.7 μm, 1.0 × 100 mm high-performance liquid chromatography column C4 BEH (Waters), delivered at a flow rate of approximately 50 μl min⁻¹. The column was expanded over 20 minutes using a gradient of acetonitrile in 0.1% v / v formic acid (2–80% v / v). The analytical column outlet was directly connected to a hybrid quadrupole time-of-flight mass spectrometer (Xevo G2, Waters) via an electrospray ionization source. Data were acquired in cation mode with a cone voltage of 30 V over the m / z range of 300–2,000. Scans were manually combined and deconvoluted using MaxEnt1 (Masslynx, Waters). The theoretical molecular weight of the protein containing ncAA was calculated by first calculating the theoretical molecular weight of the wild-type protein using an online tool (http: / / web.expasy.org / protparam / ), and then manually correcting the theoretical molecular weight of ncAA. [Examples]

[0223] PylRS table for clustering analysis

[0224] [Table 1] JPEG2026518309000003.jpg252170 JPEG2026518309000004.jpg252170 JPEG2026518309000005.jpg247170 JPEG2026518309000006.jpg250170 JPEG2026518309000007.jpg252170 JPEG2026518309000008.jpg252170 JPEG2026518309000009.jpg252170 JPEG2026518309000010.jpg78170

[0225] (References) 1 Chin, J. W. Expanding and reprogramming the genetic code. Nature 550, 53 (2017). https: / / doi.org:10.1038 / nature24031 2 De La Torre, D. & Chin, J. W. Reprogramming the genetic code. Nature Reviews Genetics 22, 169 - 184 (2021). https: / / doi.org:10.1038 / s41576-020-00307-7 3 Robertson, W. E. et al. Sense codon reassignment enables viral resistance and encoded polymer synthesis. Science 372, 1057 - 1062 (2021). https: / / doi.org:10.1126 / science.abg3029 4 Spinck, M. et al. Genetically programmed cell-based synthesis of non-natural peptide and depsipeptidemacrocycles. Nature Chemistry 5 Cervettini, D. et al. Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase-tRNA pairs. Nat. Biotechnol., 1 - 11 (2020). https: / / doi.org:10.1038 / s4158�-020-0479-2 6 Dunkelmann, D. L., Willis, J. C. W., Beattie, A. T. & Chin, J. W. Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids. Nature Chemistry 12, 535-544 (2020). https: / / doi.org:10.1038 / s41557-020-0472-x 7 Willis, J. C. W. & Chin, J. W. Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs. Nature Chemistry 10, 831-837 (2018). https: / / doi.org:10.1038 / s41557-018-0052-5 8 Srinivasan, G., James, C. M. & Krzycki, J. A. Pyrrolysine Encoded by UAG in Archaea: Charging of a UAG-Decoding Specialized tRNA. Science 296, 1459--1462 (2002). https: / / doi.org:10.1126 / science.1069588 9 Krzycki, J. A. The direct genetic encoding of pyrrolysine. Curr. Opin. Microbiol. 8, 706--712 (2005). https: / / doi.org:10.1016 / j.mib.2005.10.009 10 Neumann, H., Peak-Chew, S. Y. & Chin, J. W. Genetically encoding Nε-acetyllysine in recombinant proteins. Nat. Chem. Biol. 4, 232 (2008). https: / / doi.org:10.1038 / nchembio.73 11 Wang, L., Brock, A., Herberich, B. & Schultz, P. G. Expanding the Genetic Code of Escherichia coli. Science 292, 498--500 (2001). https: / / doi.org:10.1126 / science.1060077 12 Borrel, G. et al. Unique Characteristics of the Pyrrolysine System in the 7th Order of Methanogens: Implications for the Evolution of a Genetic Code Expansion Cassette. Archaea 2014, 11 (2014). https: / / doi.org:10.1155 / 2014 / 374146 13 Park, H.-S. et al. Expanding the Genetic Code of Escherichia coli with Phosphoserine. Science 333, 1151--1154 (2011). https: / / doi.org:10.1126 / science.1207203 14 Rogerson, D. T. et al. Efficient genetic encoding of phosphoserine and its nonhydrolyzable analog. Nat. Chem. Biol. 11, 496 (2015). https: / / doi.org:10.1038 / nchembio.1823 15 Hughes, R. A. & Ellington, A. D. Rational design of an orthogonal tryptophanyl nonsense suppressor tRNA. Nucleic Acids Research 38, 6813-6830 (2010). https: / / doi.org:10.1093 / nar / gkq521 16 Chatterjee, A., Sun, S. B., Furman, J. L., Xiao, H. & Schultz, P. G. A Versatile Platform for Single- and Multiple-Unnatural Amino Acid Mutagenesis in Escherichia coli. American Chemical Society (2013). https: / / doi.org:10.1021 / bi4000244 17 Italia, J. S. et al. Mutually Orthogonal Nonsense-Suppression Systems and Conjugation Chemistries for Precise Protein Labeling at up to Three Distinct Sites. J. Am. Chem. Soc. (2019). https: / / doi.org:10.1021 / jacs.8b12954 18 Neumann, H., Wang, K., Davis, L., Garcia-Alai, M. & Chin, J. W. Encoding multiple unnatural amino acids via evolution of a quadruplet-decoding ribosome. Nature 464, 441 (2010). https: / / doi.org:10.1038 / nature08817 19 Wang, K. et al. Optimized orthogonal translation of unnatural amino acids enables spontaneous protein double-labelling and FRET. Nat. Chem. 6, 393 (2014). https: / / doi.org:10.1038 / nchem.1919 20 Anderson, J. C. et al. An expanded genetic code with a functional quadruplet codon. Proceedings of the National Academy of Sciences 101, 7566-7571 (2004). https: / / doi.org:10.1073 / pnas.0401517101 21 Dunkelmann, D. L., Oehm, S. B., Beattie, A. T. & Chin, J. W. A 68-codon genetic code to incorporate four distinct non-canonical amino acids enabled by automated orthogonal mRNA design. Nature Chemistry 13, 1110-1117 (2021). https: / / doi.org:10.1038 / s41557-021-00764-5 22 Malyshev, D. A. et al. A semi-synthetic organism with an expanded genetic alphabet. Nature 509, 385-388 (2014). https: / / doi.org:10.1038 / nature13314 23 Fischer, E. C. et al. New codons for efficient production of unnatural proteins in a semisynthetic organism. Nature Chemical Biology 16, 570-576 (2020). https: / / doi.org:10.1038 / s41589-020-0507-z 24 Zhang, Y. et al. A semi-synthetic organism that stores and retrieves increased genetic information. Nature 551, 644 (2017). https: / / doi.org:10.1038 / nature24659 25 Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514--518 (2019). https: / / doi.org:10.1038 / s41586-019-1192-5 26 Wang, K. et al. Defining synonymous codon compression schemes by genome recoding. Nature 539, 59 (2016). https: / / doi.org:10.1038 / nature20124 27 Neumann, H., Slusarczyk, A. L. & Chin, J. W. De Novo Generation of Mutually Orthogonal Aminoacyl-tRNA Synthetase / tRNA Pairs. American Chemical Society (2010). https: / / doi.org:10.1021 / ja9068722 28 Beranek, V., Willis, J. C. W. & Chin, J. W. An Evolved Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase / tRNA Pair Is Highly Active and Orthogonal in Mammalian Cells. Biochemistry 58, 387-390 (2019). https: / / doi.org:10.1021 / acs.biochem.8b00808 29 Chatterjee, A., Xiao, H. & Schultz, P. G. Evolution of multiple, mutually orthogonal prolyl-tRNA synthetase / tRNA pairs for unnatural amino acid mutagenesis in Escherichia coli. Proc. Natl. Acad. Sci. U.S.A. 109, 14841--14846 (2012). https: / / doi.org:10.1073 / pnas.1212454109 30 Italia, J. S. et al. An orthogonalized platform for genetic code expansion in both bacteria and eukaryotes. Nature Chemical Biology 13, 446-450 (2017). https: / / doi.org:10.1038 / nchembio.2312 31 Chin, J. W. Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu. Rev. Biochem. 83, 379--408 (2014). https: / / doi.org:10.1146 / annurev-biochem-060713-035737 32 Ambrogelly, A. et al. Pyrrolysine is not hardwired for cotranslational insertion at UAG codons. Proc. Natl. Acad. Sci. U.S.A. 104, 3141--3146 (2007). https: / / doi.org:10.1073 / pnas.0611634104 33 Elliott, T. S. et al. Proteome labeling and protein identification in specific tissues and at specific developmental stages in an animal. Nat. Biotechnol. 32, 465 (2014). https: / / doi.org:10.1038 / nbt.2860 34 Suzuki, T. et al. Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 13, 1261 (2017). https: / / doi.org:10.1038 / nchembio.2497 35 Kobayashi, T., Yanagisawa, T., Sakamoto, K. & Yokoyama, S. Recognition of Non-α-amino Substrates by Pyrrolysyl-tRNA Synthetase. J. Mol. Biol. 385, 1352-1360 (2009). https: / / doi.org:10.1016 / j.jmb.2008.11.059 36 Polycarpo, C. R. et al. Pyrrolysine analogues as substrates for pyrrolysyl-tRNA synthetase. FEBS Letters 580, 6695-6700 (2006). https: / / doi.org:10.1016 / j.febslet.2006.11.028 37 Bindman, N. A., Bobeica, S. C., Liu, W. R. & Van Der Donk, W. A. Facile Removal of Leader Peptides from Lanthipeptides by Incorporation of a Hydroxy Acid. J. Am. Chem. Soc. 137, 6975-6978 (2015). https: / / doi.org:10.1021 / jacs.5b04681 38 Li, Y.-M. et al. Ligation of Expressed Protein α-Hydrazides via Genetic Incorporation of an α-Hydroxy Acid. ACS Chemical Biology 7, 1015-1022 (2012). https: / / doi.org:10.1021 / cb300020s 39 Ohtake, K. et al. Engineering an Automaturing Transglutaminase with Enhanced Thermostability by Genetic Code Expansion with Two Codon Reassignments. ACS Synthetic Biology 7, 2170-2176 (2018). https: / / doi.org:10.1021 / acssynbio.8b00157 40 Polycarpo, C. et al. An aminoacyl-tRNA synthetase that specifically activates pyrrolysine. Proc. Natl. Acad. Sci. U. S. A. 101, 12450-12454 (2004). https: / / doi.org:10.1073 / pnas.0405362101 41 Nozawa, K. et al. Pyrrolysyl-tRNA synthetase-tRNAPyl structure reveals the molecular basis of orthogonality. Nature 457, 1163 (2008). https: / / doi.org:10.1038 / nature07611 42 Herring, S. et al. The amino-terminal domain of pyrrolysyl-tRNA synthetase is dispensable in vitro but required for in vivo activity. FEBS Lett. 581, 3197--3203 (2007). https: / / doi.org:10.1016 / j.febslet.2007.06.004 43 Jiang, R. & Krzycki, J. A. PylSn and the homologous N-terminal domain of pyrrolysyl-tRNA synthetase bind the tRNA that is essential for the genetic encoding of pyrrolysine. J. Biol. Chem., jbc.M112.396754 (2012). https: / / doi.org:10.1074 / jbc.M112.396754 44 Meineke, B., Heimgartner, J., Eirich, J., Landreh, M. & Elsasser, S. J. Site-Specific Incorporation of Two ncAAs for Two-Color Bioorthogonal Labeling and Crosslinking of Proteins on Live Mammalian Cells. Cell Reports 31, 107811 (2020). https: / / doi.org:10.1016 / j.celrep.2020.107811 45 Meineke, B., Heimgartner, J., Lafranchi, L. & Elsasser, S. J. Methanomethylophilus alvus Mx1201 Provides Basis for Mutual Orthogonal Pyrrolysyl tRNA / Aminoacyl-tRNA Synthetase Pairs in Mammalian Cells. ACS Chemical Biology 13, 3087-3096 (2018). https: / / doi.org:10.1021 / acschembio.8b00571 46 Zhang, H. et al. The tRNA discriminator base defines the mutual orthogonality of two distinct pyrrolysyl-tRNA synthetase / tRNAPyl pairs in the same organism. Nucleic Acids Research 50, 4601-4615 (2022). https: / / doi.org:10.1093 / nar / gkac271 47 Fischer, J. T., Soll, D. & Tharp, J. M. Directed Evolution of Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase Generates a Hyperactive and Highly Selective Variant. Front. Mol. Biosci. 0 (2022). https: / / doi.org:10.3389 / fmolb.2022.850613 48 Tharp, J. M., Vargas-Rodriguez, O., Schepartz, A. & Soll, D. Genetic Encoding of Three Distinct Noncanonical Amino Acids Using Reprogrammed Initiator and Nonsense Codons. ACS Chemical Biology 16, 766-774 (2021). https: / / doi.org:10.1021 / acschembio.1c00120 49 Laslett, D. ARAGORN, a program to detect tRNA genes and tmRNA genes in nucleotide sequences. Nucleic Acids Research 32, 11-16 (2004). https: / / doi.org:10.1093 / nar / gkh152 50 Fan, C., Xiong, H., Reynolds, N. M. & Soll, D. Rationally evolving tRNAPyl for efficient incorporation of noncanonical amino acids. Nucleic Acids Research 43, e156-e156 (2015). https: / / doi.org:10.1093 / nar / gkv800 51 Katayama, H., Nozawa, K., Nureki, O., Nakahara, Y. & Hojo, H. Pyrrolysine Analogs as Substrates for Bacterial Pyrrolysyl-tRNA Synthetase in Vitro and in Vivo. Bioscience, Biotechnology, and Biochemistry 76, 205-208 (2012). https: / / doi.org:10.1271 / bbb.110653 52 Varani, G. & McClain, W. H. The G·U wobble base pair. EMBO reports 1, 18-23 (2000). https: / / doi.org:10.1093 / embo-reports / kvd001 53 Sievers, F. et al. Fast, scalable generation of high-quality protein multiple sequence alignments using Clustal Omega. Molecular Systems Biology 7, 539 (2011). https: / / doi.org:10.1038 / msb.2011.75 54 Guan, Y., Haroon, M. F., Alam, I., Ferry, J. G. & Stingl, U. Single-cell genomics reveals pyrrolysine-encoding potential in members of uncultivated archaeal candidate division MSBL1. Environmental Microbiology Reports 9, 404-410 (2017). https: / / doi.org:10.1111 / 1758-2229.12545 55 Cock, P. J. A. et al. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics 25, 1422-1423 (2009). https: / / doi.org:10.1093 / bioinformatics / btp163 56 Gruber, A. R., Lorenz, R., Bernhart, S. H., Neubock, R. & Hofacker, I. L. The Vienna RNA Websuite. Nucleic Acids Research 36, W70-W74 (2008). https: / / doi.org:10.1093 / nar / gkn188 57 Lorenz, R. et al. ViennaRNA Package 2.0. Algorithms for Molecular Biology 6, 26 (2011). https: / / doi.org:10.1186 / 1748-7188-6-26 58 Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12, 2825-2830 (2011). 59 Salis, H. M., Mirsky, E. A. & Voigt, C. A. Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol. 27, 946-950 (2009). https: / / doi.org:10.1038 / nbt.1568

Claims

1. It contains an exogenous class C acyl-tRNA synthetase (aRS), and Exogenous class A aRS, Exogenous class B aRS, Exogenous class N aRS, and Exogenous class S aRS Includes one, two, three, or four members of the group It is a cell, The cell wherein each aRS is either a pyrrolidyl-tRNA synthetase (PylRS) or a variant of PylRS modified to alter its acylation specificity.

2. It includes exogenous class S aRS, and, Exogenous class A aRS, Exogenous class B aRS, Exogenous class C aRS, and Exogenous class N aRS Includes one, two, three, or four members of the group It is a cell, Each aRS is either a PylRS or a variant of a PylRS modified to alter its acylation specificity. The aforementioned cells.

3. i) A class A aRS is a PylRS that does not have an N-terminal domain, or a PylRS in which the N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with a lower threshold of sequence identity of 55%, forms a cluster with any of SEQ ID NOs. 4, 35, 203-245, and / or ii) A class B aRS is a PylRS that does not have an N-terminal domain, or a PylRS in which the N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with the aligned sequences of the C-terminal domain, forms a cluster with any of SEQ ID NOs. 6, 36, 188-196, 375, 392, 399, or 400, with a lower threshold of 55% for sequence identity. iii) A class C aRS is a PylRS that does not have an N-terminal domain, or a PylRS in which the N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of the aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with a lower threshold of 55% for sequence identity, forms a cluster with any of SEQ ID NOs. 8, 11, 14, 17, 20, 23, or 247. iv) A class N aRS is a PylRS that, after mean concatenation clustering of the aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with a lower threshold of 55% for sequence identity, forms a cluster with any of SEQ ID NOs. 295-373, and / or v) A class S aRS is a PylRS in which, after mean concatenation clustering of aligned C-terminal domain sequences of sequence numbers 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with the aligned sequences of those C-terminal domains, it forms a cluster with one of sequence numbers 26, 29, 31, 33, 65-187, 197-202, 246, 248-294, 374, 376-391, or 393-398, with a lower threshold of 55% for sequence identity. The cell according to claim 1 or 2.

4. The cell according to any one of claims 1 to 3, wherein the class N aRS is a PylRS derived from an archaeal species, or a variant of PylRS modified to alter its acylation specificity, and the aRS includes the N-terminal domain of aRS as part of the same polypeptide as the C-terminal domain of aRS.

5. The cell according to any one of claims 1 to 4, wherein the aRS of class S is a PylRS derived from a bacterial species, or a variant of PylRS modified to alter the acylation specificity, and the cell expresses a separately encoded N-terminal domain associated with the aRS of class S.

6. Class A aRS induces the expression of a marker gene at 40% or more of the expression level of the control gene, and this expression is measured in cells that express the aRS, express tRNA according to SEQ ID NO: 37, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. Class A aRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the aRS and tRNA mediated by SEQ ID NO: 38, ii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 39, iii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 40, and iv) cells expressing the aRS and tRNA mediated by SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The cell according to any one of claims 1 to 5.

7. Class B aRS induces the expression of a marker gene at 40% or more of the expression of the control gene, and this expression is measured in cells that express the aRS, express tRNA according to SEQ ID NO: 38, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. Class B aRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the aRS and tRNA mediated by SEQ ID NO: 37, ii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 39, iii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 40, and iv) cells expressing the aRS and tRNA mediated by SEQ ID NO: 41, and the cells also express the marker gene which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The cell according to any one of claims 1 to 6.

8. Class C aRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the aRS, express tRNA according to SEQ ID NO: 39, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. Class C aRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the aRS and tRNA mediated by SEQ ID NO: 37, ii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 38, iii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 40, and iv) cells expressing the aRS and tRNA mediated by SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The cell according to any one of claims 1 to 7.

9. The cell according to claim 7 or 8, wherein the class B aRS and / or class C aRS is a bacterial PylRS or a variant modified to alter the acylation specificity of PylRS, and the cell does not express the associated N-terminal domain.

10. Class N aRS induces the expression of a marker gene at 40% or more of the expression level of the control gene, and this expression is measured in cells that express the aRS, express tRNA according to Sequence ID No. 40, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. Class N aRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the aRS and tRNA mediated by SEQ ID NO: 37, ii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 38, iii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 39, and iv) cells expressing the aRS and tRNA mediated by SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The cell according to any one of claims 1 to 9.

11. Class S aRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the aRS, express tRNA according to SEQ ID NO: 41, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. Class S aRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the aRS and tRNA mediated by SEQ ID NO: 37, ii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 38, iii) cells expressing the aRS and tRNA mediated by SEQ ID NO: 39, and iv) cells expressing the aRS and tRNA mediated by SEQ ID NO: 40, and the cells also express the marker gene which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The cell according to any one of claims 1 to 10.

12. The aRS for Class A is 1R26PylRS. Class B aRS is Lum1PylRS or S Δ - This is ClosePylRS. Class C aRS is NitraPylRS or S Δ - It is I2PylRS. Class N aRS is MmPylRS and / or Class S aRS is DebPylS. The cell according to any one of claims 1 to 11.

13. 1R26PylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:

35. Lum1PylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:

36. S Δ - ClosPylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of sequence numbers 26, 27, or 28. NitraPylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of sequence numbers 11, 12, or 13. S Δ - I2PylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 33 or 34. MmPylRS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 1 or 3, and / or DebPylS has an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 29 or 30, and an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:

62. The cell according to claim 12.

14. The cell according to claim 12 or 13, expressing one, two, three, four, or five tRNAs, which are any combination of SEQ ID NO: 37, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, or SEQ ID NO:

59.

15. The cell according to any one of claims 1 to 14, wherein each exogenous aRS is part of an exogenous aRS-tRNA pair.

16. The cell according to claim 15, wherein each exogenous aRS-tRNA pair is orthogonal to each other.

17. In experimental assays, each aRS-tRNA pair induces the expression of a marker gene at more than 40% of the expression level of the control gene, and the expression is measured in cells containing a marker gene that can only be decoded when the aRS-tRNA pair is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. All other pairings of exogenous aRSs and exogenous tRNA induce the expression of the marker gene at less than 20% of the expression of the control gene in the assay. The quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity is greater than 2.

5. The cell according to claim 16.

18. The cell according to any one of claims 1 to 17, wherein all one, two, three, four, or five of the aRS enzymes of class A, B, C, N, and / or S are modified to alter their acylation specificity.

19. A cell according to any one of claims 1 to 18, which is a prokaryotic cell, a bacterial cell, or an Escherichia coli cell.

20. A method for producing cells containing at least two exogenous acyl-tRNA synthetases, i) Screening one or more PylRS to identify a first PylRS belonging to class C or S, ii) Screening one or more PylRS to identify a second PylRS belonging to class A, B, C, N, or S, Optionally, iii) modifying the first PylRS and / or the second PylRS to change the acylation specificity, and iv) Generating cells expressing the first PylRS and the second PylRS, wherein the first PylRS and the second PylRS are not of the same class. The method, including the method described above.

21. A Class A PylRS is a PylRS that lacks an N-terminal domain or in which the N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with an alignment of the C-terminal domain sequence, forms a cluster with any of SEQ ID NOs. 4, 35, 203-245. A Class B PylRS is a PylRS that either lacks an N-terminal domain or whose N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with an aligned sequence of the C-terminal domain, forms a cluster with any of SEQ ID NOs. 6, 36, 188-196, 375, 392, 399, or 400, with a lower threshold of 55% for sequence identity. A Class C PylRS is a PylRS that either lacks an N-terminal domain or whose N-terminal domain is not expressed in cells, and which, after mean concatenation clustering of aligned C-terminal domain sequences of SEQ ID NOs. 4, 6, 8, 11, 14, 17, 20, 23, 26, 29, 31, 33, 35, 36, 64-400 with an aligned sequence of the C-terminal domain, forms a cluster with any of SEQ ID NOs. 8, 11, 14, 17, 20, 23, or 247, with a lower threshold of 55% for sequence identity. A PylRS of class N is a PylRS derived from an archaeal species, wherein the N-terminal domain of the PylRS is included as part of the same polypeptide as the C-terminal domain of the PylRS, and / or The PylRS of class S is a PylRS derived from a bacterial species, and the cell expresses separately encoded N-terminal domains associated with the PylRS of class S. The method according to claim 20.

22. Class A PylRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the PylRS, express tRNA according to Sequence ID No. 37, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. The Class A PylRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the PylRS and tRNA mediated by Sequence ID No. 38, ii) cells expressing the PylRS and tRNA mediated by Sequence ID No. 39, iii) cells expressing the PylRS and tRNA mediated by Sequence ID No. 40, and iv) cells expressing the PylRS and tRNA mediated by Sequence ID No. 41, wherein the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The method according to claim 20 or 21.

23. Class B PylRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the PylRS, express tRNA according to Sequence ID No. 38, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. The Class B PylRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the PylRS and tRNA of SEQ ID NO: 37, ii) cells expressing the PylRS and tRNA of SEQ ID NO: 39, iii) cells expressing the PylRS and tRNA of SEQ ID NO: 40, and iv) cells expressing the PylRS and tRNA of SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The method according to any one of claims 20 to 22.

24. Class C PylRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the PylRS, express tRNA according to Sequence ID No. 39, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. The class C PylRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the PylRS and tRNA of SEQ ID NO: 37, ii) cells expressing the PylRS and tRNA of SEQ ID NO: 38, iii) cells expressing the PylRS and tRNA of SEQ ID NO: 40, and iv) cells expressing the PylRS and tRNA of SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The method according to any one of claims 20 to 23.

25. The method according to claim 23 or 24, wherein the class B PylRS and / or class C PylRS is a bacterial PylRS or a variant modified to alter the acylation specificity of PylRS, and the resulting cells do not express the associated N-terminal domain.

26. Class N PylRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the PylRS, express tRNA according to Sequence ID No. 40, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. The class N PylRS induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the PylRS and tRNA mediated by SEQ ID NO: 37, ii) cells expressing the PylRS and tRNA mediated by SEQ ID NO: 38, iii) cells expressing the PylRS and tRNA mediated by SEQ ID NO: 39, and iv) cells expressing the PylRS and tRNA mediated by SEQ ID NO: 41, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The method according to any one of claims 20 to 25.

27. Class S PylRS induces the expression of a marker gene at more than 40% of the expression of the control gene, and this expression is measured in cells that express the PylRS, express tRNA according to Sequence ID No. 41, and contain the marker gene, and the marker gene can only be decoded when the tRNA is charged, and the control gene is identical to the control gene except that it can be decoded by endogenous tRNA. The PylRS of class S induces the expression of a marker gene at less than 20% of the expression of the control gene, and this expression is measured in i) cells expressing the PylRS and tRNA of SEQ ID NO: 37, ii) cells expressing the PylRS and tRNA of SEQ ID NO: 38, iii) cells expressing the PylRS and tRNA of SEQ ID NO: 39, and iv) cells expressing the PylRS and tRNA of SEQ ID NO: 40, and the cells also express the marker gene, which can only be decoded when the tRNA is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. The method according to any one of claims 20 to 26.

28. Screening one or more PylRS to identify a third PylRS belonging to class A, B, C, N, or S, Optionally, modify the third PylRS to change the acylation specificity, and To generate cells expressing a first PylRS, a second PylRS, and the third PylRS, wherein each PylRS is of a different class. The method according to any one of claims 20 to 27, further comprising:

29. Screening one or more PylRS to identify a fourth PylRS belonging to class A, B, C, N, or S, Optionally, modify the fourth PylRS to change the acylation specificity, and To generate cells expressing a first PylRS, a second PylRS, a third PylRS, and the fourth PylRS, wherein each PylRS is of a different class. The method according to claim 28, further comprising:

30. Screening one or more PylRS to identify a fifth PylRS belonging to class A, B, C, N, or S, Optionally, modify the fifth PylRS to change the acylation specificity, and The method of generating cells that express a first PylRS, a second PylRS, a third PylRS, a fourth PylRS, and the fifth PylRS, wherein each PylRS is of a different class. The method according to claim 29, further comprising:

31. The method according to any one of claims 20 to 30, comprising the steps of determining whether the identified PylRS is active in the cell type of cells to be produced, and if the PylRS is inactive, discarding the PylRS and rescreening to identify a suitable class of PylRS.

32. The method according to any one of claims 20 to 31, wherein the PylRS enzyme expressed by the cell is modified to be capable of loading tRNA with non-natural amino acids or monomers other than alpha amino acids.

33. The method according to any one of claims 20 to 32, wherein in the generated cells, each exogenous aRS is part of an exogenous aRS-tRNA pair.

34. The method according to claim 33, wherein each exogenous aRS-tRNA pair is orthogonal to each other.

35. In experimental assays, each aRS-tRNA pair induces the expression of a marker gene at more than 40% of the expression level of the control gene, and the expression is measured in cells containing a marker gene that can only be decoded when the aRS-tRNA pair is charged, and the control gene is the same gene except that it can be decoded by endogenous tRNA. All other pairings of exogenous aRSs and exogenous tRNAs induce the expression of the marker gene at less than 20% of the expression of the control gene in the assay. The quotient obtained by dividing the lowest intrapair activity by the highest interpair cross-reactivity is greater than 2.

5. The method according to claim 34.

36. Cells obtained or obtainable by the method described in any of claims 20 to 35.

37. A nucleic acid sequence encoding an exogenous protein having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 29 or SEQ ID NO: 30, and a protein having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 62, or Nucleic acid sequences encoding an exogenous protein having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 31 or SEQ ID NO: 32, and a protein having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 63 Cells containing this substance.

38. The cell according to claim 36 or 37, which is a prokaryotic cell, a bacterial cell, or an Escherichia coli cell.

39. Use of the cells according to any one of claims 1 to 19 or 36 to 38 for the production of polymers containing at least one non-natural amino acid or monomers that are not alpha amino acids.

40. The use according to claim 39, wherein the monomer that is not an alpha amino acid is an alpha hydroxy acid or a beta amino acid.

41. The use according to claim 39 or 40, wherein the polymer comprises at least one standard amino acid.

42. A method for producing a polymer containing at least one non-natural amino acid or a monomer that is not an alpha-amino acid, To culture the cells according to any one of claims 1 to 19 or 36 to 38, To supply the gene encoding the polymer to the cell, and To obtain the polymer The method, including the method described above.

43. The method according to claim 42, wherein the monomer that is not an alpha amino acid is an alpha hydroxy acid or a beta amino acid.

44. The method according to claim 43, wherein the polymer comprises at least one standard amino acid.