P ylrs / TRNA pairs for producing proteins of interest
New PylRS-tRNA pairs from archaea enable the incorporation of non-canonical amino acids into proteins in bacterial and eukaryotic cells, addressing the need for orthogonal systems in protein engineering and enhancing protein diversity and functionality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-04-02
AI Technical Summary
There is a need for PylRS-tRNA pairs that are orthogonal to heterologous cells such as bacterial and/or eukaryotic cells for applications in protein engineering, as existing systems are not effectively utilized in these cells and the translation table used by certain genomes has not been firmly established.
The development of new PylRS-tRNA pairs derived from archaea, specifically from Euryarchaeota and Thermoplasmatota, which can incorporate pyrrolysine into proteins, including non-canonical amino acids, and are orthogonal in bacterial and eukaryotic cells, allowing for the site-specific incorporation of non-natural monomers into proteins.
These PylRS-tRNA pairs enable the biosynthesis of proteins containing non-canonical amino and a-hydroxy acids in bacterial and eukaryotic cells, demonstrating high substrate side chain promiscuity and tolerance for mutations, thus expanding the diversity and functionality of proteins.
Smart Images

Figure IMGF000078_0001 
Figure IMGF000099_0001 
Figure IMGF000029_0001
Abstract
Description
PYLRS / TRNA PAIRS FOR PRODUCING PROTEINS OF INTERESTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 698,833, filed September 25, 2024, which application is incorporated herein by reference in its entirety.STATEMENT FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with Government support under grant numbers 2002182 and 2334028 awarded by the National Science Foundation. The Government has certain rights in this invention.INCORPORATION BY REFERENCE OF SEQUENCE LISTING PROVIDED AS AN XM L FILE
[0003] A Sequence Listing is provided herewith as a Sequence Listing XML, “BERK- 542PRV_SEQ_LIST.xml” created on September 16, 2024 and having a size of 108,944 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in their entirety.I. INTRODUCTION
[0004] Pyrrolysine (Pyl), a derivative of lysine and formally the 22ndamino acid, features a pyrroline ring and is encoded by TAG, normally a stop codon. Its synthesis and incorporation in proteins requires Pyl B,C,D genes as well as a pyrrolysyl-tRNA synthetase (PylRS) and the Pyl tRNA with a CUA anticodon (UAG codon). Pyl has only been experimentally confirmed in mono-, di-, tri- methylamine methyltransferases, tRNA guanylyltransferase, and PylB in archaea. In the archaeon Methanosarcina acetivorans, five additional proteins are also presumed to contain Pyl because their detected molecular masses are consistent with TAG readthrough. Pyl is also hypothesized to occur in other proteins when internal TAG codons were observed during in silico analysis of protein-coding genes. However, these observations remain unverified, and the translation table used by these genomes has not been firmly established. It has also been suggested that the TAG codon has a dual meaning in Pyl-encoding archaea, although later analyses revealed that the proposed requirement for signaling Pyl incorporation is absent in a subset of known Pyl-containing proteins.
[0005] To date, Pyl has been reported within few bacteria and 11 major groups of archaea, including members of Methanosarcinae, Archaeoglobacea, MSBL-1 clade (Perse phonarchaea), Methanonatronarchaeia, Thermoplasmatota, Asgardarchaeota, Hydrothermarchaeota Verstraetearchaeota, Nitrososphaerota, Bathyarchaeota, and (tentatively) Korarchaeota.
[0006] There is a need for PylRS-tRNAPylpairs, e.g., pairs that are orthogonal to heterologous cells such as bacterial and / or eukaryotic cells, for applications such as protein engineering. PylRS-tRNAPylpairs are provided herein.II. SUMMARY
[0007] Genetic code expansion (GCE) has gained considerable interest in protein engineering, as it allows the site-specific incorporation of non-natural monomers (NNMs) (e.g., non-canonical amino acids (ncAAs) as well as a-hydroxy acid, p2- hydroxy, p3-amino acid, and p2-amino acid analogs of a-amino acids) into proteins. NNMs have gained significant attention in protein engineering and drug development owing to their ability to introduce new properties and functionalities to proteins, including in eukaryotes. GCE is a powerful tool, e.g., for cell or animal imaging, the monitoring of protein interactions in target cells, drug development, and switch regulation. The use of NNMs through GCE offers a promising avenue for expanding the diversity and functionality of proteins (see, e.g., Guo et al., Protein Cell. 2024 May 7;15(5):331-363; Wan et al., Biochim Biophys Acta. 2014 Jun; 1844(6): 1059-70; and Young et al., ACS Chem Biol. 2018 Apr 20; 13(4): 854-870).
[0008] The work described in the experimental examples below led to the surprising finding by the inventors that some Euryarchaeota (Halobacteriota) and Thermoplasmatota use a new genetic code in which TAG, normally a stop codon, consistently encodes pyrrolysine. Pyrrolysine incorporation was confirmed via proteomics and many new enzymes containing Pyl residues were identified. New PylRS - tRNAPylpairs were identified from organisms with this newly discovered archaeal genetic code. Given the high prevalence of TAG / UAG codon usage in these organisms - the newly identified PylRS / tRNA pairs are expected to function effectively regardless of where the codon (e.g. UAG / TAG) is located along the length of the protein-coding sequence. The work described in the working examples below demonstrated that each one can be used to generate proteins containing non-canonical a-amino and a-hydroxy acids.
[0009] PylRS / tRNAPylcan be used in bacterial and eukaryotic cells. Owing to its high substrate side chain promiscuity and high tolerance for mutations in the substratebinding pocket, PylRS, e g., derived from two genomes of archaea in the family Methanosarcinaceae, possesses flexible active sites and is orthogonal in bacteria and eukaryotes. Thus, PylRS has become a frequently applied aaRS in eukaryotes. Several PylRS-tRNAPylpairs are orthogonal in both bacterial and mammalian cells and can be employed to biosynthesize proteins containing non- canonical a-amino, a-hydroxy, p2-hydroxy, and3-amino acids using genetic code expansion.
[0010] Provided are compositions and methods for producing proteins with non-natural monomers (NNMs) using newly identified PylRS - tRNAPyl pairs. In some embodiments, the PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77. In some embodiments, the PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10. In some cases, the PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1. In some cases, the tRNA includes a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18 and 81-97. In some cases, the tRNA includes a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, the tRNA includes a nucleotides sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11.
[0011] The methods include providing a subject PylRS, a tRNA, and a NNM or salt thereof, to a bacterial cell, eukaryotic cell, or an in vitro translation system, where the tRNA can be acylated with the NNM by the PylRS. Such methods can also include providing a DNA or mRNA that encodes a protein of interest (POI) - where the nucleotides encoding the POI include one or more in-frame codons recognized by the tRNA (i.e. , the in-frame codon(s) is the reverse complement of the anticodon of the tRNA). In some cases, the codon is a stop codon for the bacterial or eukaryotic cell (e.g., anamber (UAG), ochre (UAA) or opal (UGA) stop codon, e.g., in some cases the codon is UAG and the anti-codon of the tRNA is therefore CUA).
[0012] In some cases, the NNM is selected from those depicted in Figures 15-22. In some cases, the NNM is selected from the lysine derivatives depicted in Figures 15-22. In some cases, the NNM is selected from those depicted in Figure 15. In some cases, the NNM is an a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog of any of the molecules depicted in Figures 15-22. In some cases, the NNM is an a-hydroxy acid, P2-hydroxy acid, p3-amino acid analog, or p2-amino acid analog of any of the molecules depicted in Figure 15.
[0013] Provided are systems that include a nucleic acid encoding a subject PylRS (e.g., one that includes an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, one that includes an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10 and 25-77), where the nucleotide sequence encoding the PylRS is: a) operably linked to a bacterial or eukaryotic promoter; and / or (b) codon-optimized for expression in a bacterial or eukaryotic host cell. In some cases, the nucleic acid is a plasmid or a viral vector. In some cases, the system further includes a tRNA (or a nucleic acid encoding it), where, e.g., the tRNA can be acylated by the PylRS (e.g., they are an orthogonal pair). In some cases, the nucleotide sequence encoding the tRNA is present on the same nucleic acid as the nucleotide sequence encoding the PylRS (and in some cases they are on different nucleic acids). In some cases, the nucleotide sequence encoding the tRNA is operably linked to a bacterial or eukaryotic promoter, and / or is codon-optimized for expression in a bacterial or eukaryotic host cell. In some cases, a subject system also includes one or more NNMs (that can be used by the PylRS to acylate the tRNA). See, e.g., the NNMs of Figures 15-22.
[0014] Also provided are cell, e.g., bacterial or eukaryotic cells, that include a subject PylRS or a nucleic acid encoding same; and a subject tRNA or a nucleic acid encoding same.
[0015] Reagents, compositions, and kits / systems that find use in practicing the subject methods are provided.III. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The following detailed description of embodiments of the invention will be better understood when read in conjunction with the appended drawings. It should beunderstood that the invention is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.
[0017] FIG. 1 Comparison of phylogeny and distribution of the Pyl amino acid in archaea. The range of TAG frequency of Pyl genomes in each taxonomic group is shown.
[0018] FIG. 2 Phylogeny of PylRS system and providing support for lateral gene transfer.
[0019] FIG. 3 Categories of TAG codons in Pyl-encoding archaea.
[0020] FIG. 4 Phylogenetic tree - Comparison of percent of TAG codons in each of the categories and the TAG frequency for Pyl-encoding archaea, sorted by taxonomy.
[0021] FIG. 5A-5B (A) MS / MS spectrum of Pyl-containing tryptic peptides for M. burtonii Deacetylase. The blue peaks represent observed b-ions and the yellow-peaks represent observed y-ions. Black peaks were not identified as b- or as y-ions. Observed b- and y-ions are indicated in the peptide sequence ladder as well as in the MS / MS spectrum. (B) Top: Summary of expressed Pyl proteins in proteomic data; Coverage of Pyl-proteins with Trypsin and Glu-Cin M. burtonii. Bottom: Summary of expressed Pyl proteins in proteomic data; Coverage of Pyl proteins with Trypsin and Glu-Ci in M. alvus.
[0022] FIG. 6 The alpha carbons of the two predicted Pyl residues are marked with large orange and magenta spheres for Methanosarcina sp., and Methanohalophilus sp., respectively, for PDB: 4AY7, Corrinoid:Coenzyme M Methyltransferase. In yellow is a view into the binding cavity of coenzyme M methyltransferase MtaA.
[0023] FIG. 7 The alpha carbon of the predicted Pyl residue is shown in orange, and is located in the bridge helix (yellow) of IscB (PDB 7UTN), and the bound nucleic acid is shown in brown.
[0024] FIG. 8A-8C (A) Chemical structures of established PylRS substrates: pyrrolysine (Pyl, 1) the natural substrate for PylRS; Boc-Lysine 2 (BocK), a synthetic a-amino acid substrate, and OH-Boc-Lysine 3 (BocK-OH), an a-hydroxy acid substrate. (B) Shown is the F528 signal normalized to GD600 for bacterial cultures of BL21 cells doubly transformed with pMega plasmids harboring the PylRS / tRNAPyl from candidate organisms and a pET22A reporter plasmid containing sfGFP with an in-frame amber codon at position 3. Cells were grown to OD600 = ~0.6 and then supplemented with 1 mM of the indicated PylRS substrate and sfGFP expression was induced with 1 mM IPTG. F528 / OD600 signal was acquired for over a 20 h time frame and the final signal at T = 20 h is shown here. (C) Intact MS of sfGFP-3TAG expressed in the presence of a control PylRS / tRNAPyl pair (from M. alvus) or in the presence of the best performingcandidate PylRS / tRNAPyl pair (M. levihalophilus) as determined by expression tests in panel B. sfGFP was expressed in the presence of 1 mM 2 or 3.
[0025] FIG. 9 Steps for assigning TAG codon into different categories.
[0026] FIG. 10 Different types of misannotation for Pyl proteins resulting from misinterpretation of TAG as a stop codon.
[0027] FIG. 11 Intact MS of rest of the constructs.
[0028] FIG. 12 Full gel of 3TAG-sfGFP, including a positive control (maPyIRS).
[0029] FIG. 13 Time course data for expression test.
[0030] FIG. 14 Overview of the methods I workflow used in Example 1 .
[0031] FIG. 15 NNMs that serve as substrates for native (wild type) PylRS.
[0032] FIG. 16 Lysine derivatives that have been genetically incorporated into proteins using engineering PylRS mutants in coordination with tRNAPyl.
[0033] FIG. 17 Phe derivatives that have been genetically incorporated into proteins using engineering PylRS mutants in coordination with tRNAPyl.
[0034] FIG. 18 a-Hydroxy acids that have been genetically incorporated into proteins using PylRS mutants in coordination with tRNAPyl.
[0035] FIG. 19 Lys derivatives that can be incorporated in vivo using PylRS. The depicted NNMs are Lys-derivatives containing a carbamate group starting at the Lys NE. Most of these NNMs can be incorporated with the wild-type or PylRS(Y306A:Y384F, M. mazei notation) variant. Photo- or chemically caged post-translational modifications (blue), NNMs containing bioorthogonal groups and thus are possible targets for sitespecific bioconjugation (red), photo- or chemically caged NNMs (which include canonical a-amino acids that are photo- or chemically caged) (magenta), NNMs containing functional groups which are useful for spectroscopic applications (orange), photo- or proximity triggered cross-linking NNMs (violet), fluorescent NNMs (cyan).
[0036] FIG. 20 Lys derivatives that can be incorporated in vivo using the PylRS. The depicted NNMs are Lys-derivatives with a Lys NE amide, or one with an ester (146). 140 and 141 are exceptions. Caged or noncaged post-translational modifications (blue), NNMs containing bioorthogonal groups and are thus possible targets for sitespecific bioconjugation (red), photo- or proximity triggered cross-linkable NNMs (violet).
[0037] FIG. 21 Short and bulky or larger bulky non-Lys NNMs that can be incorporated in vivo using the PylRS. Post-translational modifications (blue), NNMs containing bioorthogonal groups and are thus possible targets for site-specific bioconjugation (red), photocaged NNMs (which include canonical a-amino acids that are photocaged)(magenta), NNMs containing functional groups which are useful for spectroscopic applications (orange), photo- or proximity triggered cross-linkable NNMs (violet), fluorescent NNMs (cyan).
[0038] FIG. 22 Small and bulky His analogs, small aliphatic and small caged NNMs. This group also contains the unusual a- and p-hydroxy acids, and all of them can be incorporated in vivo using the PylRS. Some of the Lys derivatives contain ester which are not that common in the GCE field. Also, the in vivo incorporable a-hydroxy acids are very unusual substrates. They especially highlight the unique substrate promiscuity of the PylRS. Post-translational modifications (blue), NNMs containing bioorthogonal groups and thus are possible targets for site-specific bioconjugation (red), photocaged NNMs (which include canonical a-amino acids that are photocaged) (magenta), NNMs containing functional groups which are useful for spectroscopic applications (orange), photo- or proximity triggered cross-linkable NNMs (violet), fluorescent NNMs (cyan).
[0039] FIG. 23 Distribution of the Pyl machinery in Archaea is paraphyletic, with the Pyl machinery variably present among representatives of the same family (pie charts). The range of TAG frequency in Pyl genomes (salmon violins) shows that low TAG frequency is not a hallmark of Pyl archaea, in contrast to previous reports. Grey violins correspond to genomes without Pyl. The highest TAG% values for Thermoproteia (up to 66.9%), Bathyarchaeia-other (up to 52.2%) and Nitrososphaeria-other (up to 57.1%) are not shown. The “+” symbols indicate that Pyl genomes represent less than 5% of the genomes in the displayed lineage. The total number of genomes per lineage displayed in the tree is indicated between brackets next to the lineage name.
[0040] FIG. 24 Phylogenetic tree showing evidence for the sporadic distribution of the Pyl code in archaea. Genomes inferred to be fully recoded (black boxes, Group C) have a large fraction of internal TAG codons (Category 1) and very few cases where TAG readthrough leads to large out-of-frame gene overlaps (Category 4). In Group A genomes, TAG content is generally high (>15%), while Group B genomes typically have 5 - 15% TAG content and internal TAG in genes, in addition to the methylamine methyltransferases. Within Group C, the TAG content is variable, probably in part due to the time since code transition. Tree construction used 40 universally conserved proteins. 9816 positions, ultrafast bootstrap values >90% are displayed.
[0041] FIG. 25A-25B. Proteomic coverage confirms Pyl residues and extensions past the TAG codons for proteins expressed by( A) M. burtonii and (B) M. alvi. Coveragebased on enzymatic cleavage using Trypsin is shown in blue and Glu-C is shown in green. The TAG location is shown with a vertical red line and Pyrrolysine is denoted by the one-letter designation, “O”.
[0042] FIG. 26A-26D. PylRS variants from organisms that use the Pyl code support the incorporation of non-canonical a-amino and a-hydroxy acids into the model protein sfGFP. A, Chemical structures of three established PylRS substrates: pyrrolysine (Pyl, 1) the natural substrate for PylRS; Boc-Lysine 2 (BocK), a synthetic a-amino acid substrate, and HO-Boc-Lysine 3 (BocK-HO), an a-hydroxy acid substrate. B, Shown is the F528 signal normalized to ODeoo for bacterial cultures of BL21(DE3) cells doubly transformed with a pMega plasmid harboring the PylRS / tRNAPylfrom candidate organisms A-J and a pET22A reporter plasmid containing sfGFP with an in-frame amber codon at position 3. Cells were grown to OD6oo = ~0.6, supplemented with 1 mM of the indicated PylRS substrate, and sfGFP expression was induced with 1 mM IPTG. F528 / OD6OO signal was acquired for over a 20 h time frame and the final signal at T = 20 h is shown here. Note: the presence of separately encoded PylSc and PylSn domains of the PylRS protein in the JDFR19 genus. As prior studies have found PylRS enzymes naturally lacking PylSn to be active, 2 JDFR19 PylRS homologs with only their PylSc domain were also evaluated. C, SDS-PAGE of purified sfGFP-3TAG constructs in the presence of either 1 mM 2 or 3 and the corresponding PylRS / tRNAPylpair. D, Intact MS of sfGFP-3TAG expressed in the presence of a control PylRS / tRNAPylpair (from M. alvi) or in the presence of the best performing candidate PylRS / tRNAPylpair (M levihalophilus) as determined by expression tests in panel B. sfGFP was expressed in the presence of 1 mM 2 or 3.
[0043] FIG. 27 Maximum likelihood phylogeny of the PylBCDS enzymes. Stars on the branch indicate the status of PylSn (missing, fused with PylSc or separated from PylSn). 1 ,241 positions, ultrafast bootstrap values >90% are displayed.
[0044] FIG. 28 Steps for assigning TAG codon into different categories. Genes in blue contain a TAG codon, while other genes are shown in yellow.
[0045] FIG. 29 TAG-containing genes in Pyl-encoding archaea are sorted into one of four categories based on the result of readthrough of the TAG codon. Category 1 (internal) genes contain an internal TAG codon and encode proteins for which the amino acid sequence before and after the TAG codon has >30% sequence identity with sequences from the UniProt databases. Category 2 includes short extensions, when sequence extension after TAG results in a <15 aa extension and no sense codon overlap. Category 3 includes genes with extension past TAG resulting in an in-framefusion, a short overlap with an adjacent gene, or a long extension. Large-out-of frame overlaps (>60 nt) are assigned to Category 4.
[0046] FIG. 30 Alignment of protein (GAF domain-containing / PAS family protein); Pyl (O) is at position 269 in multiple genera.
[0047] FIG. 31 The positions of the Pyl residue in homologs of methylcobalamin:CoM methyltransferase (MtaA) from A, Methanosarcina sp. (orange) and B, Methanohalophilus sp. (magenta) are marked on the structure of MtaA (PDB 4AY7). Both positions are near the opening of the binding cavity (yellow, shown in complex with Zn2+and coenzyme M). The actual positions may be even closer together due to potential differences in the structures of the two homologs.
[0048] FIG. 32 The positions of the Pyl residue in each of five putative homologs of IscB from different organisms (A, D, Methanosarcina lacustris sp., B, Methanonatronarchaeum sp. Amet6_2, C, Candidatus Methanohalarchaeum thermophilum, E, Methanohalobium evestigatum) are marked in orange on the bridge helix A, and HNH endonuclease domain B,C,D,E, of IscB (PDB 7UTN). Although these positions are not conserved, they are all located near the paired guide RNA (yellow) and target DNA (pink) strands.
[0049] FIG. 33 MS / MS spectrum of two Pyl-containing proteins. The blue peaks represent observed b-ions and the yellow-peaks represent observed y-ions. Black peaks were not identified as b- or as y-ions. Observed b- and y-ions are indicated in the peptide sequence ladder as well as in the MS / MS spectrum. (Upper panel) MS / MS spectrum of Pyl-containing tryptic peptide of M. burtonii deacetylase from the global DDA measurements. (Lower panel) MS / MS spectrum of Pyl-containing tryptic peptide for M. bt / rton / 7 trimethylamine methyltransferase 2 from the global DDA measurements.
[0050] FIG. 34 Different types of misannotation for Pyl proteins resulting from misinterpretation of TAG as a stop codon.
[0051] FIG. 35 Time-dependent changes in 528 nm emission (F528) and cell density (ODeoo) of BL21 (DE3) E. coli harboring a pMega plasmid expressing the indicated PylRS / tRNApylpair and a reporter plasmid encoding sfGFP-3TAG and grown in the presence of 1 mM BocK (monomer 2) or a-HO-BocK (monomer 3). Growth and expression conditions are described further in Supplementary Information.
[0052] FIG. 36 Uncropped SDS-PAGE gel of lysates from BL21(DE3) cells transformed with a pMega plasmid encoding the indicated PylRS / tRNAPylpair and a reporter plasmid encoding sfGFP-3TAG.
[0053] FIG. 37 LC-MS characterization of sfGFP-3TAG produced in BL21(DE3) cells supplemented with either Boc-Lys (monomer 2) or a-HO-Boc-Lys (monomer 3). Shown are the deconvolved mass spectra of purified sfGFP-3TAG isolated from BL21(DE3) E. co / / transformed with a pMega plasmid encoding the indicated PylRS / tRNApylpair and a reporter plasmid encoding sfGFP-3TAG.
[0054] FIG. 38 A schematic representation of the proposed mechanism for the evolution of genome-wide TAG recoding. Organisms that have recently acquired the Pyl cassette still retain their prior genomic distribution of TAG, with frequent use of TAG as a stop codon and little to no incorporation of TAG into genes (column 1). The presence of the Pyl cassette exerts a strong evolutionary pressure against the use of TAG as a stop codon, leading to a large decrease in TAG frequency. At the same time, TAG coding for Pyl begins to appear inside genes, potentially through neutral mutation (column 2). Finally, organisms well-adapted to the Pyl cassette no longer use TAG as a stop codon but their TAG frequency increases despite this due to the accumulation of TAG coding for Pyl inside genes (column 3).
[0055] FIG. 39A-39C The relationship between TAG% and Category 1 TAG genes (A) and Category 4 TAG genes (B) and the environment type and methanogen or nonmethanogen status of the organisms (C). Saline environments (env. saline) correspond to hypersaline lakes and marine environments.IV. DEFINITIONS
[0056] The meaning of “orthogonality” (e.g., as used in the phrase “orthogonal pair”) in the context of PylRS / tRNA pairs as used herein means that the PylRS cannot acylate other tRNAs that are present (e.g., other tRNAs present in a host cell, e g., endogenous tRNAs present in a bacterial or eukaryotic cell); and that the tRNA is not a substrate for other RS enzymes that are present (e.g., other RS enzymes present in a host cell, e.g., endogenous RS enzymes present in a bacterial or eukaryotic cell). The members of an orthogonal pair therefore have a unique correspondence with each other. The meaning of orthogonality of aaRS / tRNA pairs is understood in the art, and can be found, e.g., in Wang L, Schultz P G. Expanding the genetic code [J], Angewandte chemie international edition, 2005, 44(1): 34-66.
[0057] The terms “polypeptide,” “peptide,” and “protein”, used interchangeably herein, refer to a polymeric form of amino acids of any length, which can include genetically coded and non-genetically coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. Theterm includes fusion proteins, including, but not limited to, fusion proteins with a heterologous amino acid sequence, fusions with heterologous and homologous leader sequences, with or without N-terminal methionine residues; immunologically tagged proteins; and the like.
[0058] As used herein the term “isolated” is meant to describe a compound of interest that is in an environment different from that in which the compound naturally occurs. “Isolated” is meant to include compounds that are within samples that are substantially enriched for the compound of interest and / or in which the compound of interest is partially or substantially purified. For example, an "isolated" plasmid, nucleic acid, vector, virus, virion, host cell, protein, or other substance refers to a preparation of the substance devoid of at least some of the other components present where the substance or a similar substance naturally occurs or from which it is initially prepared. Thus, for example, an isolated substance may be prepared by using a purification technique to enrich it from a source mixture. Enrichment can be measured on an absolute basis, such as weight per volume of solution, or it can be measured in relation to a second, potentially interfering substance present in the source mixture. An isolated plasmid, nucleic acid, vector, virus, host cell, protein, or other substance is in some embodiments purified, e.g., from about 80% to about 90% pure, at least about 90% pure, at least about 95% pure, at least about 98% pure, or at least about 99%, or more, pure. In some cases, a nucleic acid of the disclosure (e.g., a nucleic acid encoding a subject PylRS, a nucleic acid encoding a tRNA, or a nucleic acid encoding both) is isolated. In some cases, a tRNA of the disclosure is isolated. In some cases, a PylRS of the present disclosure is isolated. In some cases, a composition is isolated (e.g., a composition that includes: a PylRS, or a nucleic acid encoding it; and a tRNA or a DNA encoding it).
[0059] As used herein, the term “purified” or “substantially purified” refers to a compound that is removed from its natural environment and is at least 60% free, at least 75% free, at least 80% free, at least 85% free, at least 90% free, at least 95% free, at least 98% free, or more than 98% free, from other components with which it is naturally associated.
[0060] The term “vector” as used herein particularly refers to plasmids, cosmids, viruses, bacteriophages and other vectors commonly used in genetic engineering. In one embodiment of the present disclosure, the vectors are suitable for the transformation, transduction and / or transfection of host cells, e.g., prokaryotic cells (e.g., (eu)bacteria,archaea) or eukaryotic cells (e.g., mammalian cells, insect cells, human cells, mouse cells, non-human primate cells, fungal cells, yeast, and the like).
[0061] “Heterologous,” as used herein, means a nucleotide or polypeptide sequence that is not found in the native nucleic acid or protein, respectively, or is not found together in nature. For example, the archaeal PylRS proteins of SEQ ID NOs: 1-10 and 25-77 would be considered to be heterologous to a given bacterial cell or eukaryotic cell. Likewise, if a given PylRS were to be fused with an expression tag, e.g., a His tag, the PylRS amino acid sequenced could be considered to be heterologous to the His tag sequence.
[0062] The terms “polynucleotide” and “nucleic acid,” used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxynucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The terms “polynucleotide” and “nucleic acid” should be understood to include, as applicable to the embodiment being described, single-stranded (such as sense or antisense) and double-stranded polynucleotides. A nucleic acid may be a single-stranded or double-stranded deoxyribonucleotide, or ribonucleotide of any length, and include coding and noncoding sequences of a gene, exons, introns, sense and anti-sense complimentary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acids, isolated and / or purified DNA and / or RNA, synthetic DNA and RNA, fragments, primers and nucleic acid probes. The skilled artisan is aware that the nucleotide sequences of RNA are identical to the DNA sequences that encode them, with the difference of thymine (T) being replaced by uracil (U).
[0063] An “isolated nucleic acid” or “isolated nucleic acid sequence” relates to a nucleic acid or nucleic acid sequence that is in an environment different from that in which the nucleic acid or nucleic acid sequence occurs naturally and can include those that are substantially free from contaminating endogenous material.
[0064] Percent complementarity between particular stretches of nucleotide sequences within nucleic acids and / or amino acid sequences within proteins can be determined routinely, e.g., using convenient BLAST programs (basic local alignment search tools) and PowerBLAST programs known in the art (e.g., Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) (e.g., BLASTP, BLASTN) (e.g., blast.ncbi.nlm.nih.gov / Blast.cgi).
[0065] As used herein, the term “endogenous nucleic acid” refers to a nucleic acid that is normally found in and / or produced by a given bacterium, eukaryotic cell, organism, or cell in nature. An “endogenous nucleic acid” is also referred to as a “native nucleic acid” or a nucleic acid that is “native” to a given bacterium, eukaryotic cell, organism, or cell.
[0066] The terms “DNA regulatory sequences,” “control elements,” and “regulatory elements,” used interchangeably herein, refer to transcription and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, protein degradation signals, and the like, that provide for and / or regulate expression, e.g., of a coding sequence and / or production of an encoded polypeptide in a host cell.
[0067] An “expression cassette” comprises a DNA coding sequence operably linked to a promoter.
[0068] The term “transformation” is used interchangeably herein with “genetic modification” and refers to a permanent or transient genetic change induced in a cell following introduction of new nucleic acid (i.e. , DNA exogenous to the cell). Genetic change (“modification”) can be accomplished either by incorporation of the new DNA into the genome of the host cell, or by transient or stable maintenance of the new DNA as an episomal element. Where the cell is a eukaryotic cell, a permanent genetic change is generally achieved by introduction of the DNA into the genome of the cell. In prokaryotic cells, permanent changes can be introduced into the chromosome or via extrachromosomal elements such as plasmids and expression vectors, which may contain one or more selectable markers to aid in their maintenance in the recombinant host cell. Suitable methods of genetic modification include viral infection, transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, and the like. The choice of method is generally dependent on the type of cell being transformed and the circumstances under which the transformation is taking place (i.e. in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al, Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.
[0069] “Operably linked” refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For instance, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression (The coding sequence can also be said to be operably linked to the promoter). As used herein, the terms “heterologous promoter” and “heterologous control regions” refer to promoters and other control regions that arenot normally associated with a particular nucleic acid in nature. For example, a “transcription control region heterologous to a coding region” is a transcription control region that is not normally associated with the coding region in nature.
[0070] A “host cell,” or “transformed cell” as used herein, denotes an in vivo or in vitro eukaryotic cell, a prokaryotic cell (e.g., bacterial cell), or a cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, which eukaryotic cells (e.g., a fungal cell, a plant cell, an insect cell, a mammalian cell) or prokaryotic cells (e.g., bacterial) can be, or have been, used as recipients for a nucleic acid (e.g., an expression vector that comprises a nucleotide sequence encoding a subject PylRS and / or tRNA), and include the progeny of the original cell which has been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement as the original parent, due to natural, accidental, or deliberate mutation. The host cell may contain a recombinant gene which has been integrated into the nuclear or organelle genomes of the host cell. Alternatively, the host may contain the recombinant gene extra-chromosomally.
[0071] The term “conservative amino acid substitution” refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, a group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids having amide-containing side chains consists of asparagine and glutamine; a group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains consists of lysine, arginine, and histidine; and a group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0072] A particular organism or cell is meant to be “capable of producing a POI” when it produces a POI naturally or when it does not produce said POI naturally but is transformed (e.g., with an exogenous nucleic acid) to produce said POI.
[0073] Functional equivalents of the polypeptides according to the disclosure can be produced by mutagenesis, e.g., by point mutation, lengthening or shortening of the proteins or nucleic acids.
[0074] The term “nuclear export signal” (abbreviated as “NES”) refers to an amino acid sequence which can direct a polypeptide containing it (such as a NES-containingPylRS) to be exported from the nucleus of a eukaryotic cell. NESs are known in the art. For example, NES databases such as NESbase(services. healthtech. dtu.dk / datasets / NESbase-1.0 / ); see, e.g., Le Cour et al., Nucl Acids Res 31 (1), 2003; Cour et al., La Cour et al., Protein Eng Des Sei 17(6):527-536, 2004; and Fu et al., Nucl Acids Res 41 :D338-D343, 2013. As such an NES can be added to a subject PylRS. As an illustrative example, a NES can be attached at or near (e.g., within 50 amino acids of) the C-terminus and / or can be attached at or near (e.g., within 50 amino acids of) the N-terminus.
[0075] The term “nuclear localization signal” (abbreviated as “NLS”, also referred to in the art as “nuclear localization sequence”) refers to an amino acid sequence which can direct a polypeptide containing it (e.g., a PylRS) to be imported into the nucleus of a eukaryotic cell. NLSs are known in the art, and a multitude of NLS databases and tools for NLS prediction are available. PylRSs of the disclosure can be modified by removing an NLS present in the naturally occurring PylRS sequence. In some cases, at least one NES is also introduced. The removal of a NLS from and / or the introduction of a NES into a PylRS or mutant thereof, can change the localization of the modified polypeptide when expressed in a eukaryotic cell, and in particular can avoid or reduce accumulation of the polypeptide in the nucleus of the eukaryotic cell. Thus, the localization of a PylRS of the disclosure expressed in a eukaryotic cell can be altered. See, e.g., WO2018 / 06948, the disclosure of which is hereby incorporated by reference.
[0076] Unless indicated otherwise, “tRNAPyl”, can be used to refer to a tRNA that can be acylated (essentially selectively and in particular selectively) by a PylRS of the disclosure. Such a tRNA is also referred to herein simply as a “tRNA.” The tRNAPyldescribed herein in the context of the present disclosure may be a wildtype tRNA that can be acylated by a PylRS with pyrrolysine (or any NNM), or a mutant of such tRNA. In some embodiments, the anticodon of the tRNAPylrecognizes the amber stop codon (UAG), however variants can be generated such that the tRNA will recognize another desired codon (e.g., ocher (UAA), opal (UGA)).
[0077] A polynucleotide sequence encoding a “polypeptide of interest” or “POI” can comprise one or more, e.g., two or more, more than three, etc., codons (e.g., selector codons) which are recognized by the anticodon of a subject tRNA. The term “translation system” generally refers to a set of components necessary to incorporate an amino acid into a polypeptide chain (protein). Components of a translation system can include, e.g., ribosomes, tRNAs, aminoacyl tRNA synthetases (RS), mRNA and thelike. Translation systems include artificial mixture of said components, cell extracts and living cells, e.g., living bacterial cells, living eukaryotic cells.
[0078] Preferably, the pair of PylRS and tRNA used for preparing a POI according to the present disclosure will be orthogonal to the translation system being used. Incorporation occurs in a site-specific manner, e.g., the tRNAPylrecognizes a codon (e.g., a selector codon such as an amber stop codon) in the mRNA coding for the POI.
[0079] In some cases, a subject PylRS can be referred to as preferentially acylating a tRNAPyl. The term “preferentially acylates” refers to an efficiency of, e.g., about 50% efficient, about 70% efficient, about 75% efficient, about 85% efficient, about 90% efficient, about 95% efficient, or about 99% or more efficient, at which the PylRS acylates the tRNAPylwith an NNM compared to an endogenous tRNA or amino acid of a host cell, e.g., a bacterial or eukaryotic cell. The NNM is then incorporated into a growing polypeptide chain with high fidelity, e.g., at greater than about 75%, greater than about 80%, greater than about 90%, greater than about 95%, or greater than about 99% or more efficiency for a given codon (e.g., selector codon) that is recognized by the anticodon of the tRNA.
[0080] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0081] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0082] Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to orapproximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0083] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.
[0084] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.
[0085] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. As such, the articles “a” and “an” are used herein to refer to one or to more than one (i.e. , to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Thus, for example, reference to “a cell” includes a plurality of such cells and reference to “the polypeptide” includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0086] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. For example, it is appreciated that certain features of the invention, which are, for clarity, described in the context ofseparate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0087] While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112.V. DETAILED DESCRIPTION
[0088] Genetic code expansion (GCE) is a versatile tool for the site-specific incorporation of non-natural monomer (NNMs) into proteins. These NNMs can be used for protein engineering like fluorescent labeling of proteins or conjugation of toxic payloads with antibodies. For the incorporation of the NNMs into the protein during translation, an aminoacyl-tRNA synthetase / tRNA (aaRS / tRNA) pair can be used, which recognizes / incorporates the NNM and otherwise is orthogonal to the host organism. There are over 200 known NNMs that can be used in combination with PylRS / tRNAPylpairs. See, e.g., Liu et al., Annu Rev Biochem 83:379-408, 201 ; Lemke, ChemBioChem 15:1691-1694, 2014; Wan et al., Biochim Biophys Acta. 2014 Jun; 1844(6): 1059-70; Koch et al., Chem Rev. 2024 Jul 2; and Guo et al., Protein Cell. 2024 May 7;15(5):331-363.Compositions (e.g., systems, kits) and Methods
[0089] As noted above, provided are composition and methods for producing proteins with NNMs using the newly identified PylRS - tRNAPyl pairs. A subject method includes providing to a bacterial cell, a eukaryotic cell, or an in vitro translation system: (i) apyrrolysyl-tRNA synthetase (PylRS) (wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, or wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10 and 25-77); (ii) a NNM or a salt thereof; and (iii) a tRNA that can be acylated with the NNM by the PylRS (i.e., the tRNA is a cognate tRNA of the PylRS). In some embodiments, a subject method also includes providing a DNA or mRNA comprising a nucleotide sequence encoding a POI, wherein the nucleotide sequence encoding the POI comprises one or more in-frame codons that are the reverse complement of the tRNA’s anticodon. As such, the POI that is produced includes one or more of the NNM that was inserted by the tRNA during translation (at a position(s) determined by the position(s) of the codon / anti-codon). The following will describe the various elements from above (e.g., the PylRS, the tRNA, the NNM, the codon / anti-codon), and such descriptions apply to both the methods and compositions of matter disclosed herein. pyrrolysyl-tRNA synthetase (PylRS)
[0090] The disclosure pertains to new pyrrolysyl-tRNA synthetase (pylRS) and tRNA pairs (see, e.g., Table 1), which facilitate the efficient incorporation of non-natural monomers (referred to herein as NNMs) into proteins in response to specific codons, e.g., in a host organism such as a bacterial cell or eukaryotic cell.
[0091] Aminoacyl tRNA synthetases (RSs) are enzymes capable of acylating (e.g., aminoacylating) (also referred to as ‘charging’) a tRNA with an amino acid or amino acid analog. Pyrrolysyl tRNA synthetases (PylRSs) are RS enzymes capable of acylating a tRNA (tRNAPyl) with pyrrolysine, but can also recognize and charge a cognate tRNA with many other NNMs.
[0092] In some embodiments, the subject pylRS-tRNA pairs are orthogonal, meaning they do not cross-react with the host's endogenous tRNA or synthetases, allowing for precise and site-specific incorporation of NNMs. Utility of the new PylRS proteins and tRNAs spans various applications, including protein engineering, the development of novel therapeutics, and the creation of proteins with enhanced or novel functions, such as increased stability, altered binding affinities, or introduction of post-translational modifications.
[0093] There are several amino acid tRNA synthetase / tRNA (aaRS / tRNA) pairs known in the art. Examples of PylRS / tRNAPylpairs from archaea include those from, e.g., Methanosarcina maze!, Methanosarcina barker! and Methanomethylophilusalvus. PylRS proteins contain a binding site that naturally recognizes pyrrolysine in its natural cellular setting, but can also recognize many other NNMs. PylRS enzymes have high substrate side chain promiscuity, low selectivity toward a-amine, and low selectivity toward the tRNA anticodon, making them very useful for many applications, such as protein engineering. For example, there are over 200 NNMs known that can be used with PylRS / tRNAPylpairs.
[0094] The present disclosure provides new pyrrolysyl-tRNA synthetase (pylRS) proteins and cognate tRNAs (i.e., pylRS / tRNA pairs) - see Table 1.
[0095] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77. In some cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs: 1-10 and 25-77.
[0096] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10. In some cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs: 1-10.
[0097] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1 . In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1. Insome cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs: 10, 6, and 1.
[0098] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 10.
[0099] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 6.
[0100] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 1.
[0101] In some embodiments, a subject PylRS includes two separately encoded polypeptides, e.g., a PylSc and a PylSn (see, e.g., Table 1). As such, in some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% ormore, 99% or more, or 100%) identical to SEQ ID NO: 4. In some cases, a PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4. In some cases, a PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4. In some cases, a PylRS includes the amino acid sequence of SEQ ID NO: 3 and the amino acid sequence of SEQ ID NO: 4.
[0102] Likewise, in some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2. In some cases, a PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2. In some cases, a PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2. In some cases, a PylRS includes the amino acid sequence of SEQ ID NO: 1 and the amino acid sequence of SEQ ID NO: 2.
[0103] In some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3. In some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4. In some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1. In some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% ormore, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2.Variants
[0104] Generally, a “mutant” (also referred to as a “variant”) PylRS differs from the corresponding wildtype PylRS by having an amino acid mutation (addition, substitution and / or deletion) at one or more residues. A variant PylRS can improve function, e.g., improve stability, alter substrate specificity and / or enhance enzymatic activity. Variant PylRS proteins are known in the art and the mutations used to enhance the activities of other PylRS proteins can be used to inform mutations to enhance the activities of a subject PylRS (e.g., a PylRS of any one of SEQ ID NOs: 1- 10, or any one of SEQ ID NOs: 1-10 and 25-77). See, e.g., US patent publication Nos. US20230287383, 20220010296, 20220154237, 20180171321 , 20150148525, and 20220154237, as well as international patent publication Nos. WO 2018185222, and WO2023031445 (which are hereby incorporated by reference for their disclosures related to PylRS / tRNA pairs, including those related to PylRS variants) for non-limiting examples of PylRS variants. Also see, e.g., Koch et al., Front. Bioeng. Biotechnol. 2021, 9, 1-14; Hu et al., ACS Synth. Biol. 2020, 9 (10), 2723-2736; Wan et al., Biochim Biophys Acta. 2014 Jun; 1844(6): 1059-70; Koch et al., Chem Rev. 2024 Jul 2; and Guo et al., Protein Cell. 2024 May 7;15(5):331-363; Koch et al., Int. J. Mol. Sci. 2021, 22 (20), 11194; Koch et al., Methods in molecular biology (Clifton, N.J.) 2023, 2676, 3-19; and Gong et al., J Bacteriol. 2023 Feb 22;205(2):e0038522. Examples include, but are not limited to:• / WmPyIRS: (A302T / V346A / W348A / L401 / W417T); (L301M / Y306A / L309A / C348F); (A302I / N346T / C348I / Y384L / W471 K); (Y306L / L309A / N346A / C348M / W417T); (C348W / W417S); (Y306A / Y384F); (A302T / N346A / C348G / Y384F / W417T);(R91 K / G131 E / Y306A / Y384F); (Y306V / L309A / C348F / Y384F); (L309G / N346A / C348I / V401 K / W417I); (Y306A / Y384F);(N346Q / C348S / V401 G / W417T / G419G) ; (R61K / M300L / A302G / L309F / C348A / E444G); (F271L / F313C, F271 L / F313M, and F271N / F313I); (Y306L / C348I / Y384F);(A302S / L309M / I322L / N346A / W348G / W417T); (N346D / C348S / Y384F); (N346S / C348M / V401G); or (N346G / C348M / W417L)• / WbPyIRS: (L274
[0309] A / C313
[0348] F / Y349
[0384] F);(Y271
[0306] A / L274
[0309] A / C313
[0348] A) ;(Y271
[0306] C / N311
[0346] Q / Y349
[0384] F / V366
[0401] C); (C323
[0358] W / W382
[0417] T); (N311
[0346] Q / C313
[0348] S); (L270
[0305] F / L274
[0309] M / N311
[0346] A / C313
[0348] G); (L270
[0305] G / N311
[0346] G / C313
[0348] A);(N 311
[0346] S / C313
[0348] G / V366
[0401] A / W382
[0417] T) ;(A267
[0302] Q / N311
[0346] S / C313
[0348] W); (N311
[0346] Q / Y349
[0384] F);(L266
[0301] M / L270
[0305] l / L274
[0309] A / C313
[0348] F); (Y271
[0306] M / C313(348)T); (C313[348)T / Y349
[0384] F); (Y271
[0306] A); or(T 13
[0048] l / l 36
[0071] V / C313
[0348] W / W382
[0417] S)• / WaPyIRS:(Y126
[0306] G / M 129
[0309] A / V168
[0348] F / H227
[0405] T / Y228
[0406] P / L229
[0407] l) ; (Y126
[0306] L / M129
[0309] A / N166
[0346] A / V168
[0348] M / W239
[0417] T);(L125
[0305] F / N 166
[0346] A / V168
[0348] G) ; (Y126
[0306] A) ;(V168
[0348] G / A223
[0401] C / Y206
[0384] F);(Y126
[0306] T / M129
[0309] R / V168
[0348] H / H227
[0405] l / Y228
[0406] P; or Y126
[0306] T / M129
[0309] R / V168
[0348] H / Y206
[0384] W)• GIPyIRS (PylRS from methanogenic archaeon ISO4-G1): (V167
[0348] G / A221
[0401] C / Y204
[0384] F); or(L124
[0305] G / Y125
[0306] F / N 165
[0346] G / V167
[0348] F / Y204
[0384] W / A221
[0401] G / W237
[0417] Y)• DfiPyIRS (PylRS from Desulfitobacterium hafniense): (N176
[0367] A / T178
[0368] G)• ChPyIRS (chimeric PylRS consisting of PylSn / WbPyIRS and / WmPyIRS PylSc): (C348G / V401C / Y384F); or (V401 K / Y384F)
[0105] One of ordinary skill in the art would understand that the sequences of known variants such as those listed above can be used to determine the corresponding amino acid residue(s) in another PylRS (e.g., in any one of the PylRSs of SEQ ID NOs: 1-10, or any one of the PylRSs of SEQ ID NOs: 1-10 and 25-77). For example, the amino acid sequences can be aligned using convenient alignment software based on primary sequence and / or structural information. For example, identification of a corresponding amino acid residue in another PylRS can be determined by an alignment of multiple polypeptide sequences using several computer programs including, but not limited to, MUSCLE (multiple sequence comparison by log-expectation; version 3.5 or later; Edgar, Nucleic Acids Res (2004), 32 (5): 1792-1797 Edgar, Nucleic Acids Res (2004), 32 (5): 1792-1797), MAFFT (version 6.857 or later; Katoh et al., Nucleic Acids Res (2002), 30 (14): 3059-3066; Katoh et al., Nucleic Acids Res (2005), 33 (2): 511-518;Katoh et al., Bioinformatics (2007), 23 (3): 372-374; Katoh et al., Methods Mol Biol (2009), 537: 39-64; Katoh et al., Bioinformatics (2010), 26 (15): 1899-1900), EMBOSS EMMA employing ClustalW (1.83 or later; Thompson et al., Nucleic Acids Res (1994), 22 (22): 4673-4680), and the like.
[0106] In some cases, a PylRS is a variant of any one of SEQ ID NOs: 1-10 and 25-77, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, or 99% or more) identical to any one of SEQ ID NOs: 1-10 and 25-77, but it includes one or more amino acid mutations relative that corresponding wild type sequence.
[0107] In some cases, a PylRS is a variant of any one of SEQ ID Nos: 1-10, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, or 99% or more) identical to any one of SEQ ID NOs: 1-10, but it includes one or more amino acid mutations relative that corresponding wild type sequence.
[0108] For example, in some cases, a PylRS is a variant of SEQ ID No: 1 , meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 1 , but it includes one or more amino acid mutations relative to SEQ ID NO: 1 . In some cases, a PylRS is a variant of SEQ ID No: 2, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 2, but it includes one or more amino acid mutations relative to SEQ ID NO: 2. In some cases, a PylRS is a variant of SEQ ID No: 3, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 3, but it includes one or more amino acid mutations relative to SEQ ID NO: 3. In some cases, a PylRS is a variant of SEQ ID No: 4, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 4, but it includes one or more amino acid mutations relative to SEQ ID NO: 4. In some cases, a PylRS is a variant of SEQ ID No: 5, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 5, but it includes one or more amino acid mutations relative to SEQ ID NO: 5. In some cases, a PylRS is a variant of SEQ ID No: 6, meaning it includes an amino acid sequence that is 80% or more (e.g.,85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 6, but it includes one or more amino acid mutations relative to SEQ ID NO: 6. In some cases, a PylRS is a variant of SEQ ID No: 7, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 7, but it includes one or more amino acid mutations relative to SEQ ID NO: 7. In some cases, a PylRS is a variant of SEQ ID No: 8, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 8, but it includes one or more amino acid mutations relative to SEQ ID NO: 8. In some cases, a PylRS is a variant of SEQ ID No: 9, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 9, but it includes one or more amino acid mutations relative to SEQ ID NO: 9. In some cases, a PylRS is a variant of SEQ ID No: 10, meaning it includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more) identical to SEQ ID NO: 10, but it includes one or more amino acid mutations relative to SEQ ID NO: 10. tRNA
[0109] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18 and 81-97. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18 and 81-97. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11- 18 and 81-97. In some cases, a subject tRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 11-18 and 81-97.
[0110] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%)identical to any one of SEQ ID NOs: 11-18. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, a subject tRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 11-18.
[0111] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject tRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 18, 14, and 11.
[0112] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject tRNA comprises the nucleotide sequence of SEQ ID NO: 18.
[0113] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject tRNA comprises the nucleotide sequence of SEQ ID NO: 14.
[0114] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, asubject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a subject tRNA comprises the nucleotide sequence of SEQ I D NO: 11.
[0115] In some embodiments, a subject tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a subject tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a subject tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a subject tRNA comprises the nucleotide sequence of SEQ ID NO: 12.
[0116] Table 1. PylRS / tRNA pairs of the disclosure
[0117] The term “selector codon” as used herein refers to a codon (of an mRNA) that is recognized by (i.e., that hybridizes with) a subject tRNA (via the tRNA’s anticodon) in the translation process and is not recognized by endogenous tRNAs, e.g., of a bacterial or eukaryotic cell. The term is also used for the corresponding codons in polypeptide-encoding sequences of polynucleotides that are not mRNAs (e.g., DNA such as plasmid DNA). The anticodon of a subject tRNA binds to a selector codonwithin an mRNA and thus incorporates the NNM site-specifically into the growing chain of the polypeptide encoded by said mRNA.
[0118] The known 64 genetic (triplet) codons code for 20 amino acids and three stop codons. Because only one stop codon is needed for translational termination, the other two can in principle be used to encode for NNM insertion. For example, the amber codon, UAG, has been successfully used as a selector codon in in vitro translation systems and in cells to direct the incorporation of NNMs. Selector codons utilized in methods of the present disclosure expand the genetic codon framework of the protein biosynthetic machinery of the translation system used. Specifically, selector codons include, but are not limited to, nonsense codons, such as stop codons, e.g., amber (UAG), ocher (UAA), and opal (UGA) codons; codons consisting of more than three bases (e.g., four base codons); and codons derived from natural or unnatural base pairs. For a given system, a selector codon can also include one of the natural three base codons (i.e. , natural triplets), wherein the translation system being used does not use (or only scarcely uses) said natural triplet, e.g., a system that is lacking a tRNA that recognizes the natural triplet or a system in which the natural triplet is a rare codon.
[0119] A subject tRNA may be mutated / engineered to exhibit a desired feature, e.g., an altered anticodon-codon such that it will recognize a given codon with reduced or increased efficiency, or such that it will recognize a different codon. Such altered anticodon-codons may be such that specific codons are recognized more specifically by the tRNA's anticodon (i.e. higher base-to-base binding specificity, less wobblebase pairing abilities), e.g., either by altering the anticodon itself or by adapting the special structure of the tRNA such as to alter the binding behavior. Methods for engineering tRNAs in this context are known in the art and include those as described in, e.g., Liu et al., PNAS (1997), 94: 10092-10097; Wang and Schultz, Chem Biol (2001), 8: 883-890; and Maranhao et al., ACS Synth Biol (2017), 6 (1): 108-119. Also, tRNAs may be engineered such as to enhance the affinity of acylated tRNA to the EF (elongation factor) Tu (cf., e.g., Schrader et al., PNAS (2011), 108 (13): 5215-5220).
[0120] In some embodiments, the anticodon of a subject tRNA recognizes the codon UAG. In some embodiments, the anticodon of a subject tRNA recognizes the codon UAA (e.g., via mutation of any one of SEQ ID NOs: 11-18, or any one of SEQ ID NOs: 11-18 and 81-97). In some embodiments, the anticodon of a subject tRNA recognizes the codon UGA (e.g., via mutation of any one of SEQ ID NOs: 11-18, or any one of SEQ ID NOs: 11-18 and 81-97).
[0121] In some cases, a subject PylRS / tRNA pair is orthogonal not only to the translation system of interest (e.g., a host cell of interest such as a bacterial or eukaryotic host cell), but is also orthogonal relative to another PylRS / tRNA pair, e.g., one known in the art, or e.g., two subject PylRS / tRNA pairs can be used simultaneously. In such cases, the two different PylRS / tRNA pairs that are used can be referred to as mutually orthogonal. In such cases, one PylRS / tRNA pair would recognize one particular codon (e.g., UAA) and the other would recognize a different codon (e.g., UAG). As noted above, either (or both) of the tRNAs can be engineered to recognize the desired codon.
[0122] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18 and 81-97. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77 and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11- 18 and 81-97. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and 25-77 and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18 and 81-97. In some cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs: 1-10 and 25-77 and the tRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 11-18 and 81-97.
[0123] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% ormore, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 1-10 and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 11-18. In some cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs:I-10 and the tRNA comprises the nucleotide sequence of any one of SEQ ID NOs:I I-18.
[0124] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1 , and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1 , and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 10, 6, and 1 , and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to any one of SEQ ID NOs: 18, 14, and 11. In some cases, a subject PylRS includes the amino acid sequence of any one of SEQ ID NOs: 10, 6, and 1 and the tRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 18, 14, and 11 .
[0125] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject PylRS includes an aminoacid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10, and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 10, and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 18. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 10 and the tRNA comprises the nucleotide sequence of SEQ ID NO: 18.
[0126] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6, and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 6, and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 14. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 6 and the tRNA comprises the nucleotide sequence of SEQ ID NO: 14.
[0127] In some embodiments, a subject PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 , and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a subject PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 , and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% ormore, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a subject PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 , and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a subject PylRS includes the amino acid sequence of SEQ ID NO: 1 and the tRNA comprises the nucleotide sequence of SEQ ID NO: 11.
[0128] In some embodiments, a subject PylRS includes two separately encoded polypeptides, e.g., a PylSc and a PylSn (see, e.g., Table 1). As such, in some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4 and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 3 and an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 4, and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 12. In some cases, a PylRS includes the amino acid sequence of SEQ ID NO: 3 and the amino acid sequence of SEQ ID NO: 4 and the tRNA comprises the nucleotide sequence of SEQ ID NO: 12.
[0129] in some cases, a PylRS includes an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% ormore, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2, and the tRNA comprises a nucleotide sequence that is 80% or more (e.g., 85% or more, 90% for more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a PylRS includes an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2, and the tRNA comprises a nucleotide sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a PylRS includes an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 1 and an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 2, and the tRNA comprises a nucleotide sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 11. In some cases, a PylRS includes the amino acid sequence of SEQ ID NO: 1 and the amino acid sequence of SEQ ID NO: 2 and the tRNA comprises the nucleotide sequence of SEQ ID NO: 11.Non-natural monomers (NNM)
[0130] The term “non-canonical amino acid” (abbreviated ncAA) is used herein to refer to an amino acid that is not one of the 20 canonical a-amino acids incorporated into proteins during translation. Thus, Selenocysteine (Sec) and pyrrolysine (Pyl) would be considered herein to be ncAAs (they are non-canonical a-amino acids). [3-amino acids such as p3-amino acids and p2-amino acids are also considered herein to be ncAAs.
[0131] The terms “non-natural monomer” (abbreviated “NNM”) and “unnatural amino acid” (abbreviated “UNAA”) are used interchangeably herein to refer to any molecule that: (1) is not one of the 20 canonical a-amino acids incorporated into proteins during translation, and (2) can be used as a substitute for an amino acid during translation (via acylation to a tRNA by a tRNA synthetase). As an example, in some cases, an NNM is not an amino acid, but can be used as a substitute for an amnio acid during translation (i.e. , the molecule can be conjugated a tRNA via acylation by a tRNA synthetase). Examples of such NNMs include, but are not limited to hydroxy acids (e.g., a-hydroxy acids, p2-hydroxy acids). As another example, in some cases, an NNM is an amino acid, but is not an a-amino acid, and can be used as a substitute foran amnio acid during translation. Examples of such NNMs include, but are not limited to p-amino acids (e.g., p3-amino acids, p2-amino acids). Yet another example is a derivatized canonical amino acid, e.g., one that is photo- or chemically caged.
[0132] Thus, the term “NNM” encompasses ncAAs as well as analogs of a-amino acids (e.g., a-hydroxy acids, p2-hydroxy acids, p3-amino acids, and p2-amino acids). When translationally incorporated into a protein of interest (POI), an NNM can yield residues which are different from the amino acid residues corresponding to the 20 canonical amino acids. An NNM can be recognized and used as a substrate by a subject PylRS. In other words, a PylRS of the present disclosure can acylate its cognate tRNA (the tRNA with which it forms an orthologous pair) with an NNM, and the NNM can thereby be incorporated into a POI.
[0133] Examples of NNMs useful in methods and compositions (e.g., systems, kits, and the like) of the present invention will be known to one of ordinary skill in the art (see, e.g., Liu et al., Annu Rev Biochem 83:379-408, 2010, Lemke, ChemBioChem 15:1691- 1694, 2014). Figures 15-22 depict examples of NNMs that have been successfully used to acylate tRNAs using PylRS / tRNA pairs. See, e.g., Wan et al., Biochim Biophys Acta. 2014 Jun;1844(6):1059-70 (e.g., figures 7-10 of this reference); Koch et al., Chem Rev. 2024 Jul 2 (e.g., figures 5-8 of this reference); and Guo et al., Protein Cell. 2024 May 7;15(5):331-363.
[0134] Amino acid analogs such as a-hydroxy acid, p2-hydroxy acid, p3-amino acid, and p2- amino acid analogs are known in the art to substitute for amino acids for tRNA acylation by tRNA synthetases, e.g., PylRS enzymes. See, e.g., Fricke et al., Nat. Chem. 15, 960-971 (2023); Hamlish et al., ACS Cent Sci. 2024 Apr 23; 10(5): 1044- 1053; Dunkelmann et al., Nature 625, 603-610 (2024); Schissel et al., ChemRxiv. 2024; doi:10.26434 / chemrxiv-2024-bkzp3; and Schepartzet al., bioRxiv 2023.10.03.560714; doi: doi.org / 10.1101 / 2023.10.03.560714.
[0135] In some cases, the NNM is selected from those depicted in Figures 15-22. In some cases, the NNM is selected from those depicted in Figures 15-22, or is a a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figures 15-22, or is a a-hydroxy acid, p2-hydroxy acid, P3-amino acid, or p2-amino acid analog thereof. In some cases, the NNM is a lysine derivative selected from those depicted in Figures 15-22. In some cases, the NNM is a lysine derivative selected from those depicted in Figures 15-22, or is a a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. In some cases, the NNM is a lysine derivative selected from those depicted in Figures 15-22, or is a a-hydroxy acid,2-hydroxy acid, p3-amino acid, or p2-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figure 15. In some cases, the NNM is selected from those depicted in Figure 15, or is a a-hydroxy acid, p2-hydroxy acid, or P3-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figure 15, or is a a-hydroxy acid, p2-hydroxy acid, p3-amino acid, or p2- amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figures 15-18. In some cases, the NNM is selected from those depicted in Figures 15-18, or is a a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figures 15-18, or is a a- hydroxy acid, p2-hydroxy acid, p3-amino acid, or p2-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figures 19-22. In some cases, the NNM is selected from those depicted in Figures 19-22, or is a a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figures 19-22, or is a a-hydroxy acid, p2-hydroxy acid, P3-amino acid, or p2-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figure 19. In some cases, the NNM is selected from those depicted in Figure 19, or is a a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. In some cases, the NNM is selected from those depicted in Figure 19, or is a a-hydroxy acid, p2-hydroxy acid, p3-amino acid, or p2-amino acid analog thereof.
[0136] Examples of NNMs that can be used include, but are not limited to: post-translational modification mimics (e.g., photo or chemically caged or noncaged), NNMs containing bioorthogonal groups that are possible targets for site-specific bioconjugation, photo or chemically caged NNMs or NAAs (Natural amino acids), NNMs containing functional groups which are useful for spectroscopic applications, photo- or proximity triggered cross-linkable NNMs, fluorescent NNMs.
[0137] Examples include, but are not limited to: photoactivatable amino acids (amino acids that can be activated by light, facilitating controlled photo-crosslinking or photocaging), fluorescent amino acids (amino acids containing fluorescent moieties, e.g., for real-time tracking and imaging of proteins), bioorthogonal amino acids (amino acids that facilitate bioorthogonal chemistry, e.g., for selective protein labeling or modification without interfering with native cellular processes). unnatural side chain amino acids (amino acids with unique side chains that provide, e.g., novel chemical functionalities not found in the 20 canonical amino acids), post-translational modification mimics (amino acids that mimic natural post-translational modifications such as phosphorylation, acetylation, or glycosylation), metal-binding amino acids(amino acids designed to chelate metal ions, e.g., useful in metalloenzyme engineering or as catalytic sites in proteins), redox-active amino acids (amino acids capable of participating in redox reactions, e.g., useful for studying oxidative stress and redox biology), isotopically labeled amino acids (amino acids containing stable isotopes (e.g.,A13C,A15N) for use, e.g., in NMR spectroscopy or mass spectrometry), sterically restricted amino acids (amino acids with bulky side chains that induce conformational constraints in proteins, e.g., aiding in the study of protein folding and dynamics), cleavable amino acids (amino acids that can be cleaved under specific conditions, allowing, e.g., for controlled degradation or activation of proteins), and enzyme inhibitor amino acids (amino acids designed to inhibit specific enzymes when incorporated into a protein, useful, e.g., for studying enzyme function and regulation).
[0138] Examples of lysine derivatives that can be used as a NNM include, but are not limited to those depicted in Figure 15. Examples of lysine derivatives that can be used as a NNM include, but are not limited to: 1. NE-Acetyl-L-lysine (AcK) (a lysine derivative that mimics the acetylation of lysine residues, useful in studying protein acetylation and its effects on function); 2. Ns-Methyl-L-lysine (MeK) (a methylated lysine analog, employed in the study of histone modifications and protein-protein interactions); 3. NE- Benzoyl-L-lysine (BzK) (a lysine derivative with a benzoyl group, useful for sitespecific protein modification and studying lysine modifications); 4. Ns-Carbamoyl-L- lysine (CbK) (a lysine analog with a carbamoyl group, which can be used to mimic lysine carbamylation in proteins, which can be used for studying post-translational modifications); 5. Ns-Biotinyl-L-lysine (BioK) (a biotinylated lysine derivative that facilitates site-specific biotinylation of proteins for affinity purification and detection); 6. Ns-Azidoacetyl-L-lysine (AzK) (a lysine derivative with an azido group, facilitating bioorthogonal click chemistry for protein labeling and conjugation); 7. NE- Trifluoroacetyl-L-lysine (TfaK) (a trifluoroacetylated lysine analog, which can be used in studies of protein stability and interactions involving fluorinated residues); 8. NE- Alkyne-L-lysine (AlkK) (a lysine derivative with an alkyne group, facilitating bioorthogonal conjugation via click chemistry); 9. Ns-Malony-L-lysine (MalK) (a lysine analog with a malonyl group, which can be used to study the effects of lysine malonylation on protein function); 10. Nc-Trimethyl-L-lysine (TML) (a trimethylated lysine derivative, which is valuable in researching the role of lysine methylation in epigenetic regulation); 11. Ns-Crotonyl-L-lysine (CrK) (a lysine derivative that mimics crotonylation, which can be used in the study of this post-translational modificationand its impact on protein function); 12. Nc-Boc-L-lysine (BocK) (a lysine analog with a Boc (tert-butyloxycarbonyl) protecting group, useful for controlled deprotection and functionalization in protein engineering); 13. Nc-Formyl-L-lysine (ForK) (a lysine derivative that mimics formylation, useful in studying protein-DNA interactions and formylation's role in cellular processes); 14. Nc-Nitro-L-lysine (NO2K) (a lysine derivative with a nitro group, which can be used in studies involving nitration and its effects on protein structure and function); 15. Nc-Dansyl-L-lysine (DnsK) (a fluorescent lysine analog, useful for studying protein structure, dynamics, and interactions through fluorescence spectroscopy); 16. Nc-Palmitoyl-L-lysine (PalmK) (a lysine derivative that mimics palmitoylation, allowing for the study of lipid modifications on proteins and their functional implications); 17. Ns-Hydroxy-L-lysine (HylK) (a lysine derivative with a hydroxy group, involved in collagen formation and which can be used in the study of protein hydroxylation); 18. Nc-Succinyl-L-lysine (SucK) (a lysine analog with a succinyl group, which can be used for studying the effects of lysine succinylation on protein function); 19. Nc-Malonyl-L-lysine (MalK) (a lysine derivative with a malonyl group, which can be used in research on lysine malonylation as a post- translational modification); 20. Nc-Farnesyl-L-lysine (FarK) (a lysine derivative that mimics farnesylation, which can be used to study the role of lipid modifications in protein localization and function); and 21. BocK-OH.
[0139] In other words, Nc-Acetyl-L-lysine (AcK), Ns-Methyl-L-lysine (MeK), Nc-Benzoyl-L- lysine (BzK), Nc-Carbamoyl-L-lysine (CbK), Nc-Biotinyl-L-lysine (BioK), NE- Azidoacetyl-L-lysine (AzK), NE-Trifluoroacetyl-L-lysine (TfaK), NE-Alkyne-L-lysine (AlkK), Ns-Malony-L-lysine (MalK), Ns-Trimethyl-L-lysine (TML), Ns-Crotonyl-L-lysine (CrK), Ns-Boc-L-lysine (BocK), BocK-OH, Nc-Formyl-L-lysine (ForK), Ns-Nitro-L-lysine (NO2K), Ns-Dansyl-L-lysine (DnsK), Ns-Palmitoyl-L-lysine (PalmK), Ns-Hydroxy-L- lysine (HylK), Ns-Succinyl-L-lysine (SucK), Ns-Malonyl-L-lysine (MalK), and NE- Farnesyl-L-lysine (FarK).
[0140] Examples of lysine derivatives that can be used as a NNM include, but are not limited to: H-Lys(Boc)-OH (“Boc”), Cyclooctyne - Lysine (“SCO”), Bicyclo [6.1.0] nonyne - Lysine (“BCN”), trans-Cyclooct-2-en - Lysine (“TCO*A”) and trans-Cyclooct-4-en - Lysine (“TCO-E”), for example described in Reinkemeier et al., Eur J Chem (2021) 27 (19) 6094-6099. These NNMs can react via click chemistry with tetrazines, which is a very fast and biorthogonal reaction.
[0141] An NNM may include a group (sometimes referred to as “labeling group”) that facilitates reaction with a suitable group (herein referred to as “docking group”) ofanother molecule (herein termed a “conjugation partner’’) so as to covalently attached the conjugation partner molecule to the NNM. When a NNM comprising a labeling group is translationally incorporated into a POI, the labeling group becomes part of the POI. Accordingly, a POI prepared according to the method of the present disclosure can be reacted with one or more than one conjugation partner molecule such that the conjugation partner molecules bind covalently to the (labeling groups of the) NNM residue(s) of the POI. Such conjugation reactions may be used for in situ coupling of POIs within a cell or tissue expressing the POI, or for site-specific conjugation of isolated or partially isolated POIs.
[0142] Particular choices useful for combinations of labeling groups and docking groups (of conjugation partner molecules) are those, which can react by metal-free click reactions. Such click reactions include strain-promoted inverse-electron-demand Diels-Alder cycloadditions (SPIEDAC; see, e.g., Devaraj et al., Angew Chem Int Ed Engl 2009, 48:7013)) as well as cycloadditions between strained cycloalkynyl groups, or strained cycloalkynyl analog groups having one or more of the ring atoms not bound by the triple bond substituted by amino groups), with azides, nitrile oxides, nitrones and diazocarbonyl reagents (see, e.g., Sanders et al., J Am Chem Soc 2010, 133:949; Agard et al., J Am Chem Soc 2004, 126:15046), for example strain promoted alkyne-azide cycloadditions (SPAAC). Such click reactions allow for ultrafast and biorthogonal covalent site-specific coupling of NNM labeling groups of POIs with suitable groups of coupling partner molecule.
[0143] Pairs of docking and labeling groups which can react via the above-mentioned click reactions are known in the art. Examples of suitable NNMs including docking groups include, but are not limited to, the NNMs described, e.g., in WO 2012 / 104422 and WO 2015 / 107064.
[0144] Examples of particular suitable pairs of docking groups (understood by the conjugation partner molecule) and labeling groups (understood by the NNM residue(s) of the POI) include but are not limited to: (a) a docking group comprising (or essentially consisting of) a group selected from an azido group, a nitrile oxide functional group (i.e. a radical of formula, a nitrone functional group or a diazocarbonyl group, combined with a labeling group comprising (or essentially consisting of) an optionally substituted strained alkynyl group (such groups can react covalently in a copper-free strain promoted alkyne-azide cycloaddition (SPAAC)); (b) a docking group comprising (or essentially consisting of) an optionally substituted strained alkynyl group, combined with a labeling group comprising (or essentiallyconsisting of) a group selected from an azido group, a nitrile oxide functional group (i.e. a radical of formula, a nitrone functional group or a diazocarbonyl group (such groups can react covalently in a copper-free strain promoted alkyne-azide cycloaddition (SPAAC)); (c) a docking group comprising (or essentially consisting of) a group selected from optionally substituted strained alkynyl groups, optionally substituted strained alkenyl groups and norbornenyl groups, combined with a labeling group comprising (or essentially consisting of) an optionally substituted tetrazinyl group ( such groups can react covalently in a copper-free strain promoted inverseelectron-demand Diels-Alder cycloaddition (SPIEDAC)); (d) a docking group comprising (or essentially consisting of) an optionally substituted tetrazinyl group, combined with a labeling group comprising (or essentially consisting of) a group selected from optionally substituted strained alkynyl groups, optionally substituted strained alkenyl groups and norbornenyl groups ( such groups can react covalently in a copper-free strain promoted inverse-electron-demand Diels-Alder cycloaddition (SPIEDAC)).
[0145] Optionally substituted strained alkynyl groups include, but are not limited to, optionally substituted trans-cyclooctenyl groups. Optionally substituted strained alkenyl groups include, but are not limited to, optionally substituted cyclooctynyl groups, such as those described in WO 2012 / 104422 and WO 2015 / 107064. Optionally substituted tetrazinyl groups include, but are not limited to, those described in WO 2012 / 104422 and WO 2015 / 107064.
[0146] The NNMs used in the context of the present disclosure can be used in the form of their salt. Salts of an NNM as described herein mean acid or base addition salts, especially addition salts with physiologically tolerated acids or bases. Physiologically tolerated acid addition salts can be formed by treatment of the base form of an NNM with appropriate organic or inorganic acids. NNMs containing an acidic proton may be converted into their non-toxic metal or amine addition salt forms by treatment with appropriate organic and inorganic bases. The NNMs and salts thereof described in the context of the present disclosure also include the hydrates and solvent addition forms thereof, e.g., hydrates, alcoholates and the like.
[0147] Physiologically tolerated acids or bases are in particular those which are tolerated by the translation system used for preparation of POI with NNM residues, e.g. are substantially non-toxic to living bacterial cells or eukaryotic cells.
[0148] NNMs, and salts thereof, useful in the context of the present disclosure can be prepared by analogy to methods, which are well known in the art and are described, e.g., in the various publications cited herein.
[0149] The nature of the coupling partner molecule depends on the intended use. For example, the POI may be coupled to a molecule suitable for imaging methods or may be functionalized by coupling to a bioactive molecule. For instance, in addition to the docking group, a coupling partner molecule may comprise a group that selected from, but are not limited to, dyes (e.g. fluorescent, luminescent, or phosphorescent dyes, such as dansyl, coumarin, fluorescein, acridine, rhodamine , silicon-rhodamine, BODIPY, or cyanine dyes), molecules able to emit fluorescence upon contact with a reagent, chromophores (e.g., phytochrome, phycobilin, bilirubin, etc.), radiolabels (e.g. radioactive forms of hydrogen, fluorine, carbon, phosphorous , sulfur, or iodine, such as tritium,18F,11C,14C,32P,33P,33S,35S,11In,1251,1231,1311,212B,90Y or186Rh), MRI-sensitive spin labels, affinity tags (e.g. biotin, His-tag, Flag-tag, strep- tag, sugars, lipids, sterols, PEG-linkers, benzylguanines, benzylcytosines, or cofactors), polyethylene glycol groups (e.g., a branched PEG, a linear PEG, PEGs of different molecular weights, etc.), photocrosslinkers (such as p-azidoiodoacetanilide), NMR probes, X-ray probes, pH probes, IR probes, resins, solid supports and bioactive compounds (e.g. synthetic drugs). Suitable bioactive compounds include, but are not limited to, cytotoxic compounds (e.g., cancer chemotherapeutic compounds), antiviral compounds, biological response modifiers (e.g., hormones, chemokines, cytokines, interleukins, etc.), microtubule affecting agents, hormone modulators, and steroidal compounds. Specific examples of useful coupling partner molecules include, but are not limited to, a member of a receptor / ligand pair; a member of an antibody / antigen pair; a member of a lectin / carbohydrate pair; a member of an enzyme / substrate pair; biotin / avidin; biotin / streptavidin and digoxin / antidigoxin.
[0150] The ability of certain (labeling groups of) NNM residues to be coupled covalently in situ to (the docking groups of) conjugation partner molecules, in particular by a click reaction as described herein, can be used for detecting a POI having such NNM residue( s) within a eukaryotic cell or tissue expressing the POI, and for studying the distribution and fate of the POIs. Specifically, the method of the present disclosure for preparing a POI by expression in bacterial cells or eukaryotic cells can be combined with super-resolution microscopy (SRM) to detect the POI within the cell or a tissue of such cells.
[0151] Several SRMs methods are known in the art and can be adapted so as to utilize click chemistry for detecting a POI expressed by a eukaryotic cell of the present disclosure. Specific examples of such SRM methods include DNA-PAINT (DNA point accumulation for imaging in nanoscale topography; described, e.g., by Jungmann et al., Nat Methods 11:313-318, 2014), dSTORM (direct stochastic optical reconstruction microscopy) and STED (stimulated emission depletion) microscopy.Nucleic Acids
[0152] In some embodiments, a subject PylRS is provided (e.g., to a bacterial or eukaryotic cell or to an in vitro translation system) as a protein (i.e., not as a nucleic acid encoding the protein). However, in some embodiments, a subject PylRS is provided (e.g., to a bacterial or eukaryotic cell or to an in vitro translation system) as a nucleic acid encoding the PylRS. In some such cases, the nucleic acid is a DNA (e.g., a vector such as a plasmid, a viral vector, a minicircle, a bacteriophage, and the like), e.g., which can be introduced into a cell. In some cases, the nucleic acid is an mRNA, e.g., an mRNA can be provided to an in vitro translation system or can be introduced into a cell.
[0153] Likewise, in some embodiments, a subject tRNA is provided (e.g., to a bacterial or eukaryotic cell) as a DNA encoding the tRNA (e.g., a vector such as a plasmid, a viral vector, a minicircle, and the like), in some embodiments, a subject tRNA is provided (e.g., to a bacterial or eukaryotic cell or to an in vitro translation system) as an RNA.
[0154] Thus, the present disclosure provides a nucleic acid (e.g., DNA or mRNA) comprising a nucleotide sequence encoding a subject pyrrolysyl-tRNA synthetase (PylRS) (e.g., one that includes an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, one that includes an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10 and 25-77). In some cases the nucleic acid is a DNA (e.g., a vector such as a plasmid, a viral vector, a bacteriophage, a minicircle, and the like). In some cases the nucleic acid is an mRNA. The present disclosure provides a nucleic acid (e.g., DNA) comprising a nucleotide sequence encoding a subject tRNA (e.g., one that includes a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18, one that includes a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18 and 81-97). In some cases, the nucleic acid is a vector such as a plasmid, a viral vector, a bacteriophage, a minicircle, and the like.
[0155] Suitable expression vectors include viral expression vectors (e g. viral vectors based on vaccinia virus; poliovirus; adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:25432549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS 92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:1088 1097, 1999; WO 94 / 12649, WO 93 / 03769; WO 93 / 19191 ; WO 94 / 28938; WO 95 / 11984 and WO 95 / 00655); adeno-associated virus (AAV) (see, e.g., Ali et al., Hum Gene Ther 9:81 86, 1998, Flannery et al., PNAS 94:69166921 , 1997; Bennett et al., Invest Opthalmol Vis Sci 38:28572863, 1997; Jomary et al., Gene Ther 4:683 690, 1997, Rolling et al., Hum Gene Ther 10:641 648, 1999; Ali et al., Hum Mol Genet 5:591 594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828; Mendelson et al., Virol. (1988) 166:154-165; and Flotte et al., PNAS (1993) 90:10613-10617); SV40; herpes simplex virus; human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319 23, 1997; Takahashi et al., J Virol 73:7812 7816, 1999); a retroviral vector (e.g., Murine Leukemia Virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus); and the like. In some cases, a recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some cases, a recombinant expression vector of the present disclosure is a recombinant lentivirus vector. In some cases, a recombinant expression vector of the present disclosure is a recombinant retroviral vector.
[0156] Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. may be used in the expression vector. The transcription control element can be a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell type-specific promoter. In some cases, the transcription control element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcription control element can be functional in eukaryotic cells.
[0157] The transcription control element can be a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter isa tissue-specific promoter. In some cases, the promoter is a cell type-specific promoter. In some cases, the transcription control element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcription control element can be functional in eukaryotic cells.
[0158] A promoter can be a constitutively active promoter (i.e. , a promoter that is constitutively in an active / ”ON” state), it may be an inducible promoter (i.e., a promoter whose state, active / ”ON” or inactive / “OFF”, is controlled by an external stimulus, e.g., the presence of a particular temperature, compound, or protein.), it may be a spatially restricted promoter (i.e., transcription control element, enhancer, etc.)(e.g., tissue specific promoter, cell type specific promoter, etc.), and it may be a temporally restricted promoter (i.e., the promoter is in the “ON” state or “OFF” state during specific stages of embryonic development or during specific stages of a biological process, e.g., hair follicle cycle in mice).
[0159] Promoters can be derived from viruses and can therefore be referred to as viral promoters, or they can be derived from any organism, including prokaryotic or eukaryotic organisms. Promoters can be used to drive expression by any convenient RNA polymerase (e.g., pol I, pol II, pol III). In general, for eukaryotic systems, a pol II promoter is usually used to drive transcription an mRNA (one encoding a PylRS). Example promoters include, but are not limited to the SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); a herpes simplex virus (HSV) promoter, a cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CMVIE), a rous sarcoma virus (RSV) promoter, a beta-actin promoter, an EF1a promoter, an estrogen receptor-regulated promoter, and the like.
[0160] Examples of bacterial promoters include, but are not limited to: cos, tac, trp, tet, trp- tet, Ipp, lac, Ipp-lac, lacl q , T7, T5, T3, gal, tre, ara, rhaP (rhaP BAD )SP6, lambda- PR, the lambda-PL promoter, PenP, SPO1, and SPO2. Thus, promoters suitable for use with bacterial hosts include the beta-lactamase and lactose promoter systems, alkaline phosphatase, a tryptophan (trp) promoter system, and hybrid promoters such as the tac promoter. Yeast or fungal promoters include ADC1 , MFalpha, AC, P-60, CYC1 , Gal1 / 10, PHO5, GAP, TEF, rp28, ADH, PGK.
[0161] tRNA expression is usually under the control of Pol III promoter (e.g., a tRNA promoter). There are three types of pol III promoters: type 1 , type 2, and type 3. Type 1 pol III promoters include 5S rRNA promoters and initiation involves TFIIIA binding to the C Block intragenic 5S rRNA control sequence, serving as a platform to positionTFIIIC, and subsequent assembly of TFIIIB. Type 2 pol III promoters include tRNA promoters and transcription initiation involves TFIIIC binding to A and B Block intragenic regions, serving as a platform to position TFIIIB, which in turn, assembles Pol III at the transcription start site. Type 3 pol III promoters solely utilize upstream regulatory elements and do not utilize intragenic elements. Type 3 pol III promoters generally comprise a proximal sequence element (PSE), a TATA box, and an upstream distal sequence element (DSE). The DSE comprises at least one of an octamer (OCT) sequence and a Sphl postoctamer homology (SPH) sequence. The transcription factor Oct-1 binds the OCT sequence in the DSE and a SBF (SPH- binding factor) (also referred to as STAF (selenocysteine tRNA gene transcription activating factor) / SBF or zinc finger protein 143 (ZNF143)) transcription factor binds to the SPH sequence in the DSE. Assembly of the snRNA activating protein complex (SNAPc) to the PSE is stimulated by Octi and / or SBH binding to the DSE. SNAPc acts to assemble TFIIIB, which consists of a TATA-binding protein (TBP), a TFIIB- related factor (BRF2), and B-double-prime (BDP1), at the TATA box. TFIIIB, in turn, assembles Pol III at the start site of transcription. Type 3 pol III promoters, such as 7SK, U6, and H1 , can be used for the expression of small RNA, including short hairpin RNA (shRNA) and guide RNA (gRNA).
[0162] In some cases, a subject tRNA is operably linked to a Pol III promoter. In some cases, a subject tRNA is operably linked to a tRNA promoter. In some cases, a subject tRNA is operably linked to a U6 or H1 promoter.
[0163] In some embodiments, a nucleotide sequence encoding a subject PylRS is present on the same nucleic acid as the nucleotide sequence encoding a subject tRNA (e.g., plasmid, virus, bacteriophage, and the like). In some cases, the nucleotide sequences encoding a subject PylRS and a subject tRNA are present on different nucleic acids (e.g., different plasmids, viruses, bacteriophages, and the like).
[0164] Examples of inducible promoters include, but are not limited toT7 RNA polymerase promoter, T3 RNA polymerase promoter, Isopropyl-beta-D-thiogalactopyranoside (IPTG)-regulated promoter, lactose induced promoter, heat shock promoter, Tetracycline-regulated promoter, Steroid-regulated promoter, Metal-regulated promoter, estrogen receptor-regulated promoter, etc. Inducible promoters can therefore be regulated by molecules including, but not limited to, doxycycline; estrogen and / or an estrogen analog; IPTG; etc.
[0165] In some cases, a nucleotide sequence encoding a subject PylRS is operably linked to a promoter. In some cases, a nucleotide sequence encoding a subject tRNA isoperably linked to a promoter. In some cases, the promoter (for the PylRS and / or for the tRNA) is functional in bacterial cells (e.g., E. coli). In some cases, the promoter (for the PylRS and / or for the tRNA) is functional in eukaryotic cells (e.g., an insect cell, a mammalian cell, a mouse cell, a non-human primate cell, a human cell, and the like).
[0166] Inducible promoters suitable for use include any inducible promoter described herein or known to one of ordinary skill in the art. Examples of inducible promoters include, without limitation, chemically / biochemically-regulated and physically-regulated promoters such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracyclineresponsive promoter systems, which include a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)), steroid-regulated promoters (e.g., promoters based on the rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptors, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from metallothionein (proteins that bind and sequester metal ions) genes from yeast, mouse and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene or benzothiadiazole (BTH)), temperature / heat- inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light responsive promoters from plant cells).
[0167] In some cases, the promoter is a spatially restricted promoter (i.e., cell type specific promoter, tissue specific promoter, etc.) such that in a multi-cellular organism, the promoter is active (i.e., “ON”) in a subset of specific cells. Spatially restricted promoters may also be referred to as enhancers, transcription control elements, control sequences, etc. Any convenient spatially restricted promoter may be used as long as the promoter is functional in the targeted host cell (e.g., eukaryotic cell; prokaryotic cell).
[0168] In some cases, the promoter is a reversible promoter. Suitable reversible promoters, including reversible inducible promoters are known in the art. Such reversible promoters may be isolated and derived from many organisms, e.g., eukaryotes and prokaryotes. Modification of reversible promoters derived from a first organism for use in a second organism, e.g., a first prokaryote and a second a eukaryote, a first eukaryote and a second a prokaryote, etc., is well known in the art. Such reversible promoters, and systems based on such reversible promoters but also comprising additional control proteins, include, but are not limited to, alcohol regulated promoters(e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator proteins (AlcR), etc.), tetracycline regulated promoters, (e.g., promoter systems including TetActivators, TetON, TetOFF, etc.), steroid regulated promoters (e.g., rat glucocorticoid receptor promoter systems, human estrogen receptor promoter systems, retinoid promoter systems, thyroid promoter systems, ecdysone promoter systems, mifepristone promoter systems, etc.), metal regulated promoters (e.g., metallothionein promoter systems, etc.), pathogenesis-related regulated promoters (e.g., salicylic acid regulated promoters, ethylene regulated promoters, benzothiadiazole regulated promoters, etc.), temperature regulated promoters (e.g., heat shock inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter, etc.), light regulated promoters, synthetic inducible promoters, and the like.
[0169] Methods of introducing a nucleic acid (e.g., a nucleic acid encoding a subject PylRs and / or tRNA) into a host cell are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a cell. Suitable methods include e.g., viral infection, transfection, lipofection, nucleofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, transformation, bacterial conjugation, and the like (e.g., for viral vector delivery, cells can be contacted with viral particles comprising the subject viral expression vectors). For example, nucleic acids may be provided to the cells using well-developed transfection techniques; see, e.g. Angel and Yanik (2010) PLoS ONE 5(7): e11756, and the commercially available TransMessenger® reagents from Qiagen, Stemfect™ RNA Transfection Kit from Stemgent, TranslT®-mRNA Transfection Kit from Mirus Bio LLC, nucleofection, and the like. See also Beumer et al. (2008) PNAS 105(50):19821-19826.
[0170] Introducing a vector into cells can occur in vivo or can occur in any culture media and under any culture conditions that promote the survival of the cells. A vector into a target cell can be carried out in vivo or ex vivo or in vitro.
[0171] Retroviruses, for example, lentiviruses, are suitable for use in methods of the present disclosure. Commonly used retroviral vectors are “defective”, i.e. unable to produce viral proteins required for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate viral particles comprising nucleic acids of interest, the retroviral nucleic acids comprising the nucleic acid are packaged into viral capsids by a packaging cell line. Different packaging cell lines provide adifferent envelope protein (ecotropic, amphotropic or xenotropic) to be incorporated into the capsid, this envelope protein determining the specificity of the viral particle for the cells (ecotropic for murine and rat; amphotropic for most mammalian cell types including human, dog and mouse; and xenotropic for most mammalian cell types except murine cells). The appropriate packaging cell line may be used to ensure that the cells are targeted by the packaged viral particles. Methods of introducing subject vector expression vectors into packaging cell lines and of collecting the viral particles that are generated by the packaging lines are well known in the art. Nucleic acids can also introduced by direct micro-injection (e.g., injection of RNA).
[0172] Vectors used for providing the nucleic acids encoding a subject PylRS and / or a subject tRNA (e.g., to a target host cell or to an in vitro translation system) can include suitable promoters for driving the expression, that is, transcription activation, of the nucleic acid of interest. In other words, in some cases, the nucleic acid of interest will be operably linked to a promoter. This may include ubiquitously acting promoters, for example, the CMV-p-actin promoter, or inducible promoters, such as promoters that are active in particular cell populations or that respond to the presence of drugs such as tetracycline. In addition, vectors used for providing a nucleic acid encoding a subject PylRS and / or a subject tRNA to a cell may include nucleic acid sequences that encode for selectable markers in the target cells, so as to identify cells that have taken up the PylRS and / or a tRNA.
[0173] Nucleic acids may be prepared in order to adapt a nucleotide sequence to a specific expression system. As such, in some cases, a nucleotide sequence encoding a PylRS of the present disclosure is codon optimized. For example, bacterial expression systems or mammalian expression systems are known to more efficiently express polypeptides if amino acids are encoded by particular codons (which can be determined, e.g., using codon usage tables). Where appropriate, the nucleic acid sequences encoding the polypeptides described herein may be optimized for increased expression in a given host cell of interest. For example, nucleic acids of an embodiment herein may be synthesized using codons particular to a host for improved expression. This type of optimization can entail a mutation of a proteincoding nucleotide sequence (e.g., one encoding a subject PylRS) to mimic the codon preferences of the intended host organism or cell while encoding the same protein. Thus, the codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell was a mammalian cell (e.g., a human or mouse cell), a mammalian, human, or mouse (and the like) codon-optimized protein-encoding nucleotide sequence can be used. As another non-limiting example, if the intended host cell were a bacterial cell, then a codon-optimized protein-encoding nucleotide sequence can be generated that is designed for the particular bacterial type of interest, e.g., E. coli. As another non-limiting example, if the intended host cell were an insect cell, then an insect codon-optimized protein-encoding nucleotide sequence can be generated.
[0174] A protein of the present disclosure (e.g., a subject PylRS or a POI) may be produced in vitro (e.g., using an in vitro translation system or via chemical synthesis) or by eukaryotic cells or by bacterial cells, and may be isolated or purified therefrom. A protein may be further processed by unfolding, e.g. heat denaturation, dithiothreitol reduction, etc. and may be further refolded, using methods known in the art.
[0175] In some cases, a protein of interest (POI) is produce via an in vitro translation system as opposed to being produced in a cell. A subject PylRS and a subject tRNA would therefore be provided to the in vitro translation system (e.g., in DNA, RNA, or protein form). Any convenient translation system can be used, and various systems will be known to one of ordinary skill in the art. Examples include, but are not limited to: ribosome or mRNA display, E. coli cell-free expression systems, Rabbit Reticulocyte expression systems, HeLa expression systems, wheat germ extract expression systems, linked transcription / translation systems, coupled transcription / translation systems, and the like.
[0176] Modifications of interest that do not alter primary sequence include chemical derivatization of polypeptides, e.g., acylation, acetylation, carboxylation, amidation, etc. Also included are modifications of glycosylation, e.g. those made by modifying the glycosylation patterns of a polypeptide during its synthesis and processing or in further processing steps; e.g. by exposing the polypeptide to enzymes which affect glycosylation, such as mammalian glycosylating or deglycosylating enzymes. Also embraced are sequences that have phosphorylated amino acid residues, e.g. phosphotyrosine, phosphoserine, or phosphothreonine.
[0177] A protein of the present disclosure (e.g., a subject PylRS) may be prepared by in vitro synthesis, using conventional methods as known in the art. Various commercial synthetic apparatuses are available, for example, automated synthesizers by Applied Biosystems, Inc., Beckman, etc. By using synthesizers, naturally occurring amino acids may be substituted with unnatural amino acids. The particular sequence and the manner of preparation will be determined by convenience, economics, purity required, and the like.
[0178] If desired, various groups may be introduced into the peptide during synthesis or during expression, which allow for linking to other molecules or to a surface. Thus cysteines can be used to make thioethers, histidines for linking to a metal ion complex, carboxyl groups for forming amides or esters, amino groups for forming amides, and the like.
[0179] Also suitable for inclusion in embodiments of the present disclosure are nucleic acids and proteins of the present disclosure that have been modified using ordinary molecular biological techniques and synthetic chemistry so as to improve their resistance to proteolytic degradation, to change the target sequence specificity, to optimize solubility properties, to alter protein activity (e.g., transcription modulatory activity, enzymatic activity, etc.) or to render them more suitable. Analogs of such polypeptides include those containing residues other than naturally occurring L-amino acids, e.g. D-amino acids or non-naturally occurring synthetic amino acids. D-amino acids may be substituted for some or all of the amino acid residues.Cells
[0180] In some embodiments, a subject PylRS (e.g., as a protein or an mRNA encoding it) and a subject tRNA are provided to an in vitro translation system in order to generate an NNM-containing protein of interest. In some embodiments, a subject PylRS (e.g., as a protein, as an mRNA encoding it, or as a DNA encoding it) and a subject tRNA (e.g., as an RNA or as a DNA encoding it) are provided to a cell, e.g., introduced into a cell. Thus, provided are cells comprising a subject PylRS and / or a nucleic acid encoding the PylRS. Also provided are cells comprising a subject tRNA and / or a nucleic acid encoding it. Also provided are cells comprising (a) a subject PylRS and / or a nucleic acid encoding the PylRS; and (b) a subject tRNA and / or a nucleic acid encoding it. In some cases, the host cell is a bacterial cell, e.g., a bacterial cell can be used to produce an NNM-containing POI. A bacterial cell can also be used to propagate nucleic acids, e.g., one or more nucleic acid encoding a PylRS and / or a subject tRNA. In some cases, the host cell is a eukaryotic cell (e.g., an insect cell, a mammalian cell), e.g., such a cell can be used to produce an NNM-containing POI.
[0181] A host cell of the present disclosure may generally be any convenient host cell. Depending on the context, the term “host” can mean the wild-type host or a genetically altered, recombinant host or both. In principle, all bacterial or eukaryotic cells may be considered as host cells.
[0182] Suitable cells include, but are not limited to: a bacterial cell; a cell of a single-cell eukaryotic organism; a plant cell; an algal cell, e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like; a fungal cell (e.g., a yeast cell); an animal cell; an invertebrate animal cell (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.); an insect cell (e.g., a mosquito; a bee; an agricultural pest; etc.); an arachnid cell (e.g., a spider; a tick; etc ); a vertebrate animal cell (e.g., fish, amphibian, reptile, bird, mammal); a mammalian cell (e.g., rodent, mouse, rat, primate, non-human primate, human, and the like), a lagomorph cell (e.g., a rabbit); an ungulate cell (e.g., a cow, a horse, a camel, a llama, a vicuna, a sheep, a goat, etc.); a marine mammal cell (e.g., a whale, a seal, an elephant seal, a dolphin, a sea lion; etc.), and the like.
[0183] For example, microorganisms such as bacteria can be used as host organisms. In some cases, the host cell is capable of stably expressing a subject PylRS and / or a subject tRNA as described and provided herein. Such host cells may include, e.g., bacterial cells or eukaryotic cells (e.g., mammalian cells, algal cells, insect cells, fungal cells, yeast, and the like). In one embodiment of the present disclosure, the host cell is not archaea. Examples of bacterial host cells in context with the present disclosure include Gram negative and Gram positive bacterial cells. Specific examples for suitable host cells include, but are not limited to: E. coli, Corynebacterium glutamicum, Mycoplasma capricolum, CHO cells (Chinese Hamster cells), SF9 cells (insect cells), C. elegans cells, S. cerevisiae cells, Schizosaccharmyces pombe cells, Micrococcus luteus cells, Pichia pastoris cells (also known as Komagataella pastoris or Komagataella phaffii), plant cells, and Bombyx mori cells. Suitable prokaryotes include, but are not limited to, eubacteria, such as Gram-negative or Gram-positive organisms, for example, Enterobacteriaceae such as E. coli. Various E. coli strains are publicly available, such as E. coli K12 strain MM294 (ATCC 31 , 446); E. 0011 X1776 (ATCC 31 ,537); E. coli strain W3110 (ATCC 27,325); and K5772 (ATCC 53,635). Other suitable prokaryotic host cells include Enterobacteriaceae such as Escherichia, e.g., E. coli, Enterobacter, Erwinia, Klebsiella, Proteus, Salmonella, e.g., Salmonella typhimurium, Serratia, e.g., Serratia marcescans, and Shigella, as well as Bacilli such as B. subtilis and B. licheniformis, Pseudomonas such as P. aeruginosa, and Streptomyces. The examples listed herein are illustrative rather than limiting. For example, gram-positive or gram-negative bacteria can be used, e.g., bacteria of the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae,Streptococcaceae or Nocardiaceae, e.g., of the genera Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium or Rhodococcus. The genus and species Escherichia coli is one specific example.
[0184] As noted above, eukaryotic cells may be used as hosts. Eukaryotic cells of the present invention can be selected from, but are not limited to, mammalian cells, insect cells, yeast cells and plant cells. The eukaryotic cells of the invention may be present as individual cells or may be part of a tissue (e.g. a cell in a (cultured) tissue, organ or entire organism). As non-limiting examples the following plants or cells derived therefrom can be sued from the genera Nicotiana, in particular Nicotiana benthamiana and Nicotiana tabacum (tobacco); as well as Arabidopsis, in particular Arabidopsis thaliana.
[0185] Particular non-limiting examples of insect cells are Sf21, Sf9 and High Five cells. Particular non-limiting examples of mammalian cells are HEK293, HEK293T, HEK293F, CHO, CHO-S, COS, and HeLa cells. Alternatively, yeasts such as those of families like Saccharomyces or Pichia can also be used as host cells.
[0186] Depending on the host cell type, the cells can be grown or cultured in a manner known by a person skilled in the art.
[0187] Using the vectors according to the invention, prokaryotic or eukaryotic recombinant hosts can be produced, which are for example transformed with at least one vector according to the invention and can be used for producing NNM-containing POIs. Vectors / constructs according to the disclosure can be introduced into a suitable host cell / system and expressed using any convenient method - discussed elsewhere herein. Particularly common cloning and transfection methods, known by a person skilled in the art, include, for example: co-precipitation, protoplast fusion, electroporation, liposome-mediated transfection, retroviral transfection, lipid nanoparticle delivery, nanoparticle delivery, and the like, for expressing the stated nucleic acids in the respective expression system. Example systems are described for example in Current Protocols in Molecular Biology, F. Ausubel et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual. 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.Protein of Interest (POI)
[0188] As noted elsewhere herein, in some embodiments an NNM-containing protein of interest (POI) is produced by generating the protein via translation (e.g., in a bacterial cell, in a eukaryotic cell, in an in vitro translation system) in the presence of a subject PylRS and tRNA. The nucleic acid encoding the POI can be DNA (e.g., a DNA encoding a POI can be provided to a bacterial or eukaryotic cell, or to an in vitro translation system, e.g., one that is transcription coupled / linked), or can be an mRNA (e.g., an mRNA can be provided to a bacterial or eukaryotic cell, or to an in vitro translation system). The nucleic acid encoding the POI will include one or more inframe codons that are recognized by the subject tRNA’s anticodon (e.g., in some cases, one or more amber (UAG), ocher (UAA), or opal (UGA) codons).
[0189] The nucleotide sequence encoding a POI can be present on the same nucleic acid (e.g., an expression vector) as the nucleotide sequence encoding a subject PylRS and / or a subject tRNA - or the nucleotide sequence encoding a POI can be present on a separate nucleic acid from the PylRS or the tRNA (e.g., present on a separate DNA molecule, present on an mRNA, etc.). As such, in some cases, a nucleic acid (e.g., an expression vector) includes a nucleotide sequence encoding a subject PylRS and a nucleotide sequence encoding a subject tRNA. In some cases, a nucleic acid (e.g., an expression vector) includes a nucleotide sequence encoding a subject PylRS and a nucleotide sequence encoding a POI. In some cases, a nucleic acid (e.g., an expression vector) includes a nucleotide sequence encoding a subject PylRS, a nucleotide sequence encoding a subject tRNA, and a nucleotide sequence encoding a POI. In some cases, a nucleic acid (e.g., an expression vector) includes a nucleotide sequence encoding a subject tRNA and a nucleotide sequence encoding a POI.
[0190] The POI can be of any convenient length and can be any polypeptide. The POI can be any form of polypeptide or protein molecule, which may be recombinantly produced in any suitable translation system (e.g., an in vitro, i.e. , cell-free, expression system, a bacterial cell, a eukaryotic cell) in the presence of at least one NMM and a subject PylRS and tRNA. For example, the POI can be any bacterial or eukaryotic protein (e.g., any mouse protein, and human protein), or can be a non-naturally existing protein (i.e., can include an amino acid sequence that does not exist in nature). In some cases, the POI is an antibody or antibody component.
[0191] In some embodiments, the NNM-containing POI that is produced can be referred to as functionalized because the incorporated NNM provides a function to the POI that itwould not have had in the absence of the NNM. In some cases, the POI is functionalized such that it can be conjugated / linked (covalently or non-covalently, reversibly or irreversibly) to another molecule, sometimes referred to herein as a conjugation partner molecule. As an illustrative example, in some cases, the POI is an antibody or antibody component, and the conjugation partner molecule is drug molecule, e.g., a small molecule drug, a toll-like receptor (TLR) agonist or antagonist, and the like. For example, in some cases, an NNM is click chemistry compatible, e.g., is tetrazine-functionalized or can react with a tetrazine-functionalized conjugation partner molecule. Non-limiting examples of NNMs that can be used to functionalize a POI are discussed above.
[0192] As noted above, a POI can be any polypeptide, e.g., any bacterial or eukaryotic protein (e.g., any mouse protein, and human protein), or a non-naturally existing protein (i.e. , can include an amino acid sequence that does not exist in nature). For example, a POI can be an antibody / an immunoglobulin. Immunoglobulins from any of the classes or subclasses may be selected, e.g. IgG, IgA, IgM , IgD and IgE. In some cases, the immunoglobulin is of the class IgG including but not limited to IgG subclasses (IgG 1 , 2, 3 and 4) or the class IgM which is able to specifically bind to a specific epitope on an antigen. Antibodies can be intact immunoglobulins derived from natural sources or from recombinant sources and can be immunoreactive portions of intact immunoglobulins. Antibodies may exist in a variety of forms including, for example, polyclonal antibodies, monoclonal antibodies, camelized single domain antibodies, recombinant antibodies, anti-idiotype antibodies, multispecific antibodies, antibody fragments, such as, Fv, VHH, Fab, F(ab )2, Fab', Fab'-SH, F(ab')2, single chain variable fragment antibodies (scFv), tandem / bis-scFv, Fc, pFc', scFv-Fc, disulfide Fv (dsFv), bispecific antibodies (bc-scFv) such as BiTE antibodies, trispecific antibody derivatives such as tribodies, camelid antibodies, minibodies, nanobodies, resurfaced antibodies, humanized antibodies, fully human antibodies, single domain antibodies (sdAb, also known as Nanobody™), chimeric antibodies, chimeric antibodies comprising at least one human constant region, dual-affinity antibodies such as dual-affinity retargeting proteins (DART™), and multimers and derivatives thereof, such as divalent or multivalent single-chain variable fragments (e.g. di-scFvs, tri-scFvs) including but not limited to minibodies, diabodies, triabodies, tribodies, tetrabodies, and the like, and multivalent antibodies. Reference is made to [Trends in Biotechnology 2015, 33, 2, 65], [Trends Biotechnol.2012, 30, 575-582], and [Cane. Gen. Prot.201310, 1-18], and [BioDrugs 2014, 28, 331-343], the contents of which arehereby incorporated by reference. "Antibody fragment" refers to at least a portion of the variable region of the immunoglobulin that binds to its target, i.e. the antigenbinding region. Other embodiments use antibody mimetics as POIs, such as but not limited to Affimers, Anticalins, Avimers, Alphabodies, Affibodies, DARPins, and multimers and derivatives thereof; see, e.g., Geering et al., Trends Biotechnol. 2015 Feb;33(2):65-79. In the context of this disclosure the term "antibody" is meant to encompass all of the antibody variations, fragments, derivatives, fusions, analogs and mimetics outlined in this paragraph, unless specified otherwise.
[0193] Typical non-limiting examples of antibody molecules that can produced to form an NNM-containing POI of the present disclosure are selected form biologically, in particular pharmacologically active antibody molecules. Non-limiting examples include, but are not limited to: trastuzumab, bevacizumab, cetuximab, panitumumab, ipilimumab, rituximab, alemtuzumab, ofatumumab, gemtuzumab, brentuximab, ibritumomab, tositumomab, pertuzumab, adecatumumab, IGN101 , INA01 labetuzumab, hua33, pemtumomab, , minretumomab (CC49), cG250, J591 , MOv-18, farletuzumab (MGRAb-003), 3F8, ch14,18, KW-2871 , hu3S193, lgN311 , IM-2C6, CDP-791 , etaracizumab, volociximab, nimotuzumab, MM- 121 , AMG 102, METMAB, SCH 900105, AVE1642, IMC-A12, MK-0646, R1507, CP 751871, KB004, III A4, mapatumumab, HGS-ETR2, CS-1008, denosumab, sibrotuzumab, F19, 81 C6, pinatuzumab, lifastuzumab, glembatumumab, coltuximab, lorvotuzumab, indatuximab, anti-PSMA, MLN-0264, ABT-414, milatuzumab, ramucirumab, abagovomab, abituzumab, adecatumumab, afutuzumab, altumomab pentetate, amatuximab, anatumomab, anetumab, apolizumab, arcitumomab, ascrinvacumab, zumab, bavituximab , bectumomab, belimumab, bivatuzumab, brontictuzumab, cantuzumab, capromab, catumaxomab, citatuzumab, cixutumumab, clivatuzumab, codrituzumab, conatumumab, dacetuzumab, dallotuzumab, daratumumab, demcizumab, denintuzumab, depatuxizumab, derlotuximab, detumomab, dinutuximab, itumab, duligotumab, durvalumab, dusigitumab, ecromeximab, edrecolomab, elgemtumab, emactuzumab, enavatuzumab emibetuzumab, enfortumab, enoblituzumab, ensituximab, epratuzumab, ertumaxomab, etaracizumab, farletuzumab, ficlatuzumab, figitumumab, flanvotumab, futuximab, galiximab, ganitumab, icrucumab, igovomab, imalumab, imgatuzumab, indusatumab, inebilizumab, intetumumab, iratumumab, isatuximab, lexatuzumab, lilotomab, lintuzumab, lirilumab, lucatumumab, lumretuzumab.margetuximab, matuzumab, mirvetuximab, mitumomab, mogamulizumab, moxetumomab, naptumomab, narnatumab, necitumumab,nesvacumab, nimotuzumab, nivolumab, nofetumomab, obinutuzumab, ocaratuzumab, ofatumumab, olaratumab, onartuzumab, ontuxizumab, oportuzumab, oregovomab, otlertuzumab, pankomab, parsatuzumab, pasotuxizumab, patritumab, pembrolizumab, pemtumomab, pidilizumab, pintumomab, polatuzumab, prisumumab, quilizumab, racotumomab, ramucirumab, rilotumumab, robatumumab, sacituzumab, samalizumab, satumomab, seribantumab, siltuximab, sofituzumab, tacatuzumab, taplitumomab, tarextumab, tenatumomab, teprotumumab, tetulomab, ticilimumab, tositumomab, tovetumab, tremelimumab, tucotuzumab, ublituximab, ulocuplumab, urelumab, utomilumab, vadastuximab, vandortuzumab, vantictumab, vanucizumab, varlilumab, veltuzumab, vesencumab, volociximab, vorsetuzumab votumumab, zalutumumab, zatuxima, and combinations and derivatives thereof.
[0194] After translation, the NNM-containing POI prepared according to the present disclosure may optionally be recovered and purified, either partially or substantially to homogeneity, according to procedures generally known in the art. Unless the POI is secreted into the culture medium, recovery usually requires cell disruption. Methods of cell disruption are well known in the art and include physical disruption, e.g., by (ultrasound) sonication, liquid-sheer disruption (e.g., via French press), mechanical methods (such as those utilizing blenders or grinders) or freeze-thaw cycling, as well as chemical lysis using agents which disrupt lipid-lipid, protein-protein and / or protein- lipid interactions (such as detergents), and combinations of physical disruption techniques and chemical lysis. Standard procedures for purifying polypeptides from cell lysates or culture media are also well known in the art and include, e.g., ammonium sulfate or ethanol precipitation, acid or base extraction, column chromatography, affinity column chromatography, anion or cation exchange chromatography, phosphocellulose chromatography, hydrophobic interaction chromatography, hydroxylapatite chromatography, lectin chromatography, gel electrophoresis and the like. Protein refolding steps can be used, as desired, in making correctly folded mature proteins. High performance liquid chromatography (HPLC), affinity chromatography or other suitable methods can be employed in final purification steps where high purity is desired. Antibodies made against the polypeptides of the disclosure can be used as purification reagents, i.e. , for affinitybased purification of the polypeptides. A variety of purification / protein folding methods are well known in the art and any convenient method can be used, including, e.g., those set forth in Scopes, Protein Purification, Springer, Berlin (1993); and Deutscher,Methods in Enzymology Vol. 182: Guide to Protein Purification, Academic Press (1990); and the references cited therein.
[0195] As noted, those of skill in the art will recognize that, after synthesis, expression and / or purification, polypeptides can possess a conformation different from the desired conformations of the relevant polypeptides. For example, polypeptides produced by prokaryotic systems often are optimized by exposure to chaotropic agents to achieve proper folding. During purification from, e.g., lysates derived from E. coli, the expressed polypeptide is optionally denatured and then renatured. This is accomplished, e.g., by solubilizing the proteins in a chaotropic agent such as guanidine HCI. In general, it is occasionally desirable to denature and reduce expressed polypeptides and then to cause the polypeptides to re-fold into the preferred conformation. For example, guanidine, urea, DTT, DTE, and / or a chaperonin can be added to a translation product of interest. Methods of reducing, denaturing and renaturing proteins are well known to those of skill in the art. Polypeptides can be refolded in a redox buffer containing, e.g., oxidized glutathione andL-arginine.Conjugation partners
[0196] Conjugation partners can be any convenient molecule - in some cases a conjugation partner is selected from bioactive compounds, labeling agents, and chelators. Examples of conjugation partners include, but are not limited to: small organic molecule drugs, steroids, lipids, proteins, aptamers, oligopeptides, oligonucleotides, oligosaccharides, as well as peptides, peptoids, amino acids, nucleotides, oligo- or polynucleotides, nucleosides, DNA, RNA, toxins, glycans and immunoglobulins.
[0197] Examples of conjugation partners include, but are not limited to: hormones, cytotoxins, antiproliferative / antitumor agents, antiviral agents, antibiotics, cytokines, anti-inflammatory agents, antihypertensive agents, chemosensitizing, photosensitizing and radiosensitizing agents, anti-AIDS substances, anti-viral agents, immunosuppressants, immunostimulants, enzyme inhibitors, anti-Parkinson agents, neurotoxins, channel blockers, modulators of cell-extracellular matrix interactions including cell growth inhibitors and anti-adhesion molecules, inhibitors of DNA, RNA or protein synthesis, steroidal and non-steriodal anti-inflammatory agents, anti - angiogenic factors, anti-Alzheimer agents. In some embodiments, a conjugation partner is a low to medium molecular weight compound (e.g. about 200 to 5000 Da, about 200 to about 1500 Da, preferably about 300 to about 1000 Da).
[0198] Examples of conjugation partners include cytotoxic drugs, e.g., for cancer therapy. Such drugs include, in general, DNA damaging agents, anti-metabolites, natural products and their analogs, enzyme inhibitors such as dihydro folate reductase inhibitors and thymidylate synthase inhibitors, DNA binders, DNA alkylators, radiation sensitizers, DNA intercalators, DNA cleavers, microtubule stabilizing and destabilizing agents, topoisomerases inhibitors. Examples include but are not limited to platinumbased drugs, the anthracycline family of drugs, the vinca drugs, the mitomycins, the bleomycins, the cytotoxic nucleosides, taxanes, lexitropsins, the pteridine family of drugs, diynenes, the podophyllotoxins, dolastatins, maytansinoids, differentiation inducers, and taxols. Particularly useful members of those classes include, for example, auristatins, maytansines, maytansinoids, calicheamicins, dactinomycins, duocarmycins, CC1065 and its analogs, camptothecin and its analogs, SN-38 and its analogs; DXd, tubulysin M, cryptophycins, pyrrolobenzodiazepines and pyrrolobenzodiazepine dimers (PBDs), pyridinobenzodiazepines (PDDs) and indolinobenzodiazepines (IBDs) (see US20210206763A1), methotrexate, methopterin, dichloromethotrexate, 5-fluorouracil, DNA minor groove binders, - mercaptopurine, cytosine arabinoside, melphalan, EUROSIN, EUROSIDEIN, actinomycin, anthracyclines (doxorubicin, epirubicin, idarubicin, daunorubicin, PNU- 159682 (see US 10,288,745 B2.) and its analogs, mitomycin C, mitomycin A, caminomycin, aminopterin, tallysomycin, podophyllotoxin and ;podophyllotoxin derivatives such as etoposide or etoposide phosphate, vinblastine, vincristine, vindesine, taxol, taxotere retinoic acid, butyric acid, N8-acetyl spermidine, staurosporin, colchicine, camptothecin, esperamicin, ene-diynes, and their analogues, hemiasterlin and its analogues.
[0199] Other example drug classes are angiogenesis inhibitors, cell cycle progression inhibitors, P13K / m-TOR / AKT pathway inhibitors, MAPK signaling pathway inhibitors, kinase inhibitors, protein chaperones inhibitors, HDAC inhibitors, PARP inhibitors, Wnt / Hedgehog signaling pathway inhibitors, RNA polymerase inhibitors, and protein degraders (see pubs.acs.org / doi / 10.1021 / acschembio.0cQ0285).
[0200] Examples of auristatins include dolastatin 10, monomethyl auristatin E (MMAE), auristatin F, monomethyl auristatin F (MMAF), auristatin F hydroxypropylamide (AF HPA), auristatin F phenylene diamine (AFP), monomethyl auristatin D (MMAD), auristatin PE, auristatin EB, auristatin EFP, auristatin TP and auristatin AQ. Suitable auristatins are also described in US Publication Nos.2003 / 0083263, 2011 / 0020343, and 2011 / 0070248; PCT Application Publication Nos. WO09 / 117531 ,W02005 / 081711 , W004 / 010957; W002 / 088172 and WO01 / 24763, and US Patent Nos. 7,498,298; 6,884,869; 6,323,315; 6,239,104; 6,124,431 ; 6,034,065; 5,780,588; 5,767,237; 5,665,860; 5,663,149; 5,635,483; 5,599,902; 5,554,725; 5,530,097; 5,521 ,284; 5,504,191; 5,410,024; 5,138,036; 5,076,973; 4,986,988; 4,978,744; 4,879,278; 4,879,278; 4,816,444; and 4,486,414, the disclosures of which are incorporated herein by reference in their entirety.
[0201] Examples of drugs include the dolastatins and analogues thereof including: dolastatin A (US Pat No. 4,486,414), dolastatin B (US Pat No. 4,486,414), dolastatin 10 (US Pat No. 4,486,444, 5,410,024, 5,504,191 , 5,521 ,284, 5,530,097, 5,599,902, 5,635,483, 5,663,149, 5,665,860, 5,780,588, 6,034,065, 6,323,315), dolastatin 13 (US Pat No.4, 986, 988), dolastatin 14 (US Pat No. 5,138,036), dolastatin 15 (US Pat No. 4,879,278), dolastatin 16 (US Pat No. 6,239,104), dolastatin 17 (US Pat No. 6,239,104), and dolastatin 18 (US Pat No. 6,239,104), each patent incorporated herein by reference in their entirety. Examples of maytansines, maytansinoids, such as DM-1 and DM-4, or maytansinoid analogs, including maytansinol and maytansinol analogs, are described in US Patent Nos. 4,424,219; 4,256,746; 4,294,757; 4,307,016; 4,313,946; 4,315,929; 4,331,598; 4,361 ,650; 4,362,663; 4,364,866; 4,450,254; 4,322,348; 4,371 ,533; 5,208,020; 5,416,064; 5,475,092; 5,585,499; 5,846,545; 6,333,410; 6,441 ,163; 6,716,821 and 7,276,497.
[0202] Other examples include mertansine and ansamitocin. Pyrrolobenzodiazepines (PBDs), which expressly include dimers and analogs, include but are not limited to those described in [Denny, Exp. Opinion. Then Patents, 10(4):459-474 (2000)], [Hartley et al., Expert Opin Investig Drugs. 2011 , 20(6):733-44], Antonow et al., Chem Rev. 2011 , 111(4), 2815-64],
[0203] Calicheamicins include, e.g. enediynes, esperamicin, and those described in US Patent Nos.5,714,586 and 5,739,116.
[0204] Examples of duocarmycins and analogs include CC1065, duocarmycin SA, duocarmycin A, duocarmycin B I, duocarmycin B2, duocarmycin Cl, duocarmycin C2, duocarmycin D, DU-86, KW-2189, adozelesin, bizelesin, carzelesin, seco-adozelesin. Other examples include those described in, for example, US Patent No. 5,070,092; 5,101 ,092; 5,187,186; 5,475,092; 5,595,499; 5,846,545; 6,534,660; 6,548,530; 6,586,618; 6,660,742; 6,756,397; 7,049,316; 7,553,816; 8,815,226; US20150104407; 61 / 988,011 filed May 2, 2014 and 62 / 010,972 filed June 11, 2014; the disclosure of each of which is incorporated herein in its entirety.
[0205] Examples of vinca alkaloids include vincristine, vinblastine, vindesine, and navelbine, and those disclosed in US Publication Nos. 2002 / 0103136 and 2010 / 0305149, and in US Patent No. 7,303,749, the disclosures of which are incorporated herein by reference in their entirety.
[0206] Examples of epothilone compounds include epothilone A, B, C, D, E, and F, and derivatives thereof. Suitable epothilone compounds and derivatives thereof are described, for example, in US Patent Nos.6, 956, 036; 6,989,450; 6,121 ,029; 6,117,659; 6,096,757; 6,043,372; ; 5,969,145; and 5,886,026; and WO97 / 19086; WO98 / 08849; W098 / 22461 ; W098 / 25929; W098 / 38192; WO99 / 01124; WO99 / 02514; WO99 / 03848; WO99 / 07692; WO99 / 27890; and W099 / 28324; the disclosures of which are incorporated herein by reference in their entirety.
[0207] Examples of cryptophycin compounds are described in US Patent Nos. 6,680,311 and 6,747,021 ; the disclosures of which are incorporated herein by reference in their entirety. Examples of platinum compounds include cisplatin, carboplatin, oxaliplatin, iproplatin, ormaplatin, tetraplatin. Examples of DNA binding or alkylating drugs include CC-1065 and its analogs, anthracyclines, calicheamicins, dactinomycins, mitromycins, pyrrolobenzodiazepines, and the like. Examples of microtubule stabilizing and destabilizing agents include taxane compounds, such as paclitaxel, docetaxel, tesetaxel, and carbazitaxel; maytansinoids, auristatins and analogs thereof, vinca alkaloid derivatives, epothilones and cryptophycins.
[0208] Examples of topoisomerase inhibitors include camptothecin and camptothecin derivatives, camptothecin analogs and non-natural camptothecins, such as, for example, CPT-11, SN-38, topotecan, 9-aminocamptothecin, rubitecan, gimatecan, karenitecin, silatecan, lurtotecan, exatecan, DXd, diflometotecan, belotecan, lurtotecan and S39625. Other camptothecin compounds that can be used in the present disclosure include those described in, for example, J. Med. Chem., 29:2358- 2363 (1986); J. Med. Chem., 23:554 (1980); J. Med Chem., 30: 1774 (1987).
[0209] Examples of angiogenesis inhibitors include, but are not limited to, MetAP2 inhibitors, VEGF inhibitors, PIGF inhibitors, VGFR inhibitors, PDGFR inhibitors, MetAP2 inhibitors. Exemplary VGFR and PDGFR inhibitors include sorafenib, sunitinib and vatalanib. Exemplary MetAP2 inhibitors include fumagillol analogs, meaning compounds that include the fumagillin core structure. Examples of cell cycle progression inhibitors include CDK inhibitors such as, for example, BMS-387032 and PD0332991 ; Rho-kinase inhibitors such as, for example, AZD7762; aurora kinase inhibitors such as, for example, AZD1152, MLN8054 and MLN8237; PLK inhibitorssuch as, for example, Bl 2536, BI6727, GSK461364, ON-01910; and KSP inhibitors such as, for example, SB 743921 , SB 715992, MK-0731 , AZD8477, AZ3146 and ARRY-520.
[0210] Examples of P13K / m-T0R / AKT signaling pathway inhibitors include phosphoinositide 3- kinase (P13K) inhibitors, GSK-3 inhibitors, ATM inhibitors, DNA-PK inhibitors and PDK-1 inhibitors. Examples of P13 kinases are disclosed in US Patent No.6, 608, 053, and include BEZ235, BGT226, BKM120, CAL263, demethoxyviridin, GDC-0941 , GSK615, IC87114, LY294002, Palomid 529, perifosine, PF-04691502, SAR245408, SAR245409, SF1126, Wortmannin, XL147 and XL765. Examples of AKT inhibitors include, but are not limited to AT7867. Examples of MAPK signaling inhibitor pathways include MEK, Ras, JNK, B-Raf and p38 MAPK inhibitors. Examples of MEK inhibitors are disclosed in US Patent No.7, 517, 944 and include GDC-;0973, GSKI 120212, MSC1936369B, AS703026, R05126766 and R04987655, PD0325901 , AZD6244, AZD8330 and GDC-0973. Examples of B-raf inhibitors include CDC-0879, PLX-4032, and SB590885. Examples of B p38 MAPK inhibitors include BIRB 796, LY2228820 and SB 202190. Exemplary receptor tyrosine kinases inhibitors include but are not limited to AEE788 (NVP- AEE 788), BIBW2992 (Afatinib), Lapatinib, Erlotinib (Tarceva), Gefitinib (Iressa), AP24534 (Ponatinib), ABT-869 (linifanib), AZD2171 , CHR-258 (Dovitinib), Sunitinib (Sutent), Sorafenib (Nexavar), and Vatalinib.
[0211] Examples of protein chaperone inhibitors include HSP90 inhibitors. Exemplary inhibitors include 17AAG derivatives, BIIB021 , BIIB028, SNX-5422, NVP-AUY-922 and KW-2478. Examples of HDAC inhibitors include Belinostat (PXD101), CUDC-101 , Droxinostat, ITF2357 (Givinostat, Gavinostat), JNJ-26481585, LAQ824 (NVP- LAQ824, Dacinostat), LBH-589 (Panobinostat), MC1568, MGCD0103 (Moceti nostat), MS -275 (Entinostat), PCI-24781 , Pyroxamide (NSC 696085), SB939, Trichostatin A and Vorinostat (SAHA). Exemplary PARP inhibitors include iniparib (BSI 201), olaparib (AZD-2281), ABT-888 (Veliparib), AG014699, CEP9722, MK 4827, KU- 0059436 (AZD2281), LT-673, 3-aminobenzamide, A-966492, and AZD2461. Examples of Wnt / Hedgehog signaling pathway inhibitors include vismodegib, cyclopamine and XAV-939. Examples of RNA polymerase inhibitors include amatoxins. Exemplary amatoxins include alpha-amanitins, beta amanitins, gamma amanitins, eta amanitins, amanullin, amanullic acid, amanisamide, amanon, and proamanullin. Examples of cytokines include IL-2, IL-7, IL-10, IL-12, IL-15, IL-21 , TNF.
[0212] As non-limiting examples of particular drugs there may be mentioned Auristatins, Maytansinoids, PBDs, topoisomerase inhibitors, anthracyclines. In anotherembodiment, a combination of two or more different drugs as described above are used.
[0213] According to another embodiment, the bioactive compound may be selected from any synthetic or naturally occurring compounds including one or more natural and / or nonnatural, proteinogenic and / or non-proteinogenic amino acid residues, such as in particular oligo- or polypeptides or proteins.
[0214] Examples of conjugation partners also include labeling agents. Labeling agents which may be used according to the disclosure can include any type of label known in the art which does not inhibitor negatively affect reactivity of the conjugation moiety, the. a tetrazine moiety.
[0215] Labels include, but are not limited to, dyes (e.g. fluorescent, luminescent, or phosphorescent dyes, such as dansyl, coumarin, fluorescein, acridine, rhodamine, silicon-rhodamine, BODIPY, or cyanine dyes), chromophores (e.g., phytochrome, phycobilin, bilirubin, etc.), radiolabels (e.g. radioactive forms of hydrogen, fluorine, carbon, phosphorous, sulfur, or iodine, such as tritium, fluorine-18, carbon-11, carbon- 14, phosphorous-32, phosphorous -33, sulphur-33, sulphur-35, iodine-123, or iodine- 125), MRI-sensitive spin labels, affinity tags (e.g. biotin, His-tag, Flag-tag, strep-tag, sugars, lipids, sterols , PEG-linkers, benzylguanines, benzylcytosines, or co-factors), polyethylene glycol groups (e.g., a branched PEG, a linear PEG, PEGs of different molecular weights, etc.), photocrosslinkers (such as p-azidoiodoacetanilide), NMR probes, X-ray probes , pH probes, IR probes, resins, solid supports.
[0216] In some embodiments, example dyes can include an NIR contrast agent that fluoresces in the near infrared region of the spectrum. Exemplary near-infrared fluorophores can include dyes and other fluorophores with emission wavelengths (e.g., peak emission wavelengths) between about 630 and 1000 nm, e.g., between about 630 and 800 nm, between about 800 and 900 nm, between about 900 and 1000 nm , between about 680 and 750 nm, between about 750 and 800 nm, between about 800 and 850 nm, between about 850 and 900 nm, between about 900 and 950 nm, or between about 950 and 1000 nm. Fluorophores with emission wavelengths (e.g., peak emission wavelengths) greater than 1000 nm can also be used in the methods described herein.
[0217] In some embodiments, example fluorophores include 7-amino-4-methylcoumarin-3- acetic acid (AMCA), TEXAS RED™ (Molecular Probes, Inc., Eugene, Oreg.), 5-(and - 6)-carboxy-X -rhodamine, lissamine rhodamine B, 5-(and -6)-carboxyfluorescein, fluorescein-5-isothiocyanate (FITC), 7-diethylaminocoumarin-3-carboxylic acid,tetramethylrhodamine-5-(and -6)-isothiocyanate, 5 -( and -6)- carboxytetramethylrhodamine, 7-hydroxycoumarin-3-carboxylic acid, 6-[fluorescein 5- (and -6)-carboxamido]hexanoic acid, N-(4,4-difluoro-5,7-dimethyl-4- bora-3a,4a diaza- 3-indacenepropionic acid, eosin-5-isothiocyanate, erythrosin-5-isothiocyanate, and CASCADE™ blue acetylazide (Molecular Probes, Inc., Eugene, Oreg.) and ATTO dyes. Further labeling agents are 177-Lutetium, 89-Zirkonium, 131-lod, 68-Gallium, 99m-Technecium, 225-Actinium, 213-Bismut, 90-Ytrium and 212-Plumbum.
[0218] Examples of conjugation partners also include chelators. Lists of typically applicable chelators and their short names are given below; Corresponding salts thereof are also applicable.
[0219] Acetyl acetone (ACAC), ethylene diamine (EN), 2-(2-aminoethylamino)ethanol (AEEA), diethylene triamine (DIEN), iminodiacetate (IDA), triethylene tetramine (TRIEN), triaminotriethylamine, nitrilotriacetate (NTA) and its saltslike Na3NTA or FeNTA, ethylenediaminotriacetate (TED), ethylenediamine tetraacetate (EDTA) and its salts like Na2EDTA and CaNa2EDTA, diethylene triaminpentaacetate (DTPA), 1 ,4,7,10-ztetraazacyclododecane-1 ,4,7,10-tetraacetate ( DOTA), Oxalate (OX), tartrate (TART), citrate (CIT), dimethylglyoxime (DMG), 8-hydroxyquinoline, 2,2'- bipyridine (BPY), 1 ,10-phenanthroline (PHEN), dimercapto succinic acid (DMSA), 1 ,2- bis(diphenylphosphino)ethane (DPPE), sodium salicylate, methoxy salicylates, British anti-Lewisite or 2,3-dimercaprol (BAL), meso-2,3- dimercaptosuccinic acid (DMSA); Siderophores secreted by microorganisms, as for example desferrioxamine or deferoxamine B, also known as Deferral (Novartis), produced by Streptomyces spp.; deferoxamine (DFO), a trihydroxamic acid secreted by Streptomyces pilosus; phytochemicals like curcuminoids and derivatives of mugineic acid, like 3-hydroxy- mugineic acid and 2 -deoxy-mugineic acid; synthetically produced chelators, like Ibuprofen; derivatives of catechol, hydroxamate and hydroxypyridinone, like hydroxamate desferal and hydroxypyridinone deferiprone; deferiprone (L1 or 1 ,2- dimethyl-3-hydroxypyrid-4-one); D-penicillamine (DPA or D-PEN) which is |3-p- dimethylcysteine or 3-mercapto-D-valine; tetraethylenetetraamine (TETA) or trientine and its two major metabolites N1 -acetyltriethylenetetramine (MAT) and N1.N10 - diacetyltriethylenetetramine (DAT); hydroxyquinolines; clioquinol, which is a halogenated derivative of 8-hydroxyquinoline; and 5,7-dichloro-2- [(dimethylamino)methyl]quinolin-8-ol (PBT2).Kits
[0220] Provided are kits / systems for carrying out a subject method. Kits can include Various combinations of components useful in any of the methods described elsewhere herein. A kit can further include one or more additional reagents, where such additional reagents can be any convenient reagent. Components of a subject kit can be in separate containers; or can be combined in a single container. In some cases one or more of a kit’s components are pharmaceutically formulated for administration to a human.
[0221] In addition to above-mentioned components, a subject kit can further include instructions for using the components of the kit to practice the subject methods (e.g., dosing instructions, instructions to administer the component(s) to an individual. The instructions for practicing the subject methods are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e. , associated with the packaging or subpackaging) etc. In some embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g. CD-ROM, diskette, flash drive, etc. In some embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g. via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate.EXEMPLARY NON-LIMITING ASPECTS OF THE DISCLOSURE
[0222] Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure are provided below. As will be apparent to those of ordinary skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below. It will be apparentto one of ordinary skill in the art that various changes and modifications can be made without departing from the spirit or scope of the invention.1. A method of producing a protein of interest (POI) that comprises one or more non-natural monomer (NNMs), the method comprising: providing to a bacterial cell, a eukaryotic cell, or an in vitro translation system:(i) a pyrrolysyl-tRNA synthetase (PylRS), wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10,(ii) a NNM or a salt thereof,(iii) a tRNA that can be acylated with the NNM by the PylRS, and(iv) a DNA or mRNA comprising a nucleotide sequence encoding a POI, wherein the nucleotide sequence encoding the POI comprises one or more in-frame codons that are recognized by the tRNA’s anticodon; thereby producing a NNM-containing POI.2. The method of 1 , wherein the PylRS is provided to the bacterial cell, the eukaryotic, or the in vitro translation system as a nucleic acid encoding the PylRS.3. The method of 1 or 2, wherein the tRNA is provided to the bacterial cell, the eukaryotic, or the in vitro translation system as a nucleic encoding the tRNA.4. The method of any one of 1-3, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1.5. The method of any one of 1-4, wherein the tRNA’s anticodon recognizes the codon UAG.6. The method of any one of 1-5, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.7. The method of any one of 1-5, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.8. The method of any one of 1-7, wherein the NNM is any one of the compounds depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof.9. The method of any one of 1-7, wherein the NNM is any one of the lysine derivatives depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or [33-amino acid analog thereof.10. The method of any one of 1-7, wherein the NNM is any one of the compounds depicted in Figure 15.11. The method of any one of 1-7, wherein the NNM is Pyl, Boc-Lys, or HO-Boc- Lys.12. The method of any one of 1-11 , wherein the cell is a mammalian cell.13. The method of any one of 1-11 , wherein the cell is an E. coli cell.14. The method of any one of 1-13, further comprising reacting the NNM- containing POI with a conjugation partner molecule such that the conjugation partner molecule binds covalently to at least one NNM of the NNM- containing POI.15. A system comprising: a nucleic acid comprising a nucleotide sequence encoding a pyrrolysyl- tRNA synthetase (PylRS), wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, wherein the nucleotide sequence encoding the PylRS is:(a) operably linked to a bacterial or eukaryotic promoter, and / or(b) codon-optimized for expression in a bacterial or eukaryotic host cell.16. The system of 15, wherein the nucleic acid is a plasmid or a viral vector.17. The system of 15 or 16, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1.18. The system of any one of 15-17, further comprising a tRNA or a nuclei acid comprising a nucleotide sequence encoding said tRNA, wherein the tRNA can be acylated by the PylRS.19. The system of 18, wherein the tRNA’s anticodon recognizes the codon UAG.20. The system of 18 or 19, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.21. The system of 18 or 19, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.22. The system of any one of 18-21 , wherein the nucleotide sequence encoding the tRNA is operably linked to a bacterial or eukaryotic promoter.The system of any one of 18-22, wherein the nucleotide sequence encoding the tRNA is present on the same nucleic acid as the nucleotide sequence encoding the PylRS. The system of any one of 18-22, wherein the nucleotide sequence encoding the tRNA and the nucleotide sequence encoding the PylRS are present on different nucleic acids. The system of any one of 18-24, further comprising one or more non-natural monomers (NNMs), wherein said tRNA can be acylated with said one or more NNMs by said PylRS. The system of 25, wherein the NNM is any one of the compounds depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof. The system of 25, wherein the NNM is any one of the lysine derivatives depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or3-amino acid analog thereof. The system of 25, wherein the NNM is any one of the compounds depicted in Figure 15. The system of 25, wherein the NNM is Pyl, Boc-Lys, or HO-Boc-Lys. A bacterial or eukaryotic cell comprising:(a) a pyrrolysyl-tRNA synthetase (PylRS), or a nuclei acid comprising a nucleotide sequence encoding the PylRS, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, and(b) a tRNA that can be acylated by the PylRS, or a nuclei acid comprising a nucleotide sequence encoding the tRNA. The bacterial or eukaryotic cell of 30, wherein the cell is a mammalian cell. The bacterial or eukaryotic cell of 30, wherein the cell is an E. coli cell. The bacterial or eukaryotic cell of any one of 30-32, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1. The bacterial or eukaryotic cell of any one of 30-33, wherein the tRNA’s anticodon recognizes the codon UAG. The bacterial or eukaryotic cell of any one of 30-34, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.36. The bacterial or eukaryotic cell of any one of 30-34, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.37. The bacterial or eukaryotic cell of any one of 30-36, wherein the nucleotide sequence encoding the tRNA is present on the same nucleic acid as the nucleotide sequence encoding the PylRS.EXPERIMENTAL EXAMPLES
[0223] The following examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following examples, but rather should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.
[0224] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize the present invention and practice the claimed methods. The following working examples therefore are not to be construed as limiting in any way the remainder of the disclosure.
[0225] General methods in molecular and cellular biochemistry can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., HaRBor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference. Reagents, cloning vectors, cells, and kits for methods referred to in, or related to, this disclosure are available from commercial vendors such as BioRad, Agilent Technologies, Thermo Fisher Scientific, Sigma-Aldrich, New England Biolabs (NEB), Takara Bio USA, Inc., and the like, as well as repositories such as e.g., Addgene, Inc., American Type Culture Collection (ATCC), and the like.Example 1: A new genetic code with all TAG codons as pyrrolysine in diverse archaea
[0226] The use of TAG codons was analyzed in Archaea that encode Pyl to determine if these organisms use a previously unidentified genetic code. Surprisingly, organisms within two major groups of Archaea encoded numerous genes with internal in-frame TAG codons, which implied the consistent incorporation of Pyl at TAG codons in these organisms. Targeted proteomic analysis was applied to human gut and marine archaeal isolates suspected of genome-wide re-encoding of TAG to Pyl. The analysis showed that TAG in expressed proteins encodes Pyl. Moreover, five PylRS-tRNAPylpairs were identified from organisms previously unexplored for genetic code expansion. The work here demonstrated that each one supports the cellular biosynthesis of proteins containing non-canonical a-amino and a-hydroxy acids. On this basis, the existence of a previously unrecognized genetic code, currently confined to archaea, is proposed here and is referred to herein as Genetic Code 34. The phylogenetic distribution of Code 34 is not monophyletic, hinting that code transition may be relatively facile. Based on analysis of the frequency and distribution of TAG codons in archaeal genomes, features indicative of recent vs. ancient re-coding were inferred. A series of steps that may lead to the adoption of this new genetic code is herein proposed.Unusual phylogenetic distribution of Pyl
[0227] To investigate the phylogenetic distribution of Pyl-encoding archaea (Pyl archaea), genomes for representatives of all families known to encode Pyl machinery were collected and evidence was sought for new taxa. It was determined that the majority of Pyl archaea belong to the Halobacteriota and Thermoplasmatota phyla, with few Pyl-encoding genomes within Asgard, Thermoproteota, and Hydrothermarchaeota archaea (Fig. 1). Notably, organisms that encode Pyl machinery have a discontinuous phylogenetic distribution, sometimes even at the genus level (Fig. 1). For example, the Pyl pathway is encoded and actively utilized by a single member of the Archaeoglobus. The proteins from this genome share 98.6% average amino acid identity (AAI) with Archaeoglobus LCB024-003 yet LCB024-003 does not encode Pyl. Pyl is limited to only a single family within either the Asgard and Hydrothermarcheaota phyla, and in the Ca. Methanomethylicus genus, only two of five archaea encode Pyl.
[0228] Given the intermixing of closely related archaea that do and do not use Pyl, the phylogenies of the Pyl machinery were compared to that of the organisms (Fig. 2).Consistent with a prior suggestion, discordance between the tree structures provides evidence for lateral transfer of the Pyl genes.TAG frequency varies among Pyl archaea
[0229] Given that organisms almost always interpret TAG as a stop codon, it was reasoned that there may be detrimental consequences of the acquisition of Pyl machinery, such as the erroneous incorporation of Pyl and readthrough of the normal gene end. It is also possible that TAG in these organisms have a dual meaning, however no mechanism for deciphering TAG-stop from TAG-Pyl has been validated. Initial studies suggested that the presence of an RNA stem loop structure adjacent to the TAG codon (a Pyl insertion sequence, or PYLIS) is required for cells to interpret the codon as Pyl. However, the PYLIS was later shown to have no effect on levels of Pyl incorporation in E. coli and no PYLIS or PYLIS-like structures could be detected in a subset of known Pyl-containing methylamine methyltransferase genes in archaea. Similarly, here sequences of genes encoding the Pyl-containing tRNA guanylyltransferase and PylB were examined, and no PYLIS was detected.
[0230] One strategy that avoids or minimizes the need for dual interpretation of a TAG sequence is to eliminate its use as a stop codon. In fact, it has long been thought that low TAG frequency ( _5%) represents a hallmark of archaea that encode Pyl-related genes, yet a few outliers have been reported. Thus, the percentage of genes with an in-frame TAG codon in the genomes of Pyl archaea was calculated. In taxa that intermix phylogenetically with non-Pyl genomes, the genome-wide TAG frequency was always high (>15%) (Fig. 1). Thus, low TAG frequency is not a core feature across Pyl archaea. The discrepancy from expected TAG content in many Pyl archaea highlights the need to determine the distribution of TAG codons and investigate how they are used.TAG is recoded genome-wide as a sense codon
[0231] To determine how TAG is used by Pyl archaea, it was first predicted the genes in each organism using genetic code 11 (standard for bacteria and archaea). All TAG codons were then re-predicted as sense codons and the resulting genes were classified into four categories (Fig. 3, Fig. 9). Genes in Category 1 contain one or more TAG codons such that the sequences before and after the TAG codon possess >30% sequence identity to a known protein (see Methods). In addition, readthrough of the TAG codon in Category 1 genes results in an extension of >15 amino acids (aa).Based on this analysis, it was predicted that genes in Category 1 interpret the TAG as an internal sense codon. Genes in Category 2 were characterized by an extension of <15 aa prior to sense codon overlap with another gene. Genes in Category 3 were characterized by an extension of >15 aa and / or led to a short (<60 bp) gene overlap and / or to an in-frame gene fusion. For Category 2 and 3, it was not possible to predict whether TAG is a stop or sense codon. For Category 4, read-through led to a large (>60 nt) out-of-frame overlap. In this case, there is no support for TAG reassignment, as large gene overlaps (>60 bp) are rare in prokaryotes. The prevalence of Category 1 and absence or extreme rarity of Category 4 cases in a genome may be an indicator of genome-wide TAG reassignment.
[0232] Contrary to the expectation that TAG readthrough would primarily lead to short gene extensions (Category 2) due to backup stop codons, Pyl archaea from six genera had numerous Category 1 TAG codons (Fig. 4). In fact, in a subset of the Thermoplasmatota and Euryarchaeota, most TAG codons fall into Category 1 and there are very few, or no, TAG codons that could lead to a large gene overlap (Category 4). This pattern is consistent with reliance only on TGA and TAA as stop codons. It is inferred that in these genomes, TAG is completely recoded as a sense codon (Pyl). Thus, experimental evidence was sought to confirm that these archaea interpret the TAG codon as Pyl.Proteomic validation of Genetic Code 34
[0233] To test for recoding of TAG, isolates of two archaea were grown with large fractions of Category 1 TAG codons and mass spectrometry was used to identify peptides with the exact masses predicted if Pyl is present. Methanococcoides Burtonii is a psychrophilic Halobacteriota from an Antarctic Lake and Methanomethylophilus alvus is a Thermoplasmatota from the human gut. To increase the probability of detection of the specific peptides of interest both trypsin and endoproteinase Glu-C, were used which cleave at different amino acids and thus increased the number of peptides suitable for measurement. Although comprehensive detection of all peptides is not possible, targeted peptide detection greatly improved the ability to identify Pyl- containing peptides (see Methods). Peptides with the mass / charge ratio expected for a Pyl peptide were manually validated by analysis of the MS / MS spectra.
[0234] For both M. burtonii and M. alvus, the proteomic data confirm genome-wide recoding of TAG as a sense codon. The detection of >2 unique peptide matches in the global tryptic or Glu-C datasets was used to confirm expression of proteins. Notably, for allexpressed Pyl proteins, peptide coverage after Pyl (encoded by TAG) was detected for proteolytic peptides (>6aa and <50aa). For M. burtonii, peptides containing Pyl were detected in 27 of 39 expressed Pyl proteins. Similarly, for M. alvus, Pyl was detected in 12 of the 15 expressed Pyl proteins. Expressed proteins for which Pyl was not detected had low coverage levels and / or had a proteolytic peptide that was too short or too long for detection. A database was constructed to recognize any of the standard amino acids at the TAG=Pyl / stop position, and peptides with another amino acid in the TAG position were not detected. Pyl was also detected in 24 additional proteins that had low levels of expression (< 2 unique peptides matches in the global datasets). In total, Pyl was detected in 56 proteins not previously shown to contain Pyl.
[0235] Examples of new proteins found to contain Pyl in M. burtonii include 12,18- didecarboxysiroheme deacetylase, thiouridine synthase, DNA helicase, his tRNA synthetase, isocitrate dehydrogenase, epimerase, glycosyltransferase, DNA polymerase, and multiple hypothetical proteins (Fig. 5A, 5B). New Pyl proteins in M. alvus include transacetylase-like protein, a methyltransferase, kinase, acetyltransferase, hypothetical and DUF-containing proteins, AAA family associated protein, and others (Fig. 5C). Pyl was also identified in trimethylamine methyltransferase as previously reported. Thus, it was concluded that M. burtonii and M. alvus archaea have adopted a previously undefined genetic code, with 62 sense codons encoding 21 amino acids, and only two stop codons. It is probable that the other Pyl-encoding archaea for which evidence was detected for numerous in-frame TAG codons (Category 1) also use this genetic code.Over a thousand proteins were predicted to contain Pyl
[0236] After establishing that select archaea use TAG as a dedicated sense codon, 1022 unique Pyl-containing proteins were predicted in genomes inferred to use this new genetic code. Of these 1022 proteins, 276 have TAG in a position of the gene inferred to be a sense codon based on alignment of profile hidden Markov models of conserved proteins. The largest number of predicted Pyl proteins of a single organism, 182 in total, occurs in Methanomicrobia JDFR-19. This archaeon is from crustal fluids collected from oceanic crust near the Juan de Fuca Ridge, and is a recoding outlier with a 13% TAG frequency.
[0237] Pyl occurs in numerous protein families, but the only Pyl-containing proteins that archaea have in common are the methylamine methyltransferases. The proposedfunction of Pyl is to activate and orient the methylamine for methyl transfer to a cognate corrinoid protein, via the reactivity of the electrophilic imine bond. It remains to be established what role Pyl plays in the enzymology of the new Pyl-containing proteins. Outside of the methyltransferases, Pyl proteins have a limited phylogenetic range, such that Pyl occurs in a protein within a single genome or genus, as a singleton. Pyl singletons may arise due to neutral mutations, as described for Pyl in tRNA-His guanylyltransferase. In methylcobalamin: coenzyme M methyltransferase, Pyl is in a different position in the amino acid sequence in organisms belonging to two genera (positions 90 and 289 in Methanosarcina and Methanohalophilus respectively). However, both of these residues occur in the same region of the folded protein, in or near the active site (Fig. 6).
[0238] Pyl is in many proteins involved in energy metabolism and DNA processing, including citrate synthase, isocitrate dehydrogenase, aminoacyl-tRNA synthetases, DNA polymerases, helicases. Additionally, transposases and CRISPR Cas proteins contain Pyl. Pyl was previously predicted in transposases and the Cas1 and Cas3 proteins. Here, Pyl was observed in Cas1 via targeted proteomic analysis, and 12 new transposase families with Pyl were identified, with some occurring in many copies per genome. One example is Methanosarcina mazei strain TMA, which has 49 copies of the IS66 family transposase, with each copy containing Pyl. Pyl is also in Cas2, Cas4, Cas8e, Casio, and there is a region with multiple recoded proteins in a Cas1 Casposon. Casposons are a superfamily of mobile elements that encode Cas1 , family B DNA polymerase, and conserved, uncharacterized protein domains.
[0239] Archaeal TnpB and IscB proteins both are predicted to contain Pyl. TnpB and IscB are RNA-guided nucleases that are the likely ancestors of Cas12 and Cas9, respectively. In IscB, the Pyl residue occurs within the active site (bridge helix) (Fig. 7). The Pyl- containing nucleases identified in this study represent a diverse and overlooked set of potential gene-editing proteins. In studies of IscB and TnpB nucleases, the Pyl- containing variants are always either mis-annotated as fragments, hypothetical proteins, or missed altogether (Fig. 10), and the Pyl-casposon was assumed to be inactivated due to the internal TAG codons. The findings of the current study motivate experimental work to test the functions of these and other newly defined Pyl- containing proteins.PylRS — tRNApylpairs from Genetic Code 34 organisms support genetic code expansion in E. coli
[0240] The archaeal Pyl system is widely used for genetic code expansion (GCE), a versatile synthetic biology tool that introduces non-canonical a- and [3-hydroxy and amino acids site-specifically into ribosomal products during translation to expand, probe, or control protein function. GCE relies on orthogonal aminoacyl tRNA synthetase / tRNA pairs that are selective for the monomer of interest. The PylRS — tRNApylpair is uniquely suited for GCE because it is naturally orthogonal to the endogenous aaRS / tRNA pairs found in bacteria, yeast, and mammalian cells. Additionally, PylRS enzymes remain active, even in vivo, towards substrates whose side chains and / or - a-substituent differ substantially from that of Pyl itself. This feature allows many monomers that are structurally similar to Pyl 1 to be accepted by the enzyme (Fig. 8A).
[0241] The high incidence of TAG codon(s) encoding Pyl within genes of some archaeal genomes suggests a robust system for monomer incorporation. It was reasoned that PylRS — tRNAPylpairs from such organisms should support GCE. To test this idea, it was asked how well the PylRS — tRNAPylpairs from Genetic Code 34 organisms with the highest levels of natural TAG would incorporate Boc-Lys 2 or HO-Boc-Lys 3 at an in-frame TAG codon located at position 3 of super folder green fluorescent protein (3TAG-sfGFP). Boc-Lys 2 and HO-Boc-Lys 3 are convenient and commonly employed analogs for Pyl 1 (Fig. 8A). At least one representative sequence from each recoded genus was evaluated. In addition, the presence of separately encoded PylSc and PylSn domains in the JDFR19 genus. As prior studies have found PylRS enzymes naturally lacking PylSn to be active, 2 JDFR19 PylRS homologs with only their PylSc domain were also evaluated. All in all, 8 new PylRS — tRNAPpylpairs were tested for their ability to support GCE in E. coli.
[0242] All PylRS — tRNAPylhomologs were introduced into BL21 E. coli cells on a pMEGA plasmid alongside a pET22b plasmid encoding 3TAG-sfGFP. Initial screening was performed in 96-well plate format where the change in OD6oo and fluorescence at 528 nm (F52s) were measured over 20 h (Fig. 8B, Fig. 13) as a proxy to evaluate amber suppression for given PylRS - tRNAPylhomologs. When cells containing one of the 8 PylRS — tRNAPylpairs tested were supplemented with 1 mM of monomer 2, two pairs - those from LMO and levihalophilus - supported a 4- or 10-fold improvement increase in F528 / OD600 signal over background, respectively. Notably, five homologs show evidence of activity when supplemented with 1 mM of monomer 3, suggesting a preference for HO-Boc-Lys 3 over Boc-Lys 2, consistent with previously reportedsubstrate preferences for other PylRS enzymes. LC-MS analysis of intact sfGFP purified from preparative-scale growths of all cells showing > 2-fold improvement in sfGFP expression over background revealed, in all cases, a mass corresponding to the introduction of Boc-Lys 2 or HO-Boc-Lys 3 at a single position (Fig. 8C, Fig. 11, Fig. 12). PylRS originating from Methanohalophilus levihalophilus performed comparably to / WaPyIRS.Clues to the pathway to genome-wide TAG reassignment
[0243] After investigating the Pyl proteins and the PylRS — tRNApylsystems of archaea with genome-wide TAG recoding, possible clues to the code being adopted were examined. The adoption of Genetic Code 34 may follow several stages.
[0244] First, the Pyl system is acquired via lateral transfer of a gene cassette and is associated primarily or exclusively with methylamine methyltransferases.
[0245] Since the PYLIS is not required for Pyl incorporation, the Pyl machinery may incorporate Pyl at normal stop sites in these genomes. It has been suggested that Pyl incorporation at stop positions would not be highly detrimental and protein extension would be limited due to backup stop codons (and non-functional proteins would be degraded)5. However, release factor 1 (RF1 ; recognizes TAG and TAA) is present in all complete and near-complete archaeal genomes examined and may compete with the Pyl tRNA for the TAG codon, reducing the instances of Pyl incorporation at stop positions. This is analogous to the 'ambiguous intermediate' evolutionary model for recoding proposed for CPR bacteria and eukaryotes. The subset of genomes with many Category 4 and one or very few Category 1 cases (Group A) likely recently acquired the Pyl machinery (all have > 15% TAG codons) thus may represent the first stage of code transition. Notably, Group A is phylogenetically very diverse, often with a single example within a phylum.
[0246] Over time, the ability to incorporate Pyl at TAG codons would lead to internal TAG codons introduced by mutation, and a decrease in deleterious cases of TAG as a sense codon. A subset of genomes have a small number of Category 1 TAG codons, while still containing Category 4 cases, so it was inferred they may not be fully recoded. Genomes in this category typically have intermediate (>5-15%) fractions of TAG codons. Genomes of this type were designated as Group B. Code transition, defined as where TAG is never a stop codon, may be inhibited by the fraction of Category 4 codons. For this reason, TAG frequency would decrease to avoid deleterious effects of TAG reassignment and the genomes most able to undergotransition to code 34 would be those with overall lower frequencies of TAG stop codons.
[0247] Subsequently, random mutations would lead to an accumulation of internal (Category 1) TAG codons in genomes and use of TAG as a stop codon would be eliminated. These genomes have low TAG frequencies (< 5%) and Category 1 cases dominate. This was designated as Group C. Proteomic data confirm that TAG is consistently interpreted as Pyl. Further, cases designated as ambiguous Category 2 or 3 correspond to Pyl. For the two organisms investigated using proteomics, this is confirmed.
[0248] Finally, mutation would allow the proteome to explore the potential value of incorporation into proteins and the instances of Pyl proteins would increase. These genomes would have the majority of cases classified as Category 1 , few to no Category 4 (and the ambiguous Category 2 and 3 cases would correspond to Pyl). This stage, representing organism that underwent code transition long ago, is designated as Group D, and exemplified by JDFR-19 genome, which has up to four Pyl within a single protein. This genome has 13% TAG because there has been sufficient time to reincorporate the TAG codon at positions that are neutral or beneficial.
[0249] TAG re-coding is likely a one way street, especially for archaea that have adopted code 34 (Group C and D). Loss of the Pyl system would result in loss of not only the ability to use methylamines, but also function of Pyl proteins. For code 34 archaea, loss of the ability to synthesize and incorporate Pyl would be far more catastrophic, as regain of function of the Pyl proteins would require numerous codon adjustments. Thus, it was concluded that gain of the ability to use Pyl can be readily acquired, so long as code transition can occur without major detriment, but reversal of full code transition may be unlikely.MethodsIdentification and annotation of Pyl- encoding Archaeal genomes
[0250] Genomes encoding Pyl were identified based on the presence of cellular machinery required for production and incorporation of this amino acid. NCBI BlastP was searched against the non-redundant protein database with the pyrrolysyl-synthetase (PylSc) sequence from Methanosarcina Barkeri on December 1 , 2023 and included genomes with >40% sequence identity to M. barkeri PylS sequence. Additionally, all genomes from each genus predicted to encode Pyl were downloaded based on taxonnumber from the NCBI website at yyw on Decern ber 1 ,2023. For these genomes, Prodigal V.2.6.3 was used to predict genes using default parameters (single genome mode, translation table 11), and hmmer V3.3.1 was used via hmmscan against the NCBI-fam database V.13. This HMM gene annotation used to identify the Pyl pathway are as follows: TIGR03912 (PylSn), TIGR02367 (PylSc), TIGR03910 (PylB), TIGR03909 (PylC) and TIGR0391 (PylD). Genomes identified as encoding Pyl contain at least 2 of the Pyl-specific HMMs. However, genomes with the following criteria were excluded from analysis: more than 50 % gaps, and / or completeness estimate of <60% and / or contamination (redundancy) estimate of >10%. Archaeal families containing genomes that putatively encode the Pyl machinery but meet the exclusion criteria (Persephonarchaeia, Korarchaeia, and Wukongarchaeia) are indicated by an asterisk in the Pyl-encoding genome distribution Fig. X. A recent Methanoglobus genome encoding Pyl was reported and released on February 25, 2024 during the preparation of this manuscript. This Methanoglobus genome was added to the compilation of Pyl-encoding archaea and included for genomic analyses. For phylogeny, NCBI taxonomy is included for reference but GTDB taxonomy is used because it contains a narrower classification (family instead of phylum level) for some of the Pyl genomes.Phylogenetic analysis of Pyl distribution
[0251] The reference species tree was built from a concatenation of the sequences of 36 universally conserved proteins from the phylosift dataset together with sequences of L30 and S4 ribosomal proteins and of the RNA polymerase subunit A and B. Proteins were retrieved from the selected proteomes with an HMM search (hmmer v3.3.2 [hmmer.org / ]). They were aligned with MAFFT (v7.453) with the accuracy-oriented methods, L-INS-i, trimmed with BMGE (v2) with the Blosum30 parameter, and concatenated. A maximum likelihood tree was built with IQ-TREE (v2.0.6) using the LG+F+R10 model. This tree was used as a guide tree to run the Posterior Mean-Site Frequency (PMSF) model under the LG + C60 + F + G4 model LG+C60+F+G implemented in IQ-TREE using a 1 ,000-replicate, ultra-fast bootstrap. Genome taxonomy was determined using GTDBTk 2.1.1. Taxonomy, TAG categories and TAG% were mapped on the species tree using ITOL server.Quantification of ORF TAG frequency
[0252] To quantify the mean TAG frequency of archaeal taxa, Pyl-encoding genomes were dereplicated for each phylum with dRep v.3.4.1 with default parameters. Then, Prodigal v2.6.3 was used to predict ORFs using default parameters and the -d option used for nucleotide output. The python script “find_tag_end.py” was used to retrieve all genes with an in-frame TAG codon. The TAG frequency was calculated as the total number of predicted ORFs containing an in-frame TAG codon, divided by the total number of predicted ORFs.Identification and quantification of genes by TAG codon category
[0253] To quantify TAG codon categories, genes with in-frame TAG codons were first extracted. NCBI genome coding sequence (CDS) files were used for gene identification. Prokka V.1.14.5 was used for gene identification for genomes without NCBI CDS files. Each TAG codon was treated as a sense codon, and sequences containing TAG were extended until the nearest in-frame TAA or TGA stop codon. The resulting genes were then translated and annotated using the UniProt- SwissProt database version 2023_05 and the Uniprot Archaeal and Bacterial Reference Proteome databases version 2023_05 (November 8, 2023 release). TAG- containing genes were then grouped into one of the four designated categories, based on the location of the TAG codon: (1) Internal (2) Short extension (3) Ambiguous and (4) Likely stop. Each gene was assigned to only a single category, using the lowest Category number for which criteria was satisfied. Genes that were truncated due to being located at the end of a contig were excluded when extension length could not be determined. Category 1 (Internal) genes contain an internal TAG codon, and encode proteins for which the amino acid sequence before and after the TAG codon has >30% sequence identity with sequences from the UniProt databases. Additionally, for Category 1, the gene must also encode a protein that is >15 amino acids longer than if TAG was a stop. For Category 2 (short extension), readthrough of TAG results in a protein sequence extension of <15 aa. Sequences overlapping an adjacent gene by >1 sense codon were excluded from Category 2. The Category 3 (Ambiguous) designation was assigned for genes if TAG readthrough resulted in (a) gene overlap of <60 bp with an adjacent gene, and / or (b) genes encoding a protein with >15 aa extension and / or (c.) gene fusion (in-frame overlap) with an adjacent gene. Finally, TAG readthrough in Category 4 (“Likely stop”) genes results in overlap of an adjacentgene by >60 bp, including overlap of an entire gene in a different reading frame or oriented in the opposite direction.Identification of Pyl- containing proteins
[0254] The genetic code prediction tool, Codetta, was initially used to identify Pyl proteins. By default, Codetta does not detect the recoding of TAG to Pyl. First, Codetta completes a six-frame translation of a genome nucleotide sequence that is aligned to profile HMMs. The software then recognizes recoding events when a codon has a specific nonstandard interpretation based on profile HMM alignments. For Codetta to complete this task, there must be a single or primary meaning that can be deduced from the alignments, for example, Codetta detected reassignment of AGG (a canonical arginine codon) to methionine'^. When TAG is recoded to Pyl, TAG occurs in place of different codons (there is no consensus for which codon it replaces). Thus, Codetta does not determine a single amino acid interpretation for TAG, so it does not detect recoding by default. Identification of known Pyl proteins may be enabled with Codetta, but Pyl detection is in this case limited exclusively to methylamine methyltransferases because they are the only proteins with Pyl- containing profile HMMs available for Codetta. However, the alignment output produced by Codetta includes gene matches (profile HMMs) for each codon, including TAG. Using this Codetta alignment output, Pyl-containing proteins were identified. The alignment step was implemented using the Codetta V.2.0 codetta_align.py command with a compatible profile HMM database prepared with the hmmer V.3.1 hmmbuild command (-enone option) for the PGAP v.15 database, which includes TIGRFAM v.15.0 entries.
[0255] Identification of Pyl-containing proteins was also performed by manual curation of genomes. Annotation with TAG as a sense codon was imported into Geneious v.2023-07-20. Then a secondary annotation for ORFs was performed by modifying the protein prediction software, Prodigal, to include TAG as a sense codon in the archaeal genetic code. The modified version of Prodigal is available at github(dot)com / VeronikaKivenson / Prodigal. Then Geneious was used to visualize and overlay the standard genetic code annotation and the TAG readthrough annotation. Sequences with in-frame TAG codon(s) were searched against Blast to identify matches as previously described.Structural prediction of Pyl- containing proteins
[0256] Structures for homologous proteins were obtained from RCSB PDB. Structure 4AY8 was used for MtaA and structure 7UTN was used for IscB. The Visual Molecular Dynamics software package was used for visualization. Homology modeling was performed using Swiss-Model.Anaerobic cultivation
[0257] M. alvus was grown as previously described, with methanol and ruminal fluid. M. burtonii was obtained from Deutsche Sammlung fur Mikroorganismen und Zellkulturen (DSMZ) and was grown on DSMZ Methanococcoides medium 141c (with 5 mM trimethylamine hydrochloride), with the following modifications: instead of Wolin’s mineral solution and Wolin’s vitamin solution, ATCC trace mineral supplement (MD- TMS) and ATCC Vitamin Supplement (MD-VS) were used. M. burtonii was grown at room temperature in 15 mL Hungate tubes with a headspace of 85% N2, 10% CO2, 5% H2gas in Anaerobe Systems anaerobic chamber (AS-150).Sample preparation for LC-MS / MS
[0258] 20 mg cell pellets were resuspended in 100 mM ammonium bicarbonate, pH 8.0 and subjected to mechanical lysis by bead beating with 0.15 mm zirconium oxide beads for five min (Geno / Grinder 2010; SPEX). Resulting cell lysates were adjusted to 4% sodium dodecyl sulfate (SDS) / 10 mM dithiothreitol (DTT), incubated at 90°C for 10 min to denature and reduce proteins, and pre-cleared by centrifugation at 21 ,000 x g. Samples were adjusted to 30 mM iodoacetamide (Sigma) and incubated in the dark at room temperature for 20 min to alkylate cysteine residues. The alkylated proteins were further cleaned and prepared for proteolytic digestion via the protein aggregation capture (PAC) method. In brief, 300 pg of hydrophobic magnetic beads (1 micron, SpeedBead Magnetic Carboxylate; GE Healthcare UK) and acetonitrile (ACN) were added to protein samples for a final concentration of 80% ACN to induce protein aggregation on the beads. Samples were incubated for 20 min at room temperature. The aggregated proteins were placed on a magnet and the supernatant was removed. The pelleted beads were subjected to a wash with 1 mL of ACN, followed by 1 mL of 70% ethanol. The final wash buffer was removed from tubes and the beads were resuspended in 200 pL of 100 mM ammonium bicarbonate for in-solution proteolytic digestion. Protein amounts were quantified by corrected absorbance (Scopes) at 205 nm (NanoDrop OneC; Thermo Fisher). After protein quantitation, samples weredigested using either MS-grade GluC Endoproteinase or trypsin (Thermo Scientific) with a 1 :75 protease: protein (wt:wt) ratio at 37°C, shaking overnight at 600 rpm. An additional round of protease was added for a second 3-h digestion period at 37°C, shaking at 600 rpm. The resulting tryptic peptides were filtered on 10 kDa MWCO filter plate (AcroPrep Advance, Omega 10 K MWCO) at 1 ,500 * g and adjusted to 1% formic acid before quantification by NanoDrop OneC.LC-MS / MS
[0259] 3 pg of peptide solution were analyzed by automated 1 D LC-MS / MS analysis using a Vanquish ultra-HPLC (UHPLC) system plumbed directly in-line with a QExactive- Plus mass spectrometer (Thermo Scientific) with an in-house-pulled 100 pm inner diameter nanospray emitter (packed to 15 cm with 1.7 pm Kinetex C18 reverse-phase resin (Phenomenex)). Peptides were loaded, desalted, and separated by uHPLC under the following conditions: sample injection followed by 100% solvent A (98% H2O, 2% acetonitrile, 0.1% formic acid) from 0 to 30 min to load and desalt, a linear gradient from 0 to 30% solvent B (70% acetonitrile, 30% water, 0.1 % formic acid) from 30 to 220 min for separation, a column wash from 0 to 100% solvent B from 220 to 255 min, and 100% solvent A from 255 to 275 min for column re-equilibration. Eluting peptides were analyzed with the following MS settings for global searches: positive mode, data-dependent acquisition, top-20 method; mass range 400-1500 m / z; MS and MS / MS resolution 70 and 17.5 K, respectively; MS / MS loop count 20; isolation window 1.8 m / z; isolation offset 0.3 m / z; charge state exclusion of unassigned, +1 , +6-8 charges.
[0260] For targeted searches, parallel reaction monitoring inclusions lists were constructed to detect peptides of interest in the predicted pyrrolysine containing proteins. An in-silico digest of the proteins was conducted with UniPept using either GluC or trypsin as the protease with a minimum peptide length of six amino acids. Based on the predicted peptides from the in silica digest, peptides were selected for the PRM measurements if they were classified as any of the following peptide types: (1) the first proteolytic peptide in the protein, (2) the proteolytic peptide containing the pyrrolysine residue,(3) the first proteolytic peptide downstream of the pyrrolysine-containing peptide, and(4) the last proteolytic peptide in the sequence. For the peptides of interest, the PRM inclusion list included the monoisotopic precursor m / z of the peptide including a static modification of carbamidomethylation (+57.021) of cysteine and a charge state (z) of +2. No variable modifications of oxidation (+15.995) of methionine, deamidation of(+0.984) of asparagine and glutamine, or methyl loss (-14.016) for pyrrolysine were considered for the inclusion list. A maximum of 50 entries were monitored for each PRM measurement. The same LC-MS / MS method was used as the global measurements.Proteomics data analysis
[0261] Global database searches were conducted with Proteome Discoverer 3.0 (Comet / Percolator) against custom organism specific protein databases constructed from all predicted proteins in the respective, as well as common mass spectrometry contaminants. Protein sequences were included with the predicted pyrrolysine residue represented with “X” as well as the sequence version that terminated in a stop codon. For all database searches, the peptide mass tolerance was set at 20 ppm. Accepted modifications included static modifications of carbamidomethylation (+57.021) of cysteine residues and pyrrolysine (+237.148) of “X” residues. Variable modifications of oxidation (+15.995) of methionine, deamidation of (+0.984) of asparagine and glutamine residues, and a methyl loss (-14.016) for pyrrolysine (X) residues. A maximum of three variable modifications allowed per peptide. Semi- proteolytic digestion of peptides was allowed for either trypsin or GluC depending on the sample, and a 1% peptide-level FDR threshold was applied.
[0262] Preliminary confirmation of pyrrolysine expression was based on (1) a unique peptide match in the global datasets that passed the 1 % peptide-level FDR threshold for the pyrrolysine containing proteins or (2) a detected precursor mass in the PRM measurement that was within 10ppm of the monoisotopic precursor m / z of the peptides of interest. Peptides were considered unique if the amino acid sequence was not predicted in any other protein for the respective organism. To more confidently confirm pyrrolysine expression, MS / MS spectra from all peptides of pyrrolysine-containing proteins were manually inspected and validated. A MS / MS spectrum of a peptide was considered for evidence of pyrrolysine expression if each of following criteria was met: (1) Precursor ion (MS1) mass error <10 ppm, (2) Fragmentation ion (MS / MS) mass error <0.02 Da and (3) Location of the fragment ion along the peptide sequence, with either direct sequencing of the fragment ending in pyrrolysine or fragment ions that were longer than the pyrrolysine-containing fragment.
[0263] In addition to the database searches, the data were searched using a de novo assisted database search workflow with a database constructed to recognize any ofthe standard amino acids at the TAG=pyrrolysine / stop position. This workflow will recognize the 20 standard amino acids if they were detected at the Pyl position in the MS / MS spectrum. This workflow did not detect any peptides with another amino acid in the pyl position.Construction of plasmids encoding PylRS homologs into pMega vector
[0264] For in vivo GCE, the previously reported pMEGA vector (parent backbone Addgene Plasmid #200225) was utilized for aaRS and tRNA production. pMEGA plasmids containing sequences encoding PylRS homologs (Table A) were ordered as circular, dsDNA plasmids from Twist Bioscience. Upon attaining pMEGA containing PylRS homologs, corresponding pylT sequences (Table A) were cloned into each vector as follows. Circular pMEGA vectors were linearized using oligonucleotides P1 and P2 (Table B). PCR using the Q5® High-Fidelity 2X Master Mix (NEB M0492S) was performed according to manufacturer protocols for 30 cycles. The PCR reaction was subject to a PCR clean up with the QIAquick PCR Purification Kit (Qiagen, catalog # 28104). The concentration of purified, linearized products was determined by absorbance at 260 nm using a NanoDrop ND-1000 Spectrophotometer. Next, 33.3 ng of purified, linearized pMEGA and 100 ng of a gBIock containing the appropriate pylT sequence (Table A, Table C) were combined in a 20 pL Gibson Assembly reaction using HiFi DNA Assembly Master Mix (NEB, catalog #E2621 L) and incubated at 50°C for 1 h to generate a circular pMEGA vector containing the coding sequence for the appropriate pylT. The circularized plasmid from the previous step was transformed into NEB 5-alpha competent E. coli (NEB, catalog # C2987H) as follows. Frozen stocks of cells were thawed on ice for 10 min. Upon thawing, the entirety of the previous Gibson Assembly reaction was added to cells and incubated on ice for 30 min. Cells incubated with plasmid were then subjected to heat shock at 42°C for 30 sec and placed on ice for 2 min. 350 pL of Super Optimal broth with Catabolite repression (S.O.C.) outgrowth medium (NEB, catalog # B9020S) was added to cells and cells were incubated at 37°C for 1 h with shaking at 220 rpm. Agar plates containing spectomycin were inoculated with 50 pL of transformed cells and grown overnight at 37°C. 3 single colonies per construct were picked and inoculated into liquid cultures containing 5 mL LB + spectomycin and grown for 16 h at 37°C. Pure plasmid was isolated from 5 mL cultures using Qiaprep Spin Miniprep Kit (Qiagen, catalog # 27106) and sequences were confirmed by whole plasmid sequencing with Primordium Labs. The new pMEGA plasmids were double transformed with apET22b-3TAG-sfGFP into BL21 (DE3) cells following transformation protocol detailed below.Transformation protocol
[0265] BL21(DE3) E. coli (NEB, catalog # C2987H and C2527H, respectively) were transformed in accordance with manufacturer protocols with some modifications as follows. Frozen stocks were thawed on ice. Upon thawing, 100 ng of each relevant plasmid was added. After a 30 min incubation on ice, cells were heat-shocked for 30 sec at 42°C and allowed to recover on ice for 2 min. Following, 350 pL of Super Optimal broth with Catabolite repression (S.O.C.) was immediately added and cells were recovered at 37°C shaking for 1 h before plating 50 pL on LB agar plates containing appropriate antibiotics. Plates were incubated overnight at 37°C.Expression assays of 3T AG sfGFP
[0266] Starter E. coli cultures were grown overnight in 5 mL of LB Miller (AmericanBio, Catalog #AB01201) in 15 mL culture tubes supplemented with antibiotics at 37°C. Prior to cultures reaching OD, the appropriate amount of each monomer (stored as 25 mM stocks) were allotted into the wells of a black, clear bottom, 96-well plate (1 mM final concentration). Once cultures reached an OD6oo of 0.6 (roughly 3 h), protein expression was induced by addition of 1 mM IPTG. The culture was transferred into the wells of the plate to a total volume of 200 pL in each well. A Breathe-Easy® sealing membrane (Sigma Z380059) was placed over the 96-well plate and the plate was loaded into a BioTek Synergy HTX microplate reader with no lid. ODeoo and F528 values (Aex = 485 nm) were measured every 10 min for 20 h. The plate was maintained at 37°C and was shaken during this time. Points represented in Fig. 8B represent three biological replicates where three random colonies were picked on a plate from a single transformation.Expression and Purification 3TAG-GFP containing Boc-Lys or HO-Boc-Lys
[0267] Starter cultures of 5 mL of Miller’s LB Broth (AmericanBio, catalog # AB01201) supplemented with carbomycin and spectinomycin were inoculated with a single colony of BL21 (DE3) E. coli cells harboring pET22b-3TAG-sfGFP-6xHis alongside pMEGA-JDFR19-PylRS-pylT, pMEGA-JDFR19-PylRSCTD-pylT, pMEGA-B74G9- PylRSCTD-pylT, pMEGA-LMO1-PylRS-pylT, pMEGA-Bin14-PylRS-PylRS-pylT, pM EGA-1 D-PyIRS-pyIT, pMEGA-Levihalophilus-PyIRS-pyIT, or pMEGA-MaPyIRSand grown overnight for 16 h at 37°C with shaking at 220 rpm until the culture was saturated. The starter culture was used to inoculate a 100 mL expression culture in a 1 :100 dilution of Miller’s LB Broth supplemented with 1 mM Boc-Lys (Combi-Blocks, CAS # 2418-95-3) or 1 mM HO-Boc-Lys (Accela, CAS # 111223-31-5) with carbomycin and spectinomycin. The expression culture was grown at 37°C with shaking at 220 rpm to an OD600 of 0.6 at which point it was induced with 1 mM IPTG and grown for 16 h under the same conditions. The expression culture was harvested by centrifugation at 4,300 x g at 4°C for 30 min. The resulting cell pellet was suspended in 10 mL of Lysis Buffer (50 mM sodium phosphate pH 6.8 and 300 mM NaCI) containing half a tablet of complete, mini EDTA-free ULTRA protease inhibitor cocktail (Sigma-Aldrich, St. Louis, MO). The cell suspension was disrupted by sonication on ice (Branson Sonifier 250, 5 cycles of 30 sec pulse at 50% duty cycle and microtip limit of 5 followed by 30 sec pause). The cell lysate was cleared by centrifugation at 10,000xg at 4°C for 20 min. TALON® Metal Affinity Resin (1 mL) (Takara Biosciences, catalog # 635504) was equilibrated with Lysis Buffer, added to the cleared cell lysate, and incubated on a rotisserie at 4°C for 1 h. The TALON® resin-lysate mixture was then passed through a gravity flow Poly-Prep Chromatography Column (Bio-Rad Laboratories, Hercules, CA). Non-specifically bound proteins were removed by washing the TALON® resin with 10 mL of Lysis Buffer. The 6xHis-tagged protein was eluted by washing the TALON® resin with 2 mL of Elution Buffer (50 mM sodium phosphate pH 6.8 and 250 mM imidazole). The purified protein was quantified using absorbance at 280 nm, snap frozen as singleuse aliquots, and stored at -80°C. Protein yields are reported in Table D.Intact Protein LC-MS
[0268] LC-MS analysis of all protein samples were performed on an Agilent 1290 Infinity II HPLC connected to an Agilent 6530B QTOF AJS-ESI. The mobile phase for LC-MS was water and acetonitrile with 0.1% (v / v) formic acid at a flow rate of 0.4 mL / min. Each protein sample was injected onto an Poroshell 300SB-C8 column (2.1 x 75 mm, 5 pM, room temp, Agilent) and separated using a linear gradient from 5% acetonitrile for 0 to 2 min and ramping to 95% acetonitrile over 7.5 min, and then washing with 95% acetonitrile for 2 min. The following parameters were used during acquisition: Fragmentor voltage 225 V, gas temperature 300 °C, drying gas flow 10 L / min, sheath gas temperature 350 °C, sheath gas flow 11 L / min, nebulizer pressure 35 psi, skimmer voltage 65 V, Vcap 5000 V, 1 spectra / s.Example 2: An archaeal genetic code with all TAG codons as pyrrolysine
[0269] This example includes some information from the above example (Example 1), but also includes additional information.
[0270] Multiple genetic codes developed during the evolution of Eukaryotes and Bacteria, yet no alternative genetic code is known for Archaea. Proteomics were used to test if certain Archaea consistently incorporate pyrrolysine (Pyl) at TAG codons, supporting an alternative archaeal genetic code, designated herein the Pyl code. This genetic code has 62 sense codons encoding 21 amino acids. By contrast with monophyletic genetic code distributions in bacteria, the Archaeal Pyl code occurs sporadically, indicating that it arose independently in multiple lineages. It was discovered herein that over 1800 archaeal proteins contain Pyl, increasing the number of such proteins by two orders of magnitude. Additionally, five Pyl tRNA - synthetase pairs from Pyl code archaea were used to introduce pyrrolysine analogs into proteins in Escherichia coli.
[0271] The genetic code is neither universal nor immutable; subsequent to its initial discovery, analysis of genetic material from diverse lineages revealed substantial code variation. More recently, alternative genetic codes in which either TGA or TAG stop codons incorporate a standard amino acid were identified in bacteria and bacteriophages. Interestingly, these stop codons can also be selectively interpreted as either selenocysteine (the 21stamino acid, at TGA) or pyrrolysine (the 22ndamino acid, at TAG). Pyrrolysine (Pyl) is a derivative of lysine that features a pyrroline ring. Its biosynthesis requires the Pyl B,C,D genes, while its incorporation into proteins requires a pyrrolysyl-tRNA synthetase (PylRS) and a Pyl tRNA carrying a CUA anticodon. The Pyl B,C,D genes and PylRS / tRNAPylpair are referred to here as the Pyl machinery. The PylRS / tRNAPylpairs are widely used tools for genetic code expansion (GCE), in which non-canonical a-amino acids or their analogs are incorporated into ribosomal products to modify protein function. PylRS / tRNAPylpairs act on a-amino acid substrates with diverse side chains and are orthogonal in many bacterial and mammalian systems, showing limited reactivity with canonical a-amino acids, and neither interfere nor integrate with prokaryotic or eukaryotic aminoacyl- tRNA pairs.
[0272] In organisms that naturally encode the Pyl machinery, understanding of the mechanism that signals Pyl incorporation is incomplete and the extent to which Pyloccurs in proteins remains unclear. It has been suggested that a PYLIS, a structure analogous to the SECIS required to selectively signal selenocysteine incorporation, enables dual encoding of TAG. However, in a subset of genes that encode Pyl proteins, the PYLIS structure has not been detected. In bacteria, Pyl has been reported in one L-serine dehydratase and a few methylamine methyltransferases and otherwise, TAG is read as a stop codon. Pyl has been predicted to occur in proteins of 11 major groups of archaea. All archaea without Pyl machinery use TAG as a stop codon, and those with Pyl machinery are believed to also use TAG as a stop codon, except in the very few enzymes in which Pyl occurs. Additionally, Pyl has been hypothesized to occur within proteins when internal TAG codons were observed during in silica analysis (22, 33, 34). However, incorporation has been experimentally confirmed only in five proteins (mono-, di-, tri- methylamine methyltransferases, tRNA guanylyltransferase, and PylB). The archaeon Methanosarcina acetivorans has five additional proteins presumed to contain Pyl because their detected molecular masses are consistent with TAG readthrough. The latter authors suggest that TAG may not be selectively reassigned in M. acetivorans, but genome-wide incorporation of Pyl at this codon has not been established.
[0273] It was asked if some archaea broadly incorporate Pyl into proteins using an alternative genetic code, and whether these genomes would yield biotechnologically relevant PylRS / tRNAPylpairs. The analysis revealed that organisms within two Archaeal phyla contain multiple genes with internal in-frame TAG codons. Using global and targeted proteomics, two such Archaea were analyzed, a human gut and a marine isolate, and in both cases the genome-wide incorporation of Pyl at TAG codons was confirmed, including within multiple enzyme classes. On this basis the existence of a previously unrecognized genetic code is proposed herein, currently confined to archaea, and designated the Pyl code. It is confirmed that five PylRS / tRNAPylpairs predicted by this analysis efficiently introduce Pyl analogs into proteins in E. coli. This work demonstrates that Pyl code organisms can add distinct “parts” to the genetic code expansion toolbox and facilitate the synthesis of heteropolymers. The distribution of Pyl code among archaea is not monophyletic, which is consistent with multiple independent code transitions and hints that codon reassignment to Pyl could be facile. Features indicative of recoding are inferred that occurred recently compared with the distant past and propose a series of steps that may lead to genetic code transition.ResultsPhylogenetic distribution of Pyl
[0274] To investigate the phylogenetic distribution of archaea with Pyl machinery (Pyl- encoding archaea; Pyl archaea), a collection of genomes from all families known to incorporate Pyl was augmented, together with taxa identified that carry the Pyl machinery. The majority of Pyl archaea belong to the Halobacteriota (Methanosarcinia, Archaeoglobi and Methanonatronarchaia) and Thermoplasmatota (Methanomassiliicocci) phyla, with few others within the phyla of Asgardarcheota, Thermoproteota, and Hydrothermarchaeota (FIG. 23). Notably, organisms that encode Pyl machinery have a discontinuous phylogenetic distribution (FIG. 23), sometimes even at the genus level. For example, the Pyl pathway is encoded and actively utilized by a single member of the Archaeoglobus (Ca. Methanoglobus hypatiae LCB24). The proteins from this genome share 98.6% average amino acid identity with Archaeoglobi LCB024-003, yet LCB024-003 does not encode Pyl. Similarly, Pyl is limited to a single family within the Asgardarcheota and Hydrothermarchaeota phyla, and in the Ca. Methanomethylicus genus, only two of five archaea encode Pyl.
[0275] Given the intermixing of closely related archaea that do and do not encode the Pyl machinery, the phylogenies of the Pyl machinery (FIG. 27) were compared to that of the organisms (FIG. 23). Consistent with prior suggestions, discordance between the tree structures provides evidence for lateral transfer of the Pyl genes. For example, MSBL1 archaea (Hadarchaeota) may have gained their Pyl machinery through horizontal gene transfer from Methanonatronarchaeia, and “Ca. Methanoglobus” from a Thermoproteota.TAG frequency varies among Pyl archaea
[0276] Given that organisms almost always interpret TAG as a stop codon, it was reasoned that there may be detrimental consequences associated with the acquisition of Pyl machinery, such as the erroneous incorporation of Pyl and readthrough of the normal gene end. It is also possible that TAG codons possess a dual encoding in organisms that carry a Pyl machinery; however, there exist no validated mechanisms for recognizing a dual-message TAG codon. Initial studies suggested that an RNA stem loop downstream of the TAG codon (a Pyl insertion sequence, or PYLIS) is required to interpret a TAG codon as Pyl. However, consistent with subsequent reports that the PYLIS is not required for Pyl incorporation, the PYLIS motif flanking TAG codons ofgenes encoding Pyl-containing proteins was not found, except for a small subset of the methylamine methyltransferases. While there may be an unknown means by which TAG is interpreted independently for each gene, there is no evidence for this process at present. Regulatory sequences may be selected against due to the cost of maintaining additional contextual requirements.
[0277] Loss of the TAG codon would eliminate readthrough. It has long been thought that low TAG frequency (< 5%) is a hallmark of archaea that encode Pyl machinery, yet a few outliers have been reported. The percentage of genes with a TAG codon in the genomes of Pyl archaea was calculated, and it was found that the TAG frequencies are always high (>15%) among Pyl archaea that are intermixed phylogenetically with genomes lacking Pyl machinery (FIG. 23). The discrepancy from expected TAG content in many Pyl archaea highlights the need to determine the distribution of TAG codons and investigate their function.TAG is recoded genome-wide as a sense codon
[0278] To determine how TAG is used by Pyl archaea, the genes in each organism using Genetic Code 11 , which is standard for bacteria and archaea, were first predicted. All TAG codons were then re-predicted as sense codons and the resulting genes were classified into four categories (FIG. 28, FIG. 29). Category 1 genes contain one or more TAG codons in which the sequences before and after the TAG codon possess >30% sequence identity to a known protein. In addition, readthrough of the TAG codon in Category 1 genes extends the protein sequence by >15 amino acids (aa). Based on this analysis, it is predicted that genes in Category 1 read the TAG as an internal sense codon. Category 2 genes were characterized by short extensions of <15 aa, and had no overlap with sense codon(s) of an adjacent gene. This scenario was originally thought to be the main outcome of TAG readthrough. Genes in Category 3 were characterized by an extension of >15 aa after TAG readthrough (the sequence extension lacks similarity with known genes) and / or a short (<60 nt) gene overlap and / or an in-frame gene fusion. For Category 2 and 3, it is not possible to predict whether TAG is a stop or sense codon based on sequence analysis alone. For Category 4, readthrough led to a large (>60 nt) out-of-frame overlap. In this case, there is no support for TAG reassignment, as large gene overlaps are rare in prokaryotes. The prevalence of Category 1 and absence, or extreme rarity, of Category 4 cases in a genome may be an indicator of genome-wide TAG reassignment.
[0279] Notably, Pyl archaea from nine genera (Methanomethylophilus, Methanogranum, Methanoplasma, Methanimicrococcus, Methanosarcina, Methanohalobium, Methanohalophilus, Methanococcoides, and JDFR-19) have numerous Category 1 genes and very few, or no, Category 4 genes (Fig. 2). This pattern is consistent with reliance only on TGA and TAA as stop codons. It is inferred that in these genomes, TAG is completely recoded as a sense codon (Pyl).
[0280] It was predicted that 1903 Pyl-containing proteins in genomes inferred to use the Pyl code. Contrary to its initial discovery in a narrow class of enzymes, Pyl is incorporated into a functionally diverse set of proteins. The largest number of predicted Pyl proteins of a single organism, 191 in total, occurs in JDFR-19 (Methermicoccaceae). This archaeon is a recoding outlier with a 13% TAG frequency. The position of Pyl is conserved in a GAF domain-containing PAS family protein in genomes of three archaeal genera (FIG. 30). Pyl also occupies a position close to RNA and / or DNA binding sites in the CRISPR-Cas9-related nuclease IscB and in the binding site of methylcobalamin:CoM methyltransferase (FIG. 31 8. FIG. 32). However, while Pyl- containing proteins are-functionally enriched in Gene Ontology (GO) terms associated with methylamine methyltransferases, the functions of other proteins that incorporate Pyl are not consistent within or between evolutionary clades. All complete genomes with Pyl code encode one or more of the methylamine methyltransferases, the only Pyl-containing proteins that archaea have in common. Outside of the methylamine methyltransferase, proteins found to incorporate Pyl have a limited phylogenetic range, such that Pyl is in a protein within a single genome or genus, as a singleton. This is consistent with a neutral role for Pyl in nearly all proteins where it occurs, and suggests that the necessity of methylamine metabolism alone drove the evolution of the Pyl code. Next, experimental evidence was sought to confirm that these archaea read the TAG codon as Pyl.Proteomic validation of the Pyl code
[0281] To test for recoding of TAG, two isolates were chosen from different lineages and environments with genomes that are predicted to contain a large fraction of Category 1 genes and used mass spectrometry to identify proteolytic peptides and their subsequent fragmentation whose exact masses confirmed the presence of a Pyl residue. Methanococcoides burtonii is a psychrophilic Halobacteriota from an Antarctic lake and Methanomethylophilus alvi (formerly M. alvus) is a Thermoplasmatota from the human gut. To increase the probability of detection of thespecific peptides of interest, both trypsin and endoproteinase Glu-C were used, which cleave between different amino acids and thus increased the number of peptides suitable for analysis. Pyl-containing tryptic or GluC-fragments will be observed only if two criteria are met: the Pyl-containing protein is sufficiently abundant, and the Pyl- containing fragment(s) possesses the appropriate mass and charge to be detected. Although comprehensive detection of all peptides is not possible owing to their variable abundance, searching specifically for peptides of interest based on their predicted masses (targeted peptide detection) greatly improved the ability to identify Pyl-containing peptides (see Methods'). Peptides with the mass-to-charge ratio (m / z) expected for a Pyl-containing tryptic or GluC-product were manually validated by analysis of the MS / MS spectra.
[0282] For both M. burtonii and M. alvi, the proteomic data confirm genome-wide recoding of TAG as a sense codon. The detection of >2 unique peptide matches in the global tryptic or Glu-C datasets was used to confirm protein expression. Notably, for all proteins that were expressed and predicted to contain Pyl, and for which a proteolytic peptide (>6aa and <50aa) spanning the region predicted to contain Pyl was detected, the presence of Pyl was confirmed. For M. burtonii, peptides containing Pyl were detected in 27 of 39 expressed Pyl proteins. Similarly, for M. alvi, Pyl was detected in 12 of the 15 expressed Pyl proteins. Expressed proteins in which Pyl was not detected had low sequence coverage and / or had a Pyl-containing proteolytic peptide that was too short or too long for detection. As an additional confirmation of Pyl incorporation at TAG codons, de-novo-assisted database searches were conducted to enable identification of any of the standard amino acids at the TAG codon. No peptides were detected with another amino acid in the TAG position in the de novo- assisted database searches. Pyl was detected in 24 additional proteins that had low levels of expression (<2 unique peptides matches in the global datasets). In total, Pyl was detected in 54 proteins not previously shown to contain Pyl.
[0283] Examples of proteins found to contain Pyl in M. burtonii include 12,18- didecarboxysiroheme deacetylase, thiouridine synthase, DNA helicase, histidine tRNA synthetase, isocitrate dehydrogenase, epimerase, glycosyltransferase, DNA polymerase, and multiple hypothetical proteins (FIG. 25A, FIG. 33A). Additionally, Pyl in CRISPR Cas1 in the M. burtonii proteomic data was observed, as were other archaeal genomes encode both Pyl- containing CRISPR Cas proteins and RNA- guided nucleases which have been missed or mis-annotated in previous studies (FIG. 34). Pyl proteins in M. alvi include transacetylase-like protein, a methyltransferase,kinase, acetyltransferase, hypothetical and DUF-containing proteins, and a AAA family associated protein (FIG. 25B). Pyl was also identified in trimethylamine methyltransferase (FIG. 33B), as previously reported. Thus, it was concluded that M. burtonii and M. alvi archaea have adopted a nonstandard genetic code with 62 sense codons encoding 21 amino acids, and only two stop codons. It is probable that the other Pyl-encoding archaea for which evidence was detected for numerous internal TAG codons (Category 1) also use the Pyl code.PylRS / tRNAPylpairs from Pyl code organisms support genetic code expansion (GCE) in E. coli
[0284] PylRS / tRNAPylpairs enable genetic code expansion, yet their widespread use has been hindered by unpredictable stop codon suppression. Pyl code organisms appear to have naturally overcome this challenge through their PylRS / tRNAPylsystems. The high incidence of TAG codon(s) encoding Pyl within genes of some archaeal genomes suggests a robust system for Pyl incorporation that may be leveraged for the site-specific incorporation of unnatural monomers into ribosomal products. Indeed the PylRS / tRNApylpairs from M. mazei, M. barken, and M. alvi are already widely used for GCE in bacterial, yeast, and mammalian cells. Here, it was tested whether PylRS / tRNAPylpairs from Pyl code archaea support the incorporation of two widely used analogs of Pyl (1), the a-amino acid Boc-Lys (2) and the a-hydroxy acid HO- Boc-Lys (3) (FIG. 26A), into a model protein, super folder green fluorescent protein (sfGFP), containing an in-frame TAG codon located at position 3 (sfGFP-3TAG). At least one representative sequence from each recoded genus was evaluated. Eight PylRS / tRNAPylpairs were tested for their ability to support genetic code expansion in E. coli, using the PylRS / tRNAPylpair from M. alvi as a benchmark for current GCE efficiency levels. Initial screenings evaluated whether organisms supplemented with Boc-Lys (2) or HO-Boc-Lys (3) developed higher levels of sfGFP fluorescence (F528) as a function of cell growth (ODeoo) over 20 hours (h) (FIG. 26B, FIG. 35). When BL21(DE3) E. coli were supplemented with Boc-Lys (2), the PylRS / tRNA^1pairs from Methanomethylophilus alvi and Methanohalophilus levihalophilus developed substantial levels of sfGFP compared to non-supplemented growths, while the PylRS / tRNAPylpair from Methanococcocoides LMO-1 supported moderate sfGFP expression. SDS-PAGE analysis of sfGFP-3TAG proteins expressed in BL21 (DE3) E. coli supplemented with either 1 mM of Boc-Lys (2) or HO-Boc-Lys (3) indicates that the major protein product has a molecular mass of approximately 27 kDa, theexpected molecular weight of sfGFP (FIG. 26C, FIG. 36). LC-MS analysis of intact sfGFP-3TAG purified from cells expressing the PylRS / tRNAPylpair from Methanomethylophilus alvi and Methanohalophilus levihalophilus, confirmed the introduction of Boc-Lys (2) or HO-Boc-Lys (3) into sfGFP-3TAG (FIG. 26D).The activity of the PylRS / tRNAPylpair from Methanohalophilus levihalophilus is sufficient to support robust genetic code expansion with a-amino acid and a-hydroxy acid Pyl analogs in E. coli.
[0285] A surprising result was the greater effectiveness of some PylRS / tRNAPylpairs for incorporation of a-hydroxy acid HO-Boc-Lys (3) compared to the a-amino acid Boc- Lys (2). In particular, the PylRS / tRNAPylpairs from Methanomicrobia JDFR-19 and Methanogranum 1 D supported sfGFP expression only in the presence of HO-Boc-Lys (3). This finding is important because, unlike Boc-Lys (2), HO-Boc-Lys (3) generates an ester linkage within the protein backbone. Site-specific ester linkages can be exploited to post-translationally edit the protein backbone to introduce backbone elements that cannot otherwise be introduced in cells, such as p2, y- and b-amino acid linkages and reactive 1,3-diketones. LC-MS analysis of intact sfGFP-3TAG purified from large-scale E. coli growths expressing PylRS / tRNAPylpairs from JDFR 19 and Methanogranum 1 as well as Methanohalobium bin 14 and supplemented with a- hydroxy acid HO-Boc-Lys (3) showed unambiguous evidence for the ester linkage (FIG. 26C, FIG. 37). The selectivity for an a-hydroxy acid substrate by the PylRS / tRNAPylpairs identified here avoids well documented problems associated with metabolic conversion of a-hydroxy acids into a-amino acids that would circumvent post-translational backbone editingDiscussion
[0286] The patterns of TAG use in archaea that incorporate Pyl suggest steps in the process by which the Pyl code is established (FIG. 38). An early stage, represented by Group A, an intermediate stage by Group B, and a final stage, with genome-wide reassignment of the TAG codon, by Group C (Fig. 2), is herein proposed.
[0287] Group A is characterized by internal TAG codons exclusively in the methylamine methyltransferases and high TAG codon frequencies, similar to those in archaea lacking the Pyl system (FIG. 23). Additional features of Group A are patchy taxonomic distribution of the Pyl machinery, and incongruence between Pyl machinery and organismal phylogenies. These traits are consistent with recent acquisition of the Pyl cassette, potentially by lateral gene transfer. Upon acquisition of the Pyl system, TAGcodons that were used exclusively as stop codons may be interpreted as Pyl. Unintended stop codon readthrough can be mitigated or prevented by tandem stop codons that can reduce extension after TAG readthrough. Substrate dependent transcriptional regulation can minimize TAG readthrough, as some organisms downregulate Pyl machinery when not grown on methylamines. Finally, release factor 1 (recognizes TAG and TAA) is present in all complete and near-complete archaeal genomes examined and may compete with the Pyl tRNA for the TAG codon.
[0288] A second, transitory stage in the evolution of the Pyl code, exemplified by Group B genomes, involves the progressive reduction of TAG codons to reduce stop codon readthrough, intermediate TAG codon frequencies, and a higher prevalence of internal (Category 1) TAG codons compared to Group A (Fig. 2, FIG. 39). The organisms in this group are obligate methylotrophic or methyl-reducing methanogens in environments where methylamine is highly available, such as saline environments where they are produced from os mo protecta nt degradation and the animal gut where methylamine precursors are abundant in certain diets.
[0289] At the third and final stage, Group C genomes have fully recoded TAG to a sense codon. The defining feature of Group C archaea are abundant internal (Category 1) TAG codons, usually coupled with low TAG frequencies. Like Group B, Group C archaea are methanogens specialized in methyl-compound utilization and generally live in methylamine-rich environments. The constant requirement for Pyl in methylamine methyltransferases may enable loss of Pyl regulation, leading to insertion of TAG in many genes, including those essential for function. Critically, the insertion of TAG in essential genes would require constitutive expression of the Pyl machinery, even if cells grow on compounds other than methylamines. Presence of TAG in essential genes and constitutive expression of the Pyl machinery would establish the Pyl code.
[0290] Once the Pyl code has been established, code reversal would result in loss of methylamine metabolism and premature truncation of essential genes. These barriers to reversal and the sporadic distribution of the Pyl code suggests multiple independent code transitions. Ancestral character reconstruction indicates that Pyl code may have emerged independently five times (FIG. 24). Thus, it was inferred that transition to Pyl code is facile and more easily gained than lost. Interestingly, this parallels inferences of ease of code switching in some phages and contrasts with the monophyletic pattern of bacterial recoding. A metabolic process driving the adoptionof an unusual genetic code may explain why the Pyl code is not observed in bacteria, as they are generally metabolically versatile and never methanogens.
[0291] Despite genome-wide TAG recoding in Group C archaea, Pyl is an order of magnitude less abundant in any proteome than the next rarest amino acid. This rarity may be because Pyl is bulky, chemically reactive, and has a ring nitrogen which may introduce a positive charge into proteins at physiological pH. Thus, most mutations that introduce Pyl into organisms with Pyl machinery would be selected against. Two Methermicoccaceae (JDFR-19) genomes have up to four Pyl within a single protein and represent an exception regarding the low TAG frequency among the Group C genomes. Given that it takes time for the proteome to accumulate Pyl in positions where it is either neutral or beneficial, it is suspected that JDFR-19 underwent code transition long ago.In summary, it was found that Archaea, previously not known to use more than a single genetic code, contains clades in which a genome-wide code switch has occurred such that the TAG stop codon is interpreted as pyrrolysine. These clades, comprised of Group C genomes, require a translation table interpreting all TAG codons as coding for Pyl for correct protein prediction. An evolutionary path to code switch is proposed and evidence is provided that this code transition is relatively facile, occurring in unrelated lineages. Finally, the biotechnological value of the pyrrolysine incorporation machinery from recoded genomes was demonstrated.MethodsIdentification and annotation of Pyl- encoding Archaeal genomes
[0292] Genomes encoding Pyl were identified based on the presence of cellular machinery required for production and incorporation of this amino acid. A database of 14257 archaeal genomes from Genbank with more than 60% completeness and less than 5% contamination (as estimated by CheckM was assembled in April 2024. For these genomes, Prodigal v.2.6.3 was used to predict genes using default parameters (single genome mode, translation table 11), and hmmer v.3.3.1 was used via hmmscan against the NCBI-fam database v.13. The HMM gene annotation used to identify the Pyl pathway is as follows: TIGR03912 (PylSn), TIGR02367 (PylSc), TIGR03910 (PylB), TIGR03909 (PylC) and TIGR0391 (PylD), and Pyl-containing methyltransferases (MtmB, MtbB and MttB) using manually curated HMM profiles. For pylT a HMM profile was constructed using the few previously identified pylT genes. Sequences were recovered together with a 10 bp flanking region. They were alignedwith known pylT, checked for the presence of an aligned CTA (CUA) and trimmed to remove 5’ and 3’ extra bases. Secondary structure was checked using RNAFold and an additional trimming was performed, if needed.Taxonomic distribution and evolution of the Pyl machinery and Pyl code in Archaea
[0293] The reference species tree of Archaea was built from a concatenation of the sequences of 36 universally conserved proteins from the phylosift dataset together with sequences of L30 and S4 ribosomal proteins and of the RNA polymerase subunit A and B. Proteins were retrieved from the selected proteomes with HMM search (hmmer v3.3.2 [hmmer(dot)org / ]). They were aligned with MAFFT (v7.453) with the accuracy-oriented methods, L-INS-i, trimmed with BMGE (v2) with the Blosum30 parameter, and concatenated. A maximum likelihood tree was built with IQ-TREE (v2.0.6) using the LG+F+R10 model. This tree was used as a guide tree to run the Posterior Mean-Site Frequency (PMSF) model under the LG + C60 + F + G4 model implemented in IQ-TREE using a 1 ,000-replicate, ultra-fast bootstrap. Genome taxonomy was determined using GTDBTk 2.1.1. Taxonomy, TAG categories and TAG% were mapped on the species tree using the ITOL server.
[0294] PylBCDS proteins from archaea and bacteria were separately aligned using MAFF and trimmed with BMGE using the above mentioned parameters, and then concatenated. A maximum likelihood tree was built with IQ-TREE (v2.0.6) using the LG+R9. Presence / absence of PylSn and its fusion were mapped on the three, together with taxonomy using ITOL server. Horizontal gene transfers were predicted by comparing PylBCDS topology with the reference species tree topology.
[0295] Ancestral presence and emergence of Pyl code in archaea was determined using GLOOME. The input tree of archaea comprises 428 taxa covering the phylogenetic diversity of the domain and includes archaea inferred to use Pyl code. This species tree was built as described above.Quantification of ORF TAG frequency
[0296] To quantify the mean TAG frequency of archaeal taxa, Pyl-encoding genomes were dereplicated for each phylum with dRep v.3.4.1 with default parameters. Then, Prodigal v2.6.3 was used to predict ORFs using default parameters and the -d option used for nucleotide output. The Python script “find_tag_end.py” was used to retrieve all genes with an in-frame TAG codon. The TAG frequency was calculated as the totalnumber of predicted ORFs containing an in-frame TAG codon, divided by the total number of predicted ORFs.Identification and quantification of genes by TAG codon category
[0297] To quantify TAG codon categories, genes with in-frame TAG codons were first extracted. NCBI genome coding sequence (CDS) files were used for gene identification. Prokka V.1.14.5 was used for gene identification for genomes without NCBI CDS files. Each TAG codon was treated as a sense codon, and sequences containing TAG were extended until the nearest in-frame TAA or TGA stop codon. The resulting genes were then translated and annotated using the UniProt- SwissProt database version 2023_05 and the Uniprot Archaeal and Bacterial Reference Proteome databases version 2023_05 (November 8, 2023 release). TAG- containing genes were then grouped into one of the four designated categories, based on the location of the TAG codon: (1) Internal (2) Short extension (3) Long extension / short overlap / fusion and (4) Large out-of-frame overlap. Each gene was assigned to only a single category, using the lowest category number for which criteria were satisfied. Genes that were truncated due to being located at the end of a contig were excluded when extension length could not be determined. Category 1 (Internal) genes contain an internal TAG codon, and encode proteins for which the amino acid sequence before and after the TAG codon has >30% sequence identity with sequences from the UniProt databases. Additionally, for Category 1 , the gene must also encode a protein that is >15 amino acids longer than if TAG was a stop. For Category 2 (short extension), readthrough of TAG results in a protein sequence extension of <15 aa. Sequences overlapping an adjacent gene by >1 sense codon were excluded from Category 2. The Category 3 designation was assigned for genes if TAG readthrough resulted in (a) gene overlap of <60 nt with an adjacent gene, and / or (b) genes encoding a protein with >15 aa extension and / or (c) gene fusion (in-frame overlap) with an adjacent gene. Finally, TAG readthrough in Category 4 genes results in overlap of an adjacent gene by >60 nt, including overlap of an entire gene in a different reading frame or oriented in the opposite direction.Prediction and analysis of Pyl- containing proteins
[0298] The genetic code prediction tool, Codetta, was initially used to identify Pyl proteins. By default, Codetta does not detect the recoding of TAG to Pyl. First, Codetta completes a six-frame translation of a genome nucleotide sequence that is aligned to profileHMMs. The software then recognizes recoding events when a codon has a specific nonstandard interpretation based on profile HMM alignments. For Codetta to complete this task, there must be a primary meaning that can be deduced from the alignments, for example, Codetta detected reassignment of AGG (a canonical arginine codon) to methionine. When TAG is recoded to Pyl, TAG occurs in place of different codons such that there is no consensus for which codon it replaces. Thus, Codetta does not determine a single amino acid (re)interpretation for TAG. Identification of known Pyl proteins may be enabled with Codetta, but Pyl detection in this case is limited exclusively to methylamine methyltransferases because they are the only proteins with Pyl- containing profile HMMs that are available for Codetta. However, the alignment output produced by Codetta includes gene matches (profile HMMs) for each codon, including TAG. Using this Codetta alignment output, Pyl- containing proteins when TAG is internal in a gene in a Pyl code genome were identified. The alignment step was implemented using the Codetta v.2.0 codetta_align.py command with a compatible profile HMM database prepared with the hmmer v.3.1 hmmbuild command (--enone option) for the PGAP v.15 database, which includes TIGRFAM v.15.0 entries.
[0299] A secondary annotation for ORFs was performed by modifying the protein prediction software, Prodigal, to include TAG as a sense codon in the archaeal genetic code. The modified version of Prodigal is available atcan be invoked using the -g 34 (translation table 34) option. Geneious was used to visualize and overlay the standard genetic code annotation and the TAG readthrough annotation. Sequences with inframe TAG codon(s) were searched against Blast to identify protein matches. For gene ontology (GO) enrichment analysis, GOATOOLS v. 1.2.3 was used after protein functional annotation was performed using InterProScan v.5.69. Fisher’s Exact Test was used to determine the statistical significance of enrichment for each GO term in the study set relative to the background. The resulting p-values were corrected for multiple hypothesis testing using the Bonferroni, Sidak, and Holm methods. Searches for PYLIS structures were performed with Infernal v.1.1.4 using covariance model(s) via cmsearch with the RFAM v.14.10 database. Regions of 200 nt flanking TAG- containing genes were examined to determine if matches to the RFAM database were present. Prediction and annotation of Pyl- containing proteins was performed using the Bridges-2 computing resource.Structural prediction of Pyl- containing proteins
[0300] Structures for homologous proteins were obtained from RCSB PDB. Structure 4AY8 was used for MtaA and structure 7UTN was used for IscB. The Visual Molecular Dynamics software package was used for visualization.Anaerobic cultivation
[0301] Methanomethylophilus alvi (formerly M. alvus) strain Mx-05T(JCM 31474T) was grown as previously described, with methanol and ruminal fluid. Methanococcoides burtonii strain ACE-MT(DSM 6242T) was obtained from Deutsche Sammlung fur Mikroorganismen und Zellkulturen (DSMZ) and grown on DSMZ Methanococcoides medium 141c (with 5 mM trimethylamine hydrochloride), with the following modifications: instead of Wolin’s mineral solution and Wolin’s vitamin solution, ATCC trace mineral supplement (MD-TMS) and ATCC Vitamin Supplement (MD-VS) were used. M. burtonii was grown at room temperature in 15 mL Hungate tubes with a headspace of 85% N2, 10% CO2, 5% H2 gas in Anaerobe Systems anaerobic chamber (AS- 150).Sample preparation for LC-MS / MS
[0302] 20 mg cell pellets were resuspended in 100 mM ammonium bicarbonate, pH 8.0 and subjected to mechanical lysis by bead beating with 0.15 mm zirconium oxide beads for five min (Geno / Grinder 2010; SPEX). Resulting cell lysates were adjusted to 4% sodium dodecyl sulfate (SDS) / 10 mM dithiothreitol (DTT), incubated at 90°C for 10 min to denature and reduce proteins, and pre-cleared by centrifugation at 21 ,000 x g. Samples were adjusted to 30 mM iodoacetamide (Sigma) and incubated in the dark at room temperature for 20 min to alkylate cysteine residues. The alkylated proteins were further cleaned and prepared for proteolytic digestion via the protein aggregation capture (PAC) method. In brief, 300 pg of hydrophobic magnetic beads (1 micron, SpeedBead Magnetic Carboxylate; GE Healthcare UK) and acetonitrile (ACN) were added to protein samples for a final concentration of 80% ACN to induce protein aggregation on the beads. Samples were incubated for 20 min at room temperature. The aggregated proteins were placed on a magnet and the supernatant was removed. The pelleted beads were subjected to a wash with 1 mL of ACN, followed by 1 mL of 70% ethanol. The final wash buffer was removed from tubes and the beads were resuspended in 200 pL of 100 mM ammonium bicarbonate for in-solution proteolytic digestion. Protein amounts were quantified by corrected absorbance (Scopes) at205 nm (NanoDrop OneC; Thermo Fisher). After protein quantitation, samples were digested using either MS-grade GluC Endoproteinase or trypsin (Thermo Scientific) with a 1 :75 protease: protein (wt:wt) ratio at 37°C, shaking overnight at 600 rpm. An additional round of protease was added for a second 3-h digestion period at 37°C, shaking at 600 rpm. The resulting tryptic peptides were filtered on 10 kDa MWCO filter plate (AcroPrep Advance, Omega 10 K MWCO) at 1 ,500 g and adjusted to 1% formic acid before quantification by NanoDrop OneC.LC-MS / MS
[0303] 3 pg of peptide solution were analyzed by automated 1 D LC-MS / MS analysis using a Vanquish ultra-HPLC (uHPLC) system plumbed directly in-line with a QExactive- Plus mass spectrometer (Thermo Scientific) with an in-house-pulled 100 pm inner diameter nanospray emitter (packed to 15 cm with 1.7 pm Kinetex C18 reverse-phase resin (Phenomenex)). Peptides were loaded, desalted, and separated by uHPLC under the following conditions: sample injection followed by 100% solvent A (98% H2O, 2% acetonitrile, 0.1% formic acid) from 0 to 30 min to load and desalt, a linear gradient from 0 to 30% solvent B (70% acetonitrile, 30% water, 0.1 % formic acid) from 30 to 220 min for separation, a column wash from 0 to 100% solvent B from 220 to 255 min, and 100% solvent A from 255 to 275 min for column re-equilibration. Eluting peptides were analyzed with the following MS settings for global searches: positive mode, data-dependent acquisition, top-20 method; mass range 400-1500 m / z; MS and MS / MS resolution 70 and 17.5 K, respectively; MS / MS loop count 20; isolation window 1.8 m / z; isolation offset 0.3 m / z; charge state exclusion of unassigned, +1 , +6-8 charges.
[0304] For targeted searches, parallel reaction monitoring inclusions lists were constructed to detect peptides of interest in the predicted pyrrolysine containing proteins. An in-silico digest of the proteins was conducted with UniPept using either GluC or trypsin as the protease with a minimum peptide length of six amino acids. Based on the predicted peptides from the in silica digest, peptides were selected for the PRM measurements if they were classified as any of the following peptide types: (1) the first proteolytic peptide in the protein, (2) the proteolytic peptide containing the pyrrolysine residue,(3) the first proteolytic peptide downstream of the pyrrolysine-containing peptide, and(4) the last proteolytic peptide in the sequence. For the peptides of interest, the PRM inclusion list included the monoisotopic precursor m / z of the peptide including a static modification of carbamidomethylation (+57.021) of cysteine and a charge state (z) of+2. No variable modifications of oxidation (+15.995) of methionine, deamidation of (+0.984) of asparagine and glutamine, or methyl loss (-14.016) for pyrrolysine were considered for the inclusion list. A maximum of 50 entries were monitored for each PRM measurement. The same LC-MS / MS method was used as the global measurements.Proteomics data analysis
[0305] Global database searches were conducted with Proteome Discoverer 3.0 (Comet / Percolator)against custom organism specific protein databases constructed from all predicted proteins in the respective, as well as common mass spectrometry contaminants. Protein sequences were included with the predicted pyrrolysine residue represented with “X” as well as the sequence version that terminated in a stop codon. For all database searches, the peptide mass tolerance was set at 20 ppm. Accepted modifications included static modifications of carbamidomethylation (+57.021) of cysteine residues and pyrrolysine (+237.148) of “X” residues. Database searches were performed with variable modifications of oxidation (+15.995) of methionine, deamidation of (+0.984) of asparagine and glutamine residues, and a methyl loss (-14.016) for pyrrolysine (X) residues, and a maximum of three variable modifications allowed per peptide. Semi-proteolytic digestion of peptides was allowed for either trypsin or GluC depending on the sample, and a 1% peptide-level FDR threshold was applied.
[0306] Preliminary confirmation of pyrrolysine expression was based on (1) a unique peptide match in the global datasets that passed the 1 % peptide-level FDR threshold for the pyrrolysine containing proteins or (2) a detected precursor mass in the PRM measurement that was within 10ppm of the monoisotopic precursor m / z of the peptides of interest. Peptides were considered unique if the amino acid sequence was not predicted in any other protein for the respective organism. To more confidently confirm pyrrolysine expression, MS / MS spectra from all peptides of pyrrolysine- containing proteins were manually inspected and validated. A MS / MS spectrum of a peptide was considered for evidence of pyrrolysine expression if each of following criteria was met: (1) Precursor ion (MS1) mass error <10 ppm, (2) Fragmentation ion (MS / MS) mass error <0.02 Da and (3) Location of the fragment ion along the peptide sequence, with either direct sequencing of the fragment ending in pyrrolysine or fragment ions that were longer than the pyrrolysine-containing fragment.
[0307] In addition to the database searches, the data were searched using a de novo- assisted database search workflow using PEAKS StudioX to recognize any of the standard amino acids at the TAG=pyrrolysine / stop position (represented as “X” in the database). This workflow will recognize the 20 standard amino acids if they were detected at the Pyl position in the MS / MS spectrum. The same modifications as previous database searches were included and the parent and fragment ion mass error tolerances were set to ± 10 ppm and ±0.02 Da, respectively. De novo only parameters were left at default settings with average local confidence (ALC) scores of >80% and de novo sequence tags displayed if at least six amino acids were shared with the database sequence.Construction of plasmids encoding PylRS / tRNAPylpairs
[0308] Plasmids encoding the PylRS / tRNAPylpairs evaluated in FIG. 26 were derivatives of the previously reported pMEGA vector (parent backbone Addgene Plasmid #200225). pMEGA plasmids containing each PylRS homolog were ordered as circular, dsDNA plasmids from Twist Bioscience. Each one was then modified to eliminate the M. alvi pylT gene and replace it with the one corresponding to each PylRS homolog. Briefly, each circular pMEGA vector obtained from Twist was linearized using oligonucleotides P1 and P2. PCR using the Q5® High-Fidelity 2X Master Mix (NEB M0492S) was performed for 30 cycles according to manufacturer protocols. The PCR reaction was subject to a PCR clean up with the QIAquick PCR Purification Kit (Qiagen, catalog # 28104). The concentration of purified, linearized products was determined by absorbance at 260 nm using a NanoDrop ND-1000 Spectrophotometer. Next, 33.3 ng of each purified, linearized pMEGA plasmid was combined with 100 ng of a gBIock containing the appropriate pylT sequence in a 20 pL Gibson Assembly reaction using HiFi DNA Assembly Master Mix (NEB, catalog #E2621 L). The reaction mixture was incubated at 50°C for 1 h to generate a family of circular pMEGA vectors containing the coding sequence for the appropriate pylT.
[0309] The circularized pMega plasmids from the previous step were then transformed into NEB 5-alpha competent E. coli (NEB, catalog # C2987H) as follows. Frozen stocks of cells were thawed on ice for 10 min. Upon thawing, the entirety of the previous Gibson Assembly reaction was added to cells and incubated on ice for 30 min. Cells incubated with plasmid were then subjected to heat shock at 42°C for 30 sec and placed on ice for 2 min. 350 pL of Super Optimal broth with Catabolite repression (S.O.C.) outgrowth medium (NEB, catalog # B9020S) was added to cells and cellswere incubated at 37°C for 1 h with shaking at 220 rpm. Agar plates containing spectinomycin were inoculated with 50 pL of transformed cells and grown overnight at 37°C. 3 single colonies per construct were picked and inoculated into liquid cultures containing 5 mL LB + spectinomycin and grown for 16 h at 37°C. Pure plasmid was isolated from 5 mL cultures using Qiaprep Spin Miniprep Kit (Qiagen, catalog # 27106) and sequences were confirmed by whole plasmid sequencing with Primordium Labs. The new pMEGA plasmids were double transformed with a pET22b-3TAG-GFP into BL21 (DE3) cells following the transformation protocol detailed below.Transformation protocol
[0310] BL21(DE3) E. coll (NEB, catalog # C2987H and C2527H, respectively) were transformed in accordance with manufacturer protocols with some modifications as follows. Frozen stocks were thawed on ice. Upon thawing, 100 ng of each relevant plasmid was added. After a 30 min incubation on ice, cells were heat-shocked for 30 sec at 42°C and allowed to recover on ice for 2 min. Following, 350 pL of Super Optimal broth with Catabolite repression (S.O.C.) was immediately added and cells were allowed to recover at 37°C with shaking for 1 h before plating 50 pL on LB agar plates containing the appropriate antibiotics. Plates were incubated overnight at 37°C.Expression assays of sfGFP-3TAG containing Boc-Lys or HO-Boc-Lys
[0311] Starter E. coli cultures were grown overnight in 5 mL of LB Miller (AmericanBio, Catalog #AB01201) in 15 mL culture tubes supplemented with carbomycin and spectinomycin at 37°C. The next morning, 500 pL of the saturated overnight culture was added to 50 mL LB in a 250 mL baffled flask with carbomycin and spectinomycin and grown at 37°C to an OD6oo of 0.6. Prior to cultures reaching QD600 0.6, the appropriate amount of each monomer (stored as 25 mM stocks) were allotted into the wells of a black, clear bottom, 96-well plate (1 mM final concentration). Once cultures reached an OD6oo of 0.6 (roughly 3 h), protein expression was induced by addition of 1 mM IPTG. The culture was transferred into the wells of the plate to a total volume of 200 pL in each well. A Breathe-Easy® sealing membrane (Sigma Z380059) was placed over the 96-well plate and the plate was loaded into a BioTek Synergy HTX microplate reader with no lid. OD6oo and F528 values (Aex= 485 nm) were measured every 10 min for 20 h. The plate was maintained at 37°C and was shaken during this time. Points represented in FIG. 26B represent three biological replicates where three random colonies were picked on a plate from a single transformation.Expression and Purification sfGFP-3TAG containing Boc-Lys or HO-Boc-Lys
[0312] Starter cultures of 5 mL of Miller’s LB Broth (AmericanBio, catalog # AB01201) supplemented with carbomycin and spectinomycin were inoculated with a single colony of BL21 (DE3) E. coli cells harboring pET22b-sfGFP-3TAG-6xHis alongside pMEGA-JDFR19-PylRS-pylT, pMEGA-JDFR19-PylRSCTD-pylT, pMEGA-B74G9- PylRSCTD-pylT, pMEGA-LMO1-PylRS-pylT, pMEGA-Bin14-PylRS-PylRS-pylT, pMEGA-1 D-PylRS-pylT, pMEGA-Levihalophilus-PyIRS-pyIT, or pMEGA-MaPyIRS and grown overnight for 16 h at 37°C with shaking at 220 rpm until the culture was saturated. The starter culture was used to inoculate a 100 mL expression culture in a 1 :100 dilution of Miller’s LB Broth supplemented with 1 mM Boc-Lys (Combi-Blocks, CAS # 2418-95-3) or 1 mM HO-Boc-Lys (Accela, CAS # 111223-31-5) with carbomycin and spectinomycin. The expression culture was grown at 37°C with shaking at 220 rpm to an OD6oo of 0.6 at which point it was induced with 1 mM IPTG and grown for 16 h under the same conditions. The expression culture was harvested by centrifugation at 4,300 x g at 4°C for 30 min. The resulting cell pellet was suspended in 10 mL of Lysis Buffer (50 mM sodium phosphate pH 6.8 and 300 mM NaCI) containing half a tablet of complete, mini EDTA-free ULTRA protease inhibitor cocktail (Sigma-Aldrich, St. Louis, MO). The cell suspension was disrupted by sonication on ice (Branson Sonifier 250, 5 cycles of 30 sec pulse at 50% duty cycle and microtip limit of 5 followed by 30 sec pause). The cell lysate was cleared by centrifugation at 10,000xg at 4°C for 20 min. TALON® Metal Affinity Resin (1 mL) (Takara Biosciences, catalog # 635504) was equilibrated with Lysis Buffer, added to the cleared cell lysate, and incubated on a rotisserie at 4°C for 1 h. The TALON® resin-lysate mixture was then passed through a gravity flow Poly-Prep Chromatography Column (Bio-Rad Laboratories, Hercules, CA). Non-specifically bound proteins were removed by washing the TALON® resin with 10 mL of Lysis Buffer. The 6xHis-tagged protein was eluted by washing the TALON® resin with 2 mL of Elution Buffer (50 mM sodium phosphate pH 6.8 and 250 mM imidazole). The purified protein was quantified using absorbance at 280 nm, snap frozen as singleuse aliquots, and stored at -80°C.Intact Protein LC-MS for GCE
[0313] LC-MS analysis of all protein samples were performed on an Agilent 1290 Infinity II HPLC connected to an Agilent 6530B QTOF AJS-ESI. The mobile phase for LC-MSwas water and acetonitrile with 0.1% (v / v) formic acid at a flow rate of 0.4 mL / min. Each protein sample was injected onto a Poroshell 300SB-C8 column (2.1 x 75 mm, 5 pM, room temp, Agilent) and separated using a linear gradient from 5% acetonitrile for 0 to 2 min and ramping to 95% acetonitrile over 7.5 min, and then washing with 95% acetonitrile for 2 min. The following parameters were used during acquisition: Fragmentor voltage 225 V, gas temperature 300 °C, drying gas flow 10 L / min, sheath gas temperature 350 °C, sheath gas flow 11 L / min, nebulizer pressure 35 psi, skimmer voltage 65 V, Vcap 5000 V, 1 spectra / s.
[0314] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparent to those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
[0315] Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e. , any elements developed that perform the same function, regardless of structure. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.
[0316] The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of present invention is embodied by the appended claims. In the claims, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) is expressly defined as being invoked for a limitation in the claim only when the exact phrase "means for" or the exact phrase "step for" is recited at the beginning of such limitation in the claim; if such exact phrase is not usedin a limitation in the claim, then 35 U.S.C. § 112 (f) or 35 U.S.C. §112(6) is not invoked.
Claims
CLAIMSWhat is claimed is:
1. A method of producing a protein of interest (POI) that comprises one or more non-natural monomer (NNMs), the method comprising: providing to a bacterial cell, a eukaryotic cell, or an in vitro translation system:(i) a pyrrolysyl-tRNA synthetase (PylRS), wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10,(ii) a NNM or a salt thereof,(iii) a tRNA that can be acylated with the NNM by the PylRS, and(iv) a DNA or mRNA comprising a nucleotide sequence encoding a POI, wherein the nucleotide sequence encoding the POI comprises one or more in-frame codons that are recognized by the tRNA’s anticodon; thereby producing a NNM-containing POI.
2. The method of claim 1 , wherein the PylRS is provided to the bacterial cell, the eukaryotic, or the in vitro translation system as a nucleic acid encoding the PylRS.
3. The method of claim 1 or claim 2, wherein the tRNA is provided to the bacterial cell, the eukaryotic, or the in vitro translation system as a nucleic encoding the tRNA.
4. The method of any one of claims 1-3, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1.
5. The method of any one of claims 1-4, wherein the tRNA’s anticodon recognizes the codon UAG.
6. The method of any one of claims 1-5, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.
7. The method of any one of claims 1-5, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.
8. The method of any one of claims 1-7, wherein the NNM is any one of the compounds depicted in Figures 15-22, or is a-hydroxy acid,2-hydroxy acid, or |33-amino acid analog thereof.
9. The method of any one of claims 1-7, wherein the NNM is any one of the lysine derivatives depicted in Figures 15-22, or is a-hydroxy acid,2-hydroxy acid, or3-amino acid analog thereof.
10. The method of any one of claims 1-7, wherein the NNM is any one of the compounds depicted in Figure 15.
11. The method of any one of claims 1-7, wherein the NNM is Pyl, Boc-Lys, or HO-Boc-Lys.
12. The method of any one of claims 1-11, wherein the cell is a mammalian cell.
13. The method of any one of claims 1-11, wherein the cell is an E. coli cell.
14. The method of any one of claims 1-13, further comprising reacting the NNM-containing POI with a conjugation partner molecule such that the conjugation partner molecule binds covalently to at least one NNM of the NNM-containing POI.
15. A system comprising: a nucleic acid comprising a nucleotide sequence encoding a pyrrolysyl-tRNA synthetase (PylRS), wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, wherein the nucleotide sequence encoding the PylRS is:(a) operably linked to a bacterial or eukaryotic promoter, and / or(b) codon-optimized for expression in a bacterial or eukaryotic host cell.
16. The system of claim 15, wherein the nucleic acid is a plasmid or a viral vector.
17. The system of claim 15 or claim 16, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1.
18. The system of any one of claims 15-17, further comprising a tRNA or a nuclei acid comprising a nucleotide sequence encoding said tRNA, wherein the tRNA can be acylated by the PylRS.
19. The system of claim 18, wherein the tRNA’s anticodon recognizes the codon UAG.
20. The system of claim 18 or claim 19, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.
21. The system of claim 18 or claim 19, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.
22. The system of any one of claims 18-21 , wherein the nucleotide sequence encoding the tRNA is operably linked to a bacterial or eukaryotic promoter.
23. The system of any one of claims 18-22, wherein the nucleotide sequence encoding the tRNA is present on the same nucleic acid as the nucleotide sequence encoding the PylRS.
24. The system of any one of claims 18-22, wherein the nucleotide sequence encoding the tRNA and the nucleotide sequence encoding the PylRS are present on different nucleic acids.
25. The system of any one of claims 18-24, further comprising one or more non-natural monomers (NNMs), wherein said tRNA can be acylated with said one or more NNMs by said PylRS.
26. The system of claim 25, wherein the NNM is any one of the compounds depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or p3-amino acid analog thereof.
27. The system of claim 25, wherein the NNM is any one of the lysine derivatives depicted in Figures 15-22, or is a-hydroxy acid, p2-hydroxy acid, or [33-amino acid analog thereof.
28. The system of claim 25, wherein the NNM is any one of the compounds depicted in Figure 15.
29. The system of claim 25, wherein the NNM is Pyl, Boc-Lys, or HO-Boc-Lys.
30. A bacterial or eukaryotic cell comprising:(c) a pyrrolysyl-tRNA synthetase (PylRS), or a nuclei acid comprising a nucleotide sequence encoding the PylRS, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 1-10, and(d) a tRNA that can be acylated by the PylRS, or a nuclei acid comprising a nucleotide sequence encoding the tRNA.31 . The bacterial or eukaryotic cell of claim 30, wherein the cell is a mammalian cell.
32. The bacterial or eukaryotic cell of claim 30, wherein the cell is an E. coli cell.
33. The bacterial or eukaryotic cell of any one of claims 30-32, wherein the PylRS comprises an amino acid sequence that is 80% or more identical to any one of SEQ ID NOs: 10, 6, and 1.
34. The bacterial or eukaryotic cell of any one of claims 30-33, wherein the tRNA’s anticodon recognizes the codon UAG.
35. The bacterial or eukaryotic cell of any one of claims 30-34, wherein the tRNA comprises a nucleotide sequence that is 80% or more identical to any one of SEQ ID NOs: 11-18.
36. The bacterial or eukaryotic cell of any one of claims 30-34, wherein the tRNA comprises a nucleotides sequence that is 80% or more identical to any one of SEQ ID NOs: 18, 14, and 11.
37. The bacterial or eukaryotic cell of any one of claims 30-36, wherein the nucleotide sequence encoding the tRNA is present on the same nucleic acid as the nucleotide sequence encoding the PylRS.
Citation Information
Patent Citations
Cell lines
US20170306381A1
A facile system for encoding unnatural amino acids in mammalian cells
WO2010114615A2
Cell lines
WO2014044872A1