Tunable tetrazine amino acids for encodable protein labeling
Tunable tetrazine amino acids with orthogonal encoding systems enable rapid and stable bioorthogonal reactions, addressing the limitations of existing systems by achieving high reaction rates and maintaining protein stability for efficient protein labeling.
Patent Information
- Application Number
- PCT/US2025/014193
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
Existing bioorthogonal ligation systems face challenges in achieving high reaction rates and maintaining protein stability due to structural alterations of functional groups, particularly with tetrazine-alkene cycloaddition reactions, which compromise protein reactivity and stability when reaction rates exceed certain thresholds.
Development of tunable tetrazine amino acids (Tet ncAAs) with specific substituents and orthogonal tRNA/synthetase pairs to encode proteins, enabling rapid and stable bioorthogonal reactions through [4+2] cycloaddition, allowing for site-specific labeling and modification.
The Tet ncAAs achieve exceptionally high reaction rates (up to 104 M^-1 s^-1) while maintaining protein stability, facilitating efficient and selective labeling of proteins with diverse molecules in live cells and media.
Smart Images

Figure US2025014193_07082025_PF_FP_ABST
Abstract
Description
[0001] TUNABLE TETRAZINE AMINO ACIDS FOR ENCODABLE PROTEIN LABELING
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of Application No. 63 / 627687, filed January 31 , 2024, Application No. 63 / 627674, filed January 31, 2024, Application No. 63 / 627665, filed January 31, 2024, and Application No. 63 / 627625, filed January 31, 2024, the disclosure of each of which is hereby expressly incorporated by reference in its entirety.
[0004] STATEMENT OF GOVERNMENT LICENSE RIGHTS
[0005] This invention was made with government support under RM1-GM144227 and 1R01GM131168-01 awarded by National Institutes of Health and NSF-2054824 awarded by National Science Foundation. Ute government has certain rights in the invention.
[0006] STATEMENT REGARDING SEQUENCE LISTING
[0007] The Sequence Listing XML associated with this application is provided in XML format and is hereby incorporated by reference into the specification. The name of the XML file containing the sequence listing is 3014-P37WO Sequence Listing. xml. The XML file is 31,552 bytes; was created on January 31, 2025; and is being submitted electronically via Patent Center with the filing of the specification.
[0008] BACKGROUND
[0009] Encodable bioorthogonal ligations have become an indispensable and powerful tool for studying living process in cells because one of the two chemical needed functional groups can be encoded into specific locations on biopolymers and the second functionality can be used to deliver diverse small molecule probes, materials, or cellular components. Encodable bioorthogonal ligation systems are challenging systems to develop because three key comprising elements; the bioorthogonal chemistry, the functional group encoding and the delivery of the label into the cell, are interconnected and each need to be optimized for the labeling tool to work effectively. For this reason, many catalyst-free bioorthogonal reactions (Staudinger ligation, strain promoted azide-alkyne cycloaddition, oxime ligation, photo induced tetrazole- alkene ligation, and strain promoted tetrazine-alkene cycloaddition) have been explore for ability in encodable bioorthogonal ligations. Bioorthogonal ligations that used small functional groups were critical in launching the field because the traditional functional group encoding step relied on the natural cellular enzymes which could tolerate only small deviations in structure from natural substrates. Include a sentence here about tetrazine encoding.
[0010] Genetic code expansion (GCE) can now be engineered to encode a wide variety of bioorthogonal functional groups. In GCE, a reactive functional group is converted into a noncanonical amino acid (ncAA) which then enters protein translation using an evolved tRNA / RS pair specific for the ncAA. Using this approach, many bioorthogonal functional groups components have been encoded site-specifically into proteins and evaluated for labeling in cell and on cell surfaces. The extensive work on developing and using labeling reactions in live cells make clear that high biomolecular rate constants are critical to advance the ability to monitor biological process in live cells, in real time. Exceptionally high reaction rates are the only way to access the uncatalyzed labeling of low intracellular protein concentrations with low label concentrations in short time frames. Unfortunately, very few encodable bioorthogonal functional groups can be structurally altered to increase bioorthogonal reaction rates.
[0011] Recently, the 1,2,4,5-tetrazine-alkene cycloaddition reaction (tetrazine reaction), also known as the strain-promoted inverse electron demand Diels-Alder (IEDDA) reaction has been explored extensively because of its extraordinarily fast bioorthogonal reaction rates. In addition to exceptionally high reaction rates for biorthogonal reactions, the tetrazine reaction also has an usually broad functional reaction rate range (1 to IO6M ’s1). The broad reaction rate range results from the electron donating or withdrawing substituent effects on the tetrazine diene and the variety of ring strain imposed on the alkene dienophile. There have been numerous studies that explore the chemical range of functional groups, theoretical limits, and their use in bioorthogonal applications. Thus far many ncAA comprising of the ring strained alkene dienophile have been encoded into proteins to determine the reaction limits of genetically encoded tetrazine reactions but very few tetrazine ncAA structures have been explored. The methanogenic archaea derived pyrrolysyl-tRNA / tRNA-synthetase systems (PylRS / tRNAPyl) which can function efficiently as an orthogonal GCE tool across most domains of life has been used extensively to encode strained alkenes attached as a lysine derivative. The ease at which diverse large hydrophobic rings can be encoded as lysine derivatives with the PylRS / tRNA1^1has allowed on protein and in cell reactivity and stability assessment of strained alkenes having different ring size, attachment chemistry, and levels of ring size. Obtaining reaction rates with 103M ’s1is possible by encoding strain alkenes without losing significant amino acid stability. Unfortunately, once reaction rates are pushed above 104M1s1using the more reactive transcyclooctene (TCO) ncAAs a significant portion of the ncAA protein is no longer reactive to Tet-labels due to isomerization of the trans bond from cellular conditions.
[0012] As an alternative approach to maximize encoded bioorthogonal reaction rates, encoding Tet ncAAs and reacting with sTCO labels has been reported. The synthesis and encoding of first-generation Tet ncAAs using an Mj tRNA synthetase MjRS / tRN AMjpair was demonstrated in 2012. A required part of the synthesis the amine linkage in first- generation Tet ncAAs resulted in a slow tetrazine reaction (102M-1s-1) and compromised the stability of the final labeled protein due to low levels of amine elimination. Newer synthetic routes to make tetrazines provided second- and third-generation Tet ncAAs, which were encoded with the Mj-RS / tRNAM|and PylRS / tRNAPylpairs. These newer Tet ncAA-containing proteins were able to achieve reaction rates of 2 to 8 xlO4M-1s-1in live E. coli and mammalian cells while maintaining complete reactivity and product stability.
[0013] Despite the advancement in the development of Tet ncAAs for encoding into proteins for rapid bioorthogonal labeling, a need exists for improved Tet ncAAs. The present disclosure seeks to fulfill this need and provide further related advantages. SUMMARY
[0014] In one aspect, the disclosure provides a tetrazine amino acid having formula (1) or (II): or a stereoisomer or salt thereof, wherein
[0015] R is selected from the group consisting of:
[0016] (a) a phenyl group substituted with a group selected from O-C1-C4 alkyl, m- C1-C4 alkyl, O-C1-C3 alkoxy, m-Cl-C3 alkoxy, o-cyano, m-cyano, o-nitro, m-nitro, o- or m-primary amino (-NH2), secondary amino (-NHRX), or tertiary amino (-NHRxRy) (wherein Rxand Ryare independently C1-C6 alkyl), o-fluoro, m- fluoro, 3,5-difluoro, 3,4,5- trifluoro, o-trifluoromethyl, m-trifluoromethyl, and o-, m-, or p-C(=O)Rz(wherein Rza counterion, hydrogen, or C1-C6 alkyl),
[0017] (b) a substituted or an unsubstituted heteroaryl group,
[0018] (c) a substituted or an unsubstituted heterocyclyl group,
[0019] (d) an amino Cl -C6 alkyl group,
[0020] (e) a thio C1-C6 alkyl group,
[0021] (f) a carboxylate group,
[0022] (g) a sulfonate group, and
[0023] (h) an amide group;
[0024] Rcis hydrogen, a counter ion, or a carboxyl protecting group; and
[0025] RNis hydrogen or an amine protecting group. In certain embodiments, the alkyl group is selected from methyl, ethyl, n-propyl, i- propyl, n-butyl, s-butyl, t-butyl, n-pentyl, and n-hexyl.
[0026] In certain embodiments, the phenyl group is substituted with a group selected from the group consisting of 3-methyl, 2-methoxy, 3-methoxy, 2-cyano, 3-cyano, 2-nitro, 3- nitro, 2-fluoro, 3-fluoro, 3,5-difluoro, 3,4,5-trifluoro, 2-hydroxy, 3-hydroxy, 2-amino, 3- amino, 2-acetyl, 3-acetyl, 2-trifluoromethyl, and 3-trifluoromethyl.
[0027] In certain embodiments, the heteroaryl group is a pyridyl group, a furanyl group, a thiophenyl group, or an oxazolyl group.
[0028] In certain embodiments, the heterocyclyl group is a pyrrolidinyl group, a piperadinyl group, a piperazinyl group, a tetrahydrofuranyl group, a tetrahydropyranyl, a tetrahydrothiophenyl group, an oxazolydinyl group, or a dihydropyran group.
[0029] In certain embodiments, the amino C1-C6 alkyl group is an aminobutyl group.
[0030] In certain embodiments, the thio C1-C6 alkyl group is a thiomethyl group.
[0031] In certain embodiments, the carboxylate group is a C1-C6 alkyl carboxylate group (e.g., methyl carboxylate (-CO2CH3), ethyl carboxylate (-CO2CH2CH3)).
[0032] In certain embodiments, the amide group is a C1-C6 alkyl amide group (e.g., methyl amide (-C(=O)NHCH3), butyl amide (-C(=O)NH(CH2)3CH3)).
[0033] In another aspect, the disclosure provides a method for making a protein or a polypeptide of interest, comprising incorporating a tetrazine amino acid or a stereoisomer or salt thereof, as described herein, into a protein or polypeptide.
[0034] In a related aspect, the disclosure provides a method for genetically encoding a protein or a polypeptide of interest, comprising incorporating a tetrazine amino acid or a stereoisomer or salt thereof, as described herein, into a protein or polypeptide by genetic encoding.
[0035] In a further aspect, the disclosure provides a protein or polypeptide, comprising at least one tetrazine amino acid residue, wherein the tetrazine amino acid residue is derived from a tetrazine amino acid or a stereoisomer or salt thereof, as described herein. In a related aspect, the disclosure provides a protein or polypeptide, comprising at least one tetrazine amino acid residue, wherein the tetrazine amino acid residue is incorporated into the protein or polypeptide by genetic encoding of the protein or polypeptide using a tetrazine amino acid or a stereoisomer or salt thereof, as described herein.
[0036] In another aspect, the disclosure provides a composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from a tetrazine amino acid or a stereoisomer or salt thereof, as described herein, and wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid comprising the first reactive group.
[0037] In a related aspect, the disclosure provides a composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from genetic encoding of the protein or polypeptide using a tetrazine amino acid or a stereoisomer or salt thereof, as described herein, and wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid comprising the first reactive group.
[0038] In a further aspect, the disclosure provides a kit for in cellulo production of a tetrazine-labeled protein or a tetrazine-labeled polypeptide, comprising (a) a tRNA; (b) an aminoacyl-tRNA synthetase; and (c) a tetrazine amino acid or a stereoisomer or salt thereof, as described herein, wherein the tRNA and aminoacyl-tRNA synthetase are an orthogonal tRNA / orthogonal aminoacyl-tRNA pair effective for incorporating the compound into a protein or polypeptide to provide a tetrazine-labeled protein. DESCRIPTION OF THE DRAWINGS
[0039] The foregoing aspects and many of the attendant advantages of this disclosure will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings.
[0040] FIG. 1 illustrates site-specific protein labeling with representative tetrazine amino acid (Tet ncAA) derivatives. Rapid and quantitative protein labeling using genetically encoded tetrazine amino acids Tet-v2.0, Tet-v3.0 and Tet-v4.0 with sTCO reagents and tuning their size, reactivity and stability by attaching various electron donor and acceptor substituents R.
[0041] FIGS. 2A and 2B illustrate representative tetrazine amino acid derivatives (Tet-v2.0 and Tet-v3.0 ncAAa). FIG. 2A is a schematic illustration of the preparation of representative tetrazine ncAAs. FIG. 2B illustrates representative tetrazine ncAAs described herein (synthetic yields in parenthesis).
[0042] FIG. 3A and 3B illustrate genetic incorporation of representative Tet-ncAAs. Permissivity and efficiency of evolved synthetases measured by florescence of expressed GFP-TAG150 in presence and absence of (A) 0.5 mM Tet-v2.0 derivatives using D12-Tet- v2.0 synthetase (FIG. 3A). 0.5 mM Tet-v3.0 amino acids using R284-Tet-v3.0 synthetase (FIG. 3B). Asterisks (*) represents tetrazine amino acid used to evolve orthogonal aaRS / tRNACUApair.
[0043] FIG. 4 compares Tet ncAA reaction rates with sTCO inside protein: increasing order of second-order rate constants of encoded Tet-v2.0, Tet-v3.0 and Tet-v4.0 derivatives into GFP-150 with sTCO in PBS (pH 7.1) at room temperature.
[0044] FIG. 5 illustrates the SDS-PAGE mobility shift assay: labeling efficiency of GFP- Tet-v3.0 (Ph, Pyr, 4-F-Ph, 4-NH2-Ph) verified by SDS-PAGE mobility shift upon reaction with sTCO-PEG5k.
[0045] FIGS. 6A and 6B verify stability and labeling efficiency of Tet-v3.0 on protein for long time incubation in PBS (pH 7.1) at 4 °C and room temperature (RT). SDS-PAGE mobility shift assay by reacting sTCO-PEG5k, verified stability and reaction ability of encoded Tet-v3.0 (Ph, and Pyr) stored at (A) 4 °C and (B) RT for eight days in presence of 100 mM imidazole.
[0046] FIGS. 7A-7D compare ESI-Q-TOF mass spectrometry analyzed incorporation efficiency of representative Tet-v3.0Ph derivatives and amplitude of labeling reaction with sTCO. ESI mass spectrometry analysis of GFP-Tet-v3.0Ph / 4-F-Ph / 4-NH2-Ph / Pyr and reactions with sTCO-OH. The purified GFP-Tet3.0-ncAAs (black) and reacted with 5-fold molar excess of sTCO-OH for 10 minutes (grey) in PBS (pH 7.1). The reacted GFP-Tet- v3.0 proteins showed as expected 124 Da increase in mass corresponding to the addition of sTCO-OH and loss of molecular nitrogen. No unreacted GFP-Tet-v3.0Ph / 4-F-Ph / Pyr was detected, verifying the reactivity of encoded Tet-v3.0Ph / 4-F-Ph / Pyr with sTCO-OH was quantitative. The lower mass peak labeled with * is a loss of n-terminal methionine and upper mass peaks are salt sodium and potassium adducts. The unreacted lower peak at 27841 Da avg. observed due to near-cognate suppression of amber codon. Near-cognate suppression predominantly observed for GFP-Tet-v3.0_4-NH2-Ph due to lower incorporation efficiency of working synthetase R284. Cal. Mass of GFP-wt: 27827.02 Da avg; (FIG. 7A) GFP-Tet-v3.0Ph observed: 28017.28 Da avg, (expected: 28016.17 Da avg); GFP-Tet-v3.0Ph + sTCO-OH observed: 28141.20 Da avg, (expected: 28140.18 Da avg). (FIG. 7B) GFP-Tet-v3.0_4-F-Ph observed: 28035.31 Da avg, (expected: 28034.16 Da avg); GFP-Tet-v3.0_4-F-Ph + sTCO-OH observed: 28159.29 Da avg, (expected: 28158.17 Da avg). (FIG. 7C) GFP-Tet-v3.0_4-NH2-Ph observed: 27841.1 Da avg (near-cognate suppression); 28030.75 Da avg, (expected: 28031.18 Da avg); GFP-Tet-v3.0_4-NH2-Ph + sTCO-OH observed: 27841.1 Da avg; 28153.64 Da avg, (expected: 28155.19 Da avg). (FIG. 7D) GFP150-Tet-v3.0Pyr observed 28017.66 Da avg (expected: 28017.1 Da avg). GFP-Tet-v3.0Pyr + sTCO-OH observed: 28141.8 Da avg. (expected: 28141.1 Da avg.)
[0047] FIGS. 8 A and 8B verify stability and reactivity of Tet-v3.0Ph and Tet-v3.0Pyr inside protein by ESI-Q-TOF mass spectrometry after 8 days incubation at room temperature. ESI mass spectrometry analysis of pure GFP-Tet-v3.0Ph / Pyr in absence and presence of sTCO after 8 days incubation in PBS (pH 7. 1) at room temperature under basic conditions (100 mM imidazole was added). Purified GFP-Tet3.0-ncAAs (black) and upon reaction with 5-fold molar excess of sTCO-OH for 10 minutes (grey). The GFP-Tet-v3.0Ph showed quantitative labeling by the addition of sTCO-OH. No unreacted GFP-Tet-v3.0Ph was detected, that verified the genetically encoded Tet-v3.0Ph is stable enough for long time incubation. No reaction was observed for GFP-Tet-v3.0Pyr. The lower mass peak labeled with * is a loss of n-terminal methionine and upper mass peaks are salt sodium and potassium adducts. The lower peak at 27842.2 Da avg. observed due to near-cognate suppression of amber codon. (FIG. 8A) GFP-Tet-v3.0Ph observed: 28017.21 Da avg, (expected: 28016.17 Da avg); GFP-Tet-v3.0Ph + sTCO-OH observed: 28142.74 Da avg, (expected: 28140.18 Da avg). (FIG. 8B) In presence and absence of sTCO it shows an unreactive single major peak at 28006.5 Da avg. which is 11 Da avg. unit lower than expected. (GFP-Tet-v3.0Pyr expected: 28017.1 Da avg, and GFP-Tet-v3.0Pyr+sTCO expected: 28141.1 Da avg). The MS result indicated that the GFP-Tet-v3.0Pyr degraded under the following conditions and converted to its oxadiazole derivative which is 12 Da lower molecular mass than Tet-v3.0Pyr.
[0048] FIGS. 9A-9D compare measured reaction rates for Tet-v2.0 and Tet-v3.0 derivatives: Plot of pseudo first order rate constant (k ) against concentration of sTCO to determine the second order rate constant (fa) for reaction of GFP-Tet-v2.0 and GFP-Tet- v3.0 with sTCO.
[0049] DETAILED DESCRIPTION
[0050] The present disclosure provides compositions and methods for producing translational components that genetically encoded tetrazine amino acids in cells that meet the attributes needed for ideal bioorthogonal ligations. The components include orthogonal tRNAs, orthogonal aminoacyl-tRNA synthetases, orthogonal pairs of tRNAs / synthetases and tetrazine amino acids. Proteins containing tetrazine amino acids and methods of producing proteins with tetrazine amino acids in cells are also provided. In order to generate an ideal bioorthogonal reaction with the properties described the tetrazine amino acid reactivity is key to control. The tetrazine amino acid needs to be non-reactive to all biological conditions, in live cells, and in media, but highly reactive with its TCO partner. Orthogonal tRNAs / synthetases then need to be engineered to accept the tetrazine amino acid that contains these reactive properties. This results in tetrazine amino acid tRNA / synthetase pairs that function in cells to produce proteins with site-specifically incorporated tetrazine amino acids. The tetrazine amino acid containing proteins can then be used in cellulo, in vivo, or in vitro for ideal bioorthogonal ligations. The present disclosure provides composition and methods for producing tetrazine amino acid tRNA / synthetase pairs that are orthogonal in prokaryotic cells and eukaryotic cells. In certain embodiments, the methods for producing proteins with site-specifically incorporated tetrazine amino acids described herein are in cellulo methods. In other embodiments, the methods for producing proteins with site-specifically incorporated tetrazine amino acids described herein are cell-free methods.
[0051] The disclosure provides cells with translation components (e.g., pairs of orthogonal aminoacyl-tRNA synthetases (O-RSs) and orthogonal tRNAs (O-tRNAs) and individual components thereof, that are used in protein biosynthetic machinery to incorporate a tetrazine amino acid in a growing polypeptide chain, in a cell.
[0052] Compositions of the disclosure include a cell comprising an orthogonal aminoacyl- tRNA synthetase (O-RS), where the O-RS preferentially aminoacylates an orthogonal tRNA (O-tRNA) with at least one tetrazine amino acid (i.e., a tetrazine non-canonical amino acid as described herein) in the cell.
[0053] The cell also optionally includes the tetrazine amino acid(s). The cell optionally includes an orthogonal tRNA (O-tRNA), where the O-tRNA recognizes a selector codon and is preferentially aminoacylated with the tetrazine amino acid by the O-RS. In one aspect, the O-tRNA mediates the incorporation of the tetrazine amino acid into a protein with, for example, at least 45%, at least 50%, at least 60%, at least 75%, at least 80%, at least 90%, at least 95%, or 99% efficiency.
[0054] In another embodiment, the cell comprises a nucleic acid that comprises a polynucleotide that encodes a polypeptide of interest, where the polynucleotide comprises a selector codon that is recognized by the O-tRNA. In one aspect, the yield of the polypeptide of interest comprising the tetrazine amino acid is, e.g., at least 2.5%, at least 5%, at least 10%, at least 25%, at least 30%, at least 40%, 50% or more, of that obtained for the naturally occurring polypeptide of interest from a cell in which the polynucleotide lacks the selector codon. In another aspect, the cell produces the polypeptide of interest in the absence of the tetrazine amino acid, with a yield that is, e.g., less than 35%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 2.5%, of the yield of the polypeptide in the presence of the tetrazine amino acid.
[0055] The disclosure also provides a cell comprising an orthogonal aminoacyl-tRNA synthetase (O-RS), an orthogonal tRNA (O-tRNA), the tetrazine amino acid, and a nucleic acid that comprises a polynucleotide that encodes a polypeptide of interest. The polynucleotide comprises a selector codon that is recognized by the O-tRNA. In addition, the O-RS preferentially aminoacylates the orthogonal tRNA (O-tRNA) with the tetrazine amino acid in the cell, and the cell produces the polypeptide of interest in the absence of the tetrazine amino acid, with a yield that is, e.g., less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 2.5% of the yield of the polypeptide in the presence of the tetrazine amino acid.
[0056] Compositions that include a cell comprising an orthogonal tRNA (O-tRNA) are also a feature of the disclosure. Typically, the O-tRNA mediates incorporation of the tetrazine amino acid into a protein that is encoded by a polynucleotide that comprises a selection codon that is recognized by the O-tRNA in vivo. In one embodiment, the O-tRNA mediates the incorporation of the tetrazine amino acid into the protein with at least 45%, at least 50%, at least 60%, at least 75%, at least 80%, at least 90%, at least 95%, or even 99% efficiency.
[0057] In one aspect, the disclosure comprises a composition comprising a protein, wherein the protein comprises at least one tetrazine amino acid and at least one post-translational modification, wherein the at least one post- translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition to the at least one tetrazine amino acid comprising a first reactive group.
[0058] Thus, proteins (or polypeptides of interest) with at least one tetrazine amino acid are also a feature of the disclosure. In certain embodiments of the disclosure, a protein with at least one tetrazine amino acid includes at least one post-translational modification. In one embodiment, the at least one post-translational modification comprises attachment of a molecule (e.g., a dye, a polymer [e.g., a derivative of polyethylene glycol], a photocrosslinker, a cytotoxic compound, an affinity label, a derivative of biotin, a resin, a second protein or polypeptide, a metal chelator, a cofactor, a fatty acid, a carbohydrate, a polynucleotide (e.g., DNA, RNA) comprising a second reactive group by a [4+2] cycloaddition to the at least one tetrazine amino acid comprising a first reactive group. For example, the first reactive group is tetrazine moiety (e.g., a tetrazine moiety of a tetrazine non-canonical amino acid as described herein) and the second reactive group is an alkenyl moiety (e.g., a trans-cycloalkenyl moiety, such as a trans-cyclooctene (sTCO)). In certain embodiments, a protein of the disclosure includes at least one tetrazine amino acid comprising at least one post- translational modification. In certain embodiments, the post- translational modification is made in vivo in a cell.
[0059] Examples of a protein (or polypeptide of interest) include, but are not limited to, a cytokine, a growth factor, a growth factor receptor, an interferon, an interleukin, an inflammatory molecule, an oncogene product, a peptide hormone, a signal transduction molecule, a steroid hormone receptor, erythropoietin (EPO), insulin, human growth hormone, an alpha- 1 antitrypsin, an angiostatin, an antihemolytic factor, an antibody, an apolipoprotein, an apoprotein, an atrial natriuretic factor, an atrial natriuretic polypeptide, an atrial peptide, a C-X-C chemokine, T39765, NAP-2, ENA-78, a Gro-a, a Gro-b, a Gro- c, an IP-10, a GCP-2, an NAP-4, an SDF-1, a PF4, a MIG, a calcitonin, a c-kit ligand, a cytokine, a CC chemokine, a monocyte chemoattractant protein- 1, a monocyte chemoattractant protein-2, a monocyte chemoattractant protein-3, a monocyte inflammatory protein- 1 alpha, a monocyte inflammatory protein- 1 beta, RANTES, 1309, R83915, R91733, HCC1, T58847, D31065, T64262, a CD40, a CD40 ligand, a C-kit Ligand, a collagen, a colony stimulating factor (CSF), a complement factor 5a, a complement inhibitor, a complement receptor 1, a cytokine, DHFR, an epithelial neutrophil activating peptide-78, a GRO alpha / MGSA, a GRO beta, a GRO gamma, a MIP-1 alpha, a MIP-1 delta, a MCP-1, an epidermal growth factor (EGF), an epithelial neutrophil activating peptide, an erythropoietin (EPO), an exfoliating toxin, a Factor IX, a Factor VII, a Factor VIII, a Factor X, a fibroblast growth factor (FGF), a fibrinogen, a fibronectin, a G-CSF, a GM-CSF, a glucocerebrosidase, a gonadotropin, a growth factor, a growth factor receptor, a hedgehog protein, a hemoglobin, a hepatocyte growth factor (HGF), a hirudin, a human serum albumin, an ICAM-1, an ICAM-1 receptor, an LFA-1, an LFA-1 receptor, an insulin, an insulin-like growth factor (IGF), an IGF-I, an IGF-II, an interferon, an IFN- alpha, an IFN-beta, an IFN-gamma, an interleukin, an IL-1, an IL-2, an IL-3, an IL-4, an IL-5, an IL-6, an IL-7, an IL-8, an IL-9, an IL-10, an IL-11, an IL-12, a keratinocyte growth factor (KGF), a lactoferrin, a leukemia inhibitory factor, a luciferase, a neurturin, a neutrophil inhibitory factor (NIF), an oncostatin M, an osteogenic protein, an oncogene product, a parathyroid hormone, a PD-ECSF, a PDGF, a peptide hormone, a human growth hormone, a pleiotropin, a protein A, a protein G, a pyrogenic exotoxins A, B , or C, a relaxin, a renin, an SCF, a soluble complement receptor I, a soluble I-CAM 1 , a soluble interleukin receptors, a soluble TNF receptor, a somatomedin, a somatostatin, a somatotropin, a streptokinase, a superantigen, a staphylococcal enterotoxin, an SEA, an SEB, an SEC1, an SEC2, an SEC3, an SED, an SEE, a steroid hormone receptor, a superoxide dismutase (SOD), a toxic shock syndrome toxin, a thymosin alpha 1, a tissue plasminogen activator, a tumor growth factor (TGF), a TGF-alpha, a TGF-beta, a tumor necrosis factor, a tumor necrosis factor alpha, a tumor necrosis factor beta, a tumor necrosis factor receptor (TNFR), a VLA-4 protein, a VCAM-1 protein, a vascular endothelial growth factor (VEGEF), a urokinase, a Mos, a Ras, a Raf, a Met; a p53, a Tat, a Fos, a Myc, a Jun, a Myb, a Rel, an estrogen receptor, a progesterone receptor, a testosterone receptor, an aldosterone receptor, an LDL receptor, a SCF / c-Kit, a CD40L / CD40, a VLA-4 / VCAM-1, an ICAM-l / LFA-1, a hyalurin / CD44, a corticosterone, a protein present in Genebank or other available databases, and / or a portion thereof. In one embodiment, the polypeptide of interest includes a transcriptional modulator protein (e.g., a transcriptional activator protein (such as GAL4), or a transcriptional repressor protein) or a portion thereof.
[0060] The disclosure also provides methods for producing, in a cell, at least one protein comprising at least one tetrazine amino acid (as well as proteins produced by such methods). The methods include growing, in an appropriate medium, a cell that comprises a nucleic acid that comprises at least one selector codon and encodes the protein. The cell also comprises an orthogonal tRNA (O-tRNA) that functions in the cell and recognizes the selector codon and an orthogonal aminoacyl tRNA synthetase (O-RS) that preferentially aminoacylates the O-tRNA with the tetrazine amino acid, and the medium comprises a tetrazine amino acid. In one embodiment, the O-RS aminoacylates the O-tRNA with the tetrazine amino acid (e.g., at least 45%, at least 50%, at least 60%, at least 75%, at least 80%, at least 90%, at least 95%, or even 99%).
[0061] In one embodiment, the method further includes incorporating into the protein the tetrazine amino acid, where the tetrazine amino acid comprises a first reactive group; and contacting the protein with a molecule (e.g., a dye, a polymer, [derivative of polyethylene glycol], a photocrosslinker, a cytotoxic compound, an affinity label, a derivative of biotin, a resin, a second protein or polypeptide, a metal chelator, a cofactor, a fatty acid, a carbohydrate, a polynucleotide [e.g., DNA, RNA]) that comprises a second reactive group. The first reactive group reacts with the second reactive group to attach the molecule to the tetrazine amino acid through a [4+2] cycloaddition. In one embodiment, the first reactive group is tetrazine moiety (e.g., the tetrazine moiety of a tetrazine non-canonical amino acid as described herein) and the second reactive group is an alkenyl moiety (e.g., a trans- cycloalkenyl moiety, such as sTCO).
[0062] In certain embodiments, the encoded protein comprises a therapeutic protein, a diagnostic protein, an industrial enzyme, or portion thereof. In one embodiment, the protein that is produced by the method is further modified through the tetrazine amino acid. For example, the tetrazine amino acid is modified through a [4+2] cycloaddition. In another embodiment, the protein produced by the method is modified by at least one post-translational modification (e.g., N-glycosylation, O-glycosylation, acetylation, acylation, lipid-modification, palmitoylation, palmitate addition, phosphorylation, glycolipid-linkage modification) in vivo.
[0063] In certain embodiments, the compositions and the methods of the disclosure include cells. The translation components of the disclosure can be derived from a variety of organisms (e.g., non-eukaryotic organisms, such as a prokaryotic organism, or an archaebacterium, or a eukaryotic organism).
[0064] Kits are also a feature of the disclosure. For example, a kit for producing a protein that comprises at least one tetrazine amino acid in a cell is provided, where the kit includes a container containing a polynucleotide sequence encoding an O-tRNA or an O-tRNA, and a polynucleotide sequence encoding an O-RS or an O-RS. In one embodiment, the kit further includes at least one tetrazine amino acid (i.e., a tetrazine non-canonical amino acid as described herein). In another embodiment, the kit further comprises instructional materials for producing the tetrazine-containing protein.
[0065] As used herein, the term “orthogonal” refers to a molecule (e.g., an orthogonal tRNA (O-tRNA) and / or an orthogonal aminoacyl tRNA synthetase (O-RS)) that functions with endogenous components of a cell with reduced efficiency as compared to a corresponding molecule that is endogenous to the cell or translation system, or that fails to function with endogenous components of the cell. In the context of tRNAs and aminoacyl- tRNA synthetases, orthogonal refers to an inability or reduced efficiency, e.g., less than 20% efficient, less than 10% efficient, less than 5% efficient, or less than 1% efficient, of an orthogonal tRNA to function with an endogenous tRNA synthetase compared to an endogenous tRNA to function with the endogenous tRNA synthetase, or of an orthogonal aminoacyl-tRNA synthetase to function with an endogenous tRNA compared to an endogenous tRNA synthetase to function with the endogenous tRNA. The orthogonal molecule lacks a functional endogenous complementary molecule in the cell. For example, an orthogonal tRNA in a cell is aminoacylated by any endogenous RS of the cell with reduced or even zero efficiency, when compared to aminoacylation of an endogenous tRNA by the endogenous RS. In another example, an orthogonal RS aminoacylates any endogenous tRNA in a cell of interest with reduced or even zero efficiency, as compared to aminoacylation of the endogenous tRNA by an endogenous RS. A second orthogonal molecule can be introduced into the cell that functions with the first orthogonal molecule. For example, an orthogonal tRNA / RS pair includes introduced complementary components that function together in the cell with an efficiency (e.g., 50% efficiency, 60% efficiency, 70% efficiency, 75% efficiency, 80% efficiency, 90% efficiency, 95% efficiency, or 99% or more efficiency) to that of a corresponding tRNA / RS endogenous pair.
[0066] One advantage of the tetrazine non-canonical amino acids described herein is that they present additional chemical moieties that can be used to add additional molecules. These modifications can be made in vivo in a eukaryotic cell, or in vitro. Thus, in certain embodiments, the post- translational modification is through the tetrazine non-canonical amino acid. For example, the post-translational modification can be through a [4+2] cycloaddition reaction. Most reactions currently used for the selective modification of proteins involve covalent bond formation between nucleophilic and electrophilic reaction partners (e.g., the reaction of alpha-haloketones with histidine or cysteine side chains). Selectivity in these cases is determined by the number and accessibility of the nucleophilic residues in the protein. In proteins of the disclosure, other more selective reactions can be used in vitro and in vivo. This allows the selective labeling of virtually any protein with a host of reagents including fluorophores, crosslinking agents, polymers, saccharide derivatives and cytotoxic molecules.
[0067] Thus, this disclosure provides another highly efficient method for the selective modification of proteins, which involves the genetic incorporation of tetrazine amino acids into proteins in response to a selector codon. These tetrazine amino acid side chains can then be modified by a [4+2] cycloaddition reaction with strained alkenyl derivatives. Because this method involves a cycloaddition rather than a nucleophilic substitution, proteins can be modified with extremely high selectivity. This reaction can be carried out at room temperature in dilute aqueous conditions with excellent regio selectivity. In one aspect, the disclosure provides tetrazine non-canonical amino acids (Tet ncAAs) useful for genetically encoding the tetrazine non-canonical amino acid in a polypeptide of interest. The tetrazine non-canonical amino acids of the disclosure are referred to herein as Tet-v2.0 and Tet-v3.0 amino acids. As used herein the term “tetrazine non-canonical amino acid (Tet ncAA)” refers to a 1,2,4,5-tetrazine having an amino acid moiety at C3 (e.g., -(C6H4)CH2CH(NHRN)(CO2RC)) and optionally a substituent (other than hydrogen) at C6 (see, for example, formula (I) and (IT)).
[0068] The present disclosure provides tetrazine amino acids having formula (I) and (II): or a stereoisomer or salt thereof, wherein
[0069] R is selected from the group consisting of:
[0070] (a) a phenyl group substituted with a group selected from O-C1-C4 alkyl, m- C1-C4 alkyl, O-C1-C3 alkoxy, m-Cl-C3 alkoxy, o-cyano, m-cyano, o-nitro, m-nitro, o- or m-primary amino (-NH2), secondary amino (-NHRX), or tertiary amino (-NHRxRy) (wherein Rxand Ryare independently C1-C6 alkyl), o-fluoro, m-fluoro, 3,5-difluoro, 3,4,5- trifluoro, o-trifluoromethyl, m-trifluoromethyl, and o-, m-, or p-C(=O)Rz(wherein Rza counterion, hydrogen, or C1-C6 alkyl),
[0071] (b) a substituted or an unsubstituted heteroaryl group,
[0072] (c) a substituted or an unsubstituted heterocyclyl group,
[0073] (d) an amino C1-C6 alkyl group,
[0074] (e) a thio C 1-C6 alkyl group, (f) a carboxylate group,
[0075] (g) a sulfonate group, and
[0076] (h) an amide group;
[0077] Rcis hydrogen, a counter ion, or a carboxyl protecting group; and
[0078] RNis hydrogen or an amine protecting group.
[0079] In certain embodiments, R is a substituted phenyl group. Suitable substituents include C1-C6 alkyl (e.g., methyl), C1-C3 haloalkyl (trifluoromethyl), halo (e.g., fluoro, chloro), hydroxy, C1-C3 alkoxy (e.g., methoxy), cyano, nitro, -CCAR7, where z is a counterion, hydrogen, or C1-C6 alkyl group, and primary (-NH2), secondary (-NHRX), and tertiary amine (-NRXRV) (where Rxand Ry are independently C1-C6 alkyl). The phenyl groups may include one or more substituents at the o-, m-, and / or p- (i.e., 2-, 3-, and / or 4-) positions. Representative substituted phenyl groups include phenyl groups substituted with m-methyl (3-methyl), o-methoxy (2-methoxy), m-methoxy (3-methoxy), o-cyano (2- cyano), m-cyano (3-cyano), o-nitro (2-nitro), m-nitro (3-nitro), o-fluoro (2-fluoro), m- fluoro (3-fluoro), 3,5-difluoro, 3,4,5-trifluoro, o-hydroxy (2-hydroxy), m-hydroxy (3- hydroxy), 0-NH2, (2-NHz), 1U-NH2 (3-NHz), m-C(=0)CH3 (2-acetyl), o-trifluoromethyl (2-trifluoromethyl), and m-trifluoromethyl (3-trifluoromethyl).
[0080] In further embodiments, R is an unsubstituted heteroaryl group or a substituted heteroaryl group. As used herein, the term “heteroaryl” refers to a monocyclic heteroaryl group that is a 3- to 6-membered carbocyclic group containing one or more nitrogen, oxygen, or sulfur atoms in the carbocyclic ring. Suitable heteroaryl groups include pyridyl (e.g., 2-, 3-, and 4-pyridyl), furanyl (2- and 3-furanyl), thiophenyl (2- and 3-thiophenyl), and oxazolyl (e.g., 2-, 3-, 4-, and 5-oxazolyl) groups. Suitable substituents include C1-C6 alkyl (e.g., methyl), C1-C3 haloalkyl (trifluoromethyl), halo (e.g., fluoro, chloro), hydroxy, C1-C3 alkoxy (e.g., methoxy), cyano, nitro, and primary (-NH2), secondary (-NHRX), and tertiary amine (-NRXR ) (where Rxand Ry are independently Cl -C6 alkyl). The heteroaryl groups may include one or more substituents (e.g., at one or more of each ring position). In other embodiments, R is a substituted or an unsubstituted heterocyclyl group. As used herein, the term “heterocyclyl” refers to a monocyclic heterocyclyl group that is a 3- to 6-membered carbocyclic group containing one or more nitrogen, oxygen, or sulfur atoms in the carbocyclic ring. Suitable heterocyclyl groups include pyrrolidinyl (e.g., 2- and 3- pyrrolidinyl), piperadinyl (e.g., 2-, 3-, and 4-piperadinyl), piperazinyl (e.g., 2- and 3- piperazinyl), tetrahydrofuranyl (e.g., 2- and 3-tetrahydrofuranyl), tetrahydropyranyl (e.g., 2-, 3-, and 4-tetrahydropyranyl), tetrahydrothiophenyl (e.g., 2- and 3- tetrahydro thiophenyl), and oxazolydinyl (e.g., 2-, 3-, 4-, and 5-oxazolydinyl) groups. Suitable heterocyclyl groups include dihydropyran (DHP) groups. Suitable substituents include C1-C6 alkyl (e.g., methyl), C1-C3 haloalkyl (trifluoromethyl), halo (e.g., fluoro, chloro), hydroxy, C1-C3 alkoxy (e.g., methoxy), cyano, nitro, and primary (-NH2), secondary (-NHRX), and tertiary amine (-NRxRy) (where Rxand Ry are independently Cl- C6 alkyl).
[0081] In another embodiment, R is an amino C1-C6 alkyl group. Representative amino C1-C6 alkyl groups include aminomethyl, aminoethyl, aminopropyl, aminobutyl, aminopentyl, and aminohexyl groups. In one embodiment, the amino C1-C6 alkyl group is an aminobutyl (-NH(CH2)3CH3) group.
[0082] In a further embodiment, R is a thio C1-C6 alkyl group. Representative thio Cl- C6 alkyl groups include thiomethyl, thioethyl, thiopropyl, thiobutyl, thiopentyl and thiohexyl groups. In one embodiment, the thio C1-C6 alkyl group is a thiomethyl (-SCH3) group.
[0083] In another embodiment, R is a carboxylate group (e.g., -CO2RZ, where z is a counterion, hydrogen, or C1-C6 alkyl group). Representative carboxylate groups include methyl carboxylate (-CO2CH3) and ethyl carboxylate (-CO2CH2CH3) groups. In one embodiment, the carboxylate group is an ethyl carboxylate group.
[0084] In another embodiment, R is a sulfonate group (e.g., -SO3R7. where z is a counterion, hydrogen, or C1-C6 alkyl group). Representative sulfonate groups include sulfonate (-SO3) and methyl sulfonate (-SO2OCH3) and ethyl carboxylate (-CO2CH2CH3) groups. In a further embodiment, R is an amide group. Representative amide groups include C1-C6 alkyl amide groups (e.g., methyl amide ( C(=O)NHCH3), butyl amide (- C(=O)NH(CH2)3CH3)).
[0085] Carboxyl protecting groups, amine protecting groups, and counter ions include those known in the art. Suitable carboxyl protecting groups and amine protecting groups are described in Protective Groups in Organic Synthesis, T.W. Greene, John Wiley & Sons, 1981, expressly incorporated herein by reference in its entirety.
[0086] Suitable carboxyl protecting groups include ester, amide, and hydrazine groups. Representative ester groups include substituted methyl esters (e.g., methoxymethyl, methylthiomethyl, tetrahydropyranyl, tetrahydrofuranyl, methoxyethoxymethyl, benzyloxy methyl, phenacyl, p-bromophenacyl, a-methylphenacyl, p-methoxyphenacyl, diacylmethyl, N-phthalimidomethyl, and ethyl), 2-substituted ethyl esters (e.g., 2,2,2-trichloroethyl, 2-haloethyl, co-chloroalkyl, 2-(trimethylsilyl)ethyl, 2-methylthioethyl, 2-(p-nitrophenylsulfenyl)ethyl, 2-(p-toluenesulfonyl)ethyl, 1-methyl-l- phenylethyl, t-butyl, cyclopentyl, cyclohexyl, allyl, cinnamyl, phenyl, p-methylthiophenyl, and benzyl), substituted benzyl esters (e.g., triphenylmethyl, diphenylmethyl, bis(o- nitrophenyl)methyl, 9-anthrylmethyl, 2-(9,10-dioxo)anthrylmethyl, 5-dibenzosuberyl, 2,4,6-trimethylbenzyl, p-bromobenzyl, o-nitrobenzyl, p-nitrobenzyl, p-methoxybenzyl, piperonyl, and 4-picolyl), silyl esters (e.g., trimethylsilyl, triethylsilyl, t-butyldimethylsilyl, i-propyldimethylsilyl, and phenyldimethylsilyl), activated esters (e.g., S-t-butyl, S-phenyl, S-2-pyridyl, N-hydroxypiperidinyl, N-hydroxysuccinimidoyl, N-hydroxyphthalimidoyl, and N-hydroxybenzotriazolyl), stannyl esters (e.g., triethylstannyl, tri-n-butylstannyl), and other esters (e.g., O-acyl oximes, 2,4-dinitrophenylsulfenyl, 2-alkyl-l,3-oxazolidines, 4- alkyl-5-oxo-l,3-oxazolidines, and 5-alkyl-4-oxo-l,3-dioxolanes). Representative amides include N,N-dimethyl, pyrrolidinyl, piperidinyl, o-nitrophenyl, 7-nitroindolyl, and 8- nitrotetrahydroquinolyl. Representative hydrazines include N-phenylhydrazide and N,N’- diisopropylhy drazide . Suitable amino protecting groups include carbamate and amide groups. Representative carbamate groups include alkyl and aryl carbamate groups (e.g., methyl and substituted methyl, substituted ethyl, substituted propyl and isopropyl, t-butyl, cyclobutyl, cyclopentyl, cyclohexyl, 1-adamantyl, vinyl, allyl, cinnamyl, phenyl, and benzyl). Representative amide groups include N-formyl, N-acetyl, substituted N-propionyl, cyclic imides, N-alkyl amide (e.g., N-allyl, N-phenacyl), amino acetals, N-benzyl amides, imine derivatives, enamine derivatives, N-heteroatom derivatives, N-metal derivatives (e.g., N- borane, N-copper, N-zinc), N-N derivatives (e.g., N-nitro, N-nitroso), N-P derivatives (e.g., phosphinyl, phosphoryl) N-Si derivatives (e.g., N-trimethylsilyl), and N-S derivatives (e.g., N-sulfenyl, N-sulfonyl).
[0087] In certain embodiments, the compounds of the disclosure are amino acids and maybe exist in neutral (e.g., -NH2and -CO2H) or ionic form (e.g., -NH3+ and -CO2) depending on the pH of the environment. It will be appreciated that the compounds of the disclosure include a chiral carbon center and that the compounds of the disclosure can take the form of a single stereoisomer (e.g., L or D isomer) or a mixture of stereoisomers (e.g., a racemic mixture or other mixture). It will be appreciated that the individual stereoisomers and mixtures of isomers are useful in methods of the disclosure for incorporating tetrazinecontaining residues into proteins and polypeptides.
[0088] The tetrazine amino acids described herein are useful in the methods for genetically encoding a polypeptide of interest.
[0089] Representative tetrazine amino acids of the present disclosure include those shown in FIGS. 1, 2B, and 4.
[0090] The preparation of representative tetrazine amino acids of the present disclosure is illustrated in FIG. 2A.
[0091] In another aspect, the disclosure provides methods for making proteins or polypeptides that include tetrazine-containing residues derived from tetrazine amino acids of formula (I) or (II). In one embodiment, the method includes incorporating a tetrazine amino acid of the disclosure into the protein or polypeptide. Tetrazine amino acids can be incorporated into a protein or polypeptide by conventional synthetic techniques (e.g., peptide synthesis, such as solid phase peptide synthesis). Alternatively, the tetrazine amino acid can be incorporated into a protein or polypeptide by genetic encoding, as described herein in detail. In one embodiment, the disclosure provides a method for genetically encoding a protein or polypeptide of interest that includes incorporating a tetrazine amino acid of the disclosure into the protein or polypeptide by genetic encoding.
[0092] In a further aspect, the disclosure provides a protein or polypeptide that includes at least one tetrazine amino acid residue is provided. The tetrazine amino acid residue is derived from a tetrazine amino acid of the disclosure. In one embodiment, the disclosure provides a protein or polypeptide, comprising at least one tetrazine amino acid residue, wherein the tetrazine amino acid residue is incorporated into the protein or polypeptide by genetic encoding of the protein or polypeptide using a tetrazine amino acid of the disclosure.
[0093] In another aspect, the disclosure provides a post-translationally modified composition protein or polypeptide (e.g., composition) comprising a protein or polypeptide that comprises at least one tetrazine amino acid residue and at least one post-translational modification, wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition to the at least one tetrazine amino acid residue comprising a first reactive group.
[0094] In certain embodiments, the disclosure provides a composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid residue comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from a tetrazinecontaining compound, wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid residue comprising the first reactive group.
[0095] In other embodiments, the disclosure provides a composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid residue comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from genetic encoding of the protein or polypeptide with a tetrazine-containing compound, wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid residue comprising the first reactive group.
[0096] In certain of the embodiments, the first reactive group is the tetrazine group of the tetrazine amino acid residue and the second reactive group is a suitably reactive group reactive (alkyne or alkene, such as a strained alkyne or strained alkene, or a cyclic alkene or cyclic alkyne).
[0097] In another aspect of the disclosure, a kit for in cellule production of a tetrazine- labeled protein or a tetrazine-labeled polypeptide is provided. In certain embodiments, the kit includes:
[0098] (a) a tRNA;
[0099] (b) an aminoacyl-tRNA synthetase; and
[0100] (c) a tetrazine amino acid as described herein, wherein the tRNA and aminoacyl-tRNA synthetase are an orthogonal tRNA / orthogonal aminoacyl-tRNA pair effective for incorporating the tetrazine amino acid into a protein to provide a tetrazine-labeled protein.
[0101] The kit can include additional components to facilitate in cellule protein production using the tRNA, the aminoacyl-tRNA synthetase, and the tetrazine amino acid. Methods for use of tetrazine amino acids including their genetic encoding to provide a tetrazinemodified protein and bioorthogonal ligation using the tetrazine-modified protein are described in PCT / US2016 / 030469, entitled “Reagents and Methods for Bioorthogonal Labeling of Biomolecules in Living Cells”, expressly incorporated herein by reference in its entirety. These methods are applicable to the tetrazine amino acids of the disclosure. The following describes methods for genetic encoding representative Tet ncAAs described herein to provide a tetrazine-modified protein and bioorthogonal ligation using the tetrazine-modified protein.
[0102] The reactivity of the tetrazine functional group of the encoded tetrazine amino acid in the tetrazine- alkene cycloaddition is tunable due to the tetrazine substituent (FIG. 1). The synthetic accessibility and reaction rates of representative Tet ncAAs, their encoding ability into proteins with different evolved tRNA / RS pairs, and assessment of their reaction extent, rate constants, and stability in proteins is described herein. Eukaryotic compatible Tet ncAAs encoding systems were also evaluated for encoding efficiency, selectivity, and in-cell labeling in mammalian cells. Attributes for Tet ncAAs for developing and implementing GCE encoded bioorthogonal ligations include synthetic accessibility, encodability, stability, and labeling rate.
[0103] Synthetic Accessibility of Tet ncAAs
[0104] Growing interest in the use of tetrazine biorthogonal chemistry in living system is reflected through the massive serge in tetrazine utilizing publications over the past decade. Initially, the lack of synthetic accessibility of asymmetrical 1,2, 4, 5 -tetrazines was a roadblock for selecting GCE machinery that would site-specifically encode the tetrazine functionality. The present disclosure provides a generalizable synthetic scheme for the preparation of tetrazine compounds containing ncAA for genetic code expansion (see FIG. 2A). The tetrazine ncAAs prepared by the scheme have sufficiently small size and correct shape to fit into the active site of engineered tRNA / RS pairs in addition to being stable for the 24-36 hrs in culture media with cells at 37 °C. In 2012, the first tetrazine ncAA was synthesized by coupling an electrophilic thiomethyl containing 1,2,4,5-tetrazine with and boc -protected 4-amino phenyl alanine (referred to herein as “Tet-1”) having an amine linkage and prepared in moderate synthetic yield 33%. Having sufficient quantities of Tet- 1 ncAA enabled the evolution of a Methanococcus jannaschii (Mj) tyrosyl tRNA synthetase to encode Tet-1, which was a demonstration that encoding the tetrazine functionality was possible. This first system showed moderate reaction kinetics (fa of 880 M-1s-1) and demonstrated low levels of TCO label loss via elimination at the amine linkage. Subsequent efforts for the development of improved tetrazine-containing ncAA for genetic encoding are described herein.
[0105] A representative Tet ncAA (Tet-v2.0Me also referred to herein as “Tet-2Me”) was synthesized by the addition of hydrazine and boc-4-nitrile phenyl alanine in presence of acetonitrile using NifOTfh as a catalyst provided 57% yield, which was later optimized 78% yield. The removal of the amine bridge from Tet-1 significantly altered the ncAA shape requiring the selection of a new MJ Tet-RS / tRNA pair for its genetic encoding. Methanococcus jannaschii (Mj) tyrosyl tRNA synthetase for Tet-v2.0Me (D12-Tet- v2.0Me) Tet-2Me inside protein showed excellent stability and fast reactivity with sTCO reagents (8 x 104M 's1) that greatly advances the tetrazine-based quantitative Tet-protein labeling and demonstrate sub-stochiometric labeling in live E. coli cells. Unfortunately, the Tet-2Me structure shape, which was efficiently encoded by the prokaryotic restrained MJ system, was not accepted in the active site by more universal PylRS / tRNA. Using the Tet-2 synthetic approach described herein, Tet- 3 was generated using boc-3-nitrile phenyl alanine where the synthesized tetrazine attached to the meta position of phenyl ring instead of para position in Tet-v2.0 (FIGS. 2A and 2B). This meta Tet-3-methyl derivative was accepted into selected PylRS / tRNA active sites and provided access to Tet protein encoding and fast labeling in eukaryotic cells.
[0106] To achieve effective labeling under biological conditions exceptionally fast and stable tetrazine amino acid is highly desirable. The balance of stability and reactivity of Tet ncAAs was attained by the generation of 1,2,4,5-Tet ncAAs by attaching different donor and acceptor substituents following the synthetic methods described herein.
[0107] Toxicity of Tet ncAAs
[0108] Stable and nontoxicity of Tet ncAAs is essential for incorporation of the Tet ncAAs into recombinant proteins. The toxicity of synthesized Tet ncAAs to the E.coli cells was tested by measuring the cell growth in presence of Tet ncAAs with 0.25 mM, 0.5 mM, 0.75 mM and 1 mM concentrations. A fresh and saturated DH10B E. coli cells stock was inoculated into 5 mL auto inducing media (AIM) with and without Tet ncAAs and allowed to grow for 40 hrs. at 37 °C. Measured optical density (OD) of cultured cells at 24 and 40 hrs. clearly indicated that the cells are well behaved in presence of 1 mM Tet-v2.0Me and Tet-v3.0Me with minimal effect on cell growth whereas the pyridyl substituents are significantly reduces the cells growth. Tet-v4.0 alkyl substituents were not toxic to the cells.
[0109] Therefore, to minimize the Tet ncAAs toxicity the concentration of the Tet ncAA can be reduced by improving the efficiency of GCE machinery. Cell toxicity can be avoided some extent if Tet ncAA is supplemented into the media after 4 - 6 hrs. of cell inoculation when the cells are healthy enough to grow (ODcen about 0.5 to 0.7).
[0110] Genetic incorporation of Tet ncAAs into proteins
[0111] To incorporate Tet ncAAs into the protein, a GCE machinery was developed to encode Tet ncAAs onto the protein in response to TAG codon suppression with evolving orthogonal aaRS / tRNAcuA pair. The permissibility of developed synthetases for structurally parallel Tet ncAA derivatives was evaluated in 25 mL scale in presence and absence of 0.5 mM Tet ncAA.
[0112] Devolved Methanococcus jannaschii (Mj) tyrosyl tRNA synthetase for Tet-v2.0Me (D12-Tet-v2.0Me) is permissive for Tet-v2.0Et, Tet-v2.0Ip but unable to incorporate aromatic substituted Tet-v2.0Ph and Tet-v2.0Pyr (FIG. 3A). While a eukaryotic orthogonal pyrrolysyl amino-acyl-tRNA synthetase / pyrrolysyl-tRNAcuA (PylRS / Pyl-tRNAcuA) pair from Methanosarcina barkeri (Mb) system was unable to incorporate Tet-v2.0Me due to steric hinderance between the tetrazine ring and the peptide backbone of PylRS active sites, structurally distinct Tet-v3.0Me is efficiently encoded with evolved Mb-PylRS R284-Tet- v3.0Me by bypassing the steric hinderance. The synthetase R284-Tet-v3.0Me is well permissive and high fidelity for Tet-v3.0Et, Tet-v3.0Ip, Tet-v3.0Bu and moderately permissive for Tet-v3.0Ph, Tet-v3.0Pyr, Tet-v3.0_4-F-Ph, and poorly permissive for Tet- v3.0_4-NH2-Ph (FIG. 3B). Tet-v3.0Bu is encodable efficiently into eukaryotic proteins and labeled quantitatively with sTCO reagents. The incorporation efficiency and fidelity of the selected synthetases to representative Tet ncAAs were verified by mass spectra analysis. The His6-tag purified GFP150-Tet-v3.0 (Ph, 4-F-Ph, Pyr) characterized by a single major peak corresponding to Tet ncAA incorporation (FIG. 7A, 7B, and 7D, respectively). Due to the moderate permissibility it shows some amount of natural amino acid (glutamic acid) incorporation (i.e., known as near-cognate suppression, NCS) identified with less intense mass peak at 27841 Da. Similarly, the poorly permissive for Tet-v3.0_4-NH2-Ph shows very intense peak at 27841 due to larger amount of near-cognate suppression (FIG. 7C).
[0113] Tet ncAA stability and reactivity inside protein
[0114] To achieve clean and quantitative Tet-protein labeling, stability of incorporated Tet ncAA is essential. Recombinant protein production through biosynthetic pathway requires long time incubation of Tet ncAA into the expression media. Whereas the reactive tetrazine scaffolds have a tendency to lose their reactivity either by reversible reduction or crossreactivity with biological nucleophiles, tetrazine stability can be tuned by attaching substituents with different steric and electronic properties. The inherent stability-reactivity balance of Tet ncAA, which is inversely proportional to each other, is highly desirable for in-cell protein labeling.
[0115] Several studies have shown that the highly reactive H-substituted Tet-v2.0 degrades under biological environment. However, Tet ncAA stability incredibly increased by replacing the hydrogen of tetrazine C6 with alkyl or aryl substituents. As described herein, electron-withdrawing groups reduce stability and electron-donating group enhance stability of the Tet ncAA relative to unsubstituted Tc ncAA. The electron-withdrawing dipyridyl- and pyrimidyl- attached tetrazines are unstable where 60-80% of tetrazines are degraded after 12 hrs. incubation in PBS at 37°C. Phenyl- substituted and unsubstituted Tet ncAA remain intact (about 75%) at the same conditions and the electron -rich hydroxylsubstituted tetrazines show exceptional stability with marginal degradation. Therefore, in the context of Tet-ncAA development, as described herein, the amine linkage of Tet-vl.O was replaced with weak electron-donating methylene or phenyl groups to provide the tetrazine amino acid compounds (e.g., Tet-v2.0, Tet-v3.0) that avoid the hydrolysis and improve the stability.
[0116] The stability and reaction efficiency of protein-incorporated Tet ncAAs was investigated. The Tet-proteins were stored at 4 °C in 50 mM PBS (pH 7.1) for two days after protein expressions (32 hrs. incubation in AIM at 37 °C) and purification. The SDS- PAGE gel mobility shift assay and mass spectrometry analysis of the incorporated Tet-v2.0 and Tet-v3.0 derivatives confirmed their stability and reactivity with sTCO.
[0117] Mobility shift assay. Labeling efficiency and purity of Tet-containing proteins have been verified by reacting with 10 eqv. sTCO-PEG5k under PBS (pH 7.1) for 10 minutes at room temperature and examined the mobility shift of the reacted proteins in SDS-PAGE gel. The mobility shift assay provides insight into the reaction efficiency of Tet-containing proteins and also quantifies the amount of reactive Tet ncAA present inside all portion of protein and measures if there is any misincorporation or Tet-degradation by showing unreactive protein. As described herein, nearly all of GFP150-Tet reacted, and a minimal unreacted band was observed for efficiently Tet-encoded proteins which measures the incorporated Tet ncAAs are stable and reacts quantitatively with sTCO. However, the less efficiently or poorly incorporated Tet-containing proteins are not able react quantitatively due to near-cognate suppression and resulted considerable amount of unreacted proteins were observed for Tet-v3.0 (Ph, 4-F-Ph, Pyr) and larger amount remains unreactive for Tet- v3.0_4-NH2-Ph (FIG. 5).
[0118] In addition, to assess their stability and reactivity more rigorously, GFP-Tet-v3.0 (Ph, Pyr) were incubated at 4 °C and room temperature (RT) under 50 mM PBS (pH 7.1) with 100 mM imidazole and studied their reaction efficiency for another 8 days. The reactivity remains same at 4 °C and at RT, shows efficient labeling up to 3 days then a minimal degradation was observed. However, the electron-withdrawing pyridyl attached GFP-Tet-3.0Pyr degraded with time and no reaction was observed after 3 days due to loss of tetrazine functionality at RT (FIGS. 6 A and 6B). Mass spectrometry. The site-specific labelling reaction between Tet ncAA and sTCO is quantitative inside protein was confirmed by ESI-Q-TOF mass spectrometry. A single major peak has been observed in mass spectra corresponding to single site Tet- incorporation into GFP150. Therefore, to study their rection efficiency the GFP-Tet proteins were exposed with 10 eqv. sTCO for 10 minutes under PBS (pH 7.1) at room temperature, then desalted the sample for mass spectrometry. As a result, the major single peaks were completely shifted to the 124.2 Da higher masses for the addition of sTCO with loss of N2 and verified the Tet-sTCO adduct formation. The peak at 27841 Da corresponds to natural amino acids incorporation remains unreactive (FIGS. 7A-7D and Table 1 below). No other peak has been detected for sTCO exposed samples which verified that the GFP- Tet reacts with sTCO quantitatively without involving side reactions.
[0119] GFP-Tet- v 3. OPh / Pyr reaction efficiency with sTCO after 8 days incubation at 4 °C and RT under 50 mM PBS (pH 7.1) with 100 mM imidazole were verified by MS analysis. A single major mass peak was observed correspond to GFP-Tet-v3.0Ph without any degradation and as expected, upon reaction with sTCO, it completely shifted to 124.2 Da higher mass. No side reactions were detected (FIG. 8A).
[0120] However, at RT, the GFP-Tet-v3.0Pyr showed a major peak which is 11 Da. unit less than expected mass and incubation with sTCO exhibited a small portion was reacted where the major mass remains unreactive (FIG. 8B). The MS data for Tet-v3.0Pyr suggested that the major mass at 28006.5 / 27930.5 Da avg. corresponds to tetrazine degradation which was converted to its oxadiazole derivative and is therefore not reactive to the sTCO. All above findings demonstrated that the methyl, ethyl, isopropyl, butyl, phenyl and substituted phenyl derivatives of Tet ncAA are stable and reactive inside protein however, the Pyr-Tet ncAA degrades under harsh conditions.
[0121] Tet ncAA kinetics with sTCO inside protein and their limits
[0122] In application of labeling reaction under biological system, a fast kinetics (fe > 104M-1s-1) is necessary to compete with biological reaction. A high-speed labeling approach can reduce the labeling time and able to react at low concentrations of labels which is beneficial for low abundant protein labeling. Considering genetically encoded Tet ncAA based IEDAA reaction is under the mid-range of reported conjugation reaction as the fastest Tet-amino acids are less stable in aqueous system. While there are many factors for tuning the reactivity, such as electronic properties and steric hindrance of the substituents, varying the Tet-attached substituents significantly impacts balancing the reaction rate and stability of encoded Tet ncAAs. As described herein, tetrazine reactivity enhances by attaching electron-withdrawing substituents while rendering them more electrophilic and susceptible to nucleophilic attack and stimulating self-degradation. Electron-donating substituents slow the reaction rate but improved the stability. Therefore, the aqueous stability vs. fast reactivity needs to be optimized to develop an ideal system for tetrazine based efficient protein labeling both in vitro and in-cell.
[0123] Reaction kinetics of representative Tet-ncAAs adorned with electron-donating and electron-withdrawing alkyl and aryl substituents are described herein (FIG. 4). Free Tet ncAA reactivity with sTCO using stopped-flow spectroscopy was evaluated. In PBS (pH 7.1 ), Tet-v2.0Me and Tet-v3.0(Me, Et, Ip, Bu) compounds showed moderate reaction rates, 5000 -29000 M’'s’'. Tet-v3.0Ip displayed decreased reactivity due to steric interference. Electron- withdrawing pyridyl for Tet-v3.0Pyr showed enhanced reactivity (fe = 96000 M’ 's’1) and Tet-v2.0Pyr derivatives showed the greatest reactivity, 204000 M 's’1.
[0124] Tet-reaction kinetics on protein were determined by measuring the regenerated fluorescence of GFP-Tet upon reaction with sTCO. The sfGFP fluorescence get quenched 4 -6 times when Tet-amino acids were incorporated at TAG site 150 and it can be recovered by reacting with sTCO due to the adduct formation and loss of tetrazine functionality. Protein encoded Tet-amino acid reactivity was enhanced three-fold compared to free amino acids due to the hydrophobic effect at the protein surface. However, the protein encoded Tet-vl.O and Tet-v2.0-H showed modest to moderate reactivity at 880 M 's’1and 3 x 104M 's’1, respectively, with sTCO and TCO* but these amino acids are electronically unfavorable for long time incubation. Tet-v2.0 and Tet-v3.0 derivatives exhibited moderate to high reactivity with second-order rate constant 2 x 104- 8 xlO4M 's’1, which is several orders of magnitude greater than Tet-vl.O. Though Tet-v2.0Me and Tet-v3.0Me are structurally different, their encoded protein reactivity and stability are similar. As with free-Tet amino acid reactivity, steric interference of the isopropyl (Ip) group reduces GFP- Tet-v2.0Ip / Tet-v3.0Ip (fa= 0.9 x 104M_1s-1 / 2.2 x 104M ’s1) reactivity. GFP-Tet-v3.0Pyr provides super-fast kinetics (fa=23 x 104M 's1). Butyl, phenyl, and 4-F-phenyl derivatives Tet-v3.0Bu, Tet-v3.0Ph, and Tet-v3.0_4-F-Ph exhibited high reactivity (fe = 7.4 -8.6 x 104M ’s1) with excellent stability for biological study (FIG. 4 and FIGS. 9A-9D).
[0125] Tet nCAA - sTCO product linkage stability
[0126] A stable linkage between the complementary reagents of ligation reaction is essential for labeling. The stability and biocompatibility of the tetrazine-sTCO adduct under physiological condition is proven for various applications in different research areas. While both reagents sTCO and tetrazine are sensitive to acidic and basic conditions, the conjugated product 1 ,4-dihydropyridazine formed by carbon-carbon bond formation is stable and inert to biological milieu. Sometimes the conjugated product 1 ,4- dihydropyridazine aromatize to stable pyridazine for electron rich-tetrazines under ambient conditions. Tetrazine-sTCO adduct stability inside protein was tested by SDS-page gel mobility shift assay and MS analysis.
[0127] Eukaryotic protein labeling
[0128] Eukaryotic protein labeling under cellular environments is a useful tool for studying protein dynamics, function, and localization in mammalian cells. To achieve effective intracellular labeling, the labeling reagent and the adduct formed after conjugation must be stable and orthogonal to the cellular compartments, applied conjugation reaction rate should be rapid enough (> 105M 's1) for quantitative labeling within short timeframe at low concentration level (nM to pM) that reduces the amount of loaded probe for quantitative labeling and minimizes the probability of background reaction. Moreover, most of the labeling model used side chain modified cyclooctynes (BCNs) and cyclooctenes (TCOs) ncAAs for genetic incorporation and tetrazine attached probes for fast labeling. The present disclosure provides a eukaryotic compatible GCE system for Tet ncAA incorporation for effective protein expression in mammalian cells. HEK293T cells were healthy with 0.3 mM Tet ncAA and cell viability was compromised above 0.3 mM tetrazine amino acid and efficient protein production was observed at concentration level of 30 to 300 pM. Suppression efficiency increases about 30 - 40% after the addition of a nuclear export sequences (NES) to the aaRS. Genetically incorporated Tet-v3.0Bu by eukaryotic orthogonal pyrrolysyl-tRNA synthetase PylRS / tRNAcuA pairs showed stability in HEK293T cells expressed proteins and enabled reaction with sTCO reagents without any detectable degradation products. Fast reaction kinetics of Tet-v3.0Bu (fa = 8 x 104M 's1) in protein allows quantitative and sub-stoichiometric labeling where the sTCO reagent is limited.
[0129] Tablel. Summary table of representative tetrazine amino acids (Compounds 1-29) for both prokaryotic and eukaryotic protein labelling and their reactivity inside protein.
[0130] General Synthetic Methods
[0131] All purchased chemicals were used without further purification. 3-Amino- PROXYL was purchased from Toronto Research Chemicals. Anhydrous dichloromethane (DCM) and dimethyl sulfoxide were used after overnight stirring with calcium hydride and distillation under argon atmosphere. Thin layer chromatography (TLC) was performed on silica 60F-254 plates. The TLC spots of alkene were identified by potassium permanganate staining. Flash chromatographic purification was performed using silica gel 60 (230-400 mesh size).]H NMR spectra were recorded on Broker at 400MHz and 700 MHz and13C NMR spectra were recorded at 175 MHz. The chemical shifts were shown in ppm and are referenced to the residual non-deuterated solvent peak CDCh (5 =7.26 in]H NMR, 5 = 77.23 in13C NMR), CD3OD (5 =3.31 in!H NMR, 6 = 49.2 in13C NMR), d6-DMSO (5 =2.5 in!H NMR, 5 = 39.5 in13C NMR) as an internal standard. Splitting patterns of protons are designated as follows: s-singlet, d-doublet, t-triplet, q-quartet, m-multiplet, bs- broad singlet, dd- doublet of doublets.
[0132] Tetrazine amino acid rate constant measurements. The solutions of tetrazine amino acids (0.2 mM) and sTCO-OH (1.0-6.0 mM) were made in PBS (137 mM NaCl, 2.7 mM KC1, 10 mM NaiHPO-i, 1.8 mM KH2PO4, pH 7.4) with 2% methanol to ensure solubility of Tet ncAAs. The measured loss of tetrazine absorbance at 270 nm was used to determine the reaction rate of Tet ncAAs. Pseudo-first order conditions were employed with sTCO-
[0133] OH in 5- to 60- fold excess of tetrazine amino acids at 25 °C. All measurements were performed in triplicate and the resulted decay curves were fit to a single exponential equation. The mean values of pseudo first order rate constants (k ) were plotted against different concentration of sTCO-OH to obtain second order rate constants (fo) from the slope of the plot.
[0134] Efficiency and fidelity of selected synthetases. The efficiency and fidelity of the selected synthetases were measured by expressing GFP150 in 50 mL AIM with and without 0.5 mM Tet ncAA. DHIOb cells containing the selected pBK-RSs and the pALS- GFP150TAG plasmids, were used to inoculate 5 mL of NIM containing kanamycin (50 pg / mL) and tetracycline (25 pg / mL). Cells were grown for 16 hours at 37 °C shaking at 250 rpm. The saturated NIM cultures (500 pL) were used to inoculate in 50 mL AIM containing kanamycin (50 pg / mL) and tetracycline (25 pg / mL) and measured fluorescence using a Turner Biosystems Picofluor fluorimeter diluting 100 pL cell culture in 1.9 mL water. Two synthetases D12 and R248 were identified to have good efficiency and fidelity.
[0135] Expression and purification of GFP150-TAG-Tet ncAAs. Using expression conditions above for efficiency and fidelity measurements, GFP150-Tet ncAAs were expressed in 50 mL AIM and all cells were harvested after 36 hrs by centrifugation at 5000 ref for 10 min. Media was removed, and cell pellets were stored at -80 °C. Cells were resuspended in wash buffer (NaCl 300 mM, NaH2PO4 15.5 mM, Na2HPO4 34.5 mM, imidazole 5 mM, pH 7.1). Cells were lysed using a Microfluidics M-110P microfluidizer (18,000 psi) and the lysate was collected in wash buffer. The lysate was clarified by centrifugation (21000 ref, 30 mins.) and to the clarified supernatant TALON resin (300 pL bed volume) was added. Lysate was incubated with the resin for 1-2 hours gently rocking at 4° C. Resin and lysate were applied to a column and flow through was discarded. Resin was washed 5 times with 10 mL wash buffer. Protein was eluted with 250 pL elution buffer (NaCl 300 mM, NaH2PO415.5 mM, Na2HPO434.5 mM, imidazole 250 mM, pH 7.0). Protein concentration was determined by measuring absorbance at 280 nm. Protein purity was assessed using SDS-PAGE. Mobility Shift Assay. The purified GFP150-Tet ncAAs were diluted to 50 pM in PBS. The protein was reacted with excess sTCO-PEG5k (500 pM) for 5-10 minutes in PBS at room temperature. Protein was denatured through the addition of Laemmli buffer and heated at 95 °C for 10 minutes. Samples were then analyzed using a 12% SDS-PAGE gel.
[0136] Mass spectra of GFP-Tet ncAAs. Purified GFP-TAG150-Tet ncAAs were diluted to 50 pM and desalted using Zeba™ spin desalting column and analyzed using an FT LTQ mass spectrometer at the Oregon State University mass spectrometry facility. Waters SYNAPT G2 HDMS with a Waters Acquity I class UPLC mass spectrometer was used to verify reaction of purified GFP-TAG150-Tet ncAAs with sTCO. Samples were run 45 minutes gradient with H2O:ACN:0.1% formic acid using a Thermo Scientific- MAbPac™RP column (2.1x100 mm and a 0.2 ml / min flow rate). Spectra were deconvoluted using the Maximum Entropy deconvolution algorithm (MaxEnt3) in Waters MassLynx software.
[0137] Sequence Listing
[0138] Tet2 aaRS / tRNA pairs tRNA (DNA sequence)
[0139] TCCCGGCGGTAGTTCAGCAGGGCAGAACGGCGGACTCTAAATCCGCAT GGCGCTGGTTCAAATCCGGCCCGCCGGACCACTGCAGAT (SEQ ID NO:1)
[0140] E7RS (DNA Sequence)
[0141] ATGGACGAATTTGAAATGATAAAGAGAAACACATCTGAAATTATCAG CGAGGAAGAGTTAAGAGAGGTTTTAAAAAAAGATGAAAAATCTGCTGGGAT AGGTTTTGAACCAAGTGGTAAAATACATTTAGGGCATTATCTCCAAATAAAA AAGATGATTGATTTACAAAATGCTGGATTTGATATAATTATAGCTTTGGCTGA TTTAATGGCCTATTTAAACCAGAAAGGAGAGTTGGATGAGATTAGAAAAATA GGAGATTATAACAAAAAAGTTTTTGAAGCAATGGGGTTAAAGGCAAAATATG TTTATGGAAGTGAATCTGAGCTTGATAAGGATTATACACTGAATGTCTATAGA TTGGCTTTAAAAACTACCTTAAAAAGAGCAAGAAGGAGTATGGAACTTATAG
[0142] CAAGAGAGGATGAAAATCCAAAGGTTGCTGAAGTTATCTATCCAATAATGCA
[0143] GGTTAATGGTATTCATTATAATGGCGTTGATGTTGCAGTTGGAGGGATGGAGC
[0144] AGAGAAAAATACACATGTTAGCAAGGGAGCTTTTACCAAAAAAGGTTGTTTG
[0145] TATTCACAACCCTGTCTTAACGGGTTTGGATGGAGAAGGAAAGATGAGTTCTT
[0146] CAAAAGGGAATTTTATAGCTGTTGATGACTCTCCAGAAGAGATTAGGGCTAA
[0147] GATAAAGAAAGCATACTGCCCAGCTGGAGTTGTTGAAGGAAATCCAATAATG
[0148] GAGATAGCTAAATACTTCCTTGAATATCCTTTAACCATAAAAAGGCCAGAAA
[0149] AATTTGGTGGAGATTTGACAGTTAATAGCTATGAGGAGTTAGAGAGTTTATTT
[0150] AAAAATAAGGAATTGCATCCAATGGATTTAAAAAATGCTGTAGCTGAAGAAC
[0151] TTATAAAGATTTTAGAGCCAATTAGAAAGAGATTATAA (SEQ ID N0:2)
[0152] E7RS (Protein Sequence)
[0153] MDEFEMIKRNTSEIISEEELREVLKKDEKSAGIGFEPSGKIHLGHYLQIKKM
[0154] IDLQNAGFDIIIALADLMAYLNQKGELDEIRKIGDYNKKVFEAMGLKAKYVYGSE
[0155] SELDKDYTLNVYRLALKTTLKRARRSMELIAREDENPKVAEVIYPIMQVNGIHYN
[0156] GVDVAVGGMEQRKIHMLARELLPKKVVCIHNPVLTGLDGEGKMSSSKGNFIAV
[0157] DDSPEEIRAKIKKAYCPAGVVEGNPIMEIAKYFLEYPLTIKRPEKFGGDLTVNSYE
[0158] ELESLFKNKELHPMDLKNAVAEELIKILEPIRKRL* (SEQ ID NOG)
[0159] D12RS (DNA Sequence)
[0160] ATGGACGAATTTGAAATGATAAAGAGAAACACATCTGAAATTATCAG
[0161] CGAGGAAGAGTTAAGAGAGGTTTTAAAAAAAGATGAAAAATCTGCTGGGAT
[0162] AGGTTTTGAACCAAGTGGTAAAATACATTTAGGGCATTATCTCCAAATAAAA
[0163] AAGATGATTGATTTACAAAATGCTGGATTTGATATAATTATACAGTTGGCTGA
[0164] TTTACACGCCTATTTAAACCAGAAAGGAGAGTTGGATGAGATTAGAAAAATA
[0165] GGAGATTATAACAAAAAAGTTTTTGAAGCAATGGGGTTAAAGGCAAAATATG
[0166] TTTATGGAAGTGAATCTGATCTTGATAAGGATTATACACTGAATGTCTATAGA
[0167] TTGGCTTTAAAAACTACCTTAAAAAGAGCAAGAAGGAGTATGGAACTTATAG
[0168] CAAGAGAGGATGAAAATCCAAAGGTTGCTGAAGTTATCTATCCAATAATGCA GGTTAATTCTATTCATTATAATGGCGTTGATGTTGCAGTTGGAGGGATGGAGC
[0169] AGAGAAAAATACACATGTTAGCAAGGGAGCTTTTACCAAAAAAGGTTGTTTG TATTCACAACCCTGTCTTAACGGGTTTGGATGGAGAAGGAAAGATGAGTTCTT CAAAAGGGAATTTTATAGCTGTTGATGACTCTCCAGAAGAGATTAGGGCTAA GATAAAGAAAGCATACTGCCCAGCTGGAGTTGTTGAAGGAAATCCAATAATG GAGATAGCTAAATACTTCCTTGAATATCCTTTAACCATAAAAAGGCCAGAAA
[0170] AATTTGGTGGAGATTTGACAGTTAATAGCTATGAGGAGTTAGAGAGTTTATTT AAAAATAAGGAATTGCATCCAATGGATTTAAAAAATGCTGTAGCTGAAGAAC TTATAAAGATTTTAGAGCCAATTAGAAAGAGATTATAA (SEQ ID N0:4)
[0171] D12RS (Protein Sequence)
[0172] MDEFEMIKRNTSEIISEEELREVLKKDEKSAGIGFEPSGKIHLGHYLQIKKM
[0173] IDLQNAGFDIIIQLADLHAYLNQKGELDEIRKIGDYNKKVFEAMGLKAKYVYGSE S DLDKD YTLN V YRLALKTTLKR ARRSMELIAREDENPKV AEVI YPIMQ VNS IH YN GVDVAVGGMEQRKIHMLARELLPKKVVCIHNPVLTGLDGEGKMSSSKGNFIAV DDSPEEIRAKIKKAYCPAGVVEGNPIMEIAKYFLEYPLTIKRPEKFGGDLTVNSYE ELESLFKNKELHPMDLKNAVAEELIKILEPIRKRL* (SEQ ID N0:5)
[0174] Tet3 aaRS / tRNA pairs
[0175] DhPyl tRNA (DNA sequence) ggggggtggatcgaatagatcacacggactctaaatccgtgcaggcgggtgaaactcccgcaccccccg
[0176] (SEQ ID NO: 6)
[0177] MmPyl tRNA (DNA sequence) ggaaacctgatcatgtagatcgaacggactctaaatccgttcagccgggttagattcccggggtttccg
[0178] (SEQ ID NO:7)
[0179] NES-FLAG-R284 (DNA Sequence)
[0180] ATGGCGTGTCCGGTTCCTTTGCAGTTGCCTCCACTGGAGCGCCTCACAC
[0181] TCGACGACTACAAGGACGACGACGACAAGGACAAGAAACCCCTGGACGTGC TGATCAGCGCCACCGGCCTGTGGATGAGCCGGACCGGCACCCTGCACAAGAT CAAGCACCACGAGGTGTCAAGAAGCAAAATCTACATCGAGATGGCCTGCGGC GACCACCTGGTGGTGAACAACAGCAGAAGCTGCCGGACCGCCAGAGCCTTCC
[0182] GGCACCACAAGTACAGAAAGACCTGCAAGCGGTGCCGGGTGTCCGACGAGG
[0183] ACATCAACAACTTTCTGACCAGAAGCACCGAGAGCAAGAACAGCGTGAAAGT
[0184] GCGGGTGGTGTCCGCCCCCAAAGTGAAGAAAGCCATGCCCAAGAGCGTGTCC
[0185] AGAGCCCCCAAGCCCCTGGAAAACAGCGTGTCCGCCAAGGCCAGCACCAACA
[0186] CCAGCCGCAGCGTGCCCAGCCCCGCCAAGAGCACCCCCAACAGCTCCGTGCC
[0187] CGCCTCTGCTCCTGCTCCCAGCCTGACACGGTCCCAGCTGGACAGAGTGGAG
[0188] GCCCTGCTGTCCCCCGAGGACAAGATCAGCCTGAACATGGCCAAGCCCTTCC
[0189] GGGAGCTGGAACCCGAGCTGGTGACCCGGCGGAAGAACGACTTCCAGCGGCT
[0190] GTACACCAACGACCGGGAGGACTACCTGGGCAAGCTGGAACGGGACATCAC
[0191] CAAGTTCTTCGTGGACCGGGGCTTCCTGGAAATCAAGAGCCCCATCCTGATCC
[0192] CCGCCGAGTACGTGGAGCGGATGGGCATCAACAACGACACCGAGCTGTCCAA
[0193] GCAGATTTTCCGGGTGGACAAGAACCTGTGCCTGCGGCCTATGCTGGCCCCC
[0194] ACCGGCTACAACTACCTGCGGAAACTGGACAGAATCCTGCCTGGCCCCATCA
[0195] AGATTTTCGAAGTGGGACCCTGCTACCGGAAAGAGAGCGACGGCAAAGAGC
[0196] ACCTGGAAGAGTTTACAATGGTGGGCTTTGCCCAGATGGGCAGCGGCTGCAC
[0197] CCGGGAGAACCTGGAAGCCCTGATCAAAGAGTTCCTGGATTACCTGGAAATC
[0198] GACTTCGAGATCGTGGGCGACAGCTGCATGGTGTACGGCGACACCCTGGACA
[0199] TCATGCACGGCGACCTGGAACTGAGCAGCGCCGTGGTGGGACCCGTGTCCCT
[0200] GGACCGGGAGTGGGGCATCGACAAGCCCTGGATCGGAGCCGGCTTCGGCCTG
[0201] GAACGGCTGCTGAAAGTGATGCACGGCTTCAAGAACATCAAGCGGGCCAGCA
[0202] GAAgcgagagctactacaacggcatcagcaccaacctgtga (SEQ ID N0:8)
[0203] NES-FLAG-R284 (Protein Sequence)
[0204] MACPVPLQLPPLERLTLDDYKDDDDKDKKPLDVLISATGLWMSRTGTLH
[0205] KIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDI
[0206] NNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRS
[0207] VPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVT
[0208] RRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINND TELSKQIFRVDKNLCLRPMLAPTGYNYLRKLDRILPGPIKIFEVGPCYRKESDGKE
[0209] HLEEFTMVGFAQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMH
[0210] GDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESY
[0211] YNGISTNL* (SEQ ID NO: 9)
[0212] NES-FLAG-R274 (DNA Sequence)
[0213] ATGGCGTGTCCGGTTCCTTTGCAGTTGCCTCCACTGGAGCGCCTCACAC
[0214] TCGACGACTACAAGGACGACGACGACAAGGACAAGAAACCCCTGGACGTGC
[0215] TGATCAGCGCCACCGGCCTGTGGATGAGCCGGACCGGCACCCTGCACAAGAT
[0216] CAAGCACCACGAGGTGTCAAGAAGCAAAATCTACATCGAGATGGCCTGCGGC
[0217] GACCACCTGGTGGTGAACAACAGCAGAAGCTGCCGGACCGCCAGAGCCTTCC
[0218] GGCACCACAAGTACAGAAAGACCTGCAAGCGGTGCCGGGTGTCCGACGAGG
[0219] ACATCAACAACTTTCTGACCAGAAGCACCGAGAGCAAGAACAGCGTGAAAGT
[0220] GCGGGTGGTGTCCGCCCCCAAAGTGAAGAAAGCCATGCCCAAGAGCGTGTCC
[0221] AGAGCCCCCAAGCCCCTGGAAAACAGCGTGTCCGCCAAGGCCAGCACCAACA
[0222] CCAGCCGCAGCGTGCCCAGCCCCGCCAAGAGCACCCCCAACAGCTCCGTGCC
[0223] CGCCTCTGCTCCTGCTCCCAGCCTGACACGGTCCCAGCTGGACAGAGTGGAG
[0224] GCCCTGCTGTCCCCCGAGGACAAGATCAGCCTGAACATGGCCAAGCCCTTCC
[0225] GGGAGCTGGAACCCGAGCTGGTGACCCGGCGGAAGAACGACTTCCAGCGGCT
[0226] GTACACCAACGACCGGGAGGACTACCTGGGCAAGCTGGAACGGGACATCAC
[0227] CAAGTTCTTCGTGGACCGGGGCTTCCTGGAAATCAAGAGCCCCATCCTGATCC
[0228] CCGCCGAGTACGTGGAGCGGATGGGCATCAACAACGACACCGAGCTGTCCAA
[0229] GCAGATTTTCCGGGTGGACAAGAAcctgtgcctgcggcctatgctggcccccaccGGCtacaactacc tgcggaaactggacagaatcctgcctggccccatcaagattttcgaagtgggaccctgctaccggaaagagagcgacggcaaa gagcacctggaagagtttacaatggtgGGCtttAGCcagatgggcagcggctgcACCCGGGAGAACCTGG
[0230] AAGCCCTGATCAAAGAGTTCCTGGATTACCTGGAAATCGACTTCGAGATCGT
[0231] GGGCGACAGCTGCATGGTGTACGGCGACACCCTGGACATCATGCACGGCGAC
[0232] CTGGAACTGAGCAGCGCCGTGGTGGGACCCGTGTCCCTGGACCGGGAGTGGG
[0233] GCATCGACAAGCCCTGGATCGGAGCCGGCTTCGGCCTGGAACGGCTGCTGAA AGTGATGCACGGCTTCAAGAACATCAAGCGGGCCAGCAGAAgcgagagctactacaac ggcalcagcaccaacctgtga (SEQ ID NO: 10)
[0234] NES-FLAG-R274 (Protein Sequence)
[0235] MACPVPLQLPPLERLTLDDYKDDDDKDKKPLDVLISATGLWMSRTGTLH
[0236] KIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDI
[0237] NNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRS
[0238] VPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVT
[0239] RRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINND
[0240] TELSKQIFRVDKNLCLRPMLAPTGYNYLRKLDRILPGPIKIFEVGPCYRKESDGKE
[0241] HLEEFTMVGFSQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMH
[0242] GDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESY
[0243] YNGISTNL* (SEQ ID NO: 11)
[0244] NES-R284 (DNA Sequence)
[0245] ATGGCGTGTCCGGTTCCTTTGCAGTTGCCTCCACTGGAGCGCCTCACAC
[0246] TCGACGACAAGAAACCCCTGGACGTGCTGATCAGCGCCACCGGCCTGTGGAT
[0247] GAGCCGGACCGGCACCCTGCACAAGATCAAGCACCACGAGGTGTCAAGAAG
[0248] CAAAATCTACATCGAGATGGCCTGCGGCGACCACCTGGTGGTGAACAACAGC
[0249] AGAAGCTGCCGGACCGCCAGAGCCTTCCGGCACCACAAGTACAGAAAGACCT
[0250] GCAAGCGGTGCCGGGTGTCCGACGAGGACATCAACAACTTTCTGACCAGAAG
[0251] CACCGAGAGCAAGAACAGCGTGAAAGTGCGGGTGGTGTCCGCCCCCAAAGT
[0252] GAAGAAAGCCATGCCCAAGAGCGTGTCCAGAGCCCCCAAGCCCCTGGAAAA
[0253] CAGCGTGTCCGCCAAGGCCAGCACCAACACCAGCCGCAGCGTGCCCAGCCCC
[0254] GCCAAGAGCACCCCCAACAGCTCCGTGCCCGCCTCTGCTCCTGCTCCCAGCCT
[0255] GACACGGTCCCAGCTGGACAGAGTGGAGGCCCTGCTGTCCCCCGAGGACAAG
[0256] ATCAGCCTGAACATGGCCAAGCCCTTCCGGGAGCTGGAACCCGAGCTGGTGA
[0257] CCCGGCGGAAGAACGACTTCCAGCGGCTGTACACCAACGACCGGGAGGACTA
[0258] CCTGGGCAAGCTGGAACGGGACATCACCAAGTTCTTCGTGGACCGGGGCTTC
[0259] CTGGAAATCAAGAGCCCCATCCTGATCCCCGCCGAGTACGTGGAGCGGATGG GCATCAACAACGACACCGAGCTGTCCAAGCAGATTTTCCGGGTGGACAAGAA cctgtgcctgcggcctatgctggcccccaccGGCtacaactacctgcggaaactggacagaatcctgcctggccccatcaag attttcgaagtgggaccctgctaccggaaagagagcgacggcaaagaGCACCTGGAAGAGTTTACAATG GTGGGCTTTGCCCAGATGGGCAGCGGCTGCACCCGGGAGAACCTGGAAGCCC TGATCAAAGAGTTCCTGGATTACCTGGAAATCGACTTCGAGATCGTGGGCGA CAGCTGCATGGTGTACGGCGACACCCTGGACATCATGCACGGCGACCTGGAA CTGAGCAGCGCCGTGGTGGGACCCGTGTCCCTGGACCGGGAGTGGGGCATCG ACAAGCCCTGGATCGGAGCCGGCTTCGGCCTGGAACGGCTGCTGAAAGTGAT GCACGGCTTCAAGAACATCAAGCGGGCCAGCAGAAgcgagagctactacaacggcatcagc accaacctgtga (SEQ ID NO: 12)
[0260] NES-R284 (Protein Sequence)
[0261] MACPVPLQLPPLERLTLDDKKPLDVLISATGLWMSRTGTLHKIKHHEVSR SKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTE SKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKSTPN SSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRRKNDFQRL YTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINNDTELSKQIFRV DKNLCLRPMLAPTGYNYLRKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMVG FAQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELSSAV VGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL* (SEQ ID NO: 13)
[0262] Ma / RNA (top strand) for Tet3.0-butyl encoding with MaB5RS (DNA Sequence) 5’_AGATCTGGGGGACGGTCCGGCGACCAGCGGGTCTCTAAAACCTAGCATAG CGGGGTTCGACaCCCCGGTCTCTCG_3’ (SEQ ID NO: 14)
[0263] DNA sequence (top strand) for AfaB5RS selected for Tet3.0-butyl encoding
[0264] 5’_ATGACAGTGAAATACACAGATGCCCAGATCCAGCGCCTGCGGGAG TATGGCAACGGCACCTATGAGCAGAAGGTGTTTGAAGATCTGGCCTCTAGAG ATGCAGCCTTCTCCAAGGAGATGTCCGTGGCTTCCACAGACAACGAGAAAAA GATCAAGGGCATGATTGCCAACCCCAGCCGCCATGGGCTGACCCAGCTGATG AATGACATCGCCGACGCCCTGGTGGCCGAGGGCTTCATCGAGGTCAGGACCC CCATCTTCATTTCTAAGGACGCGCTGGCTCGCATGACCATCACCGAGGACAA GCCCCTGTTCAAGCAGGTGTTCTGGATCGATGAGAAGAGGGCTCTGAGGCCC ATGCTCGCCCCCAACCATTACTCCGTGATGCGGGACCTGcgcGACCACACGGA CGGCCCTGTGAAAATTTTCGAAATGGGCTCCTGCTTTAGGAAAGAAAGCCAC AGCGGAATGCACCTGGAGGAGTTCACCATGCTGAATCTGGCGGACATGGGGC CAAGAGGAGATGCCACAGAAGTGCTGAAGAACTACATCTCAGTGGTCATGAA GGCTGCTGGACTGCCCGACTATGATTTGGTGCAGGAAGAGAGCGATGTCTAC AAAGAaACCATTGATGTGGAGATCAATGGCCAGGAGGTGTGCTCTGCTGCGG TGGGCCCCCACTACCTGGACGCCGCCCACGACGTGCATGAACCCTGGAGTGG AGCGGGCTTTGGCCTGGAGAGGCTGCTGACCATAAGAGAAAAGTACAGCACT GTGAAGAAAGGCGGCGCCTCCATCTCCTACTTGAATGGAGCCAAGATCAACA
[0265] GCGGCTGA_3’ (SEQ ID NO: 15)
[0266] Amino acid sequence of AfaB5RS selected for Tet3.0-butyl encoding
[0267] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEK
[0268] KIKGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLF KQVFWIDEKRALRPMLAPNHYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMH LEEFTMLNLADMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETID
[0269] VEINGQEVCSAAVGPHYLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGAS
[0270] ISYLNGAKINSG* (SEQ ID NO:16)
[0271] Ma / R A sequence (top strand) for Tet3.0-butyl encoding with AfaF9RS (DNA Sequence)
[0272] 5 ’ _AGATCTGGGGG ACGGTCCGGCGACC AGCGGGTCTCT A A A ACCTAGC AT AG
[0273] CGGGGTTCGACACCCCGGTCTCTCG_3’ (SEQ ID NO: 17)
[0274] DNA sequence (top strand) for AfaF9RS selected for Tet3.0-butyl encoding
[0275] 5’_ATGACAGTGAAATACACAGATGCCCAGATCCAGCGCCTGCGGGAG
[0276] TATGGCAACGGCACCTATGAGCAGAAGGTGTTTGAAGATCTGGCCTCTAGAG ATGCAGCCTTCTCCAAGGAGATGTCCGTGGCTTCCACAGACAACGAGAAAAA GATCAAGGGCATGATTGCCAACCCCAGCCGCCATGGGCTGACCCAGCTGATG
[0277] AATGACATCGCCGACGCCCTGGTGGCCGAGGGCTTCATCGAGGTCAGGACCC
[0278] CCATCTTCATTTCTAAGGACGCGCTGGCTCGCATGACCATCACCGAGGACAA
[0279] GCCCCTGTTCAAGCAGGTGTTCTGGATCGATGAGAAGAGGGCTCTGAGGCCC
[0280] ATGCTCGCCCCCAACTTTTACTCCGTGATGCGGGACCTGcgcGACCACACGGAC
[0281] GGCCCTGTGAAAATTTTCGAAATGGGCTCCTGCTTTAGGAAAGAAAGCCACA
[0282] GCGGAATGCACCTGGAGGAGTTCACCATGCTGAATCTGGCTGACATGGGGCC
[0283] AAGAGGAGATGCCACAGAAGTGCTGAAGAACTACATCTCAGTGGTCATGAAG
[0284] GCTGCTGGACTGCCCGACTATGATTTGGTGCAGGAAGAGAGCGATGTCTACA
[0285] AAGAaACCATTGATGTGGAGATCAATGGCCAGGAGGTGTGCTCTGCTGCGGT
[0286] GGGCCCCCACTACCTGGACGCCGCCCACGACGTGCATGAACCCTGGAGTGGA
[0287] GCGGGCTTTGGCCTGGAGAGGCTGCTGACCATAAGAGAAAAGTACAGCACTG
[0288] TGAAGAAAGGCGGCGCCTCCATCTCCTACTTGAATGGAGCCAAGATCAACAG
[0289] CGGCTGA_3’ (SEQ ID NO: 18)
[0290] Amino acid sequence of AfaF9RS selected for Tet3.0-butyl encoding
[0291] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEK
[0292] KIKGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLF
[0293] KQVFWIDEKRALRPMLAPNFYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHL
[0294] EEFTMLNLADMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDV
[0295] EINGQEVCSAAVGPHYLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASIS
[0296] YLNGAKINSG* (SEQ ID NO: 19)
[0297] Ma / RNA sequence (top strand) for Tet3.0-H encoding with AfaG9RS (DNA
[0298] Sequence)
[0299] 5’_AGATCTGGGGGACGGTCCGGCGACCAGCGGGTCTCTAAAACCTAGC
[0300] ATAGCGGGGTTCGACACCCCGGTCTCTCG_3’ (SEQ ID NO:20)
[0301] DNA sequence (top strand) for AA / G9RS selected for Tet3.0-H encoding
[0302] 5’_ATGACAGTGAAATACACAGATGCCCAGATCCAGCGCCTGCGGGAG
[0303] TATGGCAACGGCACCTATGAGCAGAAGGTGTTTGAAGATCTGGCCTCTAGAG ATGCAGCCTTCTCCAAGGAGATGTCCGTGGCTTCCACAGACAACGAGAAAAA GATCAAGGGCATGATTGCCAACCCCAGCCGCCATGGGCTGACCCAGCTGATG
[0304] AATGACATCGCCGACGCCCTGGTGGCCGAGGGCTTCATCGAGGTCAGGACCC
[0305] CCATCTTCATTTCTAAGGACGCGCTGGCTCGCATGACCATCACCGAGGACAA GCCCCTGTTCAAGCAGGTGTTCTGGATCGATGAGAAGAGGGCTCTGAGGCCC ATGCTCGCCCCCAACTTTTACTCCGTGATGCGGGACCTGcgcGACCACACGGAC GGCCCTGTGAAAATTTTCGAAATGGGCTCCTGCTTTAGGAAAGAAAGCCACA
[0306] GCGGAATGCACCTGGAGGAGTTCACCATGCTGAATCTGTGTGACATGGGGCC AAGAGGAGATGCCACAGAAGTGCTGAAGAACTACATCTCAGTGGTCATGAAG
[0307] GCTGCTGGACTGCCCGACTATGATTTGGTGCAGGAAGAGAGCGATGTCTACA AAGAaACCATTGATGTGGAGATCAATGGCCAGGAGGTGTGCTCTGCTGCCGT
[0308] GGGCCCCCACTACCTGGACGCCGCCCACGACGTGCATGAACCCTGGAGTGGA GCGGGCTTTGGCCTGGAGAGGCTGCTGACCATAAGAGAAAAGTACAGCACTG
[0309] TGAAGAAAGGCGGCGCCTCCATCTCCTACTTGAATGGAGCCAAGATCAACAG
[0310] CGGCTGA_3’ (SEQ ID N0:21)
[0311] Amino acid sequence of AfaG9RS selected for Tet3.0-H encoding
[0312] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEK
[0313] KIKGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLF KQVFWIDEKRALRPMLAPNFYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHL EEFTMLNLCDMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDV EINGQEVCSAAVGPHYLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASIS YLNGAKINSG* (SEQ ID NO:22)
[0314] While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the disclosure.
Claims
CLAIMSThe embodiments of the disclosure in which an exclusive property or privilege is claimed are defined as follows:
1. A tetrazine amino acid having formula (I) or (II):or a stereoisomer or salt thereof, whereinR is selected from the group consisting of:(a) a phenyl group substituted with a group selected from O-C1-C4 alkyl, m- C1-C4 alkyl, O-C1-C3 alkoxy, m-Cl-C3 alkoxy, o-cyano, m-cyano, o-nitro, m-nitro, o- or m-primary amino (-NH2), secondary amino (-NHRX), or tertiary amino (-NHRxRy) (wherein Rxand Ryare independently C1-C6 alkyl), o-fluoro, m- fluoro, 3,5-difluoro, 3,4,5- trifluoro, o-trifluoromethyl, m-trifluoromethyl, and o-, m-, or p-C(=O)Rz(wherein Rza counterion, hydrogen, or C1-C6 alkyl),(b) a substituted or an unsubstituted heteroaryl group,(c) a substituted or an unsubstituted heterocyclyl group,(d) an amino Cl -C6 alkyl group,(e) a thio C1-C6 alkyl group,(f) a carboxylate group,(g) a sulfonate group, and(h) an amide group;Rcis hydrogen, a counter ion, or a carboxyl protecting group; andRNis hydrogen or an amine protecting group.
2. The tetrazine amino acid of Claim 1, wherein the alkyl group is selected from methyl, ethyl, n-propyl, i-propyl, n-butyl, s-butyl, t-butyl, n-pentyl, and n-hexyl.
3. The tetrazine amino acid of Claim 1 , wherein the phenyl group is substituted with a group selected from the group consisting of 3-methyl, 2-methoxy, 3-methoxy, 2- cyano, 3-cyano, 2-nitro, 3-nitro, 2-fluoro, 3-fluoro, 3,5-difluoro, 3,4,5-trifluoro, 2-hydroxy, 3-hydroxy, 2-amino, 3-amino, 2-acetyl, 3-acetyl, 2-trifluoromethyl, and 3 -trifluoromethyl.
4. The tetrazine amino acid of Claim 1, wherein the heteroaryl group is a pyridyl group, a furanyl group, a thiophenyl group, or an oxazolyl group.
5. The tetrazine amino acid of Claim 1, wherein the heterocyclyl group is a pyrrolidinyl group, a piperadinyl group, a piperazinyl group, a tetrahydrofuranyl group, a tetrahydropyranyl, a tetrahydrothiophenyl group, an oxazolydinyl group, or a dihydropyran group.
6. The tetrazine amino acid of Claim 1, wherein the amino C1-C6 alkyl group is an aminobutyl group.
7. The tetrazine amino acid of Claim 1, wherein the thio C1-C6 alkyl group is a thiomethyl group.
8. The tetrazine amino acid of Claim 1, wherein the carboxylate group is a Cl- C6 alkyl carboxylate group (e.g., methyl carboxylate (-CO2CH3), ethyl carboxylate (- CO2CH2CH3)).
9. The tetrazine amino acid of Claim 1, wherein the amide group is a C1-C6 alkyl amide group (e.g., methyl amide (-C(=O)NHCH3), butyl amide (- C(=O)NH(CH2)3CH3)).
10. A method for making a protein or a polypeptide of interest, comprising: incorporating a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof, into a protein or polypeptide.
11. A method for genetically encoding a protein or a polypeptide of interest, comprising:incorporating a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof, into a protein or polypeptide by genetic encoding.
12. A protein or polypeptide, comprising at least one tetrazine amino acid residue, wherein the tetrazine amino acid residue is derived from a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof.
13. A protein or polypeptide, comprising at least one tetrazine amino acid residue, wherein the tetrazine amino acid residue is incorporated into the protein or polypeptide by genetic encoding of the protein or polypeptide using a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof.
14. A composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof, and wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid comprising the first reactive group.
15. A composition comprising a protein or polypeptide, wherein the protein or polypeptide comprises at least one tetrazine amino acid comprising a first reactive group and at least one post-translational modification, wherein the tetrazine amino acid residue is derived from genetic encoding of the protein or polypeptide using a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof, and wherein the at least one post-translational modification comprises attachment of a molecule comprising a second reactive group by a [4+2] cycloaddition reaction to the at least one tetrazine amino acid comprising the first reactive group.
16. A kit for in cellulo production of a tetrazine-labeled protein or a tetrazinelabeled polypeptide, comprising:(a) a tRNA;(b) an aminoacyl-tRNA synthetase; and(c) a tetrazine amino acid of any one of Claims 1-9, or a stereoisomer or salt thereof, wherein the tRNA and aminoacyl-tRNA synthetase are an orthogonal tRNA / orthogonal aminoacyl-tRNA pair effective for incorporating the compound into a protein or polypeptide to provide a tetrazine-labeled protein.
Citation Information
Patent Citations
Reagents and methods for bioorthogonal labeling of biomolecules in living cells
US20190077776A1
Anti-CD74 antibody conjugates, compositions comprising Anti-CD74 antibody conjugates and methods of using Anti-CD74 antibody conjugates
US20190144546A1
Bio-orthogonal drug activation
US20210162060A1
Amino acids bearing a tetrazine moiety
WO2022248587A1
Hydrophilic tetrazine-functionalized payloads for preparation of targeting conjugates
WO2023104941A1