Orthogonal translation systems and methods for identifying components thereof
Patent Information
- Application Number
- PCT/US2025/010451
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2025-01-06
- Publication Date
- 2026-01-29
AI Technical Summary
The identification and discovery of orthogonal aminoacyl-tRNA synthetases (aaRS) and tRNAs for incorporating non-canonical amino acids (ncAAs) into peptides and proteins is a bottleneck, limiting the range of chemistries that can be incorporated, with existing methods heavily relying on Methanocaldococcus jannaschii TyrRS:tRNA and PylRS:tRNA pairs.
Development of an orthogonal translation system (OTS) comprising a polynucleotide encoding a tRNA and an aminoacyl-tRNA synthetase with specific sequence identities, capable of incorporating noncanonical amino acids at UGA or UAG codons in bacteria, using a high-throughput cell-free protein synthesis system to identify and characterize tRNA and aaRS pairs.
Enables the incorporation of noncanonical amino acids into proteins by identifying functional OTSs, expanding the range of chemistries that can be used in genetic code expansion and providing a powerful tool for protein engineering.
Smart Images

Figure US2025010451_29012026_PF_FP_ABST
Abstract
Description
ORTHOGONAL TRANSLATION SYSTEMS AND METHODS FORIDENTIFYING COMPONENTS THEREOFCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 618,182 filed on January’ 5, 2024. The content of which is incorporated by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under grant number AI092531 awarded by the National Institutes of Health, and grant number W911NF-18-1-0200 awarded by the Department of Defense. The government has certain rights in the invention.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] The application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on January' 6, 2025, is named “702581.02600.xml” and is 733,325 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety.BACKGROUND
[0004] Orthogonal aminoacyl-tRNA synthetases (aaRS) and tRNAs are essential tools for expanding the genetic code and incorporating non-canonical amino acids (ncAAs). However, the identification and discovery of these translation systems is a bottleneck that limits the range of chemistries that can be incorporated into peptides and proteins. Most ncAAs that have been successfully incorporated are mostly chemical derivatives of tyrosine and lysine. This reflects the fact that the genetic code expansion field has heavily relied on the Methanocaldococcus jannaschii TyrRS:tRNA and the PylRS:tRNA pairs for ncAA incorporation. Incorporating other unnatural chemistries into proteins using orthogonal tRNAs and aminoacyl-tRNA synthetases remains a central challenge in synthetic biology.SUMMARY
[0005] In an aspect, provided herein is an orthogonal translation system in a bacteria, the system comprising: a polynucleotide encoding a tRNA, the polynucleotide having at least 80 % sequence identity to the polynucleotide sequence of SEQ ID NO: 1; and an aminoacyl-tRNA synthetase having at least 80% sequence identity to the polypeptide sequence of SEQ ID NO: 2, or a polynucleotide encoding the aminoacyl-tRNA synthetase. In embodiments, the bacteria is E. coll.
[0006] The polynucleotide encoding the tRNA; and the aminoacyl-tRNA synthetase or the polynucleotide encoding the aminoacyl-tRNA synthetase may be present in an E. coll lysate or an E. coll cell.
[0007] In embodiments, the tRNA and the aminoacyl-tRNA synthetase can incorporate a tryptophan residue at a UGA codon. In embodiments, the tRNA and the aminoacyl-tRNA synthetase can incorporate a noncanonical amino acid at a UGA codon.
[0008] The polynucleotide encoding the tRNA may have at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the polynucleotide sequence of SEQ ID NO: 1. The aminoacyl- tRNA synthetase may have at least 85%. at least 90%, at least 95%, or at least 99% sequence identity to the polypeptide sequence of SEQ ID NO: 2.
[0009] In another aspect, provided herein is a method for incorporating an amino acid at a UGA stop codon, the method comprising: providing a translation template to any of the orthogonal translation systems described herein. The translation template may be expressed from a transcription template in an E. coll lysate. The E. coll lysate may be part of a cell-free protein synthesis system.
[0010] In another aspect, provided herein is an orthogonal translation system in a bacteria, the system comprising: a polynucleotide encoding a tRNA, the polynucleotide having at least 80% sequence identity to the polynucleotide sequence of SEQ ID NO: 3; and an aminoacyl-tRNA synthetase having at least 80% sequence identity to the polypeptide sequence of SEQ ID NO: 4, or a polynucleotide encoding the aminoacyl-tRNA synthetase. In embodiments, the bacteria is E. coll.
[0011] The polynucleotide encoding the tRNA and the aminoacyl-tRNA synthetase or the polynucleotide encoding the aminoacyl-tRNA synthetase may be present in an E. coll lysate or E. coll cell.
[0012] In embodiments, the tRNA and the aminoacyl-tRNA synthetase can incorporate a glutamine residue at a UAG codon. In embodiments, the tRNA and the aminoacyl-tRNA synthetase can incorporate a noncanonical amino acid at a UAG codon.
[0013] The polynucleotide encoding the tRNA may have at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the polynucleotide sequence of SEQ ID NO: 3. The aminoacyl- tRNA synthetase may have at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the polypeptide sequence of SEQ ID NO: 4.
[0014] In another aspect, provided herein is a method for incorporating an amino acid at a UAG stop codon, the method comprising: providing a translation template to any of the orthogonal translation systems described herein. The translation template may be expressed from a transcription template in an / ■’ coli lysate. The E. coll lysate may be part of a cell-free protein synthesis system.
[0015] In another aspect, provided herein is a kit for identifying a candidate orthogonal tRNA to an organism, the kit comprising: one or more transcription templates, wherein each of the one or more transcription templates comprises a polynucleotide encoding a suppressor tRNA; a cell-free protein synthesis (CFPS) system derived from the organism; and a translation template encoding a reporter protein, wherein the translation template comprises a premature stop codon. The reporter protein may be 216X-sfGFP, wherein X is UAG, UAA, or UAG. The suppressor tRNA may comprise an RNase P tag; and wherein the CFPS comprises RNase P. The CFPS system may comprise an E. coli lysate.
[0016] In another aspect, provided herein is a method for identifying a candidate orthogonal tRNA to an organism, the method comprising: transcribing one or more transcription templates in vitro to produce one or more transcription products, wherein each of the one or more transcription templates comprises a polynucleotide encoding a premature suppressor tRNA; and incubating the one or more transcription products with a cell-free protein synthesis (CFPS) system derived from the organism, and a translation template encoding a reporter protein having a premature stop codon; performing an assay that detects the reporter protein; wherein the CFPS system comprises RNase P; and wherein if the reporter protein is not detected, the suppressor tRNA is a candidate orthogonal tRNA to the organism. Each of the one or more transcription templates may be linear.
[0017] Transcribing the one or more transcription templates may be performed with a T7 RNA polymerase. The reporter protein may be 216X-sfGFP, wherein X is UAG, UAA, or UAG.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The patent or patent application file contains at least one drawing in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.
[0019] FIGS. 1 A-1C. tRNAs can be expressed, matured, and evaluated with an in vitro expression and maturation process. (A) Schematic of decoupled expression and maturation process. tRNAs are purchased as short linear DNA sequences, transcribed in vitro with T7 RNA polymerase, and matured into functional tRNAs by endogenous RNase P in cell extracts. The mature tRNA can be incorporated into a cell-free protein synthesis reaction to assess activity7and orthogonality. (B) The M. cilvus tRNAPylcuA can be cleaved into a mature tRNA by the action of endogenous RNase P in cell extracts. (C) Evaluation of a panel of metagenomically-identified tRNAscuv shows efficient expression of all tRNAs when using the RNase P tag compared to plasmid-based expression and constructs lacking the RNase P tag.
[0020] FIGS. 2A-2B. A high-throughput screen of bioinformatically identified suppressor tRNAs. (A) Overall workflow for evaluation of suppressor tRNAs uses a decoupled tRNA expression platform to generate a premature tRNA. The tRNA is then processed and evaluated in CFPS through endogenous RNase P and through suppression of a stop codon in 216X-sfGFP, respectively. (B) A panel of suppressor tRNAs can be tested for activity7against each stop codon. Suppressor tRNAs are grouped by codon along with internal controls of the M. jannaschii tRNATyr, the AT barkeri tRNAPyl. the AT. alvus tRNAPyl, and the A17 VC10 Int tRNAPyl for each codon.
[0021] FIGS. 3A-3B. tRNA orthogonality7and activity7can be robustly detected. (A) Cell-free reactions do not synthesize 216UAG-sfGFP in the presence of an orthogonal tRNA or in the absence of a suppressor tRNA. The M.jannaschii tRNATyrCUA is orthogonal to the chimeric PylRS (IPYE) and therefore is not able to readthrough a premature stop codon. M.j is M. j annaschii and M.b is M. barkeri. Data is from a high-throughput screen in which n = 1 but is representative of data from independent experiments (B) T216 of sfGFP is robust to all 20 canonical amino acid mutations Each bar represents the average of n = 3 data points and the error bar represents standard deviation.
[0022] FIG. 4. Representative ESI-MS of purified 216X-sfGFP purified proteins from CFPS reactions with non-orthogonal tRNAs. Reactions using tRNAscuv are shown in blue. 28 / 29 show either a Gin or Lys incorporation in 216UAG-sfGFP, while 1 / 29 shows incorporation of Ala at UAG. Reactions using tRNAsucA are shown in gold, and all 216UGA-sfGFP show incorporation of Trp.
[0023] FIGS. 5A-5C. Analysis of products from tRNA expression and maturation support conclusion of orthogonality. (A) DNA templates were correctly amplified from commercial templates. (B) in vitro transcription products are correctly synthesized from PCR-amplified templates. (C) In vitro transcription products are processed by RNase P to remove the tag.
[0024] FIGS. 6A-6D. Characterization and discovery of orthogonal TrpRS:tRNA systems. (A) A putative TrpRS found in phage genomes shows strong structural homology to E. coli TrpRS (B) The tRNA specificity of these putative TrpRSs are not mutually orthogonal. (C) These TrpRSs enable incorporation of Trp at the UGA codon, confirming their activity. (D) Identification of OTSs. in vitro aminoacylation reactions highlight orthogonality of aaRSs. The API TrpRS:tRNATrpUCA pairs appear functionally orthogonal in E. coli, while the HF2 TrpRS:tRNATrpUCA pair is not orthogonal in E. coh. n.d. = not detected. Control is E. coli TrpRS with total E. coli tRNA. Spectra were collected twice; a representative plot is shown.
[0025] FIG. 7. Comparison of tRNA secondary structures (SEQ ID NOs: 454-457) as predicted by R2DT.1Identity elements for A. coli tRNATrpccA are highlighted in red, and the corresponding nucleotides for the API, HF2, and TGA-29 tRNAsucA are highlighted in red.
[0026] FIG. 8. HF2 TrpRS recognizes E. coli tRNATroccA nonspecifically.
[0027] FIGS. 9A-9D. Orthogonal aaRS and tRNA pairs can be identified from metagenomic data. (A) Screening a panel of 16 aaRSs found in metagenomic data against orthogonal tRNAscuA identifies several active aaRS: tRNA pairs. Data shown here are curated from the full dataset presented in Figure 10A-10B (B) aaRS activity7is confirmed by analyzing purified 216UAG-sfGFP from CFPS reactions. GlnRS enables incorporation of either Gin or Lys and TyrRS enables incorporation of Tyr. (C)TyrRS strongly recognizes total E. coli tRNA, disqualifying it from being a functional OTS. (D) GlnRS is weakly active against total E. coli tRNA, suggesting that it could be an effective OTS.
[0028] FIGS. 10A-10B. Full screening results from CFPS reactions to identify functional aaRS:tRNA pairs using metagenomically-identified aaRSs and orthogonal tRNAscuA.
[0029] FIG. 11. Screening results from CFPS reactions to identify7functional aaRS:tRNA pairs using metagenomically-identified aaRSs and orthogonal tRNAsuuA.
[0030] FIG. 12. Screening results from CFPS reactions to identify functional aaRS:tRNA pairs using metagenomically-identified aaRSs and orthogonal tRN ASUCA.
[0031] FIG. 13. A library of suppressor tRNA sequences from metagenomic data provides candidate orthogonal tRNA sequences. A phylogenetic tree of suppressor tRNA sequences was built. Colors denote predicted codon for each tRNA. and shapes indicate confidence scores. Specific tRNA sequences highlighted with red circle on the branches are those identified to be orthogonal based on FIG. 2B. M.j = M. jannaschii, M.a = M. alvus and M.b = M. barkeri.DETAILED DESCRIPTION
[0032] Orthogonal translation systems (OTSs). kits, and methods for using the OTSs and kits are disclosed, in various aspects. In certain aspects, systems and methods are disclosed for identifying one or more components of OTSs. The orthogonal translation systems can include orthogonal tRNA and aminoacyl-tRNA sy nthetase pairs that can add an amino acid at the UGA or UAG stop codon, in various aspects. The methods for identifying one or more components of OTSs can include identifying candidate orthogonal tRNAs and candidate orthogonal aminoacyl-tRNA synthetases in vitro.
[0033] As described herein, a high-throughput, cell-free method was developed to characterize the activity and orthogonality of putative tRNAs and aaRSs responsible for stop codon reassignment from phage genomes to be used in E. coli systems. A set of >200 suppressor tRNAs for all three stop codons were curated and characterized for their activity and orthogonality in E. coli cell lysates. A panel of aaRSs were characterized and thousands of combinations of aaRS:tRNA pairs were tested. As described herein, the first natural OTS capable of reassigning the UGA codon as Trp in E. coli is disclosed, as well as a novel GlnRS:tRNA pair.
[0034] Orthogonal translation systems, orthogonal tRNAs. and aminoacyl tRNA synthetases
[0035] In a first aspect, orthogonal translation systems (OTSs) are disclosed. The OTSs can include one or more tRNAs and / or one or more aminoacyl-tRNA synthetases. The OTSs may be orthogonal to a bacteria. The OTSs may be orthogonal to E. coli, E. coli extracts, and / or E. coli lysates. The OTSs may be orthogonal to B. subtilis, B. subtilis extracts, and / or B. subtilis lysates.
[0036] In embodiments, the one or more tRNAs are encoded by a polynucleotide having at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% sequence identify to the polynucleotide sequence of SEQ ID NO: 1. The one or more tRNAs may be encoded by a polynucleotide that consists essentially of, or consists of the polynucleotide of SEQ ID NO: 1. (CCTGGAGTAGCGCAATTGGAAGAGCCGCGGTCTTCAAAACCGAGAGTTGTGAGTTC GAGTCTTACCTCCAGGGCCA).
[0037] In embodiments, the one or more aminoacyl-tRNA synthetases has at least 80%, at least 85%, at least 90%, at least 95%. at least 99%, or 100% sequence identify to the polypeptide sequence of SEQ ID NO: 2(MKERMLTGIKPTGNSITLGNYIGGLLPLIKYQDQFDLFLFVADLHALTVYQKDLTLGNN IENLVATYLAAGIDPNKVTIFKQSEIPEHTQLEWVLTCTTDLPDLLKMPQYKNYKEINKN KAVPAGMLMYPSLMNADILLYNTDYIPVGIDQKPHVNLCHDIAMKFNARYGETFKIPKP IVPETGAKIMSLTTPTKKMSKSESDNGTIYLLEDVEITRRKIMKAITDSENKVYFNPETKP GVSNLLSIYSALSEIPISELEEKYANTSNYGVFKKDLADLVCDKMAQIQARIKRIKELGIIHTVLQNGANKASCEAHNMLETVYKKVGLK). The one or more aminoacyl-tRNA synthetases may have a polypeptide sequence that consists essentially of, or consists of the polypeptide sequence of SEQ ID NO: 2.
[0038] In embodiments, the OTS comprises the tRNA polynucleotide having at least 80% sequence identity' to SEQ ID NO: 1 and the aminoacyl-tRNA synthetase having at least 80% identity to SEQ ID NO: 2.
[0039] In embodiments, the one or more tRNAs are encoded by a polynucleotide having at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% sequence identity to the polynucleotide sequence of SEQ ID NO: 3(TGCCCTATGATTTGTAATGGTAGCAAGGGAGGCTCTAACCCTCCAAGTCTGGGTTCG AATCCTAGTGGGGCGACCA). The one or more tRNAs may be encoded by a polynucleotide that consists essentially of, or consists of the polynucleotide of SEQ ID NO: 3.
[0040] In embodiments, the one or more aminoacyl-tRNA synthetases have at least 80% at least 85%, at least 90%, at least 95%, at least 99%, or 100% sequence identity7to the polypeptide sequence of SEQ ID NO: 4(MAITEEKPVEEKKSLSFVEQLVEQDLAEGKNGGRIQTRFPPEPNGYLHIGHAKAICMDF GVAEKYNGVCNLRFDDTNPTKENNEYVENILNDISWLGFKWGNIYYASDYFDKLWEFA IWMIKNGHAYIDEQTAEQIAEQKGTPTTPGIASPYRDRPIEENLELFNKMNTPEAVEGSM VLRAKLDMANPNMHFRDPIIYRIIHTPHHRTGTKWNAYPMYDFAHGQSDFFEGVTHSIC TLEFVPHRPLYDKFIDFLKEMRGETENIHDFRPRQIEFNRLNLTYTMMSKRKLLALVNEG VVAGWDDPRMPTLSGMRRRGYSPESIRKFIDSIGYTKFDALNDMALLEAAVRDDLNKR SLRVSAVLDPVKLVITNYPEGQTEEMVAINNPENEADGTHTITFSKNLWIERGDFMENAS KKFFRMTPGKEVRLKNAYIVKCTGCTKDADGNVIEIQAEYDPISKSGMEGANRKVKGTL HWVSADHCKKAQVRVYDRLFTVEDLGADEREFHELLNPDSLKTIDNCYVEEYAAERKP GEYLQFQRIGYFMADLDSTPDNLIFNKTVGLKDTWAKKNK). The one or more aminoacyl- tRNA synthetases may have a polypeptide sequence that consists essentially or, or consists of the polypeptide sequence of SEQ ID NO: 4.
[0041] In embodiments, the OTS comprises the tRNA polynucleotide having at least 80% sequence identity to SEQ ID NO: 3 and the aminoacyl-tRNA synthetase having at least 80% identity to SEQ ID NO: 4.
[0042] An “orthogonal translation system” is an aaRS:tRNA pair that can be used in a host organism that can be used to incorporate noncanonical amino acids into the host proteins during translation, where the aaRS and tRNA do not interact with the host’s native translation machinery.
[0043] A ‘‘transfer RNA” or “tRNA” is an adaptor molecule composed of RNA, typically 76-90 nucleotides in length in eukaryotes. It is a necessary component of translation from the genetic code in messenger RNA (mRNA) to the amino acid sequence of proteins. Each three-nucleotide codon in mRNA is complemented by a three-nucleotide anticodon in tRNA. Each tRNA carries an amino acid, and along with the protein-synthesizing machinery', the ribosome, it attaches the amino acid to a growing polypeptide chain. An “aminoacyl-tRNA synthetase” or “aaRS” is an enzyme that attaches the appropriate amino acid onto its corresponding tRNA.
[0044] The OTS may incorporate a tryptophan residue at a stop codon UGA in a polypeptide being translated. The OTSs may incorporate a noncanonical amino acid at a stop codon UGA in a polypeptide being translated.
[0045] The OTS may be in a cell or a cell lysate. The OTS may be a component of a cell-free protein synthesis (CFPS) system, which may include any or all of the components and properties of CFPS systems and reactions discussed herein. In exemplary embodiments the OTS is provided in an E. coli lysate or E. coli cell. In other embodiments, the OTS is provided in a B. subtilis lysate or B. subtilis cell.
[0046] Cell-free protein synthesis (CFPS) is known and has been described in the art. (See, e.g., U.S. Patent No. 6,548,276; U.S. PatentNo. 7,186,525; U.S. Patent No. 8,734,856; U.S. Patent No. 7,235,382; U.S. Patent No. 7,273,615; U.S. Patent 7,008,651; U.S. Patent 6,994,986 U.S. Patent 7,312,049; U.S. Patent No. 7,776,535; U.S. PatentNo. 7,817,794; U.S. PatentNo. 8,298,759; U.S. Patent No. 8,715,958; U.S. Patent No. 9.005,920; U.S. Publication No. 2014 / 0349353, U.S. Publication No. 2016 / 0060301 , U.S. Publication No. 2018 / 0016612, and U.S. Publication No. 2018 / 0016614, the contents of which are incorporated herein by reference in their entireties). A “cell-free system” or “CFPS reaction mixture” typically contains a crude or partially -purified cell extract, an RNA translation template, and a suitable reaction buffer for promoting cell-free protein synthesis from the RNA translation template. In other embodiments, the cell-free system can include a DNA expression template encoding an open reading frame operably linked to a promoter element for a DNA-dependent RNA polymerase. In these other aspects, the cell-free system can also include a DNA-dependent RNA polymerase to direct transcription of an RNA translation template encoding the open reading frame. In these other aspects, additional NTP’s and divalent cation cofactor can be included in the cell-free system. A reaction mixture is referred to as complete if it contains all reagents necessary' to enable the reaction, and incomplete if it contains only a subset of the necessary reagents. It will be understood by one of ordinary skill in the art that reaction components are routinely stored as separate solutions, each containing a subset of the total components, for reasons of convenience, storage stability, or to allow for application-dependentadjustment of the component concentrations, and that reaction components are combined prior to the reaction to create a complete reaction mixture. Furthermore, it will be understood by one of ordinary skill in the art that reaction components are packaged separately for commercialization and that useful commercial kits may contain any subset of the reaction components of the invention.
[0047] The term “reaction mixture,” as used herein, refers to a solution containing reagents necessary to carry out a given reaction. A reaction mixture is referred to as complete if it contains all reagents necessary to perform the reaction. Components for a reaction mixture may be stored separately in separate container, each containing one or more of the total components. Components may be packaged separately for commercialization and useful commercial kits may contain one or more of the reaction components for a reaction mixture.
[0048] The disclosed cell-free systems may utilize components that are crude and / or that are at least partially isolated and / or purified. As used herein, the term “crude” may mean components obtained by disrupting and lysing cells and, at best, minimally purifying the crude components from the disrupted and lysed cells, for example by centrifuging the disrupted and lysed cells and collecting the crude components from the supernatant and / or pellet after centrifugation. The term “isolated or purified” refers to components that are removed from their natural environment, and are at least 60% free, preferably at least 75% free, and more preferably at least 90% free, even more preferably at least 95% free from other components with which they are naturally associated.
[0049] The cell-free system disclosed herein may comprise a cellular extract from a host strain. Because CFPS exploits an ensemble of catalytic proteins prepared from the crude lysate of cells, the cell extract (whose composition is sensitive to growth media, lysis method, and processing conditions) is an important component of extract-based CFPS reactions. A variety of methods exist for preparing an extract competent for cell-free protein synthesis, including those disclosed in U.S. Patent Application Publication No. 2014 / 0295492 and U.S. Patent Application Publication No. 2016 / 0060301, the contents of which are incorporated by reference in their entireties. The cellular extract of the platform may be prepared from a cell culture of a prokaryote (e.g., E. coll). While E. coli is exemplified herein, the bacterial species is not intended to be limiting. Other bacterial species suitable for the compositions and methods disclosed herein include but are not limited to (e.g, Bacillus species such as Bacillus subtilis. Vibrio species such as Vibrio natrigens, Pseudomonas species, etc.). A eukary otic cell culture (e.g. Chinese Hamster Ovary cells) may be used. In some embodiments, the cell culture is in stationary phase. In some embodiments, stationary phase may be defined as the cell culture having an OD600 of greater than about 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, or having an OD600 within a range bounded by any ofthese values. Further methods for preparing a cell-free system are disclosed in International Patent Application Publication No. WO2020185451A2, the contents of which are incorporated by reference in its entirety.
[0050] The cell extract may be prepared by lysing the cells of the cell culture and isolating a fraction from the lysed cells. For example, the cell extract may be prepared by lysing the cells of the cell culture and subjecting the lysed cells to centrifugal force, and isolating a fraction after centrifugation.
[0051] In embodiments, physiologically compatible (but not necessarily natural) ions and buffers are utilized for transcription, translation, and / or glycosylation, e.g., potassium glutamate, ammonium chloride and the like. Physiological cytoplasmic salt conditions are well-known to those of skill in the art.
[0052] Methods for using orthogonal translation systems
[0053] In a second aspect, methods are disclosed for use of an OTS. The methods may comprise incorporating an amino acid in a growing polypeptide chain at an mRNA UGA stop codon, a UAG stop codon, or both by providing a translation template to any of the OTSs described herein. As used herein, "translation template’7refers to an RNA product of transcription from an expression template that can be used by ribosomes to synthesize polypeptides or proteins.
[0054] The translation template may be expressed from a transcription template, and therefore the methods may further comprise providing a translation template and the components necessary for transcription of the transcription template into the translation template. As used herein, “transcription template” and “expression template” refer to a nucleic acid that serves as substrate for transcribing at least one RNA that can be translated into a sequence defined biopolymer (e.g., a polypeptide or protein). Expression templates include nucleic acids composed of DNA or RNA. Suitable sources of DNA for use a nucleic acid for an expression template include genomic DNA, cDNA and RNA that can be converted into cDNA. Genomic DNA, cDNA and RNA can be from any biological source, such as a tissue sample, a biopsy, a swab, sputum, a blood sample, a fecal sample, a urine sample, a scraping, among others. The genomic DNA, cDNA and RNA can be from host cell or virus origins and from any species, including extant and extinct organisms. As used herein, “expression template” and “transcription template” have the same meaning and are used interchangeably.
[0055] The translation template may be formed by expression of the transcription template in a cell lysate or a cell. The cell lysate may be provided in a CFPS. The translation template may be formed in a bacteria lysate. In exemplars’ embodiments, the translation template is formed in an E. coli cell or E. coli lysate.
[0056] Methods and compositions / kits for identifying one or more components of OTSs
[0057] In a third aspect, a composition or kit for identifying a candidate orthogonal tRNA to an organism is provided. The composition or kit comprises: (a) one or more transcription templates, each transcription template comprising a polynucleotide encoding a suppressor tRNA; (b) one or more CFPS systems derived from the organism; and (c) one or more translation templates, each translation template encoding a reporter protein, and each translation template comprising a premature stop codon. In any of the kits described herein, one or all of the components (a), (b), and (c) are provided in separate containers for storage until use. Furthermore, each of the one or more transcription templates, CFPS systems, and translation templates may be provided in separate containers. The terms “container” and “vessel,” as used herein, refer to any container suitable for holding on or more of the reactants (e.g., for use in one or more transcription, translation, and / or glycosylation steps) described herein. Examples of vessels include, but are not limited to, a microtitre plate, a test tube, a microfuge tube, a beaker, a flask, a multi-well plate, a cuvette, a flow system, a microfiber, a microscope slide and the like.
[0058] The suppressor tRNA may be identified from a metagenomic analysis, e.g., of one or more viruses, such as for example one or more bacteriophages. The transcription template may comprise a promoter. The promoter may be any type of promoter suitable for use in in vitro transcription systems, such as those disclosed herein. In an embodiment, the promoter is a T7 promoter. In embodiments, the composition or kit further comprises a T7 polymerase. In preferred embodiments, the transcription template is linear. The suppressor tRNA may be a premature tRNA including an RNase P tag (5’ leader sequence) so that RNase P in the CFPS system can cleave the RNase P tag, thereby processing the premature tRNA into mature tRNA. It should be understood that alternative systems are also contemplated for processing the premature tRNA into mature tRNA. The CFPS system is derived from the organism, meaning the components of the system, including the aaRSs, are the same as those that are native to the organism. In preferred embodiments, the CFPS system comprises a cell lysate. In exemplary embodiments, the organism is E. coli.
[0059] The reporter protein may be any reporter protein suitable for use in the methods and systems disclosed herein. The term “reporter” or “reporter protein” refers to a protein that can be detected in a reaction mixture, typically in response to the presence of an analyte in the reaction mixture. By incorporating a premature stop codon in the translation template encoding the reporter protein, the reporter protein will only be expressed and detected if a tRNA is able to incorporate an alternate amino acid at the premature stop codon, thereby suppressing or preventing premature stopping of translation. In order for the tRNA to function, it must be aminoacylated by endogenousaaRS in the system. Therefore, if the reporter protein is not expressed, the suppressor tRNA is orthogonal to the system. Numerous reporter proteins and reporter protein / substrate combinations are well known in the art. By way of example, but not by way of limitation, exemplary reporter proteins include fluorescent proteins, enzymes that create colored products, and luciferase. Nonlimiting examples include Green Fluorescent Protein, Red Fluorescent Protein, Y ellow Fluorescent Protein, catechol 2,3-dehydrogenase (C23DO). lacZ, derivatives thereof, and the like. In exemplary embodiments, the reporter protein comprises 216X-sfGFP. wherein X refers to a premature stop codon of UAG, UAA, or UAG. The 216X-sfGFP may be encoded by a polynucleotide sequence comprising or consisting of: atgagcaaaggtgaagaactgtttaccggcgttgtgccgattctggtggaactggatggcgatgtgaacggtcacaaattcagcgtgcgtgg tgaaggtgaaggcgatgccacgattggcaaactgacgctgaaatttatctgcaccaccggcaaactgccggtgccgtggccgacgctggt gaccaccctgacctatggcgttcagtgttttagtcgctatccggatcacatgaaacgtcacgatttctttaaatctgcaatgccggaaggctatg tgcaggaacgtacgattagctttaaagatgatggcaaatataaaacgcgcgccgttgtgaaatttgaaggcgataccctggtgaaccgcattg aactgaaaggcacggattttaaagaagatggcaatatcctgggccataaactggaatacaactttaatagccataatgtttatattacggcgga taaacagaaaaatggcatcaaagcgaattttaccgttcgccataacgttgaagatggcagtgtgcagctggcagatcattatcagcagaatac cccgattggtgatggtccggtgctgctgccggataatcattatctgagcacgcagaccgttctgtctaaagatccgaacgaaaaaggctagc gggaccacatggttctgcacgaatatgtgaatgcggcaggtattacgtggagccatccgcagttcgaaaaataa (SEQ ID NO:451); atgagcaaaggtgaagaactgtttaccggcgttgtgccgattctggtggaactggatggcgatgtgaacggtcacaaattcagcgtgcgtgg tgaaggtgaaggcgatgccacgattggcaaactgacgctgaaatttatctgcaccaccggcaaactgccggtgccgtggccgacgctggt gaccaccctgacctatggcgttcagtgttttagtcgctatccggatcacatgaaacgtcacgatttctttaaatctgcaatgccggaaggctatg tgcaggaacgtacgattagctttaaagatgatggcaaatataaaacgcgcgccgttgtgaaatttgaaggcgataccctggtgaaccgcattg aactgaaaggcacggattttaaagaagatggcaatatcctgggccataaactggaatacaactttaatagccataatgtttatattacggcgga taaacagaaaaatggcatcaaagcgaattttaccgttcgccataacgttgaagatggcagtgtgcagctggcagatcattatcagcagaatac cccgattggtgatggtccggtgctgctgccggataatcattatctgagcacgcagaccgttctgtctaaagatccgaacgaaaaaggctaac gggaccacatggttctgcacgaatatgtgaatgcggcaggtattacgtggagccatccgcagttcgaaaaataa (SEQ ID NO:452); or atgagcaaaggtgaagaactgtttaccggcgttgtgccgattctggtggaactggatggcgatgtgaacggtcacaaattcagcgtgcgtgg tgaaggtgaaggcgatgccacgattggcaaactgacgctgaaatttatctgcaccaccggcaaactgccggtgccgtggccgacgctggt gaccaccctgacctatggcgttcagtgttttagtcgctatccggatcacatgaaacgtcacgatttctttaaatctgcaatgccggaaggctatg tgcaggaacgtacgattagctttaaagatgatggcaaatataaaacgcgcgccgttgtgaaatttgaaggcgataccctggtgaaccgcattg aactgaaaggcacggattttaaagaagatggcaatatcctgggccataaactggaatacaactttaatagccataatgtttatattacggcgga taaacagaaaaatggcatcaaagcgaattttaccgttcgccataacgttgaagatggcagtgtgcagctggcagatcattatcagcagaatac cccgattggtgatggtccggtgctgctgccggataatcattatctgagcacgcagaccgttctgtctaaagatccgaacgaaaaaggctgacgggaccacatggttctgcacgaatatgtgaatgcggcaggtattacgtggagccatccgcagttcgaaaaataa (SEQ ID NO: 453).
[0060] 216X-sfGFP is a superfolder green fluorescent protein (sfGFP) comprising a mutation at position T216, wherein mutation is encoded by a stop codon. This premature stop codon prevents translation and expression of the sfGFP unless it can be suppressed by an amber codon-based tRNA.
[0061] In a fourth aspect, a method for identifying a candidate orthogonal tRNA to an organism is provided, the method comprising transcribing one or more transcription templates in vitro with to produce one or more transcription products, wherein each transcription template comprises a polynucleotide encoding a premature suppressor tRNA; incubating the one or more transcription products with a CFPS system derived from the organism, and a translation template encoding a reporter protein having a premature stop codon under conditions that promote protein synthesis; and performing an assay that detects the report protein; wherein the CFPS system comprises RNase P; and wherein if the reporter protein is not detected, the suppressor tRNA is a candidate orthogonal tRNA to the organism. In embodiments, if the reporter protein is detected, but is detected at significantly lower levels than what is observed when the assay is performed with a known non- orthogonal tRNA, then the suppressor tRNA is a candidate orthogonal tRNA. The method may be used in a high-throughput format to examine multiple suppressor tRNAs, as described herein.
[0062] In preferred embodiments, the transcription templates are linear. In embodiments, transcribing the transcription templates is done with a T7 RNA polymerase.
[0063] In exemplary embodiments, the reporter protein comprises 216X-sfGFP, wherein X refers to UAG, UAA, or UAG. An assay that detects the reporter protein may comprise, for example, a fluorescence microscopy assay. Any assays for detecting reporter proteins known in the art may be used.
[0064] Once a candidate orthogonal tRNA to an organism is identified, functionality and orthogonality may be verified with cognate aaRS for the candidate. The candidate aaRS:tRNA pair may also be tested for functionality in incorporating non-canonical amino acids at non-stop codons.
[0065] The terms “polynucleotide”, “nucleic acid” and “oligonucleotide,” as used herein, refer to polydeoxyribonucleotides (containing 2-deoxy-D-ribose). polyribonucleotides (containing D- ribose), and to any other type of polynucleotide that is an N glycoside of a purine or pyrimidine base. There is no intended distinction in length between the terms “nucleic acid”, “oligonucleotide” and “polynucleotide”, and these terms will be used interchangeably. These terms refer only to the primary structure of the molecule. Thus, these terms include double- and single-stranded DNA, as well as double- and single-stranded RNA. For use in the present methods, an oligonucleotide alsocan comprise nucleotide analogs in which the base, sugar, or phosphate backbone is modified as well as non-purine or non-pyrimidine nucleotide analogs.
[0066] Oligonucleotides can be prepared by any suitable method, including direct chemical synthesis by a method such as the phosphotriester method of Narang et al., 1979, Meth. Enzymol. 68:90-99; the phosphodiester method of Brown et al., 1979, Meth. Enzy mol. 68: 109-151; the diethylphosphoramidite method of Beaucage et al., 1981, Tetrahedron Letters 22: 1859-1862; and the solid support method of U.S. Pat. No. 4.458,066, each incorporated herein by reference. A review of synthesis methods of conjugates of oligonucleotides and modified nucleotides is provided in Goodchild, 1990, Bioconjugate Chemistry 1(3): 165-187, incorporated herein by reference.
[0067] Oligonucleotides and polynucleotides may optionally include one or more non-standard nucleotide(s), nucleotide analog(s) and / or modified nucleotides. Examples of modified nucleotides include, but are not limited to diaminopurine, S2T, 5 -fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xantine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl)uracil, 5- carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine. N6-isopentenyladenine. 1 -methylguanine, 1 -methylinosine. 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46- isopentenyladenine, uracil-5 -oxy acetic acid (v), wybutoxosine. pseudouracil, queosine, 2- thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5 -oxy acetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2- carboxypropyl) uracil, (acp3)w, 2,6-diaminopurine and the like. Nucleic acid molecules may also be modified at the base moiety (e g., at one or more atoms that typically are available to form a hydrogen bond with a complementary nucleotide and / or at one or more atoms that are not ty pically capable of forming a hydrogen bond with a complementary nucleotide), sugar moiety or phosphate backbone.
[0068] A “recombinant nucleic acid” is a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two or more otherwise separated segments of sequence. This artificial combination is often accomplished by chemical synthesis or, more commonly, by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques known in the art. The term recombinant includes nucleic acids that have been altered solely by addition, substitution, or deletion of a portion of the nucleic acid. Frequently, a recombinant nucleic acid may include a nucleic acid sequence operably linked to a promotersequence. Such a recombinant nucleic acid may be part of a vector that is used, for example, to transform a cell. The nucleic acids disclosed herein may be ‘"substantially isolated or purified / ’ The term '‘substantially isolated or purified” refers to a nucleic acid that is removed from its natural environment, and is at least 60% free, preferably at least 75% free, and more preferably at least 90% free, even more preferably at least 95% free from other components with which it is naturally associated.
[0069] As used herein, the terms “peptide,” “polypeptide.” and “protein.” refer to molecules comprising a chain a polymer of amino acid residues j oined by amide linkages. The term “amino acid residue,” includes but is not limited to amino acid residues contained in the group consisting of alanine (Ala or A), cysteine (Cys or C). aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H). isoleucine (He or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gin or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Vai or V), tryptophan (Trp or W), and tyrosine (Tyr or Y) residues. The term “amino acid residue” also may include nonstandard or unnatural amino acids. The term “amino acid residue” may include alpha-, beta-, gamma-, and delta-amino acids.
[0070] In some embodiments, the term “amino acid residue” may include noncanonical or nonstandard or unnatural amino acid residues. Examples of nonstandard or unnatural amino acids include, but are not limited, to homocysteine, 2-Aminoadipic acid. N-Ethylasparagine, 3- Aminoadipic acid. Hydroxylysine, (3-alanine. 0-Amino-propionic acid, allo-Hydroxylysine acid, 2-Aminobutyric acid, 3-Hydroxyproline, 4-Aminobutyric acid, 4-Hydroxyproline, piperidinic acid, 6-Aminocaproic acid, Isodesmosine, 2-Aminoheptanoic acid, allo-Isoleucine, 2- Aminoisobutyric acid, N-Methylglycine, sarcosine, 3-Aminoisobutyric acid, N-Methylisoleucine, 2-Aminopimelic acid, 6-N-Methyllysine, 2.4-Diaminobutyric acid, N-Methylvaline, Desmosine, Norvaline, 2,2'-Diaminopimelic acid. Norleucine, 2,3-Diaminopropionic acid. Ornithine, and N- Ethylglycine, a p-acetyl-L-phenylalanine, a p-iodo-L-phenylalanine, an O-methyl-L-tyrosine, a p- propargy loxyphenylalanine, a p-propargyl-phenylalanine. an L-3-(2-naphthyl)alanine, a 3-methyl- phenylalanine, an O-4-allyl-L-tyrosine, a 4-propyl-L-tyrosine, a tri-O-acetyl-GlcNAcp0-serine, an L-Dopa, a fluonnated phenylalanine, an isopropyl-L-phenylalanine, a p-azido-L-phenylalanine, a p-acyl-L-phenylalanine, a p-benzoyl-L-phenylalanine, an L-phosphoserine, a phosphonoserine, a phosphonotyrosine, a p-bromophenylalanine, a p-amino-L-phenylalanine, an isopropyl-L- phenylalanine, an unnatural analogue of a tyrosine amino acid: an unnatural analogue of a glutamine amino acid; an unnatural analogue of a phenylalanine amino acid; an unnatural analogue of a serine amino acid; an unnatural analogue of a threonine amino acid; an unnatural analogue ofa methionine amino acid; an unnatural analogue of a leucine amino acid; an unnatural analogue of a isoleucine amino acid; an alkyl, aryl, acyl, azido, cyano, halo, hydrazine, hydrazide, hydroxyl, alkenyl, alkyl, ether, thiol, sulfonyl, seleno, ester, thioacid, borate, boronate, ufa hor, phosphono, phosphine, heterocyclic, enone, imine, aldehyde, hydroxylamine, keto, or amino substituted amino acid, or a combination thereof; an amino acid with a photoactivatable cross-linker; a spin-labeled amino acid; a fluorescent amino acid; a metal binding amino acid; a metal-containing amino acid; a radioactive amino acid; a photocaged and / or photoisomenzable amino acid; a biotin or biotinanalogue containing amino acid; a keto containing amino acid; an amino acid comprising polyethylene glycol or poly ether; a heavy atom substituted amino acid; a chemically cleavable or photocleavable amino acid; an amino acid with an elongated side chain; an amino acid containing a toxic group; a sugar substituted amino acid; a carbon-linked sugar-containing amino acid; a redox-active amino acid; an a-hydroxy containing acid; an amino thio acid; an a, a disubstituted amino acid; a P-amino acid; a y-amino acid, a cyclic amino acid other than proline or histidine, and an aromatic amino acid other than phenylalanine, ty rosine or try ptophan.
[0071] The term “amino acid residue” may include L isomers or D isomers of any of the aforementioned amino acids.
[0072] As used herein, a “peptide” is defined as a short polymer of amino acids, of a length ty pically of 20 or less amino acids, and more ty pically of a length of 12 or less amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110). In some embodiments, a peptide as contemplated herein may include no more than about 2, 3, 4, 5. 6, 7, 8. 9, 10. 11. 12. 13. 14, 15, 16, 17, 18, 19, or 20 amino acids. A polypeptide, also referred to as a protein, is ty pically of length > 100 amino acids (Garrett & Grisham, Biochemistry, 2nd edition, 1999, Brooks / Cole, 110). A polypeptide, as contemplated herein, may comprise, but is not limited to, 100, 101, 102, 103. 104, 105, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 525, about 550, about 575, about 600, about 625, about 650, about 675, about 700, about 725, about 750, about 775, about 800, about 825, about 850, about 875, about 900, about 925, about 950, about 975, about 1000. about 1100, about 1200. about 1300, about 1400, about 1500, about 1750, about 2000, about 2250, about 2500 or more amino acid residues.
[0073] A peptide as contemplated herein may be further modified to include non-amino acid moieties. Modifications may include but are not limited to acylation (e.g., O-acylation (esters), N- acylation (amides), S-acylation (thioesters)), acetylation (e.g.. the addition of an acetyl group, either at the N-terminus of the protein or at lysine residues), formylation lipoylation (e.g.,attachment of a lipoate, a C8 functional group), myristoylation (e.g., attachment of myristate, a C14 saturated acid), palmitoylation (e.g., attachment of palmitate, a C16 saturated acid), alkylation (e.g., the addition of an alkyl group, such as an methyl at a lysine or arginine residue), isoprenylation or prenylation (e.g., the addition of an isoprenoid group such as famesol or geranylgeraniol), amidation at C-terminus, glycosylation (e g., the addition of a glycosyl group to either asparagine, hydroxylysine, serine, or threonine, resulting in a glycoprotein), . Distinct from glycation, which is regarded as a nonenzymatic attachment of sugars, polysialylation (e.g., the addition of polysialic acid), glypiation (e.g., glycosylphosphatidylinositol (GPI) anchor formation, hydroxylation, iodination (e.g., of thyroid hormones), and phosphorylation (e.g., the addition of a phosphate group, usually to serine, tyrosine, threonine or histidine).
[0074] The term "amplification reaction” refers to any chemical reaction, including an enzymatic reaction, which results in increased copies of a template nucleic acid sequence or results in transcription of a template nucleic acid. Amplification reactions include reverse transcription, the polymerase chain reaction (PCR), including Real Time PCR (see U.S. Pat. Nos. 4,683,195 and 4,683,202; PCR Protocols: A Guide to Methods and Applications (Innis et al., eds, 1990)), and the ligase chain reaction (LCR) (see Barany et al.. U.S. Pat. No. 5.494,810). Exemplary “amplification reactions conditions” or “amplification conditions” typically comprise either two or three step cycles. Two-step cycles have a high temperature denaturation step followed by a hybridization / elongation (or ligation) step. Three step cycles comprise a denaturation step followed by a hybridization step followed by a separate elongation step.
[0075] The terms “target,” “target sequence”, “target region”, and “target nucleic acid,” as used herein, are synonymous and refer to a region or sequence of a nucleic acid which is to be amplified, sequenced, or detected.
[0076] The term “hybridization,” as used herein, refers to the formation of a duplex structure by two single-stranded nucleic acids due to complementary base pairing. Hybridization can occur between fully complementary nucleic acid strands or between “substantially complementary” nucleic acid strands that contain minor regions of mismatch. Conditions under which hybridization of fully complementary nucleic acid strands is strongly preferred are referred to as “stringent hybridization conditions” or “sequence-specific hybridization conditions”. Stable duplexes of substantially complementary sequences can be achieved under less stringent hybridization conditions; the degree of mismatch tolerated can be controlled by suitable adjustment of the hybridization conditions. Those skilled in the art of nucleic acid technology can determine duplex stability empirically considering a number of variables including, for example, the length and base pair composition of the oligonucleotides, ionic strength, and incidence of mismatched base pairs, 1following the guidance provided by the art (see, e.g., Sambrook et al., 1989, Molecular Cloning- A Laboratory’ Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York; Wetmur, 1991, Critical Review in Biochem. and Mol. Biol. 26(3 / 4):227-259; and Owczarzy et al., 2008, Biochemistry, 47: 5336-5353, which are incorporated herein by reference).
[0077] The term “primer,” as used herein, refers to an oligonucleotide capable of acting as a point of initiation of DNA synthesis under suitable conditions. Such conditions include those in which synthesis of a primer extension product complementary to a nucleic acid strand is induced in the presence of four different nucleoside triphosphates and an agent for extension (for example, a DNA polymerase or reverse transcriptase) in an appropriate buffer and at a suitable temperature.
[0078] A primer is preferably a single-stranded DNA. The appropriate length of a primer depends on the intended use of the primer but typically ranges from about 6 to about 225 nucleotides, including intermediate ranges, such as from 15 to 35 nucleotides, from 18 to 75 nucleotides and from 25 to 150 nucleotides. Short primer molecules generally require cooler temperatures to form sufficiently stable hybrid complexes with the template. A primer need not reflect the exact sequence of the template nucleic acid, but must be sufficiently complementary’ to hybridize with the template. The design of suitable primers for the amplification of a given target sequence is well known in the art and described in the literature cited herein.
[0079] Primers can incorporate additional features which allow7for the detection or immobilization of the primer but do not alter the basic property of the primer, that of acting as a point of initiation of DNA synthesis. For example, primers may contain an additional nucleic acid sequence at the 5' end which does not hybridize to the target nucleic acid, but which facilitates cloning or detection of the amplified product, or w hich enables transcription of RNA (for example, by inclusion of a promoter) or translation of protein (for example, by inclusion of a 5’-UTR, such as an Internal Ribosome Entry’ Site (IRES) or a 3’-UTR element, such as a poly(A)n sequence, w here n is in the range from about 20 to about 200). The region of the primer that is sufficiently complementary’ to the template to hybridize is referred to herein as the hybridizing region.
[0080] As used herein, a primer is “specific,” for a target sequence if, when used in an amplification reaction under sufficiently stringent conditions, the primer hybridizes primarily to the target nucleic acid. Typically, a primer is specific for a target sequence if the primer-target duplex stability’ is greater than the stability of a duplex formed between the primer and any other sequence found in the sample. One of skill in the art will recognize that various factors, such as salt conditions as w ell as base composition of the primer and the location of the mismatches, will affect the specificity’ of the primer, and that routine experimental confirmation of the primer specificity will be needed in many cases. Hybridization conditions can be chosen under which theprimer can form stable duplexes only with a target sequence. Thus, the use of target-specific primers under suitably stringent amplification conditions enables the selective amplification of those target sequences that contain the target primer binding sites.
[0081] As used herein, a '‘polymerase” refers to an enzy me that catalyzes the polymerization of nucleotides. “DNA polymerase” catalyzes the polymerization of deoxyribonucleotides. Known DNA polymerases include, for example, Pyrococcus furiosus (Pfu) DNA polymerase. E. coli DNA polymerase I, T7 DNA polymerase and Thermus aquaticus (Taq) DNA polymerase, among others. “RNA polymerase” catalyzes the polymerization of ribonucleotides. The foregoing examples of DNA polymerases are also known as DNA-dependent DNA polymerases. RNA-dependent DNA polymerases also fall within the scope of DNA polymerases. Reverse transcriptase, which includes viral polymerases encoded by retroviruses, is an example of an RNA-dependent DNA polymerase. Known examples of RNA polymerase (“RNAP”) include, for example, T3 RNA polymerase, T7 RNA polymerase, SP6 RNA polymerase and E. coli RNA polymerase, among others. The foregoing examples of RNA polymerases are also known as DNA-dependent RNA polymerase. The polymerase activity of any of the above enzymes can be determined by means well known in the art.
[0082] The term “promoter” refers to a cis-acting DNA sequence that directs RNA polymerase and other trans-acting transcription factors to initiate RNA transcription from the DNA template that includes the cis-acting DNA sequence.
[0083] As used herein, the term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a '‘plasmid,” which refers to a circular double stranded DNA loop into which additional DNA segments can be ligated. Such vectors are referred to herein as “expression vectors.” In general, expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. In the present specification, ■‘plasmid” and '‘vector” can be used interchangeably.
[0084] In certain exemplar}' embodiments, the recombinant expression vectors comprise a nucleic acid sequence (e.g., a nucleic acid sequence encoding one or more rRNAs or reporter polypeptides and / or proteins described herein) in a form suitable for expression of the nucleic acid sequence in one or more of the methods described herein, which means that the recombinant expression vectors include one or more regulator}' sequences which is operatively linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, “operably linked” is intended to mean that the nucleotide sequence encoding one or more rRNAs or reporter polypeptides and / or proteins described herein is linked to the regulatory sequence(s) in a manner which allows for expression of the nucleotide sequence (e.g., in an in vitro transcription and / or translation system). The term“regulatory sequence” is intended to include promoters, enhancers and other expression control elements (e.g., polyadenylation signals). Such regulatory sequences are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990).
[0085] As utilized herein, a “deletion” means the removal of one or more nucleotides relative to the native polynucleotide sequence. The engineered strains that are disclosed herein may include a deletion in one or more genes (e.g., a deletion in gmd and / or a deletion in waaL). Preferably, a deletion results in a non-functional gene product. As utilized herein, an “insertion” means the addition of one or more nucleotides to the native polynucleotide sequence. The engineered strains that are disclosed herein may include an insertion in one or more genes (e.g., an insertion in gmd and / or an insertion in waaL). Preferably, a deletion results in a non-functional gene product. As utilized herein, a “substitution” means replacement of a nucleotide of a native polynucleotide sequence with a nucleotide that is not native to the polynucleotide sequence. The engineered strains that are disclosed herein may include a substitution in one or more genes (e.g., a substitution in gmd and / or a substitution in waaL). Preferably, a substitution results in a non-functional gene product, for example, where the substitution introduces a premature stop codon (e.g.. TAA, TAG, or TGA) in the coding sequence of the gene product. In some embodiments, the engineered strains that are disclosed herein may include two or more substitutions where the substitutions introduce multiple premature stop codons (e.g., TAATAA, TAGTAG, or TGATGA).
[0086] Engineered strains be engineered to include and express one or more heterologous genes. As would be understood in the art, a heterologous gene is a gene that is not naturally present in the engineered strain as the strain occurs in nature. A gene that is heterologous to E. coli is a gene that does not occur in E. coli and may be a gene that occurs naturally in another microorganism or a gene that does not occur naturally in any other known microorganism (i.e., an artificial gene).
[0087] Miscellaneous
[0088] As used in this specification and the claims, the singular forms “a,” “an,” and “the” include plural forms unless the context clearly dictates otherwise. For example, the term “a substituent” should be interpreted to mean “one or more substituents,” unless the context clearly dictates otherwise.
[0089] As used herein, “about”, “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary' skill in the art and will vary' to some extent on the context in which they are used. If there are uses of the term which are not clear to persons of ordinary skill in the art given the context in which it is used, “about” and “approximately” will mean up to plusor minus 10% of the particular term and “substantially"’ and “significantly” will mean more than plus or minus 10% of the particular term.
[0090] As used herein, the terms "‘include” and “including” have the same meaning as the terms “comprise” and “comprising.” The terms “comprise” and “comprising” should be interpreted as being “open” transitional terms that permit the inclusion of additional components further to those components recited in the claims. The terms “consist” and “consisting of’ should be interpreted as being “closed” transitional terms that do not permit the inclusion of additional components other than the components recited in the claims. The term “consisting essentially of’ should be interpreted to be partially closed and allowing the inclusion only of additional components that do not fundamentally alter the nature of the claimed subject matter. Embodiments recited as “including.” “comprising,” or “having” certain elements are also contemplated as “consisting essentially of’ and “consisting of’ those certain elements.
[0091] The phrase “such as” should be interpreted as “for example, including.” Moreover, the use of any and all exemplary' language, including but not limited to “such as”, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed.
[0092] Furthermore, in those instances where a convention analogous to “at least one of A, B and C, etc.” is used, in general such a construction is intended in the sense of one having ordinary' skill in the art would understand the convention (e.g., “a system having at least one of A, B and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description or figures, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of "‘A” or ‘B or “A and B.”
[0093] All language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can subsequently be broken down into ranges and subranges. A range includes each individual member. Thus, for example, a group having 1-3 members refers to groups having 1, 2. or 3 members. Similarly, a group having 6 members refers to groups having 1, 2, 3, 4, or 6 members, and so forth.
[0094] The modal verb “may” refers to the preferred use or selection of one or more options or choices among the several described embodiments or features contained within the same. Where no options or choices are disclosed regarding a particular embodiment or feature contained in the same, the modal verb “may” refers to an affirmative act regarding how to make or use and aspectof a described embodiment or feature contained in the same, or a definitive decision to use a specific skill regarding a described embodiment or feature contained in the same. In this latter context, the modal verb ‘'may” has the same meaning and connotation as the auxiliary verb “can.”EXAMPLE
[0095] The following Example is illustrative and should not be interpreted to limit the scope of the claimed subject matter.
[0096] The genetic code, which defines the link between DNA and protein sequences, is almost universally conserved throughout biology7because it enables the faithful translation of proteins. Proteins are encoded by a set of 61 sense codons, which code for 20 canonical amino acids, and 3 stop codons, which serve as termination signals. Due to this redundancy, the genetic code can be changed and engineered for both natural and synthetic purposes.
[0097] Synthetic biologists have engineered the genetic code to incorporate non-canonical amino acids (ncAAs). ncAAs are important tools for protein engineering that are incorporated through a process called genetic code expansion. In genetic code expansion, a heterologous aminoacyl-tRNA synthetase (aaRS) and suppressor tRNA are introduced into a host of interest to reassign a stop codon as a ncAA. Importantly, the aaRS:tRNA pair must be orthogonal, meaning that they operate in parallel without interacting with native translation machinery7. When an aaRS:tRNA pair satisfies this condition, it is called an orthogonal translation system (OTS). OTSs have become powerful tools, enabling the incorporation of > 200 ncAAs.1
[0098] The identification and discovery of OTSs is a bottleneck in the genetic code expansion field. While many ncAAs have been successfully incorporated, they mostly represent chemical derivatives of tyrosine and lysine. This reflects the fact that the genetic code expansion field has heavily relied on the Methanocaldococcus jannaschii TyrRS:tRNATyrcuA and the PylRS:lRNAI< i A pairs for ncAA incorporation. There has been some work on expanding the scope of OTSs to enable the incorporation of a wider range of chemistries.2 7In most cases, these works follow the general strategy7of using aaRS:tRNA pairs from evolutionarily distant organisms, which takes advantage of divergent aaRS:tRNA interaction mechanisms between domains of life. Yet, there are still only two commonly used OTSs.
[0099] Recently , there has been a growing appreciation for the widespread stop codon reassignments found in bacteriophages.8,9These bacteriophages are predicted to use stop codon reassignment as a regulation strategy and can encode their own machinery (tRNAs, aaRSs, or both) for stop codon reassignment. In contrast to those found in evolutionarily distant organisms, thesetRNAs and aaRSs represent a vast but untapped and uncharacterized catalog of translational machinery for stop codon reassignment.
[0100] A high-throughput, cell-free method was developed to characterize tRNAs and aaRSs responsible for stop codon reassignment from phage genomes. We curated a set of >200 suppressor tRNAs for all three stop codons and characterized their activity7and orthogonality in E. coli cell lysates. We then characterize a panel of aaRSs and rapidly test thousands of combinations of aaRS:tRNA pairs. We discovered the first natural OTS capable of reassigning the UGA codon as Trp in E. coli, as well as a novel GlnRS:tRNA pair that is a candidate for further development. In total, we developed a powerful method to discover orthogonal translation systems and highlight a vast and untapped source for OTSs in bacteriophages.
[0101] Identification of suppressor tRNAs in phage genomes
[0102] A library of suppressor tRNAs from bacteriophages that were identified in metagenomic datasets (Table 1 below) were curated. These were identified using tRNAscan-SE, a bioinformatics algorithm to detect and classify tRNA sequences from genomes.10In total, we identified a library of > 200 suppressor tRNAs. This library included suppressor tRNAs for all three stop codons, UAG, UAA, and UGA, and a range of confidence scores, to capture novel sequences that may act as tRNAs. For ease of nomenclature, they are referred to by an identifier that includes their predicted codon and a unique number throughout the example (Table 1 below).
[0103] Development of a high throughput tRNA expression platform
[0104] A high-throughput technology was required to screen through the suppressor tRNA dataset. Suppressor tRNAs can be expressed and processed into mature tRNAs during CFPS. This plasmid-based expression system used a proK promoter and terminator cassette for tRNA expression but could not activate expression when used as a linear template. The requirement for plasmid-based expression hindered high-throughput evaluation of these tRNAs, particularly for the large numbers that we identified computationally and wished to evaluate.
[0105] Instead, we envisioned a high throughput tRNA evaluation platform that decouples expression from maturation (FIG. 1 A). A premature tRNA could first be transcribed in vitro using T7 RNA polymerase from a linear DNA sequence. These linear DNA sequences are typically less than 300 bp. which makes them compatible with commercial gene synthesis at cheap costs. The template DNA is designed to contain a 5’ purine rich sequence that can be efficiently cleaved by RNase P11and is terminated by the conserved 3 ’-CCA sequence. This 5’ purine rich RNA sequence circumvents several issues, including inefficient transcription from sequences containing A, C, or T at the +1 position when using T7 RNA polymerase12and inefficient ribozyme cleavage when using 5’ self-cleaving hammerhead ribozymes.13After transcription, crude products could then besupplemented directly into CFPS reactions, where endogenous RNase P could process the premature tRNA into mature tRNA that is capable of suppressing a premature stop codon in a protein of interest.
[0106] We first tested whether the decoupled expression and maturation platform would be amenable to evaluating tRNAs in high-throughput. We used the M. alvus tRNA^'cuA as a test case because it contains a native +1 G and therefore does not require processing to be efficiently transcribed. First, we analyzed RNase P activity in cell extracts by incubating premature M. alvus tRNA^'cuA containing the 5’ RNase P tag with serial dilutions of cell extract at 37 °C for 1 hour. After incubation with a 0.1X volumes of cell extract, the band corresponding to the premature tRNA was consumed and a band corresponding to the cleaved RNase P tag appeared as observed by denaturing Urea-PAGE (FIG. IB). Mature tRNA bands were observed, which likely contains a mixture of E. coll total tRNA and the mature M. alvus IRNA^’CUA. The sizes of these tRNAs are consistent with a standard of the M. alvus IRNA^'CUA lacking the RNase P tag. These data suggest efficient cleavage of the RNase P tag from premature tRNAs. Next, we tested to see if the cleaved M. M. alvus tRNA^'cuA was functional in CFPS by assaying its ability to suppress a UAG codon in sfGFP at position 216 (hereafter referred to as 216UAG-sfGFP) when supplemented with its cognate M. alvus IRNA^'CUA and Azidolysine (AzK), a ncAA accepted by the M. alvus tRNAl>vlc( JA. The M. alvus tRNAPvlcr.\ efficiently supported the synthesis of 216UAG-sfGFP both in the absence and presence of the RNase P tag. suggesting that the RNase P tag is efficiently removed in situ during CFPS and that the resulting tRNA is functional for UAG suppression (FIG. 1 C, see M. alvus column). In addition, products from control in vitro transcription reactions where T7 RNA polymerase was omitted were not able to support 216UAG-sfGFP, suggesting that the linear DNA template used for tRNA expression is not sufficient for tRNA expression in CFPS and that the active molecule is the tRNA product produced during in vitro transcription (FIG. 1C). In total, this suggests that a decoupled tRNA expression and maturation platform could be amenable to high-throughput testing of suppressor tRNAs.
[0107] To evaluate the generalizability of this platform, we tested a previously-described panel of metagenomically identified IRNASCUA that were found in the genomes of huge phages (FIG. IC).14We showed that this panel of tRNAscuA could be expressed from plasmids and were substrates for E. coli GlnRS.15When supplemented into CFPS reactions as crude in vitro transcription products without the RNase P tag, we observed that 3 / 6 supported little to no 216UAG-sfGFP synthesis (FIG. IC). Of these three, two are likely not efficiently transcribed due to pyrimidine rich sequences at the 5’ tRNA end (AOTPERU CU-CL-ST 331-34: 5 -UCCG and A6_CU-CL_34-1: 5’-UGCC). However, in the presence of the RNase P tag, all tRNAscuA supportefficient 216UAG-sfGFP synthesis (FIG. 1C). Protein yields using RNase P based tRNA templates were robust across all tRNAs tested and were comparable or better to yields from plasmid-based expression templates. In total, this suggests that our decoupled tRNA expression and maturation platform could be amenable to high-throughput workflows.
[0108] High-throughput screening of suppressor tRNAs in phage genomes
[0109] With the successful development of a method to express, mature, and evaluate suppressor tRNA activity, we next designed a workflow to evaluate tRNAs that was entirely cell-free, requiring no cloning or transformation steps, rapid, able to be completed within a single day (PCR: 1.5 hours, in vitro transcription (IVT): 4 hours, CFPS: 2-6 hrs), and amenable to liquid-handling systems (FIG. 2A). Linear DNA templates for bioinformatically identified tRNAs (Table 1 below) were purchased within an expression cassette consisting of a T7 promoter and RNase P tag and were amplified by PCR to prepare in vitro transcription templates. These crude PCR products were used as templates for in vitro transcription reactions, and these crude in vitro transcription products were directly supplemented into CFPS reactions containing a 216X-sfGFP template, which serves as a readout for tRNA suppression activity’.
[0110] Using this CFPS-based assay, we could rapidly identify orthogonal tRNA candidates from large sets of suppressor tRNAs. An orthogonal tRNA would not be aminoacylated by endogenous E. coli aaRSs, resulting in little to no 216X-sfGFP synthesis. On the other hand, a tRNA that is not orthogonal is charged with an amino acid by an E. coli aaRS, resulting in active 216X-sfGFP synthesis. To account for all possible amino acid mutations that could occur from suppression at position 21 , we first confirmed that sfGFP is robust to mutations at T216. All 20 canonical amino acids are tolerated at this position, preventing misidentification of orthogonal tRNAs due to synthesis of an inactive sfGFP (FIG. 3B). FIG. 3A shows that our assay functions as expected with known tRNAs (ALjTyrand ALAPyl). We then used this workflow to evaluate suppressor tRNA activity and orthogonality in our assay against all three stop codons: UAG, UAA, and UGA (FIG. 2B). As internal controls, we also included a set of known orthogonal tRNAs, the M. jannaschii tRNATyr, the AL barkeri tRNAPvl. the AL alvus tRNAP l. and theA17VC 10 / « / tRNA1''1. along with a cognate aaRS and compatible ncAA, to suppress each of the stop codons. Within the UAG and UGA codons, we found that suppressor tRNAs exhibited a wide range of activity; some exhibited near WT-levels of sfGFP synthesis of 216X-sfGFP, while others exhibited no 216X-sfGFP synthesis consistent with a no tRNA control (FIG. 2B). These represent non-orthogonal and orthogonal tRNA candidates, respectively. Interestingly, tRNAsuuA appeared to be non-functional in our assay, which could reflect insufficient expression, improper folding, or lack of essentialpost-transcriptional modifications. In total, this assay rapidly assessed tRNA function and could identify orthogonal tRNA candidates.
[0111] We were first interested in characterizing the activity of tRNAs that were not orthogonal. To identify' the aminoacyl-tRNA synthetases responsible for charging these tRNAs, we purified 216UAG-sfGFP and 216UGA-sfGFP from a representative set of these reactions and analyzed them by intact protein electrospray ionization mass spectrometry’ (ESI-MS). These tRNAs were chosen across a range of activities, spanning 216X-sfGFP synthesis yields from 10-100% of WT sfGFP. ESI-MS of 29 tRNAscuA revealed that 28 / 29 enabled incorporation of either glutamine or lysine at the UAG codon (FIG. 4). This is consistent with literature showing recoding of UAG to glutamine in bacteriophages.16Interestingly, we found one tRNAcuA (TAG-151) that directed the incorporation of alanine (FIG. 4). Rare examples of UAG recoding to alanine in the mitochondria of green algae have been reported in the literature,17although this has been disputed.18Whether this UAG > Ala stop codon reassignment event is present in the phage-host pair has yet to be determined. Analysis of 13 tRNAsucA shows incorporation of tryptophan in 13 / 13 cases (FIG. 4), which is a common stop codon reassignment mechanism.19This data set shows that active tRNAs generally follow known stop codon reassignment events found in nature and could elucidate design rules for determining tRNA orthogonality.
[0112] We also found many tRNAs that are predicted to be orthogonal based on their inability’ to suppress the stop codon in 216X-sfGFP (FIG. 2B). We ruled out potential issues in the tRNA expression and maturation workflow for a set of these tRNAs, including DNA template amplification (FIG. 5A), in vitro transcription (FIG. 5B), and processing by RNase P in cell extracts (FIG. 5C) as explanations for lack of activity . Other alternative explanations, such as improper folding, incompatibility' with auxiliary translation factors like EF-Tu or the ribosome, or lack of essential post transcriptional modifications could also play a role but were not investigated further. We hypothesized that the activity of an orthogonal tRNA could be restored by supplementing the reaction with its cognate aaRS and that these translation systems might operate orthogonally in E. coli. We therefore set out to find these aaRSs.
[0113] Characterization of phage-encoded API and HF2 TrpRS:tRNATrpucA translation systems
[0114] Two IRNASUCA were identified in the high-throughput screen. TGA-11 (JS_APl_S143_scaffold_136784, hereafter referred to as the API tRNAucA) and TGA-13 (JS_HF2_S141_scaffold_159238, hereafter referred to as the HF2 IRNAUCA), that were predicted to be orthogonal based on near-background levels of 216UGA-sfGFP. Previous work implicated these tRNAs as part of a code change operon in bacteriophages which contained the IRNAUCA, RF2, and a putative TrpRS (sequence provided in Table 2 below).9We refer to the TrpRS thatcolocalizes with the API tRNAUCA as the API TrpRS. following a similar convention for the HF2 TrpRS. A SWISS-MODEL homology model of both putative TrpRSs showed strong structural homology to E. coli TrpRS (FIG. 6A). We predicted that these TrpRS:tRNA pairs mediate reassignment of UGA to Trp and that these systems may be functional in E. coli extracts.
[0115] First, we purified and added these putative TrpRS enzymes into CFPS reactions in the presence of putative orthogonal IRNASUCA to identify a functional pair (FIG. 6B). While no 216UGA-sfGFP synthesis is observed in the presence of only tRNA or aaRS, we observe highly efficient 216UGA-sfGFP synthesis in the presence of a functional aaRS:tRNAUCApair. As expected, the API TrpRS and HF2 TrpRS yielded the highest 216UGA-sfGFP yields in the presence of their cognate tRNA (FIG. 6B). These TrpRSs do not have completely mutually orthogonal tRNA specificity, as they each display activity for both the API and HF2 tRNAsucA. However, the API TrpRS is active with TGA-29 while HF2 TrpRS is not, suggesting that it may have a distinct tRNA recognition mechanism. Indeed, the API tRNAucA and TGA-29 both contain the G73 discriminator base, an identity element used by bacterial TrpRSs, while the HF2 IRNAUCA contains the A73 discriminator base, an identity element typically used by eukaryotic and archaeal TrpRS (FIG. 7).20This difference may play a role in the observed tRNA substrate specificities. These results show the identification of active TrpRS:tRNAucA pairs that reassign the UGA stop codon as a sense codon.
[0116] We sought to better characterize the activity of the API and HF2 TrpRSs. Purified 216UGA-sfGFP from CFPS reactions containing the API-IRNAUCA: API -TrpRS pair and the HF2- tRNAucA:HF2-TrpRS pair were analyzed by ESI-MS. Compared to a WT-sfGFP control, 216UGA-sfGFP from both reactions displayed a +85 Da shift, consistent with a T216W substitution (FIG. 6C). We also tested the aminoacylation activity of these enzy mes using in vitro aminoacylation assays in which a tRNA is aminoacylated by an aaRS of interest, digested by RNase A to liberate an aminoacylated adenosine (aa-A), and analyzed by LC-MS (FIG. 6D). The data confirm that the enzymes are functional TrpRSs that enable a UGA > W codon reassignment.
[0117] We then assessed the orthogonality of the API -IRNAUCA: API -TrpRS pair and the HF2- tRNAucA:HF2-TrpRS pair in E. coli cell extracts to determine if they are suitable for genetic code expansion. As noted previously, addition of tRNA alone to CFPS reactions did not result in observable 216UGA-sfGFP synthesis, suggesting that the tRNA is orthogonal (FIG. 6B). The orthogonality of the API TrpRS and HF2 TrpRS was assessed by using a mixture of total E. coli tRNA as a substrate in aminoacylation assays. The HF2 TrpRS synthesized Trp-A in the presence of total E. coli tRNA, suggesting that it was not orthogonal (FIG. 6D, purple graph). We found that it displayed activity towards E. coli tRNATrpCCA, though we did not rule out other tRNAs whichmay also be non-specifically recognized by the HF2 TrpRS (FIG. 8). In contrast, the API TrpRS did not synthesize detectable Trp-A in the presence of total E. coll tRNA. This shows that the API TrpRS: API tRNAucA is an orthogonal translation system suitable for genetic code expansion. This is the first reported discovery of a natural UGA-suppressing OTS in E. coll.
[0118] Identification and characterization of host-encoded aaRSs
[0119] While the cognate aaRSs for the API and HF2 tRNAsucA could be identified by genomic proximity with the tRNA as part of a code change operon in bacteriophage genomes, it is likely that many phage-encoded suppressor tRNAs are aminoacylated by a host-encoded aaRS. We searched metagenomic datasets for aaRSs based on TAG- 103 (labelled PHAGE-A2— js4906-20- 3_S2_Complete_Phage_26_29_curated scaffold). This suppressor tRNA is sourced from the curated genome of a bacteriophage infecting a host from the Prevotella genus. We identified 15 predicted aaRSs from this metagenomic dataset, including MetRS, LeuRS, GluRS, AspRS, CysRS, TyrRS, GlnRs, LysRS, PheRS, AsnRS, GlyRS, and SerRS (sequences provided in Table 2). Interestingly, the SerRS contained an in-frame UGA codon, suggesting that the host also reassigns its stop codon. In future experiments, we included only a SerRS version recoding the UGA as UGG (SerRS-TGG). based on the common UGA > W reassignment.
[0120] To identify candidate aaRS:tRNA pairs that may be functional and orthogonal in E. coll, we coexpressed these aaRSs genes with orthogonal tRNAscuA, tRNAsuuA, and IRNASUCA that were identified in FIG. 2 (FIG. 9A, FIG. 10. FIG. 11, FIG. 12). When possible, we included a positive control using a tRNA that was identified to be active with E. coli aaRSs (from FIG. 2) and a negative control without tRNA. In total, this resulted in > 1 ,200 combinations of aaRS:tRNA pairs. We observed no 216X-sfGFP synthesis when using any of the tRNAsuuA or the tRNASucA as substrates for this panel of aaRSs (FIG. 11, FIG. 12). However, a putative GlnRS and TyrRS were able to support 216UAG-sfGFP synthesis when paired with several tRNAscuA. (FIG. 9A). We chose to study the GlnRS:TAG-59 and TyrRS:TAG-27 pairs, which were the highest performing aaRS:tRNA pair for each (FIG. 9A, red boxes). The GlnRS:TAG-59 pair was able to recover efficient 216UAG-sfGFP synthesis, while the TyrRS:TAG-27 IRNACUA pair was only able to moderately support 216UAG-sfGFP synthesis. This low activity may reflect a variety of factors, including that TAG-27 may not be the cognate tRNA for this TyrRS, that the TyrRS may not be strongly active in these reaction conditions, or that the tRNA may be weakly compatible with E. coli translation machinery7. These results show how cell-free systems can be leveraged to quickly identify functional UAG suppressing aaRS:tRNA pairs from metagenomic datasets.
[0121] We next confirmed the assignment of these aaRSs by purification and ESI-MS analysis of 216UAG-sfGFP from reactions containing each pair. As expected, the GlnRS resulted in a productdisplaying +27 Da mass shift, consistent with incorporation of Gin or Lys (FIG. 9B). The TyrRS enabled incorporation of Tyr at the UAG codon, as observed by a +61 Da mass shift (FIG. 9B). In the presence of the TyrRS: TAG-27 pair, a small amount of product corresponding to Gin or Lys readthrough is visible. This product likely corresponds to readthrough of the UAG codon by endogenous E. coli tRNAGln / Lys rather than aminoacylation of TAG-27 with Gin. In previous work, non-specific readthrough of the UAG codon by Gin has been observed.21,22
[0122] To assess the orthogonality of these host-encoded aaRSS, we then evaluated their activity against a mixture of total E. coli tRNA. The TyrRS, consistent with its low yields of 216UAG- sfGFP, only produced a small amount of Tyr- A product in the presence of TAG-27 (FIG. 9C, see inset). However, it showed significant cross reactivity with total E. coli tRNA, which suggests that the TyrRS is not orthogonal (FIG. 9C). TAG-27, however, remains orthogonal when used as a substrate with E. coli TyrRS (FIG. 9C).
[0123] Orthogonality was then assessed with the GlnRS:TAG-59 pair. The GlnRS produced a product corresponding with Gln-A in the presence of TAG-59 and Gin, confirming its assignment as a GlnRS (FIG. 9D). Yields of Gln-A product were significantly higher when TAG-59 was used as substrate compared to total E. coli tRNA (FIG. 9D). We speculate that this small amount of nonspecific aminoacylation activity is insufficient to break orthogonality and may be an artifact of the assay, which ignores competition by endogenous E. coli aaRSs that would outcompete for the pool of E. coll tRNAs. In support of this idea, we show that although TAG-59 is non-specifically aminoacylated by E. coli GlnRS in this assay, this non-specific activity is not observed in CFPS reactions where E. coli tRNAs likely outcompete TAG-59 for binding to E. coli GlnRS (FIG. 9D, see no aaRS row in FIG. 9A). We therefore tentatively hypothesize that this GlnRS:TAG-59 pair may also function as an OTS in E. coli, although further testing will be necessary to confirm this hypothesis.
[0124] Discussion
[0125] The disclosure herein includes a high-throughput pipeline to rapidly identify and discover OTSs from metagenomic datasets. This pipeline screens hundreds of suppressor tRNA candidates to identify orthogonal tRNAs and thousands of combinations of aaRS: orthogonal tRNA pairs to find OTSs. Using this workflow, we discovered a novel aaRS and tRNA pair, the API TrpRS:APl IRNAUCA, that functions as an OTS in E. coli. While recent works have reported the development of the Saccharomyces cerevisiae TrpRS:tRNATrpcuA and the E. coli TrpRS: RNATrpucA as OTSs, the API system is a unique discovery because it naturally and orthogonally reassigns the UGA codon as Trp in the context of native E. coli translation.5,6,23This work therefore presents the first discovery of an OTS in E. coli that naturally uses the UGA codon, though other OTSs can bereprogrammed to do so6-24 26The API OTS is a powerful new tool in genetic code expansion and could be used to incorporate multiple, distinct nAAs.
[0126] This platform provides several key advantages compared to state-of-the-art methods in OTS discover}7. Firstly, the speed and throughput of our entirely cell-free workflow far exceeds in vivo methods. In vivo workflows would require cloning ~ 1000s of plasmid constructs, transforming them all, and screening them in 96-well plates, which is far more laborious and time-consuming than the methods presented here. Secondly, the cell-free workflow circumvents issues related to toxicity. We believe that the API OTS would likely be toxic to E. coli due to efficient reassignment of UGA codons throughout the genome and that this would have prevented its identification using in vivo workflows, though this has not yet been thoroughly tested. Finally, we deliberately chose to screen a natural and widespread source for stop codon reassignment, bacteriophages.9Because typical OTS discovery efforts begin with aaRS:tRNA pairs that suppress sense codons, engineering the tRNAs to reassign stop codons later in the process often makes them incompatible for genetic code expansion.2This can result from either breaking important anticodon identity elements for aaRS recognition or by introducing anticodon identity elements that render the tRNA non- orthogonal. On the other hand, we screened natural suppressor tRNAs that can be directly evaluated for their orthogonality and activity. In total, these results therefore present new, powerful tools for OTS discovery7.
[0127] We predict that the results from our in vitro workflow will correlate to in vivo performance. We note that all OTSs so far. including the AT jannaschii TyrRS:tRN ATyrcuA and the pyrrolysyl systems, were discovered in vivo and have functioned as published when used in cell- free systems. While we generally expect that the reverse is true, we anticipate that stop codon reassignment in the E. coli genome will result in growth defects. Efforts in synonymous codon compression to liberate codons may enable the use of aaRS:tRNA pairs that would otherwise be deleterious.27-29
[0128] In conclusion, we report the development of a cell-free pipeline to identify orthogonal tRNAs and screen thousands of aaRS:tRNA pairs, as well as the discovery of an orthogonal API TrpRS:APl tRNAucA pair that enables a UGA > W reassignment. The results presented here open many interesting lines of further investigation. This platform may be useful to identify rare deviations from the standard genetic code as was observed using TAG-151, w7hich mediated a UAG > Ala reassignment event. While it is still unknown if this exact code change is biologically relevant in the phage-host pair, it raises fascinating questions about the conservation and structure of the genetic code throughout biology. For the genetic code expansion field, efforts to understand and predict tRNA orthogonality in E. coli could leverage our dataset, which describes the activityand orthogonality of hundreds of suppressor tRNAs, to derive general design rules of orthogonality. This would be a large advance in designing and engineering OTSs. Finally, we are actively interested in using the AP 1 OTS for ncAA incorporation.
[0129] Methods
[0130] In vitro transcription of RNase-P-based tRNA templates
[0131] DNA templates encoding the tRNAs were purchased as eBlocks from IDT in a linear template:
[0132] tttcgccacctctgacttgagcgtcgatttttgtgatgctcgtcaggggggcggagcctatggaaaaacgccagcaacgcgatc ccgcgaaattaatacgactcactatagggagaccacaacggtttccctctaga[tRNA]gtcgaccggctgctaacaaagcccgaaagg aagctgagttggctgctgccaccgctgagcaataactagcataaccccttggggcctctaaacgg (SEQ ID NO: 445)
[0133] where [tRNA] denotes the tRNA sequence as listed in Table 1. Templates were amplified using Q5 DNA Polymerase (NEB) following manufacturer’s protocols using a Tmof 50 °C, with a T7-specific forward primer (5 -TAATACGACTCACTATAGGG-3’) (SEQ ID NO: 446) and a tRNA-specific reverse primer (Table 3). Reverse primers were purchased with a C2’ -methoxy modification on the penultimate 5 ’ base to reduce 3’ non-templated nucleotide addition by T7 RNA polymerase.30Primers were purchased from IDT. To verify amplification, sample was mixed with 2 pL of PCR product, 1 pL of 6X Gel Loading Dye (NEB), and 3 pL of nuclease free (NF) water (Ambion) and loaded into a 1% w / v agarose gel containing SYBR Safe Dy e (Apex Bio). The gel was run in IX TAE buffer at 120V for 20 minutes and imaged using a Bio-Rad Gel Doc XR+.
[0134] tRNA templates were in vitro transcribed using HiScribe T7 RNA Synthesis Kit (NEB). 1.8 pL of crude PCR product was mixed with 0.75 pL each of 10X Buffer, 100 mM NTPs, T7 RNA Polymerase, and 3.7 pL NF water. Reactions were set up at 0.75X concentration as recommended by NEB for products < 0.3 kb. Reactions were incubated for 6 hours at 37 °C and stored at -20 °C or -80 °C. Crude in vitro transcription products were either analyzed by denaturing PAGE or used in CFPS reactions.
[0135] Denaturing PAGE
[0136] 12% Urea-PAGE gels were cast using the SequaGel UreaGel 19: 1 Denaturing Gel System (National Diagnostics) using appropriate spacers (1 mm for analytical gels, 3 mm for preparative gels) and an appropriate comb. Gels were allowed to polymerize for 1 - 2 hours at room temperature. In vitro transcription products or cleaved tRNAs (see tRNA aminoacylation assays section) were mixed with an equal volume of 2X RNA Loading Dye (8 M Urea, 2 mM Tris pH 7.5, 2 mM EDTA, 0.004% Bromophenol Blue), denatured at 70 °C for 15 minutes, and cooled on ice for 5 minutes. Samples were loaded onto the gel. along with a low range ssRNA ladder (NEB). The gel was run in IX TBE buffer at 230 V for 2.5 hours or until the dye front had reached thebotom of the gel. For imaging, gels were stained with IX SYBR Gold (Thermo Fisher) for 10 min with gentle shaking.
[0137] CFPS reaction set up
[0138] CFPS reactions w ere set up in 5 pL reactions consisting of 8 mM magnesium glutamate, 10 mM ammonium glutamate, 130 mM potassium glutamate, 1.2 mM ATP, 0.85 mM GTP, 0.85 mM UTP, 0.85 mM CTP, 0.03 mg / mL folinic acid, 0. 17 mg / mL total E. coll tRNA, 0.4 mM NAD, 0.27 mM CoA. 4 mM oxalic acid. 1 mM putrescine, 1.5 mM spermidine. 57 mM HEPES pH 7.2, 2 mM total amino acids, 0.03 M phosphoenolpyruvate, 5 ng / pL plasmid DNA (purified by Zymo DNA midiprep), 30% v / v cell extract, and water. To express tRNAs from crude IVT reactions, 10% v / v crude tRNA product was added. For the Al jannaschii tRNATyr, the BpyRS was added at 1 mg / mL and Bpy was added at 1 mM. For tRNAsPyl, the corresponding PylRSs were added at 1 mg / mL and AzK w as added at 1 mM. The AL barker! tRNA1’5'1was paired with the IPYE chimeric PylRS. ’1the AL. Alvus tRNAl>vlwas paired with the AL alvus PylRS,32and the ' 17VC 10 / / t tRNA1*'1was paired with the Luml PylRS.33API TrpRS and HF2 TrpRS were purified and added at concentrations of 1 mg / mL. aaRS purification protocols are described below. CFPS reactions were incubated at 37 °C for UAG suppression or at 30 °C for UAA and UGA suppression.
[0139] Cell extract preparation
[0140] 759. T7 was used as the chassis strain for cell extract preparation for all experiments because it has been optimized for efficient ncAA incorporation at the UAG codon.34759.T7 was streaked onto an LB-agar plate from a glycerol stock and incubated at 34 °C for ~ 20 hrs. A single colony was inoculated into 100 mL of LB and incubated at 34 °C for ~ 20 hrs wi th 220 rpm shaking. The next day, 5 L of 2xYTPG (16 g / L tryptone. 10 g / L yeast extract, 5 g / L NaCl, 7 g / L dibasic potassium phosphate, 3 g / L monobasic potassium phosphate, 18 g / L glucose) was inoculated at ODeoo = 0.075 with the overnight culture and grown at 34 °C with 220 rpm shaking. At ODeoo = 0.6, IPTG was added to afinal concentration of 1 mM to induce expression of T7 RNA polymerase. Cells were grown until ODeoo = 3.0. Cells were pelleted by centrifugation at 5k x g for 10 min at 4 °C and were washed three times with cold S30 buffer (10 mM Tris-Acetate pH 8.2, 14 mM Mg Acetate, 60 mM K Acetate, 2 mM DTT). Cells were flash frozen in liquid nitrogen and stored at - 80 °C.
[0141] Cells were thawed on ice and resuspended in 0.8 mL S30 buffer / g cells. Cells were sonicated in 1.4 mL aliquots using three 45 sec on / 59 sec off cycles at 50% amplitude for a total of 950 J in a Q125 Sonicator (Qsonica). Lysed cells were centrifuged for 10 min at 12 k x g at 4 °C. Supernatant was removed and a run-off reaction was performed by incubating the supernatant at 37 °C for 1 hour. Extracts were centrifuged for 10 min at 10 k x g at 4 °C to remove insolublecomponents. Supernatant was transferred to a Slide-a-Lyzer (10 kDa MWCo) and dialyzed into 200X volumes of S30 buffer for three hours. Clarified extract was isolated by centrifuging for 10 min at 10 k x g at 4 °C, put into single use aliquots, flash-frozen in liquid nitrogen, and stored at - 80 °C.
[0142] sfGFP quantification
[0143] WT sfGFP and 216X-sfGFP fluorescence was measured by a Bio-Rad Synergy 2 plate reader. Yields were calculated using a standard curve of sfGFP fluorescence vs. sfGFP yields as measured by14C-leucine scintillation counts. WT sfGFP containing14C-leucine was synthesized by adding14C-leucine (Perkin Elmer) at a concentration of 10 pM in CFPS. An equal volume of 0.5 M KOH was added to CFPS reactions to hydrolyzel4C-leucine-acylated tRNA. Samples were spotted onto two Filtermat A (Perkin Elmer) fiberglass paper sheets. After drying, one filtermat was washed 3X with 5% w / v trichloroacetic acid and IX with 100% ethanol. Meltilex A (Perkin Elmer) was applied to both sheets, and scintillation counts were measured using the MicroBeta2. Soluble yields were calculated as described previously.35Dilutions were made in IX PBS and a standard curve was built using linear regression in Microsoft Excel.
[0144] aaRS purification
[0145] Overexpression vectors containing aaRSs with a C-terminal 10X His tag were either purchased from Twist Biosciences in the pET.BCS backbone or were cloned in-house. aaRS sequences were purchased as gBlocks from IDT containing appropriate overhangs for Gibson Assembly into a pET.BCS overexpression vector. Linearized pET.BCS vector was synthesized by PCR using Q5 DNA Polymerase following the manufacturer’s instructions using forward primer (5’-GCAGTAGTGGTCATCATC-3’) (SEQ ID NO: 447) and reverse primer (5’- atgTCCTCCTTATGTGTG-3 ) (SEQ ID NO: 448) at a Tmof 59 °C. PCR amplification was confirmed by agarose gel electrophoresis as described previously. PCR products were column purified using Zymo DNA Clean and Concentrate, resuspended in IX CutSmart Buffer, and digested with Dpnl (NEB) overnight at 37 °C. Gibson Assembly reactions were set up with 25 ng of backbone and a 3X molar excess of the insert, along with 75 mM Tris-HCl pH 7.5, 7.5 mM MgC12, 0.15 mM dNTPs, 7.5 mM DTT, 0.75 mM NAD, 0.004 U / pL T5 Exonuclease, 0.025 U / pL Phusion Polymerase, 4 U / pL Taq DNA Ligase, and 3.125 pg / mL ET SSB. Reactions were incubated at 50 °C for one hour. The entire reaction was transformed into chemically competent NEB 5a cells following the manufacturer’s protocol, plated onto LB-Carb
[0100] agar plates, and incubated overnight at 37 °C. Single colonies were inoculated into 5 mL of LB-Carb
[0100] and grown overnight at 37 °C with 250 RPM shaking. Plasmid DNA was purified using Zymo Miniprep Kits and sequence confirmed.
[0146] Plasmids encoding aaRSs of interest were transformed into BL21 (DE3) Star following manufacturer's protocols and plated onto LB-Carb
[0100] , The next day. a single colony was inoculated into 3 mL of LB-Carb
[0100] and grown at 37 °C until saturated. 250 mL of Overnight Express TB Media (Millipore) was prepared by mixing 15 g of powder, 2.5 mL of glycerol, and 250 mL of water and was sterilized by microwaving until bubbles started to appear. After cooling, 250 pL of Carb
[0100] was added and mixed. 250 pL of saturated culture was mixed into the Overnight Express TB Media and incubated at 37 °C overnight with 250 rpm shaking.
[0147] Cells were pelleted by centrifugation at 5k x g for 10 min in an Avanti J-25 centrifuge and washed with Buffer 1 (300 mM NaCl, 50 rnM monobasic sodium phosphate pH 8.0) containing 10 mM imidazole pH 8.0. Cells were resuspended in Buffer 1 + 10 rnM imidazole pH 8.0 by vortexing and supplemented with Benzonase (Thermo Fisher). Cells were pulled through an 18- gauge syringe needle and lysed by homogenization at -20,000 PSI in an Avestin B3 homogenizer. Cellular debris was pelleted by centrifugation at 20k x g. Supernatant was added to pre-equilibrated Ni-NTA resin (Qiagen) and incubated with end-over-end shaking at 4 °C for 1 hour. Supernatant was removed by centrifugation, and the resin was washed five times with Buffer 1 + 20 mM imidazole pH 8.0. Resin was then packed into a gravity column. Proteins were eluted with 20 mL of Buffer 1 + 0.5 M imidazole pH 8.0 in 1 mL fractions. Protein-containing fractions were identified by measuring A280 on nanodrop and analyzed by SDS-PAGE. 1 uL of eluted protein was mixed with 3.75 pL 4X LDS Sample Buffer. 1.5 pL IM DTT, and water to 15 pL. Samples were denatured at 95 °C for 10 minutes, and 10 pL of sample were loaded onto a 4-12% Bis-Tris NuPAGE gel (Invitrogen). Gel was run in IX MES Buffer at 180 V for 45 minutes. Gels were stained in AcquaStain Protein Gel (Bulldog Bio) for 15 minutes and then imaged to confirm protein size and purity . aaRS-containing fractions were pooled and dialyzed into Buffer 1 + 40% v / v glycerol, with three buffer changes. After dialysis, proteins were quantified by Nanodrop (using molecular weights and exctinction coefficients calculated by ExPasy ProtParam), put into singleuse aliquots, flash-frozen in liquid nitrogen, and stored in -80C.
[0148] Intact protein ESI-MS
[0149] Proteins were synthesized in CFPS reactions scaled up to 50 pL and were purified with Strep-Tactin XT Resin and spin columns as recommended by the manufacturer (IBA Life Sciences). Eluted proteins w ere buffer exchanged into 100 mM ammonium acetate using Amicon Ultra 0.5 mL Centrifugal Filters (10 kDa MWCO). Protein purity7was confirmed by SDS-PAGE.
[0150] Samples were injected on a 1200 HPLC System (Agilent Technologies Inc., Santa Clara, California, USA) onto a Thermo Hypersil-C18 column (3.0 pm. 30 x 2.1 mm) for reverse-phase separation which was maintained at 35 °C with a constant flow rate at 0.400 ml / min, using agradient of mobile phase A (water, 0. 1 % formic acid (v / v)) and mobile phase B (acetonitrile, 0.1% formic acid (v / v)). The gradient program was as follows: 0 - 0.5 min, 1%B; 0.5 - 5 min, 1 - 100%B; 5 - 7.25 min, 100%B; 7.25-7.5 mins, 100 - 1%B; 7.5-10 min, 1%B. ’‘MS-Only”, positive ion mode acquisition was utilized on an Agilent 6230 time-of-flight mass spectrometer equipped Electrospray ionization source (Agilent Technologies Inc., Santa Clara, California, USA). The source conditions were as follows: Gas Temperature, 320 °C; Drying Gas flow, 5 L / min; Nebulizer, 20 psi; VCap. 4500 V; Fragmentor, 210 V; Skimmer, 65 V; and Oct 1 RF, 750 V. The acquisition rate in MS-Only mode was 3 spectra / second, utilizing m / z 922.009798 as reference masses. Data was plotted using a custom python script.
[0151] tRNA aminoacylation assays
[0152] To prepare for tRNA aminoacylation assays, tRNAs were first synthesized in scaled-up in vitro transcription reactions, described in Section 3.5.1, from a DNA template containing a 5’ hammerhead ribozyme construct.13DNA templates were purified using the Zymo DNA Clean and Concentrate kit to ensure efficient transcription. After IVT, crude IVT products were cleaved by adding 5X volumes of NF water and incubating at 60 °C for between 2-8 hours. 0.1X volumes of 3M NaCl and 2.5X volumes of 100% ice-cold ethanol were added to cleavage reactions and incubated at -20 °C for at least 2 hours to precipitate products. Crude product was pelleted by centrifugation at 21k x g for 10 min at 4 °C. Products were resuspended in 200 pL IX RNA Loading Dye, denatured, loaded onto a denaturing Urea-PAGE gel (made with 3 mm spacers), and run in IX TBE at 230V for 2.5 hours. Bands containing cleaved tRNA were visualized by UV shadowing, were excised from the gel, and were crushed and passive eluted into 0.3 M NaCl overnight with end-over-end shaking at 4 °C. Supernatant from passive elution was isolated by centrifugation and was ethanol precipitated. tRNA concentrations were quantified by Nanodrop 2000c.
[0153] For the API and HF2 aaRSs, aminoacylation reactions were set up by in 30 pL reactions consisting of 100 mM HEPES pH 7.4, 4 mM DTT, 10 mM MgC12, 10 mM ATP, 7.5 nM aaRS, 2.5 mM amino acid, 4 U / rnL PPIase (NEB), and 15.6 pM tRNA candidate or 106 pM total E. coli tRNA.36 We chose a concentration of 15.6 pM tRNA because it represents a 10-fold excess of the in vivo concentration of tRNATrp in E. coli cells (943 tRNATrp molecules / cell37 assuming a cell volume of 1 fL (https: / / ecmdb.ca / e_coli_stats)). A 10-fold excess was chosen to ensure observable signal by LC-MS and because in vivo tRNA expression is typically done with a pEVOL vector containing a pl 5 A origin of replication, which has a copy number ~ 10. 106 pM total E. coli tRNA was calculated by assuming 64,274 tRNA molecules / cell37and a cell volume of 1 fL (ecmdb.ca / e_coli_stats). For the metagenomically-identified GlnRS and TyrRS, aminoacylationreactions were set up in 30 pL reactions mimicking CFPS conditions, which consisted of 130 mM potassium glutamate, 1.2 mM ATP. 0.85 mM GTP, 0.85 mM UTP, 0.85 mM CTP, 0.03 mg / mL folinic acid, 0.4 mM NAD, 0.27 mM CoA, 4 mM oxalic acid, 1 mM putrescine, 1.5 mM spermidine, 57 mM HEPES pH 7.2, 2 mM total amino acids, 0.03 M phosphoenolpyruvate, 15.6 pM tRNA or 106 pM E. coli tRNA, 7.5 nM aaRS, and 30% v / v cell extract that was fdtered through an Amicon Ultra 0.5 mL Centrifugal Filters (3 kDa MWCO) to remove large residual aaRSs and tRNAs. The GlnRS and TyrRS were found to be inactive in the previous conditions. For kinetics, 5 pL timepoints were removed and quenched with 1.1X volumes of RNase A solution (200 mM sodium acetate pH 5.2, 1.5 U / pL RNase A (NEB) and incubated for 5 min at room temperature. Reactions were precipitated with 0.1X volumes of 50% w / v trichloroacetic acid and incubated at - 80 °C for at least 30 min. Samples were spun down at 21.000 x g for 10 min at 4 °C, and supernatant was transferred to autosampler vials and analyzed by LC-MS.
[0154] Samples were injected on a 1290 Infinity II UHPLC System (Agilent Technologies Inc., Santa Clara, California, USA) onto a Poroshell 120 EC-C18 column (1.9 pm, 50 x 2.1 mm) (Agilent Technologies Inc., Santa Clara, California, USA) for reverse-phase separation which was maintained at 30 °C with a constant flow rate at 0.500 ml / min, using a gradient of mobile phase A (water, 0.1 % formic acid (v / v)) and mobile phase B (acetonitrile, 0.1% formic acid (v / v)). The gradient program was as follows: 0 - 1 min, 2%B; 1 - 5 min, 2 - 40%B; 5 - 6 min, 40 - 99%B; 6 - 8 mins, 99%B; 8 - 8.10 min, 99 - 2%B; 8.10 - 14 min, 2%B. "MS-Only", positive ion mode acquisition was utilized on an Agilent 6545 quadrupole time-of-flight mass spectrometer equipped with a JetStream ionization source (Agilent Technologies Inc., Santa Clara, California, USA). The source conditions were as follows: Gas Temperature, 300 °C; Drying Gas flow, 12 L / min; Nebulizer, 45 psi; Sheath Gas Temperature, 350 °C; Sheath Gas Flow-, 12 L / min; VCap. 3500 V; Fragmentor, 110 V; Skimmer, 65 V; and Oct 1 RF, 750 V. The acquisition rate in MS-Only mode was 3 spectra / second, utilizing m / z 121.050873 and m / z 922.009798 as reference masses.
[0155] Characterizing sfGFP mutations
[0156] Mutations at T216 in sfGFP were installed using a previously described method, e.g., as described in PCT application no. PCT / US2023 / 069097. 20 sets of mutagenic primers were designed for site-directed mutagenesis of sfGFP T216 and ordered from IDT. All following Tms are calculated by Benchling. The primers were designed to overlap with a Tm of 40-45°C. The forward primer containing the mutation w as designed to have a Tmof 60-62 °C and the reverse primer was designed to have a Tmof 58 °C. Linear templates containing the desired mutation were amplified by Q5 DNA Polymerase in 10 pL PCR reactions using a touchdown PCR, starting at Tmof 72 °C and stepping down -1 °C / cycle until a Tmof 62 °C was reached. Cycles were repeated for a total of 25 cycles.
[0157] PCR products were digested by adding 1 pL of Dpnl to the crude PCR reactions and were incubated for 2 hrs at 37 °C. Reactions were diluted 4X, and 1 pL was added into a 4 pL Gibson Assembly Reaction, as described in Section 3.5.4. Gibson Assembly reactions were incubated at 50 °C for one hour. Reactions were diluted 10X and used in a second PCR reaction to prepare linear expression templates using Q5 DNA Polymerase following the manufacturer's protocol using forward primer (5’- ctgagatacctacagcgtgagc -3') (SEQ ID NO: 449) and reverse primer (5’- cgtcactcatggtgatttctcacttg-3’) (SEQ ID NO: 450).
[0158] Linear templates were then added into CFPS reactions along with 1 pM GamS to protect linear templates and incubated at 37 °C. sfGFP was quantified as previously described.
[0159] RNase P cleavage assays
[0160] In vitro transcription reactions were set up as previously described. 1 pL of IVT reaction was diluted with 80 pL NF water and was split into 2 x 9 pL aliquots. One aliquot was treated with 1 pL of cell extract, and the other was treated with 1 pL of water. Reactions were incubated at 37 °C for one hour and were quenched by addition of RNA Loading Dye. Denaturing Urea PAGE was run as previously described.
[0161] tRNA structure predictions
[0162] tRNA structure predictions were using R2DT (macentral.org / r2dt).38All tRNA structures except TGA-29 were predicted using the default settings. TGA-29 structure was predicted using the Constrained Folding setting on the Full Molecule.
[0163] Table 1: Bioinformatically identified suppressor tRNAs found in the genomes of bacteriophages.
[0165] Table 3: Primers used for IVT template amplification.
[0166] References
[0167] (1) Dumas. A.: Lercher. L.; Spicer, C. D.; Davis, B. G. Designing Logical Codon Reassignment - Expanding the Chemistry in Biology. Chem. Sci. 2015, 6 (1), 50-69.
[0168] (2) Cervettini, D.; Tang, S.; Fried, S. D.; Willis, J. C. W.; Funke, L. F. H.; Colwell, L. J.; Chin, J. W. Rapid Discovery and Evolution of Orthogonal Aminoacyl-TRNA Synthetase- TRNA Pairs. Nat. Biotechnol. 2020. 38 (8), 989-999.
[0169] (3) Zambaldo, C.; Koh, M.; Nasertorabi. F.; Han, G. W.: Chatterjee. A.; Stevens, R. C.; Schultz, P. G. An Orthogonal Seryl-TRNA Synthetase / TRNA Pair for Noncanonical Amino Acid Mutagenesis in Escherichia Coli. Bioorg. Med. Chem. 2020, 28 (20), 115662.
[0170] (4) Chatteijee, A.; Xiao, H.; Schultz, P. G. Evolution of Multiple, Mutually Orthogonal Prolyl-TRNA Synthetase / TRNA Pairs for Unnatural Amino Acid Mutagenesis in Escherichia Coli. Proc. Natl. Acad. Sci. U. S. A. 2012, 109 (37), 14841-14846.
[0171] (5) Ellefson, J. W.; Meyer, A. J.; Hughes, R. A.; Cannon, J. R.; Brodbelt, J. S.; Ellington, A. D. Directed Evolution of Genetic Parts and Circuits by Compartmentalized Partnered Replication. Nat. Biotechnol. 2014, 32 (1), 97-101.
[0172] (6) Italia. J. S.: Addy, P. S.; Wrobel. C. J. J.; Crawford, L. A.; Lajoie, M. J.; Zheng. Y.; Chatterjee, A. An Orthogonalized Platform for Genetic Code Expansion in Both Bacteria and Eukary otes. Nat. Chem. Biol. 2017, 13 (4), 446-450.
[0173] (7) Andrews, J.; Gan, Q.; Fan, C. "Not-So-Popular" Orthogonal Pairs in Genetic Code Expansion. Protein Sci. 2022, e4559.
[0174] (8) Ivanova, N. N.; Schwientek, P.; Tripp, H. J.; Rinke, C ; Pati, A.; Huntemann, M.; Visel, A.; Woyke, T.; Kyrpides, N. C.; Rubin, E. M. Stop Codon Reassignments in the Wild. Science 2014, 344 (6186), 909-913.
[0175] (9) Borges. A. L.; Lou, Y. C.; Sachdeva, R.; ALShayeb, B.; Penev, P. L: Jaffe, A. L.; Lei, S.; Santini, J. M.; Banfield, J. F. Widespread Stop-Codon Recoding in Bacteriophages May Regulate Translation of Lytic Genes. Nature Microbiology 2022, 1-10.
[0176] (10) Chan, P. P.; Lin, B. Y.; Mak, A. J.; Lowe, T. M. TRNAscan-SE 2.0: Improved Detection and Functional Classification of Transfer RNA Genes. Nucleic Acids Res. 2021, 49 (16), 9077-9096.
[0177] (11) Fukunaga, J.-L; Gouda, M.; Umeda, K ; Ohno, S.; Yokogawa, T.; Nishikawa, K. Use of RNase P for Efficient Preparation of Yeast TRNATyr Transcript and Its Mutants. J. Biochem. 2006, 139 (1), 123-127.
[0178] (12) Imburgio. D.; Rong, M.; Ma, K.; McAllister, W. T. Studies of Promoter Recognition and Start Site Selection by T7 RNA Polymerase Using a Comprehensive Collection of Promoter Variants. Biochemistry 2000, 39 (34), 10419-10430.
[0179] (13) Fechter, P.; Rudinger, J.; Giege, R.; Theobald-Dietrich, A. Ribozyme Processed TRNA Transcripts with Unfriendly Internal Promoter for T7 RNA Polymerase: Production and Activity. FEBS Lett. 1998, 436 (1), 99-103.
[0180] (14) Al-Shayeb. B.; Sachdeva. R.: Chen. L.-X.; Ward, F.; Munk, P.; Devoto, A.; Castelle, C. J.; Olm, M. R.; Bouma-Gregson, K.; Amano, Y.; He, C ; Meheust, R.; Brooks, B.; Thomas, A.; Lavy, A.; Matheus-Camevali, P.; Sun, C.; Goltsman, D. S. A.; Borton, M. A.; Sharrar, A.; Jaffe, A. L.; Nelson, T. C.; Kantor, R.; Keren, R.; Lane, K. R.; Farag, I. F.; Lei, S.; Finstad, K.; Amundson, R.i Anantharaman, K.; Zhou, J.; Probst, A. J.; Power, M. E.; Tringe. S. G.; Li, W.-J.; Wrighton, K.; Harrison, S.; Morowitz, M.; Reiman, D. A.; Doudna, J. A.; Lehours, A.-C.; Warren, L.; Cate, J. H. D.; Santini, J. M.; Banfield, J. F. Clades of Huge Phages from across Earth’s Ecosystems. Nature 2020, 578 (7795), 425-431.
[0181] (15) Seki, K.; Galindo, J. L.; Karim, A. S.; Jewett, M. C. A Cell-Free Gene Expression Platform for Discovering and Characterizing Stop Codon Suppressing TRNAs. ACS Chem. Biol. 2023. https: / / doi.org / 10.1021 / acschembio.3c00051.
[0182] (16) Peters, S. L.; Borges, A. L.; Giannone, R. J.; Morowitz, M. J.; Banfield, J. F.; Hettich, R. L. Experimental Validation That Human Microbiome Phages Use Alternative Genetic Coding. Nat. Commun. 2022, 13 (1), 5710.
[0183] (17) Hayashi-Ishimaru, Y.; Ohama, T.; Kawatsu, Y.; Nakamura, K; Osawa, S. UAG Is a Sense Codon in Several Chlorophycean Mitochondria. Curr. Genet. 1996, 30 (1), 29-33.
[0184] (18) Fucikova. K.; Lewis, P. O.; Gonzalez-Halphen. D.; Lewis, L. A. Gene Arrangement Convergence, Diverse Intron Content, and Genetic Code Modifications in Mitochondrial Genomes of Sphaeropleales (Chlorophyta). Genome Biol. Evol. 2014, 6 (8), 2170-2180.
[0185] (19) Ambrogelly, A.; Palioura, S.; Soil, D. Natural Expansion of the Genetic Code. Nat. Chem. Biol. 2007, 3 (1), 29-35.
[0186] (20) Giege, R ; Eriani, G. The TRNA Identity Landscape for Aminoacylation and Beyond. Nucleic Acids Res. 2023. https: / / doi.org / 10.1093 / nar / gkad007.
[0187] (21) Seki, K.; Galindo, J. L.; Jewett, M. C. Orthogonal TRNA Expression Using Endogenous Machinery in Cell-Free Systems. bioRxiv, 2022, 2022.10.04.510903. https : / / doi. org / 10.1101 / 2022.10.04.510903.
[0188] (22) Beyer, J. N.: Hosseinzadeh. P.; Gottfried-Lee. L; Van Fossen, E. M.; Zhu, P.; Bednar, R. M.; Karplus, P. A.; Mehl, R. A.; Cooley, R. B. Overcoming Near-Cognate Suppressionin a Release Factor 1-Deficient Host with an Improved Nitro-Tyrosine TRNA Synthetase. J. Mol. Biol. 2020. 432 (16). 4690-4704.
[0189] (23) Hughes, R. A.; Ellington, A. D. Rational Design of an Orthogonal Tryptophanyl Nonsense Suppressor TRNA. Nucleic Acids Res. 2010, 38 (19), 6813-6830.
[0190] (24) Wan, W.; Huang, Y.; Wang, Z ; Russell, W. K.; Pai, P.-J.; Russell, D. H.; Liu, W. R. A Facile System for Genetic Incorporation of Two Different Noncanonical Amino Acids into One Protein in Escherichia Coli. Angew. Chem. Int. Ed Engl. 2010, 49 (18), 3211—3214.
[0191] (25) O’Donoghue, P.; Prat, L.; Heinemann, I. U.; Ling, J.; Odoi, K.; Liu, W. R.; Soil, D. Near-Cognate Suppression of Amber, Opal and Quadruplet Codons Competes with Aminoacyl- TRNAPyl for Genetic Code Expansion. FEBS Lett. 2012, 586 (21), 3931-3937.
[0192] (26) Hu. Z.; Liang, J.; Su, T.; Zhang, D.; Li. H.; Gao, X.; Yao, W.; Song, X. Minimizing the Anticodon-Recognized Loop of Methanococcus Jannaschii Tyrosyl-TRNA Synthetase to Improve the Efficiency of Incorporating Noncanonical Amino Acids. Biomolecules 2023, 13 (4), 610.
[0193] (27) Lajoie, M. J.; Rovner, A. J.; Goodman. D. B.; Aemi, H.-R.; Haimovich, A. D.; Kuznetsov. G.; Mercer, J. A.; Wang, H. H.; Carr. P. A.; Mosberg, J. A.; Rohland. N.: Schultz, P. G.; Jacobson, J. M.; Rinehart, J.; Church, G. M.; Isaacs, F. J. Genomically Recoded Organisms Expand Biological Functions. Science 2013, 342 (6156), 357-360.
[0194] (28) Ostrov, N.; Landon, M.; Guell, M.; Kuznetsov, G.; Teramoto, J.; Cervantes, N.; Zhou, M.; Singh, K.; Napolitano, M. G.; Moosbumer, M.; Shrock, E.; Pruitt. B. W.; Conway, N.; Goodman, D. B.; Gardner, C. L.; Tyree, G ; Gonzales, A.; Wanner, B. L.; Norville, J. E.; Lajoie, M. J.; Church, G. M. Design, Synthesis, and Testing toward a 57-Codon Genome. Science 2016, 353 (6301), 819-822.
[0195] (29) Fredens, J.; Wang, K.; de la Torre, D.; Funke, L. F. H.; Robertson, W. E.; Christova, Y.; Chia, T.; Schmied, W. H.; Dunkelmann, D. L.; Beranek, V.; UttamapinanL C.; Llamazares, A. G.; Elliott, T. S.; Chin, J. W. Total Synthesis of Escherichia Coli with a Recoded Genome. Nature 2019, 569 (7757), 514-518.
[0196] (30) Kao, C.; Zheng, M.; Rudisser, S. A Simple and Efficient Method to Reduce Nontemplated Nucleotide Addition at the 3 Terminus of RNAs Transcribed by T7 RNA Polymerase. RNA 1999, 5 (9), 1268-1272.
[0197] (31) Bryson, D. L; Fan, C.; Guo, L.-T.; Miller, C.; Soil, D.; Liu, D. R. Continuous Directed Evolution of Aminoacyl-TRNA Synthetases. Nat. Chem. Biol. 2017, 13 (12), 1253-1260.
[0198] (32) Willis, J. C. W.; Chin, J. W. Mutually Orthogonal PyrrolysyLTRNA Synthetase / TRNA Pairs. Nat. Chem. 2018, 10 (8), 831-837.
[0199] (33) Dunkelmann, D. L.; Willis, J. C. W.; Beatie, A. T.; Chin, J. W. Engineered Triply Orthogonal Pyrrolysyl-TRNA Synthetase / TRNA Pairs Enable the Genetic Encoding of Three Distinct Non-Canonical Amino Acids. Nat. Chem. 2020, 12 (6), 535-544.
[0200] (34) Des Soye, B. J.; Gerbasi, V. R.; Thomas, P. M.; Kelleher, N. L.; Jewet, M. C. A Highly Productive, One-Pot Cell-Free Protein Synthesis Platform Based on Genomically Recoded Escherichia Coli. Cell Chem Biol 2019, 26 (12), 1743-1754.e9.
[0201] (35) Swartz. J. R.; Jewet, M. C.; Woodrow, K. A. Cell-Free Protein Synthesis With Prokaryotic Combined Transcription-Translation. In Recombinant Gene Expression: Reviews and Protocols; Baibas, P., Lorence, A., Eds.; Humana Press: Totowa, NJ, 2004; pp 169-182.
[0202] (36) McMurry, J. L.; Chang, M. C. Y. Fluorothreonyl-TRNA Deacylase Prevents Mistranslation in the Organofluorine Producer Streptomyces Catleya. Proc. Natl. Acad. Sci. U. S. A. 2017, 114 (45), 11920-11925.
[0203] (37) Dong, H.; Nilsson, L.; Kurland, C. G. Co-Variation of TRNA Abundance and Codon Usage in Escherichia Coli at Different Growth Rates. J. Mol. Biol. 1996, 260 (5), 649-663.
[0204] (38) Sweeney. B. A.; Hoksza, D.; Nawrocki, E. P.; Ribas, C. E.; Madeira, F.; Cannone, J. J.; Gutell, R.; Maddala, A.; Meade, C. D.; Williams, L. D.; Petrov. A. S.; Chan, P. P.; Lowe, T. M.; Finn, R. D.; Petrov, A. I. R2DT Is a Framework for Predicting and Visualising RNA Secondary Structure Using Templates. Nat. Commun. 2021, 12 (1), 3494.
[0205] Preferred aspects of this invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations of those preferred aspects may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect a person having ordinary' skill in the art to employ such variations as appropriate, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject mater recited in the claims appended hereto as permited by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.
[0206] The steps of the methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The steps may be repeated or reiterated any number of times to achieve a desired goal unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
CLAIMSWe claim:
1. An orthogonal translation system in a bacteria, the system comprising: a polynucleotide encoding a tRNA, the polynucleotide having at least 80 % sequence identity to the polynucleotide sequence of SEQ ID NO: 1; and an aminoacyl-tRNA synthetase having at least 80% sequence identity to the polypeptide sequence of SEQ ID NO: 2, or a polynucleotide encoding the aminoacyl-tRNA synthetase.
2. The method of claim 1 , wherein the bacteria is E. coli.
3. The orthogonal translation system of claim 2, wherein: i) the polynucleotide encoding the tRNA; and ii) the aminoacyl-tRNA synthetase or the polynucleotide encoding the aminoacyl-tRNA synthetase are present in an E. coli lysate or an E. coli cell.
4. The orthogonal translation system of any one of claims 1-3, wherein the tRNA and the aminoacyl-tRNA synthetase can incorporate a tryptophan residue at a UGA codon.
5. The orthogonal translation system of any one of claims 1-3, wherein the tRNA and the aminoacyl-tRNA synthetase can incorporate a noncanonical amino acid at a UGA codon.
6. The orthogonal translation system of any one of claims 1-5. wherein the polynucleotide encoding the tRNA has at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the polynucleotide sequence of SEQ ID NO: 1.
7. The orthogonal translation system of any one of claims 1-6, wherein the aminoacyl-tRNA synthetase has at least 85%, at least 90%, at least 95%, or at least 99% sequence identity' to the polypeptide sequence of SEQ ID NO: 2.
8. A method for incorporating an amino acid at a UGA stop codon, the method comprising: providing a translation template to the orthogonal translation system of any one of claims 1-7.
9. The method of claim 8, wherein the translation template is expressed from a transcription template in an E. coli lysate.
10. The method of claim 9, wherein the E. coli lysate is part of a cell-free protein synthesis system.
11. An orthogonal translation system in a bacteria, the system comprising: a polynucleotide encoding a tRNA, the polynucleotide having at least 80% sequence identity' to the polynucleotide sequence of SEQ ID NO: 3; and an aminoacyl-tRNA synthetase having at least 80% sequence identity' to the polypeptide sequence of SEQ ID NO: 4, or a polynucleotide encoding the aminoacyl-tRNA synthetase.
12. The method of claim 11, wherein the bacteria is E. coli.
13. The orthogonal translation system of claim 12. wherein: i) the polynucleotide encoding the tRNA; and ii) the aminoacyl-tRNA synthetase or the polynucleotide encoding the aminoacyl-tRNA synthetase are present in an E. coli lysate or E. coli cell.
14. The orthogonal translation system of any one of claims 11-13, wherein the tRNA and the aminoacyl-tRNA synthetase can incorporate a glutamine residue at a UAG codon.
15. The orthogonal translation system of any one of claims 11-13. wherein the tRNA and the aminoacyl-tRNA synthetase can incorporate a noncanonical amino acid at a UAG codon.
16. The orthogonal translation system of any one of claims 11-15, wherein the polynucleotide encoding the tRNA has at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the polynucleotide sequence of SEQ ID NO: 3.
17. The orthogonal translation system of any one of claims 11-15. wherein the aminoacyl-tRNA synthetase has at least 85%, at least 90%, at least 95%, or at least 99% sequence identity' to the polypeptide sequence of SEQ ID NO: 4.
18. A method for incorporating an amino acid at a UAG stop codon, the method comprising: providing a translation template to the orthogonal translation system of any one of claims 11-17.
19. The method of claim 18, wherein the translation template is expressed from a transcription template in an E. coli lysate.
20. The method of claim 19, wherein the E. coli lysate is part of a cell-free protein synthesis system.21 . A kit for identifying a candidate orthogonal tRNA to an organism, the kit comprising: one or more transcription templates, wherein each of the one or more transcription templates comprises a polynucleotide encoding a suppressor tRNA; a cell-free protein synthesis (CFPS) system derived from the organism; and a translation template encoding a reporter protein, wherein the translation template comprises a premature stop codon.
22. The kit of claim 21, wherein the reporter protein is 216X-sfGFP, wherein X is UAG, UAA. or UAG.
23. The kit of claim 21 or 22, wherein the suppressor tRNA comprises an RNase P tag; and wherein the CFPS comprises RNase P.
24. The kit of any one of claims 21-23, wherein the CFPS system comprises an E. coliIvsate.
25. A method for identifying a candidate orthogonal tRNA to an organism, the method comprising: transcribing one or more transcription templates in vitro to produce one or more transcription products, wherein each of the one or more transcription templates comprises a polynucleotide encoding a premature suppressor tRNA; and incubating the one or more transcription products with a cell-free protein synthesis (CFPS) system derived from the organism, and a translation template encoding a reporter protein having a premature stop codon; performing an assay that detects the reporter protein; wherein the CFPS system comprises RNase P; and wherein if the reporter protein is not detected, the suppressor tRNA is a candidate orthogonal tRNA to the organism.
26. The method of claim 25, wherein each of the one or more transcription templates are linear.
27. The method of claim 25 or 26, wherein transcribing the one or more transcription templates is performed with a T7 RNA polymerase.
28. The method of any one of claims 25-27, wherein the reporter protein is 216X- sfGFP, wherein X is UAG, UAA, or UAG.