A translation-independent directed evolution strategy to engineer aminoacyl-TRNA synthetases
The START method addresses the challenge of incorporating noncanonical monomers into proteins by using a translation-independent directed evolution strategy to identify mutant aaRSs capable of charging these monomers, achieving efficient and selective incorporation into proteins in living cells.
Patent Information
- Application Number
- PCT/US2024/059491
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
Current directed evolution strategies for altering the substrate specificity of aminoacyl-tRNA synthetases (aaRSs) rely on ribosomal translation, which is incompatible with noncanonical monomers that are poor substrates for the ribosome, limiting the incorporation of structurally divergent monomers into proteins in living cells.
A translation-independent directed evolution strategy, termed START, which selectively identifies mutant aaRSs capable of charging noncanonical amino acids by correlating sequence barcodes on tRNAs with their cognate aaRSs, allowing for the enrichment and identification of aaRS mutants without relying on ribosomal translation.
The START method enables the systematic engineering of aaRSs to efficiently charge a wider variety of noncanonical monomers with high fidelity and efficiency, overcoming the limitations of existing technologies and facilitating the incorporation of exotic monomers into proteins in living cells.
Smart Images

Figure US2024059491_19062025_PF_FP_ABST
Abstract
Description
A TRANSLATION-INDEPENDENT DIRECTED EVOLUTION STRATEGY TO ENGINEER AMINOACYL-TRNA SYNTHETASES RELATED APPLICATIONS
[0001] This application claims the benefit under 35 USC 119(e) of U.S. Provisional Application No.63 / 609,135, filed on December 12, 2023, which is incorporated herein by reference in its entirety. GOVERNMENT SUPPORT
[0002] This invention was made with Government support under contract number CHE 2002182, awarded by National Science Foundation. The Government has certain rights in the invention. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] A Sequence Listing conforming to the rules of WIPO Standard ST.26 is hereby incorporated by reference. Said Sequence Listing has been filed as an electronic document via Patent Center encoded as XML in UTF-8 text. The electronic document, created on December 9, 2024, is entitled “0342.0016US1_ST26.xml”, and is 137,419 bytes in size. BACKGROUND OF THE INVENTION
[0004] Over the last two decades, the genetic code of cells in various domains of life has been expanded to include hundreds of noncanonical amino acids (ncAAs) with diverse chemical structures.1-4This has been accomplished by developing engineered aminoacyl-tRNA synthetase / tRNA pairs that selectively incorporate various ncAAs in response to a unique codon (e.g., a repurposed nonsense codon).1-5The ability to site-specifically incorporate enabling ncAAs into proteins using this technology has unlocked powerful new ways to probe and manipulate protein function for both basic science and biotechnology applications.1-3, 6, 7
[0005] Despite such exciting progress, this technology has been largely restricted to the incorporation of simple α-amino acids and, to a lesser extent, some α-hydroxy acids.8-12Pioneering early work using in vitro translation systems has demonstrated that the ribosome has the potential to accept a much wider variety of monomers for translation such as D-α-amino acids,13β- and γ- amino acids,14-17long-chain amino acids,18aminobenzoic acid,19α-aminoxy and α-hydrazino acids,20α-thio acids,21and others.22-24The remarkable tolerance of the ribosome for such noncanonical monomers, and further expansion of its noncanonical substrate scopethrough engineering25-28offer a possible path for repurposing mRNA-templated ribosomal translation to synthesize non-peptide polymers.
[0006] Although incorporation of structurally divergent monomers into peptides has been broadly explored in vitro, examples of achieving the same in living cells remain scarce. The key bottleneck limiting progress on this front is the lack of engineered aminoacyl-tRNA synthetases (aaRSs) capable of efficiently charging such monomers. In the handful of examples, where such exotic monomers were incorporated into proteins in E. coli, such as β3-p-Bromo- homophenylalanine26and a β2-hydroxy acid,29tRNA acylation relied on promiscuous recognition that certain wild-type aaRSs exhibited fortuitously for these substrates. The ability to systematically engineer aaRSs to charge a wider variety of noncanonical monomers with high fidelity and efficiency is needed to overcome this limitation. However, current directed evolution strategies available to alter the substrate specificity of aaRSs rely on the expression of reporter proteins with a selectable phenotype (such as antibiotic resistance, or fluorescence; Figure 1A).30-39Unfortunately, such translation-dependent selection schemes are incompatible with noncanonical monomers that are poor substrates for the ribosome. SUMMARY OF THE INVENTION
[0007] Although incorporation of structurally divergent monomers into peptides has been broadly explored in vitro, examples of achieving the same in living cells remain scarce. The key bottleneck limiting the progress on this front is the lack of engineered aaRSs capable of efficiently charging such monomers. In the handful of examples, where such exotic monomers were incorporated into proteins in E. coli, such as β3-p-Bromo-homophenylalanine26and a β2- hydroxy acid,29tRNA acylation relied on promiscuous recognition that certain wild-type aaRSs fortuitously exhibited for these substrates. The ability to systematically engineer aaRSs to charge a wider variety of noncanonical monomers with high fidelity and efficiency is needed to overcome this limitation.
[0008] However, established directed evolution strategies for altering the substrate specificity of aaRSs rely on the ribosomal translation of reporter proteins with a selectable phenotype (such as antibiotic resistance, or fluorescence; Figure 1A).30-39Unfortunately, such translation-dependent selection schemes are incompatible with noncanonical monomers that are poor substrates for the ribosome. Although the ribosome could be engineered to accept novel monomers, such efforts would require access to engineered aaRSs capable of generating monomer-charged tRNAs. Engineering of both the aaRS and the ribosome at the same time ischallenging, as it would require co-localization of two rare mutants from each library in the same cell, which is associated with low statistical odds. Consequently, a two-step solution is needed, where the aaRS is first engineered to charge the noncanonical monomer without relying on the translational readout, which can then be used for ribosome engineering.
[0009] Described herein is a general approach for genetically-engineering aminoacyl-tRNA synthetases (aaRSs) to modify or alter their substrate specificity for amino acids. Specifically, these engineered aminoacyl-tRNA synthetases (also referred to herein as variant, or mutated / mutant aaRSs) have the enzymatic activity to use noncanonical amino acids as their substrates and to preferentially and selectively charge corresponding tRNA molecules with those noncanonical amino acids. Noncanonical amino acids (ncAAs) or Unnatural amino acids (UAAs / Uaa) are non-proteogenic amino acids that are either found naturally in organisms or are synthetically made in a laboratory. As used herein the term noncanonical or unnatural amino acids includes non-α-amino acids. In non-alpha amino acids the amine group is displaced further from the carboxylic acid end of the amino acid molecule.
[0010] As described herein, the general method of altering substrate specificity of the aaRS involves direct selection of tRNA acylation without ribosomal translation. The method of the present invention is described in terms of the “START” process as shown in Figure 1B, and Figure 21. (START is the acronym for Selects tRNA-Acylation without Ribosomal Translation). The novel approach of the present invention builds on a reported strategy to enrich acylated tRNA molecule pools from living cells, where the 2′,3′-dihydroxy functionality at the 3′- terminus of uncharged tRNAs is selectively oxidized using a suitable oxidizing agent, such as periodate.40-47In the reported strategy, the acylated tRNAs are protected from this damage, and can be subsequently tagged with a unique oligonucleotide sequence at its intact 3′-terminus, allowing their subsequent identification through selective reverse-transcription and PCR (RT- PCR).42-47Although this strategy is effective for enriching and characterizing acylated tRNA sequences, it is incomplete as it does not reveal the identity of the aaRS responsible for the acylation. To identify the enzymatically active, acylating aaRS, it is essential to somehow connect the identity of each aaRS variant to its cognate tRNA sequence.
[0011] The present invention described herein achieves this result by introducing a unique / distinct sequence barcode within a permissive site of the tRNA. As described herein, a “permissive site” is a site that does not affect (e.g., interfere, inhibit or compromise) the ability of the engineered tRNA molecule to recognize its cognate mRNA codon, nor affect the site of enzymatic aminoacylation of the engineered tRNA molecule by aminoacyl-tRNA synthetase(aaRS). Aminoacylation, as used herein, is the attachment / ligation / charging of an amino acid to a tRNA as catalyzed by aminoacyl-tRNA synthetases.
[0012] Importantly, the genetically-engineered, nucleotide sequence barcoded tRNA molecule retains its biological activity to recognize a specific codon for an amino acid, or a nonsense codon, and maintains the accessibility of site required to covalently bind an amino acid (e.g., specifically a non-naturally-occurring / noncanonical amino acid (ncAA) or non-α-amino acid) via aminoacyl-tRNA synthetase. Using Methanomethylophilus alvus pyrrolysyl- tRNA / ((tRNAMaPyl) as the model system, the feasibility of inserting a sequence barcode in a tRNA anticodon loop with limited impact on tRNA expression and charging is demonstrated. Next, the START selection scheme was optimized to allow robust enrichment of an acylated barcoded tRNAMaPyl, from a mixed population that also contained an uncharged tRNAMaPylwith a distinct barcode. An active site library of M. alvus pyrrolysyl-tRNA synthetase (MaPylRS) was created, where each unique mutant is encoded by associated tRNAMaPylbarcodes. This library was subjected to the START scheme described herein to identify MaPylRS mutants capable of charging different ncAAs. Finally, it is demonstrated that the START selection scheme is compatible with non-α-amino acid monomers. START represents the first translation- independent directed evolution platform for aaRS engineering, which will be invaluable for developing incorporation systems for unique noncanonical monomers.
[0013] It is important to note that the novel method described herein is a general approach and not limited to the exemplified M. alvus pyrrolysyl-tRNA (tRNAMaPyl). This novel method is designed to identify mutated aminoacyl-tRNA synthetases using tRNA / aaRS pairs derived from other organisms such as E.coli.
[0014] In general, according to one aspect of the present invention, is a method to identify mutant / variant aminoacyl-tRNA synthetases (aaRS), wherein the method is independent of ribosomal translation (also referred to herein as a ribosomal translation-independent method), that selectively incorporate non-canonical amino acids into proteins, wherein each mutated / variant aaRS comprises a nucleotide sequence barcode corresponding to, or correlated with a tRNA molecule containing the same nucleotide sequence barcode. This sequence barcode serves as a unique identifier characteristic to identify the enzymatically active aaRSAs capable of charging cognate tRNAs with noncanonical amino acids. As used herein, the terms “mutated”, “mutant” or” variant” all refer to proteins or nucleic acids with alterations, differences, substitutions, insertions, deletions or deviations from the naturally-occurring or wild-type sequence of the protein or nucleic acid, for example, in the nucleotide sequence encoding theprotein, or amino acid sequence of the protein). The term “derived from” as used herein refers to the mutant proteins or nucleic acids that originate from wild-type proteins or nucleic acids, i.e., the original wild-type proteins or nucleic acids have mutations such as alterations, differences, substitutions, insertions, deletions or deviations introduced into their wild-type sequences, but essentially retain the 80%, 85%, 90%, 95%, 98% or 99% of the wild-type sequence.
[0015] One step of the present method comprises acylation / charging of the barcoded tRNA molecule by the mutant / variant aminoacyl-tRNA synthetase resulting in protection of the barcoded tRNA molecules from oxidation treatment. If the barcoded tRNA molecule is acylated by its cognate aminoacyl-tRNA synthetase it is protected from oxidation damage when contacted with a suitable oxidizing reagent such as sodium periodate. Other suitable oxidizing reagents are known to those of skill in the art.
[0016] After the oxidative treatment, the undamaged, acylated barcoded tRNA molecules are de-acylated to remove the aaRS from the tRNAs, and the 3’ end of barcoded tRNA molecules is then extended by annealing an extension nucleotide sequence and subjected to reverse transcription PCR (RT-PCR) to enrich the tRNA molecule population. Because each mutant aaRS is tagged with a unique sequence barcode, as is each tRNA, the enriched tRNA population and the mutant aaRSs are sequenced by Next Generation Sequencing (NGS) and the sequence of the barcodes matched to identify the aaRS mutant that acylated the corresponding / correlated tRNA molecules.
[0017] This method can identify aminoacyl-tRNA synthetases that have the specific activity to charge noncanonical amino acids such as pyrrolysine, lysine or phenylalanine analogs or a non-α-amino acids or non-α-amino acid analogs.
[0018] Specifically, the method of the present invention includes the steps comprising combining a library of sequence barcoded tRNA molecules and a library of mutant / variant aminoacyl-tRNA synthetases (aaRS) into a single DNA molecule, wherein the sequences of the mutant / variant aminoacyl-tRNA synthetases correspond to the sequence barcodes of the tRNA molecules under suitable conditions in a cell for expression of the tRNA molecules and the aaRSs, wherein the tRNA molecules are acylated by the mutant aminoacyl-tRNA synthetases in the presence of the non-canonical amino acid; subjecting the resulting combination of acylated tRNA molecules and unacylated tRNA molecules to oxidation conditions whereby the acylated tRNA molecules are protected from oxidation; de-acylating the unoxidized, acylated tRNA molecules in the combination; extending the 3’ end of the deacylated tRNA molecules to attachan adapter sequence; annealing a complementary nucleotide sequence to this 3’-adapter sequence of the tRNA molecules; selectively reverse transcribing and PCR amplifying the 3’-adapter- containing tRNA sequences; and sequencing the amplified tRNA sequences and the mutant aaRSs to determine the barcode sequence and identify and match the sequence barcode of the tRNA with the corresponding aaRS mutant with specific activity to aminoacylate noncanonical amino acids.
[0019] More specifically, the present invention encompasses a method of identifying specific active mutant / variant aminoacyl-tRNA synthetase (aaRS) molecules in the absence of ribosomal translation, which can acylate a cognate tRNA with noncanonical monomers, the steps comprising: a.) Providing a gene library of cognate tRNA molecules, wherein the tRNA molecules are tagged with a unique oligonucleotide sequence (barcoded) in a manner that does not jeopardize their expression, folding, and aminoacylation by the aaRS; b.) Providing a gene library of mutant / variant aminoacyl-tRNA synthetase (aaRS) molecules, in the same plasmid as the tRNA library from step a.), wherein each aaRS variant is correlated to a barcode sequence of a tRNA molecule of the tRNA library of step a.); c.) Co-expressing the library of aaRS molecules with the correlated barcode-containing tRNA molecules under conditions specific for charging the tRNA molecules with a cognate aaRS molecule to form a combination of tRNA molecules charged with the desired noncanonical monomer and uncharged tRNA molecules; d.) Contacting the combination of step c.) with a periodate oxidation buffer solution under conditions suitable for the oxidation of uncharged tRNA molecules thereby forming a mixture of oxidized uncharged tRNA molecules and unoxidized charged tRNA molecules; e.) Deacylating the mixture of step d.) to obtain tRNA molecules with an intact 3’-terminus; f.) Introducing a unique DNA oligonucleotide sequence at the intact 3’-terminus of the tRNA from step e.), either by ligation, or by DNA polymerase mediated extension of the 3’-terminus of the tRNA after hybridizing it with a suitable DNA template;g.) Isolating the tRNA molecules containing the unique DNA oligonucleotide sequence at the 3’- terminus of step f.) by non-native PAGE or other methods (e.g., oligonucleotide or streptavidin mediated pulldown); h.) Amplifying the isolated tRNA molecules containing the unique DNA oligonucleotide sequence at the 3’-terminus of step f.) by RT PCR using an oligonucleotide primer that selectively binds the unique DNA oligonucleotide sequence at the 3’-terminus; i.) Sequencing the tagged, amplified tRNA molecules to obtain the sequence of the tRNA molecules and the corresponding barcodes; and j.) Correlating the sequence of the barcodes of the enriched tRNA population with the sequence of the variant aaRS molecule to identify specific active mutant / variant aminoacyl-tRNA synthetase (aaRS).
[0020] In one embodiment of the present invention, to prevent the different aaRS and tRNA mutants from cross-reacting with each other in step c.) of the present method, each unique DNA molecule may be spatially separated before expressing them. This can be accomplished by the step of putting them in E. coli, which picks up a distinct plasmid with a unique aaRS-tRNA combination. However, one can also potentially do this using in vitro translation, where each plasmid is encapsulated in a microscopic droplet with translation machinery.
[0021] The present invention also encompasses a genetically-engineered tRNA molecule, wherein the tRNA molecule comprises or contains a unique barcode sequence introduced or inserted into the tRNA molecule nucleotide sequence at a permissive site in the tRNA molecule, wherein the barcode sequence does not interfere with aminoacylation activity by aminoacyl- tRNA synthetases or the binding of the tRNA to its designated codon sequence of the mRNA molecule.
[0022] In the present invention, the tRNA anti-codon loop is mutated / modified to specifically bind to (e.g., recognize) a codon i.e., an amber codon (UAG / TAG), ochre codon (TAA / TAG) or opal codon (UGA / TGA) of a mRNA molecule. These codons, also known as nonsense codons or stop codons typically signal the end of protein synthesis. However, tRNA molecules that bind to the amber, ochre or opal codons act as suppressor molecules and suppressthe stop signal so that protein synthesis can continue and a noncanonical amino acid can be incorporated instead.
[0023] Specifically, the barcode sequence is then introduced into the anti-codon loop of the tRNA molecule, thus expanding the anti-codon loop sequence region of the tRNA beyond the recognition codon. Importantly, the sequence barcode dose not interfere or inhibit the binding of the tRNA to the mRNA codon, nor affect aminoacylation or charging of the tRNA by its cognate aaRS. The barcode sequence inserted into the anti-codon loop of the tRNA is between about 5 nucleotides and about 20 nucleotides in length. Typically, the barcode sequence is about 15 nucleotides in length. Examples of sequences suitable for use as a barcode are described herein, see for example SEQ ID NOs.65 to 72.
[0024] Any tRNA molecule is suitable for use in the START procedure. As exemplified herein, the genetically-engineered barcoded tRNA molecule is derived from tRNAMaPyl, wherein the tRNAMaPylcomprises at least about 80%, 85%, 90%, 95%, 98% or 99% sequence identity with the full-length nucleotide sequence of SEQ ID NO:61. The sequence barcode is inserted into the anti-codon region of the tRNA molecule. As used herein, the term “derived from” means that the wild-type tRNA sequence is mutated in manner that inserts the sequence barcode in the anti-codon region of the wild-type tRNA, but the barcode tRNA molecule retains essentially the same wild-type sequence comprising at least about 80%, 85%, 90%, 95%, 98% or 99% sequence identity with the full-length sequence of the wild-type tRNA , for example SEQ ID NO:61, but with the barcode incorporated therein.
[0025] The present invention further encompasses a variant / mutant M. alvus pyrrolysyl- tRNA aminoacyl synthetase (MaPylRS), or a composition comprising a variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS), wherein the variant MaPylRS preferentially aminoacylates a M. alvus pyrrolysyl-tRNA (tRNAMaPyl) with a noncanonical amino acid over the naturally-occurring pyrrolysine, lysine or phenylalanine amino acid, wherein the wild-type M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) is derived from SEQ ID NO:58. The variant MaPylRS comprises the amino acid sequence of SEQ ID NO:58, or an amino acid sequence with at least about 80%, 85%, 90%, 95%, 98% or 99% sequence identity with the full- length wild-type SEQ ID NO:58, wherein the MaPylRS is mutated relative to SEQ ID NO:58 at amino acid residues leucine (L or Leu) 125; asparagine (N or Asn) 166, and valine (V or Val) 168. Also encompassed are nucleotide sequences that encode the mutant MaPylRSs described herein.
[0026] Specifically, the mutant MaPylRS comprises SEQ ID NO:58, or an amino acid sequence with at least about 80%, 85%, 90% , 95%, 98% or 99% sequence identity with the full- length SEQ ID NO:58, wherein the leucine (L) at position 125 is replaced with valine (V) or alanine (A) or threonine (T); the asparagine (N) at position 166 is replaced with serine (S) or alanine (A) or threonine (T); and the valine (V) at position 168 is replaced with cysteine (C) or tryptophan (W) or threonine (T).
[0027] More specifically, the mutant MaPylRS comprises an amino acid sequence selected from the group consisting of : mutant VSV SEQ ID NO: 75; mutant LSV SEQ ID NO: 76; mutant LAC SEQ ID NO: 77; mutant LAV SEQ ID NO:78; mutant AAV SEQ ID NO: 79; mutant ASV SEQ ID NO:80; mutant TTW SEQ ID NO: 81, or mutant AST SEQ ID NO:82.
[0028] The variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) of the present invention has biological enzymatic activity to charge its cognate / corresponding tRNA with a noncanonical amino acid, wherein the noncanonical amino acid is a pyrrolysine, lysine or phenylalanine analog or a non-α-amino acid or non-α-amino acid analog. Specifically, the noncanonical amino acids are selected from the group consisting of: Nℇ-((tert-butoxy)carbonyl)- L-lysine / Nℇ-Boc-L-Lysine (BocK); OH-BocK; H-BocK; m-iodo-L-phenylalanine (mIF); o-nitro- L-phenylalanine (cNF) or o-cyano-L-phenylalanine (2-CNF).
[0029] Also encompassed by the present invention is a cell(s) comprising the genetically- engineered t-RNA molecule, specifically an M. alvus pyrrolysyl-tRNA (tRNAMaPyl. Also encompassed is a cell(s) comprising variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS). Specific sequences of M. alvus pyrrolysyl-tRNA (tRNAMaPyland variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) are found herein. The cell can be a prokaryotic cell such as E.coli, or can be an eukaryotic cell, and in particular a mammalian cell. Choice of cell for expression is based on the specific aaRS-tRNA pair to be used. Cells encompassed by the present invention also include the additional components or nucleic acid sequences required fort the cells to produce / express proteins comprising the noncanonical amino acids described herein. Such components are described herein and are known to those of skill in the art.
[0030] Also encompassed by the present invention are methods of producing a protein or peptide of interest in a prokaryotic (e.g., E.coli) eukaryotic or mammalian cell with one, or more, noncanonical amino acid analogs (e.g., pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs) at specified amino acid residue positions in the protein or peptide. The protein orpeptide is genetically-engineered to incorporate stop / nonsense codons (e.g., amber, ochre or opal) into the desired specific site or sites of incorporation in the protein or peptide. A nucleic acid encoding the protein is introduced into a competent cell (e.g., via plasmids) growing in suitable culture medium. A nucleic acid encoding, for example, a tRNAMaPylthat recognizes the nonsense codon(s) in the protein, and a nucleic acid encoding, for example, a mutant MaPylRS that aminoacylates the tRNAMaPyl;are also introduced into the cell and cultured under suitable conditions for growth and expression of the proteins.
[0031] The cell culture medium is contacted with one, or more, pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs under conditions suitable for incorporation of the one, or more, pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs into the protein at the site or sites of the stop / nonsense codon(s), and the protein / peptide is expressed with the noncanonical amino acid at the specified sites, thereby producing the protein or peptide of interest with one, or more noncanonical amino acid analogs (e.g., pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs) at specified sites / positions in the protein or peptide.
[0032] In one embodiment noncanonical amino acid analog is a pyrrolysine, lysine, phenylalanine or non-α-amino acid analog selected from the group consisting of: Nε-Boc-L- lysine (BocK), m-iodo-L-phenylalanine (mIF), OH-BocK, H-BocK, o-nitro-L-phenylalanine (oNF) and o-cyano-L-phenylalanine (2-CNF). Other structurally similar noncanonical amino acid analogs are known to those of skill in the art and are suitable for use in the present invention.
[0033] In a further embodiment, the mutant MaPylRS is selected from the group consisting of: mutant VSV SEQ ID NO: 75; mutant LSV SEQ ID NO: 76 mutant LAC SEQ ID NO: 77; mutant LAV SEQ ID NO: 78; mutant AAV SEQ ID NO: 79; mutant ASV SEQ ID NO:80; mutant TTW SEQ ID NO:81, or mutant AST SEQ ID NO:82.
[0034] Further encompassed by the present invention are kits comprising variant aaRS and tRNA molecules as described herein. Specifically encompassed herein is a kit for producing a protein or peptide of interest in a cell, wherein the protein or peptide comprises one, or more noncanonical amino acid analogs. The components of the kit comprise a container containing a polynucleotide sequence encoding an tRNAMaPylthat recognizes a nonsense / stop codon; and a container containing a polynucleotide encoding a variant MaPylRS that preferentially aminoacylates the tRNAMaPylwith a noncanonical amino acid analog, wherein the mutant MaPylRS comprises SEQ ID NO:58, or an amino acid sequence with at least about 80%, 85%,90% , 95%, 98% or 99% sequence identity with the full-length SEQ ID NO:58, wherein the leucine (L) at position 125 is replaced with valine (V) or alanine (A) or threonine (T); the asparagine (N) at position 166 is replaced with serine (S) or alanine (A) or threonine (T); and the valine (V) at position 168 is replaced with cysteine (C) or tryptophan (W) or threonine (T).
[0035] More specifically, the mutant MaPylRS comprises an amino acid sequence selected from the group consisting of : mutant VSV SEQ ID NO: 75; mutant LSV SEQ ID NO: 76; mutant LAC SEQ ID NO: 77; mutant LAV SEQ ID NO:78; mutant AAV SEQ ID NO: 79; mutant ASV SEQ ID NO:80; mutant TTW SEQ ID NO: 81, or mutant AST SEQ ID NO:82.
[0036] The variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) of the kit has biological enzymatic activity to charge its cognate / corresponding tRNA with a noncanonical amino acid, wherein the noncanonical amino acid is a pyrrolysine, lysine or phenylalanine analog or a non-α-amino acid or non-α-amino acid analog. Specifically, the noncanonical amino acids are selected from the group consisting of: Nℇ-((tert-butoxy)carbonyl)- L-lysine / Nℇ-Boc-L-Lysine (BocK); OH-BocK; H-BocK; m-iodo-L-phenylalanine (mIF); o-nitro- L-phenylalanine (cNF) or o-cyano-L-phenylalanine (2-CNF).
[0037] Also included in the kit are the desired noncanonical amino acid analogs, buffers, diluents and other components necessary for producing a protein or peptide of interest with site- specific incorporation of one, or more noncanonical amino acid analogs as described herein. The kit can further comprise instructions for producing the protein or peptide of interest.
[0038] As produced as described herein, the proteins or peptides of interest with noncanonical amino acid analogs incorporated into the proteins or peptides can further comprise bioconjugation handles that allow site-specific labeling of recombinant proteins. These can be used to generate therapeutically relevant protein conjugates (e.g., antibody-drug conjugates) or research reagents (e.g., antibody-fluorophore conjugates). Others are electrophilic amino acids that can create crosslink with natural nucleophilic amino acids, which is useful for many applications, including the ability to make genetically encoded cyclic peptides in cells. Such genetically encoded cyclic peptides can be evolved to develop novel therapeutics.
[0039] The above and other features of the invention including various novel details of construction and combinations of parts, and other advantages, will now be more particularly described with reference to the accompanying drawings and pointed out in the claims. It will be understood that the particular method and device embodying the invention are shown by way of illustration and not as a limitation of the invention. The principles and features of this inventionmay be employed in various and numerous embodiments without departing from the scope of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In the accompanying drawings, reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale; emphasis has instead been placed upon illustrating the principles of the invention. Of the drawings:
[0041] FIGs.1A-B. Directed evolution strategies for engineering the substrate-specificity of aaRSs. A) Established selection strategies depend on ribosomal translation to connect the activity of aaRS to the expression of a reporter protein with a selectable phenotype. This is incompatible with noncanonical monomers that are poor substrates for the ribosome. B) In START, members in an aaRS mutant library are individually correlated to a sequence barcode within the cognate tRNA. Successful charging by a mutant aaRS protects its cognate barcoded tRNA from a periodate-mediated oxidation, enabling their subsequent enrichment via templated 3′-extension and RT-PCR.
[0042] FIGs.2A-E. The anticodon loop of tRNAMaPylcan be expanded to include a sequence barcode. A) The 3′-extension method used to selectively tag tRNAMaPylwith an intact 3′-end, using a 3′-Cy5-labeled DNA template. Native PAGE followed by fluorescence imaging shows that the designed Cy5-DNA template selectively hybridizes with tRNAMaPyl, and enables the formation of the extension product. B) A tRNAMaPylmutant with a severely expanded anticodon loop (SEQ ID NO:59) can be effectively charged by MaPylRS as revealed by the tRNA-extension assay. In the absence of tRNA acylation, periodate treatment prevents the formation of the extension product, but when co-expressed with MaPylRS and a cognate ncAA substrate (BocK), robust formation of extension product was observed, indicating successful acylation of the extended tRNAMaPyl. C) Scheme for characterizing the expression and potential aminoacylation of the members of the tRNAMaPyl barcode library. D) A large majority of the barcoded tRNAMaPyllibrary members are successfully acylated in the presence of MaPylRS and BocK, as revealed by protection from periodate oxidation in the tRNA-extension assay. E) Illumina sequencing of the barcoded tRNAMaPyllibrary DNA (left), after it is expressed in E. coli (middle), and members that survive periodate treatment when co-expressed with MaPylRS in the presence of BocK (right), reveal similar distribution and composition. These analyses indicate that the majority of the barcoded tRNAMaPylsequences are successfully expressed and acylated by MaPylRS.
[0043] FIGs.3A-C. START enables enrichment of an acylated barcoded tRNAMaPylfrom a defined mixed population. A) The scheme of the experiment showing two distinct barcoded tRNAMaPylare separately expressed, where one is charged with BocK (barcode sequence SEQ ID NO:59) and the other is not (barcode sequence SEQ ID NO:60). A defined mixture of these two populations are subjected to the START scheme, and their relative abundance is quantitatively measured in the presence and absence of the periodate-mediated selection using amplicon sequencing. B) PAGE followed by fluorescence imaging of the tRNA extension reaction of this mixed population in the presence or absence of periodate oxidation. C) Amplicon sequencing reveals that the acylated tRNAMaPylbarcodes are enriched approximately 11-fold upon periodate selection.
[0044] FIGs.4A-P. Selection of ncAA-selective MaPylRS mutants using START. A) The active site of MaPylRS, where the three key residues randomized to generate the library are highlighted. B) Scheme for generating and selecting the MaPylRS library, where each mutant is correlated to distinct barcoded tRNAMaPylsequences. C) The structure of BocK. D) PAGE followed by fluorescence imaging of the extension reaction on the library selection in the presence and absence of BocK. E) Top 10% most enriched sequence barcodes that show at least 2-fold higher enrichment in the presence of BocK (relative to no BocK). Enrichment of each barcode in the presence (blue) or the absence (gray) of BocK are shown. Clone VSV (SEQ ID NO:75); Clone LSV (SEQ ID NO:76); Clone LAC (SEQ ID NO:77); Clone LAV (SEQ ID NO:78); Clone AAV (SEQ ID NO:79); Clone ASV (SEQ ID NO: 80). F) WebLogo analyses of the residues observed in the active sites of the MaPylRS mutants correlated to these top 10% barcodes reveal the wild-type MaPylRS sequence as the most prevalent. G) The structure of mIF. H) PAGE followed by fluorescence imaging of the extension reaction on the library selection in the presence and absence of mIF. I) Top 10% most enriched sequence barcodes that show at least 2-fold higher enrichment in the presence of mIF (relative to no mIF). Enrichment of each barcode in the presence (blue) or the absence (gray) of mIF are shown. J) WebLogo analyses of the residues observed in the active sites of the MaPylRS mutants correlated to these top 10% barcodes reveal that the prevalent signature is distinct from that of the wild-type MaPylRS. K) Selected clones for further characterizations, based on their average enrichment levels, the number of associated barcodes that show enrichment, and similarity to the consensus sequence of the most enriched sequences (panel J). L) The mIF-charging activity of each MaPylRS mutant, upon co-expression with tRNACUAMaPyl, measured as the expression of a sfGFP-151-TAG reporter in the presence / absence of mIF and normalized to the expression of awild-type sfGFP reporter. M) ESI-MS analysis of purified sfGFP-151-TAG reporter for LSV and LAC mutants in the presence of mIF show masses consistent with the incorporation of mIF at the TAG codon. In the absence of mIF, the LAC mutant likely charges phenylalanine, as indicated by the mass of the isolated reporter protein. N) The structures of oNF and oCNF. O) MaPylRS clones selected using START (See Figures 13 and 14 for details) to be characterized for charging oNF and oCNF. Clone TTW (SEQ ID NO:81); Clone AST (SEQ ID NO:-82); Clone ASV (SEQ ID NO:80) same as ASV above. P) ESI-MS analysis of purified sfGFP-151-TAG reporters expressed using these mutants in the presence of the appropriate ncAA show masses consistent with the incorporation of oNF or oCNF at the TAG codon.
[0045] FIGs.5A-B. START scheme is compatible with non-α-amino acid monomers. A) The structures of BocK, OH-BocK, and H-BocK. B) Native PAGE followed by fluorescence imaging of tRNA-extension reaction performed on tRNAMaPylco-expressed with MaPylRS in the presence of BocK, OH-BocK, or H-BocK, or in the absence of any cognate substrate, and with or without a periodate treatment. Successful formation of extension product in the presence of each substrate, which are known substrates for MaPylRS, confirms the compatibility of this selection with non-α-amino acid monomers.
[0046] FIGs.6A-B. The design of a 3′-Cy5-labeled DNA template for selective 3′-extension of tRNAMaPyl(SEQ ID NOS: 61) is the nucleotide sequence of wild-type tRNAMaPyland (SEQ ID NO:62) is the extension primer sequence).
[0047] FIG.7. The tRNA-3’-extension strategy differentiates charged and uncharged tRNAMaPyl. When tRNAMaPylis coexpressed with MaPylRS in the presence of BocK, it is protected from periodate oxidation and can form the extension product. However, in the absence of MaPylRS, the uncharged tRNA is oxidized by periodate and is unable to form an extension product.
[0048] FIG.8. The majority of barcodes observed in the tRNAMaPylgene library are also found when the expressed tRNAs are sequenced following RT-PCR. Barcodes are organized by their observed abundance within the expressed tRNA pool (orange dots), and the corresponding abundance within the gene pool (gray dots) are co-plotted. Out of the >600,000 sequenced barcodes observed in the gene pool, >500,000 are also found within the expressed tRNA pool.
[0049] FIGs.9A-C. Optimization of the DNA template length for efficient separation of the tRNA-extension product from other tRNAMaPylspecies (lower molecular weight unextendedcounterpart, and the higher molecular weight band, which is likely a pre-tRNA that was not appropriately processed to generate mature tRNA).
[0050] FIGs.10A-B. A) Two distinct barcoded tRNAMaPylused to create a defined mixture of charged and uncharged species. One of these were co-expressed with MaPylRS in the presence of BocK (barcode sequence (SEQ ID NO:63), while the other was expressed in the absence of its cognate synthetase and BocK (barcode sequence (SEQ ID NO:64). B) The tRNA- extension assay shows that the charged tRNAMaPyl is protected from periodate oxidation, as shown by the formation of the extended product, whereas the extension of the uncharged tRNA is sensitive to periodate treatment.
[0051] FIGs.11A-B. Long-read PacBio sequencing of the tRNA-barcoded MaPylRS reveals that: A) the majority of the unique MaPylRS mutants are associated to more than one distinct barcode, with the average being ~13 barcodes / MaPylRS mutant (note: the wild-type tRNA-MaPyl sequence was removed from this analysis due to its overabundance), and B) nearly all barcodes are correlated to a unique MaPylRS.
[0052] FIG.12. Criteria used to select promising individual MaPylRS mutants from the selected pool for further characterization.
[0053] FIGs.13A-E. Selection of oNF-selective MaPylRS mutants. A) The structure of oNF. B) PAGE followed by fluorescence imaging of the extension reaction on the library selection in the presence and absence of oNF. C) Top 10% most enriched sequence barcodes that show at least 2-fold higher enrichment in the presence of oNF (relative to no oNF). Enrichment of each barcode in the presence (blue) or the absence (gray) of oNF are shown. D) WebLogo analyses of the residues observed in the active sites of the MaPylRS mutants correlated to these top 10% barcodes. E) Criteria used to select promising individual MaPylRS mutants from the selected pool for further characterization for oNF charging.
[0054] FIGs.14A-E. Selection of oCNF-selective MaPylRS mutants. A) The structure of oCNF. B) PAGE followed by fluorescence imaging of the extension reaction on the library selection in the presence and absence of oCNF. C) Top 10% most enriched sequence barcodes that show at least 2-fold higher enrichment in the presence of oCNF (relative to no oCNF). Enrichment of each barcode in the presence (blue) or the absence (gray) of oCNF are shown. D) WebLogo analyses of the residues observed in the active sites of the MaPylRS mutants correlated to these top 10% barcodes. E) Criteria used to select promising individual MaPylRS mutants from the selected pool for further characterization for oCNF charging.
[0055] FIG.15: The plasmid sequence of pBK lpp-tRNAMaPyl(lpp promoter is highlighted in blue, tRNAMaPylis highlighted in red, metG promoter is highlighted in purple, MaPylRS is highlighted in brown or BOLD, glnS promoter is highlighted in orange, sfGFP is highlighted in green). (SEQ ID NO:53)
[0056] FIG.16: The plasmid sequence of pBK lpp -tRNAMaPyl-metG-MaPylRS (lpp promoter is highlighted in blue, tRNAMaPylis highlighted in red, metG promoter is highlighted in purple, MaPylRS is highlighted in brown or BOLD, glnS promoter is highlighted in orange, sfGFP is highlighted in green.) (SEQ ID NO:54)
[0057] FIG.17: The plasmid sequence of pBK lpp-tRNAMaPyl-glnS-MaPylRS (lpp promoter is highlighted in blue, tRNAMaPylis highlighted in red, metG promoter is highlighted in purple, MaPylRS is highlighted in brown or BOLD, glnS promoter is highlighted in orange, sfGFP is highlighted in green.) (SEQ ID NO:55)
[0058] FIG.18: The plasmid sequence of pEvol T5-lac-sfGFP-151-TAG (lpp promoter is highlighted in blue, tRNAMaPylis highlighted in red, metG promoter is highlighted in purple, MaPylRS is highlighted in brown or BOLD, glnS promoter is highlighted in orange, sfGFP is highlighted in green.) (SEQ ID NO:56)
[0059] FIG.19: The wild-type M.alvus Pyrrolysyl Aminoacyl-tRNA Synthetase Nucleotide Sequence. (SEQ ID NO:57)
[0060] FIG.20: The wild-type M.alvus Pyrrolysyl Aminoacyl-tRNA Synthetase Amino Acid Sequence UniProtKB M9SC49_METAX. (SEQ ID NO:58)
[0061] FIG.21: START Synopsis: Selects tRNA-Acylation without Ribosomal Translation DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0062] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. Also, all conjunctions used are to be understood in the most inclusive sense possible. Thus, the word "or" should be understood as having the definition of a logical "or" rather than that of a logical "exclusive or" unless the context clearly necessitates otherwise. Further, the singular forms and the articles "a", "an" and "the" are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms: includes, comprises, including and / or comprising, when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations,elements, components, and / or groups thereof. Further, it will be understood that when an element, including component or subsystem, is referred to and / or shown as being connected or coupled to another element, it can be directly connected or coupled to the other element or intervening elements may be present.
[0063] It will be understood that although terms such as “first” and “second” are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. Thus, an element discussed below could be termed a second element, and similarly, a second element may be termed a first element without departing from the teachings of the present invention.
[0064] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0065] A chemical strategy to select for charged tRNAs
[0066] A translation-independent directed evolution strategy would ideally allow the selection of active aaRS mutants simply based on their ability to acylate a cognate tRNA with a desired noncanonical monomer. Upon aaRS-catalyzed acylation of a tRNA on its 3′-terminal ribose, it becomes protected from oxidation by sodium periodate, which selectively oxidizes vicinal diol groups. This method has long been leveraged to characterize the charging status of tRNAs in cells.40-47Recently, this strategy has been combined with modern enrichment and sequencing methods to perform such analyses in a more high-throughput manner.42, 43, 45-47After the selective periodate-mediated oxidation of the uncharged tRNAs, a unique oligonucleotide sequence can be introduced onto the intact 3′-terminus of the acylated population either through ligation,43, 45, 46or by polymerase mediated extension of the 3′ terminus using a DNA template that is hybridized onto the tRNA sequence.44The installed oligonucleotide sequence can be subsequently used to selectively reverse-transcribe and PCR amplify the charged tRNA sequences.
[0067] The methods described herein for selective enrichment of acylated tRNAs have the potential to serve as the foundation for a translation-independent selection strategy for tRNA acylation. The M. alvus derived pyrrolysyl-tRNA synthetase (MaPylRS) / tRNAMaPylpair is usedas a model system.48The pyrrolysyl pair represents the leading platform for expanding the genetic code with both ncAAs and non-α-amino acid monomers.1-3, 9, 24, 29, 49A 3′-fluorophore- labeled DNA template (Figure 6) was designed that will hybridize with tRNAMaPyland allow the extension of its 3′-terminus by DNA polymerase (Figure 2A). This template selectively hybridized with tRNAMaPylexpressed in E. coli, as observed by native PAGE followed by fluorescence imaging; no such complex was found when total RNA from E. coli not expressing the tRNAMaPylwas used instead (Figure 2B). Incubation of this DNA:tRNA complex with Klenow fragment of E. coli DNA polymerase and dNTPs led to successful extension of the 3′- terminus, as shown by an upward shift in the PAGE analysis. Treatment of tRNAMaPylwith sodium periodate prevented the appearance of this extension product, consistent with the oxidation of its free 3′-terminus (Figure 7). However, when tRNAMaPylwas co-expressed with MaPylRS in the presence of a known ncAA substrate (Nε-Boc-L-lysine; BocK), it was protected from periodate-mediate damage, and successfully yielded the extension product (Figure 7). These results confirm that the tRNA-extension strategy can be used to selectively introduce a oligonucleotide tag at the 3′-terminus of the acylated tRNAMaPylpopulation, which may be subsequently used to enrich these by RT-PCR.
[0068] Expanding the anticodon loop of tRNAMaPylto insert a sequence barcode
[0069] To use this strategy for evolving MaPylRS, its sequence information must somehow be encoded within tRNAMaPyl, such that the identity of variants responsible for tRNA-acylation can be retrieved by sequencing the acylated tRNAMaPylpool, following their enrichment. This can be achieved by introducing a sequence barcode within a permissive site of tRNAMaPyl. The anticodon loop of tRNAMaPylis a promising location to introduce the barcode, given pyrrolysyl synthetases do not interact with the anticodon region.50-55Indeed, the anticodon of pyrrolysyl- tRNAs have been altered to enable suppression of other nonsense codons, as well as four-base codons, indicating significant plasticity.53-56
[0070] To explore if the anticodon loop of tRNAMaPylcan tolerate more dramatic expansions, a variant was generated where the native sequence in the anticodon loop was replaced with a random 15-nucleotide sequence (Figure 2B). There is no technical limit on barcode length per se, but from a practical perspective, the diversity of the sequence space encoded in a barcode of a certain length needs to be considered. Thus, the shortest barcodes that are practically useful is probably about 5 (~1000 possibilities), while the longest would be about 20 (10^12 possibilities).
[0071] The resulting expanded tRNAMaPylmutant was co-expressed in E. coli with MaPylRS either in the presence or absence of a cognate ncAA substrate (BocK). Total RNA isolated from these cells were subjected to tRNA-extension assay with or without periodate treatment (Figure 2B). The successful expression of the expanded tRNAMaPylmutant was confirmed by the appearance of the tRNA:DNA hybrid band. When incubated with DNA polymerase and dNTPs, the expanded tRNAMaPylsuccessfully yielded the extension product in the absence of periodate treatment. MaPylRS was found to efficiently acylate the expanded tRNAMaPylvariant with BocK, as the formation of the extension product was largely insensitive to periodate when the ncAA was included in the growth medium, but almost fully disappeared in its absence (Figure 2B).
[0072] Confirmation that similar insertion of a broad variety of sequences is tolerated in the anticodon loop of tRNAMaPylwithout significantly compromising its expression and charging, which is essential to access a sufficiently diverse barcode that is capable of encoding a MaPylRS mutant library, is described herein. To this end, the anticodon loop of tRNAMaPylwas replaced with a 10 fully randomized nucleotide sequence to create a library of roughly 106mutants (Figure 2C). Illumina sequencing (approximately 5x106usable reads) of the resulting plasmid DNA encoding this tRNA library revealed the presence of >600,000 unique barcodes (Figure 2E). Each nucleotide was well-represented at each position of the barcode library (Figure 2E). Greater than 90% of barcodes were found to be present at only 21 counts or less, indicating a good diversity and distribution of the barcode sequences. To evaluate the expression pattern of the barcoded tRNAMaPylgenes, this plasmid library was transformed into E. coli, the expressed tRNAMaPylsequences were amplified by RT-PCR, and analyzed by Illumina sequencing. Over 500,000 unique sequences were identified in the expressed tRNAMaPylbarcode library, and the representation of each nucleotide in the barcode was comparable to the DNA library (Figure 2E; Figure 8). These observations confirm that the majority of the barcoded tRNAMaPylsuccessfully express in E. coli.
[0073] Finally, to explore if MaPylRS is able to charge the barcoded tRNAMaPylsequences, the library also co-expressed with MaPylRS in E. coli in the presence of BocK. The resulting barcoded tRNAMaPylpool was oxidized by periodate to remove the uncharged population, and the acylated sequences were subjected to tRNA-extension using a 3′-Cy5-labeled DNA template (Figure 2C). Analysis of this extension reaction by native PAGE followed by fluorescence imaging (Figure 2d) showed that most of the tRNAMaPylpool was protected from periodate oxidation in the presence of BocK (but not when the ncAA was absent), indicating that a largemajority of barcoded tRNAMaPylsequences are efficiently charged by MaPylRS. This was further confirmed by amplifying the charged tRNAMaPylpopulation by RT-PCR, and analyzing them by Illumina sequencing. >500,000 unique barcode sequences were identified, and their composition and distribution was found to be similar to the library at the DNA and expressed tRNA level (Figure 2E). Together, these experiments show that barcoding tRNAMaPylin the anticodon loop is a viable approach that does not drastically disrupt expression or charging by MaPylRS.
[0074] Optimization of START to enrich charged, barcoded tRNAMaPylfrom a mixed population
[0075] The basic steps of the START scheme involve periodate oxidation to damage the uncharged tRNAs, selective DNA-templated extension of the charged tRNAs to introduce a unique sequence at the 3′-terminus, followed by its use to amplify these sequences by RT-PCR. To improve the enrichment of the desired population, a gel purification step was employed to isolate the extended DNA:RNA hybrid from the unextended counterparts based on its different mobility on native PAGE. Different lengths of the DNA template were evaluated to alter the extended product size and optimize its separation from other species (Figure 9). Next, for evaluating the performance of this selection scheme, two standard plasmids were created to represent an ‘active’ or ‘inactive’ member in the library (Figure 3A). Each plasmid encoded a distinct barcoded tRNAMaPyl, but one of these also encoded the wild-type MaPylRS, which will acylate the corresponding barcoded tRNAMaPylin the presence of BocK, and represent an active library member. The other plasmid lacked MaPylRS and would represent an inactive library member. These plasmids were separately transformed into E. coli, and using the aforementioned tRNA-extension assay we confirmed that the tRNAMaPylvariant co-expressed with MaPylRS was protected from oxidative damage in the presence of BocK, whereas its counterpart lacking MaPylRS was not (Figure 10). In order to quantify the degree of enrichment afforded by START, the mock inactive and active tRNAMaPylpopulations were mixed in 10:1 ratio, and the composition of these mixed population was characterized by amplicon sequencing (~50,000 reads) before and after subjecting it to the START scheme (Figure 3A-B). The tRNAMaPylvariant co-expressed with MaPylRS was found to undergo roughly an 11-fold enrichment upon enrichment (Figure 3C), suggesting that this strategy should allow differentiation between barcoded tRNAMaPylvariants associated with active or inactive MaPylRS mutants.
[0076] Identification of ncAA-selective MaPylRS mutants using START
[0077] Next, it was determined that START can be used to identify MaPylRS mutants with altered substrate specificities from a naïve library. Three key residues (Leu125, Asn166, and Val168) were simultaneously randomized around the substrate binding pocket of MaPylRS to NNK codons (Figure 4A), creating a library with a theoretical diversity of 32,768. Into the plasmid encoding the MaPylRS library, was further introduced the barcoded tRNAMaPyllibrary containing an (N)10randomized sequence at the anticodon loop (Figure 4B). The resulting library was covered using approximately 1.5x105transformants to ensure that A) each barcode corresponds to a distinct MaPylRS mutant, and B) each MaPylRS is associated with multiple distinct barcodes. Since this library utilizes only ~15% of the possible barcode sequences, the chances of two different MaPylRS mutants receiving the same barcode are very low. Additionally, as barcode diversity in the resulting library far exceeds the diversity of the MaPylRS library (roughly 50-fold), each distinct MaPylRS mutant should be associated with multiple unique barcodes, which would be beneficial to reduce the risk of false positives.
[0078] To characterize this library, and correlate each unique MaPylRS mutant to the associated sequence barcodes, PacBio®HiFi long-read DNA sequencing was used. From about 1.2x105usable reads retrieved from the sequencing analysis, the presence of >93% of all possible MaPylRS mutants was confirmed. It was also found that >92% of unique MaPylRS mutants were represented by more than one barcode, with the average being about 13 barcodes / mutant (Figure 11A). Furthermore, approximately 93% of the barcodes were found to encode a unique MaPylRS mutant (Figure 11B), which minimizes potential confusion when establishing genotype-phenotype connections.
[0079] Following characterization, the barcoded MaPylRS library was transformed into E. coli for subsequent selection. It is important to note that many MaPylRS active site mutants may charge one of the 20 canonical amino acids, which will render the associated barcoded-tRNAs insensitive to periodate oxidation-mediated selection. Traditional aaRS engineering strategies typically employ a negative selection step to deplete such cross-reactive mutants.1, 30, 31, 33, 34However, using the deep sequence coverage from NGS, it was possible to characterize the enrichment pattern of nearly all mutants present in the MaPylRS library in response to the START scheme, and doing so in parallel in the presence and the absence of a desired ncAA would help identify mutants that were selectively enriched in the presence of the ncAA (Figure 4B). A similar approach was recently used for orthogonal tRNA evolution in mammalian cells.57
[0080] Using this strategy, the barcoded MaPylRS library was selected for its ability to charge two distinct ncAAs, BocK (Figures 4C-F) and mIF (Figures 4G-J), performed induplicates. After selection, the surviving barcodes were amplified and characterized by Illumina sequencing. To identify barcodes associated with ncAA-selective MaPylRS mutants, we filtered the NGS data based on three key criteria: A) it must be represented in the collection of barcodes obtained by long-read sequencing characterization of the original library, B) it must show at least a 2-fold enrichment upon selection in the presence of ncAA, and C) its relative enrichment in the presence of the ncAA must be at least 2-fold higher than in the absence. The surviving barcodes were arranged in a decreasing order of average enrichment in the presence of the ncAA (Figure 4E and 4I), and the top 200 barcodes were used to retrieve the sequences of the corresponding MaPylRS mutants. These sequences were used to generate a WebLogo sequence map that represents the relative abundance of observed amino acid residues at each of the three randomized positions within this selected mutant pool. The most enriched sequences for the BocK-selection matched with the wild-type MaPylRS (Figure 4F), which is expected, since it is a great substrate for this synthetase. In contrast, the most enriched mutants for the mIF-selection had a distinct sequence signature (Figure 4J). The following criteria was used to select six distinct MaPylRS mutants (Figure 4K, Figure 12) from this pool of enriched sequences for individually assessing the ability to charge mIF: A) having multiple associated barcodes that show enrichment, B) high average enrichment observed across the correlated barcodes, and C) similarity to the consensus sequence of the most enriched MaPylRS mutants. These six MaPylRS mutants were co-expressed in E. coli with tRNACUAMaPyl, and assessed for their ability to express sfGFP-151-TAG in the presence and absence of mIF (Figure 4L). Gratifyingly, all six mutants exhibited robust reporter expression in the presence of mIF (between 20-50% wild-type sfGFP). Although four mutants (VSV, LSV, AAV, and ASV) exhibited low activity in the absence of mIF, LAC and LAV maintained significant (albeit significantly less than +mIF) reporter expression. ESI-MS analysis of the purified reporter protein expressed using the LSV confirmed cleaned incorporation of mIF, when it was supplied in the growth medium (Figure 4M). Although the LAC mutant show significant reporter expression in the absence of any ncAA, it was found to selectively charge mIF when it was present during expression (Figure 4M). This behavior that has been previously documented for other engineered aaRSs as well.58-60We also isolated the reporter protein expressed using LAC in the absence of mIF, and observed a mass consistent with the incorporation of phenylalanine (Figure 4M).
[0081] Using the same MaPylRS library, the START protocol was applied to identify MaPylRS mutants for charging ncAAs with two different ortho-substituted phenylalanine derivatives, o-nitro-L-phenylalanine (oNF) and o-cyano-L-phenylalanine (2-CNF). Followingthe same selection and analysis methods as described above, we identified TTW and AST as potential top candidates for charging oNF (Figure 13A-E), and ASV for charging oCNF (Figure 14A-E). All three mutants MaPylRS mutants enabled robust expression of the sfGFP-151-TAG reporter in the presence of the appropriate ncAAs, and the ESI-MS analysis of the purified reporter protein confirmed selective incorporation of oNF (for TTW and AST) and oCNF (for ASV). These experiments further confirm that START can be used to identify novel aaRS mutants that accept noncanonical monomers as substrates from naïve libraries.
[0082] START is compatible with non-α-amino acid monomers
[0083] Since the selection system reported described herein relies solely on an aaRS acylate its cognate tRNA, it should be applicable to develop mutant aaRSs that charge noncanonical monomers beyond the α-amino acids. It has been shown that the pyrrolysyl synthetases do not strongly recognize the α-amino group of its native substrate.9This feature has enabled the use of this family of aaRSs to charge analogs of its native substrate, where the α-amino group is replaced with other functionalities, including α-hydroxyacids, desamino-acids, β2-hydroxyacids, etc., without further synthetase engineering.10-12, 24, 26, 29Using this demonstrated polyspecificity of native MaPylRS, additional experiments were performed to explore if the core selection system is compatible with compatible with such non-α-amino acid monomers. E. coli cells co- expressing MaPylRS and tRNAMaPylwere treated with BocK, OH-BocK, H-BocK (Figure 5A), or no noncanonical monomers, and the resulting tRNA population was subjected to the tRNA- extension assay with or without a periodate treatment (Figure 5B). As expected, in the absence of any cognate substrate, tRNAMaPylwas uncharged, and was susceptible to periodate oxidation, which prevented tRNA extension. The presence of BocK, OH-BocK, and H-BocK each facilitated the formation of the tRNA-extension product even with periodate treatment, confirming that the charging of this noncanonical monomers protects the tRNA from periodate oxidation, and that this strategy may be potentially used to identify aaRS mutants selective for structurally diverse noncanonical monomers.
[0084] Summary
[0085] There has been considerable interest in repurposing the mRNA-templated polypeptide synthesis by the ribosome to create novel sequence-defined polymers with noncanonical molecular architectures. However, translational incorporation of novel noncanonical monomers necessary for this purpose would require engineered variants of both aaRSs and the ribosome. Established aaRS engineering strategies are reliant on ribosomaltranslation, and vice versa, thereby creating an interdependence that prevents systematic introduction of structurally novel noncanonical monomers into the genetic code. Described herein is a solution to this stalemate by developing the first directed evolution strategy for engineering aaRSs, designated START, which does not rely on ribosomal translation. It is now demonstrated that tRNAs can be equipped with sequence barcodes, which can be correlated to individual mutants in a large aaRS libraries through long-read next-generation sequencing. Acylation of the correlated partner tRNA by a novel aaRS mutants protects it from a periodate- mediated oxidation, allowing their subsequent enrichment. The identity of the corresponding aaRS mutant can then be retrieved from the barcode sequence. The efficacy of this strategy was demonstrated by identifying novel mutants of MaPylRS selective for two ncAAs from a naïve library. The compatibility of this selection system with non-α-amino acids was further demonstrated using the polyspecificity of MaPylRS for such substrates. It should be possible to extend the strategy described here to other aaRS / tRNA pairs. Although many aaRSs recognize the anticodon of its cognate tRNA as an identity element, which may complicate the introduction of a barcode in this region, it has been possible engineer aaRS anticodon binding domains to be permissive for non-natural expanded anticodon sequences.61, 62Furthermore, such sequence barcodes can also be inserted into alternative locations within the tRNA to avoid aaRS identity elements. Indeed, it has been shown that aaRSs can effectively acylate significantly altered versions of their cognate tRNAs, including severely miniaturized versions.63-66In summary, START offers an exciting opportunity to develop aaRS mutants to charge structurally divergent noncanonical monomers, and enable the ribosomal synthesis of evolvable, sequence-defined polymers with unprecedented structure and function.
[0086] Without further elaboration, it is believed that one skilled in the art can, based on the above description, utilize the present invention to its fullest extent. The following specific embodiments and examples are, therefore, to be construed as merely illustrative, and not limitative of the remainder of the disclosure in any way whatsoever.
[0087] MATERIAL AND METHODS:
[0088] Noncanonical Amino Acids
[0089] Nε-Boc-L-Lysine was purchased from Chem-Impex International.2-Amino-3-(3- iodophenyl) propanoic acid (m-iodo-L-phenylalanine) was purchased from Ambeed Inc. Boc-6- aminohexanoic acid (H-BocK) was purchased from Sigma-Aldrich. OH-BocK was a gift from Prof. Alanna Schepartz.
[0090] General:
[0091] For all cloning, the E. coli DH10B electrocompetent strain was used for transformation as well as plasmid propagation, and the cells were cultured using Luria broth (LB) for solid and liquid media. All PCRs were carried out using Phusion Hot Start II DNA Polymerase (Thermo Scientific) according to the manufacturer’s instructions. Restriction enzymes and T4 DNA ligase were purchased from New England Biolabs. Gibson assembly HiFi mix was purchased from Fisher Scientific. All custom DNA oligos including randomized primers and tREX probes were purchased from Integrated DNA Technologies (IDT). All other oligos were purchased from GENEWIZ. Sanger sequencing was performed by GENEWIZ. Amplicon sequencing was performed by Quintara Biosciences. Whole plasmid sequencing was performed by Plasmidsaurus. Long-read PacBio sequencing was performed by Quintara Biosciences.
[0092] Construction of extended anticodon tRNA mutants:
[0093] The two extended anticodon tRNAs were cloned into pBK lpp-tRNAMaPyland pBK lpp-tRNAMaPyl-metG-MaPylRS plasmids using overlap extension cloning. The two overlapping fragments were generated with lpp_BamHI_F + MaPylT_15AC_R, MaPylT_15AC_F + tRNA_NotI_R. The two fragments were joined together with an overlap extension PCR with terminal primers lpp_BamHI_F and tRNA_NotI_R. The amplified product was purified with 1% agarose gel and digested with BamHI / NotI enzymes, followed by ligation with T4 DNA ligase into the pBK vector to generate pBK lpp- tRNAMaPyl-15AC1 and pBK lpp- tRNAMaPyl-15AC1- metG-MaPylRS. The insertion of the extended tRNA was confirmed via Sanger sequencing.
[0094] For extended anticodon 15AC2, the two fragments were generated using lpp_BamHI_F + MaPyl_15AC2_iR, MaPyl_15AC2_iF + tRNA_NotI_R. The constructs were cloned as described herein.
[0095] Construction of 10-nucleotide (nt) extended anticodon tRNA library:
[0096] The 10-nt extended anticodon tRNA library (1,048,576 or ~106variants) was constructed using overlap extension similarly as described above. The primers used to generate the overlapping fragments were lpp_BamHI_F + MaPyl_lib_R, MaPyl_lib_10N_F + tRNA_NotI_R. The amplified insert was purified with 1% agarose gel, digested with BamHI / NotI and ligated into 5 μg of pBK lpp-tRNAMaPyl-15AC1-metG-MaPylRS vector, replacing the tRNAMaPyl15AC1. The ligation product was ethanol precipitated overnight with DNA: 3M NaOAc (pH 5.2): yeast tRNA (10 mg / mL): 100% cold ethanol in 1: 0.1: 0.008: 3 parts. The mixture was spun down at 20,000 xg for 20 mins at 4℃, washed with 70% ethanol, spun down again for 10 mins. The supernatant was dumped, the pellet was air-dried for 15 mins and finally, resuspended in 10 μL of ddH2O. The ethanol precipitated DNA was transformed into 500 μL DH10B electrocompetent cells, recovered for 2 hours, spun down, resuspended in 1 mL LB, and plated on 15 cm dishes. The library was covered by ~5*109distinct cfus (~5000-fold coverage). These colonies were pooled and stored as glycerol stocks for further analyses.
[0097] Construction of tRNA barcoded-MaPylRS library:
[0098] To construct the tRNA barcoded-MaPylRS library, first the MaPylRS library (32,768 variants) was generated using overlap extension. Three different fragments A, B and C were generated with the following set of primers: metG_NotI_F + MaPylRS_123_iR, MaPylRS_L125NNK_iF + MaPylRS_L165_iR and MaPylRS_N166, V168 NNK_iF + MaPylRS_NcoI_R. The pieces A and B were joined together using a primerless-overlap extension PCR. Subsequently, another primerless overlap extension PCR was performed to join the piece AB with piece C. The amplified insert was purified with 1% agarose gel, digested with NotI / NcoI and ligated into 1.3 μg of pBK lpp-tRNAMaPyl-15AC1-metG-MaPylRS. The ligation product was ethanol precipitated as described above and later transformed into 250 μL DH10B electrocompetent cells, recovered for 2 hours, spun down, resuspended in 1 mL LB, and plated on 15 cm dishes. The library was covered with ~4*107distinct cfus (~1000-fold coverage). The diversity of the sequencing was analyzed by sequencing 20 different clones as well as Amplicon sequencing.
[0099] After characterizing the MaPylRS library, the 10-nt extended anticodon tRNA library was cloned into the pBK vector containing MaPylRS library as described previously. The 10-nt extended anticodon tRNA library was capped at 50-fold coverage (~1.6 * 106variants) to minimize the number of barcodes corresponding to multiple MaPylRS mutants.
[0100] NGS Analysis for input library:
[0101] NGS Analysis for long-read PacBio sequencing was performed with the following script: github.com / cpsoni74 / PacBio_seq_analysis / tree / d5e50606fd13cb241a1ec8d4cd27bf5c670498b1
[0102] The reads were first aligned to the template using the mapping tool minimap2 and the reads that misaligned were discarded. The reads were then filtered based on data quality using Phred score (Q-score) as the metric. Only reads in which each base of the tRNA barcode had a Q-score >30 were considered. Subsequently, all the MaPylRS sequences corresponding to these reads were extracted and the randomized regions were subjected to a minimum Q-score threshold of 20. These reads were collected and written to a comma-separated values (csv) file. The graphs for number of barcodes / unique mutant and number of unique mutants / barcode were generated using bcpermut.py and aaRSperbc.py files, respectively. For more detailed instructions on using the NGS analysis pipeline, please refer to the Readme.md file in the link above.
[0103] Total RNA isolation:
[0104] DH10B cells were transformed with pBK lpp-tRNAMaPyl-15AC1-metG-MaPylRS. Cells were grown overnight in 20 mL LB broth in the presence or absence of 1 mM ncAA and appropriate antibiotics. Cells were spun down at 5000 xg for 10 mins at 4 ℃. The pellet was resuspended in 1 mL TRIzol™ reagent (Thermo Fisher) and incubated for 5 mins at RT. 200 μL of chloroform was added to the mixture, mixed thoroughly and incubated for 3 mins at RT. The solution was spun down at 12000 xg for 15 min at 4 ℃. The clear supernatant was collected and mixed with 500 μL isopropanol. Then, the solution was incubated at 4 ℃ and spun down at 12000 xg for 10 min. The pellet was resuspended in 1 mL 70% ethanol and spun down at 7500 xg for 5 min. The pellet was air dried and resuspended in 50 μL ddH2O. The total RNA concentration was determined using a Nanodrop spectrophotometer and the integrity was assessed with A260 / A280.
[0105] NaIO4 oxidation:
[0106] 10 μg of total RNA was diluted up to 152 μL with oxidation buffer (50 mM NaOAc, 150 mM NaCl, 10 mM MgCl2, 0.1 mM EDTA, pH 4.8) and 8 μL of 100 mM NaIO4 (oxidation) or NaCl (no oxidation) was added68. The mixture was incubated at 25 ℃ for 60 mins and then quenched with 100 mM glucose for 10 mins at room temperature68. Finally, the product was recovered with ethanol precipitation by adding 400 μL of 100% ethanol and 1 μL glycogen (Thermo Scientific™ R0551). The mixture was spun down at 12,000 xg for 20 mins, followedby two washes with 95% ethanol. The pellet was air dried for 15 mins and resuspended in 40 μL ddH2O.
[0107] Deacylation
[0108] tRNAs were deacylated with 100 mM Tris base (pH 9) and cleaned up with Zymo Research Oligo Clean and Concentrator kit (D4060) as per manufacturer’s instructions. The product was eluted in 45 μL ddH2O.
[0109] tRNA-extension assay
[0110] 1 μL of 10 mM dNTPs, 1 μL of 1 μM tRNA-extension probe and 5 μL of NEBuffer 2 (B7002S) was added to the deacylated mixture. The mixture was subjected to a stepwise gradient: 95 ℃, 70 ℃, 50 ℃ for 2 min each and cooled to 4 ℃. Subsequently, 0.5 μL of Klenow fragment (3’ → 5’ exo-) (M0212S) was added and the polymerase extension was carried out at 37 ℃ for 20 minutes67.2x loading dye (4.8 g Urea, 30 μL Bromophenol blue, ~10 mL ddH2O) was added to the mixture.
[0111] 10% polyacrylamide 8M Urea denaturing gel was prepared using 14.4 g Urea, 10.3 mL 30% acrylamide / bisacrylamide solution, 37.5 :1 (Sigma-Aldrich A3699), and made up to 30 mL with 1x TBE. The solution was poured into gel cast after the addition of 200 μL 10% APS and 20 μL TEMED. The gels were pre-run at 250 V for 30 mins.30 μL of the reaction mixture was run on the gel at 200 V for 90 mins and imaged on the Cy5 channel.*indicates a Cy5 fluorophore on the 3’-end, red (unbolded)region indicates the extension sequence, bold indicates annealing to the tRNA.
[0112] Gel purification:
[0113] The desired band was cut out using the Cy5 channel and transferred to a 1.5 mL microcentrifuge tube. To extract the tRNA, crush and soak method was used as described below. The bands were weighed and 1 μL of Crush and Soak buffer (CS buffer) was added per 2 mg. The gels were crushed into fine pieces with a squisher and 10 times more CS buffer was added (10 μL per 2 mg). The mixture was frozen at -80 ℃ for 15 min, thawed at 37℃ for 15 min and incubated on a rocker at 4℃ overnight. The samples were spun down at 12,000 xg for 15 mins at 4 ℃ and the supernatant containing tRNA (~400 μL) was ethanol precipitated (>2.5 h) as follows: 400 μL RNA, 800 μL 100% ethanol, 20 μL 3M NaOAc (pH 5.2), 1 μL glycogen. Finally, the mixture was spun down at 12,000 x g for 20 min at 4 ℃, washed twice with 95% ethanol and eventually resuspended in 20 μL ddH2O.
[0114] Reverse transcription
[0115] 8 μL of the gel purified tRNA fraction was incubated with 1 μL of 10 mM dNTPs and 1 μL of 2.5 μM primer (Mapyl_polyT_R for adapter specific RT, MaPylT_RT_16nt_R for reverse transcription of tRNAMaPyl) at 65 ℃ for 5 min. Further, 4 μL of 5x Induro® RT buffer and 1 μL of Induro® RT (NEB: M0681S) were added to the mixture and made up to 20 μL with ddH2O. The reverse transcription was performed at 55 ℃ for 30 minutes and the enzyme was inactivated at 95 ℃ for 1 min. After cDNA synthesis, the tRNA was hydrolysed with 10 μL 1M NaOH, 20 μL 0.2 M EDTA (pH 8) at 65 ℃ for 15 min. The cDNA was recovered with Zymo Research Oligo Clean and Concentrator kit (D4060) and eluted in 20 μL ddH2O.
[0116] Amplification and sequencing
[0117] The amplicon sequencing adapters were installed via PCR with primers QB_5_MaPylT_polyT_div_F and amplicon_tREX_15A_R. The NGS analysis for the selection outputs was performed using the following script: github.com / cpsoni74 / Illumina_barcode_analysis.git69. The NGS Analysis for Selection section contains more details about the script.
[0118] Illumina sequencing
[0119] Illumina adapter sequences were attached to the DNA samples from the selection via two rounds of PCR. The primers were designed to add Illumina adapters provided by the TruSeq DNA HT Sample Prep Kit (Illumina). The forward adapter is AATGATACGGCGACCACCGAGATCTACAC[i5]ACACTCTTTCCCTACACGACGCTCTT CCGATCT, in which [i5] is an eight-nucleotide barcode sequence (SEQ ID NOS:---), and the reverse adapter is GATCGGAAGAGCACACGTCTGAACTCCAGTCAC[i7]ATCTCGTATGCCGTCTTCTGCT TG, in which [i7] is an eight-nucleotide barcode sequence (SEQ ID NOS:--). Both the i5 and i7 barcode sequences are provided below (SEQ ID NOS: 65-72) respectively, top to bottom.
[0120] The first set of primers, Ill-PylT-X1-F (X1a, X1b, X1c and X1d mixed in equimolar ratio) and Ill-PylT-X2-R (X2a, X2b, X2c and X2d mixed in equimolar ratio), consist of half of the TruSeq adapters, beginning immediately after the i5 or i7 barcode, followed by primer- binding sites to anneal to the sequences surrounding the tRNA library (for the forward primer: GGGGGACGGTCCGGCGAC (SEQ ID NO:73) ; for the reverse primer: AAAAAAAAAAAAAAATGGCGAG (SEQ ID NO:74)). The second set of primers, a series of Illumina-i5-F and Illumina-i7-R variants containing different barcodes, consists of the 5′ half of the TruSeq adapters, followed by an i5 or i7 barcode, followed by primer-binding sites to annealto the first PCR (for the forward primer: ACACTCTTTCCCTACACGACGC (SEQ ID NO:6) for the reverse primer: GTGACTGGAGTTCAGACGTGTGCTC (SEQ ID NO:7).
[0121] For the first PCR, samples were prepared using the primers Ill-PytlT-X1-F and Ill- PylT-X2-R and Phusion Hot Start II DNA Polymerase per the manufacturer’s instructions. PCR samples were purified with 1% agarose gel and a second round of PCR using Illumina-i5-F and Illumina-i7-R primer variants was performed to attach the region of the adapter sequences that includes the barcodes. Unique combinations of i5 and i7 barcode sequences were applied to each sample to enable multiplexing.
[0122] Samples were prepared for sequencing using the 150-cycle NextSeq500 / 550 Mid Output kit v2.5 (Illumina) according to the manufacturer’s instructions. Sequencing was carried out on an Illumina NextSeq500 System, with 40% Illumina PhiX Control.
[0123] NGS Analysis for selection:
[0124] The analysis was done using the following script: github.com / cpsoni74 / Illumina_barcode_analysis.git. Refer to the Readme.md file for detailed instructions on how the data was processed.
[0125] In summary, the reads were subjected to a Q-score filter (>30) and a mismatch filter, to discard any reads that are misaligned. Finally, the abundance and the fraction of total or relative abundance was calculated for each library member.
[0126] For each selection, three samples were sequenced. One for the input library (reference) and two for duplicates of the output library (selection). For each selection, the fold enrichment was determined by the ratio of relative abundance in output library vs input library. The enrichment factor was obtained by normalizing the fold enrichment of each member with respect to the most enriched member. Finally, each library member was ranked based on the average enrichment factor.
[0127] Only the tRNA barcodes present in the Pacbio sequencing data were collected and written to a csv file along with the associated MaPylRS mutants. The reads were further filtered such that the average enrichment in the presence of ncAA is at least 2-fold. Moreover, the difference in the average enrichment in +ncAA and -ncAA should be at least 2-fold. After filtering the reads, the tRNA barcodes were ranked based on the average enrichment in +ncAA selection. The average enrichment for each barcode in +ncAA and -ncAA selection were co-plotted on a scatter plot. The frequency plot weblogos were generated using weblogo.berkeley.edu / logo.cgi or using Logomaker.py
[0128] Cloning hits from the selection:
[0129] The MaPylRS hits from the selection were cloned using Gibson assembly. The hits were generated using overlap extension PCR similar to the MaPylRS library cloning using MaPylRS_NcoI_R_gib and MaPylRS_NdeI_F_gib. The final amplified insert was incubated with pBK-lpp-MaPylT-glnS-MaPylRS along with Gibson assembly HiFi mix at 55 ℃ for 60 min and transformed into DH10B electrocompetent cells. The hits were sequenced with Sanger Sequencing as well as whole plasmid sequencing.
[0130] Assessment of MaPylRS mutants activity using sfGFP-TAG reporter:
[0131] The pEvol-sfGFP-151-TAG reporter was co-transformed with a pBK plasmid containing tRNAMaPyland MaPylRS variant into DH10B cells. Cultures were grown in 5 mL LB overnight and diluted to an OD 0.05 in 20 mL LB. Upon reaching an OD of 0.6, expression of the reporter was induced with 1 mM IPTG and appropriate ncAA. The cultures were incubated at 30 ℃ with shaking (250 rpm) for 16 hours. Subsequently, the cultures were spun down (4500 xg, 10 min, 4 ℃), the media was removed and the pellet was resuspended in 1 mL 1x PBS. The resuspended cells were diluted 10-fold (15 μL in 135 μL PBS) and the fluorescence was measured in a black, clear-bottom 96-well plate using a plate reader (ex = 488 nm, em = 534 nm). Mean of two independent experiments were reported and the error bars represent standard deviation.
[0132] sfGFP-151-TAG expression and purification:
[0133] After inducing expression of a 20 mL culture similarly as described above, the cells were spun down (4500 xg, 10 min, 4 ℃), the media was removed and the pellets were resuspended in lysis buffer consisting of 1 mL B-Per™ (Thermo Scientific™ 78243), 10 μL Halt protease inhibitor cocktail (Thermo Scientific™ 78429) and 0.1 μL Pierce™ universal nuclease (Thermo Scientific™ 88702). The solution was incubated for 30 min at 4 ℃ and spun down at maximum speed. sfGFP was purified from the supernatant using HisPur™ Ni-NTA resin (Thermo Scientific™ 88221) as per manufacturer’s instructions. Protein purity was characterized using SDS-PAGE and whole protein ESI-MS.
[0134] Oligonucleotide sequences:
[0135] Plasmid sequences:
[0136] lpp promoter is highlighted in blue, tRNAMaPylis highlighted in red, metG promoter is highlighted in purple, MaPylRS is highlighted in brown or BOLD, glnS promoter is highlighted in orange, sfGFP is highlighted in green.
[0137] pBK lpp-tRNAMaPyl(SEQ ID NO:53) GGCTGGCCTGTTGAACAAGTCTGGAAAGAAATGCATAAGCTTTTGCCATTCTCACCG GATTCAGTCGTCACTCATGGTGATTTCTCACTTGATAACCTTATTTTTGACGAGGGGA AATTAATAGGTTGTATTGATGTTGGACGAGTCGGAATCGCAGACCGATACCAGGAT CTTGCCATCCTATGGAACTGCCTCGGTGAGTTTTCTCCTTCATTACAGAAACGGCTTTTTCAAAAATATGGTATTGATAATCCTGATATGAATAAATTGCAGTTTCATTTGATGC TCGATGAGTTTTTCTAATCAGAATTGGTTAATTGGTTGTAACACTGGCAGAGCATTA CGCTGACTTGACGGGACGGCGGCTTTGTTGAATAAATCGAACTTTTGCTGAGTTGAA GGATCCCGCCGCTTCTTTGAGCGAACGATCAAAAATAAGTGGCGCCCCATCAAAAA AATATTCTCAACATAAAAAACTTTGTGTAATACTTGTAACGCTGAATTCGGGGGACG GTCCGGCGACCAGCGGGTCTCTAAAACCTAGCCAGCGGGGTTCGACGCCCCGGTCT CTCGCCACTGCAGATCCTTAGCGAAAGCTAAGGATTTTTTTTAGTCGACCGAATTTC TGCCATTCATCCGCGGCCGCCTGCAGTTTCAAACGCTAAATTGCCTGATGCGCTACG CTTATCAGGCCTACATGATCTCTGCAATATATTGAGTTTGCGTGCTTTTGTAGGCCGG ATAAGGCGTTCACGCCGCATCCGGCAAGAAACAGCAAACAATCACATGTGAGCAAA AGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATA GGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGA AACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGC TCTCCTGTTCCGACCCTGCCGCTTACCGGATACCTGTCCGCCTTTCTCCCTTCGGGAA GCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTCGTTC GCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTAT CCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAG CAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTC TTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGGACAGTATTTGGTATCTGCGCT CTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGTTCTTGATCCGGCAAACAA ACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCAGAAAA AAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACGCTCAGTGGAAC GAAAACTCACGTTAAGGGATTTTGGTCATGAACAATAAAACTGTCTGCTTACATAAA CAGTAATACAAGGGGTGTTATGAGCCATATTCAACGGGAAACGTCTTGCTCGAGGC CGCGATTAAATTCCAACATGGATGCTGATTTATATGGGTATAAATGGGCTCGCGATA ATGTCGGGCAATCAGGTGCGACAATCTATCGATTGTATGGGAAGCCCGATGCGCCA GAGTTGTTTCTGAAACATGGCAAAGGTAGCGTTGCCAATGATGTTACAGATGAGAT GGTCAGACTAAACTGGCTGACGGAATTTATGCCTCTTCCGACCATCAAGCATTTTAT CCGTACTCCTGATGATGCATGGTTACTCACCACTGCGATCCCCGGGAAAACAGCATT CCAGGTATTAGAAGAATATCCTGATTCAGGTGAAAATATTGTTGATGCGCTGGCAGT GTTCCTGCGCCGGTTGCATTCGATTCCTGTTTGTAATTGTCCTTTTAACAGCGATCGC GTATTTCGTCTCGCTCAGGCGCAATCACGAATGAATAACGGTTTGGTTGATGCGAGT GATTTTGATGACGAGCGTAATpBK lpp -tRNAMaPyl-metG-MaPylRS (SEQ ID NO:54) GTCGACCGAATTTCTGCCATTCATCCGCGGCCGCTTTACTTAACATTTTCCCATTTGG TACTATCTAACCCCTTTTCACTATTAAGAAGTAATGCCTACTCATATGACGGTGAAA TACACTGACGCACAGATCCAGCGCCTTCGCGAATATGGGAATGGCACGTATGAACA GAAAGTGTTCGAAGATTTGGCTTCGCGCGACGCAGCCTTTAGCAAAGAAATGAGTG TTGCCTCAACCGACAATGAGAAAAAAATTAAGGGCATGATTGCCAACCCGTCACGT CATGGACTTACGCAACTTATGAACGACATTGCCGACGCATTAGTCGCTGAGGGATTT ATCGAGGTCCGCACGCCAATCTTTATCTCAAAAGACGCGCTTGCCCGTATGACGATT ACAGAAGACAAGCCCCTGTTCAAGCAAGTATTCTGGATCGACGAGAAGCGTGCCTT ACGCCCAATGTTGGCTCCAAATTTATATTCCGTTATGCGTGATTTGCGTGACCACAC CGACGGCCCAGTGAAGATTTTCGAGATGGGGAGCTGTTTTCGCAAGGAAAGTCACA GTGGCATGCATTTGGAGGAGTTCACGATGCTGAACCTTGTGGATATGGGACCGCGTG GTGATGCGACAGAGGTTTTAAAAAATTACATTAGTGTTGTGATGAAAGCAGCGGGA TTGCCCGATTATGATTTAGTCCAGGAAGAGAGTGACGTCTACAAAGAAACTATCGAT GTTGAGATTAACGGGCAAGAAGTATGTAGCGCTGCTGTCGGACCCCATTATCTGGAT GCTGCCCATGATGTGCATGAACCTTGGTCTGGTGCTGGTTTCGGTTTGGAGCGCTTA TTAACCATTCGTGAGAAATATTCCACAGTAAAGAAAGGGGGGGCAAGTATCTCGTA CCTGAACGGTGCAAAAATTAATTAACCATGGCTGCAGTTTCAAACGCTAAATTGCCT GATGCGCTACGCTTATCAGGCCTACATGATCTCTGCAATATATTGAGTTTGCGTGCTT TTGTAGGCCGGATAAGGCGTTCACGCCGCATCCGGCAAGAAACAGCAAACAATCCA AAACGCCGCGTTCAGCGGCGTTTTTTCTGCTTTTCTTCGCGAATTAATTCCGCTTCGC ACATGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCT GGCGTTTTTCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAG TCAGAGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCCTGGAA GCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGATACCTGTCCGCCTT TCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCG GTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGAC CGCTGCGCCTTATCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTA TCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTAGGCGG TGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGGACAGTATT TGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGTTCTTG ATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGAT TACGCGCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGA CGCTCAGTGGAACGAAAACTCACGTTAAGGGATTTTGGTCATGAACAATAAAACTGTCTGCTTACATAAACAGTAATACAAGGGGTGTTATGAGCCATATTCAACGGGAAAC GTCTTGCTCGAGGCCGCGATTAAATTCCAACATGGATGCTGATTTATATGGGTATAA ATGGGCTCGCGATAATGTCGGGCAATCAGGTGCGACAATCTATCGATTGTATGGGA AGCCCGATGCGCCAGAGTTGTTTCTGAAACATGGCAAAGGTAGCGTTGCCAATGAT GTTACAGATGAGATGGTCAGACTAAACTGGCTGACGGAATTTATGCCTCTTCCGACC ATCAAGCATTTTATCCGTACTCCTGATGATGCATGGTTACTCACCACTGCGATCCCC GGGAAAACAGCATTCCAGGTATTAGAAGAATATCCTGATTCAGGTGAAAATATTGT TGATGCGCTGGCAGTGTTCCTGCGCCGGTTGCATTCGATTCCTGTTTGTAATTGTCCT TTTAACAGCGATCGCGTATTTCGTCTCGCTCAGGCGCAATCACGAATGAATAACGGT TTGGTTGATGCGAGTGATTTTGATGACGAGCGTAATGGCTGGCCTGTTGAACAAGTC TGGAAAGAAATGCATAAGCTTTTGCCATTCTCACCGGATTCAGTCGTCACTCATGGT GATTTCTCACTTGATAACCTTATTTTTGACGAGGGGAAATTAATAGGTTGTATTGAT GTTGGACGAGTCGGAATCGCAGACCGATACCAGGATCTTGCCATCCTATGGAACTG CCTCGGTGAGTTTTCTCCTTCATTACAGAAACGGCTTTTTCAAAAATATGGTATTGAT AATCCTGATATGAATAAATTGCAGTTTCATTTGATGCTCGATGAGTTTTTCTAATCAG AATTGGTTAATTGGTTGTAACACTGGCAGAGCATTACGCTGACTTGACGGGACGGCG GCTTTGTTGAATAAATCGAACTTTTGCTGAGTTGAAGGATCCCGCCGCTTCTTTGAG CGAACGATCAAAAATAAGTGGCGCCCCATCAAAAAAATATTCTCAACATAAAAAAC TTTGTGTAATACTTGTAACGCTGAATTCGGGGGACGGTCCGGCGACCAGCGGGTCTC TAAAACCTAGCCAGCGGGGTTCGACGCCCCGGTCTCTCGCCACTGCAGATCCTTAGC GAAAGCTAAGGATTTTTTTTA pBK lpp-tRNAMaPyl-glnS-MaPylRS (SEQ ID NO:55) GTCGACCGAATTTCTGCCATTCATCCGCGGCCGCTCGGGTTGTCAGCCTGTCCCGCTT ATAAGATCATACGCCGTTATACGTTGTTTACGCTTTGAGGAATCCCATATGACGGTG AAATACACTGACGCACAGATCCAGCGCCTTCGCGAATATGGGAATGGCACGTATGA ACAGAAAGTGTTCGAAGATTTGGCTTCGCGCGACGCAGCCTTTAGCAAAGAAATGA GTGTTGCCTCAACCGACAATGAGAAAAAAATTAAGGGCATGATTGCCAACCCGTCA CGTCATGGACTTACGCAACTTATGAACGACATTGCCGACGCATTAGTCGCTGAGGGA TTTATCGAGGTCCGCACGCCAATCTTTATCTCAAAAGACGCGCTTGCCCGTATGACG ATTACAGAAGACAAGCCCCTGTTCAAGCAAGTATTCTGGATCGACGAGAAGCGTGC CTTACGCCCAATGTTGGCTCCAAATTTATATTCCGTTATGCGTGATTTGCGTGACCAC ACCGACGGCCCAGTGAAGATTTTCGAGATGGGGAGCTGTTTTCGCAAGGAAAGTCACAGTGGCATGCATTTGGAGGAGTTCACGATGCTGAACCTTGTGGATATGGGACCGC GTGGTGATGCGACAGAGGTTTTAAAAAATTACATTAGTGTTGTGATGAAAGCAGCG GGATTGCCCGATTATGATTTAGTCCAGGAAGAGAGTGACGTCTACAAAGAAACTAT CGATGTTGAGATTAACGGGCAAGAAGTATGTAGCGCTGCTGTCGGACCCCATTATCT GGATGCTGCCCATGATGTGCATGAACCTTGGTCTGGTGCTGGTTTCGGTTTGGAGCG CTTATTAACCATTCGTGAGAAATATTCCACAGTAAAGAAAGGGGGGGCAAGTATCT CGTACCTGAACGGTGCAAAAATTAATTAACCATGGCTGCAGTTTCAAACGCTAAATT GCCTGATGCGCTACGCTTATCAGGCCTACATGATCTCTGCAATATATTGAGTTTGCG TGCTTTTGTAGGCCGGATAAGGCGTTCACGCCGCATCCGGCAAGAAACAGCAAACA ATCCAAAACGCCGCGTTCAGCGGCGTTTTTTCTGCTTTTCTTCGCGAATTAATTCCGC TTCGCACATGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCG TTGCTGGCGTTTTTCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGC TCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCC TGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGATACCTGTCC GCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCA GTTCGGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGC CCGACCGCTGCGCCTTATCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACG ACTTATCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGAGCGAGGTATGTA GGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGGAC AGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAG TTCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCA GCAGATTACGCGCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGG GGTCTGACGCTCAGTGGAACGAAAACTCACGTTAAGGGATTTTGGTCATGAACAAT AAAACTGTCTGCTTACATAAACAGTAATACAAGGGGTGTTATGAGCCATATTCAACG GGAAACGTCTTGCTCGAGGCCGCGATTAAATTCCAACATGGATGCTGATTTATATGG GTATAAATGGGCTCGCGATAATGTCGGGCAATCAGGTGCGACAATCTATCGATTGTA TGGGAAGCCCGATGCGCCAGAGTTGTTTCTGAAACATGGCAAAGGTAGCGTTGCCA ATGATGTTACAGATGAGATGGTCAGACTAAACTGGCTGACGGAATTTATGCCTCTTC CGACCATCAAGCATTTTATCCGTACTCCTGATGATGCATGGTTACTCACCACTGCGA TCCCCGGGAAAACAGCATTCCAGGTATTAGAAGAATATCCTGATTCAGGTGAAAAT ATTGTTGATGCGCTGGCAGTGTTCCTGCGCCGGTTGCATTCGATTCCTGTTTGTAATT GTCCTTTTAACAGCGATCGCGTATTTCGTCTCGCTCAGGCGCAATCACGAATGAATA ACGGTTTGGTTGATGCGAGTGATTTTGATGACGAGCGTAATGGCTGGCCTGTTGAAC AAGTCTGGAAAGAAATGCATAAGCTTTTGCCATTCTCACCGGATTCAGTCGTCACTCATGGTGATTTCTCACTTGATAACCTTATTTTTGACGAGGGGAAATTAATAGGTTGTA TTGATGTTGGACGAGTCGGAATCGCAGACCGATACCAGGATCTTGCCATCCTATGGA ACTGCCTCGGTGAGTTTTCTCCTTCATTACAGAAACGGCTTTTTCAAAAATATGGTAT TGATAATCCTGATATGAATAAATTGCAGTTTCATTTGATGCTCGATGAGTTTTTCTAA TCAGAATTGGTTAATTGGTTGTAACACTGGCAGAGCATTACGCTGACTTGACGGGAC GGCGGCTTTGTTGAATAAATCGAACTTTTGCTGAGTTGAAGGATCCCGCCGCTTCTT TGAGCGAACGATCAAAAATAAGTGGCGCCCCATCAAAAAAATATTCTCAACATAAA AAACTTTGTGTAATACTTGTAACGCTGAATTCGGGGGACGGTCCGGCGACCAGCGG GTCTCTAAAACCTAGCCAGCGGGGTTCGACGCCCCGGTCTCTCGCCACTGCAGATCC TTAGCGAAAGCTAAGGATTTTTTTTA pEvol T5-lac-sfGFP-151-TAG (SEQ ID NO: 56) CTCCATTTTAGCTTCCTTAGCTCCTGAAAATCTCGATAACTCAAAAAATACGCCCGG TAGTGATCTTATTTCATTATGGTGAAAGTTGGAACCTCTTACGTGCCGATCAACGTCT CATTTTCGCCAAAAGTTGGCCCAGGGCTTCCCGGTATCAACAGGGACACCAGGATTT ATTTATTCTGCGAAGTGATCTTCCGTCACAGGTATTTGTTCGGCGCAAAGTGCGTCG GGTGATGCTGCCAACTTACTGATTTAGTGTATGATGGTGTTTTTGAGGTGCTCCAGT GGCTTCTGTTTCTATCAGCTGTCCCTCCTGTTCAGCTACTGACGGGGTGGTGCGTAAC GGCAAAAGCACCGCCGGACATCAGCGCTAGCGGAGTGTATACTGGCTTACTATGTT GGCACTGATGAGGGTGTCAGTGAAGTGCTTCATGTGGCAGGAGAAAAAAGGCTGCA CCGGTGCGTCAGCAGAATATGTGATACAGGATATATTCCGCTTCCTCGCTCACTGAC TCGCTACGCTCGGTCGTTCGACTGCGGCGAGCGGAAATGGCTTACGAACGGGGCGG AGATTTCCTGGAAGATGCCAGGAAGATACTTAACAGGGAAGTGAGAGGGCCGCGGC AAAGCCGTTTTTCCATAGGCTCCGCCCCCCTGACAAGCATCACGAAATCTGACGCTC AAATCAGTGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCCTG GCGGCTCCCTCGTGCGCTCTCCTGTTCCTGCCTTTCGGTTTACCGGTGTCATTCCGCT GTTATGGCCGCGTTTGTCTCATTCCACGCCTGACACTCAGTTCCGGGTAGGCAGTTC GCTCCAAGCTGGACTGTATGCACGAACCCCCCGTTCAGTCCGACCGCTGCGCCTTAT CCGGTAACTATCGTCTTGAGTCCAACCCGGAAAGACATGCAAAAGCACCACTGGCA GCAGCCACTGGTAATTGATTTAGAGGAGTTAGTCTTGAAGTCATGCGCCGGTTAAGG CTAAACTGAAAGGACAAGTTTTGGTGACTGCGCTCCTCCAAGCCAGTTACCTCGGTT CAAAGAGTTGGTAGCTCAGAGAACCTTCGAAAAACCGCCCTGCAAGGCGGTTTTTTCGTTTTCAGAGCAAGAGATTACGCGCAGACCAAAACGATCTCAAGAAGATCATCTTA TTAATCAGATAAAATATTTCTAGATTTCAGTGCAATTTATCTCTTCAAATGTAGCACC TGAAGTCAGCCCCATACGATATAAGTTGTAATTCTCATGTTTGACAGCTTATCATCG ATAAGCTTGCAATTTATCTCTTCAAATGTAGCACCTGAAGTCAGCCCCATACGATAT AAGTTGTAATTCTCATGTTAGTCATGCCCCGCGCCCACCGGAAGGAGCTGACTGGGT TGAAGGCTCTCAAGGGCATCGGTCGAGATCCCGGTGCCTAATGAGTGAGCTAACTT ACATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCAG CTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCCA GGGTGGTTTTTCTTTTCACCAGTGAGACGGGCAACAGCTGATTGCCCTTCACCGCCT GGCCCTGAGAGAGTTGCAGCAAGCGGTCCACGCTGGTTTGCCCCAGCAGGCGAAAA TCCTGTTTGATGGTGGTTAACGGCGGGATATAACATGAGCTGTCTTCGGTATCGTCG TATCCCACTACCGAGATGTCCGCACCAACGCGCAGCCCGGACTCGGTAATGGCGCG CATTGCGCCCAGCGCCATCTGATCGTTGGCAACCAGCATCGCAGTGGGAACGATGC CCTCATTCAGCATTTGCATGGTTTGTTGAAAACCGGACATGGCACTCCAGTCGCCTT CCCGTTCCGCTATCGGCTGAATTTGATTGCGAGTGAGATATTTATGCCAGCCAGCCA GACGCAGACGCGCCGAGACAGAACTTAATGGGCCCGCTAACAGCGCGATTTGCTGG TGACCCAATGCGACCAGATGCTCCACGCCCAGTCGCGTACCGTCTTCATGGGAGAA AATAATACTGTTGATGGGTGTCTGGTCAGAGACATCAAGAAATAACGCCGGAACAT TAGTGCAGGCAGCTTCCACAGCAATGGCATCCTGGTCATCCAGCGGATAGTTAATGA TCAGCCCACTGACGCGTTGCGCGAGAAGATTGTGCACCGCCGCTTTACAGGCTTCGA CGCCGCTTCGTTCTACCATCGACACCACCACGCTGGCACCCAGTTGATCGGCGCGAG ATTTAATCGCCGCGACAATTTGCGACGGCGCGTGCAGGGCCAGACTGGAGGTGGCA ACGCCAATCAGCAACGACTGTTTGCCCGCCAGTTGTTGTGCCACGCGGTTGGGAATG TAATTCAGCTCCGCCATCGCCGCTTCCACTTTTTCCCGCGTTTTCGCAGAAACGTGGC TGGCCTGGTTCACCACGCGGGAAACGGTCTGATAAGAGACACCGGCATACTCTGCG ACATCGTATAACGTTACTGGTTTCACATTCACCACCCTGAATTGACTCTCTTCCGGGC GCTATCATGCCATACCGCGAAAGGTTTTGCGCCATTCGATGGTGTCCGGGATCTCGA CGCTCTCCCTTATGCGACTCCTGCATTAGGCTCACTATAGGGGAATTGTGAGCGGAT AACAATTCCCCTCTAGAGTTTGACAGCATTGTCATCGATCTCGAGAAATCATAAAAA ATTTATTTGCTTTGTGAGCGGATAACAATTATAATAGATTCAATTGTGAGCGGATAA CAATTTCACACAGAATTCATTAAAGAGGAGAAATTACATATGAGCAAAGGAGAAGA ACTTTTCACTGGAGTTGTCCCAATTCTTGTTGAATTAGATGGTGATGTTAATGGGCAC AAATTTTCTGTCCGTGGAGAGGGTGAAGGTGATGCTACAAACGGAAAACTCACCCT TAAATTTATTTGCACTACTGGAAAACTACCTGTTCCGTGGCCAACACTTGTCACTACTCTGACCTATGGTGTTCAATGCTTTTCCCGTTATCCGGATCACATGAAACGGCATGAC TTTTTCAAGAGTGCCATGCCCGAAGGTTATGTACAGGAACGCACTATATCTTTCAAA GATGACGGGACCTACAAGACGCGTGCTGAAGTCAAGTTTGAAGGTGATACCCTTGT TAATCGTATCGAGTTAAAGGGTATTGATTTTAAAGAAGATGGAAACATTCTTGGACA CAAACTCGAGTACAACTTTAACTCACACAATGTATAGATCACGGCAGACAAACAAA AGAATGGAATCAAAGCTAACTTCAAAATTCGCCACAACGTTGAAGATGGTTCCGTTC AACTAGCAGACCATTATCAACAAAATACTCCAATTGGCGATGGCCCTGTCCTTTTAC CAGACAACCATTACCTGTCGACACAATCTGTCCTTTCGAAAGATCCCAACGAAAAGC GTGACCACATGGTCCTTCTTGAGTTTGTAACTGCTGCTGGGATTACACATGGCATGG ATGAGCTCTACAAAGGATCCCACCACCACCACCACCACTAAAAGCTTAATTAGCTG AGCTTGGACTCCTGTTGATAGATCCAGTAATGACCTCAGAACTCCATCTGGATTTGT TCAGAACGCTCGGTTGCCGCCGGGCGTTTTTTATTGGTGAGAATCCAAGCTAGCTTG GCGCCTCGAGCAGCTCAGGGTCGAATTTGCTTTCGAATTTCTGCCATTCATCCGCTTA TTATCACTTATTCAGGCGTAGCAACCAGGCGTTTAAGGGCACCAATAACTGCCTTAA AAAAATTACGCCCCGCCCTGCCACTCATCGCAGTACTGTTGTAATTCATTAAGCATT CTGCCGACATGGAAGCCATCACAAACGGCATGATGAACCTGAATCGCCAGCGGCAT CAGCACCTTGTCGCCTTGCGTATAATATTTGCCCATGGTGAAAACGGGGGCGAAGA AGTTGTCCATATTGGCCACGTTTAAATCAAAACTGGTGAAACTCACCCAGGGATTGG CTGAGACGAAAAACATATTCTCAATAAACCCTTTAGGGAAATAGGCCAGGTTTTCAC CGTAACACGCCACATCTTGCGAATATATGTGTAGAAACTGCCGGAAATCGTCGTGGT ATTCACTCCAGAGCGATGAAAACGTTTCAGTTTGCTCATGGAAAACGGTGTAACAA GGGTGAACACTATCCCATATCACCAGCTCACCTTCTTTCATTGCCATACGGAATTCC GGATGAGCATTCATCAGGCGGGCAAGAATGTGAATAAAGGCCGGATAAAACTTGTG CTTATTTTTCTTTACGGTCTTTAAAAAGGCCGTAATATCCAGCTGAACGGTCTGGTTA TAGGTACATTGAGCAACTGACTGAAATGCCTCAAAATGTTCTTTAC GATGCCATTGGGATATATCAACGGTGGTATATCCAGTGATTTTTTT
[0138] Nucleotide Sequence of Wild-type M. alvus Aminoacyl Pyrrolysyl-tRNA Synthetase (SEQ ID NO:57)
[0139] ATGACGGTGAAATACACTGACGCACAGATCCAGCGCCTTCGCGAATATG GGAATGGCACGTATGAACAGAAAGTGTTCGAAGATTTGGCTTCGCGCGACGCAGCC TTTAGCAAAGAAATGAGTGTTGCCTCAACCGACAATGAGAAAAAAATTAAGGGCAT GATTGCCAACCCGTCACGTCATGGACTTACGCAACTTATGAACGACATTGCCGACGC ATTAGTCGCTGAGGGATTTATCGAGGTCCGCACGCCAATCTTTATCTCAAAAGACGCGCTTGCCCGTATGACGATTACAGAAGACAAGCCCCTGTTCAAGCAAGTATTCTGGAT CGACGAGAAGCGTGCCTTACGCCCAATGTTGGCTCCAAATTTATATTCCGTTATGCG TGATTTGCGTGACCACACCGACGGCCCAGTGAAGATTTTCGAGATGGGGAGCTGTTT TCGCAAGGAAAGTCACAGTGGCATGCATTTGGAGGAGTTCACGATGCTGAACCTTG TGGATATGGGACCGCGTGGTGATGCGACAGAGGTTTTAAAAAATTACATTAGTGTTG TGATGAAAGCAGCGGGATTGCCCGATTATGATTTAGTCCAGGAAGAGAGTGACGTC TACAAAGAAACTATCGATGTTGAGATTAACGGGCAAGAAGTATGTAGCGCTGCTGT CGGACCCCATTATCTGGATGCTGCCCATGATGTGCATGAACCTTGGTCTGGTGCTGG TTTCGGTTTGGAGCGCTTATTAACCATTCGTGAGAAATATTCCACAGTAAAGAAAGG GGGGGCAAGTATCTCGTACCTGAACGGTGCAAAAATTAATTAA
[0140] Amino Acid Sequence of Wild-Type M. alvus Pyrrolysyl-tRNA Synthetase (SEQ ID NO:58)
[0141] MTVKYTDAQI QRLREYGNGT YEQKVFEDLA SRDAAFSKEM SVASTDNEKK IKGMIANPSR HGLTQLMNDI ADALVAEGFI EVRTPIFISK DALARMTITE DKPLFKQVFW IDEKRALRPM LAPNLYSVMR DLRDHTDGPV KIFEMGSCFR KESHSGMHLE EFTMLNLVDM GPRGDATEVL KNYISVVMKA AGLPDYDLVQ EESDVYKETI DVEINGQEVC SAAVGPHYLD AAHDVHEPWS GAGFGLERLL TIREKYSTVK KGGASISYLN GAKIN
[0142] Mutant Methanomethylophilus alvus Pyrrolysyl-tRNA synthetase (MaPyRS) amino acid sequences
[0143] MaRylRS Mutant VSV (SEQ ID NO:75): MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKIKGMIANP SRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWIDEKRALRP MLAPNVYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLSLVDMGPRGDA TEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPHYLDAAHD VHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0144] MaPylRS Mutant LSV (SEQ ID NO:76):
[0145] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNLYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLSLVD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0146] MaPylRS Mutant LAC (SEQ ID NO: 77):
[0147] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNLYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLALCD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0148] MaPylRS Mutant LAV (SEQ ID NO:78):
[0149] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNLYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLALVD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0150] MaPylRS Mutant AAV (SEQ ID NO:79):
[0151] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNAYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLALVD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0152] MaPylRS Mutant ASV (SEQ ID NO:80):
[0153] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNAYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLALVD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0154] MaPylRS Mutant TTW (SEQ ID NO:81):
[0155] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNTYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLTLWD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0156] MaPylRS Mutant AST (SEQ ID NO:82):
[0157] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKI KGMIANPSRHGLTQLMNDIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWI DEKRALRPMLAPNAYSVMRDLRDHTDGPVKIFEMGSCFRKESHSGMHLEEFTMLSLTD MGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKETIDVEINGQEVCSAAVGPH YLDAAHDVHEPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN
[0158] Conclusion
[0159] Using directed evolution, aminoacyl-tRNA synthetases (aaRSs) have been engineered to incorporate numerous noncanonical amino acids (ncAAs). Until now, selection of such novel aaRS mutants has relied on the expression of a selectable reporter protein. However, such translation-dependent selections are incompatible with exotic monomers that are suboptimal substrates for the ribosome. A two-step solution was needed to overcome this limitation: A) Engineering an aaRS to charge the exotic monomer, without ribosomal translation, and B) Subsequent engineering of the ribosome to accept the resulting acyl-tRNA for translation. Described herein is a platform for aaRS engineering that directly selects for tRNA-acylation without ribosomal translation (START). In START, each distinct aaRS mutant is correlated to a cognate tRNA containing a unique sequence barcode. Acylation by an active aaRS mutant protects the associated barcode-containing tRNAs from an oxidative treatment designed to damage the 3′-terminus of the uncharged tRNAs. Sequencing of these surviving barcode- containing tRNAs is then used to reveal the identity of aaRS mutants that acylated the correlated tRNA sequences. The efficacy of START was demonstrated by identifying novel mutants of the M. alvus pyrrolysyl-tRNA synthetase from a naïve library that enables incorporation of ncAAs into proteins in living cells.
[0160] While this invention has been particularly shown and described with references to preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the invention encompassed by the appended claims.LIST OF REFERENCES The list of references are herein incorporated by reference in their entirety. 1. Chin, J. W. (2017) Expanding and reprogramming the genetic code, Nature 550, 53-60. 2. Dumas, A., Lercher, L., Spicer, C. D., and Davis, B. G. (2015) Designing logical codon reassignment - Expanding the chemistry in biology, Chemical science 6, 50-69. 3. Young, D. D., and Schultz, P. G. (2018) Playing with the molecules of life, ACS chemical biology. 4. Icking, L.-S., Riedlberger, A. M., Krause, F., Widder, J., Frederiksen, Anne S., Stockert, F., Spädt, M., Edel, N., Armbruster, D., Forlani, G., Franchini, S., Kaas, P., Kırpat Konak, B. M., Krier, F., Lefebvre, M., Mazraeh, D., Ranniger, J., Gerstenecker, J., Gescher, P., Voigt, K., Salavei, P., Gensch, N., Di Ventura, B., and Öztürk, M. A. (2023) iNClusive: a database collecting useful information on non-canonical amino acids and their incorporation into proteins for easier genetic code expansion implementation, Nucleic Acids Research. 5. Vargas-Rodriguez, O., Sevostyanova, A., Söll, D., and Crnković, A. (2018) Upgrading aminoacyl-tRNA synthetases for genetic code expansion, Current opinion in chemical biology 46, 115-122. 6. Manandhar, M., Chun, E., and Romesberg, F. E. (2021) Genetic Code Expansion: Inception, Development, Commercialization, Journal of the American Chemical Society 143, 4859-4878. 7. Sun, S. B., Schultz, P. G., and Kim, C. H. (2014) Therapeutic applications of an expanded genetic code, Chembiochem : a European journal of chemical biology 15, 1721-1729. 8. Guo, J., Wang, J., Anderson, J. C., and Schultz, P. G. (2008) Addition of an α-Hydroxy Acid to the Genetic Code of Bacteria, Angewandte Chemie International Edition 47, 722-725. 9. Kobayashi, T., Yanagisawa, T., Sakamoto, K., and Yokoyama, S. (2009) Recognition of non- alpha-amino substrates by pyrrolysyl-tRNA synthetase, Journal of molecular biology 385, 1352- 1360. 10. Li, Y. M., Yang, M. Y., Huang, Y. C., Li, Y. T., Chen, P. R., and Liu, L. (2012) Ligation of expressed protein α-hydrazides via genetic incorporation of an α-hydroxy acid, ACS chemical biology 7, 1015-1022. 11. Bindman, N. A., Bobeica, S. C., Liu, W. R., and van der Donk, W. A. (2015) Facile Removal of Leader Peptides from Lanthipeptides by Incorporation of a Hydroxy Acid, J Am Chem Soc 137, 6975-6978.12. Spinck, M., Piedrafita, C., Robertson, W. E., Elliott, T. S., Cervettini, D., de la Torre, D., and Chin, J. W. (2023) Genetically programmed cell-based synthesis of non-natural peptide and depsipeptide macrocycles, Nature chemistry 15, 61-69. 13. Katoh, T., Iwane, Y., and Suga, H. (2017) Logical engineering of D-arm and T-stem of tRNA that enhances d-amino acid incorporation, Nucleic Acids Res 45, 12601-12610. 14. Katoh, T., and Suga, H. (2020) Ribosomal Elongation of Cyclic γ-Amino Acids using a Reprogrammed Genetic Code, J Am Chem Soc 142, 4965-4969. 15. Adaligil, E., Song, A., Cunningham, C. N., and Fairbrother, W. J. (2021) Ribosomal Synthesis of Macrocyclic Peptides with Linear γ(4)- and β-Hydroxy-γ(4)-amino Acids, ACS chemical biology 16, 1325-1331. 16. Fujino, T., Goto, Y., Suga, H., and Murakami, H. (2016) Ribosomal Synthesis of Peptides with Multiple β-Amino Acids, J Am Chem Soc 138, 1962-1969. 17. Sando, S., Abe, K., Sato, N., Shibata, T., Mizusawa, K., and Aoyama, Y. (2007) Unexpected preference of the E. coli translation system for the ester bond during incorporation of backbone- elongated substrates, J Am Chem Soc 129, 6180-6186. 18. Lee, J., Schwarz, K. J., Kim, D. S., Moore, J. S., and Jewett, M. C. (2020) Ribosome- mediated polymerization of long chain carbon and cyclic amino acids into peptides in vitro, Nature communications 11, 4304. 19. Katoh, T., and Suga, H. (2020) Ribosomal Elongation of Aminobenzoic Acid Derivatives, J Am Chem Soc 142, 16518-16522. 20. Katoh, T., and Suga, H. (2021) Consecutive Ribosomal Incorporation of α-Aminoxy / α- Hydrazino Acids with l / d-Configurations into Nascent Peptide Chains, J Am Chem Soc 143, 18844-18848. 21. Takatsuji, R., Shinbara, K., Katoh, T., Goto, Y., Passioura, T., Yajima, R., Komatsu, Y., and Suga, H. (2019) Ribosomal Synthesis of Backbone-Cyclic Peptides Compatible with In Vitro Display, J Am Chem Soc 141, 2279-2287. 22. Ad, O., Hoffman, K. S., Cairns, A. G., Featherston, A. L., Miller, S. J., Söll, D., and Schepartz, A. (2019) Translation of Diverse Aramid- and 1,3-Dicarbonyl-peptides by Wild Type Ribosomes in Vitro, ACS central science 5, 1289-1294. 23. Lee, J., Coronado, J. N., Cho, N., Lim, J., Hosford, B. M., Seo, S., Kim, D. S., Kofman, C., Moore, J. S., Ellington, A. D., Anslyn, E. V., and Jewett, M. C. (2022) Ribosome-mediated biosynthesis of pyridazinone oligomers in vitro, Nature communications 13, 6322. 24. Fricke, R., Swenson, C. V., Roe, L. T., Hamlish, N. X., Shah, B., Zhang, Z., Ficaretta, E., Ad, O., Smaga, S., Gee, C. L., Chatterjee, A., and Schepartz, A. (2023) Expanding the substratescope of pyrrolysyl-transfer RNA synthetase enzymes to include non-α-amino acids in vitro and in vivo, Nature chemistry 15, 960-971. 25. Dedkova, L. M., Fahmi, N. E., Golovine, S. Y., and Hecht, S. M. (2003) Enhanced D-amino acid incorporation into protein by modified ribosomes, J Am Chem Soc 125, 6616-6617. 26. Melo Czekster, C., Robertson, W. E., Walker, A. S., Söll, D., and Schepartz, A. (2016) In Vivo Biosynthesis of a β-Amino Acid-Containing Protein, J Am Chem Soc 138, 5194-5197. 27. Maini, R., Dedkova, L. M., Paul, R., Madathil, M. M., Chowdhury, S. R., Chen, S., and Hecht, S. M. (2015) Ribosome-Mediated Incorporation of Dipeptides and Dipeptide Analogues into Proteins in Vitro, Journal of the American Chemical Society 137, 11206-11209. 28. Chen, S., Ji, X., Gao, M., Dedkova, L. M., and Hecht, S. M. (2019) In Cellulo Synthesis of Proteins Containing a Fluorescent Oxazole Amino Acid, Journal of the American Chemical Society 141, 5597-5601. 29. Hamlish, N. X., Abramyan, A. M., and Schepartz, A. (2023) Incorporation of multiple β2-backbones into a protein <em>in vivo< / em> using an orthogonal aminoacyl- tRNA synthetase, bioRxiv, 2023.2011.2007.565973. 30. Liu, D. R., and Schultz, P. G. (1999) Progress toward the evolution of an organism with an expanded genetic code, Proceedings of the National Academy of Sciences of the United States of America 96, 4780-4785. 31. Wang, L., Brock, A., Herberich, B., and Schultz, P. G. (2001) Expanding the genetic code of Escherichia coli, Science (New York, N.Y.) 292, 498-500. 32. Chin, J. W., Cropp, T. A., Anderson, J. C., Mukherji, M., Zhang, Z., and Schultz, P. G. (2003) An expanded eukaryotic genetic code, Science (New York, N.Y.) 301, 964-967. 33. Santoro, S. W., Wang, L., Herberich, B., King, D. S., and Schultz, P. G. (2002) An efficient system for the evolution of aminoacyl-tRNA synthetase specificity, Nature Biotechnology 20, 1044-1048. 34. Ellefson, J. W., Meyer, A. J., Hughes, R. A., Cannon, J. R., Brodbelt, J. S., and Ellington, A. D. (2014) Directed evolution of genetic parts and circuits by compartmentalized partnered replication, Nat Biotechnol 32, 97-101. 35. Bryson, D. I., Fan, C., Guo, L. T., Miller, C., Söll, D., and Liu, D. R. (2017) Continuous directed evolution of aminoacyl-tRNA synthetases, Nature chemical biology 13, 1253-1260. 36. Suzuki, T., Miller, C., Guo, L. T., Ho, J. M. L., Bryson, D. I., Wang, Y. S., Liu, D. R., and Söll, D. (2017) Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase, Nature chemical biology 13, 1261-1266.37. Fischer, J. T., Söll, D., and Tharp, J. M. (2022) Directed Evolution of Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase Generates a Hyperactive and Highly Selective Variant, Frontiers in molecular biosciences 9, 850613. 38. Stieglitz, J. T., and Van Deventer, J. A. (2022) High-Throughput Aminoacyl-tRNA Synthetase Engineering for Genetic Code Expansion in Yeast, ACS Synthetic Biology 11, 2284- 2299. 39. Hohl, A., Karan, R., Akal, A., Renn, D., Liu, X., Ghorpade, S., Groll, M., Rueping, M., and Eppinger, J. (2019) Engineering a Polyspecific Pyrrolysyl-tRNA Synthetase by a High Throughput FACS Screen, Scientific reports 9, 11971. 40. Dyer, J. R. (1956) Use of periodate oxidations in biochemical analysis, Methods of biochemical analysis 3, 111-152. 41. Dittmar, K. A., Sørensen, M. A., Elf, J., Ehrenberg, M., and Pan, T. (2005) Selective charging of tRNA isoacceptors induced by amino-acid starvation, EMBO reports 6, 151-157. 42. Evans, M. E., Clark, W. C., Zheng, G., and Pan, T. (2017) Determination of tRNA aminoacylation levels by high-throughput sequencing, Nucleic Acids Research 45, e133-e133. 43. Erber, L., Hoffmann, A., Fallmann, J., Betat, H., Stadler, P. F., and Mörl, M. (2020) LOTTE- seq (Long hairpin oligonucleotide based tRNA high-throughput sequencing): specific selection of tRNAs with 3'-CCA end for high-throughput sequencing, RNA biology 17, 23-32. 44. Cervettini, D., Tang, S., Fried, S. D., Willis, J. C. W., Funke, L. F. H., Colwell, L. J., and Chin, J. W. (2020) Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase– tRNA pairs, Nature Biotechnology 38, 989-999. 45. Passarelli, M. C., Pinzaru, A. M., Asgharian, H., Liberti, M. V., Heissel, S., Molina, H., Goodarzi, H., and Tavazoie, S. F. (2022) Leucyl-tRNA synthetase is a tumour suppressor in breast cancer and regulates codon-dependent translation dynamics, Nature Cell Biology 24, 307- 315. 46. Davidsen, K., and Sullivan, L. B. (2023) A robust method for measuring aminoacylation through tRNA-Seq, eLife Sciences Publications, Ltd. 47. Tsukamoto, Y., Nakamura, Y., Hirata, M., Sakate, R., and Kimura, T. (2022) i-tRAP (individual tRNA acylation PCR): A convenient method for selective quantification of tRNA charging, RNA (New York, N.Y.) 29, 111-122. 48. Willis, J. C. W., and Chin, J. W. (2018) Mutually orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs, Nature chemistry 10, 831-837.49. Wan, W., Tharp, J. M., and Liu, W. R. (2014) Pyrrolysyl-tRNA synthetase: an ordinary enzyme but an outstanding genetic code expansion tool, Biochimica et biophysica acta 1844, 1059-1070. 50. Nozawa, K., O'Donoghue, P., Gundllapalli, S., Araiso, Y., Ishitani, R., Umehara, T., Söll, D., and Nureki, O. (2009) Pyrrolysyl-tRNA synthetase-tRNA(Pyl) structure reveals the molecular basis of orthogonality, Nature 457, 1163-1167. 51. Seki, E., Yanagisawa, T., Kuratani, M., Sakamoto, K., and Yokoyama, S. (2020) Fully Productive Cell-Free Genetic Code Expansion by Structure-Based Engineering of Methanomethylophilus alvus Pyrrolysyl-tRNA Synthetase, ACS Synthetic Biology 9, 718-732. 52. Ambrogelly, A., Gundllapalli, S., Herring, S., Polycarpo, C., Frauer, C., and Söll, D. (2007) Pyrrolysine is not hardwired for cotranslational insertion at UAG codons, Proceedings of the National Academy of Sciences of the United States of America 104, 3141-3146. 53. Chatterjee, A., Sun, S. B., Furman, J. L., Xiao, H., and Schultz, P. G. (2013) A Versatile Platform for Single- and Multiple-Unnatural Amino Acid Mutagenesis in Escherichia coli, Biochemistry 52, 1828-1837. 54. Zheng, Y., Addy, P. S., Mukherjee, R., and Chatterjee, A. (2017) Defining the current scope and limitations of dual noncanonical amino acid mutagenesis in mammalian cells, Chemical science 8, 7211-7217. 55. Niu, W., Schultz, P. G., and Guo, J. (2013) An expanded genetic code in mammalian cells with a functional quadruplet codon, ACS chemical biology 8, 1640-1645. 56. Dunkelmann, D. L., Willis, J. C. W., Beattie, A. T., and Chin, J. W. (2020) Engineered triply orthogonal pyrrolysyl-tRNA synthetase / tRNA pairs enable the genetic encoding of three distinct non-canonical amino acids, Nature chemistry 12, 535-544. 57. Jewel, D., Kelemen, R. E., Huang, R. L., Zhu, Z., Sundaresh, B., Cao, X., Malley, K., Huang, Z., Pasha, M., Anthony, J., van Opijnen, T., and Chatterjee, A. (2023) Virus-assisted directed evolution of enhanced suppressor tRNAs in mammalian cells, Nature methods 20, 95-103. 58. Young, D. D., Young, T. S., Jahnz, M., Ahmad, I., Spraggon, G., and Schultz, P. G. (2011) An evolved aminoacyl-tRNA synthetase with atypical polysubstrate specificity, Biochemistry 50, 1894-1900. 59. Italia, J. S., Addy, P. S., Wrobel, C. J. J., Crawford, L. A., Lajoie, M. J., Zheng, Y., and Chatterjee, A. (2017) An orthogonalized platform for genetic code expansion in both bacteria and eukaryotes, Nature chemical biology 13, 446-450.60. Italia, J. S., Latour, C., Wrobel, C. J. J., and Chatterjee, A. (2018) Resurrecting the Bacterial Tyrosyl-tRNA Synthetase / tRNA Pair for Expanding the Genetic Code of Both E. coli and Eukaryotes, Cell chemical biology 25, 1304-1312.e1305. 61. Chatterjee, A., Xiao, H., and Schultz, P. G. (2012) Evolution of multiple, mutually orthogonal prolyl-tRNA synthetase / tRNA pairs for unnatural amino acid mutagenesis in Escherichia coli, Proceedings of the National Academy of Sciences of the United States of America 109, 14841-14846. 62. Neumann, H., Slusarczyk, A. L., and Chin, J. W. (2010) De Novo Generation of Mutually Orthogonal Aminoacyl-tRNA Synthetase / tRNA Pairs, Journal of the American Chemical Society 132, 2142-2144. 63. Buechter, D. D., Schimmel, P., and de Duve, C. (1993) Aminoacylation of RNA Minihelices: Implications for tRNA Synthetase Structural Design and Evolution, Critical Reviews in Biochemistry and Molecular Biology 28, 309-322. 64. Frugier, M., Florentz, C., and Giegé, R. (1994) Efficient aminoacylation of resected RNA helices by class II aspartyl-tRNA synthetase dependent on a single nucleotide, The EMBO journal 13, 2218-2226. 65. Shiba, K., Ripmaster, T., Suzuki, N., Nichols, R., Plotz, P., Noda, T., and Schimmel, P. (1995) Human alanyl-tRNA synthetase: conservation in evolution of catalytic core and microhelix recognition, Biochemistry 34, 10340-10349. 66. Martinis, S. A., and Schimmel, P. (1994) Small RNA Oligonucleotide Substrates for Specific Aminoacylations, In tRNA, pp 349-370. 67. Cervettini, D., Tang, S., Fried, S. D., Willis, J. C. W., Funke, L. F. H., Colwell, L. J., and Chin, J. W. (2020) Rapid discovery and evolution of orthogonal aminoacyl-tRNA synthetase–tRNA pairs, Nature Biotechnology 38, 989-999. 68. Passarelli, M. C., Pinzaru, A. M., Asgharian, H., Liberti, M. V., Heissel, S., Molina, H., Goodarzi, H., and Tavazoie, S. F. (2022) Leucyl-tRNA synthetase is a tumour suppressor in breast cancer and regulates codon-dependent translation dynamics, Nature Cell Biology 24, 307-315. 69. Jewel, D., Kelemen, R. E., Huang, R. L., Zhu, Z., Sundaresh, B., Cao, X., Malley, K., Huang, Z., Pasha, M., Anthony, J., van Opijnen, T., and Chatterjee, A. (2023) Virus-assisted directed evolution of enhanced suppressor tRNAs in mammalian cells, Nature methods 20, 95-103.70. Dunkelmann, D. L.; Piedrafita, C.; Dickson, A.; Liu, K. C.; Elliott, T. S.; Fiedler, M.; Bellini, D.; Zhou, A.; Cervettini, D.; Chin, J. W., Adding α,α-disubstituted and β-linked monomers to the genetic code of an organism. Nature 2024, 625 (7995), 603-610.
Claims
CLAIMS What is claimed is:
1. A genetically-engineered tRNA molecule, wherein the tRNA molecule comprises / contains a unique barcode sequence introduced into the tRNA molecule at a permissive site in the tRNA molecule, wherein the barcode sequence does not interfere with aminoacylation activity by aminoacyl-tRNA synthetases.
2. The genetically-engineered barcoded tRNA molecule of claim 1, wherein the barcode sequence is introduced into the anticodon loop of the tRNA molecule, thus expanding the anticodon loop region of the tRNA.
3. The genetically-engineered barcoded tRNA molecule of either of claims 1or 2, wherein the barcode sequence is between about 5 and about 20 nucleotides in length.
4. The barcoded tRNA molecule of claim 3, wherein the barcode sequence is about 15 nucleotides in length.
5. The genetically-engineered barcoded tRNA molecule of any of claims 1 to 4, wherein the barcoded tRNA is tRNAMaPyl, wherein the tRNAMaPylcomprises at least about 80%, 85%,90% ,95%, 98% or 99% sequence identity with the full-length SEQ ID NO:
61.
6. A composition comprising a variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) wherein the variant MaPylRS preferentially aminoacylates a M. alvus pyrrolysyl-tRNA (tRNAMaPyl) with a non-canonical amino acid over the naturally- occurring pyrrolysine, lysine or phenylalanine amino acid, wherein the variant MaPylRS comprises the amino acid sequence of SEQ ID NO:58 or an amino acid sequence with at least about 80%, 85%, 90%, 95%, 98% or 99% sequence identity with the full-length SEQ ID NO:58, wherein the MaPylRS is mutated relative to SEQ ID NO:58 at amino acid residues leucine (L or Leu) 125; asparagine (N or Asn) 166, and valine (V or Val) 168.
7. The composition of claim 6, wherein the mutant MaPylRS comprises SEQ ID NO:58, or an amino acid sequence with at least about 80%, 85%, 90% , 95%, 98% or 99% sequence identity with the full-length SEQ ID NO:58, wherein the leucine (L) at position 125 is replaced with valine (V) or alanine (A) or threonine (T); the asparagine (N) at position 166 is replaced with serine (S) or alanine (A) or threonine (T); and the valine (V) at position 168 is replaced with cysteine (C) or tryptophan (W) or threonine (T).
8. The composition of either claims 6 or 7, wherein the mutant MaPylRS comprises an amino acid sequence selected from the group consisting of : SEQ ID NO: 75; SEQ ID NO: 76; SEQ ID NO:77; SEQ ID NO:78; SEQ ID NO: 79; SEQ ID NO:80; SEQ ID NO: 81 or SEQ ID NO:
82.
9. A cell comprising the genetically-engineered tRNA molecule of any of claims 1 to 4.
10. A cell comprising the variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS) of any of claims 6 to 8.
11. The composition of any of claims 6 to 8, comprising a variant / mutant M. alvus pyrrolysyl-tRNA aminoacyl synthetase (MaPylRS), wherein the non-canonical amino acid is a pyrrolysine, lysine or phenylalanine analog or a non-α-amino acid or non-α- amino acid analog.
12. The composition of claim 11, wherein the non-canonical amino acids are selected from the group consisting of: Nℇ-((tert-butoxy)carbonyl)-L-lysine / Nℇ-Boc-L-Lysine (BocK); OH-BocK; H-BocK; m-iodo-L-phenylalanine (mIF); o-nitro-L-phenylalanine (cNF) or o- cyano-L-phenylalanine (2-CNF).
13. A method to identify mutated / variant aminoacyl-tRNA synthetases (aaRS) that selectively incorporate non-canonical amino acids into proteins, wherein the method is independent of ribosomal translation, and wherein each mutated / variant aaRS comprises a nucleotide sequence barcode corresponding / correlated to a tRNA molecule containing the same nucleotide sequence barcode.
14. The method of claim 13, wherein the method comprises acylation of the barcoded tRNA molecule by the mutant / variant aminoacyl-tRNA synthetase resulting in protection of the barcoded tRNA molecules from oxidation treatment.
15. The method of either of claims 13 or 14, wherein the method further comprises sequencing the barcoded tRNA molecule to identify the aaRS mutant that acylated the corresponding / correlated tRNA molecules.
16. The method of any of claims 13-15, wherein the non-canonical amino acid is a pyrrolysine, lysine or phenylalanine analog or a non-α-amino acid or non-α-amino acid analog.
17. The method of any of claims 13-16, the steps comprising combining a library of sequence barcoded tRNA molecules and a library of mutant / variant aminoacyl-tRNA synthetases (aaRS) into a single DNA molecule, wherein the sequences of the mutant / variant aminoacyl-tRNA synthetases correspond to the sequence barcodes of the tRNA molecules under suitable conditions in a cell for expression of the tRNA molecules and the aaRSs, wherein the tRNA molecules are acylated by the mutant aminoacyl-tRNA synthetases in the presence of the non-canonical amino acid; subjecting the resulting combination of acylated tRNA molecules and unacylated tRNA molecules to oxidation conditions whereby the acylated tRNA molecules are protected from oxidation; de- acylating the unoxidized, acylated tRNA molecules in the combination; extending the 3’ end of the deacylated tRNA molecules to attach an adapter sequence; annealing a complementary nucleotide sequence to the 3’-adapter sequence of the tRNA molecules; selectively reverse transcribing and PCR amplifying the 3’-adapter-containing tRNA sequences; and sequencing the amplified tRNA sequences and the aaRSs to determine the barcode sequences and identify the corresponding aaRS mutant with the specific activity to aminoacylate noncanonical amino acids.
18. A method of identifying specific active mutant / variant aminoacyl-tRNA synthetase (aaRS) molecules in the absence of ribosomal translation, which can acylate a cognate tRNA with non-canonical monomers, the steps comprising: a.) Providing a gene library of cognate tRNA molecules, wherein the tRNA moleculesare tagged with a unique oligonucleotide sequence (barcoded) in a manner that does not jeopardize their expression, folding, and aminoacylation by the aaRS; b.) Providing a gene library of mutant / variant aminoacyl-tRNA synthetase (aaRS) molecules, in the same plasmid as the tRNA library from step a.), wherein each aaRS variant is correlated to a barcode sequence of a tRNA molecule of the tRNA library of step a.); c.) Co-expressing the library of aaRS molecules with the correlated barcode-containing tRNA molecules under conditions specific for charging the tRNA molecules with a cognate aaRS molecule to form a combination of tRNA molecules charged with the desired noncanonical monomer and uncharged tRNA molecules; d.) Contacting the combination of step c.) with a periodate oxidation buffer solution under conditions suitable for the oxidation of uncharged tRNA molecules thereby forming a mixture of oxidized uncharged tRNA molecules and unoxidized charged tRNA molecules; e.) Deacylating the mixture of step d.) to obtain tRNA molecules with an intact 3’- terminus; f.) Introducing a unique DNA oligonucleotide sequence at the intact 3’-terminus of the tRNA from step e.), either by ligation, or by DNA polymerase mediated extension of the 3’-terminus of the tRNA after hybridizing it with a suitable DNA template; g.) Isolating the tRNA molecules containing the unique DNA oligonucleotide sequence at the 3’-terminus of step f.) by non-native PAGE or other methods (e.g., oligonucleotide or streptavidin mediated pulldown); h.) Amplifying the isolated tRNA molecules containing the unique DNA oligonucleotide sequence at the 3’-terminus of step f.) by RT PCR using an oligonucleotide primer that selectively binds the unique DNA oligonucleotide sequence at the 3’-terminus; i.) Sequencing the tagged, amplified tRNA molecules to obtain the sequence of thetRNA molecules and the corresponding barcodes; and j.) Correlating the sequence of the barcodes of the enriched tRNA population with the sequence of the variant aaRS molecule to identify specific active mutant / variant aminoacyl-tRNA synthetase (aaRS).
19. A method of producing a protein or peptide of interest in an E.coli, eukaryotic or mammalian cell with one, or more, noncanonical amino acid analogs at specified amino acid residue positions in the protein or peptide, the method comprising the steps of, a.) culturing the cell in a culture medium under conditions suitable for growth, wherein the cell comprises a nucleic acid encoding a protein or peptide of interest with one, or more, nonsense codons incorporated at the one, or more specified positions in the protein or peptide, wherein the cell further comprises a nucleic acid encoding a tRNAMaPylthat recognizes the nonsense codon, and a nucleic acid encoding a mutant MaPylRS that aminoacylates the tRNAMaPyl, wherein the mutant MaPylRS is selected from the group consisting of: mutant VSV SEQ ID NO: 75; mutant LSV SEQ ID NO: 76; mutant LAC SEQ ID NO: 77; mutant LAV SEQ ID NO:78; mutant AAV SEQ ID NO: 79; mutant ASV SEQ ID NO:80; mutant TTW SEQ ID NO: 81 or mutant AST SEQ ID NO:82; and b.) contacting the cell culture medium with one, or more, pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs under conditions suitable for incorporation of the one, or more, pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs into the protein at the site or sites of the nonsense codon and expression of the protein or peptide, thereby producing the protein or peptide of interest with one, or more pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs at specified positions in the protein or peptide.
20. The method of claim 19, wherein the noncanonical amino acid analogs are selected from the group consisting of: pyrrolysine, lysine, phenylalanine or non-α-amino acid analogs.
21. The method of either of claim 19 or 20, wherein the pyrrolysine, lysine, phenylalanine or non-α-amino acid analog is selected from the group consisting of: Nε-Boc-L-lysine (BocK), m-iodo-L-phenylalanine (mIF), OH-BocK, H-BocK, o-nitro-L-phenylalanine (oNF) and o-cyano-L-phenylalanine (2-CNF).
22. A kit for producing a protein or peptide of interest in a cell, wherein the protein or peptide comprises one, or more noncanonical amino acid analogs, wherein the components of the kits comprise a container containing a polynucleotide sequence encoding an tRNAMaPylthat recognizes a nonsense / stop codon; and a container containing a polynucleotide encoding a mutant MaPylRS that preferentially aminoacylates the tRNAMaPylwith a noncanonical amino acid analog, wherein the mutant MaPylRS comprises SEQ ID NO:58, or an amino acid sequence with at least about 80%, 85%, 90% , 95%, 98% or 99% sequence identity with the full-length SEQ ID NO:58, wherein the leucine (L) at position 125 is replaced with valine (V) or alanine (A) or threonine (T); the asparagine (N) at position 166 is replaced with serine (S) or alanine (A) or threonine (T); and the valine (V) at position 168 is replaced with cysteine (C) or tryptophan (W) or threonine (T).
23. The kit of claim 22, wherein the mutant MaPylRS comprises an amino acid sequence selected from the group consisting of : mutant VSV SEQ ID NO: 75; mutant LSV SEQ ID NO: 76; mutant LAC SEQ ID NO: 77; mutant LAV SEQ ID NO:78; mutant AAV SEQ ID NO: 79; mutant ASV SEQ ID NO:80; mutant TTW SEQ ID NO: 81 or mutant AST SEQ ID NO:
82.
24. The kit of either claims 22 or 23, wherein the kit further comprises one, or more of the following containers of the desired noncanonical amino acid analogs, buffers, diluents and other components necessary for producing a protein or peptide of interest; and further comprises instructions for producing the protein or peptide of interest.
Citation Information
Patent Citations
ENCODING AND EXPRESSION OF ACE-tRNAs
WO2021252354A1
PYRROLYSYL-tRNA SYNTHETASE VARIANTS AND USES THEREOF
WO2022003142A1
Hairpin oligonucleotides and uses thereof
WO2022099010A2
Aminoacyl-trna synthetase capable of efficiently introducing lysine derivatives and use thereof
WO2023011486A1