Improvements in trna selection methods
The method enhances the identification of orthogonal aaRS-tRNA pairs by expressing split tRNAs in prokaryotic cells, facilitating the efficient incorporation of non-canonical amino acids into proteins through improved isolation and analysis of active aaRS variants.
Patent Information
- Application Number
- PCT/EP2025/072671
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-08-06
- Publication Date
- 2026-02-12
AI Technical Summary
Current methods for identifying orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pairs for incorporating non-canonical amino acids into proteins are limited, restricting the efficiency and scope of genetic code expansion.
A method involving split tRNA molecules expressed in prokaryotic cells, such as E. coli, where the tRNA is divided at the anticodon loop and extended to promote assembly, coupled with a linked aaRS gene, allowing for enhanced isolation and identification of active aaRS variants using biotinylated stmRNA enrichment and qPCR analysis.
Improves the sensitivity and efficiency of selecting aaRSs capable of charging tRNAs with non-canonical monomers, enabling the incorporation of multiple distinct non-canonical amino acids into proteins.
Smart Images

Figure 00000027_0000 
Figure 00000028_0000 
Figure 00000028_0001
Abstract
Description
[0001] Improvements in tRNA Selection Methods
[0002] The present invention concerns improvements in methods for determining the aminoacylation status of a tRNA and for selecting aminoacyl tRNA synthetases capable of charging a tRNA with a desired amino acid or non-canonical monomer. In particular, the present invention provides E. coli host cells which are modified to significantly increase the efficiency of aminoacyl tRNA-synthetase mRNA isolation based on tRNA acylation.
[0003] Background of the invention
[0004] Genetic code expansion enables the cellular synthesis of modified proteins via the co-translational incorporation of non-canonical amino acids (Chin JW. Nature. 2017; 550:53-60; Young DD, Schultz PG. ACS Chem Biol. 2018; 13:854-870). Orthogonal aaRS-tRNA pairs are crucial to genetic code expansion. These pairs consist of (1 ) a synthetase that efficiently aminoacylates its cognate tRNA, but minimally aminoacylates endogenous tRNAs in the host organism, and (2) a tRNA that is a substrate for its cognate synthetase but is a poor substrate for endogenous synthetases (Liu CC, Schultz PG. Annu Rev Biochem. 2010; 79:413-444; Chin JW. Annu Rev Biochem. 2014; 83:379-408). Derivatives of orthogonal pairs that recognize blank codons (most commonly, the amber stop codon) and selectively use non-canonical amino acids (ncAAs) have been used to site-specifically incorporate numerous ncAAs into proteins (Liu CC, Schultz PG. Annu Rev Biochem. 2010;
[0005] 79:413-444; Chin JW. Annu Rev Biochem. 2014; 83:379-408). Despite many years of effort, a limited number of orthogonal aaRS-tRNA pairs have been described, and the vast majority of these are limited to the incorporation of non-canonical amino acids and a few hydroxy acids. The discovery of aminoacyl-tRNA synthetase / tRNA pairs for the incorporation of other types of non-canonical monomers, such as betaamino acids, gamma-amino acids, alpha, alpha-disubstituted amino acids, N- alkylated amino acids and other classes of monomers has been hampered due to a paucity of methods to discover and select these pairs. The current orthogonal pairs may limit the efficiency and scope of non-canonical monomer incorporation, the incorporation of multiple non-canonical monomers and progress towards the encoded cellular synthesis of non-canonical biopolymers (Chin JW. Nature. 2017; 550:53-60). The discovery of new orthogonal pairs may enable an increase in efficiency or scope of ncAA incorporation. Moreover, the discovery of additional aaRS-tRNA pairs that are orthogonal with respect to both the host synthetases and tRNAs and each other will provide one key component required for the incorporation of multiple distinct non-canonical monomers and the encoded cellular synthesis of non-canonical biopolymers.
[0006] Cervettini et al (Nat Biotechnol . 2020 August 01 ; 38(8): 989-999) have developed a technique for analytically assessing the acylation status of a tRNA, which assists in identifying orthogonal aaRS-tRNA pairs which can be used to incorporate ncAAs into proteins. This technique is known as tRNA extension or tREX.
[0007] While tREX enables the detection of the aminoacylation status of a tRNA, the tREX technique, cannot in itself provide novel aaRS capable of incorporating new non- canonical monomers into proteins or other types of biopolymers. In a recent development, Dunkelmann et al (2024 Nature 626:603-610) have coupled the tREX acylation determination technique with a split-RNA based system which selects for orthogonal aaRS which specifically acylate cognate orthogonal tRNAs with amino acids or non-canonical monomers. By providing a library of aaRS genes coupled to tRNA molecules, it is proposed to select novel orthogonal aaRS / tRNA pairs based on the ability of the aaRS to acylate a tRNA with a given substrate. Both the tRNA and the sequence of the candidate aaRS are encoded within the same mRNA construct. This construct is processed in cells to generate stmRNAs, which can then be selectively recovered on the basis of the acylation status of the tRNA portion of the stmRNA. This procedure is known as tRNA display. We have determined that the efficiency of isolation of stmRNAs encoding aaRSs capable of acylating a tRNA with a non-canonical monomer in the tRNA display method is improved several fold when the method is performed in strains in which RNAse E has been mutated or truncated. This allows for higher sensitivity of the method in the recovery of stmRNAs comprising acylated tRNAs, and improves the sensitivity of selection and isolation of aaRSs that incorporate non-canonical monomers.
[0008] Summary of the Invention
[0009] In a first aspect, there is provided a method for expressing an acylated tRNA, comprisingthe steps of:
[0010] (1 ) expressing a gene encoding a tRNA molecule in a prokaryotic host cell, wherein the tRNA molecule is divided at the anticodon loop sequence to form two tRNA parts, and the anticodon stem in each tRNA part is extended to promote annealing of the tRNA parts to form a complete tRNA molecule;
[0011] (2) exposing said tRNA to conditions in the host cell such that the tRNA is acylated; and
[0012] (3) isolating said tRNA from the host cell, wherein the host cell does not comprise a wild-type RNase E gene.
[0013] The gene encoding a tRNA molecule is, in one embodiment, a gene encoding a split tRNA, in which the tRNA molecule is divided at the anticodon loop into two halves, which are not necessarily equal halves. The two halves are designed to assemble in vivo via non-covalent interactions, such as base pairing. For example, the anticodon stem / loop structure of the tRNA molecule can be replaced or extended with a sequence which is 0-14 nucleotides in length, preferably 6 to 12 nucleotides in length, and most preferably 10 nucleotides in length. The tRNA does not require an anticodon, since the method of the invention assesses acylation of the tRNA, not its codon specificity. The two halves of the tRNA may be expressed in trans, that is from two different plasmids or different genes, and assembled in vivo. Alternatively, the two halves are expressed in cis from one transcript, using an intervening sequence loop to join the 3’ and 5’ segments of the entire tRNA. The loop is advantageously removed post- transcriptionally, and is preferably based on a prokaryotic intergenic region from a polycistronic tRNA operon. These regions are removed by the cellular machinery to produce an intact tRNA; the advantage of the cis expression setup is that the stoichiometry is equimolar, leading to more efficient tRNA assembly, and the cis expression setup may further facilitate tRNA assembly and processing through proximity effects . The split tRNA construct is referred to as stRNA.
[0014] Acylation of the tRNA is dependent on the presence of a cognate aminoacyl-tRNA synthetase (aaRS) which can charge the tRNA with the desired monomer, and the monomer itself. Accordingto Dunkelmann et al. (2024), an aaRS gene can be covalently linked to the split tRNA gene, thus providing linked expression of tRNA and aaRS; the resulting genetic construct is termed stmRNA (split tRNA and mRNA). By expressing a single stmRNA gene in a cell, that cell can be limited to expression of one tRNA and one aaRS. The genotype and phenotype of the aaRS are therefore linked in each prokaryotic cell. Preferably, the aaRS coding sequence is linked to the tRNA coding sequence via a linker. The gene construct can be placed under the control of a suitable prokaryotic promoter, such as the inducible T7 promoter.
[0015] For many applications, orthogonal pyrrolysyl-tRNA synthetases from Methanocarcina mazei (MmPylRS) and M. barker! (MbPylRS) have been used for incorporation of ncAAs into proteins. MmPylRS is used as an example in the following description. In accordance with the invention, however, any aaRS gene, and libraries of mutated aaRS genes, can be fused to its cognate split tRNA gene to generate stmRNAs, and subsequently identify novel cognate aaRS for any selected tRNA and substrate. Suitable substrates can be chosen from amongst canonical alpha-amino acids and non-canonical monomers. Generally, by the term “non-canonical monomers”, are encompassed also non-alpha-amino acids such as hydroxy acids, beta-amino acids, alpha, alpha-disubstituted amino acids, N-alkylated amino acids, and any monomer which is capable of being charged onto a tRNA and incorporated into a polypeptide.
[0016] In one aspect, the stmRNA expressed by the host cell is isolated from the host cell and analysed for acylation. Analysis can be performed by fluoro-mREX or bio-mREX methods described by Dunkelmann (2024). Preferably, stmRNA samples are subjected to selective polymerase extension such that only stmRNAs comprising acylated tRNAs are extended. Biotin can be incorporated in the extension procedure, such that only stmRNA comprising acylated tRNA is biotinylated, and selectively captured on streptavidin beads, enriching for stmRNA molecules which encode an active aaRS variant.
[0017] The stmRNA mixture, which comprises mRNA sequences encoding active aaRSs, can be reverse transcribed to cDNA. qPCR of the cDNA can determine the abundance of acylated tRNA recovered in different experimental regimes. Sequencing can reveal the structure of active aaRS variants.
[0018] In a further aspect, stmRNA genes can be expressed in the presence and absence of noncanonical monomers (ncMs). If the method is performed on a single aaRS variant that is active for an ncM of interest, the relative abundance of transcripts after enrichment between the -ncM and the +ncM samples as measured by qPCR determines the method’s efficiency for isolating stmRNA molecules comprising acylated tRNAs, and hence active aaRS mRNA sequences. If the method is performed with a pool of aaRS sequences, the relative abundances of the different aaRS sequences across the -ncM sample, the +ncM sample, and the input sample are used to identify the sequences of aaRSs that are active and selective for an ncM, as defined by two parameters: a. Sequence enrichment indicates aaRS activity, towards an AA or an ncM. It is measured as the relative abundance of an aaRS sequence in the +ncM sample relative to the input sample. Sequences that are more abundant in the +ncM sample than in the input sample will have been enriched in the biotin pulldown, and may hence be active. b. Selectivity - indicates the specificity of an aaRS for the ncM of interest, over other natural AAs that are present in the cell. It is measured as the relative abundance of an aaRS sequence in the +ncM sample relative to the -ncAA sample. Sequences that appear in the +ncM sample but not in the -ncM sample are likely to accept the ncM of interest, but not other ncM or AA.
[0019] For single active aaRSs, Dunkelmann et al 2024 describes levels of enrichment in molecule counts as measured by qPCR of around 100x or more in the absence vs presence of the cognate ncM. It is expected that the detected molecule count enrichment may differ across aaRS variants and substrates.
[0020] The prokaryotic cell is advantageously an E. coli, cell, and preferably an E. coli BL21 (DE3) cell.
[0021] The RNase E gene may be deleted, truncated or mutated; in accordance with the invention, RNase E activity is inhibited such that the stability of the stmRNA constructs in cells is enhanced, enablingthe recovery of more stmRNA molecules from cells as measured by qPCR. In a preferred embodiment, the RNase E gene comprises the me131 truncation, as in the BL21 star (DE3) strain. In another preferred embodiment, the RNase E gene comprises mutations at 1 or more, 2 or more, 3 or more, 4 or more or all of the following positions: S557, S591 , T595, V660 and T669. Advantageously, the mutations are selected from the following:
[0022] • S557A
[0023] • S591T
[0024] • T595A
[0025] • V660A
[0026] • T669A In one example, the mutated RNase E gene comprises all 5 of the above-recited mutations.
[0027] Brief Description of the Figures
[0028] Figure 1 is a schematic of the cis split tRNA-mRNA fusion (stmRNA) gene, and the production of stmRNA.
[0029] Figure 2 is a schematic of biotin mRNA extension (bio-mREX). Biotinylated stmRNAs are enriched on streptavidin beads. The mRNA of PylRS within is reverse transcribed on the beads and quantified by qPCR.
[0030] Figure 3A shows the results of a bio-mRex experiment of an smtRNA construct where the 3’ half of the split tRNA is linked to a sequence that encodes the wild-type Methanosarcina mazei PylRS, which is known to be active and selective for the non- canonical amino acid AllocK. The experiment is carried out in E. coli BL21 (DE3), commercial strain (strain 1).
[0031] Figure 3B shows an experiment analogous to that of Figure 3A, except that a mutant E. coli BL21 (DE3) is used (strain 2) with mutations in the RNase E gene.
[0032] Figure 4 shows the results of a bio-mRex experiment of an smtRNA construct where the 3’ half of the split tRNA is linked to a sequence that encodes the wild-type Methanosarcina mazei PylRS, which is known to be active and selective for the non- canonical amino acid BocK (panel A) and (S)-6-acetamido-3-aminohexanoic acid ((S)- 3AcK) (BocAhx, panel B), in strain 1 (wild type RNase E) and strain 2 (mutant RNase E).
[0033] Figure 5 shows a bio-mRex experiment of an smtRNA construct encoding the wildtype Methanosarcina mazei PylRS which is known to be active and selective for the non-canonical amino acid BocK, in strain 2 (mutant RNase E) and star (truncated RNase E). Detailed Description of the Invention
[0034] The terms “comprising”, “comprises” and “comprised of” as used herein are synonymous with “including” or “includes”; or “containing” or “contains”, and are inclusive or open-ended and do not exclude additional, non-recited members, elements or steps. The terms “comprising”, “comprises” and “comprised of” also include the term “consisting of”. aaRS / tRNA pairs
[0035] Genetic code expansion uses an orthogonal aminoacyl-tRNA synthetase (aaRS)- tRNA pair to direct the incorporation of non-proteinogenic amino acids into proteins, in response to an unassigned codon (e.g. the amber stop codon, UAG) introduced at the desired site in a gene of interest. The orthogonal synthetase does not recognize endogenous tRNAs, and specifically aminoacylates an orthogonal cognate tRNA (which is not an efficient substrate for endogenous synthetases) with the monomer provided to (or synthesized by) the cell (Chin, J.W., 2017. Nature, 550(7674), 53-60).
[0036] The present invention comprises expression of cognate aaRS / tRNA pairs in a bacterial cell. The aaRS is orthogonal to the bacterial cell. Methods have been described to identify and / or generate orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pairs (e.g. Elliott, T. S. et al., 2014. Nat Biotechnol 32, 465-472; Elliott, T. S., et al., 2016. Cell Chem Biol 23, 805-815; and Krogager, T. P. et al., 2018. Nat Biotechnol 36, 156-159). However only a limited number of orthogonal aaRS are known in the art. The prokaryotic cell of the present invention comprises one or more heterologous nucleotides (e.g. plasmids) encoding one orthogonal aminoacyl- tRNA synthetase (aaRS)-tRNA pair. In preferred embodiments the prokaryotic cell of the present invention further comprises a plasmid encoding an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair. Alternatively, the orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair may be introduced into the prokaryotic cell by incorporation into the prokaryotic cell genome. Thus, in some embodiments the prokaryotic genome encodes an orthogonal aminoacyl-tRNA synthetase (aaRS)-tRNA pair.
[0037] Orthogonal tRNA synthetase enzymes and paired tRNAs may be obtained from any suitable source. Organisms from which aaRS may be derived include Methanosarcina mazei, Archeogiobus fuigidus, Methanomethyiophiius sp, Methanocaidococcus jannaschii and Methanosarcina barken. For example, orthogonal aaRS / tRNA pairs include Methanosarcina mazei (Mm)PylRS / MmtRNAPylcGA, Archeogiobus fulgidus (Af)TyrRS(plF) / AftRNATyr(A01)cuA, Methanomethyiophiius sp. 1 R26 (1 R26)PylRS(CbzK) / AlvtRNAANPyl(8)cuA and Methanocaldococcus jannaschii (M / )TyrRS(Nap) / M / tRNATyrCUA. Methanosarcina barken PylRS has also been used to incorporate hydroxy acids.
[0038] Advantageously, the aaRS is mutagenized in order to select variants which are able to incorporate different monomers. For example, a library of aaRS can be constructed by randomising positions M300, L301 , A302, M344 and N346 in MmPylRS with degenerate codons. Further mutagenesis can be used to select mutants better able to incorporate different ncMs with different sidechains; preferably, mutations are selected for at positions C348, V401 and W417, which delimit the pocket to the enzyme which binds to the substrate sidechain. Other mutations can be introduced, based on the principle of variation of the enzyme in the region which interacts with the sidechain of the monomer.
[0039] Split tRNA
[0040] Split tRNA genes may be prepared accordingto the method of Dunkelmann et al., 2024. Advantageously, the tRNA gene is split in the anticodon loop; however splits at other locations may be envisaged.
[0041] For example, the sequence of the anticodon stem loop in each half of the tRNAcuA Pyl gene can be replaced with an extension of 0 to 14 nucleotides in length; the extensions within each pair of tRNA halves are designed to base pair with each other and form a stem to stabilize the stRNAPyl.
[0042] The pair of tRNA halves may be expressed in trans (from two different plasmids) in the presence or absence of PylRS and BocK. For stRNAPyl molecules with stems comprising more than 12 bp, gel bands consistent with degradation products are observed. Stems of 10 or 12 bp result in robust aminoacylation, which is dependent on the presence of both tRNA halves, PylRS and BocK. A 10-bp stem is sufficient to facilitate the association of the 2 tRNA halves, minimize degradation and enable robust aminoacylation of the assembled tRNA.
[0043] To ensure equimolar stoichiometry of both tRNA halves and facilitate spatial proximity, and thereby the assembly, of the two halves, the two tRNA halves can be expressed in cis from one transcript by inserting an intervening sequence between them, thereby creating a circular permutation of the parent tRNA.
[0044] Processing to remove the intervening sequence yields the stRNA. The intergenic regions of polycistronic tRNA operons in E. coli connect the 3' end of a tRNA gene to the 5' end of the following tRNA gene and are efficiently removed from the resulting transcripts. The software tRNA operon generator (Dunkelmann, et al., Nat. Chem. 13, 1110-1117 (2021 )) can be used to select E. coli intergenic regions as the intervening sequences for the circular permutation strategy.
[0045] Linkage of aaRS and stRNA genes
[0046] The linkage can be performed accordingto Dunkelmann et al., 2024. The aaRS coding sequence and a linker sequence can be fused to the 5' end of the 3' half of the cognate cis stRNA. The resulting stmRNA cassette can be placed under the control of a suitable promoter, such as an inducible promoter, optimally the inducible T7 promoter. Transcription, processing and maturation of this construct leads to a stmRNA in which the mRNA encoding the synthetase is covalently linked to the 5' end of the 3' half of the stRNA, and associated with the 5' half of the stRNA.
[0047] Translation of the synthetase mRNA within the stmRNA generates the synthetase protein, which — in the presence of its cognate ncM — would acylate the stmRNA in the same cell. This generates a covalent link between the monomer attached to the tRNA and the mRNA of the synthetase gene that catalysed the attachment.
[0048] Acylation-specific enrichment of tRNA
[0049] Methods for enrichment of tRNA which is selective for acylated molecules have been described in Cervettini et al., 2020, and Dunkelmann et al., 2024.
[0050] Selective isolation of acylated stmRNAs with respect to non-acylated stmRNAs can be carried out using bio-mREX (Dunkelmann et al., 2024), in which biotinylated stmRNAs (which result from aminoacylated stmRNAs) are selectively captured on streptavidin beads. The aaRS mRNAs are then directly reverse transcribed from the stmRNA captured on the beads, to create aarRS cDNA. The aaRS cDNA is then released by RNase H treatment and heating, and quantified by quantitative PCR (qPCR).
[0051] Prokaryotic cell
[0052] As used herein, the term “prokaryotic cell”, in the context of the invention, refers to a unicellular organism that lacks a membrane-bound nucleus, mitochondria, or any other membrane-bound organelle. Prokaryotes are divided into two domains, Archaea and Bacteria. The genome of prokaryotic organisms generally is a circular, double-stranded piece of DNA, multiple copies of which may exist at any time. Preferably, the prokaryotic cell of the present invention is a bacterial cell. Preferably, but not necessarily, the prokaryotic cell is suitable for heterologous protein production, in particular the production of polypeptides and non-peptide polymers comprising one or more canonical or non-canonical amino acids (for instance those described by Ferrer-Miralles, N. and Villaverde, A., 2013. Microbial Cell Factories, 12:113). Suitable bacterial cells include: Escherichia (e.g. Escherichia coli), caulobacteria (e.g. Cauiobacter crescentus), phototrophic bacteria (e.g. Rodhobacter sphaeroides), cold adapted bacteria (e.g. Pseudoaiteromonas haiopianktis, Shewaneiia sp. strain Ac10), pseudomonads (e.g. Pseudomonas fiuorescens, Pseudomonas putida, Pseudomonas aeruginosa), halophilic bacteria (e.g. Haiomonas elongate, Chromohalobacter salexigens), streptomycetes (e.g. Streptomyces lividans, Streptomyces griseus), nocardia (e.g. Nocardia lactamdurans), mycobacteria (e.g. Mycobacterium smegmatis), coryneform bacteria (e.g. Corynebacterium glutamicum, Corynebacterium ammoniagenes, Brevibacterium lactofermentum), bacilli (e.g. Bacillus subtilis, Bacillus brevis, Bacillus megaterium, Bacillus licheniformis, Bacillus amyloliquefaciens), and lactic acid bacteria (e.g. Lactococcus lactis, Lactobacillus plantarum, Lactobacillus casei, Lactobacillus reuteri, Lactobacillus gasseri) cells. In some embodiments the prokaryotic cell is a gram-negative bacterial cell.
[0053] Preferably, the prokaryotic cell of the present invention is an Escherichia coli, Salmonella enterica, or Shigella dysenteriae cell. These are phylogenetically related species as disclosed by Lukjancenko, O., et al., 2010. Microbial ecology, 60(4), pp.708-720; and Karberg, K.A., et al., 2011 . PNAS, 108(50), pp.20154-20159.
[0054] More preferably, the prokaryotic cell of the present invention is an E. coli cell. Any suitable E. coli is contemplated, including K-12, MG1655, BL21 , BL21 (DE3), AD494, Origami, HMS174, BLR(DE3), HMS174(DE3), Tuner(DE3), Origami2(DE3), Rosetta2(DE3), Lemo21 (DE3), NiCo21 (DE3), T7 Express, SHuffle Express, C41 (DE3), C43(DE3), and ml 5 pREP4 or derivatives thereof (Rosano, G.L. and Ceccarelli, E.A., 2014. Frontiers in microbiology, 5, p.172). Most preferably, the prokaryotic cell is MG1655 or BL21 or a derivative thereof. MG1655 is considered as the wild type strain of E coli. The GenBank ID of genomic sequence of this strain is U00096. BL21 is widely available commercially. For example, it can be purchased from New England BioLabs with catalog number C2530H (https: / / www.neb.com / products / c2530-bl21- competent-e-coli). The preferred organism is E. coli BL21 (DE3). In some embodiments, one or more tRNA or release factors may be deleted from the prokaryotic cell and the cell may remain viable. For example, a tRNA which decodes only the one or more sense codons that have been replaced (or deleted) may be dispensable. Similarly, a tRNA which decodes the one or more sense codons that have been replaced (or deleted) may be dispensable if the remaining sense codons that it decodes may also be decoded by an alternative tRNA. For example, serT, encoding tRNASeruGA, is the only tRNA that decodes TCA codons in E. coll, and is therefore normally essential. However, if the genome of the prokaryotic cell does not contain TCA codons then serT may be dispensable.
[0055] Methods for modifying bacterial cells for the production of polymers comprising non-canonical amino acids are set forth in WO2020229592. Such methods are useful for producing prokaryotic cells useful for the production of polymers comprising non-canonical monomers.
[0056] When the genome of a cell has been modified, the prokaryotic cell preferably does not display a substantially decreased growth rate. Thus, preferably the prokaryotic cell does not have a substantially decreased growth rate relative to the host cell comprising the parent genome. In some embodiments the prokaryotic cell has a doubling time less than 4 times, 3 times, 2 times, or about 1 .6 times, slower than the parent cell. The doubling time can be determined by any method known to those of skill in the art. In some embodiments the doubling time is determined at 37°C, 25°C or 42°C, in LB media.
[0057] Non-canonical amino acids
[0058] As used herein, “noncanonical amino acids” are amino acids that are not naturally encoded orfound in the genetic code. Despite the use of only 22 amino acids by the translational machinery to assemble proteins (the proteinogenic amino acids comprise 20 in the standard genetic code and an additional 2 that can be incorporated by special translation mechanisms), over 140 amino acids are known to occur naturally in proteins and thousands more may occur in nature or be synthesized in the laboratory. Thus, non-canonical amino acids may comprise any amino acid excluding L-alanine, L-cysteine, L-aspartic acid, L-glutamic acid, L- phenylalanine, glycine, L-histidine, L-isoleucine, L-lysine, L-leucine, L-methionine, L- asparagine, L-proline, L-glutamine, L-arginine, L-serine, L-threonine, L-valine, L- tryptophan and L-tyrosine. Unnatural amino acids additionally exclude L-pyrrolysine and L-selenocysteine.
[0059] In some embodiments, the non-canonical amino acids are unnatural amino acids (UAAs).
[0060] Suitable non-canonical amino acid and UAAs will be well known to those of skill in the art, for example those disclosed in Neumann, H., 2012. FEBS letters, 586(15), pp.2057-2064; and Liu, C.C. and Schultz, P.G., 2010. Annual review of biochemistry, 79, pp.413-444. In some embodiments the non-proteinogenic amino acid and / or UAAs are selected from one or more of: p-Acetylphenylalanine, m- Acetylphenylalanine, O-allyltyrosine, Phenylselenocysteine, p- Propargyloxyphenylalanine, p-Azidophenylalanine, p-Boronophenylalanine, O- methyltyrosine, p-Aminophenylalanine, p-Cyanophenylalanine, m- Cyanophenylalanine, p-Fluorophenylalanine, p-lodophenylalanine, p- Bromophenylalanine, p-Nitrophenylalanine, L-DOPA, 3-Aminotyrosine, 3- lodotyrosine, p-lsopropylphenylalanine, 3-(2-Naphthyl)alanine, Biphenylalanine, Homoglutamine, D-tyrosine, p-Hydroxyphenyllactic acid, 2-Aminocaprylic acid, Bipyridylalanine, HQ-alanine, p-Benzoylphenylalanine, o-Nitrobenzylcysteine, o- Nitrobenzylserine, 4,5-Dimethoxy-2-nitrobenzylserine, o-Nitrobenzyllysine, o- Nitrobenzyltyrosine, 2-Nitrophenylalanine, Dansylalanine, p- Carboxymethylphenylalanine, 3-Nitrotyrosine, Sulfotyrosine, Acetyllysine, Methylhistidine, 2-Aminononanoic acid, 2-Aminodecanoic acid, Pyrrolysine, Cbz- lysine, Boc-lysine and Allyloxycarbonyllysine.
[0061] Non-canonical monomers are monomers that lackthe RCH(NH2)COOH structure of an amino acid. Alpha hydroxy acids, beta-amino acids, gamma-amino acids, alpha, alpha-disubstituted amino acids, and N-alkylated amino acids are exemplary non-canonical monomers. Some non-canonical monomers are variants of canonical amino acids. Alternatively, they may be derivatives of non-canonical acids. For example, alpha-hydroxy acids include hydroxy acids with aromatic sidechain, optionally selected from F-OH, pIF-OH and NapA-OH; and alpha-hydroxy acids with an aliphatic side-chain, optionally selected from BocK-OH, PenK-OH, AllocK-OH, NorK-OH, AlkynK-OH, CbzK-OH, ButK-OH and AcK-OH. Hydroxy-acid analogues of O4BBy, O2beY, pCaaF, pVsaf, pAaF are also contemplated. See (lannuzzelli and Fasan, Chem. Sci., 2020, 11 ,6202).
[0062] Alpha-hydroxy acids
[0063] In general, alpha-hydroxy acids can be derived from canonical or noncanonical amino acids. For example, alpha-hydroxy acids include, but are not limited to, p- hydroxy-L-phenyllactic acid (the alpha-hydroxy analogue of tyrosine), leucic acid (the alpha-hydroxy analogue of leucine), lactic acid (the alpha-hydroxy analogue of alanine), 2-hydroxy-3-methylbutyric acid (the alpha-hydroxy analogue of valine), 2- hydroxy-3-phenylpropionic acid, the hydroxy derivative of phenylalanine (F-OH), and alpha-hydroxy analogues of other natural and unnatural amino acids. Derivatives of noncanonical amino acids include the hydroxy derivatives of Ns-Alloc-L-lysine (AllocK-OH), p-lodo-phenylalanine (pIF-OH), Ne-((Prop-2-yn-1-yloxy)carbonyl)-L- lysine (AlkynK-OH), p-Azido-Phenylalanine (pAzF-OH), L-3-(2-Naphthyl)alanine (NapA-OH), Ne-tert-butoxycarbonyl-L-lysine (BocK-OH), N6-Carbobenzyloxy- L- lysine (CbzK-OH) and other lysine derivatives including PenK-OH, NorK-OH, ButK-OH and AcK-OH.
[0064] Alpha, alpha-disubstituted amino acids
[0065] Alpha, alpha-disubstituted amino acids are characterized by the substitution of two alkyl or aryl groups at the alpha carbon (Co) of the amino acid. This structural modification distinguishes them from the standard amino acids, which typically have a single substituent (the side chain) at the Co position. The presence of these additional substituents at the Co position imparts distinct steric and electronic properties, which can significantly influence the conformational and functional properties of peptides and proteins incorporatin these amino acids. In general, alpha-alpha-disubstituted amino acids can be derived from canonical or non- canonical amino acids. For example, alpha, alpha-disubstituted amino acids include, but are not limited to, 2-aminoisobutyric acid (2-AIB, the alpha-methylated analogue of alanine), alpha-methyl-L-valine, alpha-methyl-L-leucine, alpha-methyl- L-isoleucine, alpha-methyl-L-phenylalanine, alpha-methyl-L-tyrosine, alpha-methyl- L-tryptophan, alpha-methyl-L-methionine, alpha-methyl-L-serine, alpha-methyl-L- cysteine, alpha-methyl-L-aspartic acid, (S)-2-amino-3-(4-iodophenyl)-2- methylpropanoic acid, alpha, alpha-diphenylglycine, alpha, alpha-dipropylglycine and other alpha, alpha-disubstituted derivatives of natural and unnatural amino acids.
[0066] Beta-amino acids
[0067] Beta-amino acids are characterized by the substitution of various groups at the beta carbon (CP) of the amino acid. In general, beta-amino acids can be derived from canonical or non-canonical amino acids. For example, beta-amino acids include, but are not limited to 3-aminopropanoic acid (a beta-amino acid analog of alanine), L-beta-homoleucine, L-beta-isoleucine, L-beta-homovaline, L-beta-homoserine, L- beta-homotyrosine, L-beta-homomethionine, L-beta-homocystene, L-beta- homolysine, 3-amino-3-phenylpropanoic acid (P3F, the beta3-amino acid analog of phenylalanine), (S)-3-amino-3-(4- bromophenyl)propanoic acid ((S)P3pBrF, the the beta3-amino acid analog of p-bromo-L-phenylalanine), 3-amino-4-(4- bromophenyl)butanoic acid (P3pBrhF), 3-amino-2-((1-ethyl-1 H-imidazol-5- yl)methyl)propanoic acid (2NeH), (S)-6-acetamido-3-aminohexanoic acid ((S)-p3AcK), (S)-3-amino-6- (((benzyloxy)carbonyl)amino)hexanoic acid ((S)-p3CbzK), (S)-3- amino-3-(3- bromophenyl)propanoic acid ((S)P3mBrF), 2-benzyl-3-hydroxypropanoic acid (P2OH-F), 3-amino-3-phenylpropanoic acid (P3F), (S)-3-amino-3- (benzo[d][1 ,3]dioxol-5- yl)propanoic acid ((S)P3MDF) and other beta-amino acid derivatives of natural and unnatural amino acids. N-methylated amino acids
[0068] N-methylated amino acids are characterized by the substitution of one or more hydrogen atoms in the amino group (NH2) with methyl groups (CH3). In general, N- methylated amino acids can be derived from canonical or non-canonical amino acids. For example, N-methylated amino acids include, but are not limited to N- methyl-L-alanine (the N-methylated derivative of alanine), N-methyl-L-valine (the N- methylated derivative of valine), N-methyl-L-leucine (the N-methylated derivative of leucine), N-methyl-L-phenylalanine, and other derivatives of natural and unnatural amino acids.
[0069] Cyclized non-alpha amino acids
[0070] Cyclized non-alpha amino acids are characterized by the formation of a cyclic structure that includes the amino group and the carboxyl group or their side chains. Examples of this class of substrates include trans-4-hydroxy-L-proline, 4-methyl-L- proline, L-pipeloic acid, L-azetidine-2-carboxylic acid, indoline-2-carboxylic acid, 1 ,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, 3,4-dehydro-L-proline, 2,3- dihydroxy-L- pro line and more.
[0071] A variety of alternative substrate classes can be envisaged by the skilled person, which can be incorporated into a polypeptide acylation of a tRNA by an aaRS. Any such substrate can be used in the method of the invention to develop novel aaRS and substrate mutants.
[0072] All of the features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0073] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made to the Examples, which are not intended to limit the invention in any way. Examples
[0074] Comparative Example 1
[0075] Acylation of MmPylRS with AllocK
[0076] The following protocol for tRNA display is described by Dunkelmann et al., 2024.
[0077] 1 . E. coli BL21 (DE3) cells are provided with a DNA construct (stmRNAs) under control of T7 promoter. The stmRNA encodes a pyrrolysine tRNA sequence (fixed) and a pyrrolysine aaRS sequence (variable), cognate to the tRNA, and there is one stmRNA construct per cell, creating a physical compartmentalisation and a genotype-phenytope linkage of the sequences encoded within the stmRNA. The constructs are designed such that: a. The gene encoding the tRNA is split into two halves b. The two halves are separated by an intergenic loop, which gets cleaved by the cell’s endogenous machinery after transcription, yielding a native acceptor stem structure (process called maturation) c. Both halves contain 10 bp extensions in the anticodon region, and these regions base-pair with each other after transcription d. One of the halves contains a further extension to the 10 bp stem region, comprisingthe sequence of an aaRS variant which gets translated in the ribosome to yield an aaRS protein
[0078] 2. Cells are incubated with a target ncM of interest (or with no ncM, as a control), and expression of the stmRNA construct is induced with IPTG.
[0079] 3. Each cell translates the mRNA of its aaRS variant, which may or may not aminoacylate the tRNA to which the aaRS mRNA is linked with a canonical amino acid (AA) or an ncM.
[0080] 4. RNA is extracted from the population of cells.
[0081] 5. The RNA sample is subject to periodate oxidation, and subsequent deacylation of the AA or ncM from the tRNA. This has different implications on tRNA molecules, depending on whether they were acylated with an AA / ncAA or not: a . tRNAs that were acylated with the ncM of interest - the presence of the ncM protects the 3’ end of the tRNA, preventing oxidation. Subsequent deacylation renders a native 3’ end, which can serve as a starting point nucleotide extension by polymerases. b . tRNAs that were NOT acylated with the ncM of interest - these tRNAs are sensitive to oxidation, and unaffected by the deacylation process. The oxidation of the 3’ in these molecules prevents the mRNA 3’ end from being extended further by a polymerase.
[0082] 6. The oxidized and deacylated tRNA sample is incubated with a DNA oligonucleotide probe, which is complementary to the 3’ half of the tRNA, but which also creates a single-stranded DNA overhang beyond the 3’ end of the tRNA.
[0083] 7. The sample is subject to nucleotide extension with a DNA polymerase (Klenow fragment), which extends the 3’ end of the tRNA and polymerizes the complementary DNA strand to the DNA overhang. This nucleotide extension reaction contains biotinylated nucleotides, which results in the covalent linkage of biotin molecules to the tRNAs. Only tRNA molecules that were acylated with the ncM of interest by their aaRS, and hence not sensitive to oxidation, are extended in this step. Consequently, this step selectively labels stmRNA molecules that encode an active aaRS variant with biotin molecules, and does not label molecules that did not encode an active aaRS variant.
[0084] 8. The biotin-labelled tRNAs are selectively captured using streptavidin beads. stmRNAs which encoded a non-functional aaRS, and which were therefore not acylated with the ncM and subject to oxidation, do not contain any biotin as they were not extended, and are thus washed away in this step. Hence, this streptavidin capture step enriches for stmRNA molecules that encode the sequence of active aaRSs that aminoacylate the tRNA with an AA or ncM, and depletes the stmRNA molecules encoding inactive aaRSs. 9. The stmRNA mixture, which should now be enriched for stmRNA molecules which encode active aaRS sequences over stmRNA molecules inactive aaRS sequences, is subject to reverse transcription to generate cDNA.
[0085] 10. As a quality control, at this stage the number of cDNA molecules in the samples may be quantified by qPCR.
[0086] If the aaRS library contains aaRSs that are active towards the ncM of interest, it would be expected that the total molecule counts in this step would be higher in an experiment where cells were incubated with an ncM, than in the negative control. The relative abundance of cDNA molecules between the +ncM and - ncM samples at this stage are what we referto as cDNA molecule count enrichment. This enrichment may be neglibible or not detectable if a very low proportion of the aaRSs within the library are active. However, if the experiment is run with a single aaRS that is known to be active, rather than a library of active / inactive aaRSs, the molecule count enrichment is expected to be large provided that the various technical steps work as expected.
[0087] We applied the foregoing protocol to MmPylRS in standard commercial E. coli BL21 (DE3), referred to as strain 1 , using the ncAA AllocK. The procedure was the same as that described for Example 2, except that AllocK was used instead of BocK. The results are shown in Figure 3A. Allock is efficiently used by MmPylRS to acylate tRNAcuAPyl in the stmRNA described above, as expected.
[0088] We recapitulated the process described in steps 1-10 with another known substrate for the wild-type MmPylRS, BocK. MmPylRS is highly efficient for Bock, though even more efficient for AllocK. In the case of BocK, the cDNA molecule count enrichment values observed were around 30x, and the cDNA molecule counts in the presence of BocK were around 1 x 106, substantially lowerthan the values obtained from the analogous AllocK experiment. This shows that the level of enrichment achieved by tRNA display can vary across substrates, and it is expected that for less efficient substrates that are not as efficiently acylated onto tRNAs by their cognate aaRSs, it may be beneficial or necessary to improve the level of cDNA molecule enrichment and total stmRNA molecule recovery achieved by tRNA display, so as to maximize the chances of identifying even low activity aaRSs for elusive substrates.
[0089] Example 2 - use of mutant E. coli host
[0090] The BocK experiment of Example 1 was performed in standard commercial BL21 (DE3) (termed strain 1). When we switched to a different strain of BL21 (DE3) (termed strain 2), we observed a >5-fold increase in the cDNA molecule count enrichment, and were able to recapitulate values in line with the AllocK experiment, obtaining cDNA molecule count enrichment values over 100x, and cDNA molecule counts close to 1 x 108in the presence of the ncAA substrate. Similar relative improvements were observed when using strain with a separate substrate, BocAhx, which is a less efficient substrate for MmPylRS. In both cases, we observed an increase in the total number of stmRNA molecules recovered from cells, as measured in the cDNA molecule counts of the ’input’ samples, when using strain 2 relative to strain 1 . When repeated using strain 2, the experiment using the efficient substrate Allock showed a mild improvement over the same experiment conducted in strain 1 , showing an increase in cDNA molecule counts in the presence of AllocK of about 1 .7x (see Figure 3B), and an increase in the level of total stmRNA of 4.4x for the -AllocK sample and 2.8x for the +AllocK sample, as measured by the total cDNA count in the input samples.
[0091] In the experiment, BL21 (DE3) strain 1 and BL21 (DE3) strain 2 cells containing a plasmid encoding the stmRNA are incubated in the presence and absence of 2 mM BocK or 2 mM BocAhx in 2xTY medium, and expression of the stmRNA is induced by addition of 1 mM IPTG. After 40 minutes, cells are spun down and mRNA is extracted from cells using an Agencourt RNAdvance Cell v2 kit. The resulting mRNA sample is oxidized using 6 mM sodium periodate, and residual DNA is removed with DNAse I. The oxidized mRNA sample is purified with a Beckman Coulter RNAclean XP beads.
[0092] Oxidized and purified RNA samples are subject to bio-mREX. RNA samples are normalized to 400-500 ng / uL, and incubated with a Cy3-labelled DNA oligonucleotide complementary to the 3’ end of the stmRNA at 65 C for 5 minutes. Primer extension of the sample is then induced via addition of Klenow fragment (3’-5’ exo-) polymerase, and a dNTP mixture where dCTP is biotinylated (Biotin-11-dCTP), followed by extension for 6 minutes at 37 C.
[0093] The extended stmRNA samples are then captured in streptavidin magnetic beads (Dynabeads MyOne C1), incubating for 1 hour at 4 C. The samples are then reverse transcribed using Superscript IV Reverse Transcriptase (ThermoFisher), and the cDNA products are quantified by qPCR using primers that target the aaRS sequence within the stmRNA.
[0094] The relative molecule counts resulting from qPCR in this experiment reflect the method’s extent of enrichment for stmRNA molecules encoding aaRSs which are active towards BocK or BocAhx. For both substrates, the results of the same experiment in standard BL21 (DE3) strain 1 and 2 show that more active molecules are captured by the method when the strain background contains chromosomal mutations, including mutations in RNAse E, and the total number of stmRNA molecules isolated in the experiment, as determined by the cDNA molecule count of the ‘input’ samples, is also higher in strain 2 relative to strain 1 . See Figure 4.
[0095] When we sequenced strain 2, we found 30 mutations in the RNAse E gene, 5 of which were non-synonymous:
[0096] • S557A
[0097] • S591T
[0098] • T595A
[0099] • V660A
[0100] • T669A
[0101] Since RNAse E is involved in tRNA processing and degradation, as well as the degradation of cellular mRNA, we hypothesized that the mutations in RNAse E may affect the both the stability and the processing of the stmRNA transcripts, potentially improving the stability of the aaRS mRNA in cells or affect the processing of the stmRNA construct into mature, split tRNA halves. This could facilitate the assembly of intact tRNA molecules available for acylation and / or allow for higher intracellular levels of aaRS expression from the mRNA sequence, improving the sensitivity of the assay.
[0102] Example 3 - truncated RNase E gene
[0103] We repeated the experiment in the BL21 Star (DE3) strain, which contains a truncated version of RNAse E (me131 ) known to increase mRNA stability, side by side with BL21 (DE3) Strain 2. In this experiment, the cDNA molecule count enrichment values were comparable in both strains (60-80x), and so were the molecule counts in the presence of BocK (~1 x 108) and the stmRNA molecule counts in the input samples. See Figure 5.
[0104] Bio-mRex was conducted analogously to Example 2, in two different strains in parallel; one is BL21 (DE3) Strain 2, which carries mutations in RNAse E, and the other is BL21 (DE3) Star, which carries a truncation of RNAse E. The enrichment levels and cDNA molecule counts for acylated smtRNA molecules are comparable across both strains after bio-mREX. This provides further evidence that RNAse E is a key contributor to the isolation of stmRNA molecules through bio-mREX.
Claims
Claims1 . A method for expressing an acylated tRNA, comprising the steps of: a. expressing a gene encoding a tRNA molecule in a prokaryotic host cell, wherein the tRNA molecule is divided at the anticodon loop sequence to form two tRNA parts, and the anticodon stem in each tRNA part is extended to promote annealing of the tRNA parts to form a complete tRNA molecule; b. exposing said tRNA to conditions in the host cell such that the tRNA is acylated; and c. isolating said tRNA from the host cell, wherein the host cell does not comprise a wild-type RNase E gene.
2. A method accordingto claim 1 , wherein the gene encoding a tRNA molecule comprises more than one coding sequence, such that the gene is split into more than one part.
3. A method according to claim 3, wherein the gene is split into two parts.
4. A method accordingto claim 2 or claim 3, wherein the gene is split at the anticodon loop sequence.
5. A method accordingto claim 4, wherein the anticodon loop sequence is extended by between 6 and 12 nucleotides; preferably wherein the anticodon loop sequence is extended by 10 nucleotides.
6. A method accordingto any one of claims 2 to 5, wherein two or more coding sequences are expressed in cis as a single transcript, and an intervening sequence connects the coding sequences.
7. A method according to any preceding claim, wherein a sequence encoding an aaRS is fused to the gene encoding the tRNA.
8. A method accordingto any preceding claim, wherein the cell is supplied with a noncanonical monomer to acylate the tRNA.
9. A method accordingto any preceding claim, wherein the tRNA is isolated from the host cell and subjected to analysis for acylation.
10. A method according to claim 9, wherein a sequence encoding an aaRS is fused to the gene encoding the tRNA and analysis for acylation comprises the steps of: a. subjecting the RNA sample comprising the tRNA-aaRS fusion to an oxidation reaction, wherein the oxidation is dependent on the absence of an acylating group on the tRNA, such that acylated tRNA moelculess are not oxidised; b. Subjectingthe RNA sample to conditions which are conducive to the deacylation of the ncM form the tRNA portion of the tRNA-aaRS fusion molecule, such that only non-oxidised tRNA are deacylated to render a native 3’ end, which can serve for nucleotide extension by polymerases; c. incubating the deacylated tRNA-aaRS fusion with a DNA oligonucleotide probe, which is complementary to the 3’ half of the tRNA portion, but which also creates a single-stranded DNA overhang beyond the 3’ end of the tRNA portion; and d. extending the probe with a DNA polymerase (Klenow fragment), which extends the 3’ end of the tRNA portion and polymerizes the complementary DNA strand to the DNA overhang with biotinylated nucleotides, such that acylated tRNA molecules are biotinylated.11 .A method accordingto any preceding claim, wherein the RNase E gene in the host cell is mutated, truncated or deleted.
12. A method accordingto claim 11 , wherein the RNase E gene is truncated, preferably wherein the truncation is the me131 truncation.
13. A method accordingto claim 11 , wherein the RNase gene is mutated at one or more of positions S557, S591 , T595, V660 and T669.
14. A method according to claim 12, wherein the mutation is selected from the group consisting of S557A, S591T, T596A, V660A and T669A.
15. A method accordingto any preceding claim, wherein the host cell is E. coli.16.A method accordingto claim 15, wherein the host cell is E. coli BL21 (DE3).
Citation Information
Patent Citations
Synthetic genome
WO2020229592A1