Compositions and methods for site-specific incorporation of unnatural amino acids into proteins

WO2026178241A1PCT designated stage Publication Date: 2026-08-27DNA TWOPOINTO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/015848
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-09-29
Filing Date
2026-02-19
Publication Date
2026-08-27

Smart Images

  • Figure IMGF000068_0001
    Figure IMGF000068_0001
  • Figure IMGF000004_0001_TABLE
    Figure IMGF000004_0001_TABLE
  • Figure IMGF000006_0001_TABLE
    Figure IMGF000006_0001_TABLE
Patent Text Reader

Abstract

The present disclosure relates to compositions and methods related to orthogonal translation systems for the site-specific incorporation of unnatural amino acids into polypeptides, for example, in in vivo systems such as in a mammalian host cell. The disclosure provides orthogonal pairs of aminoacyl-tRNA synthetases and tRNA molecules that can incorporate unnatural amino acids into proteins in mammalian cells, and in particular, lysine derivative unnatural amino acids that have side group functionalities that facilitate biocompatible conjugation or other types of modification. The disclosure further provides optimized promoter configurations for the in vivo expression of the orthogonal tRNA component of the orthogonal translation systems, improving the productivity of said systems for the production of polypeptides containing unnatural amino acids.
Need to check novelty before this filing date? Find Prior Art

Description

COMPOSITIONSAND METHODS FOR SITE-SPECIFIC INCORPORATION OF UNNATURAL AMINO ACIDS INTO PROTEINSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is related to United States Provisional Patent Applications having Serial Nos. 63 / 761,127, filed February 20, 2025, and 63 / 889,738, filed September 29, 2025, the disclosures of which are each incorporated herein by reference in their entirety for all purposes.TECHNICAL FIELD

[0002] The present disclosure relates to the use of orthogonal tRNAs, orthogonal aminoacyl-tRNA synthetases, and pairs thereof, for the incorporation of unnatural amino acids into polypeptides, for example, in an in vivo system such as in a mammalian cellular host cell.BACKGROUND

[0003] Site-specific incorporation of an unnatural amino acid (UAA) into a protein can be accomplished in host cells using a tRNA that recognizes a stop codon (the "suppressor tRNA") in an open reading frame and a suitable aminoacyl-tRNA synthetase that aminoacylates the tRNA with the unnatural amino acid. The suppressor tRNA can deliver the UAA to the translating ribosome, which "reads" the stop codon and incorporates the UAA into the growing polypeptide chain.

[0004] Systems for the site-specific and cost-effective incorporation of unnatural amino acids into proteins are of significant interest. Previously demonstrated systems which rely on tRNA synthetase-tRNA pairs that act orthogonally to host tRNA charging systems are limited by their inability to accept many potentially useful UAAs. In addition, such systems often show poor protein yields due to inefficient incorporation of the UAA and / or toxicity and / or impairment of cell growth due to the introduction of the orthogonal components.

[0005] Development of new synthetase-tRNA pairs to improve the efficiency of desired UAA incorporation can not only broaden the repertoire of chemical moieties that can be used to functionalize proteins, but also improve production yields of conjugable proteins.

[0006] Most modification reactions currently used in the art for the selective modification of proteins involve covalent bond formation between nucleophilic and electrophilic reaction partners that target naturally occurring nucleophilic residues in the protein amino acid side chains, e.g., the reaction of a-halo ketones with histidine or cysteine side chains. Selectivity in these cases isdetermined by the number and accessibility of the nucleophilic residues in the protein.Unfortunately, naturally occurring proteins frequently contain poorly positioned (e.g., inaccessible) reaction sites or multiple reaction targets (e.g., lysine, histidine and cysteine residues), resulting in poor selectivity in the modification reactions, making highly targeted protein modification by nucleophilic / electrophilic reagents difficult. Furthermore, the sites of modification are typically limited to the naturally occurring nucleophilic side chains of lysine, histidine or cysteine.Modification at other sites is difficult or impossible.

[0007] Stop codons are used by orthogonal tRNAUAAmolecules in engineered orthogonal translation systems to incorporate UAA into polypeptide chains in vivo. However, in natural endogenous systems, stop codons are also recognized by translation termination or release factors, which terminate translation and release the polypeptide chain. Efficient incorporation of UAAs by this mechanism thus requires the presence of enough of the amino-acylated suppressor tRNA to effectively compete with the endogenous levels of release factors. Expression of the suppressor tRNA that recognizes the stop codon is often limiting for the production of proteins comprising UAAs in UAA-incorporating cell lines. Cell lines with more copies of the sequence encoding the suppressor tRNA operably linked to a promoter (the "suppressor tRNA-expressing DNA") produce higher levels of UAA-incorporating proteins (see, e.g., Table 1 of Roy et. al., 2020 "Development of a high yielding expression platform for the introduction of non-natural amino acids in protein sequences" MABS 12, el684749, where hundreds of copies of the suppressor tRNA-expressing DNA were integrated). However, incorporation of many identical copies of the suppressor tRNA-expressing DNA has been found to result in genetic instability, as the host cell line recombines out concatemers comprising many identical copies of a sequence.

[0008] What is needed in the art are new strategies for incorporation of unnatural amino acids into proteins for the purpose of modifying and studying protein structure and function. What is needed in the art are new orthogonal translation systems that have the ability to co-translationally incorporate in vivo a desired UAA into the growing polypeptide chain. In these systems, UAA that have functional groups that permit modification, such as by covalent conjugation or coupling with desired moieties under physiological conditions, i.e., under biocompatible conditions, are particularly desirable. Novel chemistries for highly specific protein modification can find a wide variety of applications in the development of therapeutics, for example antibody therapeutics, and for the study of protein structure and function. There is also a need to develop orthogonaltranslation components that incorporate unnatural amino acids with novel chemical properties that allow the amino acid to serve as a target for specific modification to the exclusion of cross reactions or side reactions with other sites in the proteins.

[0009] What is needed in the art are systems and methods to reduce the amount of repetitive sequence required to achieve adequate suppressor tRNA expression in a host production cell line, thereby improving the genetic stability of the cell line and producing high levels of UAA-incorporated protein.

[0010] The present disclosure fulfills these and other needs, as will be apparent upon review of the following disclosure.SUMMARY OF THE DISCLOSURE

[0011] The present disclosure provides compositions and methods for incorporating unnatural amino acids (UAA) into a growing polypeptide chain in response to a selector codon, e.g., an amber stop codon. In some aspects, these compositions and methods are used in vivo (e.g., in a host cell). These compositions include pairs of orthogonal-tRNAs (O-tRNAs) and orthogonal aminoacyl-tRNA synthetases (O-RSs) that do not interact with, or do not impair or only minimally impact, the host cell tRNA or RS molecules. That is to say, the O-tRNA is not charged (or not charged to a significant or detectable level) with an amino acid (natural or unnatural) by an endogenous host cell aminoacyl-tRNA synthetase. Similarly, the O-RSs provided by the disclosure do not charge any endogenous tRNA with an amino acid (natural or unnatural) to a significant or detectable level. These novel compositions permit the production of large quantities of proteins having translationally incorporated unnatural amino acids. These proteins incorporating unnatural amino acids find a wide variety of uses, including therapeutics, for example, antibody therapeutics, and in biological research. In some aspects, the UAA comprises a functional group capable of biocompatible conjugation.

[0012] In some aspects, the present disclosure provides translation systems. These systems comprise an orthogonal aminoacyl-tRNA synthetase (O-RS), an orthogonal tRNA (O-tRNA), and an unnatural amino acid, where the first O-RS preferentially aminoacylates the first O-tRNA with the first unnatural amino acid. The unnatural amino acid is selected from unnatural amino acids that contain reactive side chain groups that are reactive at physiological conditions, i.e., are modifiable under conditions that preserve the primary, secondary and tertiary structure of the protein beingmodified. As a result, these proteins containing one or more UAA can be modified, e.g., conjugated with other moieties, in a manner that preserves the biological activity of the protein being modified. Examples of preferred reactive UAA that can be incorporated into proteins are provided in the chemical structures shown in FIG. 6.

[0013] Examples of O-RS and O-tRNA pairs finding use in the translation systems described herein for the incorporation of particular UAA are provided by the present disclosure; e.g., see Tables 2, 3, 4, 5, 6 and 7. Amino acid and nucleotide sequences of the orthogonal components are also provided (FIG. 15). When amino acid sequences of polypeptides (e.g., O-RS polypeptides) are provided by the present disclosure, it is understood that the disclosure will also encompass any nucleotide sequence that encodes that amino acid sequence, for example, where codon optimized nucleotide sequences find particular use.

[0014] The O-RS molecules finding use with the methods described herein can be a bacterial pyrrolysyl-RS, an Archaeal pyrrolysyl-RS, or variants derived from bacterial pyrrolysyl-RS, or derived from an Archaeal pyrrolysyl-RS. In some aspects, the bacterial or Archaeal orthogonal pyrrolysyl-RS is a native RS that also has the ability to charge a cognate tRNAPylwith a UAA in addition to pyrrolysine that occurs in the native bacterial or Archaeal context.

[0015] In some aspects, the orthogonal translation systems as described herein are crossspecies hybrid systems where the O-RS and O-tRNA are derived from different species or different Domains / Kingdoms, which can be a bacterial species or an Archaea species. For example, the disclosure provides orthogonal O-RS / O-tRNA pairs that are derived / paired as follows:

[0016] In some aspects, O-RS molecules of the disclosure can comprise directed amino acid substitutions that improve the activity of the RS molecule, such that charging of a O-tRNA is improved as compared to the charging by a native RS sequence, and as a result, improved production of protein comprising the UAA is observed as compared to an unmodified system.

[0017] Variant O-RS are also encompassed by the present disclosure, where the variant O-RS can comprise:(i) one or any number of conservative amino acid substitutions;(ii) one or more amino acid substitutions that may not be conservative substitutions, where the substitution(s) result in either an O-RS with the same or similar O-tRNAUAAcharging activity as the native RS sequence, or improved 0- tRNAUAAcharging activity compared to the native RS sequence.

[0018] In various aspects, the O-RS variants have at least 80%, or 90%, or 95%, or 98%, or 99% amino acid identity with the amino acid sequence of the parent O-RS molecule. In some aspects, the variant O-RS has O-tRNAUAAcharging activity that is at least equal to or better than the UAA charging activity of the native RS amino acid sequence.

[0019] Similarly, examples of O-tRNA molecules finding use in the translation systems described herein for the incorporation of UAA are also provided Tables 2, 3, 4, 5, 6 and 7, and sequences provided in FIG. 15. O-tRNA molecules finding use with the methods described herein can be bacterial O-tRNAUAAor Archaeal O-tRNAUAA, such as a tRNAPyl, or derived from either of those sources.

[0020] In some aspects, the translation system incorporates a second orthogonal pair (that is, a second O-RS and a second O-tRNA) that utilizes a second unnatural amino acid, so that the system is now able to incorporate at least two different unnatural amino acids at different selected sites in a polypeptide. In this dual system, the second O-RS aminoacylates the second O-tRNA with the second unnatural amino acid that is different from the first unnatural amino acid, and the second O-tRNA recognizes a selector codon that is different from the selector codon recognized by the first O-tRNA.

[0021] In some embodiments, the translation system resides in a host cell (and optionally where the translation system includes the host cell). The host cell used is not particularly limited, as long as the O-RS and O-tRNA retain their orthogonality in the host cell environment. In some aspects, the host cell is preferably a mammalian cell, such as a human cell or a rodent cell.

[0022] In some aspects, the translation system further comprises a nucleic acid encoding a protein of interest, where the nucleic acid open reading frame has at least one selector codon that is recognized by the O-tRNA, and as a result, will be recognized by an O-tRNA that results in the incorporation of a UAA into the polypeptide.

[0023] In some aspects, the disclosure provides translation systems, as described above, that recite exemplary but not limiting examples of O-RS and O-tRNA molecules for the incorporation of particular unnatural lysine-derivative UAA, and in particular, lysine-derivative UAA comprising afunctional group capable of biocompatible conjugation. This description of exemplary but not limiting O-RS / O-tRNA pairs is provided in the following Tables:

[0024] Structures of the UAA molecules listed above are provided in FIG. 6. Amino acid and nucleotide sequences of the O-RS and O-tRNA molecules listed in the Tables are provided in FIG.15.

[0025] The present disclosure also provides methods for producing proteins having one or more unnatural amino acids at selected positions. These methods utilize the translation system components described above. Generally, these methods start with the step of providing a translation system comprising: (i) an unnatural amino acid that comprises a reactive side chain moiety that is reactive a physiological conditions suitable for bioconjugation reactions; (ii) a first orthogonal aminoacyl-tRNA synthetase (O-RS); (iii) a first orthogonal tRNA (O-tRNA), for example, a bacterial O-tRNA, wherein the O-RS preferentially aminoacylates the O-tRNA with the unnatural amino acid; and, (iv) a nucleic acid encoding the protein, where the nucleic acid comprises at least one selector codon that is recognized by the first O-tRNA. The method then incorporates the unnatural amino acid at the selected position in the protein during translation of the protein in response to the selector codon, thereby producing the protein comprising the unnatural amino acid at the selected position. The method further comprises delivering the translation system components to a suitable host cell, for example, a mammalian host cell, and culturing the host cell in the presence of the lysine derivative UAA to produce the polypeptide of interest comprising the UAA. Any of the UAA molecules described in FIG. 6 find use with the methods described herein.

[0026] In one aspect, systems and methods are provided for producing a protein with a UAA incorporated at a specific site, where expression of the translation system components isoptimized. In one such aspect, the systems and methods reduce the amount of repetitive sequence required to achieve adequate suppressor tRNA expression. In one aspect, such systems and methods include: (i) the use of highly active promoters operably linked to the suppressor tRNA-coding sequence, thereby increasing suppressor tRNA expression from each gene copy and reducing the number of genomic copies required; (ii) the use of multiple different promoters operably linked to the suppressor tRNA-coding sequence, thereby reducing the overall repetitive nature of the suppressor tRNA-expressing DNA and thus also reducing its instability; and (iii) the provision of the suppressor tRNA-expressing DNA on transposons and their introduction into the cell by transposition, thereby ensuring that there are multiple independently integrated copies of the suppressor tRNA-expressing DNA, which are much more stable than long concatemers.

[0027] In one aspect, a polynucleotide is provided, the polynucleotide comprising a pig 7sk promoter operably linked to a DNA sequence encoding a heterologous tRNA. In one aspect, the nucleotide sequence of the pig 7sk promoter is SEQ ID NO:16. In one aspect, the heterologous tRNA comprises an anticodon that recognizes a stop codon. In one aspect, the stop codon is an amber codon (UAG). In one aspect, the heterologous tRNA is amino-acylated with a UAA. In one aspect, the heterologous tRNA is capable of being charged (i.e., amino-acylated) with pyrrolysine. In one aspect, the UAA is a non-naturally occurring lysine derivative, e.g., a non-naturally occurring lysine derivative as shown in FIG. 6. In one aspect, the DNA sequence encoding the heterologous tRNA comprises SEQ ID NO: 28. In one aspect, the polynucleotide comprises a transposon, the transposon further comprising left and right transposon ends such that the promoter and the DNA sequence encoding the heterologous tRNA are transposable by a corresponding transposase. In one aspect, the polynucleotide further comprises one or more additional promoters selected from a second pig 7sk promoter, a human U68cc promoter, a water buffalo U6 promoter, a human 7sk promoter, and a mouse Hl promoter, and combinations thereof. In one aspect, each additional promoter is operably linked to a separate copy of the DNA sequence encoding a heterologous tRNA. In one aspect, the nucleotide sequence of the one or more additional promoters is selected from SEQ ID NOs: 11, 14, 16, 18, 20 and 31, and any combinations thereof.

[0028] In another aspect, a method is provided for producing a protein with a UAA incorporated at a specific site, the method comprising introducing the polynucleotide encoding the protein into a mammalian cell, e.g., a CHO-K1 cell. In another aspect, a mammalian cell such as a CHO-K1 cell is provided, the mammalian cell comprising the polynucleotide.

[0029] In another aspect, a system is provided for site-specific incorporation of a non-canonical amino acid into an antibody in a mammalian host cell, e.g., a CHO-K1 cell, the system comprising: a transposon comprising multiple copies of DNA encoding pyrrolysine tRNA, each copy operably linked to a pol III promoter active in mammalian cell, wherein: (i) at least one of the pol III promoters is a pig 7sk promoter; (ii) at least one of the pol III promoters is not a pig 7sk promoter; and (iii) where there are at least two copies of DNA encoding pyrrolysine tRNA that are each operably linked to a pol III promoter that is not a pig 7sk promoter, the pol III promoters that are not a pig 7sk promoter are different from each other; a second transposon comprising an open reading frame encoding a corresponding aminoacyl tRNA synthetase operably linked to regulatory elements such that the aminoacyl tRNA synthetase is expressible in the mammalian cells, e.g., the CHO-K1 cell; and a third transposon comprising open reading frames encoding the heavy and light chains of an antibody, wherein the open reading frame encoding one chain comprises an amber stop codon (UAG) preventing expression of a full antibody unless the amber stop codon is read by the pyrrolysine tRNA to incorporate the non-canonical, e.g., unnatural, amino acid. In one aspect of the system, the first, second, and third transposons are separate. In one aspect of the system, the first, second, and third transposons are one transposon. In one aspect of the system, the first and second, the first and third, or the second and third transposons are one transposon. In one aspect of the system, any combination of the first, second, and third transposons are transposable by the same transposase or by different transposases.

[0030] In another aspect, a mammalian cell, e.g., a CHO-K1 cell, is provided whose genome comprises the system.

[0031] In another aspect, a method for site-specific incorporation of a non-canonical amino acid, more specifically an unnatural amino acid, into an antibody in a mammalian cell, such as a CHO-K1 cell, is provided, the method comprising delivering the system into the mammalian cell. The unnatural amino acid can be any unnatural amino acid. In one aspect, the unnatural amino acid is lysine derivative amino acid, for example, any of the lysine derivative amino acids as shown in FIG. 6.BRIEF DESCRIPTION OF THE FIGURES

[0032] The accompanying figures, which are incorporated in and constitute a part of the specification, are used merely to illustrate various example aspects.

[0033] FIG. 1 shows the expression of green fluorescent protein ("GFP") in various UAA-incorporating pools.

[0034] FIG. 2 shows sodium dodecyl sulfate polyacrylamide gel electrophoresis ("SDS-PAGE") analysis of trastuzumab production.

[0035] FIGS.3A and 3Bshow N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine ("Azido-Lys") incorporation in an antibody heavy chain. Panel A shows total protein stained with coomassie blue. Panel B shows fluorescence of the protein.

[0036] FIG. 4 shows expression of trastuzumab antibody in UAA-incorporating cells with different feed strategies.

[0037] FIG. 5 shows trastuzumab antibody titers from an optimized vector in UAA-incorporating cells.

[0038] FIG. 6 provides a table showing various unnatural amino acid structures finding use with the orthogonal translation systems as described herein. Also shown are the corresponding chemical names and common abbreviated names of the UAA.

[0039] FIG. 7 provides a bar graph showing trastuzumab antibody titers at day 7 following the culture of HEK293 cells in the presence of either UAA Azc-Lys or 2AzZ-Lys. The host cells were alternatively transfected with one of several synthetase and cognate tRNA pairs derived from either Archaeal or bacterial systems.

[0040] FIG. 8 provides a bar graph addressing synthetase substrate specificity and the effects of various RS point mutations. The bar graph shows trastuzumab antibody titers at day 4 following the culture of HEK293 cells alternatively in the presence of a panel of UAA, which were Azc-Lys, 2AzZ-Lys, Nbb-Lys, SpHD-Lys, and TCO-Lys. The host cells were alternatively transfected with one of several RS and tRNA pairs derived from bacterial HHW30035.1 pyrrolysyl-synthetase, and all of which included the cognate HHW30035.1 tRNAPyl. The synthetases tested were the wildtype RS sequence, as well as eight different point mutant variants. One set of HEK293 host cells were transfected with Archae M. mazei mut pyrrolysyl-synthetase (Y306A) and the Mm tRNAPyl(T25C).

[0041] FIG. 9 provides a bar graph showing the results of testing of various bacterial pyrrolysyl-RS and their cognate tRNAPylpartners identified as homologues of bacterial HHW30035.1 pyrrolysyl-RS. Each of these homologues, along with Y127A mutant variants, were tested for their ability to incorporate either UAA Azc-Lys or 2AzZ-Lys. The bar graph shows trastuzumab antibody titers at day 7 following the culture of HEK293 cells transfected with the translation components.

[0042] FIGS. 10A and 10B show the results of activity assays for orthogonal translation systems that are hybrid systems comprising both bacterial and Archaeal pyrrolysyl-RS and tRNAPylcomponents. Results are measured as antibody titer at four days post transfection. These hybrid systems were tested for their ability to incorporate various UAA. FIG. 10A shows the results of testing with five different Archaeal / bacterial combinations for the ability to incorporate 2AzZ-Lys UAA. These combinations were:i) wild type Archaea M. mazei pyrrolysyl-RS with M. mazei tRNAPyl(T25C);ii) double mutant M. mazei pyrrolysyl-RS (Y306A; Y384F) with M. mazei tRNAPyl(T25C);iii) double mutant M. mazei pyrrolysyl-RS (Y306A; Y384F) with Syntrophomonadaceae bacterial HHW30035 tRNAPyl;iv) bacterial HHW30035 pyrrolysyl-RS mutant enzyme (Y127A) with the natural cognate HHW30035 tRNAPyl;v) bacterial HHW30035 pyrrolysyl-RS mutant enzyme (Y127A) with the Archaea M. mazei tRNAPyl(T25C).

[0043] FIG. 10B shows the results of activity assays for hybrid orthogonal translation systems comprising both bacterial and Archaeal pyrrolysyl-RS and tRNAPylcomponents, using a panel of UAA. The UAA panel used included: Azc-Lys, 2AzZ-Lys, SpHD-Lys, Nrb-Lys and TCO-Lys. The two hybrid translation systems tested were:(i) the Archaea double mutant M. mazei pyrrolysyl-RS (Y306A; Y384F) with the Syntrophomonadaceae bacterial HHW30035 tRNAPyl; and(ii) the Syntrophomonadaceae bacterial HHW30035 pyrrolysyl-RS mutant enzyme (Y127A) with the M. mazei tRNAPyl(T25C).

[0044] FIGS. 11A and 11B show the results of experiments to assess the productivity of the O-RS and O-tRNA pairs in stably-integrated cell lines, and further for the ability of these RS / tRNA systems to incorporate two different UAA, namely 2AzZ-Lys and SpHD-Lys. UAA translation activity was measured by titer of the trastuzumab reporter antibody. Three different translation systems were tested in the stably transfected HEK293 cell lines. These RS / tRNA systems were:(i) the Archaea mutant M. mazei pyrrolysyl-RS (Y306A) with the M. mazei tRNAPyl(T25C), in the presence of UAA 2AzZ-Lys;(ii) the Archaea mutant M. mazei pyrrolysyl-RS (Y306A) with the bacterial Syntrophomonadaceae HHW30035 tRNAPyl, in the presence of UAA 2AzZ-Lys; and(iii) the Archaea mutant M. mazei pyrrolysyl-RS (Y306A) with the M. mazei tRNAPyl(T25C) (same as (i)), in the presence of UAA SpHD-Lys.

[0045] FIG. 11A shows the antibody titer results from these three stably integrated O-RS / O-tRNA systems using 2AzZ-Lys and SpHD-Lys. FIG. 11B shows the results of a cell viability assay at multiday timepoints from the same stably-transfected cultures and UAA treatments as described in FIG. 11A.

[0046] FIG. 12 shows results for antibody production in transiently transfected HEK293 cells using mutant synthetases to incorporate the UAA 2AzZ-Lys into trastuzumab.

[0047] FIG. 13 provides the results of experiments addressing orthogonal RS / tRNA specificity and the effects of various RS point mutations, as assessed in transiently transfected CHO-K1 cells. This experiment also addresses the translational activity of a hybrid Archaeal / bacterial translation system pair. The bar graph shows trastuzumab antibody titers at day 7 following the transfection of CHO-K1 cells, alternatively in the presence of Azc-Lys or 2AzZ-Lys. The host cells were transiently transfected with the following RS and tRNA pairs:i) mut Archaea M. mazei pyrrolysyl-RS(Y306A) with M. mazei tRNAPyl(T25C); ii) mut bacterial HHW30035 pyrrolysyl-RS mutant enzyme (Y127A) with the natural cognate HHW30035 tRNAPyl;iii) wild type M. mazei pyrrolysyl-RS with M. mazei tRNAPyl(T25C);iv) mut Archaea M. mazei pyrrolysyl-RS(Y306A) with bacterial HHW30035 tRNAPyl.

[0048] FIG. 14 provides a table of amino acid and polynucleotide sequences referenced in the present specification.

[0049] FIG. 15 provides a table of amino acid and polynucleotide sequences referenced in the present specification. Mutant / substituted positions are shown in bold underline. Other key positions are shown in bold.

[0050] FIG. 16 provides a table describing bacterial synthetase molecules that find use in the orthogonal translation systems and methods described in the present disclosure for the incorporation of unnatural amino acids into polypeptides in mammalian host cells.

[0051] FIG. 17 provides the results of experiments addressing orthogonal RS / tRNA specificity for various lysine-derivative UAA. The bar graph shows trastuzumab antibody titers at day four following transfection of HEK293 cells cultured alternatively in the presence of 0.05 mM UAA selected from a panel of five UAA, which were 2AzZ-Lys, SCO-Lys, cyclopropene-Lys, endo-BCN-Lys, or H-L-Photo-Lys. The host cells were alternatively transfected with one of three RS / tRNA pairs. The synthetases tested were the Archaeal M. mazei pyrrolysyl-synthetase double mutant (Y306A, Y384F) and the bacterial HHW30035 (Y127A) mutant. Cells transfected with the M. mazei mutant synthetase were co-transfected with either Mm tRNAPyl(T25C) or with bacterial HHW30035 tRNAPyl. Cells transfected with the bacterial HHW30035 pyrrolysyl-RS mutant (Y127A) were paired with the HHW30035 tRNAPyl.DEFINITIONS

[0052] Unless defined otherwise herein, all technical and scientific terms have the same meaning as commonly understood by one of ordinary skill in the relevant art. Singleton, et al., Dictionary of Microbiology and Molecular Biology, 2nd Ed., John Wiley and Sons, New York (1994), and Hale & Marham, The Harper Collins Dictionary of Biology, Harper Perennial, NY, 1991, provide one of skill with a general dictionary of many of the terms used herein. Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively. The terms defined below are more fully understood by reference to the specification as a whole.

[0053] As used in this specification and the appended claims, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polypeptide" includes, as a practical matter, many molecules of that polypeptide.

[0054] Where a range of values is recited, each intervening integer value, and each fraction thereof, between the recited upper and lower limits of that range is also specifically disclosed, along with each subrange between such values. The upper and lower limits of any range can independently be included in or excluded from the range, and each range where either, neither, or both limits are included is also encompassed. Where a value being discussed has inherent limits, for example where a component can be present at a concentration of from 0 to 100%, or where the pH of an aqueous solution can range from 1 to 14, those inherent limits are specifically disclosed. Where a value is explicitly recited, values that are "about" (that is, within ±10%) thesame quantity or amount as the recited value are also within the scope. Where a combination is disclosed, each sub-combination of the elements of that combination is also specifically disclosed. Conversely, where different elements or groups of elements are individually disclosed, combinations thereof are also disclosed. Where any element is disclosed as having a plurality of alternatives, examples in which each alternative is excluded singly or in any combination with the other alternatives are also hereby disclosed; more than one element can have such exclusions, and all combinations of elements having such exclusions are hereby disclosed.

[0055] As used herein, the term "about" when immediately preceding a numerical value can mean a range of plus or minus 10% of that value, e.g., "about 50" means 45 to 55, unless the context of the disclosure indicates otherwise, or is inconsistent with such an interpretation. For example, in a list of numerical values such as "about 49, about 50, about 55, "about 50" means a range extending to less than half the interval(s) between the preceding and subsequent values, e.g., more than 49.5 to less than 52.5. Furthermore, the phrases "less than about" a defined value or "greater than about" a defined value should be understood in view of the definition of the term "about" provided herein. In some aspects, use of the term "about" indicates that a particular value is intended to include a range of values that represents the limits of accuracy of the instrumentation that was used to obtain the value, or within the range of reproducibility of the value that is being measured by a particular instrument.

[0056] Terms such as "connected," "attached," "linked," "coupled," "conjugated" or the like are used interchangeably herein and encompass direct as well as indirect connection, attachment, linkage, or conjugation unless the context clearly dictates otherwise. Unless otherwise stated, there is no restriction on the particular molecules or physical structures or how the particular connection or coupling is made , other than one of ordinary skill in the art will recognize a number of functionally-equivalent alternatives that can be employed to achieve the desired result.

[0057] In some contexts, the term "operably linked" or "operably coupled" or similar expressions refer to functional linkages between two or more molecules, for example, polypeptides or polynucleotides, such that one moiety modifies the behavior of the other or results in some desired result that arises from the union of the two or more parts. For example, a first polynucleotide comprising a nucleic acid expression control sequence (such as one or more of a promoter, IRES sequence, enhancer, or array of transcription factor binding sites) and a second polynucleotide, for example an open reading frame, are operably linked if the first polynucleotideaffects transcription and / or translation of the second polynucleotide. Similarly, a first amino acid sequence comprising a trafficking signal, e.g., a secretion signal or a subcellular localization signal, and a second amino acid sequence are operably linked if the first amino acid sequence causes the second amino acid sequence to be appropriately processed, e.g., appropriately secreted or localized to the target subcellular location.

[0058] Compositions or methods "comprising" one or more recited element(s) may include other elements / components not specifically recited. For example, a composition comprising a polynucleotide most frequently comprises the polynucleotide as well as a suitable aqueous solvent or buffer solution in which the polynucleotide is dissolved or suspended. Also for example, a polynucleotide comprising an open reading frame can further include other nucleotide subsequences in addition to the open reading frame, for example, promoter elements and other types of nucleic acid regulatory elements that control the expression of the open reading frame.

[0059] The terms "nucleoside" and "nucleotide" include those moieties that contain not only the known purine and pyrimidine bases, but also other heterocyclic bases that have been modified. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, or other heterocycles. Modified nucleosides or nucleotides can also include modifications on the sugar moiety, for example, where one or more of the hydroxyl groups are replaced with halogen, aliphatic groups, or are functionalized as ethers, amines, or the like. The term "nucleotidic unit" is intended to encompass nucleosides and nucleotides.

[0060] The terms "polynucleotide," "oligonucleotide," "nucleic acid," "nucleic acid molecule," and "gene" are used interchangeably to refer to a polymeric form of nucleotides of any length that can have any of a variety of functions or activities, and may comprise ribonucleotides, deoxyribonucleotides, analogs thereof, or mixtures thereof. The length of a polynucleotide (i.e., the number of nucleotides in a contiguous polymer chain of nucleotides) is not limited in any manner. Typically, the term "oligonucleotide" refers to nucleotide sequences of 2 to 20 nucleotides in length. Nucleotide sequences of open reading frames, genes, cDNA, entire plasmids, and genomic / chromosomal sequences can be hundreds, thousands or many hundreds of thousands of nucleotides in length. A polynucleotide can be artificial / synthetic, engineered, recombinant, naturally occurring, or derived from natural sequences. Depending on context, the length of a polynucleotide is sometimes expressed in "base pairs," implying that the nucleic acid molecule is double-stranded, or would be double stranded in its in vivo or native context.

[0061] As used herein, the expressions "DNA sequence," "RNA sequence," "nucleotide sequence," "polynucleotide sequence" or similar expressions refer to the defined order of nucleotides in a contiguous sequence of nucleotides, i.e., the "primary sequence" or "primary structure" of the molecule. Thus, these terms include triple-, double-, and single-stranded DNA, as well as triple-, double-, and single-stranded RNA. The terms also encompass modified, for example by alkylation and / or by capping, and unmodified forms of the polynucleotide. More particularly, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), including tRNA, rRNA, hRNA, siRNA, and mRNA, whether spliced or unspliced, any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing nonnucleotidic backbones, for example, polyamide (for example, peptide nucleic acids ("PNAs")) and polymorpholino (commercially available from the Anti-Virals, Inc., Corvallis, Oreg., as Neugene) polymers, and other synthetic sequence-specific nucleic acid polymers providing that the polymers contain nucleobases in a configuration that allows for base pairing and base stacking, such as is found in DNA and RNA. There is no intended distinction in length between the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule," and these terms are used interchangeably herein. These terms include, for example, 3'-deoxy-2', 5'-DNA, oligodeoxyribonucleotide N3' P5' phosphoramidates, 2'-O-alkyl-substituted RNA, double- and single-stranded DNA, as well as double- and single-stranded RNA, and hybrids thereof, including for example hybrids between DNA and RNA or between PNAs and DNA or RNA, and also include known types of modifications, for example, labels, alkylation, "caps," substitution of one or more of the nucleotides with an analog, internucleotide modifications such as, for example, those with uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, or the like) with negatively charged linkages (for example, phosphorothioates, phosphorod ithioates, or the like), and with positively charged linkages (for example, aminoalkylphosphoramidates, aminoalkylphosphotriesters), those containing pendant moieties, such as, for example, proteins (including enzymes (for example, nucleases), toxins, antibodies, signal peptides, poly-L-lysine, or the like), those with intercalators (for example, acridine, psoralen, or the like), those containing chelates (of, for example, metals, radioactive metals, boron, oxidative metals, or the like), those containing alkylators, those withmodified linkages (for example, alpha anomeric nucleic acids, or the like), as well as unmodified forms of the polynucleotide or oligonucleotide.

[0062] A "promoter" means a nucleic acid sequence sufficient to direct transcription of an operably linked nucleic acid molecule. A promoter can be used together with other transcription control elements (for example, enhancers) that are sufficient to render promoter-dependent gene expression controllable in a cell type-specific, tissue-specific, or temporal-specific manner, or that are inducible by external signals or agents; such elements, may be within the 3' region of a gene or within an intron. In one aspect, the promoter may be operably linked to a nucleic acid sequence, for example, a cDNA, a gene sequence, or an effector RNA coding sequence, in such a way as to enable expression of the nucleic acid sequence, or a promoter is provided in an expression cassette into which a selected nucleic acid sequence to be transcribed can be conveniently inserted.

[0063] An "open reading frame" or "ORF" means a portion of a polynucleotide that, when translated into amino acids, contains no stop codons. The genetic code reads DNA sequences in groups of three base pairs, which means that a double-stranded DNA molecule can read in any of six possible reading frames-three in the forward direction and three in the reverse. An ORF typically also includes an initiation codon at which translation may start.

[0064] As used herein, the term "gene" generally refers to a combination of polynucleotide elements, that when operably linked in either a native or recombinant manner, provide some product or function. The term "gene" is to be interpreted broadly, and can encompass mRNA, tRNA, cDNA, cRNA and genomic DNA forms of a gene. In some uses, the term "gene" encompasses the transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons and introns. In some genes, the transcribed region will contain "open reading frames" that encode polypeptides.

[0065] In some uses of the term, a "gene" comprises only the coding sequences (e.g., an "open reading frame" or "coding region") necessary for encoding a polypeptide. In some aspects, genes do not encode a polypeptide, for example, can encode ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some aspects, the term "gene" includes not only the transcribed sequences, but in addition, also includes non-transcribed regions including upstream and downstream regulatory regions, enhancers and promoters. The term "gene" encompasses mRNA, cDNA and genomic forms of a gene.

[0066] An "amber nonsense mutation" or "amber mutation" or "amber codon" is a mutation in a polynucleotide sequence resulting in premature termination the corresponding polypeptide chain during the translation elongation. Amber mutations are the result of a base substitution that converts a codon specifying an amino acid into UAG, which signals chain termination.

[0067] An "amber suppressor tRNA" is a tRNA whose anticodon recognizes the unique nonsense codon 5'-UAG-3' (amber) in the mRNA and inserts an alternative amino acid at that position in the polypeptide chain.

[0068] The "configuration" of a polynucleotide means the functional sequence elements within the polynucleotide and the order and direction of those elements.

[0069] As used herein, a "polypeptide" or "protein" is any polymer of amino acids (natural or unnatural, or a combination thereof), of any length, typically but not exclusively joined by covalent peptide bonds. A polypeptide can be from any source, e.g., a naturally occurring polypeptide, a polypeptide produced by recombinant molecular genetic techniques, a polypeptide from a cell, or a polypeptide produced enzymatically in a cell-free system. A polypeptide can also be produced using chemical (non-enzymatic) synthesis methods. A polypeptide is characterized by the amino acid sequence in the polymer. The term "peptide" typically refers to a small polypeptide, and typically is smaller than a protein. Unless otherwise stated, it is not intended that a polypeptide be limited by possessing or not possessing any particular biological activity.

[0070] The term "polypeptide" is also intended to include the products of post-expression modifications of the naked amino acid sequence, i.e., modification of the nascent polypeptide, including without limitation and not limited to, glycosylation, acetylation, phosphorylation, amidation and derivatization by known protecting / blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids. A polypeptide may be derived from a natural biological source or produced by recombinant technology, or produced by chemical synthesis.

[0071] As used herein, the term "translation" or "protein translation" refers to the cellular process which results in the synthesis of a protein from a messenger RNA (mRNA) molecule.Translation converts the nucleotide sequence of the mRNA into a specific sequence of amino acids, where the ribosomes read the mRNA's three-nucleotide codons and recruit transfer RNA (tRNA) molecules specific for each codon, each tRNA charged with a specific amino acid, to build a growing polypeptide chain. The process has three main stages: initiation (where the ribosomeassembles on the mRNA), elongation (where the polypeptide chain grows by adding amino acids), and termination (where the completed polypeptide is released).

[0072] As used herein, an "IRES" or "internal ribosome entry site" is an RNA element that allows for translation initiation in a cap-independent manner, as part of the greater process of protein synthesis.

[0073] As used herein, the term "variant" refers to a first composition (e.g., a first molecule), that is related to a second composition (e.g., a second molecule, also termed a "parent" molecule). The variant molecule can be derived from, isolated from, based on or homologous to the parent molecule. For example, a variant aminoacyl-tRNA synthetase polypeptide can be derived from a parent aminoacyl-tRNA synthetase polypeptide by making one or more conservative amino acid substitutions in that parent polypeptide. The term variant can be used to describe either polynucleotides or polypeptides.

[0074] As used herein, a "variant" polypeptide can refer to any sequence variant where the variant molecule has an amino acid sequence that is not identical to the parent molecule, for example, because the variant molecule contains at least one targeted amino acid substitution. As applied to polypeptides, a variant molecule can, for example, have 100% amino acid sequence identity with the original parent molecule and comprise additional amino acid residues, or alternatively, can have less than 100% amino acid sequence identity with the parent molecule. For example, a variant of a parent amino acid sequence can be a second amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or any number less than 100% identical in amino acid sequence compared to the original parent amino acid sequence.Polypeptide variants also include polypeptides comprising the entire parent polynucleotide, and further comprising additional fused amino acid sequences. Polypeptide variants also include polypeptides that are portions or subsequences of the parent polypeptide, for example, unique subsequences (e.g., as determined by standard sequence comparison and alignment techniques).

[0075] In another aspect, polypeptide variants include amino acid sequences that contain minor, trivial or inconsequential changes relative to the parent sequence, for example, where the changes in the amino acid sequence do not appreciably impact the function or enzymatic activity of the polypeptide. For example, minor, trivial or inconsequential changes include changes to an amino acid sequence that (i) result in substitutions, deletions or insertions that have little or noimpact on the biological activity of the polypeptide, and / or (ii) result in the substitution of an amino acid with a chemically similar amino acid, i.e., conservative amino acid substitutions.

[0076] In some advantageous embodiments, the polypeptide variants, for example, an O-RS synthetase variant, has the ability to charge a paired O-tRNA molecule with a desired unnatural amino acid, such as a lysine UAA of FIG. 6, better than the wildtype synthetase polypeptide can charge that same tRNA.

[0077] As used herein, the term "conservative amino acid substitution" in an amino acid sequence refers to a substitution in the original parent amino acid sequence with an amino acid having a chemically similar structure. Conservative amino acid substitutions are well understood in the art, where one amino acid residue is substituted for another amino acid residue having similar chemical properties (e.g., aromatic side chains or positively charged side chains), and therefore do not substantially change the functional properties, i.e., the activity, of the resulting polypeptide molecule. A nonconservative amino acid substitution is a substitution of an amino acid from a parent sequence with an amino acid having dissimilar chemical properties.

[0078] The following are groupings of natural amino acids that contain similar chemical properties, where substitutions within a group is a "conservative" amino acid substitution.Table 1

[0079] In reference to the table above, conservative substitutions involve substitutions between amino acids in the same class. Nonconservative substitutions constitute exchanging a member of one of these classes for a member of another class.

[0080] These groupings indicated above are not rigid, as these natural amino acids can be placed in different groupings when different functional properties are considered. A skilled artisan understands and is able to readily predict whether the substitution of one amino acid in a polypeptide with a different amino acid is likely to preserve biological activity of that polypeptide.

[0081] As used herein, the term "conservative variant translation component" refers to a translation component, e.g., a conservative variant O-tRNA or a conservative variant O-RS, that functionally performs identically or similarly to a base or parent molecule, e.g., an O-tRNA or O-RS, having variations in the sequence as compared to a reference O-tRNA or O-RS. For example, an 0-RS, or a conservative variant of that O-RS, will aminoacylate a cognate O-tRNA with a specified unnatural amino acid. In this example, the O-RS and the conservative variant O-RS do not have the same amino acid sequences. The conservative variant can have, e.g., one variation, two variations, three variations, four variations, or five or more variations in sequence, as long as the conservative variant is still complementary and functions with the corresponding O-tRNA or O-RS.

[0082] In some embodiments, a conservative variant O-RS comprises one or more conservative amino acid substitutions compared to the O-RS from which it was derived. In some embodiments, a conservative variant O-RS comprises a limited number of conservative amino acid substitutions compared to the O-RS from which it was derived, for example, conservative amino acid substitutions that number not more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90 or 100 substitutions in total.

[0083] In some embodiments, a conservative variant O-RS comprises one or more conservative amino acid substitutions compared to the O-RS from which it was derived, and furthermore, retains O-RS biological activity; for example, a conservative variant O-RS that retains at least 10% of the biological activity of the parent O-RS molecule from which it was derived, or alternatively, at least 20%, at least 30%, or at least 40%. In some preferred embodiments, the conservative variant O-RS retains at least 50% of the biological activity of the parent O-RS molecule from which it was derived, or at least 60%, or at least 70%, or at least 80%, or at least 90%, or at least 95%. The amino acid substitutions of a conservative variant O-RS can occur in any domain of the O-RS, including the amino acid binding pocket.

[0084] As used herein, the terms "heterologous" or "exogenous" as applied to polynucleotides or polypeptides refers to molecules that have been rearranged or artificially supplied to a biological system and are not in a native configuration (e.g., with respect to sequence, genomic position or arrangement of parts) or are not native to that particular biological system. These terms indicate that the relevant material originated from a source other than the naturally occurring source, or refers to molecules having a non-natural configuration, genetic location orarrangement of parts. The terms "exogenous" and "heterologous" are sometimes used interchangeably with "recombinant."

[0085] As used herein, two elements are "heterologous" to one another if not naturally associated. For example, a nucleic acid sequence encoding a protein linked to a heterologous promoter means a promoter other than that which naturally drives expression of the protein. A heterologous nucleic acid flanked by transposon ends or inverted terminal repeats ("ITR"s) means a heterologous nucleic acid not naturally flanked by those transposon ends or ITRs, such as a nucleic acid encoding a polypeptide other than a transposase, including an antibody heavy or light chain. A nucleic acid is heterologous to a cell if not naturally found in the cell or if naturally found in the cell but in a different location (e.g., episomal or different genomic location) than the location described.

[0086] As used herein, the terms "native" or "endogenous" refer to molecules that are found in a naturally occurring biological system, cell, tissue, species or chromosome under study. A "native" or "endogenous" gene is generally a gene that does not include nucleotide sequences other than nucleotide sequences with which it is normally associated in nature (e.g., a nuclear chromosome, mitochondrial chromosome or chloroplast chromosome). An endogenous gene, transcript or polypeptide is encoded by its natural locus, and is not artificially supplied to the cell.

[0087] The term "host" or "host cell" means any prokaryotic or eukaryotic organism, typically a single-cell host, that can be a recipient of a heterologous nucleic acid. A "host cell" includes prokaryotic or eukaryotic organisms, for example mammalian host cells, that can be genetically engineered. For examples of such hosts, see Maniatis et al., Molecular Cloning. A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y. (1982). As used herein, the terms "host," "host cell," "host system," and "expression host" are used interchangeably.

[0088] An "isolated" polypeptide or polynucleotide means a polypeptide or polynucleotide that has been either removed from its natural environment, produced using recombinant techniques, or chemically or enzymatically synthesized. Polypeptides or polynucleotides may be purified, that is, essentially free from any other polypeptide or polynucleotide and associated cellular products or other impurities.

[0089] The term "selectable marker" means a polynucleotide segment that allows one to select for or against a molecule or a cell that contains it, often under particular conditions. These markers can encode an activity, such as, but not limited to, production of RNA, a peptide, or aprotein, or these markers can provide a binding site for RNA, peptides, proteins, inorganic and organic compounds, or compositions. Examples of selectable markers include, but are not limited to: (1) DNA segments that encode products that provide resistance against otherwise toxic compounds (e.g., antibiotics); (2) DNA segments that encode products that are otherwise lacking in the recipient cell (e.g., tRNA genes, auxotrophic markers); (3) DNA segments that encode products that suppress the activity of a gene product; (4) DNA segments that encode products that can be readily identified (e.g., phenotypic markers such as beta-galactosidase, GFP, and cell surface proteins); (5) DNA segments that bind products that are otherwise detrimental to cell survival and / or function; (6) DNA segments that otherwise inhibit the activity of any of the DNA segments described in (1) to (5) above (e.g., antisense oligonucleotides); (7) DNA segments that bind products that modify a substrate (e.g. restriction endonucleases); (8) DNA segments that can be used to isolate a desired molecule (e.g. specific protein binding sites); (9) DNA segments that encode a specific nucleotide sequence that can be otherwise non-functional (e.g., for PCR amplification of subpopulations of molecules); and / or (10) DNA segments, which when absent, directly or indirectly confer sensitivity to particular compounds.

[0090] Sequence identity can be determined by aligning sequences using algorithms, such as BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package Release 7.0 (Genetics Computer Group, 575 Science Dr., Madison, Wis.), using default gap parameters, or by inspection, and the best alignment (i.e., resulting in the highest percentage of sequence similarity over a comparison window). Percentage of sequence identity is calculated by comparing two optimally aligned sequences over a window of comparison, determining the number of positions at which the identical residues occur in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of matched and mismatched positions not counting gaps in the window of comparison (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. Unless otherwise indicated, the window of comparison between two sequences is defined by the entire length of the shorter of the two sequences. Identity or homology with respect to such sequences is defined herein as the percentage of amino acid residues in the candidate sequence that are identical with the known peptides, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent homology, and not considering any conservative substitutions as part of the sequenceidentity. N-terminal, C-terminal, or internal extensions, deletions, or insertions into the peptide sequence shall not be construed as affecting homology.

[0091] A "transposase" is a polypeptide that catalyzes the excision of a corresponding transposon from a donor polynucleotide, for example a plasmid, and (providing the transposase is not integration-deficient) the subsequent integration of the transposon into a target nucleic acid. A transposon / transposase may be a piggyBac-like transposon / transposase, or any other type of transposon system. Non-limiting, suitable transposons / transposases are disclosed in U.S. Patent Nos. 10,233,454 and 11,060,098, each of which is incorporated herein by reference in its entirety.

[0092] The term "transposition" refers to the action of a transposase in excising a transposon from one polynucleotide and then integrating it, either into a different site in the same polynucleotide, or into a second polynucleotide.

[0093] The term "transposon" means a polynucleotide that can be excised from a first polynucleotide, for instance, a plasmid, and be integrated into a second position in the same polynucleotide, or into a second polynucleotide, for instance, the genomic or extrachromosomal DNA of a cell, by the action of a corresponding trans-acting transposase. A transposon comprises a first transposon end and a second transposon end, which are polynucleotide sequences recognized by and transposed by a transposase. A transposon usually further comprises a first polynucleotide sequence between the two transposon ends, such that the first polynucleotide sequence is transposed along with the two transposon ends by the action of the transposase. Natural transposons frequently comprise DNA encoding a transposase that acts on the transposon. Transposons as used herein are "synthetic transposons," comprising a heterologous polynucleotide sequence that is transposable by virtue of its juxtaposition between two transposon ends.

[0094] The expressions "corresponding transposon" and "corresponding transposase" are used to indicate an activity relationship between a transposase and a transposon. A transposase enzyme transposes its corresponding / cognate transposon.

[0095] The term "transposon end" means the cis-acting nucleotide sequences that are sufficient for recognition by and transposition by a corresponding transposase. For example, transposon ends of piggyBac-like transposons comprise perfect or imperfect repeats such that the respective repeats in the two transposon ends are reverse complements of each other. These are referred to as terminal inverted repeats or "ITRs." A transposon end may or may not include anadditional sequence proximal to the ITR that promotes or augments transposition. Non-limiting, suitable transposon ends are disclosed in U.S. Patent Nos. 10,233,454 and 11,060,098.

[0096] A transposase protein can be introduced into a cell as a protein or as a nucleic acid encoding the transposase, for example as a ribonucleic acid, including mRNA or any polynucleotide recognized by the translational machinery of a cell; as DNA, e.g., as extrachromosomal DNA including episomal DNA; as plasmid DNA, or as viral nucleic acid. Furthermore, the nucleic acid encoding the transposase protein can be transfected into a cell as a nucleic acid vector such as a plasmid, or as a gene expression vector, including a viral vector. The nucleic acid can be circular or linear. DNA encoding the transposase protein can be stably inserted into the genome of the cell or into a vector for constitutive or inducible expression. Where the transposase protein is transfected into the cell or inserted into the vector as DNA, the transposase encoding sequence may be operably linked to a heterologous promoter. There are a variety of promoters that could be used, including constitutive promoters, tissue-specific promoters, inducible promoters, speciesspecific promoters, cell-type specific promoters, and the like. All DNA or RNA sequences encoding transposase proteins are expressly contemplated. Alternatively, the transposase may be introduced into the cell directly as protein, for example using cell-penetrating peptides (e.g., as described in Ramsey and Flynn, 2015. Pharmacol. Ther. 154: 78-86 "Cell-penetrating peptides transport therapeutics into cells"); using small molecules including salt plus propanebetaine (e.g., as described in Astolfo et. Al., 2015. Cell 161: 674-690); or electroporation (e.g., as described in Morgan and Day, 1995. Methods in Molecular Biology 48: 63-71 "The introduction of proteins into mammalian cells by electroporation").

[0097] The term "vector," "DNA vector," or "gene transfer vector" refers to a polynucleotide that is used to perform a "carrying" function for another polynucleotide. For example, vectors are often used to allow a polynucleotide to be propagated within a living cell, to allow a polynucleotide to be packaged for delivery into a cell, or to allow a polynucleotide to be integrated into the genomic DNA of a cell. A vector may further comprise additional functional elements, such as, for example, a transposon.

[0098] The disclosure refers to several genes and proteins for which it provides an example "SEQ ID NO:." Unless otherwise apparent from the context, reference to a gene or protein should be understood as including the specific SEQ ID NO, as well as allelic, species, and induced variantsthereof having at least 90, 95, or 99% identity thereto. Examples of allelic and species variants can be found in the SwissProt and other databases.

[0099] As used herein, various synthetase enzymes described herein use their GenBank Accession Number for reference. For example, the pyrrolysyl-synthetase from the bacteria Syntrophomonadaceae is referred to alternatively as HHW30035 or HHW30035.1. As used herein, these two GenBank Accession numbers are used interchangeably, and both refer to the amino acid sequence provided in the sequences found in FIG. 15. For example, the bacterial pyrrolysyl-synthetases of GenBank Accession Numbers HHW30035 or HHW30035.1 both refer to the amino acid sequence provided in SEQ ID NO: 52.

[0100] Mutations are sometimes referred to in the form XnY, wherein X is a wildtype amino acid, n is an amino acid position of X in a wildtype sequence, and Y is a replacement amino acid. If the mutation occurs in a sequence having a different number of amino acids than the wildtype sequence, it is present at the position in the sequence aligned with position n in the wildtype sequence when the respective sequences are maximally aligned.

[0101] As used herein, the term "translation system" refers to the components that incorporate an amino acid into a growing polypeptide chain (protein). Components of a translation system can include, e.g., ribosomes, tRNAs, aminoacyl-tRNA synthetases, mRNA, and the like. A translation system can be used in vitro (cell-free) or in vivo (in a host cell). In some aspects, the amino acids that are incorporated into the polypeptide can be considered part of the translation system.

[0102] As used herein, the expression "orthogonal translation system" refers to an engineered biological system that permits translation (protein synthesis) allowing for the incorporation of unnatural amino acids (UAA) into the growing polypeptide without disrupting normal cellular functions. The orthogonal translation system, as used herein, interacts with aspects of the cell's native translation machinery, such as the ribosomes, to produce polypeptides that incorporate unnatural amino acids. These orthogonal systems achieve "orthogonality" by using customized components, i.e., exogenous or heterologous components, such as orthogonal tRNA (O-tRNA), orthogonal aminoacyl-tRNA synthetase (O-RS) that has the ability to charge the O-tRNA with a particular unnatural amino acid (UAA) in a host cell. In some aspects, the orthogonal translation system includes the UAA and / or the host cell. In other aspects, the orthogonal translation system can include an artificial transcript, i.e., non-native genetic instructions, comprising an open readingframe engineered with a selector codon recognized by the O-tRNA at a position in the ORF corresponding to the insertion site of the UAA in the polypeptide.

[0103] In some aspects, the term "orthogonal" or "orthogonal system" or "orthogonal component" or the like refers generally to two or more biomolecules that are able to function independently of each other, without interfering with or being affected by each other's operations or substrates, and without causing unintended cross-talk or interference with one another. For example, a non-native tRNA that can be charged with an unnatural amino acid in a host cell is orthogonal in the host cell if none of the native, endogenous aminoacyl-tRNA synthetases is able to charge that non-native tRNA with an amino acid.

[0104] As used herein, the term "cognate" refers to components that naturally pair in an in vivo environment, and are intended to function together in vivo. With regard to orthogonal translation components, e.g., a Methanosarcina mazei pyrrolysyl tRNA synthetase (e.g., WP_011033391.1) charges its cognate tRNAPylin vivo to incorporate pyrrolysine into polypeptides.

[0105] The expression "selector codon" as used herein refers to codons recognized by the O-tRNA in the translation process and not recognized by an endogenous tRNA. The O-tRNA anticodon loop recognizes the selector codon on the mRNA and incorporates its amino acid, e.g., an unnatural amino acid, at this site in the polypeptide. Selector codons can include, e.g., nonsense codons, such as, stop codons, e.g., amber, ochre, and opal codons; four or more base codons; rare codons; codons derived from natural or unnatural base pairs and / or the like.

[0106] As used herein, a suppressor tRNA is a tRNA that alters the reading of a messenger RNA (mRNA) in a given translation system, e.g., by providing a mechanism for incorporating an amino acid into a polypeptide chain in response to a selector codon. For example, a suppressor tRNA can read through, e.g., a stop codon (e.g., an amber, ocher or opal codon), a four base codon, a rare codon, etc.

[0107] As used herein, the expression "unnatural amino acid" (UAA) refers to any artificial (i.e., synthetic) amino acid that is not one of the 20 naturally occurring, proteinogenic canonical amino acids, and further, is also not pyrrolysine or seleno-cysteine, which also occur in natural systems. Unnatural amino acids are not found in nature. FIG. 6 shows 12 unnatural amino acid structures that find use with the orthogonal translation systems described herein.

[0108] As used herein, the expressions "non-canonical amino acid" or "ncAA" is an amino acid that is not part of the standard set of 20 amino acids used in protein synthesis. Non-canonical amino acids include pyrrolysine and seleno-cysteine, which exist as natural products.

[0109] As used herein, the term "derived from" refers to a molecule that is isolated from or made using a specified starting molecule or organism, or information from the specified molecule or organism. For example, a polypeptide that is derived from a second polypeptide includes polypeptides that contain one or more amino acid substitutions relative to the second polypeptide. In the case of polypeptides, the derived species can be obtained by, for example, naturally occurring mutagenesis, artificial directed mutagenesis or artificial random mutagenesis. The mutagenesis used to derive polypeptides can be intentionally directed or intentionally random, or a mixture of each. Mutagenesis of a polypeptide typically entails manipulation of the polynucleotide that encodes the polypeptide.

[0110] As used herein, the terms "eubacteria" and "bacteria" are used synonymously, and refer to prokaryotic organisms that are distinguishable from Archaea. Similarly, the terms "Archaea" or "Archaeabacteria" refer to prokaryotes that are distinguishable from eubacteria. As understood by one of skill in the art, Eubacteria and Archaea can be distinguished from each other by a number morphological and biochemical criteria. For example, assignment of an organism to Eubacteria or Archaea can be made based on differences in ribosomal RNA sequences, RNA polymerase structure, the presence or absence of introns, antibiotic sensitivity, the presence or absence of cell wall peptidoglycans and other cell wall components, the branched versus unbranched structures of membrane lipids, and the presence / absence of histones and histone-like proteins.

[0111] The present disclosure provides discussion of mechanistic theories explaining in vivo phenomena, including protein translation, e.g., translation that utilizes orthogonal translation machinery for the incorporation of unnatural amino acids into polypeptides. However, it is not intended that the disclosure of the specification be limited in any regard to the molecular mechanism of action, and knowledge of such mechanisms is not required to make or use the compositions or methods described herein. It is not intended that the term "orthogonal translation system" and related terms as used herein be limited in any way with regard to the in vivo mechanism of such systems.DETAILED DESCRIPTIONI. General

[0112] Most broadly, the present disclosure provides compositions and methods for incorporating unnatural amino acids into a growing polypeptide chain in response to a selector codon, e.g., an amber stop codon. In some aspects, these compositions and methods are used in vivo (e.g., in a host cell). These compositions include pairs of orthogonal-tRNAs (O-tRNAs) and orthogonal aminoacyl-tRNA synthetases (O-RSs) that do not interact with, or do not impair or only minimally impact, the host cell tRNA or RS molecules. That is to say, the O-tRNA is not charged (or not charged to a significant or detectable level) with an amino acid (natural or unnatural) by an endogenous host cell aminoacyl-tRNA synthetase. Similarly, the O-RSs provided by the disclosure do not charge any endogenous tRNA with an amino acid (natural or unnatural) to a significant or detectable level. These novel compositions permit the production of large quantities of proteins having translationally incorporated unnatural amino acids, which can be targeted for further modification or conjugation. These proteins incorporating reactive unnatural amino acids find a wide variety of uses, including therapeutics, for example, antibody therapeutics, and in biological research.

[0113] Using a range of bioinformatics tools for discovery and engineering, we have developed new synthetase-tRNA pairs that show robust incorporation of desirable UAAs into proteins having potential in a variety of applications, for example, in the development of protein therapeutics, such as antibody therapeutics.II. Unnatural Amino Acids

[0114] A diverse collection of unnatural amino acids (UAA) find use when incorporated into proteins co-translationally, each for different reasons. Of particular use are UAA that have reactive groups that can be readily modified in a highly target-specific manner.

[0115] The unnatural amino acid is selected from unnatural amino acids that contain reactive side chain groups that are reactive at physiological conditions. As a result, polypeptides that incorporate these amino acids are modifiable under conditions that preserve the primary, secondary and tertiary structure of the polypeptides being modified. As a result, these proteins containing one or more UAA can be modified, e.g., conjugated with other moieties, in a manner that preserves the biological activity of the protein being modified. Examples of preferred reactiveUAA that can be incorporated into proteins are provided in the chemical structures shown in FIG.6. All of these structures are lysine derivatives, with the rationale that the lysine derivative UAAs would have cross reactivity with pyrrolysine RS and tRNAPylbecause of the common lysine structures.

[0116] A list of lysine derivative UAA with reactive side group chemistries finding use with the methods described herein includes, but not limited to:N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine,N6-(norbornene-methoxycarbonyl)-L-lysine,N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine,N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysineN6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine,N6-((2-azidoethoxy)carbonyl)-L-lysine,N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine, N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine,N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine,N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine,N-(2-(4-(6-methyl-l,2,4,5-tetrazin-3-yl)phenyl)acetyl)-L-lysine, andN6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine.The chemical structures of these UAA are shown in FIG.6.

[0117] The functional groups of these UAA that enable biocompatible conjugation include those that enable "Click chemistry." "Click chemistry" refers to a range of chemical reactions that are highly efficient, selective, and modular, allowing for quick and simple joining of molecular building blocks. "Click" reactions produce a single, high-purity product under mild conditions, minimizing waste and simplifying product purification. For example, the copper-catalyzed azidealkyne cycloaddition (CuAAC) reaction, is commonly used in bioconjugation to attach various moieties to biomolecules.

[0118] Functional groups that enable Click chemistry include but are not limited to azide (N3), alkyne, strained alkyne (eg., SCO, BCN), strained alkene (e.g., TCO, cyclopropene), pentadiene, norbornene, tetrazine, acetylphenyl and others. Click chemistry conjugations can include but are not limited to reactions between an azide and alkyne, an azide and a strained alkyne or strained alkene, a strained alkene or alkyne and tetrazine, an acetylphenyl group and a nucleophile, and a diene and dienophile. Click chemistry reactions enabled by UAAs can include but are not limited tocopper-catalyzed or copper-free conjugation, inverse-electron-demand Diels-Alder reactions, Diels-Alder cycloadditions, and oxime ligations. The UAA functional group can also allow for other methods of conjugation, including but not limited to cross-linking through a photo-activated functional group such as diazirine. The UAA may also contain functional groups for purposes other than covalent conjugation, such as biotin for strong interaction with streptavidin.III. Orthogonal Translation Systems Generally

[0119] In some aspects, the present disclosure provides translation systems. These systems comprise an orthogonal aminoacyl-tRNA synthetase (O-RS), an orthogonal tRNA (O-tRNA), and an unnatural amino acid, where the first O-RS preferentially aminoacylates the first O-tRNA with the first unnatural amino acid.

[0120] In general, when an orthogonal pair recognizes a selector codon and loads an amino acid in response to the selector codon, the orthogonal pair is said to "suppress" the selector codon. That is, a selector codon that is not recognized by the cell's endogenous translation machinery is not ordinarily translated, which can result in blocking production of a polypeptide that would otherwise be translated from the nucleic acid. An O-tRNA as provided herein recognizes a selector codon and includes at least about, e.g., a 45%, a 50%, a 60%, a 75%, a 80%, or a 90% or more suppression efficiency in the presence of a cognate synthetase in response to a selector codon as compared to the suppression efficiency of an O-tRNA comprising or encoded by a polynucleotide sequence as set forth in the sequence listing herein.

[0121] In certain desirable aspects, the cell can include an additional O-tRNA / O-RS pair, where the additional O-tRNA is loaded by the additional O-RS with a different unnatural amino acid. In that case, each of the two O-tRNA molecules must recognize a different selector codon, for example, a different stop codon, or a four base codon as known in the art.

[0122] The O-tRNA and / or the O-RS can be naturally occurring or can be, e.g., derived by mutation of a naturally occurring tRNA and / or RS.IV. Orthogonal Pairs

[0123] Examples of paired O-RS and O-tRNA molecules finding use in the translation systems described herein for the incorporation of UAA are provided in Tables 2, 3, 4, 5, 6 and 7. The O-RS and O-tRNA amino acid and nucleotide sequences of these O-RS and O-tRNA molecules are provided in FIG. 15.

[0124] These O-RS and O-tRNA are provided as matched pairs, where each ideally has specificity, or at least a preference, for reaction with the other. For example, an O-RS as used herein is specific or has a preference for charging (i.e., aminoacylating) a particular tRNA target (namely, the O-tRNA) with a defined amino acid, in this case, a defined UAA. As described herein, O-RS and O-tRNAUAAwere developed for use in mammalian cells by repurposing and / or modifying RS enzymes and tRNA molecules identified in eubacteria and Archaea that were predicted to have specificity for the amino acid pyrrolysine (i.e., pyrrolysyl-synthetase and tRNAPyl). This approach was selected because pyrrolysine tRNAs have evolved to be efficient orthogonal suppressers in a broad range of species in which pyrrolysine is utilized and some pyrrolysine synthetases have been reported to show expanded amino acid specificity beyond pyrrolysine while remaining orthogonal with respect to canonical translation systems.

[0125] Because of the utility of UAA such as those shown in FIG. 6, it was hoped that a pyrrolysyl-tRNA synthetase and tRNAPyl, or mutants or variants thereof, would have sufficient activity and specificity with such UAAs and to operate well and in an orthogonal manner in mammalian cells.

[0126] As discussed herein, a variety of synthetase / tRNA pairs derived from eubacteria and Archaea were identified that had the ability to charge the tRNA with various UAA (e.g., Azc-Lys or 2AzZ-Lys), and further, where the charged tRNA can incorporate the UAA into a polypeptide chain during protein translation. In some cases, the ability of the O-RS and O-tRNA pair to produce a protein containing the UAA could be improved by making various targeted amino acid substitutions in the RS.

[0127] O-RS molecules finding use with the methods described herein can be bacterial pyrrolysyl-RS, Archaeal pyrrolysyl-RS, or variants derived from bacterial pyrrolysyl-RS, or variants derived from an Archaeal pyrrolysyl-RS. In some aspects, the bacterial or Archaeal orthogonal pyrrolysyl-RS is a native RS that also has the ability to charge a cognate tRNAPylwith a desired UAA, in place of the pyrrolysine that occurs in the native bacterial or Archaeal context.

[0128] In some aspects, the orthogonal translation systems as described herein are crossDomain hybrid systems where the O-RS and O-tRNA are derived from different Domains, namely Eubacteria and Archaea. Most broadly, the disclosure provides orthogonal translation systems that include O-RS and O-tRNA pairs where the O-RS can aminoacylate the O-tRNA with a desired UAA. The O-RS and O-tRNA pairs can comprise any combination of Archae and bacterial (i.e.,eubacterial) components. For example, the O-RS / O-tRNA pairs can comprise: (i) an Archae O-RS and an Archae O-tRNA, (ii) a bacterial O-RS and a bacterial O-tRNA, (iii) an Archae O-RS and a bacterial O-tRNA, or (iv) a bacterial O-RS and an Archae O-tRNA. Mixed orthogonal pairs of categories (iii) and (iv) can be considered "hybrid" systems, and such hybrid systems are indicated in the listing of orthogonal pairs provided in Tables 2, 3, 4, 5, 6 and 7, below.

[0129] With regard to the use of bacterial O-RS and / or bacterial O-tRNA components, the bacterial origin can be further defined, where the bacterial O-RS and / or bacterial O-tRNA components of various sources find use in the methods of the disclosure. For example, useful molecules include, but not limited to, O-RS and / or O-tRNA molecules derived from the bacterial family Syntrophomonadaceae, or the order Thermoanaerobacterales (exemplary species Syntrophaceticus schinkii), or the family Peptococcaceae (exemplary species Paradesulfitobacterium ferrireducens), or derived from Alphaproteobacteria. Various illustrative synthetase enzymes, and the bacterial species from which they are derived, are described in the table provided in FIG. 16.

[0130] Combinations of O-RS and O-tRNA molecules finding use in the systems and methods described herein include those recited in the Tables 2, 3, 4, 5, 6 and 7, which were all shown experimentally to be effective at incorporation of the indicated UAA when paired according to these tables. The amino acid and nucleotide sequences of these O-RS and O-tRNA pairs are provided in FIG. 15.Table 2Orthogonal systems for the incorporation of Azc-Lys(N6-((2-azidoethoxy)carbonyl)-L-lysine)Table 3Orthogonal systems for the incorporation of 2AzZ-Lys(N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine)Table 4Orthogonal systems for the incorporation of SpHD-LysTable 5Orthogonal systems for the incorporation of Nrb-Lys(N6-(norbornene-methoxycarbonyl)-L-lysine)Table 6Orthogonal systems for the incorporation of TCO-Lys(N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine)Table 7Orthogonal systems for the incorporation of:SCO-Lys; N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysineCyclopropene-Lys; N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine Endo-BCN-Lys; N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine H-L-Photo-Lys; N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine

[0131] Engineered polypeptides, for example, engineered synthetase enzymes, are a feature of the disclosure, as provided in FIG. 15. An engineered O-RS polypeptide includes those RS molecules that are artificial, i.e., are not naturally occurring. Polynucleotides that encode the engineered polypeptides are also a feature of the disclosure.

[0132] Engineered tRNA molecules, for example, engineered O-tRNA molecules, and nucleic acids that encode them, are a feature of the disclosure. An engineered O-tRNA includes those 0-tRNA molecules that are artificial, i.e., are not naturally occurring. Artificial expression cassettes are also a feature of the disclosure.

[0133] Any vector or any type of nucleic acid comprising a polynucleotide disclosed herein (e.g., a tRNA disclosed herein or a polynucleotide open reading frame encoding an aminoacyl-tRNA synthetase disclosed herein) are also a feature of the disclosure. For example, a vector of the disclosure can include any type of plasmid, cosmid, transposon, phage, any type of virus, anexpression vector, and / or the like. A cell comprising a nucleic acid / vector such as these is also a feature of the disclosure.V. Selector Codons

[0134] Selector codons are inserted into an mRNA open reading frame to program the incorporation of a UAA into a polypeptide during protein translation. Selector codons expand the genetic codon framework of protein biosynthetic machinery. As known in the art, a variety of selector codons can be used in orthogonal translation systems. For example, selector codon include a unique three base codon, a nonsense codon, such as a stop codon, e.g., an amber codon (UAG), or an opal codon (UGA), an unnatural codon, at least a four base codon, a rare codon, or the like. Any number of selector codons can be introduced into a desired gene, e.g., one or more, two or more, three or more, or more than three, etc.

[0135] By using different selector codons, multiple orthogonal tRNA / synthetase pairs can be simultaneously used to allow the site-specific incorporation of multiple different unnatural amino acids.

[0136] In one embodiment, the methods involve the use of a selector codon that is a stop codon for the incorporation of an unnatural amino acid in vivo in a cell. For example, an O-tRNA is produced that is aminoacylated with an unnatural amino acid by an O-RS, and the O-tRNA recognizes the stop codon. This O-tRNA is not recognized by the naturally occurring host's aminoacyl-tRNA synthetases. When the O-RS, O-tRNA and the nucleic acid that encodes a polypeptide of interest are combined, e.g., in vivo, the unnatural amino acid is incorporated in response to the stop codon to give a polypeptide containing the unnatural amino acid at the specified position. In one embodiment of the invention, the stop codon used as a selector codon is an amber codon, UAG, and / or an opal codon, UGA. In one example, a genetic code in which UAG and UGA are both used as a selector codon can encode 22 amino acids while preserving the ochre nonsense codon, UAA, which is the most abundant termination signal.

[0137] The incorporation of unnatural amino acids in vivo can be done without significant perturbation of the host cell. In eukaryotic cells, because the suppression efficiency for the UAG codon depends upon the competition between the O-tRNA, e.g., the amber suppressor tRNA, and a eukaryotic release factor (e.g., eRF) (which binds to a stop codon and initiates release of the growing peptide from the ribosome), the suppression efficiency can be modulated by, e.g.,increasing the expression level of O-tRNA, e.g., the suppressor tRNA. In addition, additional compounds can also be present, e.g., reducing agents such as dithiothretiol (DTT).

[0138] Unnatural amino acids can also be encoded with rare codons. For example, when the arginine concentration in an in vitro protein synthesis reaction is reduced, the rare arginine codon, AGG, has proven to be efficient for insertion of Ala by a synthetic tRNA acylated with alanine. An unassigned codon AGA in Micrococcus luteus has been utilized for insertion of amino acids in an in vitro transcription / translation extract. Material and methods described herein can be adapted to use these rare codons in vivo.

[0139] Selector codons can also comprise extended codons, e.g., four or more base codons, such as, four, five, six or more base codons. Examples of four base codons include, e.g., AGGA, CUAG, UAGA, CCCU, and the like. Examples of five base codons include, e.g., AGGAC, CCCCU, CCCUC, CUAGA, CUACU, UAGGCand the like. Four-base codons have been used to incorporate unnatural amino acids into proteins using in vitro biosynthetic methods and a variety of orthogonal systems.

[0140] For a given system, a selector codon can also include one of the natural three base codons, where the endogenous system does not use (or rarely uses) the natural base codon. For example, this includes a system that is lacking a tRNA that recognizes the natural three base codon, and / or a system where the three base codon is a rare codon.VI. Variants

[0141] Variant O-RS and O-tRNA sequences are also within the teaching of the present disclosure. Variant O-RS proteins can comprise any number of targeted amino acid substitutions, which can be conservative amino acid substitutions or non-conservative amino acid substitutions.

[0142] In some aspects, the substitution(s) result in either an O-RS with the same or similar O-tRNAUAAcharging activity as the native RS sequence, or improved O-tRNAUAAcharging activity compared to the native RS sequence or the parent RS sequence.

[0143] In some aspects, variant O-RS molecules can comprise directed amino acid substitutions that improve the activity of the RS molecule, such that charging of a cognate O-tRNA is improved as compared to the charging by the native or wild type RS sequence, and as a result, improved production of protein comprising the UAA is observed as compared to an unmodified system.

[0144] I n various aspects, the O-RS variants can have at least 80%, or 90%, or 95%, or 98%, or 99% amino acid identity with the amino acid sequence of the parent O-RS molecule. In some aspects, the variant O-RS has O-tRNAUAAcharging activity that is at least equal to or better than the UAA charging activity of the native RS amino acid sequence. In other aspects, the variant O-RS has O-tRNAUAAcharging activity that is at least 50% of the UAA charging activity of the starting RS amino acid sequence.

[0145] In another aspect, O-RS variants include amino acid sequences that contain minor, trivial or inconsequential changes relative to the parent sequence. For example, minor, trivial or inconsequential changes include changes to an amino acid sequence that (i) result in substitutions, deletions or insertions that have little or no impact on the biological activity of the polypeptide, and / or (ii) result in the substitution of an amino acid with a chemically similar amino acid, i.e., a conservative amino acid substitution.

[0146] In other aspects, variants of the O-tRNA molecules as described herein is also a feature of the disclosure. In one aspect, these variants are silent variants where nucleotide substitutions can be made in the tRNA molecule that do not result in any changes in the tRNA folding structure or activity, i.e., there is no change in the ability to form a covalent coupling with the UAA in the synthetase charging process. Conversely, O-tRNA variants can also include variants the have improved charging properties where the tRNA is more readily charged with the desired UAA as compared to the parent tRNA molecule, i.e., has improved charging capacity. Any of a number of assays can be used to determine aminoacylation efficiency. These assays can be performed in vitro or in vivo. Aminoacylation can also be determined by using a reporter along with orthogonal translation components and detecting the reporter in a cell expressing a polynucleotide comprising at least one selector codon that encodes a protein.

[0147] I n various aspects, a variant O-tRNA can have a nucleotide sequence that has at least about 80%, or 90%, or 95%, or 98%, or 99% nucleotide sequence identity with the nucleotide sequence of the parent O-tRNA molecule.VII. Methods for Production of Proteins Comprising UAA

[0148] The present disclosure provides methods for the production of proteins comprising one or more UAA at selected positions. These methods utilize the translation system components described herein. Generally, these methods start with the step of providing a translation systemcomprising: (i) an unnatural amino acid that comprises a reactive side chain moiety suitable for bioconjugation reactions; (ii) a first orthogonal aminoacyl-tRNA synthetase (O-RS); (iii) a first orthogonal tRNA (O-tRNA), wherein the O-RS preferentially aminoacylates the O-tRNA with the unnatural amino acid; and, (iv) a nucleic acid encoding the protein, where the nucleic acid comprises at least one selector codon that is recognized by the first O-tRNA. The method then incorporates the unnatural amino acid at the selected position in the protein during translation of the protein in response to the selector codon, thereby producing the protein comprising the unnatural amino acid at the selected position.

[0149] The experimental details of the actual method used will vary greatly in reagents, steps and culture conditions. These variable are determined by the particular host cells used, the type of cell culture apparatus, the cell culture media that is used, the nature of the protein of interest that is being produced that contains the UAA, the production quantities required, and any purification steps that might be required in the isolation / purification of the protein containing the UAA.Methods as described herein are not limited in this regard.

[0150] The compositions and methods as described herein provide the ability to synthesize proteins that comprise reactive unnatural amino acids in large useful quantities.VIII. Systems for Incorporating a Plurality of Different UAA

[0151] In some aspects, the translation system incorporates a second orthogonal pair (that is, a second O-RS and a second O-tRNA) that utilizes a second unnatural amino acid, so that the system is now able to incorporate at least two different unnatural amino acids at different selected sites in a polypeptide. In this dual system, the second O-RS aminoacylates the second O-tRNA with the second unnatural amino acid that is different from the first unnatural amino acid, and the second O-tRNA recognizes a selector codon that is different from the selector codon recognized by the first O-tRNA.

[0152] By extension, the translation systems as described herein can be modified to incorporate a second and a third (or further) orthogonal pairs (that is, a second and third O-RS and a second and third O-tRNA) that utilize a second and third unnatural amino acid, respectively. In this manner, a polypeptide can be programmed to contain two, or three, or more, different UAA at defined positions, where the different UAA codon positions use different selector codons corresponding to the different UAA to be integrated.IX. Host Cells

[0153] In some embodiments, the translation system resides in a host cell (and optionally where the translation system includes the host cell). The host cell used is not particularly limited, as long as the O-RS and O-tRNA retain their orthogonality in their host cell environment. In some aspects, the host cell is preferably a mammalian cell, such as a human cell or a rodent cell. In some embodiments, a Chinese Hamster Ovary (CHO) cell line such as the CHO-K1 cell line or the CHO-S cell line, or any other derivative of a CHO cell line, as known in the industry or created expressly for the purpose of hosting an orthogonal translation system, is the host cell line. Other rodent cell lines can also be used as host cells. In other aspects, a Human Embryonic Kidney (HEK) cell line can be used, such as the HEK293 cell line known in the industry.

[0154] The mammalian host cell line used will preferably contain endogenous transcription and translation machinery suitable for production, preferably high levels of production, of the O-RS and O-tRNA molecules. For example, the host cell line chosen will contain suitable transcription factors that will drive the transcription of both the exogenous O-RS gene and exogenous O-tRNA gene, which can be under the transcriptional control of two different promoters and other cisacting regulatory elements that drive transcription initiation and RNA polymerase elongation.

[0155] Although orthogonal translation systems (e.g., comprising an O-RS, an O-tRNA and an unnatural amino acid) can utilize cultured host cells to produce proteins having unnatural amino acids, it is not intended that an orthogonal translation system of the invention require an intact, viable host cell. For example, an orthogonal translation system can utilize a cell-free system in the presence of a cell extract. Indeed, the use of cell free, in vitro transcription / translation systems for protein production is a well established technique. Adaptation of these in vitro systems to produce proteins having unnatural amino acids using orthogonal translation system components described herein is well within the scope of the present disclosure.X. Improved Compositions for tRNA Expression

[0156] In one aspect, systems and methods are provided for producing a protein with a UAA incorporated at a specific site, where expression of the translation system components is optimized. In one such aspect, the systems and methods reduce the amount of repetitive sequence required to achieve adequate suppressor tRNA expression. In one aspect, such systems and methods include: (i) the use of highly active promoters operably linked to the suppressortRNA-coding sequence, thereby increasing suppressor tRNA expression from each gene copy and reducing the number of genomic copies required; (ii) the use of multiple different promoters operably linked to the suppressor tRNA-coding sequence, thereby reducing the overall repetitive nature of the suppressor tRNA-expressing DNA and thus also reducing its instability; and (iii) the provision of the suppressor tRNA-expressing DNA on transposons and their introduction into the cell by transposition, thereby ensuring that there are multiple independently integrated copies of the suppressor tRNA-expressing DNA, which are much more stable than long concatemers.

[0157] In one aspect, a polynucleotide is provided, the polynucleotide comprising a pig 7sk promoter operably linked to a DNA sequence encoding a heterologous tRNA. In one aspect, the nucleotide sequence of the pig 7sk promoter is SEQ ID NO:16. In one aspect, the heterologous tRNA comprises an anticodon that recognizes a stop codon. In one aspect, the stop codon is an amber codon (UAG). In one aspect, the heterologous tRNA is amino-acylated with a UAA. In one aspect, the heterologous tRNA is naturally amino-acylated with pyrrolysine. In one aspect, the UAA is a lysine-derivative UAA that comprises a functional group capable of biocompatible conjugation, for example, any of the lysine UAA structures shown in FIG. 6 In one aspect, the DNA sequence encoding the heterologous tRNA comprises SEQ ID NO: 28, namely, the Archaeal Methanosarcina mazei tRNAPyl. In other aspects, the heterologous tRNA can be any suitable tRNA that is used in an orthogonal translation system, for example, any of the tRNA molecules recited in Tables 2, 3, 4, 5, 6 and 7, and furthermore, whose nucleotide sequences are provided in FIG. 15.

[0158] In one aspect, the polynucleotide comprises a transposon, the transposon further comprising left and right transposon ends such that the promoter and the DNA sequence encoding the heterologous tRNA are transposable by a corresponding transposase. In one aspect, the polynucleotide further comprises one or more additional promoters selected from a second pig 7sk promoter, a human U68cc promoter, a water buffalo U6 promoter, a human 7sk promoter, and a mouse Hl promoter, and combinations thereof. In one aspect, each additional promoter is operably linked to a separate copy of the DNA sequence encoding a heterologous tRNA. In one aspect, the nucleotide sequence of the one or more additional promoters is selected from SEQ ID NOs: 16, 11, 14, 18, 20, and 31, and combinations thereof.

[0159] In another aspect, a method is provided for producing a protein with a UAA incorporated at a specific site, the method comprising introducing the polynucleotide into a mammalian host cell, for example but not limited to, a CHO-K1 cell.

[0160] In another aspect, a mammalian host cell, such as a CHO-K1 cell, is provided, the host cell comprising the polynucleotide.

[0161] In another aspect, a system is provided for site-specific incorporation of a non-canonical amino acid into an antibody in a mammalian host cell such as a CHO-K1 cell, the system comprising: a transposon comprising multiple copies of DNA encoding pyrrolysine tRNA, each copy operably linked to a pol III promoter active in the mammalian host cell, wherein: (i) at least one of the pol III promoters is a pig 7sk promoter; (ii) at least one of the pol III promoters is not a pig 7sk promoter; and (iii) where there are at least two copies of DNA encoding pyrrolysine tRNA that are each operably linked to a pol III promoter that is not a pig 7sk promoter, the pol III promoters that are not a pig 7sk promoter are different from each other; a second transposon comprising an open reading frame encoding a corresponding aminoacyl tRNA synthetase operably linked to regulatory elements such that the aminoacyl tRNA synthetase is expressible in the mammalian host cell; and a third transposon comprising open reading frames encoding the heavy and light chains of an antibody, wherein the open reading frame encoding one chain comprises an amber stop codon (UAG) preventing expression of a full antibody unless the amber stop codon is read by the pyrrolysine tRNA to incorporate the non-canonical amino acid. In one aspect of the system, the first, second, and third transposons are separate. In one aspect of the system, the first, second, and third transposons are one transposon. In one aspect of the system, he first and second, the first and third, or the second and third transposons are one transposon. In one aspect of the system, any combination of the first, second, and third transposons are transposable by the same transposase or by different transposases.

[0162] In another aspect, a mammalian host cell such as a CHO-K1 cell is provided whose genome comprises the system.

[0163] In another aspect, a method for site-specific incorporation of a non-canonical amino acid, that is to say an unnatural amino acid, into any protein of interest such as an antibody in a mammalian host cell is provided, the method comprising delivering the system components into the host cell, and further optionally where some of each of those components are integrated into the host cell genome. In some embodiments, the unnatural amino acid is a lysine derivative amino acid comprising a functional group capable of biocompatible conjugation.XI. Proteins of Interest

[0164] The diversity of proteins of interest that find use with the compositions and methods described herein for the insertion UAA is not limited in any regard, and further, no attempt is made herein to recite the full spectrum of proteins that might beneficially be modified to incorporate reactive UAA, such as lysine derivative UAA. Essentially any protein (or portion thereof) can include a reactive unnatural amino acid, and can be produced using the compositions and methods described herein. No attempt is made to identify the many thousands of known proteins, any of which can be modified to include one or more unnatural amino acid.

[0165] In some aspects, antibodies find particular use as targets for incorporation of reactive UAA for modification / conjugation of the resulting antibodies. For example, in the development of therapeutic antibodies, it is advantageous to produce antibody-drug conjugates (ADCs) that can deliver a drug payload to a specific site in the body that is targeted by the binding specificity of the antibody. Further for example, ADCs can deliver a potent, cytotoxic payload directly to cancer cells by binding to specific tumor antigens while minimizing damage to healthy tissues, representing a significant advance in targeted cancer therapy.

[0166] Conjugable antibody, especially therapeutic antibodies, are required at high sequence purity and yield. Some methods for conjugation involve reaction with free amines of lysines or cysteines in the antibody sequence. Utilization of lysines is not practical for antibodies that contain lysine residues that are critical for antigen interaction. Free cysteines can also be problematic for production of monodisperse, properly assembled antibodies. For these reasons, methods that target canonical amino acid modification are limited in their use to and do not flexibly allow optimization of conjugation site placement. Efficient incorporation of UAAs in vivo allows the production of a defined ready to conjugate species with full flexibility to place the conjugation site for optimal antibody and drug conjugate function.

[0167] Following incorporation of the UAA into the polypeptide, the unnatural amino acid side chains can be specifically and regioselectively modified. Because of the unique reaction chemistries of these unnatural amino acid substituents, proteins into which they are incorporated can be modified with extremely high selectivity. In some cases, the unnatural amino acid reactive group has the advantage of being completely alien to in vivo systems, thereby improving reaction selectivity. In some aspects, the modification reactions can be conducted using relatively mildreaction conditions that permit either in vitro or in vivo conjugation reactions that preserve protein biological activity.

[0168] The nature of the material that is conjugated to the protein of interest via the unnatural amino acid linkage in the protein of interest is not particularly limited, and can be any desired entity, e.g., dyes, fluorophores, crosslinking agents, saccharide derivatives, polymers (e.g., derivatives of polyethylene glycol), photocrosslinkers, cytotoxic compounds, affinity labels, derivatives of biotin, resins, beads, a second protein or polypeptide (or more), a polynucleotide(s) (e.g., DNA, RNA, etc.), metal chelators, cofactors, fatty acids, carbohydrates, and the like.

[0169] In other aspects, the modification or conjugation of the reactive UAA linkage in the protein of interest to a particular moiety can impart novel biological properties or improved characteristics to the protein to which the moiety is attached.XII. Kits

[0170] The present disclosure also encompasses compositions such as kits to facilitate handling of the compositions of the disclosure as well as methods described in the disclosure. Most generally, the term "kit" is used herein to describe any assemblage of articles that facilitate the execution of a process, method, assay, analysis, manipulation of a sample or reagent, or the like. Kits can contain written instructions describing how to use the kit (e.g., instructions describing the methods of the present disclosure), reagents (e.g., unnatural amino acids) or enzymes required for the method, vectors / plasmids, primers, probes, buffer solutions, any type of containers (for example, containers for sample collection or sample manipulation) or reaction vessels, or any other components. A kit need not contain every component necessary to execute a method of the invention. The compositions and kits of the invention may include, e.g., one or more of any of the reaction components described above with respect to the subject methods.

[0171] In some embodiments, kits of the invention can include the components of an orthogonal translation system, that is to say, (i) at least one unnatural lysine derivative amino acid (UAA), (ii) an orthogonal pyrrolysyl aminoacyl-tRNA synthetase (O-RS), and (iii) a bacterial pyrrolysyl orthogonal tRNA (O-tRNAPyl).

[0172] With regard to the UAA, the amino acid provided in the kit can be any unnatural lysinederivative amino acid comprising a functional group capable of biocompatible conjugation. For example, the UAA can be selected from:N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine,N6-(norbornene-methoxycarbonyl)-L-lysine,N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine,N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysineN6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine,N6-((2-azidoethoxy)carbonyl)-L-lysine,N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine,N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine,N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine,N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine,N-(2-(4-(6-methyl-l,2,4,5-tetrazin-3-yl)phenyl)acetyl)-L-lysine, and N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine.

[0173] With regard to the O-RS and O-tRNA components, any suitable matched pair O-RS / O-tRNA where the O-RS has the ability to charge the O-tRNA with the desired UAA can be used. For example, the O-RS and O-tRNA pair can be any of the O-RS / O-tRNA pairs provided in Tables 2, 3, 4, 5, 6 and 7.

[0174] The O-RS provided with the kit can be supplied on any suitable nucleic acid comprising an open reading frame encoding the amino acid sequence of the synthetase enzyme. The nucleic acid can be any type of expression vector where the synthetase ORF will be expressed in the host cell to be used, that is to say, where the ORF is under the control of a promoter and other regulatory elements that result in transcription and translation of the ORF leading to production of the enzyme. In some embodiments, the ORF encoding the synthetase is delivered to the host cell on a vector comprising a transposon, and further where the transposon is transposed to integrate into the host cell genome.

[0175] The O-tRNA provided with the kit can be supplied on any suitable nucleic acid such as a vector. The nucleic acid can be any type of expression vector where the O-tRNA will be expressed in the host cell to be used, that is to say, where the O-tRNA is under the transcriptional control of a promoter and other regulatory elements that result in transcription to produce the O-tRNA molecule. In some embodiments, the tRNA nucleotide sequence is delivered to the host cell on a vector comprising a transposon, and further where the transposon is transposed to integrate into the host cell genome. In some embodiments, expression of the O-tRNA (which is a heterologoustRNA, is driven by a promoter or promoters (poll 11 type promoters) that are selected for their ability to expresses a heterologous tRNA molecule in a mammalian host cell, as described herein, for example, a 7sk promoter, for example, a pig 7sk promoter. As provided with the kits of the disclosure, a nucleic acid such as a vector encoding the O-tRNA can comprise multiple copies of the O-tRNA gene, which expression can be driven by the same or different promoters.

[0176] The kits of the disclosure can further comprise a mammalian host cell, provided as either a host cell culture, or in frozen form. In the kits of the disclosure that comprise the host cells, the host cells can comprise the nucleic acids encoding the O-RS and / or O-tRNA molecules.

[0177] Kits can also comprise a vector backbone, that is to say an empty vector, that will be used to express a polypeptide of interest, where the nucleic acid ORF encoding that polypeptide of interest will comprise at least one selector codon that is recognized by the O-tRNA. Alternatively, that vector to be used for the expression of a polypeptide of interest can be preprogrammed with the desired ORF. Alternatively still, the kit can comprise an ORF of a marker or control polypeptide that comprises at least one selector codon for incorporation of the UAA.

[0178] Kits of the disclosure can optionally comprise a second O-RS and a second O-tRNA, wherein the second O-RS preferentially aminoacylates the second O-tRNA with a second unnatural amino acid that is different from the first unnatural amino acid, and wherein the second O-tRNA recognizes a selector codon that is different from the selector codon recognized by the first O-tRNA.

[0179] The kits can further include one or more additional reagents or other materials employed in methods of the disclosure, e.g., as described above, where such reagents may include, but are not limited to: cells culture media and media supplements, cell culture flasks, plates, multiwell plates, or any other kind of suitable cell culture vessel substrate, reagents for mammalian cell transfection,

[0180] The subject kits may include, or the compositions and devices may be provided with, one or more test reagents, including e.g., control nucleic acids (e.g., control nucleic acid templates), and the like. In some instances, components of the subject compositions and / or kits may be presented as a "cocktail" where, as used herein, a cocktail refers to a collection or combination of two or more different but similar components in a single vessel. Components of the kits may be present in separate containers, or multiple components may be present in a single container, as desired. The subject compositions may be present in any suitable environment, andmay be transported or stored at a suitable temperature, e.g., room temperature, chilled or frozen. According to one embodiment, the composition is present in a reaction tube or a well. In certain aspects, the composition is present in two or more (e.g., a plurality of) reaction tubes or wells (e.g., a plate, such as a 96-well plate, a multi-well plate, e.g., containing about 1000, 5000, or 10,000 or more wells). The tubes and / or plates may be made of any suitable material, e.g., polypropylene, or the like, PDMS, or aluminum. The containers may also be treated to reduce adsorption of nucleic acids to the walls of the container. Any suitable reaction vessel(s) may be employed in the subject kits or devices and / or to contain a subject composition.

[0181] In addition to the above-mentioned components, a subject kit may further include instructions for using the components of the kit, e.g., to practice the subject methods as described above. The instructions are generally recorded on a suitable recording medium. The instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or sub-packaging) etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g., portable flash drive, CD-ROM, diskette, Hard Disk Drive (HDD) etc. In yet other embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g., via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate.EXAMPLES

[0182] Before describing various embodiments in detail, it is to be understood that this present specification is not limited to particular biological systems or reagents, for example, particular host cells, nucleotide sequences, polypeptides, or plasmid constructs, all which can, of course, vary. It is also to be understood that the terminology and particular reagents described in the example below is for the purpose of describing particular embodiments only, and is not intended to be limiting. One of ordinary skill in the art will recognize the variety of alternative reagents and protocols that can be used in place of those described herein, and where those substitutions remain within the scope and spirit of the claimed subject matter.Example 1Pol III promoters for expression of non-coding RNAs in CHO cells

[0183] To identify highly active RNA polymerase Ill-dependent promoters ("pol III promoters"), a set of ten DNA sequences were synthesized that could be transcribed to produce RNA reporters. Each reporter comprised a modified version of the human 7sk non-coding RNA, with insertion of two 24-base unique sequences in the middle of the molecule to serve as "barcodes" or unique identifiers, and an eight base poly thymidine pol III transcriptional terminator.

[0184] Each DNA sequence was operably linked to a pol III promoter. With reference to Table 8, the pol III promoters tested were human U68cc, human U6.1, human U69cc, water buffalo U6, mouse U6, pig 7sk, sheep 7sk, human 7sk, water buffalo 7sk and mouse Hl. These promoters are listed in Table 8, column A. The nucleotide sequence of each promoter is listed in Table 8, column B (SEQ ID NOs: 11-20). The nucleotide sequence of the reporter sequence operably linked to each promoter is listed in Table 8, column C (SEQ ID NOs: 1-10).Table 8

[0185] Each of the five U6 promoters operably linked to its reporter sequence were assembled into a single transposon (525865). The full nucleotide sequence of the polynucleotide comprising transposon 525865 is SEQ ID NO: 21. The four 7sk promoters and the mouse Hl promoter, each operably linked to its reporter sequence, were assembled into a single transposon (525867). The full nucleotide sequence of the polynucleotide comprising transposon 525867 is SEQ ID NO: 22. Each transposon comprised a left transposon end with nucleotide sequence SEQ ID NO: 23, a righttransposon end with nucleotide sequence SEQ ID NO: 24, and a transcriptional unit comprising an HSV-TK promoter operably linked to an open reading frame encoding puromycin-acetyl transferase and a polyadenylation signal such that the puromycin-acetyl transferase was constitutively expressible in a mammalian cell. Each transposon was separately transfected into a pool of suspension-adapted CHO-K1 cells, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence SEQ ID NO: 25. Into five million cells, 25 pg of transposon and 3 pg of transposase mRNA were electroporated; the cells were cultured in Advanced CHO Fed-batch media (Sigma) in the presence of 10 pg / ml puromycin until the cells recovered to >95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen. Cell pools were thawed, scaled up, and grown in a fed batch process in Advanced CHO Fed-batch media (Sigma) plus 4 mh / l glutamine for 14 days. On the 14th day, viability of each pool was above 90%, and the viable cell density was approximately 15 million cells per ml. Cell pellets containing approximately 10 million cells were obtained from each pool and frozen at -80°C.

[0186] RNA was extracted from each pellet using QIAGEN® RNeasy® Mini Kit (Cat. No. / ID:74104) according to the manufacturer's directions. RNA was quantified by absorbance at 260 nm, and approximately 3 pg from each pellet was combined and sequenced by direct RNA sequencing on an Oxford Nanopore Technologies ("ONT") Flongle cell according to the manufacturer's directions. A total of 67,000 reads were obtained, 31 of which comprised a barcode sequence. Column D of Table 8 shows the percentage of these ONT reads associated with each reporter sequence. Because the number of bar-code containing reads obtained was relatively low, a reverse transcriptase PCR amplification method was also used to measure transcript frequencies. The RNA pool was amplified with primers with nucleotide sequences SEQ ID NOs: 26 and 27 using New England Biolabs OneTaq One-Step RT-PCR Kit (catalog no. E5315S) according to the manufacturer's directions. The amplicon was cloned into a cloning vector, transformed into E. coli, and 400 independent colonies were picked and the cloned DNA sequenced. A total of 394 reads were obtained comprising a barcode sequence. Column E of Table 8 shows the percentage of these PCR-derived reads associated with each reporter sequence. Column F of Table 8 shows the average percentage of barcode containing reads identified by the two different methods.

[0187] As shown in Table 8, the results of the two different sequencing methods were highly correlated. The most active pol III promoter tested was the pig 7sk promoter, which was 3-fold more active in CHO-K1 cells than the human U68cc, the water buffalo U6, the mouse Hl, and the human 7sk promoter sequences. These promoters were all significantly more active in CHO-K1 cells than the mouse U6, the sheep 7sk, the human U6.1, the human U69cc, and the water buffalo 7sk promoters. It is thus most advantageous to use a pig 7sk promoter sequence operably linked to a non-coding RNA, such as a suppressor tRNA, to obtain high levels of expression of that noncoding RNA in CHO-K1 cells.

[0188] This data also shows that different 7sk promoters from different animals have very different activities in CHO-K1 cells. The pig 7sk promoter is more than 30-fold more active than the sheep 7sk promoter, more than 9-fold more active than the water buffalo 7sk promoter, and more than 3-fold more active than the human 7sk promoter in CHO-K1 cells. This is unexpected and could not be predicted from the sequences or the animal origins of the different 7sk promoters alone.Example 2Pol III promoter combinations for expression of non-coding RNAs in mammalian cells

[0189] Although the pig 7sk promoter (for example, a promoter comprising the nucleotide sequence SEQ ID NO: 16) is the most active pol III promoter tested for CHO-K1 cells, a polynucleotide comprising multiple adjacent copies of that promoter operably linked to a suppressor tRNA may suffer from genetic instability if integrated into the genome of a mammalian cell such as a CHO-K1 cell. This instability can be reduced by instead using a set of different pol III promoters within a suppressor tRNA-expressing polynucleotide, and incorporating the set of pol III promoters, each operably linked to a suppressor tRNA, onto a transposon to enable multiple independent integrations.

[0190] Based on the results from Example 1 and shown in Table 8, an advantageous transposon for expression of a non-coding RNA such as a suppressor tRNA may comprise one or more of the following pol III promoters: the pig 7sk promoter, the human U68cc promoter, the water buffalo U6 promoter, the mouse Hl promoter, and the human 7sk promoter. An example transposon (560742) comprised a left transposon end with nucleotide sequence SEQ ID NO: 23 and a right transposon end with nucleotide sequence SEQ ID NO: 24. Between these two transposon ends the example transposon further comprised a transcriptional unit comprising anopen reading frame encoding a selectable marker (glutamine synthetase) operably linked to an HSV-TK promoter and an SV40 polyadenylation signal such that the glutamine synthetase was constitutively expressible in a mammalian cell. The example transposon further comprised:(i) a pig 7sk promoter with nucleotide sequence SEQ ID NO: 16 operably linked to a DNA sequence encoding a pyrrolysine tRNA with nucleotide sequence SEQ ID NO: 28, followed by a pol III transcriptional terminator sequence with nucleotide sequence SEQ ID NO: 29;(ii) a human U68cc promoter with nucleotide sequence SEQ ID NO: 11 operably linked to a DNA sequence encoding a pyrrolysine tRNA with nucleotide sequence SEQ ID NO: 28, followed by a pol III transcriptional terminator sequence with nucleotide sequence SEQ ID NO: 29;(iii) a water buffalo U6 promoter with nucleotide sequence SEQ ID NO: 14 operably linked to a DNA sequence encoding a pyrrolysine tRNA with nucleotide sequence Seq ID NO: 28, followed by a pol III transcriptional terminator sequence with nucleotide sequence SEQ ID NO: 29;(iv) a mouse Hl promoter with nucleotide sequence SEQ ID NO: 20 operably linked to a DNA sequence encoding a pyrrolysine tRNA with nucleotide sequence SEQ ID NO: 28, followed by a pol III transcriptional terminator sequence with nucleotide sequence SEQ ID NO: 29; and(v) a human 7sk promoter with nucleotide sequence SEQ ID NO: 18 operably linked to a DNA sequence encoding a pyrrolysine tRNA with nucleotide sequence SEQ ID NO: 28, followed by a pol III transcriptional terminator sequence with nucleotide sequence SEQ ID NO: 29. The full nucleotide sequence of the polynucleotide comprising transposon 560742 is SEQ ID NO: 30.The pyrrolysine tRNA naturally has an anticodon that reads the amber stop UAG.

[0191] Transposon 560742 was transfected into a pool of suspension-adapted CHO-K1 cells engineered to be deficient in the expression of glutamine synthetase, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence of SEQ ID NO: 25. Into five million cells, 25 pg of transposon and 3 pg of transposase mRNA were electroporated; the cells were cultured in Advanced CHO Fed-batchmedia (Sigma) in the absence of glutamine until the cells recovered to greater than 95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen.

[0192] The average number of integrated copies of transposon 560742 was assessed by digital droplet PCR (ddPCR) using a BioRad QX200 AutoDG Droplet Digital PCR System according to the manufacturer's directions. The forward primer had nucleotide sequence SEQ ID NO: 33, and the reverse primer had nucleotide sequence SEQ ID NO: 34. The probe sequence had nucleotide sequence SEQ ID NO: 35 and was purchased from IDT Technologies. The probe had a 5' FAM fluorophore, a 3' Iowa Black quencher, and a ZEN middle quencher nine nucleotides from the 5' end. The dd PCR indicated that the cells in the selected pool comprised on average 28 copies of transposon 560742 per cell genome. This represents approximately 140 copies of the suppressor tRNA encoding DNA per genome, with each copy being operably linked to a promoter demonstrated to be active in a CHO-K1 cell.

[0193] Because the pig 7sk promoter is so much more active in CHO-K1 cells than the other promoters tested, it may be suitable to have more than one copy of the pig 7sk promoter operably linked to the DNA sequence encoding the suppressor tRNA. On this basis, a transposon was constructed for expression of suppressor tRNAs that was a variant of transposon 560742. The new transposon (569393) differed from transposon 560742 in the replacement of the water buffalo U6 promoter with nucleotide sequence SEQ ID NO: 14 by a modified water buffalo U6 promoter with nucleotide sequence SEQ ID NO: 31. Transposon 569393 also differed from transposon 560742 in the replacement of the human 7sk promoter with nucleotide sequence SEQ ID NO:18 by a second copy of the pig 7sk promoter with nucleotide sequence SEQ ID NO: 16. The full nucleotide sequence of the polynucleotide comprising transposon 569393 is SEQ ID NO: 32.Example 3Expression of azido-functionalized GFP in stable CHO-K1 pools

[0194] Transposon 560742 comprised five copies of the DNA encoding Methanosarcina mazei pyrrolysine tRNA, each copy operably linked to a different pol III promoter active in CHO-K1 cells. As described in Example 2, this transposon could be introduced into CHO-K1 cells to provide approximately 140 copies of the tRNA per cell. The pyrrolysine tRNA naturally recognizes a UAG (amber) stop codon.

[0195] To test whether this tRNA-carrying transposon could enable incorporation of a UAA into a protein, transposons were constructed comprising a cognate Methanosarcina mazeipyrrolysine aminoacyl tRNA synthetase and an open reading frame encoding a GFP interrupted by a single amber stop codon. The relative expression of the tRNA and its cognate aminoacyl tRNA synthetase were varied by using different strength promoters operably linked to the open reading frame encoding the aminoacyl tRNA synthetase, and by transfecting the tRNA transposon and the aminoacyl tRNA synthetase / GFP transposon into CHO-K1 cells at different ratios. Each aminoacyl tRNA synthetase transposon comprised a left transposon end with nucleotide sequence SEQ ID NO: 23 and a right transposon end with nucleotide sequence SEQ ID NO: 24. Between these two transposon ends each transposon further comprised a first transcriptional unit comprising an open reading frame encoding a selectable marker (glutamine synthetase) operably linked to an HSV-TK promoter and an SV40 polyadenylation signal such that the glutamine synthetase was constitutively expressible in a mammalian cell. Each transposon further comprised a second transcriptional unit comprising a first open reading frame encoding Methanosarcina mazei pyrrolysine aminoacyl tRNA synthetase (with amino acid sequence SEQ ID NO: 36) followed by an EMCV IRES operably linked to an open reading frame encoding GFP interrupted by a single amber stop codon within the coding region and operably linked to a rabbit globin polyadenylation signal at the 3' end. The GFP amino acid sequence is provided in SEQ ID NO: 37, and its nucleotide sequence is provided in SEQ ID NO:38. Previous data has shown that truncation (translation stop) at this codon yields no fluorescent protein. Suppression of the amber stop by incorporation of the UAA leads to full-length fluorescent GFP. GFP fluorescence can be assessed by methods standard in the art such as fluorescence activated cell sorting (FACS) or fluorescent cell counting. In transposon 560796, the second transcriptional unit comprised a human EFl promoter. In transposon 560795, the second transcriptional unit comprised a PGK promoter. In transposon 560794, the second transcriptional unit comprised an EEF2 promoter. The full nucleotide sequences of the polynucleotides comprising transposons 560796, 560795, and 560794 are provided in SEQ ID NOs: 39, 40, and 41, respectively.

[0196] Transposon 560742 was co-transfected with either transposon 560796, 560795, or 560794 into a pool of suspension-adapted CHO-K1 cells engineered to be deficient in the expression of glutamine synthetase, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence SEQ ID NO: 25. Into five million cells, a total of 25 pg of transposon DNA and 3 pg of transposase mRNA were electroporated. The ratio between the tRNA transposon 560742 and the aminoacyl tRNAsynthetase / GFP transposon was either 4:1 (i.e., 20 pg of tRNA transposon and 5 pg aminoacyl tRNA synthetase / GFP transposon) or 3:2 (i.e., 15 pg of tRNA transposon and 10 pg aminoacyl tRNA synthetase / GFP transposon). The cells were cultured in Advanced CHO Fed-batch media (Sigma) in the absence of glutamine until the cells recovered to >95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen.

[0197] Cryopreserved cell pools were thawed and cultured in Advanced CHO Fed-batch media (Sigma). Once the cells had recovered from thaw for two passages, the media was supplemented with 0.5 mM N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine (Azido-Lys). GFP production was monitored using a fluorescence cell counter (Countess) 72 hours post-addition of Azido-Lys.Relative fluorescence is shown in FIG. 1. Fluorescence was only observed in cells after treatment with Azido-Lys and was observed in all combinations. The best productivity was seen with tRNA transposon 560742 and aminoacyl tRNA synthetase / GFP transposon 560794 transfected at a ratio of 4:1.Example 4Expression of azido-functionalized antibody in stable CHO-K1 pools

[0198] Two pools were selected from the transfections described in Example 3. Pool 1 originated in the co-transfection of 20 pg tRNA transposon 560742 and 5 pg aminoacyl tRNA synthetase / GFP transposon 560794 (comprising the EEF2 promoter). Pool 2 originated in the cotransfection of 20 pg tRNA transposon 560742 and 5 pg aminoacyl tRNA synthetase / GFP transposon 560795 (comprising the PGK promoter).

[0199] Each of these pools was transfected with a third transposon, 560907. Transposon 560907 comprised a first transposon end with nucleotide sequence SEQ ID NO: 42 and a second transposon end with nucleotide sequence SEQ ID NO: 43. Between these transposon ends the transposon further comprised three transcriptional units. The first transcriptional unit comprised an open reading frame encoding puromycin acetyl transferase operably linked to an HSV-TK promoter and an SV40 polyadenylation signal such that the puromycin acetyl transferase was constitutively expressible in a mammalian cell. The second transcriptional unit comprised an open reading frame encoding an antibody light chain with mature amino acid sequence SEQ ID NO: 45 operably linked to a CMV promoter and an SV40 polyadenylation signal such that the antibody light chain was constitutively expressible in a mammalian cell. The third transcriptional unit comprised an open reading frame encoding an antibody heavy chain interrupted by a single amber(UAG) stop codon. The antibody heavy chain comprised mature amino acid sequence SEQ ID NO: 46 and nucleotide sequence SEQ ID NO: 47 and was operably linked to a CMV promoter and an SV40 polyadenylation signal such that the antibody heavy chain was constitutively expressible in a mammalian cell. These heavy and light chain sequences are from the therapeutic monoclonal antibody trastuzumab. This heavy chain sequence contains an engineered amber codon at position 118 (EU numbering). The full nucleotide sequence of the polynucleotide comprising transposon 560907 is provided in SEQ ID NO:48.

[0200] Transposon 560907 was transfected into either Pool 1 or Pool 2, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence of SEQ ID NO: 44. Into five million cells, a total of 25 pg of transposon DNA and 3 pg of transposase mRNA were electroporated. The cells were then cultured in Advanced CHO Fed-batch media (Sigma) in the absence of glutamine and with the addition of 10 pg / ml puromycin until the cells recovered to >95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen. Pool lb resulted from the transfection of transposon 560907 into cell pool 1. Pool 2b resulted from the transfection of transposon 560907 into cell pool 2.

[0201] Cryopreserved pools were thawed and cultured in Advanced CHO Fed-batch media (Sigma). Once the cells had recovered from thaw they were scaled up and grown in a 10 ml fed batch process. At day 5 of the fed batch, the media was supplemented with 0.5 mM Azido-Lys, and the culture temperature was shifted from 37°C to 32°C. A parallel fed batch for each pool was performed without addition of Azido-Lys. At day 10 of the fed batch, samples from each fed batch were clarified, and the supernatants were run on a non-reducing SDS PAGE gel, which is shown in FIG. 2. With reference to FIG. 2, in the left lane (1) are molecular weight markers, with their weights indicated (in kDa). In the next lane (2) is a control full length antibody, showing a single band at a little over 150 kDa. The fed batch supernatants from Pool 1 (lanes 3 and 4) and Pool 2 (lanes 5 and 6) are indicated above the gel, as is the presence of Azido-Lys in the culture (lanes 3 and 5).

[0202] In all samples, the light chain is clearly visible at ~ 20 kDa (indicated in FIG. 2 as "L"), with light chain dimer visible at ~40 kDa (indicated in FIG. 2 as "LD"). Full length antibody ("Full" in FIG. 2) is clearly visible in the supernatant from both Pool 1 and Pool 2, but only in cultures grown in the presence of Azido-Lys. The amount of antibody was quantified using biolayerinterferometry using protein-A derivatized tips (Forte Bio). Pool 1 supernatant contained 654 mg / L antibody. Pool 2 supernatant contained 584 mg / L antibody.

[0203] To confirm that the antibody produced contained an azido-functionalized amino acid in the heavy chain, antibody from the culture supernatants was purified by affinity chromatography with Protein A. The purified antibody was coupled with fluorescent dye DBCO-AF647 (Jenna), which specifically reacts with azido groups. Fluorescence visualized on reduced denaturing PAGE showed specific labeling of the antibody heavy chain (FIGS. 3A and 3B). Lane 1 of FIG. 3A shows molecular weight markers, with their weights indicated in kDa. Immediately above the gel images are + / - indicators showing whether the sample was coupled with DBCO-AF647. In the row above that are indicators showing whether ("+UAA") or not ("No") Azido-Lys was included in the fed batch culture. In the row above that are indicators showing whether the antibody was derived from Pool 1 or Pool 2. Also included was a control ("Ctrl") antibody, which had the same mature light chain sequence, and whose heavy chain sequence was identical except that it lacked the amber stop codon. The mature amino acid sequence of the heavy chain of the control antibody is shown in SEQ ID NO: 49. FIG. 3A shows total protein stained with coomassie blue. FIG. 3B shows fluorescence of the protein. FIG. 3A, lanes 2, 3, 6, and 7 show that in the absence of Azido-Lys in the fed batch culture media, no protein A-binding antibody was produced, as expected for a heavy chain with a stop codon incorporated prior to the Fc domain. FIG. 3A, lanes 4, 5, 8, and 9 show that when Azido-Lys was included in the fed batch culture media, protein A-binding antibody could be purified, with both heavy and light chains (indicated as H and L, respectively). The sizes of these chains were the same as the sizes of the chains of the control antibody (FIG. 3A, lanes 10 and 11) produced in the absence of Azido-Lys. FIG. 3B shows that of these proteins, only the heavy chain from Pool 1 (FIG. 3B, lane 14) and Pool 2 (FIG. 3B, lane 18) had coupled with the azido group-reactive fluorescent dye DBCO-AF647. No fluorescence was observed in the control antibody treated with DBCO-AF647 (FIG. 3B, lane 20).

[0204] Based on these results, the tRNA transposon 560742 comprising five copies of the DNA encoding Methanosarcina mazei pyrrolysine tRNA, each copy operably linked to a different pol III promoter active in CHO-K1 cells including at least one pig 7sk promoter provided enough suppressor tRNA to incorporate the unnatural amino acid Azido-Lys into an antibody to produce over 0.5 g / L of the antibody in a 5 day fed batch culture. LC-MS analysis of the purified antibody confirmed a molecular weight expected for the heavy chain with a single Azido-Lys incorporated.Example 5Increased expression of azido-functionalized antibody in stable CHO-K1 pools

[0205] Pools 1 and 2 described in Example 4 expressed open reading frames for a GFP and an antibody heavy chain, both of which were interrupted by an amber stop (UAG). Having demonstrated that these pools could produce over 0.5 g / L of UAA substituted antibody, an attempt was made to increase this yield by removing the interrupted GFP open reading frame. Modified versions of transposons 560794 and 560795 were constructed for expression of the aminoacyl tRNA synthetase. Transposon 560755 was identical to transposon 560795 (with the PGK promoter operably linked to the open reading frame encoding the aminoacyl tRNA synthetase) except that it lacked the IRES and interrupted GFP open reading frame. The full nucleotide sequence of the polynucleotide comprising transposon 560755 is provided in SEQ ID NO: 50. Transposon 560754 was identical to transposon 560794 (with the EEF2 promoter operably linked to the open reading frame encoding the aminoacyl tRNA synthetase) except that it lacked the IRES and interrupted GFP open reading frame. The full nucleotide sequence of the polynucleotide comprising transposon 560754 is provided in SEQ ID NO: 51.

[0206] Transposon 560742 (the tRNA transposon) was co-transfected with either transposon 560755 or 560754 (the aminoacyl tRNA synthetase transposons) into a pool of suspension-adapted CHO-K1 cells engineered to be deficient in the expression of glutamine synthetase, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence of SEQ ID NO: 25. Into five million cells, a total of 25 pg of transposon DNA and 3 pg of transposase mRNA were electroporated. The ratio between the tRNA transposon 560742 and the aminoacyl tRNA synthetase transposon was 4:1 (i.e., 20 pg of tRNA transposon and 5 pg aminoacyl tRNA synthetase / GFP transposon). The cells were cultured in Advanced CHO Fed-batch media (Sigma) in the absence of glutamine until the cells recovered to >95% viability, at which time banks of these selected cell pools were cryopreserved over liquid nitrogen. Pool 3a comprised transposons 560742 and 560754. Pool 4a comprised transposons 560742 and 560755.

[0207] Transposon 560907 (as described in Example 4) was transfected into either Pool 3a or Pool 4a, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence of SEQ ID NO:44. Into five million cells, a total of 25 pg of transposon DNA and 3 pg of transposase mRNA were electroporated. The cells werecultured in Advanced CHO Fed-batch media (Sigma) in the absence of glutamine and with the addition of 10 pg / ml puromycin until the cells recovered to >95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen. Pool 3 resulted from the transfection of transposon 560907 into cell pool 3a. Pool 4 resulted from the transfection of transposon 560907 into cell pool 4a.

[0208] Cryopreserved cell pools were thawed and cultured in Advanced CHO Fed-batch media (Sigma). Once the cells had recovered from thawing, they were scaled up and grown in a 10 ml fed batch process. At day 5 of the fed batch, the media was supplemented with either 0.1 or 0.5 mM Azido-Lys, and the culture temperature was shifted from 37°C to 32°C. Viability of the cells in 0.5 mM Azido-Lys fell to below 80% after day ten, but the viability of the cells in 0.1 mM Azido-Lys was maintained above 80% until day 12. At days ten and 12 of the fed batch, samples from each viable fed batch were clarified, and the concentration of antibody was measured by biolayer interferometry using protein-A derivatized tips (Forte Bio). The concentration of antibody in the culture supernatants is shown in FIG.4. FIG. 4 shows that, as in Example 4, the productivity of the cell pools was greater for the pool of cells whose genomes comprised the aminoacyl tRNA synthetase open reading frame operably linked to the EEF2 promoter (pool 3 in this Example 5, pool 1 in Example 4) (these are shown as open bars in FIG.4). The filled bars show the antibody titer in cells from pool 4. Productivity was also higher in the cultures in media containing 0.1 mM Azido-Lys than those with 0.5 mM Azido-Lys (the concentration of Azido-Lys in the culture is indicated below the bars). The day on which cultures were sampled to measure antibody concentration in the supernatant is indicated below the bars. At day 12, the pool 4 titer was about 750 mg / L, and the pool 3 titer was about 950 mg / L.Example 6Boosting expression of azido-functionalized antibody in stable CHO-K1 pools

[0209] It was observed, based on densitometry analysis of the gel shown in FIG. 2, there is considerable molar excess of light chain relative to heavy chain expressed in the culture supernatant from pools 1 and 2 treated with 0.5 mM Azido-Lys (see the band corresponding to L in FIG. 2, lanes 3 and 5). Although some excess of light chain is necessary to ensure proper folding of the heavy chain, the incorporation of a UAA into the heavy chain may limit its production.Therefore, it was tested whether boosting expression of the heavy chain relative to the light chain might provide a further boost to the production of antibody incorporating a UAA.

[0210] Into pool 3a (made as described in Example 5), transposon 568054 (a modified version of transposon 560907, configured to increase the expression of the heavy chain relative to the light chain) was transfected, together with mRNA encoding a corresponding transposase fused to a heterologous nuclear localization signal with amino acid sequence of SEQ ID NO:44. Into five million cells, a total of 25 pg of transposon DNA and 3 pg of transposase mRNA were electroporated. The cells were cultured in Advanced CHO Fed-batch media (Sigma) in the absence of glutamine and with the addition of 10 pg / ml puromycin until the cells recovered to >95% viability, at which time banks of selected cell pools were cryopreserved over liquid nitrogen. Pool 3b resulted from the transfection of transposon 568054 into cell pool 3a.

[0211] Cryopreserved cell pool 3b was thawed and cultured in Advanced CHO Fed-batch media (Sigma). Once the cells had recovered from thawing, they were scaled up and grown in a 10 mL fed batch process. At day five of the fed batch, the media was supplemented with 0.1 mM Azido-Lys, and the culture temperature was shifted from 37°Cto 32°C. At days 7, 10, 12, and 14 of the fed batch, samples were clarified, and the concentration of antibody was measured by biolayer interferometry using protein-A derivatized tips (Forte Bio). The concentration of antibody in the culture supernatants is shown in FIG. 5. By day 14, the titer of antibody with site-specific incorporation of a UAA (Azido-Lys) had reached about 1.6 g / L.

[0212] Based on these results, a CHO-K1 cell whose genome comprises the following elements is useful for site-specific incorporation of UAAs into an antibody:(i) a first transposon comprising multiple copies of the DNA encoding pyrrolysine tRNA, wherein one, and optionally two or more, copies of the DNA encoding pyrrolysine tRNA is or are operably linked to a pig 7sk promoter, and each of the remaining copies of the DNA encoding pyrrolysine tRNA are each operably linked to a pol III promoter active in CHO-K1 cells other than a pig 7sk promoter and different from each other;(ii) a second transposon comprising an open reading frame encoding a corresponding aminoacyl tRNA synthetase operably linked to regulatory elements such that the aminoacyl tRNA synthetase is expressible in the CHO-K1 cell; and(iii) a third transposon comprising open reading frames encoding the heavy and light chains of an antibody, wherein the open reading frame encoding one chain comprises an amber stop codon (UAG) preventing expression of a full antibody unless the amber stop codon is read by the pyrrolysine tRNA to incorporate a UAA,The first, second, and third transposons may be separate, or they may be combined into two or one transposon. The first, second and third transposon may be transposable by the same transposase or by different transposases.Example 7In silico discovery of pyrrolysine tRNA synthetases and associated tRNAs

[0213] An effort was undertaken to develop new orthogonal translation system components for the in vivo incorporation of unnatural amino acids into proteins in vivo, where the UAA have reactive side chain chemistries. These proteins containing the UAA can then undergo biocompatible side chain conjugations or modifications, in a manner that preserves the protein's primary, secondary and tertiary structure.

[0214] Multiple criteria were applied to the in silico screening for natural genomic sequences in bacteria (i.e., Eubacteria) and Archaea (synonymous with Archaebacteria) to identify aminoacyl-tRNA synthetase enzymes (RS) and their cognate tRNA sequences that could function in an orthogonal manner to mammalian cell tRNA charging. Among these in silico criteria was to identify proteins showing homology to known aminoacyl-tRNA synthetase enzymes that have the ability to charge a cognate tRNA with the non-canonical amino acid pyrrolysine. Some pyrrolysyl-tRNA synthetases have been shown to charge suppressor tRNA molecules with lysine derivative unnatural amino acids that contain side chain moieties enabling biocompatible reaction chemistries. It was hoped that exploration of homologues of such enzymes would discover synthetases with broadened or new UAA specificities, improved incorporation rates or efficiency in vivo, higher orthogonality to host translation, and / or lower toxicity to host cells leading to higher productivity of UAA-incorporated target proteins.

[0215] Proteins that show both structural homology to known pyrrolysyl-tRNA synthetases and, preferably, were associated with gene clusters containing suppressor tRNA genes were identified. Among those, a set of twenty-two (22) pairs spanning wide novel structural diversity were chosen for functional analysis. Of these original 22 synthetase-tRNA cognate pairs, five were of bacterial origin, and 17 were of Archaeal origin. These 22 in silico RS targets and the cognate tRNA sequences are listed in the table below. For one of the synthetases, RPJ77547.1, from Alphaproteobacteria bacterium, a tRNA from the same host genome was not tested and that synthetase was paired with a predicted pyrrolysine tRNA from a Methanohalophilus species.

[0216] As used herein, the synthetase proteins are named according to their corresponding NCBI GenBank accession numbers. The tRNA molecules are named according to the NCBI GenBank accession numbers of their corresponding synthetase. The amino acid and nucleotide sequences of these synthetase enzymes and their cognate tRNA molecules are provided in FIG. 15.Table 9

[0217] Genes for the 22 synthetases were designed for high expression in mammalian cells and synthesized along with those of their associated tRNAs. Expression of the synthetase genes was accomplished by cloning these synthetic genes into a mammalian expression vector under the control of a PGK promoter. Each of these expression vectors also contained genes for the expression of the heavy and light chains of the antibody trastuzumab, where the heavy chain gene includes an engineered amber codon (TAG) at position 118 (EU numbering). Genes for each tRNA were cloned into an expression vector such that each of two copies of the tRNA was expressed from distinct mammalian pol III promoters, a pig 7sk promoter and a human U68cc promoter.Example 8Screening of synthetases for utilization of unnatural amino acids in suppressor tRNA charging

[0218] The synthetase vectors containing the trastuzumab reporter genes and the tRNA expression vectors of Example 7 were transfected into HEK293 suspension cells by lipofection. For comparison, cells were also transfected with either wildtype Methanosarcina mazei pyrrolysyl (pyl) synthetase (WP_011033391.1) or a double mutant (Y306A; Y384F) of that enzyme along with M. mazei tRNAPyl(T25C). Previous studies have shown that this double mutation broadens the specificity of the M. mazei synthetase to accept larger UAA. Cultures were incubated at 37°C in a humidified atmosphere of 8% CO2.

[0219] Cultures were augmented with unnatural amino acids, which were either 0.1 mM N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine ("2AzZ-Lys") or 0.25 mM N6-((2-azidoethoxy)carbonyl)-L-lysine ("Azc-Lys" or alternatively "azido-lysine" or "AzK") one day after transfection. Cell viability was monitored over time using trypan blue cell staining on a Vi-CELL XR, Antibody secretion to the culture supernatant was monitored by Biolayer interferometry using probes specific for the intact antibody Fc portion. Expression of the fully assembled trastuzumab IgG was confirmed by nonreducing polyacrylamide gel electrophoresis and LC-MS. For some cultures, incorporation of the UAA was confirmed by LC-MS.

[0220] For activity screening by biolayer interferometry (BLI), positive activity was defined as significant increase in trastuzumab production relative to cultures not treated with UAA. The limit of quantitation of the trastuzumab assay was 1 pg / mL. The average BLI signal for the M mazei double mutant synthetase (Y306A; Y384F) with M. mazei tRNAPyl(T25C) in the transient HEK293 system in the absence of added UAA corresponded to titers of 3.2±0.7 pg / mL at day 4 and 4.1±1.3 pg / mL at day 7. For purposes of screening for active orthogonal translation system components, measured antibody titers below 6 pg / mL at day 7 are considered insignificant with regard to the qualification of synthetase-tRNA pair systems as active.

[0221] Several synthetase and tRNA pairs from both Archaeal and bacterial systems showed significant activity to incorporate Azc-Lys in the mammalian host cell based on high trastuzumab titers observed from cells transfected with those pairs, as shown in FIG. 7.

[0222] At the 7 day timepoint, 15 of the 22 synthetase-tRNA pairs tested showed detectable activity (>6 ug / ml trastuzumab titer). Highest productivities were observed for HHW30035.1, MBP2030157.1, RNI15967.1, NPE30940.1, MDK2891647.1, TCL11664.1, and SMH34004.1.Moderate productivities were seen for NOQ48807.1, NYT18799.1, AKB84338.1, ADE35815.1 and RPJ77547.1. Low but detectable productivities were seen for MBP1909574.1, OBZ34613.1, ADI73236.1, and MBN2322174.1. None of the 21 synthetase-tRNA pairs showed significant activity with 2AzZ-Lys. The M. mazei RS double mutant showed clear activity with both UAAs tested.Example 9Screening of synthetase mutants for altered unnatural amino acid specificity

[0223] Synthetases showing significant activity toward Azc-Lys were analyzed with phylogenetic comparison, structure prediction and molecular interaction simulation software toolsto identify amino acid substitutions that may alter the binding capacity with respect to the following unnatural amino acids, which are depicted in FIG. 6:Azc-Lys; N6-((2-azidoethoxy)carbonyl)-L-lysine2AzZ-Lys; N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysineSpHD-Lys; N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysineNrb-Lys; N6-(norbornene-methoxycarbonyl)-L-lysineTCO-Lys; N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine

[0224] Based on these analyses, substitutions were made at one or two positions in certain synthetases. These mutant synthetases were then cloned and transfected into HEK293 cells with the same tRNAs used for the wt synthetases as described in Example 8. Each mutant was then assayed for their ability to support production of the trastuzumab antibody in the presence of the UAA 2AzZ-Lys, which was not accepted by any of the wt synthetases (FIG. 7). Each culture was treated with 0.1 mM of 2AzZ-Lys 24 hours after transfection and trastuzumab production was monitored by BLI. Several of the mutant synthetases gained the ability to utilize 2AzZ-Lys (FIG. 12). Significant increases in production with 2AzZ-Lys were seen for MDK2891647.1(Y257A;Y335F), RPJ77547.1(Y269A;Y347F), HHW30035.1(Y127A), and HHW30035.1(Y127A;Y208F). Low production was seen for MBP2030157.1(Y259A;Y337F), RNI15967.1(Y255A;Y333F),TCL11664.1(Y255A;Y333F), and MBP2030157.1(Y259A).

[0225] Because of the promising activity seen with the HHW30035.1 synthetase and mutants thereof described above, additional substitutions at position 127 were explored and the resulting synthetases were tested against a range of UAAs. These substitutions at position Y127 were A, G, M, L, V, S, I and K. Genes for each of the resulting mutant synthetases were cloned and transfected into HEK293 cells with the same tRNAs used for the wt synthetase as described in Example 8. Each mutant was then assayed for their ability to support production of the trastuzumab antibody incorporating the UAA that was included in the cell culture. Each culture was treated with 0.1 mM of UAA 24 hours after transfection.

[0226] The mutant synthetases tested showed various activities with the panel of amino acids when transfected into HEK293 host cells as shown in FIG.8. Substitutions at Y127 to A,G,V and S showed improved titer with at least one of the UAA. In this assay, the combination wild type HHW30035.1 synthetase and its cognate tRNA showed strong incorporation of only Azc-Lys, but none of the other UAA tested. In contrast, the Y127A and Y127G mutants showed improvedincorporation of the full UAA panel, including the 2AzZ-Lys, Nrb-Lys, SpHD-Lys and TCO-Lys. The Y127V and Y127S mutants showed improved incorporation of SpHD-Lys. The Y127G substitution showed particularly strong activity with Nrb-Lys and SpHD-Lys, and somewhat lower activity with the other three UAA tested relative to the Y127A mutant, indicating a distinct substrate specificity profile. The other substitutions at Y127 gave generally lower activity for the panel indicating an inability to accommodate these UAA or a substrate-independent loss of specific activity.

[0227] Using similar transient transfection methods, the HHW30035.1_Y127A synthetase was introduced along with its cognate tRNA into a second cell type, namely CHO-K1 cells. For comparison, cells were also transfected with the Y306A mutant of the M. mazei synthetase and M. mazei tRNAPyl(T25C). These orthogonal paris were:(i) mut Archaea M. mazei pyrrolysyl-RS(Y306A) with M. mazei tRNAPyl(T25C);(ii) mut bacterial HHW30035 pyrrolysyl-RS mutant enzyme (Y127A) with the natural cognate HHW30035 tRNAPyl;

[0228] Cells were treated with UAA 24 hours after transfection. Significant trastuzumab production in the presence of either 0.1 mM Azc-Lys or 0.1 mM 2AzZ-Lys UAAs was seen for both synthetase-tRNA pairs (FIG. 13).Example 10Identification of active homologs of HHW30035.1

[0229] In an effort to identify still further preferred orthogonal translation components, proteins that both show structural homology to the HHW30035.1 tRNA synthetase and were associated with operons containing suppressor tRNA genes were identified in silica. Among those, a set of six additional bacterial synthetase-tRNA pairs were chosen for functional analysis. The sequences of these additional bacterial synthetase-tRNA pairs are provided in FIG. 15, which were:Table 10

[0230] Genes for each synthetase were synthesized and subcloned into vectors designed for high expression in mammalian cells. Genes for their associated cognate tRNAs were also synthesized and cloned into suitable vectors. The cloning, vectors and promoters used are as described in Example 7.

[0231] For each of these synthetases, an additional gene was also made containing a mutation with a substitution of alanine for the conserved tyrosine at the position structurally equivalent to 127 in HHW30035.1 based on predicted protein folding using AlphaFold software. These mutant sequences are also provided in FIG. 15. Synthetic genes for each of the tRNA synthetase mutants and their cognate tRNAs were cloned as described above for testing in transient HEK293 cell culture.

[0232] The synthetase and tRNA expression vectors were used to transfect HEK293 cells as described above in Example 8. Cultures were augmented alternatively with either 0.1 mM N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine ("2AzZ-Lys") or 0.25 mM N6-((2-azidoethoxy)carbonyl)-L-lysine ("Azc-Lys") one day after transfection. Cell viability and culture production of trastuzumab were monitored over time, as described above in Example 8. Antibody secretion to the culture supernatant was monitored by Biolayer interferometry using probes specific for the intact antibody Fc portion. Expression of the fully assembled trastuzumab IgG was confirmed by nonreducing polyacrylamide gel electrophoresis and LC-MS. Incorporation of the UAA was confirmed by LC-MS.

[0233] Titer of the antibody incorporating the UAA was measured and the results shown in FIG. 9. Using UAA Azc-Lys, three of the six wild-type synthetase-tRNA pairs identified in silico showed similar or greater antibody yield as compared to the yield observed when using the HHW30035.1(Y127A) pair. These were HHW28785, MDD4402188, and WP_044666330.

[0234] Using UAA 2AzZ-Lys, none of the wildtype synthetase sequences tested showed any significant production of antibody incorporating that UAA.

[0235] Using 2AzZ-Lys, four of the six mutant synthetase-tRNA pairs showed similar or greater antibody yield as compared to HHW30035.1(Y127A). These four synthetase mutants were:HHW28758(Y127A), MDD4402188(Y127A), WP_044666330(Y127A) and WP_206813094(Y127A).Example 11Productivity with cross-species synthetase-tRNA pairs

[0236] As a test of mutual orthogonality between the Archaea Methanosarcina mazei and bacterial HHW30035.1 synthetase-tRNA pairs, UAA incorporation during antibody production was assayed in HEK293 cells transiently transfected with alternative combinations of each of the two synthetases along with each of the two tRNA molecules. The methodology used was as described above in Example 8.

[0237] These results are shown in FIG. 10A. A total of five combinations were tested for the ability to incorporate 2AzZ-Lys UAA. The far left data point shows the antibody titer when using the wild type Archaea M. mazei pyrrolysyl-synthetase with M. mazei tRNAPyl(T25C), which showed no activity for the incorporation of 2AzZ-Lys. The second lane showed this same pairing, but substituted the wild type M. mazei synthetase with a double mutant M. mazei pyrrolysyl-synthetase (Y306A; Y384F). This pairing showed significant antibody production.

[0238] The third lane shows antibody production titer when the double mutant M. mazei pyrrolysyl-synthetase (Y306A; Y384F) is coupled with the Syntrophomonadaceae bacterial HHW30035 tRNAPyl. Surprisingly, significantly higher production was observed for this combination as compared to the same enzyme using the M. mazei tRNAPyl.

[0239] The two right lanes in FIG. 10A used the bacterial HHW30035 pyrrolysyl-synthetase mutant enzyme (Y127A) with either the natural cognate HHW30035 tRNAPylor the M. mazei tRNAPyl. Both of these pairings showed significant UAA incorporation, with the cognate pair HHW30035 pyrrolysyl-synthetase mutant enzyme (Y127A) / tRNAPylshowing significantly higheractivity as compared to the hybrid cross species combination of the bacterial synthetase and the Archaea

[0240] A similar experiment was also conducted using CHO-K1 cells. See FIG. 13. In the experiment, the M. mazei pyrrolysyl-synthetase (Y306A) was co-transfected with the Syntrophomonadaceae bacterial HHW30035 tRNAPylor with its cognate tRNA into CHO-K1 cells, using similar methodology as described for HEK293 cells. In CHO-K1 cells, both combinations were shown to support significant trastuzumab production (FIG. 13).

[0241] Interestingly, also as seen when using HEK293 cells, the M. mazei pyrrolysyl-synthetase (Y306A) mutant in the CHO-K1 cells when paired with the bacterial HHW30035 tRNA, outperformed the pair that used the natural M. mazei tRNAPylfor incorporation of 2AzZ-Lys in the CHO-K1 cells.

[0242] In view of the results seen in FIG. 10A showing that hybrid Archaea / bacterial synthetase-tRNA translation systems have the ability to incorporate UAA in vivo, additional testing of hybrid combinations and various UAA was undertaken. These results are shown in FIG. 10B. In these experiments, two hybrid translation systems were tested. These were:(i) the Archaea double mutant M. mazei pyrrolysyl-synthetase (Y306A; Y384F) with the Syntrophomonadaceae bacterial HHW30035 tRNAPyl; and (ii) the Syntrophomonadaceae bacterial HHW30035 pyrrolysyl-synthetase mutant enzyme (Y127A) with the M. mazei tRNAPyl.

[0243] The ability of these components to incorporate UAA was tested with a battery of UAA, identical to those tested in Example 9. These were:Azc-Lys; N6-((2-azidoethoxy)carbonyl)-L-lysine2AzZ-Lys; N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysineSpHD-Lys; N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysineNrb-Lys; N6-(norbornene-methoxycarbonyl)-L-lysineTCO-Lys; N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine

[0244] When the Archaea mutant synthetase and bacterial tRNA combination was tested with the battery of UAA (FIG. 10B, left panel), it was observed that this enzyme / tRNA combination could incorporate each of the five UAA to a significant level, and further where the enzyme / tRNA combination incorporated SpHD-Lys to a significantly higher degree compared to the other UAA.

[0245] A similar result was observed using the bacterial HHW30035 pyrrolysyl-synthetase mutant (Y127A) with the M. mazei tRNAPyl(T25C). That is, this enzyme / tRNA combination could incorporate each of the five UAA with at least a consistently detectable level, and further, where the enzyme / tRNA combination incorporated SpHD-Lys to a significantly higher degree compared to the other UAA.Example 12Productivity of stable cell lines with cross-species synthetase-tRNA pairing

[0246] To assess productivity of the non-cognate synthetase pairs in a stably-integrated cell line, genetic constructs for a mutant archae M. mazei pyrrolysyl-tRNA synthetase containing the Y306A mutation were cotransfected alternatively with M. mazei tRNAPyl(T25C) or the bacterial Syntrophomonadaceae HHW30035 tRNAPylas described in Example 6.

[0247] The M. mazei synthetase gene was cloned into a transposable element containing a glutamine synthetase selection marker gene within one vector. For each tRNA, a separate vector was constructed where five copies of the tRNA gene, each driven by a distinct pol III promoter, were cloned within a transposable element. Stable CHO-K1 cell pools containing pairs of the synthetase and tRNA elements were created co-transfection of the vector pairs with mRNA encoded transposase as described in Example 6. The recovered pools were subsequently transfected with a vector containing a transposable element with genes encoding heavy and light chains of the antibody trastuzumab (Ttz), where the heavy chain gene includes an amber codon (TAG) at position 118 (EU numbering).

[0248] The resulting cell lines were then assessed for trastuzumab production in the presence of two different UAA:2AzZ-Lys; N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysineSpHD-Lys; N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine

[0249] These UAA were provided in fed batch cultures at concentrations of 0.1 mM. Briefly, amber suppression and production of full-length antibody trastuzumab containing an engineered amber codon at position 118 (EU numbering), termed Ttz-amber (A118amb), in stable pools was determined by a fed-batch production process. Stable pools were maintained in Ex-Cell Advanced CHO Fed Batch medium (Millipore Sigma), supplemented with 0 or 4 mM L-glutamine. The cultures were maintained at 37°C with 5% CO2 and 70-80% relative humidity. For routine maintenance, cells were passaged every 3-4 days to a density of 0.3-0.5 x 10A6 cells / mL tomaintain exponential growth, ensuring viability is above 95% before initiating production. For production, cells are inoculated in the aforementioned media at 0.75 x 10A6 cells / mL with .5% w / v HyClone Cell Boost 7b supplement. A glucose bolus is added on day three to 4g / L, and additional feeds of HyClone Cell Boost 7a and 7b are provided on day 4 and, along with glucose (maintained between 3-6g / L), fed intermittently through the duration of the culture. Upon nearing the peak viable cell density, generally between days 4-6 of culture, the first bolus of UAA is added (typically 0.05-0.5 mM) and the culture is adjusted to a temperature of 32°C to shift the culture towards a production phase from the exponential growth phase. Between days 5-7 post UAA addition (days 10-12 post-inoculation), an additional bolus of the UAA is added along with the previously described hyclone feeds / glucose. Culture viability and titer is routinely monitored via Vicell (VCD / viability), Nova Flex 2 (metabolites) and titer (Octet) whereby the information obtained is used for adjusting culture conditions / feeds to optimize culture productivity.

[0250] Three different translation systems were tested in the stably transfected cell lines. These were:

[0251] As shown in FIG. 11A, strong production of the trastuzumab antibody containing the UAA 2AzZ-Lys using the Archae mutRS and with both the Archae tRNAPyland the bacterial tRNAPylthrough day 5 of the fed batch culture post UAA treatment; however, cell viability decreased dramatically for cells expressing the Archaeal Methanosarcina mazei tRNAPyl, but not for those with the bacterial HHW30035 tRNAPyl, when supplementing cultures with 0.1 mM 2AzZ-Lys (FIG.11B). Also seen in FIG. 11B, cells harboring the bacterial HHW30035 tRNAPylremained viable and showed strong antibody production for over 9 days post treatment with 2AzZ-Lys.

[0252] As shown in FIG. 11A, strong production of the trastuzumab antibody containing the UAA SpHD-Lys was also observed using the Archaeal Mm mutRS and the M. mazei tRNAPyl(T25C)through day 9 of the fed batch culture post UAA treatment. The culture fed with SpHD-Lys showed no loss of viability at least up to 9 days exposure to the UAA.Example 13Screening of synthetases for utilization of additional non-natural lysine-derivative amino acids in suppressor tRNA charging

[0253] The synthetase vectors containing the trastuzumab reporter genes and the tRNA expression vectors as described in Examples 7-11 are transfected into HEK293 suspension cells by lipofection. For comparison, cells are also transfected with either wildtype Methanosarcina mazei pyrrolysyl (pyl) synthetase (WP_011033391.1) or a double mutant (Y306A; Y384F) of that enzyme along with M. mazei tRNAPyl(T25C) or with a bacterial tRNAPyl. Cultures are incubated at 37°C in a humidified atmosphere of 8% CO2. One day following transfection, cultures are augmented with a non-natural amino acid at a concentration of 0.05-0.1 mM, with alternatively:N6-[(spiro[2.4]hepta-4,6-d ien-l-ylmethoxy)carbonyl]-L-lysine,N6-(norbornene-methoxycarbonyl)-L-lysine,N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine,N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine,N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine,N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine,N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine,N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine,N-(2-(4-(6-methyl-l,2,4,5-tetrazin-3-yl)phenyl)acetyl)-L-lysine, orN6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine.

[0254] Cell viability is monitored over time using trypan blue cell staining on a Vi-CELL XR Cell Viability Analyzer (Beckman Coulter). Antibody secretion to the culture supernatant is monitored by Biolayer interferometry using probes specific for the intact antibody Fc portion. Expression of the fully assembled trastuzumab IgG is confirmed by non-reducing polyacrylamide gel electrophoresis and LC-MS. For activity screening by biolayer interferometry (BLI), positive activity is defined as significant increase in trastuzumab production relative to cultures not treated with UAA.Example 14Testing of UAA specificity of selected orthogonal synthetase / tRNA combinations

[0255] The UAA specificity of selected orthogonal synthetase / tRNA combinations was investigated. The methods used were as described in Example 8. The incorporation of lysinederivative UAA into trastuzumab was tested using a panel of five different lysine-derivative UAA, each of which contains chemical functionalities useful for bioconjugation. These UAA were:2AzZ-LysSCO-LysCyclopropene-Lysendo-BCN-LysH-L-Photo-Lys,

[0256] The structures of each of these UAA are shown in FIG.6.

[0257] Three orthogonal combinations of synthetase and tRNA were tested. These were:

[0258] FIG. 17 provides the results of these experiments. The bar graph shows trastuzumab antibody titers at day four following transfection of HEK293 cells cultured alternatively in the presence of 0.05 mM UAA selected from a panel of five UAA, which were 2AzZ-Lys, SCO-Lys, cyclopropene-Lys, endo-BCN-Lys, or H-L-Photo-Lys. The host cells were alternatively transfected with one of the three RS and tRNA pairs listed above.

[0259] As seen in FIG. 17, all three combinations of O-RS and O-tRNA showed significant trastuzumab production when cells are cultured in each of the unnatural amino acids, relative to parallel control cultures grown in the absence of any UAA.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A translation system for the incorporation of at least one unnatural lysine derivative amino acid (UAA) in a polypeptide, the system comprising:(a) an unnatural lysine-derivative amino acid, said unnatural amino acid comprising a functional group capable of biocompatible conjugation;(b) an orthogonal aminoacyl-tRNA synthetase (O-RS), wherein said O-RS is a pyrrolysyl tRNA synthetase or derived therefrom; and(c) an orthogonal tRNA (O-tRNA), wherein the O-tRNA is a bacterial pyrrolysyl-tRNA (tRNAPyl) or derived therefrom; andwherein said O-RS is capable of aminoacylating said O-tRNA with said unnatural amino acid.

2. The system of claim 1, wherein the translation system is for the incorporation of at least one unnatural lysine derivative amino acid in a polypeptide in a mammalian host cell.

3. The system of claim 1, wherein the unnatural lysine-derivative amino acid is selected from:N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine (SpHD-Lys),N6-(norbornene-methoxycarbonyl)-L-lysine (Nrb-Lys),N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine (2AzZ-Lys),N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine (3AzZ-Lys),N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine (4AzZ-Lys),N6-((2-azidoethoxy)carbonyl)-L-lysine (Azc-Lys),N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine (endo- BCN-Lys),N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine (SCO-Lys),N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine (Cyclopropene-Lys), N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine (H-L-Photo-Lys), N-(2-(4-(6-methyl-l,2,4,5-tetrazin-3-yl)phenyl)acetyl)-L-lysine (MeTz-PhAc-Lys), andN6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine (TCO-Lys).

4. The system of claim 1, wherein the O-RS is a bacterial O-RS or is derived from a bacterial 0-RS.

5. The system of claim 1, wherein the O-RS is an Archaea O-RS or is derived from an Archaea O-RS.

6. The system of claim 1, wherein the O-tRNA is an amber suppressor tRNA.

7. The system of claim 1, further comprising a nucleic acid encoding a polypeptide of interest, said nucleic acid comprising at least one selector codon, wherein said selector codon is recognized by said O-tRNA.

8. The system of claim 7, wherein the polypeptide of interest is an antibody, an antibody fragment, or an antibody chain.

9. A host cell comprising the system components (a), (b) and (c) of claim 1.

10. The host cell of claim 9, wherein said host cell is a mammalian host cell.

11. The host cell of claim 9, wherein said host cell is a human host cell.

12. The host cell of claim 9, wherein said host cell is a rodent host cell.

13. The host cell of claim 9, wherein said host cell is an antibody-producing host cell.

14. A host cell comprising the system components (a), (b) and (c) of claim 1, and further comprising a nucleic acid encoding a polypeptide of interest, said nucleic acid comprising at least one selector codon, wherein the selector codon is recognized by said O-tRNA.

15. The host cell of claim 14, wherein said nucleic acid encodes an antibody, antibody fragment, or an antibody chain.

16. The translation system of claim 1, further comprising a second O-RS and a second O-tRNA, wherein the second O-RS preferentially aminoacylates the second O-tRNA with a second unnatural amino acid that is different from the first unnatural amino acid, and wherein the second O-tRNA recognizes a selector codon that is different from the selector codon recognized by the first O-tRNA.

17. A translation system for the incorporation of amino acid N6-((2-azidoethoxy)carbonyl)-L-lysine (Azc-Lys) in a polypeptide, the system comprising:(a) the N6-((2-azidoethoxy)carbonyl)-L-lysine amino acid; and(b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 2, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNAof an O-RS / O-tRNA pair provided in Table 2.

18. A translation system for the incorporation of amino acid N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine (2AzZ-Lys) in a polypeptide, the system comprising:(a) the N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine amino acid; and(b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 3, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an O-RS / O-tRNA pair provided in Table 3.

19. A translation system for the incorporation of amino acid N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine (SpHD-Lys) in a polypeptide, the system comprising:(a) the N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine amino acid; and (b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 4, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an O-RS / O-tRNA pair provided in Table 4.

20. A translation system for the incorporation of amino acid N6-(norbornene-methoxycarbonyl)-L-lysine (Nrb-Lys) in a polypeptide, the system comprising:(a) the N6-(norbornene-methoxycarbonyl)-L-lysine amino acid; and(b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 5, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an O-RS / O-tRNA pair provided in Table 5.

21. A translation system for the incorporation of amino acid N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine (TCO-Lys) in a polypeptide, the system comprising:(a) the N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine amino acid; and(b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 6, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an O-RS / O-tRNA pair provided in Table 6.

22. A translation system for the incorporation of an amino acid in a polypeptide, the system comprising:(a) an amino acid selected from:N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine (SCO-Lys),N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine (cyclopropene-Lys), N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine (endo-BCN- Lys), andN6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine (H-L-Photo-Lys), (b) an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with said amino acid, where the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Table 7, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an O-RS / O-tRNA pair provided in Table 7.

23. A method for producing in a translation system a polypeptide of interest comprising at least one unnatural lysine-derivative amino acid at a selected position, the method comprising:(a) providing a translation system of claim 1,(b) providing a nucleic acid encoding the polypeptide of interest, said nucleic acid comprising at least one selector codon, wherein said selector codon is recognized by said O-tRNA; and,(c) incorporating said unnatural amino acid at said selected position in said polypeptide during translation of said polypeptide in response to said selector codon, thereby producing said polypeptide comprising said at least one unnatural amino acid at the selected position.

24. The method of claim 23, wherein the system further comprises a host cell comprising theunnatural lysine derivative amino acid, the O-RS, the O-tRNA, and the nucleic acid encoding the polypeptide of interest, wherein incorporating said unnatural amino acid comprises culturing said host cell.

25. The method of claim 23, wherein the unnatural lysine-derivative amino acid is selected from:N6-[(spiro[2.4]hepta-4,6-dien-l-ylmethoxy)carbonyl]-L-lysine,N6-(norbornene-methoxycarbonyl)-L-lysine,N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine,N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysineN6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine,N6-((2-azidoethoxy)carbonyl)-L-lysine,N6-((((lR,8S,9R)-bicyclo[6.1.0]non-4-yn-9-yl)methoxy)carbonyl)-L-lysine,N6-((cyclooct-2-yn-l-yloxy)carbonyl)-L-lysine,N6-(((2-methylcycloprop-2-en-l-yl)methoxy)carbonyl)-L-lysine,N6-((2-(3-methyl-3H-diazirin-3-yl)ethoxy)carbonyl)-L-lysine,N-(2-(4-(6-methyl-l,2,4,5-tetrazin-3-yl)phenyl)acetyl)-L-lysine, and N6-((((E)-cyclooct-4-en-l-yl)oxy)carbonyl)-L-lysine.

26. A composition comprising an orthogonal tRNA synthetase (O-RS) and an orthogonal tRNA (O-tRNA) pair, wherein said O-RS is capable of aminoacylating said O-tRNA with at least one unnatural amino acid selected from the amino acids provided in FIG. 6, and wherein the O-RS and O-tRNA pair is selected from:(i) the O-RS and O-tRNA pairs provided in Tables 2, 3, 4, 5, 6 and 7, and(ii) an O-RS and O-tRNA pair comprising a variant O-RS and / or a variant O-tRNA of an 0- RS / O-tRNA pair provided in Tables 2, 3, 4, 5, 6 and 7.

27. The composition of claim 26, wherein:(i) said O-RS variants have at least 90% amino acid identity with the corresponding 0- RS amino acid sequence, and where any substituted positions, if any, in the O-RS as shown in Tables 2, 3, 4, 5, 6 and 7 are preserved, and(ii) said O-RS variants are capable of aminoacylating said O-tRNA with said unnatural amino acid.

28. A polynucleotide comprising a pig 7sk promoter operably linked to a DNA sequenceencoding a heterologous tRNA.

29. The polynucleotide of claim 28, wherein the nucleotide sequence of the pig 7sk promoter is SEQ ID NO: 16.

30. The polynucleotide of claim 28, wherein the heterologous tRNA comprises an anticodon that recognizes a stop codon.

31. The polynucleotide of claim 30, wherein the stop codon is an amber codon (DAG).

32. The polynucleotide of claim 30, wherein the DNA sequence encoding the heterologous tRNA comprises SEQ ID NO: 28.

33. The polynucleotide of claim 28, wherein the heterologous tRNA is naturally aminoacylated with pyrrolysine.

34. The polynucleotide of claim 28, wherein the heterologous tRNA is capable of being aminoacylated with an unnatural amino acid.

35. The polynucleotide of claim 34, wherein the unnatural amino acid is a lysine derivative amino acid selected from the unnatural amino acids of FIG. 6.

36. The polynucleotide of claim 28, wherein the polynucleotide comprises a transposon, the transposon further comprising left and right transposon ends such that the promoter and the DNA sequence encoding the heterologous tRNA are transposable by a corresponding transposase.

37. The polynucleotide of claim 28, wherein the polynucleotide further comprises one or more additional promoters selected from: a second pig 7sk promoter, a human U68cc promoter, a water buffalo U6 promoter, a human 7sk promoter, and a mouse Hl promoter, and combinations thereof, and wherein each additional promoter is operably linked to a separate copy of the DNA sequence encoding a heterologous tRNA.

38. The polynucleotide of claim 37, wherein the nucleotide sequence of the one or more additional promoters is selected from SEQ ID NOs: 16, 11, 14, 18, 20, and 31, and combinations thereof.

39. A mammalian host cell comprising the polynucleotide of any one of claims 28 to 38.

40. A method for producing a protein with an unnatural amino acid incorporated at a specific site, the method comprising introducing into a mammalian host cell the polynucleotide of any one of claims 28 to 38.

41. A system for site-specific incorporation of an unnatural amino acid into an antibody in a mammalian host cell, the system comprising:(i) a transposon comprising a plurality of copies of DNA encoding pyrrolysine tRNA, each copy operably linked to a pol III promoter active in the cells, wherein:(a) at least one of the pol III promoters is a pig 7sk promoter; and(b) at least one of the pol III promoters is not a pig 7sk promoter;(ii) a second transposon comprising an open reading frame encoding a corresponding aminoacyl tRNA synthetase operably linked to regulatory elements such that the aminoacyl tRNA synthetase is expressible in the host cell; and(iii) a third transposon comprising open reading frames encoding the heavy and light chains of an antibody, wherein the open reading frame encoding one chain comprises an amber stop codon (UAG) preventing expression of a full antibody unless the amber stop codon is read by the pyrrolysine tRNA to incorporate the unnatural amino acid.

42. The system of claim 41, wherein the first, second, and third transposons are separate.

43. The system of claim 41, wherein the first, second, and third transposons are one transposon.

44. The system of claim 41, wherein the first and second, the first and third, or the second and third transposons are one transposon.

45. The system of claim 41, wherein any combination of the first, second, and third transposons are transposable by the same transposase or by different transposases.

46. A mammalian cell whose genome comprises the system of any one of claims 41 to 45.

47. A method for site-specific incorporation of an unnatural amino acid into an antibody in a mammalian cell, the method comprising delivering the system of any one of claims 41 to 45 into the cell.