Novel aminoacyl-tRNA synthetase mutants for genetic code expansion in eukaryotes
Patent Information
- Application Number
- JP2024514505
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-06
- Filing Date
- 2022-09-05
- Publication Date
- 2025-09-17
AI Technical Summary
Existing aminoacyl-tRNA synthetases, such as PylRS, struggle to efficiently incorporate bulky non-standard amino acids like trans-cyclooct-2-ene-lysine (TCO) and trans-cyclooct-4-ene-lysine (TCO-E) into proteins, with known variants like PylRS AF exhibiting suboptimal incorporation efficiency.
Development of novel PylRS mutants through systematic mutation and back-mutation of key positions, resulting in variants like PylRS A1, PylRS B11, PylRS C11, PylRS G3, and PylRS H12, which exhibit improved incorporation efficiency for bulky non-standard amino acids.
The new PylRS variants demonstrate enhanced ability to incorporate bulky non-standard amino acids into proteins, outperforming previous variants in terms of incorporation efficiency.
Smart Images

Figure 2023031445000001 
Figure 2023031445000002 
Figure 2023031445000003
Abstract
Description
[Technical field]
[0001] The present invention relates to novel improved aminoacyl-tRNA synthetase variants, which are useful for genetic code expansion. The present invention also relates to the corresponding coding sequences and eukaryotic cell lines containing the coding sequences. The present invention also relates to a method for preparing a protein of interest (POI) comprising one or more unnatural amino acid residues incorporated by the novel aminoacyl-tRNA synthetase variants. The present invention also relates to a method for preparing a polypeptide conjugate, in which the POI produced by the aminoacyl-tRNA synthetase variants of the present invention is reacted with one or more binding partner molecules. Finally, the present invention relates to a kit comprising the components necessary for preparing this POI comprising one or more unnatural amino acid residues. [Background technology]
[0002] Genetic code expansion (GCE) is a versatile tool for site-specific incorporation of non-canonical amino acids (ncAAs) into proteins. These ncAAs can be used in protein engineering, such as for fluorescent labeling of proteins or conjugation of toxic payloads with antibodies. To incorporate ncAAs into proteins during translation, specialized aminoacyl-tRNA synthetase / tRNA (aaRS / tRNA) pairs are required that can recognize the ncAA and are otherwise orthogonal to the host organism.
[0003] Currently, several such aaRS / tRNA pairs are known, such as the PylRS / tRNA from archaea such as Methanosarcina mazei, Methanosarcina barkeri, and Methanomethylophilus alvus. Pyl PylRS synthase has a binding site that originally recognizes pyrrolysine. By linking specific amino acids to this binding site, various ncAAs can be made to be recognized by this synthase.
[0004] On the other hand, PylRS / tRNAPyl There are over 200 ncAA pairs that can be used in combination.
[0005] A specific group of ncAAs includes H-Lys(Boc)-OH ("Boc"), cyclooctyne-lysine ("SCO"), bicyclo[6.1.0]nonyne-lysine ("BCN"), trans-cyclooct-2-ene-lysine ("TCO"), and cyclooctyl-lysine ("TCA"). * These ncAAs, which react with tetrazines via click chemistry, are of great interest because the reaction is very fast and bioorthogonal.
[0006] Two amino acids in the binding pocket of PylRS are well known to be important for facilitating the incorporation of such bulky amino acids. With reference to the PylRS of Methanosarcina mazei, the following sequence positions were modified: tyrosine 306 had to be changed to alanine (Y306A) and tyrosine 384 had to be mutated to phenylalanine (Y384F) to be able to incorporate bulky amino acids. As a result, the PylRS mutant PylRS AF The incorporation efficiency of various ncAAs ranged from very low to high. In particular, the ncAA TCO-E was found to be highly efficient in the incorporation of PylRS. AF is not well received.
[0007] Patent Document 1 discloses mutant pyrrolysyl-tRNA synthetases from M. mazei. In particular, single mutants Y306A and Y384F, and the corresponding double mutants at positions 306 and 384, are described therein. In addition, double mutants L309A and C348A are also described therein. It is reported that the mutants are capable of aminoacylation of Boc.
[0008] Patent document 2 describes a method for producing unnatural proteins by applying a modified aminoacyl-tRNA synthetase derived from M. mazei, which contains at least one of the following substitutions: A302F, Y306A, L309A, N346S, C348V / I and Y384F.
[0009] Patent Document 3 describes a method for incorporating an amino acid containing a BCN group into a polypeptide using an orthogonal codon encoding it and an orthogonal PylRS synthetase derived from M. mazei containing a mutation selected from L301V, L305I, Y306F, L309A and C348F.
[0010] Therefore, the problem to be solved by the present invention is to prevent the addition of bulky ncAAs (particularly, trans-cyclooct-2-ene-lysine ("TCO * A) of ncAA residues selected from trans-cyclooct-4-ene-lysine ("TCO-E") and H-Lys(Boc)-OH ("Boc") and combinations thereof, (particularly the mutant PylRS AF The present invention relates to the provision of novel aminoacyl-tRNA synthetases that exhibit improved incorporation (compared to [Prior art documents] [Patent documents]
[0011] [Patent Document 1] European patent application EP-A-2192185 [Patent Document 2] European patent application EP-A-2221370 [Patent Document 3] European patent application EP-A-2804872 [Non-patent literature]
[0012] [Non-Patent Document 1] Reinkemeier et al.,Eur J Chem(2021)27(19)6094-6099 Summary of the Invention
[0013] The above mentioned problems are surprisingly solved by the prior art synthetic enzyme PylRS AF This problem can be solved by systematic mutation of specific key positions in the ribosome and subsequent back mutation at specific positions in the resulting mutants.
[0014] We have mutated five positions of PylRS from the parent gene Methanosarcina mazei to any of 20 possible amino acids. AF Library selection was performed using the PylRS AF1 gene. The gene library was selected by repeated positive and negative screening. As a result of this screening, the present inventors identified PylRS AF1, PylRS AF2, PylRS AF3, PylRS AF4, PylRS AF5, PylRS AF6, PylRS AF7, PylRS AF8, PylRS AF9, PylRS AF10, PylRS AF11, PylRS AF12, PylRS AF13, PylRS AF14, PylRS AF15, PylRS AF16, PylRS AF17, PylRS AF18, PylRS AF19, PylRS AF20, PylRS AF21, PylRS AF22, PylRS AF23, PylRS AF24, PylRS AF25, PylRS AF26, PylRS AF27, PylRS AF28, PylRS AF29, PylRS AF30, PylRS AF31, PylRS AF32, PylRS AF33, PylRS AF34, PylRS AF35, PylRS AF36, PylRS AF37, PylRS AF38, PylRS AF39, PylRS AF40, PylRS AF41, PylRS AF42, PylRS AF43, PylRS AF44, PylRS AF45, PylRS AF46, PylRS AF47, PylRS AF48, PylRS AF49, PylRS AF41, PylRS AF41, PylRS AF42, PylRS AF43, PylRS AF44, PylRS AF45, PylRS AF46, PylRS AF47, PylRS AF48, PylRS AF49, PylRS AF49, Pyl AF Several mutants of the synthase could be obtained. To examine whether the Y306A and Y384F mutations are really important for the incorporation of bulky ncAAs, these amino acids were reverted to their original amino acids by site-directed mutagenesis in the PylRS AF A1 mutant. Thus, the new mutant does not contain the Y306A and Y384F mutations and is called PylRS A1. A mutant without the Y306A and Y384F mutations was further prepared. It has the 306M 309M 348A mutations and is called PylRS MMA (Table 1).
[0015] [Table 1]
[0016] These new PylRS mutants were tested in HEK293T cells by fluorescence flow cytometry (FFC) utilizing a fluorescent reporter. The reporter contains an infrared fluorescent protein (iRFP) fused to a green fluorescent protein (GFP), with an amber stop codon at the Y39 position, iRFP-GFP. Y39TAGAll of the new mutants were able to incorporate both ncAAs tested in this study. Furthermore, when used with TCO-E, all of the new mutants were able to incorporate PylRS. AF More surprisingly, PylRS shows a stronger green signal than PylRS. AF We observed that the Y306A and Y384F mutations in PylRS A1 are indeed not important for the incorporation of bulky ncAAs. Notably, the new "non-AF" mutant PylRS A1 contains mutations that make it a better synthesizer for bulky ncAAs compared to the known PylRS AF mutants.
[0017] Similar to PylRS A1, further "non-AF" mutants were provided, namely PylRS B11, PylRS C11, PylRS G3 and PylRS H12 (see Table 2).
[0018] [Table 2] [Brief description of the drawings]
[0019] [Figure 1] The plasmid maps of the following plasmids used according to the invention are shown: reporter plasmid pCI-iRFP-EGFPY39TAG-6His (SEQ ID NO: 97) (Figure 1A; top). Plasmid pCMV-NES-PylRSAF-U6tRNArv (SEQ ID NO: 98) (Figure 1A; bottom). Plasmid pBK-PylRSWT (SEQ ID NO: 99) (Figure 1B; top). Plasmid pREP-PylT (SEQ ID NO: 100) (Figure 1B; bottom). Plasmids pYOBB2-PylT (SEQ ID NO: 101) (Figure 1C; top) and pALS-sfGFPN150TAG-MbPyl-tRNA (SEQ ID NO: 102) (Figure 1C; bottom). [Diagram 2]Data from FFC experiments of various PylRS mutants (PylRS AF, PylRS AF A1, PylRS AF B11, PylRS AF C11, PylRS AF G3, and PylRS AF H12) using various ncAAs are shown. In particular, the ability to incorporate TCO*A (Figure 2A) and TCO-E (Figure 2B) is shown. Figure 2C shows the expression profile in the absence of ncAAs. [Diagram 3] This figure shows the data obtained by FFC analysis of various PylRS mutants, PylRS AF, PylRS AF A1, PylRS A1, and PylRS MMA, in bar graphs. The ability to incorporate TCO-E (Figure 3A), TCO*A (Figure 3B), and Boc (Figure 3C) is shown. Each bar graph in the upper row shows the ratio of the average GFP signal obtained by FFC divided by the average iRFP signal, which reflects the integration efficiency. Each bar graph in the middle shows the same data, but normalized to the ratio observed for PylRS AF mutants with 100 μM ncAA. Each bar graph in the lower row shows the average GFP / average iRFP ratio normalized to the average GFP / average iRFP ratio obtained for PylRS AF mutants at the desired ncAA concentration. These bar graphs show how much higher the integration efficiency of each mutant is compared to PylRS AF. [Figure 4] A sequence alignment of various archaeal PylRS proteins is shown. [Diagram 5] FFC data are shown evaluating the incorporation efficiency of mutant PylRS A1 at various concentrations of the bulky ncAA cyclooctyne-lysine (SCO). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0020] [Detailed Description of the Invention] A. Abbreviation Bps = base pairs BCN = 2-amino-6-(9-biocyclo[6.1.0]non-4-ynylmethoxycarbonylamino)hexanoic acid BOC = 2-amino-6-(tert-butoxycarbonylamino)hexanoic acid; in the examples, "BOC" specifically refers to (2S)-2-amino-6-(tert-butoxycarbonylamino)hexanoic acid = Boc-L-Lys-OH = N-α-tert-butyloxycarbonyl-L-lysine. Cm=chloramphenicol CMV = cytomegalovirus Crm1 = chromosome region maintenance 1, also known as karyopherin exportin 1. DH10B=E. coli strain F-mcrA Δ(mrr-hsdRMS-mcrBC)φ80lacZΔM15 ΔlacX74 recA1 endA1 araD139 Δ(ara-leu)7697 galU galK λ-rpsL(Str R )nupG dsRNA = double-stranded RNA dSTORM = direct stochastic optical reconstruction microscopy E.coli BL21(DE3)AI=E.coli strain BF - ompT gal dcm lon hsdS B (r B - m B - )λ(DE3[lacI lacUV5-T7p07 ind1 sam7 nin5])[malB + ] K-12 (λ S )araB::T7RNAP-tetA FBS = fetal bovine serum GCE = Genetic Code Extension GFP = Green Fluorescent Protein sfGFP = superfolder green fluorescent protein IPTG = Isopropyl β-D-1-thiogalactopyranoside iRFP = infrared fluorescent protein Kan = Kanamycin ncAA = nonstandard amino acid NES = nuclear export signal NLS = nuclear localization signal NNK = a specific type of degenerate code used for saturation mutagenesis O-tRNA = orthogonal tRNA O-RS = Orthogonal RS PAINT = Point Integration for Imaging in Nanoscale Topography PBS = phosphate buffered saline POI=polypeptide of interest, pRS=prokaryotic RS ptRNA=prokaryotic tRNA ptRNA-ribozyme = a prokaryotic RNA molecule containing ptRNA and at least one ribozyme. PylRS = pyrrolysyl-tRNA synthetase PylRS AF = Mutant M. mazei pyrrolysyl-tRNA synthetase containing the amino acid substitutions Y306A and Y384F RP-HPLC = Reversed Phase High Performance Liquid Chromatography RS = aminoacyl-tRNA synthetase RT=room temperature SCO = 2-amino-6-(cyclooct-2-yn-1-yloxycarbonylamino)hexanoic acid
[0021] [ka]
[0022] SOC = transformation medium (e.g., 0.5% yeast extract; 2% tryptone; 10 mM NaCl, 2.5 mM KCl; 10 mM MgCl2; 10 mM MgSO4; 20 mM glucose) SPAAC = (Copper-Free) Strain-Promoted Retro-Alkyne-Azide Cycloaddition Reaction SPIEDAC = (copper-free) strain-promoted inverse electron demand Diels-Alder cycloaddition reaction SRM = Super Resolution Microscopy SV40 = simian vacuolar virus 40 T7RNAP = T7 RNA polymerase TCO-E = trans-cyclooct-4-ene-L-lysine
[0023] [ka]
[0024] TCO * A = trans-cyclooct-2-ene-L-lysine
[0025] [ka]
[0026] Tet = tetracycline tetR = tetracycline repressor tetO = tetracycline operator tRNA Pyl = a tRNA that can be acylated with pyrrolysine by wild-type or modified PylRS. To site-specifically incorporate an ncAA into the POI, it preferably has an anticodon that is the reverse complement of the selector codon. (The tRNAs used in the examples Pyl In this case, the anticodon is CUA.) UNAA = unnatural amino acid, synonymous with ncAA. U6 promoter = the promoter that normally controls the expression of U6 RNA (small nuclear RNA) in mammalian cells
[0027] B. Definition Unless otherwise defined herein, scientific and technical terms used in connection with the present invention have the meanings commonly understood by those skilled in the art. The meaning and scope of the terms should be clear. However, in the case of potential ambiguity, the definitions provided herein take precedence over any dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms include plurals and plural terms include the singular.
[0028] Pyrrolysyl-tRNA synthetase (PylRS) is an aminoacyl-tRNA synthetase (RS). An RS is an enzyme capable of acylating a tRNA with an amino acid or an amino acid analogue. Suitably, the PylRS of the invention is enzymatically active, i.e. capable of acylating a tRNA (tRNAPyl ) can be acylated with a particular amino acid or amino acid analog, preferably UNAA or a salt thereof.
[0029] As used herein, the term "archaeal pyrrolysyl-tRNA synthetase" (abbreviated as "archaeal PylRS") refers to a PylRS in which at least a segment of the PylRS amino acid sequence or the entire PylRS amino acid sequence has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or 100% sequence identity to the amino acid sequence of a native PylRS from an archaea or an enzymatically active fragment of that native PylRS.
[0030] The PylRS of the present invention can comprise a mutant archaeal PylRS, or an enzymatically active fragment thereof.
[0031] Typically, a "mutant archaeal PylRS" or "mutated archaeal PylRS" differs from the corresponding wild-type PylRS by including additions, substitutions and / or deletions of one or more amino acid residues. Preferably, these are modifications that improve PylRS stability, alter PylRS substrate specificity and / or enhance PylRS enzymatic activity. Particularly preferred "mutant archaeal PylRS" or "mutated archaeal PylRS" are described in more detail herein below.
[0032] The term "nuclear export signal" (abbreviated as "NES") refers to an amino acid sequence that can direct a polypeptide containing it (such as the NES-containing PylRS of the present invention) to be exported from the nucleus of a eukaryotic cell. The export is believed to be mediated in large part by Crm1 (chromosomal region maintenance 1; also known as karyopherin exportin 1). NESs are known in the art. For example, the database ValidNES (http: / / validness.ym.edu.tw / ) provides sequence information of experimentally verified NES-containing proteins. Furthermore, NES databases are publicly available, such as, for example, NESbase1.0 (see www.cbs.dtu.dk / databased / NESbase-1.0 / ; Le Cour et al., Nucl Acids Res 31(1), 2003), as well as tools for NES prediction, such as NetNES (see www.cbs.dtu.dk / services / NetNES / ; La Cour et al., Protein Eng Des Sel 17(6):527-536, 2004), NESpredictor (NetNES, http: / / www.cbs.dtu.dk / ; Fu et al., Nucl Acids Res 41:D338-D343, 2013; La Cour et al., Protein Eng Des Sel 17(6):527-536, 2004), and NESsential (a web interface in combination with ValidNES). Hydrophobic leucine-rich NESs are the most common and represent the best characterized group of NESs to date. Hydrophobic leucine-rich NESs are non-conserved motifs with three or four hydrophobic residues. Many of these NESs contain the conserved amino acid sequence pattern LxxLxL (SEQ ID NO: 111) or LxxxLxL (SEQ ID NO: 112), where each L is independently selected from the amino acid residues leucine, isoleucine, valine, phenylalanine, and methionine, and each x is independently selected from any amino acid (see La Cour et al., Protein Eng Des Sel 17(6):527-536, 2004).
[0033] The term "nuclear localization signal" (abbreviated as "NLS" and also referred to in the art as "nuclear localization sequence") refers to an amino acid sequence capable of directing the import of a polypeptide containing it (e.g., wild-type archaeal PylRS) into the nucleus of a eukaryotic cell. The transport is believed to be mediated by the binding of the NLS-containing polypeptide to importins (also known as karyopherins) to form a complex that translocates through the nuclear pore. NLSs are known in the art. Numerous NLS databases and tools for NLS prediction are publicly available. For example, NLSdb (see Nair et al., Nucl Acids Res 31(1), 2003), cNLS Mapper (www.nls-mapper.aib.keio.ac.jp; Kosugi et al., Proc Natl Acad Sci USA.106(25):10171-10176, 2009; Kosugi et al., J Biol Chem 284(1):478-485, 2009), SeqNLS (see Lin et al., PLoS One 8(10):e76864, 2013) and NucPred (www.sbc.su.se / ~maccallr / nucpred / ; Branmeier et al., Bioinformatics 23(9):1159-60, 2007).
[0034] The mutant archaeal PylRS of the invention as defined above can be further modified by removing any NLS present in said native PylRS from which the mutant is derived and / or by introducing at least one NES. NLSs in native PylRS can be identified using known NLS detection tools, such as, for example, cNLS mapper.
[0035] Removal of an NLS from an archaeal PylRS or a mutant thereof and / or introduction of an NES into an archaeal PylRS or a mutant thereof can alter the localization of the thus modified polypeptide when expressed in a eukaryotic cell, in particular avoiding or reducing accumulation of the polypeptide in the nucleus of the eukaryotic cell. Thus, the localization of a PylRS mutant of the invention expressed in a eukaryotic cell can be altered compared to a PylRS or PylRS mutant that differs from a PylRS mutant of the invention in that it (still) contains an NLS and lacks an NES.
[0036] If the archaeal PylRS of the invention contains an NES but (still) contains an NLS, the NES is preferably selected such that its strength overrides the NLS and prevents accumulation of the PylRS in the nucleus of a eukaryotic cell.
[0037] Removal of an NLS from a wild-type or mutant PylRS and / or introduction of an NES into a wild-type or mutant PylRS to obtain a PylRS of the present invention does not abolish PylRS enzymatic activity. Preferably, PylRS enzymatic activity is maintained at essentially the same level. That is, the PylRS of the present invention has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 91, 92, 93, 94, 95, 96, 97, 98, or 99% of the enzymatic activity of the corresponding wild-type or mutant PylRS.
[0038] The NES is appropriately positioned within the PylRS or mutant PylRS of the present invention so that the NES operates. For example, the NES can be bound to the C-terminus (e.g., C-terminus of the last amino acid residue) or N-terminus (e.g., between amino acid residue 1, the N-terminal methionine, and amino acid residue 2) of the wild-type or mutant archaeal PylRS.
[0039] The disclosure of WO2018 / 06948, which discloses mutated PylRS modified by the incorporation of an NES and / or deletion of an NLS sequence, is expressly referenced herein and incorporated by reference.
[0040] The PylRS mutant of the present invention is a tRNA Pyl The PylRS mutant is preferably UNAA or a salt thereof, and is used in combination with tRNA / PylRS(mutant). Pyl can be acylated.
[0041] Unless otherwise indicated, as used herein, "tRNA Pyl " means a tRNA that can be acylated (substantially selectively, particularly selectively) by the PylRS mutant of the present invention. In the context of the present invention, the tRNA described herein Pyl can be a wild-type tRNA that can be acylated by PylRS with pyrrolysine, or a mutant of that tRNA, such as a wild-type or mutant tRNA from an archaea, such as from the genus Methanosarcina (e.g., M. mazei or M. barkeri). The tRNA used with the PylRS of the present invention for site-specific incorporation of a UNAA into the POI Pyl The anticodon contained in the tRNA is suitably the reverse complement of the selector codon. Pyl The anticodon of is the reverse complement of the amber stop codon. For other applications, such as proteome labeling (Elliott et al., Nat Biotechnol 32(5):465-472, 2014), the tRNA used with the PylRS of the present invention may be Pyl The anticodon contained in may be a codon recognized by an endogenous tRNA of a eukaryotic cell.
[0042] As used herein, the term "selector codon" refers to a codon that is selected from a tRNA during the translation process. Pyl The term selector codon refers to a codon that is recognized by (i.e., binds to) a tRNA that is not recognized by endogenous tRNAs in a eukaryotic cell. The term is also used for the corresponding codon in a polypeptide-encoding sequence of a polynucleotide that is not a messenger RNA (mRNA), such as a DNA plasmid. Preferably, the selector codon is a minor codon in naturally occurring eukaryotic cells. tRNA PylThe anticodon of binds to the selector codon in the mRNA, thus site-specifically incorporating the UNAA into the growing chain of the polypeptide encoded by said mRNA. The 64 known genetic codons (triplets) code for 20 amino acids and three stop codons. Since only one stop codon is required for translation termination, the other two can in principle be used to code for non-proteinogenic amino acids. For example, the amber codon, UAG, has been successfully used as a selector codon in in vitro and in vivo translation systems to direct the incorporation of unnatural amino acids. The selector codons utilized in the methods of the invention expand the genetic codon framework of the protein biosynthetic machinery of the translation system used. Specifically, selector codons include, but are not limited to, nonsense codons such as stop codons (e.g., amber (UAG), ochre (UAA) and opal (UGA) codons); codons consisting of more than three bases (e.g., four-base codons); and codons derived from natural or unnatural base pairs. For a given system, the selector codon can also include one of the naturally occurring three-base codons (i.e., a naturally occurring triplet) where the endogenous translation system (e.g., a system that lacks a tRNA that recognizes the naturally occurring triplet or where the naturally occurring triplet is a rare codon) does not use (or rarely uses) the naturally occurring triplet.
[0043] In a given translation system (e.g., eukaryotic cells), a recombinant tRNA that alters the reading of mRNA so that it can read through a stop codon, a four-base codon, or a rare codon is called a suppressor tRNA. The efficiency of suppressing a stop codon (e.g., an amber codon) that acts as a selector codon is determined by the (aminoacylated) tRNA. Pyl The efficiency of such suppression of stop codons depends on the competition between tRNAs (acting as suppressor tRNAs) and release factors (e.g., RF1) that bind to the stop codon and initiate the release of the growing polypeptide chain from the ribosome. Thus, the efficiency of such suppression of stop codons can be increased by using release factor- (e.g., RF1-) deficient strains.
[0044] A polynucleotide sequence encoding a "polypeptide of interest" or "POI" can include one or more codons (e.g., selector codons), e.g., two or more, three or more, etc. Pyl In order to generate a polynucleotide sequence encoding a POI, said codon can be introduced into the polynucleotide sequence at the desired site using conventional site-directed mutagenesis.
[0045] PylRS and tRNA of the present invention Pyl is preferably orthogonal.
[0046] As used herein, the term "orthogonal" refers to a molecule (e.g., an orthogonal tRNA and / or an orthogonal RS) that is used with low efficiency by a translation system of interest (e.g., a eukaryotic cell used to express a POI described herein). "Orthogonal" refers to the inability or reduced efficiency, e.g., the efficiency of the orthogonal tRNA or orthogonal RS to interact with an endogenous RS or endogenous tRNA, respectively, of the translation system of interest is less than 20%, less than 10%, less than 5%, or, for example, less than 1%.
[0047] Thus, in certain embodiments of the invention, any endogenous RS of a eukaryotic cell of the invention acylates (orthogonal) tRNAs with low or zero efficiency when compared to acylation of endogenous tRNAs by the endogenous RS. Pyl For example, with less than 20% efficiency, less than 10% efficiency, less than 5% efficiency, or less than 1% efficiency. Alternatively or additionally, the (orthogonal) PylRS of the present invention catalyzes the acylation of tRNA by a cell's endogenous RS. Pyl In comparison to the acylation of any endogenous tRNA of a eukaryotic cell of the invention, the tRNA is acylated with low or zero efficiency, e.g., less than 20% efficiency, less than 10% efficiency, less than 5% efficiency, or less than 1% efficiency.
[0048] Unless otherwise indicated, the terms "endogenous tRNA" and "endogenous aminoacyl-tRNA synthetase" ("endogenous RS") as used herein refer to the PylRS and tRNA of the invention, respectively, as used in the context of the present invention.Pyl These terms refer to the tRNA and RS that are present in the cell ultimately to be used as a translation system prior to the introduction of the
[0049] The term "translation system" generally refers to a set of components necessary to incorporate natural amino acids into a growing polypeptide chain (protein). Translation system components can include, for example, ribosomes, tRNA, aminoacyl-tRNA synthetases (RS), mRNA, etc. Translation systems include artificial mixtures of the above components, cell extracts, and living cells (e.g., living eukaryotic cells).
[0050] PylRS and tRNA used to generate POI according to the present invention Pyl Preferably, in the eukaryotic cell used to generate the POI, the pair Pyl is orthogonal in that it is selectively acylated with a UNAA or a salt thereof (UNAA) by the PylRS of the present invention. Suitably, in the eukaryotic cell, the orthogonal pair is a UNAA acylated tRNA Pyl The POI functions to incorporate a UNAA residue into the growing polypeptide chain of the POI using, for example, a tRNA. Pyl recognizes a codon in the mRNA that encodes the POI (eg, a selector codon, such as an amber stop codon).
[0051] As used herein, the term "selectively acylated" refers to a tRNA or amino acid that is selectively acylated by a PylRS compared to the endogenous tRNA or amino acid of a eukaryotic cell. Pyl The efficiency of acylation of tRNA with UNAA means, for example, about 50% efficiency, about 70% efficiency, about 75% efficiency, about 85% efficiency, about 90% efficiency, about 95% efficiency, or about 99% or more efficiency. Pyl is incorporated into the growing polypeptide chain with high accuracy (e.g., greater than about 75%, greater than about 80%, greater than about 90%, greater than about 95%, or greater than about 99% efficiency).
[0052] tRNA suitable for producing the POI according to the present invention Pyl The tRNA / PylRS pair can be selected from a library of mutant tRNAs and PylRSs, for example, based on the results of library screening. Such selection can be performed analogously to known methods for evolving tRNA / RS pairs, for example, as described in WO02 / 085923 and WO02 / 06075. The tRNA of the present invention Pyl To generate a / PylRS pair, we start with a wild-type or mutant archaeal PylRS that (still) contains a nuclear localization signal and lacks an NES, and then ligate the appropriate tRNA Pyl Before or after identifying the / PylRS pair, the nuclear localization signal can be removed and / or a NES can be introduced.
[0053] The term "unnatural amino acid" (abbreviated as "UNAA") as used herein means an amino acid that is not one of the 20 standard amino acids or selenocysteine or pyrrolysine. The term also means an amino acid analogue, e.g., a compound that, unlike an amino acid, has an α-amino group replaced with a hydroxyl group and / or a carboxylic acid functional group to form an ester. When translationally incorporated into a polypeptide, the amino acid analogue results in an amino acid residue that is different from the corresponding amino acid residue of the 20 standard amino acids or selenocysteine or pyrrolysine. When the amino acid analogue UNAA (wherein the carboxylic acid functional group forms an ester of the formula -C(O)-OR) is used in a translation system (such as a eukaryotic cell) to generate a polypeptide, it is believed that R is removed in situ, e.g., enzymatically, in the translation system before being incorporated into the POI. Thus, R is appropriately selected to be compatible with the ability of the translation system to convert UNAA or a salt thereof into a form that is recognized and processed by the PylRS of the present invention.
[0054] UNAA useful in the methods and kits of the present invention have been described in the prior art (see, e.g., Liu et al., Annu Rev Biochem 83:379-408, 2010; Lemke, ChemBioChem 15:1691-1694, 2014).
[0055] As used herein, the term "host cell" or "transformed cell" refers to a cell (or organism) that has been modified to carry at least one nucleic acid molecule, e.g., a recombinant gene encoding a desired protein or nucleic acid sequence, which upon transcription results in a polypeptide for use as described herein. A host cell can be a prokaryotic or eukaryotic cell, such as a bacterial cell, a fungal cell, a plant cell, an insect cell, or a mammalian cell. A host cell may contain a recombinant gene integrated into the nuclear or organelle genome of the host cell. Alternatively, the host may contain the recombinant gene extrachromosomally.
[0056] It is meant that a particular organism or cell is "capable of producing a POI" if it naturally produces the POI, or if it does not naturally produce said POI but has been transformed to produce said POI.
[0057] The terms "purified", "substantially purified" and "isolated" as used herein refer to a state in which the compound of the invention is free of other distinct compounds with which it is normally associated in its natural state, and the "purified", "substantially purified" and "isolated" subject comprises at least 0.5%, 1%, 5%, 10%, or 20% by weight, or at least 50% or 75% by weight of a given sample. In one embodiment, these terms refer to a compound of the invention comprising at least 95, 96, 97, 98, 99 or 100% by weight of a given sample. As used herein, the terms "purified", "substantially purified" and "isolated" when referring to a nucleic acid or protein also refer to a state of purification or enrichment different from that which occurs naturally, for example in a prokaryotic or eukaryotic environment, such as a bacterial or fungal cell or a mammalian organism, particularly in the human body. Anything beyond the degree of purification or enrichment that occurs in nature is within the meaning of "isolated," including (1) purification from other related structures or compounds, or (2) association with structures or compounds with which it is not normally associated in said prokaryotic or eukaryotic environment. The nucleic acids or proteins, or types of nucleic acids or proteins described herein, can be isolated or associated with structures or compounds with which they are not normally associated in nature, according to a variety of methods and processes known to those of skill in the art.
[0058] In the context of the description provided herein and the appended claims, the use of "or" means "and / or" unless stated otherwise.
[0059] Similarly, the terms "comprise", "comprises", "include", "includes" and "including" are interchangeable and are not intended to be limiting.
[0060] Additionally, where the term "comprise" is used in describing various embodiments, those skilled in the art will understand that in some specific examples, the embodiments may alternatively be described using the terms "consisting essentially of" or "consisting of."
[0061] The term "about" indicates a potential variation of ±25%, particularly ±15%, ±10%, and more particularly ±5%, ±2% or ±1% of the stated value.
[0062] The term "substantially" refers to a range of values from about 80 to 100%, such as 85 to 99.9%, particularly 90 to 99.9%, more particularly 95 to 99.9%, or 98 to 99.9%, and particularly 99 to 99.9%.
[0063] "Mainly" refers to a proportion in the range of more than 50%, for example, in the range of 51-100%, particularly in the range of 75-99.9%, more particularly in the range of 85-98.5%, for example, 95-99%.
[0064] Where the present disclosure refers to features, parameters, and ranges thereof of varying degrees of preference (including general, explicitly non-preferred features, parameters, and ranges thereof), unless otherwise stated, any combination of two or more of such features, parameters, and ranges is encompassed by the disclosure herein, regardless of their respective preferences.
[0065] The term "improved activity for substrate utilization" observed for a particular enzyme or enzyme variant described herein refers to the improved utilization observed compared to a reference enzyme, in particular a non-mutated parent enzyme, or compared to an enzyme variant that differs in terms of the number and / or type of mutations contained in the variant showing said improved utilization. Suitable non-limiting parameters for expressing the change in utilization are the decrease in substrate concentration expressed as a percentage, e.g., mole %, or the increase in product concentration expressed as a percentage, e.g., mole %. A preferred "substrate" in this context is the UNAA referred to herein.
[0066] C. Specific Aspects and Embodiments of the Invention The present invention relates to the following aspects and specific embodiments thereof: A first aspect of the present invention relates to a modified archaeal pyrrolysyl-tRNA synthetase (PylRS), comprising a combination or set of sequence motifs M1 and M3; in particular, M1 is near the N-terminus and M3 is near the C-terminus of the amino acid sequence of said modified archaeal PylR; optionally, at least one further sequence motif selected from M2, M4, M5 and M6 is combined and retains PylRS activity; Where: M1, M2, M3, M4, M5 and M6 are arranged in the above order within the amino acid sequence of the modified PylRS, with M1 being closest to the N-terminus and M6 being closest to the C-terminus of the amino acid sequence of the modified archaeal PylRS, and containing the following sequences (shown in the order of N-terminus → C-terminus, respectively):
[0067] M1: LRPMX1AX2X3L(Y / M)X5X6(M / V / C)R (SEQ ID NO: 1) Similar to amino acid residues 297-310 of SEQ ID NO:56 (i.e., the wild-type PylRS of M. mazei).
[0068] M2:HLX7EFTMX8NX9(G / A)X 11 X 12 G (SEQ ID NO: 2) Similar to amino acid residues 338-351 of SEQ ID NO:56 (i.e., the wild-type PylRS of M. mazei).
[0069] M3: VYX 13 X 14 TX 15 D (SEQ ID NO: 3) Similar to amino acid residues 383-389 of SEQ ID NO:56 (i.e., the wild-type PylRS of M. mazei).
[0070] M4:SX 16 X 17 X 18GP(R / I / N)X 20 X 21 D (SEQ ID NO: 4) Similar to amino acid residues 399-408 of SEQ ID NO:56 (i.e., the wild-type PylRS of M. mazei).
[0071] M5:X 22 X 23 (I / V)X 25 X 26 PW (SEQ ID NO:5) Similar to amino acid residues 411-417 of SEQ ID NO:56 (i.e., the wild-type PylRS of M. mazei).
[0072] M6: G(A / L / I)GFGLERLL (SEQ ID NO: 6) Similar to amino acid residues 419 to 428 of SEQ ID NO: 56 (i.e., the wild-type PylRS of M. mazei), where amino acid residues X1 to X 26 are independently selected from naturally occurring amino acid residues.
[0073] Each of these conserved motifs M1 to M6 contains at least one amino acid residue and is presumed to be involved in the formation of the substrate pocket of the enzyme.
[0074] Each intervening sequence motif linking two adjacent motifs, as well as the sequence motifs forming the N- and C-terminal sequences, may be derived from the respective sequence portions of other wild-type or prior art archaeal PylRSs. As shown by the sequence alignment attached as FIG. 4, each of the intervening subsequences or the N- or C-terminal subsequences can be easily derived from other archaeal PylRSs, such as, but not limited to, SEQ ID NOs: 58, 60, 62, 64 and 66. Based on such sequence alignment, or supplementation of any similar alignment with at least one further PylRS sequence of different archaeal origin, each further modified sequence motif can be easily derived, since such less conserved portions are expected to have no or no significant effect on synthetase activity.
[0075] In certain embodiments of the modified archaeal PylRS, residues X1 to X 26 have, independently of one another, the following meanings: X1 represents an amino acid residue selected from natural amino acid residues, in particular L or H; X2 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular P or M; X3 represents an amino acid residue selected from natural amino acid residues, in particular N, T or V; X5 represents an amino acid residue selected from natural amino acid residues, in particular N, S, T or Y; X6 represents an amino acid residue selected from natural amino acid residues, in particular L, M or W; X7 represents an amino acid residue selected from natural amino acid residues, in particular E or N; X8 represents an amino acid residue selected from naturally occurring amino acid residues, in particular V or L; X9 represents an amino acid residue selected from natural amino acid residues, in particular F or L; X 11 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular Q, D or E; X 12 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular M or L; X 13 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular G, K or V; X 14 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular D, E or N; X 15 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular L, I or V; X 16 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular A or G; X 17 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular A or V; X 18represents an amino acid residue selected from the naturally occurring amino acid residues, in particular V or M; X 20 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular P, S, Y, F or V; X 21 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular L or M; X 22 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular W or H; X 23 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular G, D or E; X 25 represents an amino acid residue selected from the naturally occurring amino acid residues, in particular D, H, F or N; X 26 represents an amino acid residue selected from the natural amino acid residues, in particular K, E or D.
[0076] In further particular embodiments of the first aspect, there is provided a modified archaeal PylRS comprising a combination or set of sequence motifs M1, M3 and M2; or M1, M3, M2 and M4; or M1, M3, M2, M4 and M5; or M1, M3, M2, M4 and M6; or M1, M3, M2, M4, M5 and M6.
[0077] In further specific embodiments of the first aspect, a modified archaeal PylRS is provided, which is derived from a parent PylRS from an archaeon of the genera Methanosarcina, Methanosarcinaceae, Methanomethylophilus, Desulfitobacterium, and Candidatus Methanoplasma, in particular Methanosarcina.
[0078] In further particular embodiments of the first aspect, a modified archaeal PylRS is provided, which is derived from a parent PylRS derived from the archaea of the species Methanosarcina mazei (SEQ ID NO: 56), Methanosarcina barkeri (SEQ ID NO: 58), Methanosarcinaceae archaeon (SEQ ID NO: 60), Methanomethylophilus alvus (SEQ ID NO: 62), Desulfitobacterium hafniense (SEQ ID NO: 64) and Candidatus Methanoplasma termitum (SEQ ID NO: 66), in particular the species Methanosarcina mazei (SEQ ID NO: 56).
[0079] In a further particular embodiment of the first aspect, a modified archaeal PylRS is provided, which is derived from a parent PylRS having an amino acid sequence selected from SEQ ID NOs: 56, 58, 60, 62, 64 and 66 (particularly SEQ ID NO: 56), or a functional variant or fragment thereof, which retains pyrrolysyl-tRNA synthetase activity and has at least 60% sequence identity, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%, to a native pyrrolysyl-tRNA synthetase having an amino acid sequence selected from SEQ ID NOs: 56, 58, 60, 62, 64 and 66. This modified archaeal PylRS comprises a combination of modified sequence motifs M1 and M3; optionally combined with at least one further sequence motif selected from M2, M4, M5 and M6, each as defined above.
[0080] Functional variants or fragments of the parent PylRS having an amino acid sequence selected from SEQ ID NOs: 56, 58, 60, 62, 64 and 66 can be easily derived from the respective sequence alignment, e.g. as shown in Figure 4, which provides information on sequence motifs such as M1 to M6 important for the enzyme function, as well as intermediate or terminal sequence parts that are relatively susceptible to sequence changes and the resulting sequence mutations, e.g. conservative amino acid substitutions.
[0081] In further particular embodiments of the first aspect, a modified archaeal PylRS is provided, wherein each of the above motif combinations or sets, independently of other motifs within the set, is as follows: a) said sequence motif M1 is selected from the following sequences:
[0082] [Table 3]
[0083] b) said sequence motif M2 is selected from the following sequences:
[0084] [Table 4]
[0085] c) said sequence motif M3 is selected from the following sequences:
[0086] [Table 5]
[0087] d) said sequence motif M4 is selected from the following sequences:
[0088] [Table 6]
[0089] e) said sequence motif M5 is selected from the following sequences:
[0090] [Table 7]
[0091] f) said sequence motif M6 is selected from the following sequences:
[0092] [Table 8]
[0093] In further particular embodiments of the first aspect, a modified archaeal PylRS is provided, wherein: a) said sequence motif M1 is selected from the following sequences:
[0094] [Table 9]
[0095] b) said sequence motif M2 is selected from the following sequences:
[0096] [Table 10]
[0097] c) said sequence motif M3 is selected from the following sequences:
[0098] [Table 11]
[0099] d) said sequence motif M4 is selected from the following sequences:
[0100] [Table 12]
[0101] e) said sequence motif M5 is selected from the following sequences:
[0102] [Table 13]
[0103] f) said sequence motif M6 is selected from the following sequences:
[0104] [Table 14]
[0105] According to further particular embodiments, the following combinations or sets of sequence motifs are provided, which are selected from one of the following groups (1) to (5):
[0106] (1) M1, M3 and M2: M1a, M3a and M2a; M1b, M3b and M2b; M1c, M3c and M2c; M1d, M3d and M2d; M1e, M3e and M2e; M1f, M3f and M2f; M1a * ,M3a * and M2a * ;M1b * ,M3b * and M2b * ;M1c * ,M3c * and M2c * ;M1d * ,M3d * and M2d * ;M1e * ,M3e * and M2e * ;M1 * f,M3f * and M2f * ;
[0107] (2) M1, M3, M2 and M4: M1a, M3a, M2a and M4a; M1b, M3b, M2b and M4b; M1c, M3c, M2c and M4c; M1d, M3d, M2d and M4d; M1e, M3e, M2e and M4e; M1f, M3f, M2f and M4f; M1a * ,M3a * ,M2a * and M4a * ;M1b * ,M3b * ,M2b * and M4b * ;M1c * ,M3c * ,M2c * and M4c * ;M1d * ,M3d * ,M2d * and M4d * ;M1e * ,M3e * ,M2e * and M4e * ;M1f * ,M3f * ,M2f * and M4f * ;
[0108] (3) M1, M3, M2, M4 and M5: M1a, M3a, M2a, M4a and M5a; M1b, M3b, M2b, M4b and M5b; M1c, M3c, M2c, M4c and M5c; M1d, M3d, M2d, M4d and M5d; M1e, M3e, M2e, M4e and M5e; M1f, M3f, M2f, M4f and M5f; M1a * ,M3a * ,M2a * ,M4a * and M5a * ;M1b * ,M3b * ,M2b * ,M4b * and M5b * ;M1c * ,M3c * ,M2c * M4c *and M5c * ;M1d * ,M3d * ,M2d * ,M4d * and M5d * ;M1e * ,M3e * ,M2e * ,M4e * and M5e * ;M1f * ,M3f * ,M2f * ,M4f * and M5f * ;
[0109] (4) M1, M3, M2, M4 and M6: M1a, M3a, M2a, M4a and M6a; M1b, M3b, M2b, M4b and M6b; M1c, M3c, M2c, M4c and M6c; M1d, M3d, M2d, M4d and M6d; M1e, M3e, M2e, M4e and M6e; M1f, M3f, M2f, M4f and M6f; M1a * ,M3a * ,M2a * ,M4a * and M6a * ;M1b * ,M3b * ,M2b * ,M4b * and M6b * ;M1c * ,M3c * ,M2c * ,M4c * and M6c * ;M1d * ,M3d * ,M2d * ,M4d * and M6d * ;M1e * ,M3e * ,M2e * ,M4e * and M6e * ;M1f * ,M3f * ,M2f * ,M4f *and M6f * ;
[0110] (5) M1, M3, M2, M4, M5 and M6: M1a, M3a, M2a, M4a, M5a and M6a; M1b, M3b, M2b, M4b, M5b and M6b; M1c, M3c, M2c, M4c, M5c and M6c; M1d, M3d, M2d, M4d, M5d and M6d; M1e, M3e, M2e, M4e, M5e and M6e; M1f, M3f, M2f, M4f, M5f and M6f; M1a * ,M3a * ,M2a * ,M4a * ,M5a * and M6a * ;M1b * ,M3b * ,M2b * ,M4b * ,M5b * and M6b * ;M1c * ,M3c * ,M2c * ,M4c * ,M5c * and M6c * ;M1d * ,M3d * ,M2d * ,M4d * ,M5d * and M6d * ;M1e * ,M3e * ,M2e * ,M4e * ,M5e * and M6e * ;M1f * ,M3f * ,M2f * ,M4f * ,M5f * and M6f * ;
[0111] In further particular embodiments of the first aspect, there is provided a modified archaeal PylRS comprising a combination of sequence motifs M1a, M2a, M3a, M4a, M5a and M6a. A non-limiting example of a PylRS variant of said type is PylRS A1 (SEQ ID NO: 70).
[0112] In a further particular embodiment of the first aspect, the sequence motif M1a * , M2a * , M3a, M4a * A modified archaeal PylRS is provided that contains a combination of M5a, M6a and M7a. A non-limiting example of a PylRS variant of said type is PylRS MMA (SEQ ID NO: 72).
[0113] In a further particular embodiment of the first aspect, there is provided a modified archaeal PylRS, which is as follows:
[0114] a) PylRS A1 comprising the amino acid sequence of SEQ ID NO: 70; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 70; or a functional fragment thereof that retains PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; or b) a PylRS MMA comprising the amino acid sequence of SEQ ID NO: 72; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 72; or a functional fragment thereof that retains PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; or c) PylRS B11 comprising the amino acid sequence of SEQ ID NO: 82; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 82; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; d) PylRS C11 comprising the amino acid sequence of SEQ ID NO: 84; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 84; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; e) PylRS G3 comprising the amino acid sequence of SEQ ID NO: 86; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 86; or a functional fragment thereof that retains PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; f) PylRS H12 comprising the amino acid sequence of SEQ ID NO: 88; or an amino acid sequence having at least 60%, for example 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 88; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above.
[0115] In a further specific embodiment of the first aspect, a modified archaeal PylRS is provided, which is a functional fusion protein retaining PylRS activity, comprising at least a first protein portion having PylRS activity and a second protein portion having the same or different functionality covalently linked to the first protein.
[0116] In further particular embodiments of the first aspect, a modified archaeal PylRS is provided that exhibits at least one of the following functional characteristics: a) Altered substrate profile for non-standard amino acids (ncAA) or their salts. b) Compared to mutant PylRS AF (SEQ ID NO: 68), in particular TCO-E, TCO * Improved utilization of at least one bulky ncAA selected from A and Boc or a salt thereof. c) mutant PylRS AF A1 (SEQ ID NO: 108), respectively, in particular TCO * Improved utilization of at least one bulky ncAA selected from A and Boc or a salt thereof.
[0117] In further particular embodiments of the first aspect, there is provided a modified archaeal PylRS exhibiting at least one of the following functional characteristics: a) TCO-E, TCO-F, and TCO-F, respectively, compared to mutant PylRS AF (SEQ ID NO: 68). * Improved utilization of at least one bulky ncAA selected from A and Boc or a salt thereof. b) TCO * Improved utilization of at least one bulky ncAA selected from A and Boc or a salt thereof.
[0118] A non-limiting example of a PylRS mutant that meets these requirements is PylRS A1 (SEQ ID NO: 70).
[0119] In further particular embodiments of the first aspect, a modified archaeal PylRS is provided that exhibits at least the following functional characteristics: Improved utilization of bulky ncAA by TCO-E or its salt compared to mutant PylRS AF (SEQ ID NO: 108).
[0120] A non-limiting example of a PylRS mutant that meets these requirements is PylRS MMA (SEQ ID NO: 72).
[0121] In a further particular embodiment of the first aspect, a modified archaeal PylRS is provided which further comprises a nuclear export signal (NES). NES signal sequences are well known as described herein above. Each of such known sequences is considered as a candidate for combining said sequence with the PylRS mutant sequence of the present invention. The NES can be inserted or added at any sequence position at the N-terminus or C-terminus of the PylRS sequence, as long as it does not adversely affect the PylRS activity. "Added at the N-terminus or C-terminus" means that the NES can be added to the first or last amino acid of the PylRS mutant sequence, or inserted at any sequence position in the sequence portion corresponding to the terminal amino acid residues 30, 20, 10 or 5 of the enzyme N-terminus or C-terminus, respectively.
[0122] In particular, the NES comprises the amino acid sequence of SEQ ID NO:54, which is specifically added to the N-terminus of the enzyme.
[0123] Non-limiting examples of NES-modified PylRS mutants of the present invention include those having an amino acid sequence selected from SEQ ID NOs: 78, 80, 90, 92, 94 and 96; or an amino acid sequence having at least 60%, for example 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NOs: 78, 80, 90, 92, 94 and 96, respectively, and retaining the NES motif.
[0124] In a further particular embodiment of the first aspect, there is provided a modified archaeal PylRS, which is: a) a NES PylRS A1 comprising the amino acid sequence of SEQ ID NO: 78; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 70; or a functional fragment thereof that retains PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; or b) a NES PylRS MMA comprising the amino acid sequence of SEQ ID NO: 80; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 72; or a functional fragment thereof that retains PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; or c) NES PylRS B11 comprising the amino acid sequence of SEQ ID NO: 90; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 82; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; d) NES PylRS C11 comprising the amino acid sequence of SEQ ID NO: 92; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 84; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; e) a NES PylRS G3 comprising the amino acid sequence of SEQ ID NO: 94; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 86; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above; f) NES PylRS H12 comprising the amino acid sequence of SEQ ID NO: 96; or an amino acid sequence having at least 60%, such as 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to SEQ ID NO: 88; or a functional fragment thereof retaining PylRS activity; preferably retaining the respective pattern of amino acid residues shown in Table 2 above.
[0125] A second aspect of the present invention relates to a modified, in particular recombinant, polynucleotide or nucleic acid comprising a nucleotide sequence encoding at least one of the modified archaeal pyrrolysyl-tRNA synthetases or functional fragments as defined above according to the first aspect; as well as corresponding complementary polynucleotides; and fragments thereof which hybridize under stringent conditions as defined herein with any of the above polynucleotides.
[0126] Particular non-limiting examples of encoding polynucleotides include a nucleotide sequence selected from SEQ ID NOs: 69, 71, 77, 79, 81, 83, 85, 87, 89, 91, 93 and 95, or a nucleotide sequence having at least 60%, e.g. 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to a nucleotide sequence selected from SEQ ID NOs: 69, 71, 77, 79, 81, 83, 85, 87, 89, 91, 93 and 95, and still encode a functional PylRS mutant of the invention.
[0127] Further examples include nucleic acid sequences encoding any of the above motifs according to SEQ ID NOs: 7 to 53, or nucleic acid sequences encoding PylRS mutants containing at least one of the motifs, or more specifically, combinations of the motifs disclosed above.
[0128] In certain embodiments, the polynucleotide is a tRNA Pyl (SEQ ID NO: 113, optionally without the terminal CCA triplet), Pyl is a tRNA that can be acylated by the encoded pyrrolysyl-tRNA synthetase mutant of the present invention, as defined above.
[0129] In a further particular embodiment of the second aspect, at least one polynucleotide encoding the PylRS mutant of the invention and the tRNA Pyl A combination of polynucleotides is provided, comprising at least one polynucleotide encoding:
[0130] In a further particular embodiment, the tRNA Pyl The anticodon of is the reverse complement of a codon selected from a stop codon, a four base codon, and a rare codon.
[0131] Polynucleotides of the invention and tRNAs used in the context of the invention Pyl and / or the polynucleotide encoding the POI is preferably an expression system, such as an expression cassette.
[0132] The present invention also provides a method for transfecting eukaryotic cells with the encoded PylRS mutant, tRNA Pyl and / or a vector suitable for expressing the POI, respectively, in said cells are also provided.
[0133] In particular, such a vector is an expression vector comprising a recombinant expression system as defined above; or one of the above nucleic acids or its reverse complement; or
[0134] The expression vector of the present invention is a prokaryotic vector, a viral vector, a eukaryotic vector, or a plasmid.
[0135] A third aspect of the invention relates to a eukaryotic cell comprising: (a) a polynucleotide sequence encoding at least one of the pyrrolysyl-tRNA synthetase mutants according to the first aspect; and (b) a tRNA that can be acylated by the mutant pyrrolysyl-tRNA synthetase encoded by the sequence of (a), or a polynucleotide sequence encoding such a tRNA.
[0136] In particular, said eukaryotic cell is a mammalian cell.
[0137] The present invention particularly provides a mammalian cell capable of expressing a PylRS mutant of the present invention. In particular, the present invention provides a mammalian cell comprising a polynucleotide or a combination of polynucleotides, said polynucleotide comprising a PylRS mutant of the present invention and a tRNA Pyl Encoding tRNA Pyl is a tRNA that can be acylated (preferably selectively) by PylRS. Advantageously, the mammalian cell of the present invention comprises a tRNA Pyl and the PylRS mutant of the present invention, wherein the PylRS mutant is a tRNA Pyl can be (preferably selectively) acylated with an amino acid, such as UNAA.
[0138] A fourth aspect of the invention relates to a method for preparing a POI comprising one or more unnatural amino acid residues, the method comprising the steps of: (a) providing a eukaryotic cell comprising: (i) the pyrrolysyl-tRNA synthetase mutant of the first aspect; (ii) tRNA (tRNA Pyl ); (iii) an unnatural amino acid or a salt thereof; and (iv) a polynucleotide encoding a POI, wherein any position in the POI that is occupied by a non-natural amino acid residue is not Pyl a polynucleotide encoded by a codon that is the reverse complement of an anticodon contained in Here, the pyrrolysyl-tRNA synthetase mutant (i) is a tRNA Pyl (ii) can be acylated with an unnatural amino acid or salt (iii); and (b) allowing translation of the polynucleotide (iv) by a eukaryotic cell, thereby producing the POI.
[0139] A fifth aspect of the invention relates to a method for preparing a polypeptide conjugate comprising the steps of: (a) preparing a POI comprising one or more unnatural amino acid residues using a method according to the fourth aspect; and (b) reacting the POI with one or more binding partner molecules such that the binding partner molecules are covalently attached to the non-natural amino acid residues of the POI.
[0140] In certain embodiments, a POI of the present invention having at least one ncAA incorporated into its amino acid sequence is utilized to form a "targeting agent."
[0141] The first purpose of the targeting agent is to form a covalent or non-covalent bond with a specific "target". The second purpose of the targeting agent is the targeted delivery of a "payload molecule" to the target. To achieve the second purpose, the POI must bind (reversibly or irreversibly) to at least one payload molecule. For this purpose, the POI must be functionalized by introducing the at least one ncAA. The functionalized POI with the at least one ncAA can then be linked to the at least one payload molecule by bioconjugation via the ncAA residue. The ncAA is reactive with the payload molecule, and the payload molecule carries a corresponding site reactive with the ncAA residue of the POI. The resulting bioconjugate, i.e. the targeting agent, transfers the payload molecule to the intended target.
[0142] Specific examples of ncAAs to be incorporated into the POI include, but are not limited to, bulky ncAAs as mentioned herein. More particularly, such residues include bulky ncAAs such as TCO * A, TCO-E, Boc, and SCO * or salts thereof. Specific non-limiting examples of suitable types of POIs, potential targets, potential payloads and potential reactive groups suitable for bioconjugation with such ncAAs are described below.
[0143] In particular, the POI is an immunoglobulin, in particular an antibody molecule, more particularly a monoclonal antibody or a functional antigen binding derivative or fragment thereof, having at least one ncAA in its amino acid sequence.
[0144] In another particular embodiment, the target is a tumor-associated target and the payload molecule is an anti-tumor agent.
[0145] A sixth aspect of the invention relates to kits that can be used in the methods of preparing a UNAA residue-containing POI or conjugate thereof described herein.
[0146] Such a kit may comprise at least one unnatural amino acid, or a salt thereof; and / or a polynucleotide or a combination of polynucleotides as defined above according to the second aspect of the invention, or a eukaryotic cell according to the third aspect of the invention, wherein the encoded archaeal PylRS is a tRNA Pyl can be acylated with an unnatural amino acid or its salt.
[0147] For example, the kit includes a polynucleotide encoding a PylRS mutant of the present invention or a eukaryotic cell capable of expressing that PylRS, and further includes at least one unnatural amino acid or a salt thereof that can be used to acylate tRNA in a reaction catalyzed by the PylRS.
[0148] The specific amino acid sequences of the PylRS of the present invention referred to herein are shown below.
[0149] Specific examples of sequences of PylRS mutants of the invention: The NES sequence is double underlined. Mutations are shown in bold and underlined; numbering of mutation positions relative to the wild type sequence (SEQ ID NO:56).
[0150] Certain PylRS mutants of the present invention are derivable from M. mazei wild-type PylRS, or an enzymatically active fragment thereof.
[0151] The amino acid sequence of wild-type M. mazei PylRS is set forth in SEQ ID NO:56.
[0152] [Table 15]
[0153] In the sequences of the PylRS mutants below, the NES sequence is double underlined and the mutated amino acid residues are bolded and underlined.
[0154] [Table 16]
[0155] [Table 17]
[0156] [Table 18]
[0157] [Table 19]
[0158] D. Additional Embodiments 1. Polypeptides of the Invention In this context, the following definitions apply: "Functional variants" of the polypeptides described herein include "functional equivalents" of said polypeptides, as defined below.
[0159] An "enzyme", "protein" or "polypeptide" as described herein may be a naturally occurring or recombinantly produced enzyme, protein or polypeptide, which may be a wild-type enzyme, protein or polypeptide or may be genetically modified by suitable mutations or by C-terminal and / or N-terminal amino acid sequence extensions, such as sequences including His-tags. The enzyme, protein or polypeptide may essentially be mixed with cellular, e.g. protein, impurities, but is particularly in pure form. Suitable detection methods are described, for example, in the experimental section below or are known from the literature.
[0160] An enzyme, protein or polypeptide "in pure form" or "pure" or "substantially pure" is understood as an enzyme, protein or polypeptide having a purity, measured according to the usual methods for detecting proteins (e.g. the Biuret method or the Protein Detection according to Lowry et al.), in accordance with the present invention of more than 80, in particular more than 90, in particular more than 95 and very particularly more than 99 wt. % based on the total protein content (see RKScopes, Protein Purification, Springer Verlag, New York, Heidelberg, Berlin (1982)).
[0161] "Proteinogenic" amino acids include in particular the following (single letter code): G, A, V, L, I, F, P, M, W, S, T, C, Y, N, Q, D, E, K, R and H.
[0162] The general terms "polypeptide" or "peptide" may be used interchangeably and refer to a natural or synthetic linear chain or sequence of consecutive peptide-linked amino acid residues containing from about 10 up to 1,000 residues. Shorter chains of polypeptides, up to 30 residues, are also called "oligopeptides."
[0163] The term "protein" refers to a macromolecular structure consisting of one or more polypeptides. The amino acid sequence of a polypeptide represents the "primary structure" of the protein. The amino acid sequence also predetermines the "secondary structure" of the protein by the formation of special structural elements such as α-helices and β-sheets formed within the polypeptide chain. The arrangement of multiple such secondary structural elements defines the "tertiary structure" or spatial arrangement of the protein. When a protein comprises multiple polypeptide chains, said chains are spatially arranged to form the "quaternary structure" of the protein. The correct spatial arrangement or "folding" of a protein is a prerequisite for the function of the protein. Denaturation or unfolding destroys the function of the protein. If such destruction is reversible, the function of the protein may be restored by refolding.
[0164] A typical protein function referred to herein is an "enzymatic function", i.e., the protein acts as a biocatalyst on a substrate, e.g., a compound, catalyzing the conversion of said substrate to a product. Enzymes may have high or low substrate and / or product specificity.
[0165] Thus, a "polypeptide" referred to herein as having a particular "activity" implicitly refers to a correctly folded protein that exhibits the indicated activity, e.g., a particular enzymatic activity.
[0166] Thus, unless otherwise specified, the term "polypeptide" also includes the terms "protein" and "enzyme."
[0167] Similarly, the term "polypeptide fragment" encompasses the terms "protein fragment" and "enzyme fragment."
[0168] The term "isolated polypeptide" refers to an amino acid sequence that has been removed from its natural environment by any method or combination of methods known in the art, including recombinant, biochemical, and synthetic methods.
[0169] The present invention also relates to "functional equivalents" (also called "analogs" or "functional variants") of the polypeptides specifically described herein.
[0170] For example, a "functional equivalent" refers to a polypeptide that exhibits at least 1-10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% greater or less activity in a test used to determine enzymatic NHase activity than the activity of a polypeptide specifically described herein.
[0171] According to the invention, "functional equivalents" also encompass specific variants which have an amino acid different from the one specifically mentioned in at least one sequence position of the amino acid sequences described herein, but which nevertheless have one of the aforementioned biological activities, such as, for example, enzymatic activity. Thus, "functional equivalents" include variants obtained by addition, substitution, in particular conservative substitution, deletion and / or inversion of one or more, for example 1 to 20, in particular 1 to 15 or 5 to 10, amino acids, the changes mentioned may occur at any sequence position, provided that they result in a variant with a property profile according to the invention. Functional equivalence is in particular observed when the activity pattern is qualitatively consistent between the variant and the unchanged polypeptide, i.e. when, for example, interaction with the same agonist or antagonist or substrate is observed, but at different rates (i.e., EC 50 Value or IC 50 The amino acid substitutions (wherein the substitutions are expressed by a value of 0, 1, 2, 3, 4, 5, 6, 7, or any other parameter suitable in the current art) are also provided in Table 3 below.
[0172] [Table 20]
[0173] "Functional equivalents" in the above sense are also "precursors" of the polypeptides described herein, as well as "functional derivatives" and "salts" of the polypeptides.
[0174] In that case, a "precursor" is a natural or synthetic precursor of a polypeptide, which may or may not have the desired biological activity.
[0175] The expression "salts" refers to salts of carboxyl groups and acid addition salts of amino groups of the protein molecules according to the invention. The salts of carboxyl groups can be prepared in a known manner and include inorganic salts, such as sodium salts, calcium salts, ammonium salts, iron salts and zinc salts, as well as salts with organic bases, such as amines, for example triethanolamine, arginine, lysine, piperidine. Acid addition salts, such as salts with inorganic acids, such as hydrochloric acid or sulfuric acid, and salts with organic acids, such as acetic acid or oxalic acid, are also within the scope of the present invention.
[0176] "Functional derivatives" of the polypeptides according to the invention can also be prepared at functional amino acid side groups or at their N- or C-termini using known techniques. Such derivatives include, for example, aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups obtained by reaction with ammonia or primary or secondary amines, N-acyl derivatives of free amino groups produced by reaction with acyl groups, or O-acyl derivatives of free hydroxyl groups produced by reaction with acyl groups.
[0177] "Functional equivalents" naturally include polypeptides obtainable from other organisms as well as naturally occurring variants. For example, regions of homologous sequence regions can be established by sequence comparison and equivalent polypeptides can be determined based on the specific parameters of the invention.
[0178] "Functional equivalents" also include "fragments" such as individual domains or sequence motifs of the polypeptides according to the invention, or N- and C-terminal truncations which may or may not exhibit a desired biological function. In particular, such "fragments" retain at least a qualitative portion of the desired biological function.
[0179] Furthermore, a "functional equivalent" is a fusion protein having one of the polypeptide sequences described herein, or a functional equivalent derived therefrom, and at least one additional, functionally distinct, heterologous sequence at a functional N- or C-terminal linkage (i.e., without substantially impairing the mutual function of the fusion protein moieties). Non-limiting examples of these heterologous sequences are, for example, signal peptides, histidine anchors or enzymes.
[0180] "Functional equivalents" included according to the present invention are homologues to the specifically disclosed polypeptides. These have at least 50, 55% or 60%, in particular at least 75%, more particularly at least 80 or 85%, for example 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% homology (or identity) to one of the specifically disclosed amino acid sequences, calculated according to the algorithm of Pearson and Lipman (Proc. Natl. Acad, Sci. (USA) 85(8), 1988, 2444-2448). The percentage homology or identity of the homologous polypeptides according to the present invention means in particular the identity expressed as a percentage of amino acid residues based on the full length of one of the amino acid sequences specifically described herein.
[0181] Identity data, expressed as percentages, can also be determined by BLAST alignment, making use of the algorithm blastp (protein-protein BLAST) or by applying the Clustal settings as defined herein below.
[0182] In the case of possible glycosylation of the protein, "functional equivalents" according to the present invention include the polypeptides described herein in deglycosylated or glycosylated form, as well as modified forms which can be obtained by altering the glycosylation pattern.
[0183] Functional equivalents or homologues of the polypeptides according to the invention can be produced by mutagenesis, eg point mutations, lengthening or shortening of the protein, or as described in more detail below.
[0184] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a database of combinations of mutants, such as truncation mutants. For example, a diverse database of protein mutants can be created by combinatorial mutagenesis at the nucleic acid level, such as by enzymatic ligation of a mixture of synthetic oligonucleotides. There are numerous methods that can be used to create a database of potential homologs from degenerate oligonucleotide sequences. Chemical synthesis of degenerate gene sequences can be performed in an automatic DNA synthesizer, and the synthetic genes can then be ligated into a suitable expression vector. The use of degenerate genomes makes it possible to provide all sequences in a mixture that code for a set of desired potential protein sequences. Methods for synthesizing degenerate oligonucleotides are well known to those skilled in the art.
[0185] In the prior art, several techniques are known for screening gene products of combinatorial databases made by point mutations or truncations, and for screening cDNA libraries for gene products with selected properties. These techniques can be adapted for the rapid screening of gene banks produced by combinatorial mutagenesis of homologs according to the invention. The most frequently used technique for screening large gene banks based on high-throughput analysis involves cloning the gene bank in replicable expression vectors, transformation of suitable cells with the resulting vector database, and expression of the combined genes under conditions that facilitate the isolation of the vectors encoding the genes whose products were detected upon detection of the desired activity. Recursive ensemble mutagenesis (REM) is a technique that increases the frequency of functional variants in a database and can be used in combination with screening tests to identify homologs.
[0186] The polypeptides of the present invention include all active forms including active subsequences of the enzymes of the present invention, such as catalytic domains or active sites. In one embodiment, the present invention provides catalytic domains or active sites as defined below. In one embodiment, the present invention provides peptides or polypeptides that include or consist of active site domains predicted, for example, by use of Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) or equivalent databases. Pfam is a large collection of multiple sequence alignments and hidden Markov models covering many common protein families. [The Pfam protein family database (A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. R. Addy, S. Griffiths-Jones, K. L. Howe, M. Marshall, and E. L. Sonnhammer, Nucleic Acids Research, 30(1):276-280, 2002)] Equivalent databases are, for example, the InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl-heidelberg.de / ).
[0187] The present invention also encompasses "polypeptide variants" having a desired activity, said variant polypeptides being selected from amino acid sequences having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to a particular, particularly naturally occurring, amino acid sequence referenced by a particular SEQ ID NO, and containing at least one substitution modification relative to said SEQ ID NO.
[0188] 2. Nucleic acids and constructs 2.1 Nucleic acids In this context, the following definitions apply: The terms "nucleic acid sequence", "nucleic acid", "nucleic acid molecule" and "polynucleotide" are used interchangeably to mean a sequence of nucleotides. A nucleic acid sequence can be a single- or double-stranded deoxyribonucleotide or ribonucleotide of any length, and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified natural DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers and nucleic acid probes. Those skilled in the art will recognize that the nucleic acid sequence of RNA is identical to the DNA sequence, with the difference being that thymine (T) is replaced by uracil (U). The term "nucleotide sequence" should also be understood to include polynucleotide or oligonucleotide molecules, either in the form of a separate fragment or as a component of a larger nucleic acid.
[0189] An "isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that is in an environment other than that in which it naturally occurs, and can include one that is substantially free of contaminating endogenous material.
[0190] The term "naturally occurring" as used herein as applied to nucleic acids refers to nucleic acids found within the cells of organisms in nature and which have not been intentionally modified by man in the laboratory.
[0191] A "fragment" of a polynucleotide or nucleic acid sequence refers to a contiguous nucleotide of the polynucleotide of the embodiment herein, particularly at least 15bp, at least 30bp, at least 40bp, at least 50bp and / or at least 60bp in length.In particular, a fragment of a polynucleotide comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, more particularly at least 1000 contiguous nucleotides of the polynucleotide of the embodiment herein.Without being limited thereto, a fragment of a polynucleotide of the embodiment herein can be used as a PCR primer and / or as a probe, or for antisense gene silencing or RNAi.
[0192] As used herein, the term "hybridization" or hybridizing under certain conditions is intended to describe hybridization and washing conditions under which nucleotide sequences that are significantly identical or homologous to each other remain bound to each other. The conditions can be such that sequences that are at least about 70%, such as at least about 80%, such as at least about 85%, 90%, or 95% identical to each other remain bound to each other. Definitions of low stringency, medium stringency, and high stringency hybridization conditions are provided below. Appropriate hybridization conditions can also be selected by those skilled in the art with minimal experimentation, as exemplified by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). In addition, stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0193] A "recombinant nucleic acid sequence" is a nucleic acid sequence that results from the collection of genetic material from multiple sources using laboratory techniques (e.g., molecular cloning) to create or modify a nucleic acid sequence that does not occur in nature and is not present in a biological organism.
[0194] "Recombinant DNA techniques" refers to molecular biology procedures for preparing recombinant nucleic acid sequences, e.g., as described in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press and Sambrook et al., 1989, Cold Spring Harbor, NY, Cold Spring Harbor Laboratory Press.
[0195] The term "gene" refers to a DNA sequence that is transcribed into an RNA molecule, e.g., mRNA, in a cell and that comprises a region operably linked to an appropriate regulatory region, e.g., a promoter. Thus, a gene can include several operably linked sequences, such as a promoter, a 5' leader sequence, e.g., including sequences involved in translation initiation, a coding region of cDNA or genomic DNA, introns, exons, and / or 3' untranslated sequences, e.g., including a transcription termination site.
[0196] "Polycistronic" refers to a nucleic acid molecule, particularly an mRNA, that can separately encode multiple polypeptides within the same nucleic acid molecule.
[0197] "Chimeric gene" refers to any gene that is not normally found in nature in a species, particularly a gene in which there are one or more portions of nucleic acid sequences that are not naturally associated with each other. For example, a promoter is not naturally associated with part or all of the transcription region or with another regulatory region. The term "chimeric gene" is understood to include an expression construct in which a promoter or transcription regulatory sequence is operably linked to one or more coding sequences, or antisense, i.e., the reverse complement of the sense strand, or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms a double-stranded RNA upon transcription). The term "chimeric gene" also includes genes that are obtained by combining parts of one or more coding sequences to produce a novel gene.
[0198] "3'UTR" or "3' untranslated sequence" (also called "3' untranslated region" or "3' end") refers to nucleic acid sequences found downstream of the coding sequence of a gene, including, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal, such as AAUAAA or a variant thereof. After transcription termination, the mRNA transcript may be cleaved downstream of the polyadenylation signal and a poly(A) tail may be added, which is involved in transport of the mRNA to a site of translation, such as the cytoplasm.
[0199] The term "primer" refers to a short nucleic acid sequence that hybridizes to a template nucleic acid sequence and is used to polymerize a nucleic acid sequence that is complementary to the template.
[0200] The term "selectable marker" refers to any gene that upon expression can be used to select cells containing the selectable marker. Examples of selectable markers are provided below. One of skill in the art will know that a variety of antibiotic, fungicide, nutritional supplement or herbicide selectable markers are applicable to a variety of target species.
[0201] The present invention also relates to nucleic acid sequences which code for the polypeptides defined herein.
[0202] In particular, the present invention also relates to nucleic acid sequences (single- and double-stranded DNA as well as RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the above-mentioned polypeptides and their functional equivalents, which can be obtained, for example, using artificial nucleotide analogues.
[0203] The present invention relates both to isolated nucleic acid molecules encoding a polypeptide according to the invention or a biologically active segment thereof, and to nucleic acid fragments which can be used, for example, as hybridization probes or primers to identify or amplify the encoding nucleic acid according to the invention.
[0204] The present invention also relates to nucleic acids that have a degree of "identity" to the sequences specifically disclosed herein. "Identity" between two nucleic acids means in each case nucleotide identity over the entire length of the nucleic acid.
[0205] "Identity" between two nucleotide sequences (also peptide or amino acid sequences) is a function of the number of nucleotide residues, or the number of nucleotide residues (or amino acid residues) that are identical in the two sequences when these two sequences are aligned. An identical residue is defined as a residue that is identical in the two sequences at a given position in the alignment. As used herein, the percentage of sequence identity is calculated from the optimal alignment by taking the number of identical residues between the two sequences, dividing it by the total number of residues in the shortest sequence, and multiplying by 100. The optimal alignment is the alignment in which the percentage of identity is as high as possible. To obtain optimal alignment, gaps may be introduced in one or both sequences at one or more positions in the alignment. These gaps are considered as non-identical residues when calculating the percentage of sequence identity. Alignment for the purpose of determining the percentage of identity of amino acid or nucleic acid sequences can be achieved in a variety of ways using computer programs, such as publicly available computer programs available on the World Wide Web.
[0206] In particular, the BLAST program (Tatiana et al, FEMS Microbiol Lett., 1999, 174:247-250, 1999), available from the National Center for Biotechnology Information (NCBI) website (ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi) set to default parameters, can be used to obtain optimal alignments of protein or nucleic acid sequences and to calculate the percentage of sequence identity.
[0207] In another example, identity can be calculated using the Clustal method (Higgins DG, Sharp PM. (1989)) with the following settings using the Vector NTI Suite 7.1 program from Informax (USA):
[0208] [Table 21]
[0209] [Table 22]
[0210] Alternatively, identity can be determined by following the webpage of Chenna et al. (2003): http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the settings below.
[0211] [Table 23]
[0212] All nucleic acid sequences described herein (single- and double-stranded DNA and RNA sequences, e.g., cDNA and mRNA) can be prepared by chemical synthesis from nucleotide building blocks, for example by fragment condensation of individual overlapping complementary nucleic acid building blocks of a double helix, in a known manner. Chemical synthesis of oligonucleotides can be carried out in a known manner, for example by the phosphoramidite method (Voet, Voet, 2nd ed., Wiley Press, New York, pp. 896-897). The accumulation and gap-filling of synthetic oligonucleotides using the Klenow fragment of DNA polymerase and ligation reactions as well as general cloning techniques are described in Sambrook et al. (1989) (see below).
[0213] A nucleic acid molecule according to the present invention can further include untranslated sequences from the 3' and / or 5' ends of the coding gene region.
[0214] The present invention further relates to nucleic acid molecules which are complementary to the specifically described nucleotide sequences or segments thereof.
[0215] The nucleotide sequences according to the invention allow the production of probes and primers which can be used for the identification and / or cloning of homologous sequences in other cell types and organisms. Such probes or primers usually comprise a nucleotide sequence region which hybridizes under "stringent" conditions (as defined elsewhere herein) to at least about 12, in particular at least about 25, such as about 40, 50 or 75 consecutive nucleotides of the sense strand of a nucleic acid sequence according to the invention or the corresponding antisense strand.
[0216] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologs or paralogs, including phylogenetic methods, sequence similarity and hybridization methods, are known in the art and described herein.
[0217] "Paralogues" result from gene duplication resulting in two or more genes with similar sequence and similar function. Paralogs are usually formed by gene duplication within closely related plant species and exist in clusters. Paralogs are detected in groups of similar genes using pairwise Blast analysis or by phylogenetic analysis of gene families using programs such as CLUSTAL. In paralogs, consensus sequences may be identified that are characteristic of sequences within related genes and that are similar in function of the genes.
[0218] "Orthologs" or orthologous sequences are sequences that are similar to each other because they are found in species derived from a common ancestor. For example, plant species that share a common ancestor are known to contain many enzymes with similar sequences and functions. A person skilled in the art can identify orthologous sequences and predict their functions by constructing a multi-gene tree for a gene family of a species, for example, using CLUSTAL or BLAST programs. A method for identifying or confirming similar functions between homologous sequences is by comparing the transcriptional profiles when the related polypeptide is overexpressed or deleted (knocked out / knocked down) in a host cell or organism, such as a plant or microorganism. A person skilled in the art will understand that genes with similar transcriptional profiles having more than 50% common regulated transcripts, or more than 70% common regulated transcripts, or more than 90% common regulated transcripts have similar functions. Homologs, paralogs, orthologs and any other variants of the sequences herein are expected to function similarly by producing the enzymes of the present invention in a host cell, organism such as a plant, or microorganism.
[0219] The term "selection marker" refers to any gene that, upon expression, allows the selection of cells containing the selection marker. Examples of selection markers are listed below. Those skilled in the art will know that different antibiotic, fungicide, auxotrophic or herbicide selection markers are applicable to different target species.
[0220] The nucleic acid molecules according to the invention can be recovered using standard techniques of molecular biology and the sequence information provided according to the invention. For example, cDNA can be isolated from a suitable cDNA library using one of the specifically disclosed complete sequences or a segment thereof as a hybridization probe and standard hybridization techniques (for example as described in Sambrook (1989)).
[0221] Furthermore, a nucleic acid molecule comprising one of the disclosed sequences or a segment thereof can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on this sequence. The nucleic acid thus amplified can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be produced by standard synthesis methods, for example using an automatic DNA synthesizer.
[0222] The nucleic acid sequences according to the invention or derivatives, homologues or parts of these sequences can be isolated from other bacteria, for example via genomic or cDNA libraries, by conventional hybridization or PCR techniques, which DNA sequences hybridize under standard conditions with the sequences according to the invention.
[0223] "Hybridize" means the ability of a polynucleotide or oligonucleotide to bind to a nearly complementary sequence under standard conditions, while non-specific binding between non-complementary partners does not occur under these conditions. In this respect, the sequences can be 90-100% complementary. The property of complementary sequences to be able to bind specifically to each other is exploited, for example, in Northern or Southern blotting or in primer binding in PCR or RT-PCR.
[0224] Short oligonucleotides of the conserved regions are preferably used for hybridization. However, it is also possible to use longer fragments or complete sequences of the nucleic acids according to the invention for hybridization. These "standard conditions" vary depending on the nucleic acid used (oligonucleotide, longer fragment or complete sequence) or on the type of nucleic acid used for hybridization (DNA or RNA). For example, the melting temperature of DNA:DNA hybrids is about 10°C lower than that of DNA:RNA hybrids of the same length.
[0225] For example, depending on the specific nucleic acid, standard conditions means a temperature of 42-58° C. in an aqueous buffer with a concentration ranging from 0.1-5×SSC (1×SSC=0.15 M NaCl, 15 mM sodium citrate, pH 7.2) or even in the presence of 50% formamide, for example 42° C., 50% formamide in 5×SSC. Advantageously, the hybridization conditions for DNA:DNA hybrids are 0.1×SSC and a temperature of about 20° C.-45° C., preferably about 30° C.-45° C. For DNA:RNA hybrids, the hybridization conditions are advantageously 0.1×SSC and a temperature of about 30° C.-55° C., in particular a temperature of about 45° C.-55° C. These stated temperatures for hybridization are examples of calculated melting temperatures for a nucleic acid of about 100 nucleotides in length and with a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization are described in relevant genetics textbooks, e.g., Sambrook et al. (1989), and can be calculated using formulas known to those skilled in the art, depending on, e.g., the length of the nucleic acid, the type of hybrid, or the G+C content. Those skilled in the art can obtain further information on hybridization from the following textbooks: Ausubel et al. (eds.) (1985); Brown (ed.) (1991).
[0226] "Hybridization" can be carried out under particularly stringent conditions, such as those described in Sambrook (1989) or Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1-6.3.6.
[0227] As used herein, the term hybridization under specific conditions or hybridizing under specific conditions is intended to describe hybridization and washing conditions under which nucleotide sequences that are significantly identical or homologous to each other remain bound to each other. The conditions may be such that sequences that are at least about 70%, such as at least about 80%, and such as at least about 85%, 90%, or 95% identical remain bound to each other. Definitions of low stringency, medium, and high stringency hybridization conditions are described herein.
[0228] Appropriate hybridization conditions can be selected by one of skill in the art with minimal experimentation, as exemplified by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Further stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0229] As used herein, the defined conditions of low stringency are as follows: DNA-containing filters are pretreated for 6 hours at 40°C in a solution containing 35% formamide, 5X SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization is performed in the same solution, but with the following modifications: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and 5-20X10 632P-labeled probe is used. Filters are incubated in the hybridization mixture for 18-20 hours at 40°C, then washed for 1.5 hours at 55°C in a solution containing 2XSSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. The wash solution is replaced with fresh solution, and then incubated for an additional 1.5 hours at 60°C. Filters are blotted dry and subjected to autoradiography.
[0230] As used herein, the defined conditions of moderate stringency are as follows: DNA-containing filters are pretreated for 7 hours at 50°C in a solution containing 35% formamide, 5x SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization is performed in the same solution with the following modifications: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and 5-20x10 6 of 32P-labeled probe is used. Filters are incubated in hybridization mixture for 30 hours at 50°C and then washed for 1.5 hours at 55°C in a solution containing 2xSSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA and 0.1% SDS. The wash solution is replaced with fresh solution and incubated for an additional 1.5 hours at 60°C. Filters are blotted dry and subjected to autoradiography.
[0231] As used herein, the defined conditions of high stringency are as follows: Prehybridization of DNA-containing filters is carried out for 8 hours to overnight at 65°C in a buffer consisting of 6xSSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 μg / ml denatured salmon sperm DNA. Filters are prehybridized with 100 μg / ml denatured salmon sperm DNA and 5-20×10 6The filters are hybridized in prehybridization mixture containing 100 cpm of 32P-labeled probe at 65° C. for 48 hours. The filters are washed in a solution containing 2×SSC, 0.01% PVP, 0.01% Ficoll and 0.01% BSA at 37° C. for 1 hour, followed by a wash in 0.1×SSC at 50° C. for 45 minutes.
[0232] If the above conditions are inappropriate (e.g., for use in cross-species hybridizations), other conditions of low, medium or high stringency well known in the art (e.g., for use in cross-species hybridizations) may be used.
[0233] A detection kit for a nucleic acid sequence encoding a polypeptide of the present invention may comprise primers and / or probes specific for the nucleic acid sequence encoding the polypeptide and an associated protocol for detecting the nucleic acid sequence encoding the polypeptide in a sample using the primers and / or probes. Such a detection kit can be used to determine whether a plant, organism, microorganism or cell has been modified, i.e., transformed with a sequence encoding the polypeptide.
[0234] To test the functionality of a mutant DNA sequence according to embodiments herein, the sequence of interest is operably linked to a selectable or screenable marker gene and expression of the reporter gene is tested in transient expression assays, for example using microorganisms or protoplasts or in stably transformed plants.
[0235] The present invention also relates to derivatives of the specifically disclosed or derivable nucleic acid sequences.
[0236] Thus, further nucleic acid sequences according to the invention can be derived from the sequences specifically disclosed herein and can differ therefrom by one or more, for example 1-20, in particular 1-15 or 5-10 additions, substitutions, insertions or deletions of one or several (for example 1-10) nucleotides and can further code for a polypeptide with a desired property profile.
[0237] The present invention also encompasses nucleic acid sequences which contain so-called silent mutations or which, compared to a particular sequence, have been modified according to the codon usage of a particular original or host organism.
[0238] According to certain embodiments of the present invention, mutated nucleic acids can be prepared to adapt their nucleotide sequences to specific expression systems. For example, bacterial expression systems are known to express polypeptides more efficiently if the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, multiple codons may code for the same amino acid sequence, and multiple nucleic acid sequences may code for the same protein or polypeptide. All of these DNA sequences are encompassed by the embodiments of the present specification. If necessary, the nucleic acid sequences encoding the polypeptides described herein can be optimized to increase expression in host cells. For example, the nucleic acids of the embodiments of the present specification can be synthesized with codons specific to the host to improve expression.
[0239] The present invention also encompasses naturally occurring variants of the sequences presented herein, such as splice variants or allelic variants.
[0240] Allelic variants may have at least 60% homology at the level of amino acid derivatives, particularly at least 80%, more particularly at least 90% homology over the entire sequence range (for homology at the amino acid level, see details above for polypeptides). Advantageously, the homology may be higher in partial regions of the sequences.
[0241] The invention also relates to sequences which can be obtained by conservative nucleotide substitutions, ie whereby the amino acid in question is replaced by an amino acid of the same charge, size, polarity and / or solubility.
[0242] The present invention also relates to molecules derived from the specifically disclosed nucleic acids by sequence polymorphisms. Such genetic polymorphisms may exist in cells from different populations or within a population due to natural allelic variation. Allelic variations may also include functional equivalents. These natural variations usually result in 1-5% differences in the nucleotide sequence of a gene. The polymorphisms may result in changes in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.
[0243] Moreover, derivatives will also be understood to be homologues of the nucleic acid sequences according to the invention, such as animal, plant, fungal or bacterial homologues, truncated sequences, single stranded DNA or RNA of coding and non-coding DNA sequences, for example, homologues have at the DNA level at least 40%, particularly 60%, particularly particularly 70%, very particularly 80% homology over the entire DNA region given in the sequences specifically disclosed herein.
[0244] Furthermore, derivatives are understood to be, for example, fusions with promoters. The promoters added to the described nucleotide sequences can be modified by at least one nucleotide exchange, at least one insertion, inversion and / or deletion without impairing the functionality or effectiveness of the promoter. Furthermore, the effectiveness of promoters can be increased by changing their sequences or completely replaced with more effective promoters, even in organisms of different genera.
[0245] 2.2. Constructs for expressing the polypeptides of the invention In this context the following definitions apply: "Expression of a gene" includes "heterologous expression" and "overexpression" and includes transcription of a gene and translation of mRNA into protein. Overexpression refers to the production of a gene product, as measured by levels of mRNA, polypeptide and / or enzymatic activity, in a transgenic cell or organism that exceeds the level of production in a non-transformed cell or organism with a similar genetic background.
[0246] As used herein, "expression vector" refers to a nucleic acid molecule designed using molecular biology methods and recombinant DNA technology to deliver foreign or exogenous DNA to a host cell. An expression vector usually contains sequences necessary for proper transcription of a nucleotide sequence. The coding region usually codes for a protein of interest, but can also code for RNA, such as antisense RNA, siRNA, etc.
[0247] As used herein, "expression vector" includes any linear or circular recombinant vector, including, but not limited to, viral vectors, bacteriophages, and plasmids. Those skilled in the art can select an appropriate vector depending on the expression system. In one embodiment, the expression vector comprises a nucleic acid of an embodiment of the present invention operably linked to at least one "regulatory sequence" that controls transcription, translation, initiation, and termination, such as a transcription promoter, operator or enhancer, or mRNA ribosomal binding site, and optionally includes at least one selectable marker. A nucleotide sequence is "operably linked" when the regulatory sequence is functionally related to the nucleic acid of an embodiment of the present invention.
[0248] As used herein, an "expression system" encompasses any combination of nucleic acid molecules required for the expression or simultaneous expression of two or more polypeptides in vivo or in vitro in a given expression host. Each coding sequence may be located either on a single nucleic acid molecule or on a vector, e.g., a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed among two or more physically distinct vectors. A particular example may be an operon, comprising a promoter sequence, one or more operator sequences, and one or more structural genes, each encoding an enzyme as described herein.
[0249] As used herein, the terms "amplify" and "amplification" refer to the use of any suitable amplification method to generate or detect modifications of naturally expressed nucleic acids, as described in detail below. For example, the invention provides methods (e.g., by polymerase chain reaction, PCR) and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo dT primers) for amplifying naturally expressed nucleic acids (e.g., genomic DNA or mRNA) or recombinant nucleic acids (e.g., cDNA) of the invention in vivo, ex vivo or in vitro.
[0250] "Regulatory sequence" refers to a nucleic acid sequence that determines the expression level of a nucleic acid sequence according to an embodiment of the present specification and is capable of regulating the transcription rate of a nucleic acid sequence operably linked to the regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, and the like.
[0251] A "promoter", "nucleic acid having promoter activity" or "promoter sequence" is understood according to the invention to mean a nucleic acid which, when functionally linked to a nucleic acid to be transcribed, regulates the transcription of said nucleic acid. A "promoter" refers in particular to a nucleic acid sequence which controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors necessary for proper transcription. These other factors include, but are not limited to, transcription factor binding sites, repressor and activator protein binding sites. The meaning of the term promoter also includes the term "promoter regulatory sequence". Promoter regulatory sequences may include upstream and downstream elements that may affect the transcription, RNA processing or stability of the associated coding nucleic acid sequence. Promoters include naturally occurring and synthetic sequences. The coding nucleic acid sequence is usually located downstream of the promoter with respect to the direction of transcription starting from the transcription start site.
[0252] In this context, a "functional" or "operably" linkage is understood to mean, for example, a contiguous arrangement of one of the nucleic acids with a regulatory sequence. For example, a sequence having promoter activity and a sequence of the nucleic acid sequence to be transcribed, and optionally further regulatory elements, such as a nucleic acid sequence ensuring the transcription of the nucleic acid, such as a terminator, are linked in such a way that each of the regulatory elements can perform its function during the transcription of the nucleic acid sequence. This does not necessarily require a direct link in the chemical sense. A genetic control sequence, such as an enhancer sequence, can exert its function on a target sequence from a more distant position or even from another DNA molecule. A preferred arrangement is one in which the nucleic acid sequence to be transcribed is placed behind the promoter sequence (i.e. at the 3' end) and the two sequences are covalently linked. The distance between the promoter sequence and the nucleic acid sequence to be expressed recombinantly can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.
[0253] In addition to promoters and terminators, other exemplary regulatory elements may include targeting sequences, enhancers, polyadenylation signals, selection markers, amplification signals, replication origins, etc. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
[0254] The term "constitutive promoter" refers to an unregulated promoter that allows for continuous transcription of an operably linked nucleic acid sequence.
[0255] As used herein, the term "operably linked" refers to the linkage of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, a promoter, or rather a transcriptional regulatory sequence, is operably linked to a coding sequence if it affects the transcription of the coding sequence. Operatively linked means that the DNA sequences that are linked are usually contiguous. The nucleotide sequence associated with the promoter sequence may be of homologous or heterologous origin with respect to the plant to be transformed. The sequence may also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with the promoter sequence, after binding to the polypeptide of the embodiments herein, is expressed or repressed according to the promoter characteristics to which it is linked. The associated nucleic acid may code for a protein that is desired to be expressed or repressed at all times in the entire organism, or may code for a protein that is desired to be expressed or repressed at a particular time or in a particular tissue, cell, or cell compartment. Such a nucleotide sequence specifically codes for a protein that confers a desired phenotypic trait to the host cell or host organism modified or transformed therewith. More specifically, the relevant nucleotide sequences result in the production of a product of interest as defined herein in a cell or organism, in particular the nucleotide sequences encode a polypeptide having an enzymatic activity as defined herein.
[0256] The nucleotide sequences described herein above may be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used interchangeably. A (particularly recombinant) expression construct comprises a nucleotide sequence that encodes a polypeptide according to the invention and that is under the genetic control of a regulatory nucleic acid sequence.
[0257] In the methods applied according to the present invention, the expression cassette may be part of an "expression vector", in particular a recombinant expression vector.
[0258] An "expression unit" is understood according to the invention to mean a nucleic acid comprising a promoter as defined herein and having an expression activity which, after being functionally linked to a nucleic acid or gene to be expressed, regulates the expression of said nucleic acid or gene, i.e. the transcription and translation. It is therefore also referred to in this context as a "regulatory nucleic acid sequence". In addition to the promoter, other regulatory elements may also be present, such as enhancers.
[0259] "Expression cassette" or "expression construct" is understood according to the invention to mean an expression unit functionally linked to the nucleic acid to be expressed or to the gene to be expressed. Thus, in contrast to an expression unit, an expression cassette not only comprises nucleic acid sequences which control transcription and translation, but also nucleic acid sequences which are to be expressed as a protein as a result of transcription and translation.
[0260] The term "expression" or "overexpression" in the context of the present invention describes the production or increase of the intracellular activity of one or more polypeptides in a microorganism that are encoded by the corresponding DNA. For this purpose, it is possible, for example, to introduce genes into the organism, to replace existing genes with other genes, to increase the copy number of the genes, to use strong promoters, or to use genes that code for the corresponding polypeptides with increased activity, optionally combining these measures.
[0261] In particular, the constructs according to the invention comprise a promoter sequence 5'-upstream and a terminator sequence 3'-downstream of the respective coding sequence, and optionally other conventional regulatory elements, in each case operably linked to the coding sequence.
[0262] The nucleic acid constructs according to the invention in particular comprise a sequence encoding a polypeptide, e.g. derived from an amino acid related SEQ ID NO: described herein, or a reverse complement thereof, or derivatives and homologues thereof, advantageously operatively or functionally linked to one or more regulatory signals for controlling, e.g. increasing, gene expression.
[0263] In addition to these regulatory sequences, the natural regulation of these sequences may still be present before the actual structural gene, and may optionally be genetically altered so that the natural regulation is switched off and the expression of the gene is enhanced. However, the nucleic acid construct may also be of a simpler structure, i.e., no additional regulatory signals are inserted before the coding sequence, and the natural promoter with its regulation is not removed. Instead, the natural regulatory sequence is mutated so that it is no longer regulated, and the amount of gene expression is increased.
[0264] A preferred nucleic acid construct advantageously also comprises one or more of the previously mentioned "enhancer" sequences operably linked to the promoter, which allow enhanced expression of the nucleic acid sequence. At the 3' end of the DNA sequence, advantageous sequences such as further regulatory elements or terminators may be inserted. One or more copies of the nucleic acid according to the invention may be present in the construct. Other markers, such as genes complementing auxotrophies or antibiotic resistance, may optionally be present in the construct to select the construct.
[0265] Examples of suitable regulatory sequences are promoters, such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, lacI. q , T7, T5, T3, gal, trc, ara, rhaP(rhaP BAD )SP6, lambda-P R , or lambda-P L The regulatory sequences are present in promoters, which are preferably employed in gram-negative bacteria. Further preferred regulatory sequences are present, for example, in the promoters amy and SPO2 for gram-positive bacteria, and in the promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, ADH for yeast or fungi. Regulation can also be achieved using artificial promoters.
[0266] For expression in the host organism, the nucleic acid construct is advantageously inserted into a vector, such as a plasmid or a phage, allowing optimal expression of the gene in the host. Vector is understood to mean, in addition to plasmids and phages, all other vectors known to those skilled in the art, namely viruses such as SV40, CMV, baculovirus and adenovirus, transposons, IS elements, phasmids, cosmids and linear or circular DNA or artificial chromosomes. These vectors can replicate autonomously in the host organism or else on the chromosome. These vectors are a further development of the present invention. Binary vectors or cpo-integrating vectors are also applicable.
[0267] Suitable plasmids include, for example, pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, pIN-III of E. coli. 113 -B1, λgt11 or pBdCI, pIJ101, pIJ364, pIJ702 or pIJ361 of Streptomyces, pUB110, pC194 or pBD214 of Bacillus, pSA77 or pAJ667 of Corynebacterium, pALS1, pIL2 or pBB116 of fungi, 2alphaM, pAG-1, YEp6, YEp13 or pEMBLYe23 of yeast, or pLGV23, pGHlac of plants. + , pBIN19, pAK2004 or pDH51. The above-mentioned plasmids are only a small part of the possible plasmids. Further plasmids are well known to those skilled in the art and are described, for example, in the book "Cloning Vectors" (Eds. Pouwels PH et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).
[0268] As a further development of the vector, the nucleic acid construct according to the invention or the vector comprising the nucleic acid according to the invention can advantageously be introduced into the microorganism in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of a linearized vector, such as a plasmid, or can consist solely of the nucleic acid construct or nucleic acid according to the invention.
[0269] For optimal expression of a heterologous gene in an organism, it is advantageous to modify the nucleic acid sequence to conform to the particular "codon usage" used by the organism. This "codon usage" can be readily determined by computer evaluation of other known genes of the organism in question.
[0270] The expression cassette according to the invention is produced by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. For this purpose, conventional recombinant and cloning techniques are used, as described, for example, in T. Maniatis, EF Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989), and TJ Silhavy, ML Berman and LW Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984), and Ausubel, FM et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987).
[0271] For expression in a suitable host organism, the recombinant nucleic acid construct or gene construct is advantageously inserted into a host-specific vector that allows optimal expression of the gene in the host. Vectors are well known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels PH et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).
[0272] Alternative embodiments of the present disclosure provide methods of "altering gene expression" in a host cell. For example, a polynucleotide of the present disclosure may be enhanced or overexpressed or induced in a host cell or host organism under certain circumstances (e.g., upon exposure to certain temperatures or culture conditions).
[0273] The altered expression of the polynucleotides provided herein may result in ectopic expression, which is an expression pattern that differs between the altered organism and a control or wild-type organism. The altered expression may result from the interaction of the polypeptides of the embodiments herein with exogenous or endogenous modulators, or as a result of chemical modification of the polypeptide. This term also refers to the altered expression pattern of the polynucleotides of the embodiments herein being below detection levels or completely suppressed activity.
[0274] In one embodiment, also provided herein are isolated, recombinant, or synthetic polynucleotides that encode the polypeptides or variant polypeptides provided herein.
[0275] In one embodiment, the nucleic acid sequences encoding several polypeptides are co-expressed in a single host, particularly under the control of different promoters.In another embodiment, the nucleic acid sequences encoding several polypeptides are present on a single transformation vector or are co-transformed simultaneously using separate vectors, and transformants containing both chimeric genes can be selected.Similarly, a gene encoding one polypeptide can be expressed together with another chimeric gene in a single plant, cell, microorganism, or organism.
[0276] 3. Host Applicable to the Present Invention Depending on the context, the term "host" can refer to a wild-type host or a genetically modified recombinant host, or both.
[0277] In principle all prokaryotic or eukaryotic organisms can be considered as hosts or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the invention.
[0278] The vectors according to the invention can be used to generate recombinant hosts, prokaryotic or eukaryotic, which can be transformed, for example, with at least one vector according to the invention and can be used for the production of the polypeptides according to the invention. Advantageously, the above-mentioned recombinant constructs according to the invention are introduced into a suitable host system for expression. The described nucleic acids are expressed in the respective expression systems using in particular the common cloning and transfection methods known to those skilled in the art, such as co-precipitation, protoplast fusion, electroporation, retroviral transfection, etc. Suitable systems are described, for example, in Current Protocols in Molecular Biology, F. Ausubel et al., Ed., Wiley Interscience, New York 1997, or in Sambrook et al. Molecular Cloning: A Laboratory Manual. 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0279] For example, microorganisms such as bacteria are used as host organisms. For example, gram-positive or gram-negative bacteria are used, preferably bacteria of the family Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae or Nocardiaceae, particularly preferably bacteria of the genus Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium or Rhodococcus. The genus and species Escherichia coli are highly preferred. Advantageously, yeasts such as Saccharomyces or Pichia are also suitable hosts.
[0280] Alternatively, eukaryotic cells may be used as hosts. The eukaryotic cells of the present invention may be selected from, but are not limited to, mammalian cells, insect cells, yeast cells and plant cells. The eukaryotic cells of the present invention may be present as individual cells or may be part of a tissue (e.g., cells of a (cultured) tissue, an organ or a whole organism).
[0281] Whole plants or plant cells can also be used as natural or recombinant hosts. Non-limiting examples include the following plants or cells derived therefrom: Nicotiana, in particular Nicotiana benthamiana and Nicotiana tabacum, and Arabidopsis, in particular Arabidopsis thaliana.
[0282] Specific non-limiting examples of insect cells are Sf21, Sf9 and High Five cells.Specific non-limiting examples of mammalian cells are HEK293, HEK293T, HEK293F, CHO, CHO-S, COS, HeLa cells.
[0283] Depending on the host organism, the organisms used in the process according to the invention are grown or cultured in a manner known to those skilled in the art. The culture can be carried out batchwise, semi-batchwise or continuously. Nutrients can be present at the start of the fermentation or can be fed later semi-continuously or continuously, as will also be explained in more detail below.
[0284] 4. POIs Containing One or More UNAAs and Their Preparation 4.1 POI The POI of the present invention generally comprises at least one ncAA, a pyrrolysyl-tRNA synthetase of the present invention, and the above-mentioned tRNA. Pyl It relates to any form of polypeptide or protein molecule that can be recombinantly produced in any suitable host cell system or cell-free expression system as described above in the presence of
[0285] In certain embodiments, the POI is utilized to form a "targeting agent."
[0286] The first purpose of the targeting agent is to form a covalent or non-covalent bond with a specific "target". The second purpose of the targeting agent is the targeted delivery of a "payload molecule" to the target. To achieve the second purpose, the POI must bind (reversibly or irreversibly) to at least one payload molecule. For this purpose, the POI must be functionalized by introducing the at least one ncAA. The functionalized POI with the at least one ncAA can then be linked to the at least one payload molecule by bioconjugation via the ncAA residue. The ncAA is reactive with the payload molecule, and the payload molecule carries a corresponding site reactive with the at least one ncAA residue of the POI. The resulting bioconjugate, i.e. the targeting agent, transfers the payload molecule to the intended target.
[0287] For example, a "target" may be any molecule present in and / or on an organism, tissue, or cell. Such targets may be non-specific or specific to a particular organism, tissue, or cell. Targets include cell surface targets, e.g., receptors, glycoproteins, glycans, carbohydrates; structural proteins, e.g., amyloid plaques; abundant extracellular targets, e.g., in the interstitium, extracellular matrix targets, e.g., growth factors, proteases; intracellular targets, e.g., the surface of the Golgi apparatus, the surface of mitochondria, RNA, DNA, enzymes, components of cell signaling pathways; and / or foreign bodies, e.g., pathogens or parts thereof, e.g., viruses, bacteria, fungi, yeasts, etc.
[0288] Examples of targets include compounds such as proteins whose presence or expression levels correlate with a particular tissue or cell type, or whose expression levels are up- or down-regulated in a particular disease.
[0289] In particular, such targets are proteins such as receptors (internalizing or non-internalizing).
[0290] The target may be selected from any suitable target in the human or animal body, or on a pathogen or parasite.
[0291] Non-limiting examples of suitable targets include, but are not limited to, groups that include: cellular components such as cell membranes and cell walls, receptors such as cell membrane receptors, intracellular structures such as the Golgi apparatus or mitochondria, enzymes, receptors, DNA, RNA, viruses or virus particles, macrophages, tumor associated macrophages, antibodies, proteins, carbohydrates, monosaccharides, polysaccharides, cytokines, hormones, steroids, somatostatin receptors, monoamine oxidase, muscarinic receptors, myocardial sympathetic nervous system, leukotriene receptors, e.g. For example, leukotriene receptors on leukocytes, urokinase plasminogen activator receptor (uPAR), folate receptors, apoptosis markers, (anti)angiogenic markers, gastrin receptors, dopaminergic system, serotonergic system, GABAergic system, adrenergic system, cholinergic system, opioid receptors, GPIIb / IIIa receptors and other thrombosis-related receptors, fibrin, calcitonin receptors, tuftin receptors, P-glycoprotein, neurotensin receptors, neuropeptide receptors, substance P receptors, NK receptors body, CCK receptor, sigma receptor, interleukin receptor, herpes simplex virus tyrosine kinase, human tyrosine kinase, integrin receptor, fibronectin target, AOC3, ALK, AXL, C242, CA-125, CCL11, CCR5, CD2, CD3, CD4, CD5, CD15, CA15-3, CD18, CD19, CA19-9, CD20, CD21, CD22, CD23, CD25, CD28, CD30, CD31, CD33, CD37, CD38, CD40, CD41, CD44v6, CD45 , CD51, CD52, CD54, CD56, CD62E, CD62P, CD62L, CD70, CD72, CD74, CD79-B, CD80, CD105, CD125, CD138, CD141, CD147, CD152, CD154, CD174, CD227, CD326, CD340, VEGF / EGF and VEGF / EGF receptor, VEGF-A, VEGFR2, VEGFR1, TAG72, CEA, MUC1, MUC16, GPNMB, PSMA, Crypto, Tenascin C, Melanocortin-1 receptor, G250, HLADR, ED-B, TMEFF2, EphB2, EphB4, EphA2, FAP, mesothelin, GD2, GD3, CAIX, 5T4, clumping factor, CTLA-4, CXCR2, FGFRl, FGFR2, FGFR3, FGFR4, NaPi2b, NOTCHl, NOTCH2, NOTCH3, NOTCH4, ErbB2, ErbB3, EpCAM, FLT3, HGF, HER2, HER3, HMI24, ICAM, ICOS-L, IGF-1 receptor, TRPV1, CFTR, gdNMB, CA9, c-KIT, c-MET, ACE, APP, adrenergic receptor β2, claudin 3, RON, RORl, PD-Ll, PD-L2, B7-H3, B7-H4, IL-2 receptor, IL-4 receptor, IL-13 receptor, integrin, IFN-α, IFN-γ, IgE, IGF-1 receptor, IL-1, IL-4, IL-5, IL-6, IL-12, IL-13, IL-22, IL-23, interferon receptor, ITGB2 (CD18), LFA-1 (CD1 la), L-selectin, P-selectin, E-selectin, mucin, myostatin, NCA-90, NGF, PDGFRα, prostate cancer cells, Pseudomonas aeruginosa, rabies, RANKL, respiratory syncytial virus, rhesus factor, SLAMF7, sphingosine-1-phosphate, TGF-1, TGFβ2, TGFβ, TNFα, TRAIL-R1, TRAIL-R2, CTAA16.88, vimentin, matrix metalloproteases (MMPs) such as MMP2, MMP9, MMP14, LDL receptor, endoglin, polysialic acid and their corresponding lectins. Examples of targets of fibronectin include the alternatively spliced extra domain A (ED-A) and extra domain B (ED-B) of fibronectin. Non-limiting examples of targets in the stroma are described in V. Hofmeister, D. Schrama, J.C. Becker, Cancer Immun.; Immunother. 2008, 57, 1, the contents of which are incorporated herein by reference.
[0292] More specifically, to enable (specific) targeting of the above targets, the targeting agent can include compounds containing ncAA-functionalized peptide sequences. Such compounds include, but are not limited to, antibodies, antibody derivatives, antibody fragments, antibody (fragment) fusions (e.g., bispecific and trispecific mAb fragments or derivatives), proteins, peptides, such as octreotide and derivatives, VIP, MSH, LHRH, chemotactic peptides, bombesin, elastin, peptidomimetics, receptor agonists and antagonists, cytokines, hormones, steroids, toxins, etc.
[0293] In certain embodiments, the target is a receptor and a targeting agent capable of specifically binding to the target is employed. Suitable targeting agents include, but are not limited to, the ligand of such receptor or a portion thereof that still binds to the receptor, such as a receptor-binding peptide in the case of a receptor-binding protein ligand.
[0294] Other examples of proteinaceous targeting agents include insulin, transferrin, fibrinogen-gamma fragment, thrombospondin, claudins, apolipoprotein E, affibody molecules such as ABY-025, ankyrin repeat proteins, ankyrin-like repeat proteins, interferons, such as alpha, beta, gamma interferons, interleukins, lymphokines, colony stimulating factors, protein growth factors, such as tumor growth factors, such as alpha, beta tumor growth factors, platelet derived growth factor (PDGF), uPAR targeting proteins, apolipoproteins, LDL, annexin V, endostatin, and angiostatin.
[0295] Examples of peptide molecules, such as antibodies, used as targeting agents include LHRH receptor targeting peptides, EC-1 peptides, RGD peptides, HER2 targeting peptides, PSMA targeting peptides, somatostatin targeting peptides, bombesin. Other examples of targeting agents include lipocalins, such as anticalins.
[0296] In certain embodiments, Affibodies™ and multimers and derivatives are used.
[0297] In certain embodiments, antibodies are used to form targeting agents. Antibodies or immunoglobulins derived from IgG antibodies are particularly suitable for use in the present invention, but immunoglobulins from any class or subclass can be selected, including IgG, IgA, IgM, IgD and IgE. Suitably, the immunoglobulins can include, but are not limited to, class IgG, including IgG subclasses (IgG1, 2, 3 and 4), or class IgM immunoglobulins that can specifically bind to specific epitopes on antigens. Antibodies can be intact immunoglobulins of natural or recombinant origin, or can be the immunoreactive part of intact immunoglobulins. Antibodies include, for example, polyclonal antibodies, monoclonal antibodies, camelized single domain antibodies, recombinant antibodies, anti-idiotypic antibodies, multispecific antibodies, antibody fragments, such as Fv, VHH, Fab, F(ab)2, Fab', Fab'-SH, F(ab')2, single chain variable fragment antibodies (scFv), tandem / bis-scFv, Fc, pFc', scFv-Fc, disulfide Fv (dsFv), bispecific antibodies (bc-scFv) such as BiTE antibodies, trispecific antibody derivatives, such as tribodies, camelid antibodies, minibodies, nanobodies, surfaced antibodies, humanized antibodies, fully expressed antibodies, and the like. They can exist in a variety of forms, such as human antibodies, single domain antibodies (sdAbs, also known as Nanobodies™), chimeric antibodies, chimeric antibodies comprising at least one human constant region, dual affinity antibodies such as dual affinity retargeting proteins (DART™), and multimers such as bivalent or multivalent single chain variable fragments and their derivatives (e.g., di-scFv, tri-scFv), (including but not limited to minibodies, diabodies, tribodies, tribodies, tetrabodies, etc.), and multivalent antibodies. See [Trends in Biotechnology 2015, 33, 2, 65], [Trends Biotechnol. 2012, 30, 575-582], and [Cane. Gen. Prot. 2013 10, 1-18], and [BioDrugs 2014, 28, 331-343], the contents of which are incorporated herein by reference.
[0298] The term "antibody fragment" refers to at least a portion of an immunoglobulin variable region that binds to a target, i.e., the antigen-binding region.
[0299] In other embodiments, antibody mimetics are used as targeting agents, including but not limited to affimers, anticalins, avimers, alphabodies, affibodies, DARPins, and multimers and their derivatives. See Trends in Biotechnology 2015, 33, 2, 65, the contents of which are incorporated herein by reference.
[0300] For the avoidance of doubt, in the context of the present invention, the term "antibody" is meant to include all antibody variants, fragments, derivatives, fusions, analogues and mimetics as outlined in this paragraph, unless otherwise specified.
[0301] In preferred embodiments, the targeting agent is selected from agents derived from antibodies and antibody derivatives such as antibody fragments, fragment fusions, proteins, peptides, peptidomimetics, and the like.
[0302] In another preferred embodiment, the targeting agent is selected from agents derived from antibody fragments, fragment fusions, and other antibody derivatives that do not contain an Fc domain.
[0303] Exemplary, non-limiting examples of antibody molecules that are further modified to form the ncAA-modified POI of the present invention are selected from biologically, and in particular pharmacologically, active antibody molecules. Non-limiting examples are selected from the following group: trastuzumab, bevacizumab, cetuximab, panitumumab, ipilimumab, rituximab, alemtuzumab, ofatumumab, gemtuzumab, brentuximab, ibritumomab, tositumomab, and the like. tositumomab, pertuzumab, adecatumumab, IGN101, INA01, labetuzumab, hua33, pemtumomab, oregovomab, minretumomab (CC49), cG250, J591, MOv-18, farletuzumab (MORAb-003), 3F8, ch14,18, KW-2871, hu3S193, lgN31 1, IM-2C6, CDP-791, etaracizumab, volociximab, nimotuzumab, MM-121, AMG 102, METMAB, SCH 900105, AVE1642, IMC-A12, MK-0646, R1507, CP 751871, KB004, III A4, mapatumumab, HGS-ETR2, CS-1008, denosumab, sibrotuzumab, F19, 81 C6, pinatuzumab, rifastuzumab, glembatumumab, coltuximab, lorvotuzumab, indatuximab, anti-PSMA, MLN-0264, ABT-414, milatuzumab,Ramucirumab, abagovomab, abituzumab, adecatumumab, afutuzumab, altumomab pentetate, amatuximab, anatumomab, anetumab, apolizumab, arcitumomab, ascrinvacumab, atezolizumab, bavituximab, bectumomab, belimumab, vivax Bivatuzumab, brontictuzumab, cantuzumab, capromab, catumaxomab, citatuzumab, cixutumumab, clivatuzumab, codrituzumab, conatumumab, dacetuzumab tuzumab, dallotuzumab, daratumumab, demcizumab, denintuzumab, depatuxizumab, derlotuximab, detumomab, dinutuximab, drozitumab, durigotumab , durvalumab, dusigitumab, ecromeximab, edrecolomab, elgemtumab, emactuzumab, enavatuzumab, emibetuzumab, enfortumab, enoblituzumab,Ensituximab, epratuzumab, ertumaxomab, etaracizumab, farletuzumab, ficlatuzumab, figitumumab, flanvotumab, futuximab, galiximab, ganitumab, icrucumab, igobomab igovomab, imalumab, imgatuzumab, indusatumab, inebilizumab, intetumumab, iratumumab, isatuximab, lexatuzumab, lilotomab, lintuzumab, lirilumab, lucatumumab, lumurex Lumretuzumab, margetuximab, matuzumab, mirvetuximab, mitumomab, mogamulizumab, moxetumomab, nacolomab, naptumomab, narnatumab, necitumumab, nesvacumab, nimotuzumab uzumab, nivolumab, nofetumomab, obinutuzumab, ocaratuzumab, ofatumumab, olaratumab, onartuzumab, ontuxizumab, oportuzumab, oregovomab, otlertuzumab, pancomab,Parsatuzumab, pasotuxizumab, patritumab, pembrolizumab, pemtumomab, pidilizumab, pintumomab, polatuzumab, pritumumab, quilizumab, racotumomab, ramucirumab, rilo Rilotumumab, robatumumab, sacituzumab, samalizumab, satumomab, seribantumab, siltuximab, sofituzumab, tacatuzumab, taplitumomab, tarextumab, tenatumomab, teprotumumab teprotumumab, tetulomab, ticilimumab, tigatuzumab, tositumomab, tovetumab, tremelimumab, tucotuzumab, ublituximab, ulocuplumab, urelumab, utomilumab, vadastuximab imab, vandortuzumab, vantictumab, vanucizumab, varlilumab, veltuzumab, besencumab, volociximab, vorsetuzumab, votumumab, zalutumumab, zatuxima, combinations and derivatives thereof, and CAI 25, CAI 5-3,Other monoclonal antibodies targeting CAI 9-9, L6, Lewis Y, Lewis X, alpha-fetoprotein, CA 242, placental alkaline phosphatase, prostate-specific antigen, prostate-specific membrane antigen, prostatic acid phosphatase, epidermal growth factor, MAGE-1, MAGE-2, MAGE-3, MAGE-4, transferrin receptor, p97, MUCI, CEA, gplOO, MARTI, IL-2 receptor, CD20, CD52, CD33, CD22, human chorionic gonadotropin, CD38, CD40, mucin, P21, MPG, and Neu oncogene product.
[0304] According to further particular embodiments of the invention, targets and targeting agents are selected to provide specific or enhanced targeting of tissues or diseases such as cancer, inflammation, infection, cardiovascular disease, e.g., thrombus, atherosclerotic lesions, hypoxic sites, e.g., stroke, tumors, cardiovascular disorders, brain disorders, apoptosis, angiogenesis, organs, reporter genes / enzymes, etc. This can be achieved by selecting targets with tissue-, cell-, or disease-specific expression.
[0305] As an example, a targeting agent specifically binds to or complexes with a cell surface molecule, such as a cell surface receptor or antigen, for a certain cell population. After the targeting agent specifically binds to or complexes with the receptor, the agent enters the cell.
[0306] As used herein, a targeting agent that "specifically binds or complexes" or "targets" a cell surface molecule, an extracellular matrix target, or another target, preferentially associates with the target through intermolecular forces. For example, a ligand can preferentially associate with a target having a dissociation constant (Kd or KD) of less than about 50 nM, less than about 5 nM, or less than about 500 pM.
[0307] 4.2 Creation of POI A POI comprising one or more UNAA residues can be produced according to the invention using a eukaryotic cell that contains (e.g., is provided with) at least one unnatural amino acid or salt thereof that corresponds to the UNAA residue of the POI to be produced. The eukaryotic cell further comprises: (i) PylRS and tRNA of the present invention Pyl (PylRS is preferably selective for tRNA Pyl can be acylated with UNAA or a salt thereof); and (ii) a polynucleotide encoding a POA (wherein any position in the POI is occupied by a UNAA residue is expressed by a tRNA Pyl a codon that is the reverse complement of the anticodon of, e.g., encoded by a selector codon). The eukaryotic cells are cultured to allow translation of the polynucleotide (ii) encoding the POI, thereby producing the POI.
[0308] To produce a POI according to the methods of the invention, the translation in step (b) can be accomplished by culturing eukaryotic cells under suitable conditions, preferably in the presence of UNAA or a salt thereof (e.g., in a medium containing UNAA or a salt thereof), for a suitable time to allow translation in the ribosomes of the cells. Pyl Depending on the type of gene that encodes the POI, it may be necessary to induce expression by adding a compound that induces transcription, such as arabinose, isopropyl β-D-thiogalactoside (IPTG) or tetracycline. Pyl The mRNA (containing one or more codons that are the reverse complement of the anticodon contained in the mRNA) binds to the ribosome. A polypeptide is then formed by the stepwise binding of amino acids and UNAAs to the positions coded for by the codons recognized by each aminoacyl-tRNA. As a result, the UNAAs are bound to the tRNA Pyl is incorporated into the POI at a position encoded by a codon that is the reverse complement of the anticodon contained in
[0309] The eukaryotic cell may contain a polynucleotide sequence encoding the PylRS of the present invention that allows the cell to express the PylRS. Pyl is a tRNA contained in cells. Pyl The polynucleotide sequence encoding PylRS and tRNA can be produced by eukaryotic cells. Pyl The polynucleotide sequences encoding the can be located either on the same polynucleotide or on separate polynucleotides.
[0310] Thus, in one embodiment, the invention provides a method for making a POI that contains one or more UNAA residues, the method comprising the steps of: (a) providing a eukaryotic cell comprising a polynucleotide sequence encoding: - at least one PylRS of the invention; - at least one tRNA that can be acylated by PylRS (tRNA Pyl ); and At least one POI (any position of the POI occupied by a UNAA residue) is Pyl is encoded by a codon that is the reverse complement of the anticodon of (b) enabling translation of the polynucleotide sequence by a eukaryotic cell in the presence of UNAA or a salt thereof, thereby activating PylRS, tRNA Pyl and producing the POI.
[0311] Eukaryotic cells used to generate POIs that include one or more unnatural amino acid residues as described herein include PylRS, tRNA Pyl and a polynucleotide sequence encoding the POI into a eukaryotic (host) cell, which may be located on the same polynucleotide or on separate polynucleotides and may be introduced into the cell by methods known in the art (e.g., using viral-mediated gene delivery, electroporation, microinjection, lipofection, etc.).
[0312] After translation, the POI produced according to the invention can optionally be recovered and purified to partial or substantial homogeneity according to procedures commonly known in the art. Unless the POI is secreted into the medium, recovery usually requires cell disruption. Methods of cell disruption are well known in the art and can include physical disruption by ultrasound treatment, liquid shear disruption (e.g., using a French press), mechanical methods (e.g., using a blender or grinder), or freeze-thaw cycles, as well as chemical lysis using agents that disrupt lipid-lipid, protein-protein and / or protein-lipid interactions (e.g., detergents), and combinations of physical disruption techniques and chemical lysis. Standard procedures for purifying polypeptides from cell lysates or medium are also well known in the art. Examples include ammonium sulfate or ethanol precipitation, acid or base extraction, column chromatography, affinity column chromatography, anion or cation exchange chromatography, phosphocellulose chromatography, hydrophobic interaction chromatography, hydroxyapatite chromatography, lectin chromatography, gel electrophoresis, and the like. Protein refolding steps can be used to generate correctly folded mature proteins, if necessary. High performance liquid chromatography (HPLC), affinity chromatography or other suitable methods can be used in the final purification step where high purity is desired. Antibodies generated against the polypeptides of the invention can be used as purification reagents, i.e., affinity-based purification of the polypeptides. Various purification / protein refolding methods are well known in the art, including those described in Scopes, Protein Purification, Springer, Berlin (1993); Deutscher, Methods in Enzymology Vol. 182: Guide to Protein Purification, Academic Press (1990); and references cited therein.
[0313] As already mentioned, the skilled artisan will recognize that after synthesis, expression and / or purification, a polypeptide may have a conformation different from the desired conformation of the relevant polypeptide. For example, polypeptides produced by prokaryotic systems are often optimized by exposure to chaotropic agents to achieve proper folding. During purification, for example, from lysates from E. coli, the expressed polypeptide is denatured and then renatured as necessary. This is accomplished by solubilizing the protein in a chaotropic agent such as guanidine HCl. In general, it is sometimes desirable to denature and destroy the expressed polypeptide and then refold the polypeptide into a preferred conformation. For example, guanidine, urea, DTT, DTE and / or chaperonins can be added to the translation product of interest. Methods for reducing, denaturing and renaturing proteins are well known to those skilled in the art. The polypeptide can be refolded, for example, in a redox buffer containing oxidized glutathione and L-arginine.
[0314] 5. Payload molecules Commonly used payload molecules can be selected from biologically active compounds, labeling agents, and chelating agents, non-limiting examples of which are provided below.
[0315] 5.1 Biologically active compounds Biologically active compounds include, but are not limited to, the following: Bioactive compounds applicable according to the present invention include, but are not limited to, small organic molecule drugs, steroids, lipids, proteins, aptamers, oligopeptides, oligonucleotides, oligosaccharides, as well as peptides, peptoids, amino acids, nucleotides, oligo- or polynucleotides, nucleosides, DNA, RNA, toxins, sugar chains and immunoglobulins.
[0316] Exemplary classes of biologically active compounds that can be used in the practice of the present invention include, but are not limited to, the following: hormones, cytotoxins, anti-proliferative / anti-tumor agents, anti-virals, antibiotics, cytokines, anti-inflammatory agents, antihypertensive agents, chemosensitizers, photosensitizers and radiosensitizers, anti-AIDS substances, anti-virals, immunosuppressants, immunostimulants, enzyme inhibitors, anti-Parkinson's agents, neurotoxins, channel blockers, cell growth inhibitors and extracellular matrix interaction modulators including anti-adhesion molecules, DNA, RNA or protein synthesis inhibitors, steroidal and non-steroidal anti-inflammatory agents, anti-angiogenic factors, anti-Alzheimer's agents.
[0317] In some embodiments, the biologically active compound is a low to medium molecular weight compound (eg, about 200 to 5000 Da, about 200 to about 1500 Da, preferably about 300 to about 1000 Da).
[0318] Exemplary cytotoxic drugs are those that are used in particular for cancer treatment.Such drugs generally include DNA damaging agents, antimetabolites, natural products and their analogs, enzyme inhibitors such as dihydrofolate reductase inhibitors and thymidylate synthase inhibitors, DNA binders, DNA alkylating agents, radiosensitizers, DNA intercalators, DNA cleavage agents, microtubule stabilizing and destabilizing agents, topoisomerase inhibitors.For example, they include, but are not limited to, platinum drugs, anthracyclines, vincas, mitomycins, bleomycins, cytotoxic nucleosides, taxanes, lexitropsins, pteridines, diynenes, podophyllotoxins, dolastatins, maytansinoids, differentiation inducers, taxols, etc. Particularly useful members of these classes include, for example, auristatins, maytansines, maytansinoids, calicheamicins, dactinomycins, duocarmycins, CC1065 and its analogs, camptothecins and its analogs, SN-38 and its analogs; DXd, tubulysin M, cryptophycins, pyrrolobenzodiazepines and pyrrolobenzodiazepine dimers (PBDs), pyridinobenzodiazepines (PDDs) and indolinobenzodiazepines (IBDs) (see US20210206763A1), methotrexate, methopterin, dichloromethotrexate, 5-fluorouracil, DNA minor groove binders, 6-mercaptopurine, cytosine arabinoside, melphalan, leucosides ... cin, leurosidein, actinomycin, anthracyclines (doxorubicin, epirubicin, idarubicin, daunorubicin, PNU-159682 (see US10288745B2) and analogs thereof, mitomycin C, mitomycin A, caminomycin, aminopterin, tallysomycin, podophyllotoxin and; podophyllotoxin derivatives such as etoposide or etoposide phosphate, vinblastine, vincristine, vindesine, taxol, taxotere, retinoic acid, butyric acid, N8-acetylspermidine, staurosporine, colchicine, camptothecin, esperamicin, enediyne and analogs thereof, hemiasterlin and analogs thereof.
[0319] Other exemplary drug classes are angiogenesis inhibitors, cell cycle progression inhibitors, P13K / m-TOR / ACT pathway inhibitors, MAPK signaling pathway inhibitors, kinase inhibitors, protein chaperone inhibitors, HDAC inhibitors, PARP inhibitors, Wnt / Hedgehog signaling pathway inhibitors, RNA polymerase inhibitors, and protein degraders (see: https: / / pubs.acs.org / doi / 10.1021 / acschembio.0c00285).
[0320] Examples of auristatins include dolastatin 10, monomethylauristatin E (MMAE), auristatin F, monomethylauristatin F (MMAF), auristatin F hydroxypropylamide (AF HPA), auristatin F phenylenediamine (AFP), monomethylauristatin D (MMAD), auristatin PE, auristatin EB, auristatin EFP, auristatin TP, and auristatin AQ. Suitable auristatins also are described in U.S. Patent Publication Nos. 2003 / 0083263, 2011 / 0020343, and 2011 / 0070248; PCT Publication Nos. WO09 / 117531, WO2005 / 081711, WO04 / 010957; WO02 / 088172, and WO01 / 24763, as well as U.S. Patent Nos. 7,498,298; 6,884,869; 6,323,315; 6,239,104; 6,124,431; 6,034,065; 5,780,588; 5,767,237; Nos. 5,665,860; 5,663,149; 5,635,483; 5,599,902; 5,554,725; 5,530,097; 5,521,284; 5,504,191; 5,410,024; 5,138,036; 5,076,973; 4,986,988; 4,978,744; 4,879,278; 4,879,278; 4,816,444; and 4,486,414, the disclosures of which are incorporated herein by reference in their entireties.
[0321] Exemplary drugs include dolastatins and their analogs, including: dolastatin A (U.S. Pat. No. 4,486,414), dolastatin B (U.S. Pat. No. 4,486,414), dolastatin 10 (U.S. Pat. Nos. 4,486,444, 5,410,024, 5,504,191, 5,521,284, 5,530,097, 5,599,902, 5,635,483, 5,663,149, 5,671,149, 5,701,149, 5,711,149, 5,721,149, 5,730,160, 5,741,161, 5,751,162, 5,761,163, 5,771,164, 5,781,165, 5,791,176, 5,821,187, 5,831,189, 5,841,189, 5,851,189, 5,861,189, 5,871,189, 5,881,190, 5,891,191, 5,901,192, 5,931,193, 5,941,194, 5,951,195, 5,961,196, 5,971,197, 5,981,198, 5,991,199, 6,013,199, 6,021,199, 6,031,199, 6,041,19 Nos. 6,034,065, 6,323,315), dolastatin 13 (U.S. Pat. No. 4,986,988), dolastatin 14 (U.S. Pat. No. 5,138,036), dolastatin 15 (U.S. Pat. No. 4,879,278), dolastatin 16 (U.S. Pat. No. 6,239,104), dolastatin 17 (U.S. Pat. No. 6,239,104), and dolastatin 18 (U.S. Pat. No. 6,239,104). Each of these patents is incorporated herein by reference in its entirety.
[0322] Exemplary maytansine, maytansinoids such as DM-1 and DM-4, or maytansinoid analogs, including maytansinol and maytansinol analogs, are disclosed in U.S. Patent Nos. 4,424,219; 4,256,746; 4,294,757; 4,307,016; 4,313,946; 4,315,929; 4,331,598; 4,361,650; Nos. 6,663; 4,364,866; 4,450,254; 4,322,348; 4,371,533; 5,208,020; 5,416,064; 5,475,092; 5,585,499; 5,846,545; 6,333,410; 6,441,163; 6,716,821 and 7,276,497.
[0323] Other examples include mertansine and ansamitocin. Pyrrolobenzodiazepines (PBDs) explicitly include dimers and analogs, including, but not limited to, those described in [Denny, Exp. Opin. Ther. Patents, 10(4): 459-474 (2000)], [Hartley et al., Expert Opin Investig Drugs. 2011, 20(6): 733-44], [Antonow et al., Chem Rev. 2011, 111(4), 2815-64].
[0324] Calicheamicins include, for example, enediynes, esperamicins, and those described in US Pat. Nos. 5,714,586 and 5,739,116.
[0325] Examples of duocarmycins and analogs include CC1065, duocarmycin SA, duocarmycin A, duocarmycin Bl, duocarmycin B2, duocarmycin CI, duocarmycin C2, duocarmycin D, DU-86, KW-2189, adozelesin, bizelesin, carzelesin, and secoadzelesin. Other examples are described in, e.g., U.S. Patent Nos. 5,070,092; 5,101,092; 5,187,186; 5,475,092; 5,595,499; 5,846,545; 6,534,660; 6,548,530; 6,586,618; 6,597,712; 6,601,613; 6,601,614; 6,601,615; 6,601,616; 6,601,617; 6,601,618; 6,601,619 ... Nos. 6,660,742; 6,756,397; 7,049,316; 7,553,816; 8,815,226; US20150104407; 61 / 988,011 filed May 2, 2014, and 62 / 010,972 filed June 11, 2014, the disclosures of each of which are incorporated herein in their entirety.
[0326] Exemplary vinca alkaloids include vincristine, vinblastine, vindesine, and navelbine, and those disclosed in U.S. Patent Publication Nos. 2002 / 0103136 and 2010 / 0305149, and U.S. Patent No. 7,303,749, the disclosures of which are incorporated herein by reference in their entireties.
[0327] Exemplary epothilone compounds include epothilones A, B, C, D, E and F, and derivatives thereof. Suitable epothilone compounds and derivatives thereof are described, for example, in U.S. Pat. Nos. 6,956,036; 6,989,450; 6,121,029; 6,117,659; 6,096,757; 6,043,372; 5,969,145; and 5,886,026; as well as WO97 / 19086; WO98 / 08849; WO98 / 22461; WO98 / 25929; WO98 / 38192; WO99 / 01124; WO99 / 02514; WO99 / 03848; WO99 / 07692; WO99 / 27890; and WO99 / 28324. These disclosures are incorporated herein by reference in their entireties.
[0328] Exemplary cryptophycin compounds are described in U.S. Patent Nos. 6,680,311 and 6,747,021, the disclosures of which are incorporated herein by reference in their entireties.
[0329] Exemplary platinum compounds include cisplatin, carboplatin, oxaliplatin, iproplatin, ormaplatin, and tetraplatin.
[0330] Exemplary DNA binding or alkylating agents include CC-1065 and its analogs, anthracyclines, calicheamicins, dactinomycins, mithromycins, pyrrolobenzodiazepines, and the like.
[0331] Exemplary microtubule stabilizing and destabilizing agents include taxane compounds such as paclitaxel, docetaxel, tesetaxel, and carbazitaxel; maytansinoids, auristatins and their analogs, vinca alkaloid derivatives, epothilones, and cryptophycins.
[0332] Exemplary topoisomerase inhibitors include camptothecin and camptothecin derivatives, camptothecin analogs and non-natural camptothecins, such as, for example, CPT-11, SN-38, topotecan, 9-aminocamptothecin, rubitecan, gimatecan, karenitecin, ciratecan, lurtotecan, exatecan, DXd, diflometotecan, belotecan, lurtotecan and S39625.Other camptothecin compounds that can be used in the present invention include, for example, those described in J.Med.Chem.,29:2358-2363(1986);J.Med.Chem.,23:554(1980);J.Med.Chem.,30:1774(1987).
[0333] Angiogenesis inhibitors include, but are not limited to, MetAP2 inhibitors, VEGF inhibitors, PIGF inhibitors, VGFR inhibitors, PDGFR inhibitors, MetAP2 inhibitors.Exemplary VGFR inhibitors and PDGFR inhibitors include sorafenib, sunitinib and vatalanib.Exemplary MetAP2 inhibitors include fumagillol analogs, i.e. compounds that contain fumagillin core structure.
[0334] Exemplary cell cycle progression inhibitors include CDK inhibitors, such as, for example, BMS-387032 and PD0332991; Rho kinase inhibitors, such as, for example, AZD7762; Aurora kinase inhibitors, such as, for example, AZD1152, MLN8054 and MLN8237; PLK inhibitors, such as, for example, BI2536, BI6727, GSK461364, ON-01910; and KSP inhibitors, such as, for example, SB743921, SB715992, MK-0731, AZD8477, AZ3146, ARRY-520.
[0335] Exemplary P13K / m-TOR / ACT signaling pathway inhibitors include phosphoinositide 3-kinase (P13K) inhibitors, GSK-3 inhibitors, ATM inhibitors, DNA-PK inhibitors and PDK-1 inhibitors.
[0336] Exemplary P13 kinases are disclosed in U.S. Pat. No. 6,608,053 and include BEZ235, BGT226, BKM120, CAL263, demethoxyviridin, GDC-0941, GSK615, IC87114, LY294002, palomid 529, perifosine, PF-04691502, PX-866, SAR245408, SAR245409, SF1126, wortmannin, XL147 and XL765.
[0337] Exemplary AKT inhibitors include, but are not limited to, AT7867.
[0338] Exemplary MAPK signaling pathway inhibitors include MEK, Ras, JNK, B-Raf and p38 MAPK inhibitors.
[0339] Exemplary MEK inhibitors are disclosed in U.S. Pat. No. 7,517,944 and include GDC-0973, GSKl 120212, MSC1936369B, AS703026, R05126766 and R04987655, PD0325901, AZD6244, AZD8330 and GDC-0973.
[0340] Exemplary B-raf inhibitors include CDC-0879, PLX-4032, and SB590885.
[0341] Exemplary B p38 MAPK inhibitors include BIRB796, LY2228820 and SB202190. Exemplary receptor tyrosine kinase inhibitors include, but are not limited to, AEE788 (NVP-AEE 788), BIBW2992 (afatinib), lapatinib, erlotinib (Tarceva), gefitinib (Iressa), AP24534 (ponatinib), ABT-869 (linifanib), AZD2171, CHR-258 (dovitinib), sunitinib (sutent), sorafenib (nexavar), and vatalinib.
[0342] Exemplary protein chaperone inhibitors include HSP90 inhibitors. Exemplary inhibitors include 17AAG derivatives, BIIB021, BIIB028, SNX-5422, NVP-AUY-922 and KW-2478.
[0343] Exemplary HDAC inhibitors include belinostat (PXD101), CUDC-101, droxinostat, ITF2357 (divinostat, gabinostat), JNJ-26481585, LAQ824 (NVP-LAQ824, dacinostat), LBH-589 (panobinostat), MC1568, MGCD0103 (mosetinostat), MS-275 (entinostat), PCI-24781, pyroxamide (NSC 696085), SB939, trichostatin A, and vorinostat (SAHA). Exemplary PARP inhibitors include iniparib (BSI 201), olaparib (AZD-2281), ABT-888 (veliparib), AG014699, CEP9722, MK 4827, KU-0059436 (AZD2281), LT-673, 3-aminobenzamide, A-966492, and AZD2461.
[0344] Exemplary Wnt / hedgehog signaling pathway inhibitors include vismodegib, cyclopamine, and XAV-939.
[0345] Exemplary RNA polymerase inhibitors include amatoxins.Exemplary amatoxins include α-amanitin, β-amanitin, γ-amanitin, eta-amanitin, amanulin, amanurinic acid, amanisamide, amanone, and proamanitin.
[0346] Exemplary cytokines include IL-2, IL-7, IL-10, IL-12, IL-15, IL-21, TNF. Non-limiting examples of specific agents can include auristatins, maytansinoids, PBDs, topoisomerase inhibitors, and anthracyclines.
[0347] In another embodiment, a combination of two or more different agents as described above is used.
[0348] According to another embodiment, the biologically active compound may be selected from any synthetic or natural compound comprising one or more natural and / or non-natural, proteinogenic and / or non-proteinogenic amino acid residues, such as in particular oligo- or polypeptides or proteins.
[0349] This particular group of compounds includes immunoglobulin molecules such as, for example, antibodies, antibody derivatives, antibody fragments, antibody (fragment) fusions (e.g. bispecific and trispecific mAb fragments or derivatives), polyclonal or monoclonal antibodies, e.g. human antibodies, humanized antibodies, murine antibodies or chimeric antibodies.
[0350] Typical, non-limiting examples of antibodies for use in the present invention are selected from biologically, and in particular pharmacologically, active antibody molecules. Non-limiting examples are selected from the following group: trastuzumab, bevacizumab, cetuximab, panitumumab, ipilimumab, rituximab, alemtuzumab, ofatumumab, gemtuzumab, brentuximab, ibritumomab, tositumomab, serotoninib ... tositumomab, pertuzumab, adecatumumab, IGN101, INA01, labetuzumab, hua33, pemtumomab, oregovomab, minretumomab (CC49), cG250, J591, MOv-18, farletuzumab (MORAb-003), 3F8, ch14,18, KW-2871, hu3S193, lgN31 1, IM-2C6, CDP-791, etaracizumab, volociximab, nimotuzumab, MM-121, AMG 102, METMAB, SCH 900105, AVE1642, IMC-A12, MK-0646, R1507, CP 751871, KB004, III A4, mapatumumab, HGS-ETR2, CS-1008, denosumab, sibrotuzumab, F19, 81 C6, pinatuzumab, rifastuzumab, glembatumumab, coltuximab, lorvotuzumab, indatuximab, anti-PSMA, MLN-0264, ABT-414, milatuzumab, ramucirumab,abagovomab, abituzumab, adecatumumab, afutuzumab, altumomab pentetate, amatuximab, anatumomab, anetumab, apolizumab, arcitumomab, ascrinvacumab, atezolizumab, bavituximab, bectumomab, belimumab, bivatuzumab atuzumab, brontictuzumab, cantuzumab, capromab, catumaxomab, citatuzumab, cixutumumab, clivatuzumab, codrituzumab, conatumumab, dacetuzumab, daro Dallotuzumab, daratumumab, demcizumab, denintuzumab, depatuxizumab, derlotuximab, detumomab, dinutuximab, drozitumab, durigotumab, durvalumab mab), dusigitumab, ecromeximab, edrecolomab, elgemtumab, emactuzumab, enavatuzumab, emibetuzumab, enfortumab, enoblitzumab, ensituximab,Epratuzumab, ertumaxomab, etaracizumab, farletuzumab, ficlatuzumab, figitumumab, flanvotumab, futuximab, galiximab, ganitumab, icrucumab, igovomab, imalumab umab, imgatuzumab, indusatumab, inebilizumab, intetumumab, iratumumab, isatuximab, lexatuzumab, lilotomab, lintuzumab, lirilumab, lucatumumab, lumretuzumab, marge Margetuximab, matuzumab, mirvetuximab, mitumomab, mogamulizumab, moxetumomab, nacolomab, naptumomab, narnatumab, necitumumab, nesvacumab, nimotuzumab, nivolumab b), nofetumomab, obinutuzumab, ocaratuzumab, ofatumumab, olaratumab, onartuzumab, ontuxizumab, oportuzumab, oregovomab, otlertuzumab, pancomab, parsatuzumab,pasotuxizumab, patritumab, pembrolizumab, pemtumomab, pidilizumab, pintumomab, polatuzumab, pritumumab, quilizumab, racotumomab, ramucirumab, rilotumumab , robatumumab, sacituzumab, samalizumab, satumomab, seribantumab, siltuximab, sofituzumab, tacatuzumab, taplitumomab, tarextumab, tenatumomab, teprotumumab mab), tetulomab, ticilimumab, tigatuzumab, tositumomab, tovetumab, tremelimumab, tucotuzumab, ublituximab, ulocuplumab, urelumab, utomilumab, vadastuximab , vandortuzumab, vantictumab, vanucizumab, varlilumab, veltuzumab, besencumab, volociximab, vorsetuzumab, votumumab, zalutumumab, zatuxima, combinations and derivatives thereof, and CAI 25, CAI 5-3, CAI 9-9, L6, Lewis Y,Other monoclonal antibodies targeting Lewis X, alpha-fetoprotein, CA 242, placental alkaline phosphatase, prostate-specific antigen, prostate-specific membrane antigen, prostatic acid phosphatase, epidermal growth factor, MAGE-1, MAGE-2, MAGE-3, MAGE-4, transferrin receptor, p97, MUCI, CEA, gplOO, MARTI, IL-2 receptor, CD20, CD52, CD33, CD22, human chorionic gonadotropin, CD38, CD40, mucin, P21, MPG, and the Neu oncogene product.
[0351] 5.2 Labeling agents Labeling agents that can be used in accordance with the present invention can include any type of label known in the art that does not inhibit the reactivity of the tetrazine moiety.
[0352] Labels of the present invention can include, but are not limited to, dyes (e.g., fluorescent, luminescent, or phosphorescent dyes, such as dansyl, coumarin, fluorescein, acridine, rhodamine, silicon rhodamine, BODIPY, or cyanine dyes), chromophores (e.g., phytochromes, phycobilins, bilirubin, etc.), radiolabels (e.g., radioactive forms of hydrogen, fluorine, carbon, phosphorus, sulfur, iodine, e.g., tritium, fluorine-18, carbon-11, carbon-14, phosphorus-32, phosphorus-33, sulfur-33, etc.), and the like. , sulfur-35, iodine-123, or iodine-125), MRI-sensitive spin labels, affinity tags (e.g., biotin, His-tags, Flag-tags, strep-tags, sugars, lipids, sterols, PEG-linkers, benzylguanine, benzylcytosine, or cofactors), polyethylene glycol groups (e.g., branched PEGs, linear PEGs, PEGs of different molecular weights, etc.), photocrosslinkers (such as p-azidoiodoacetanilide), NMR probes, X-ray probes, pH probes, IR probes, resins, solid supports.
[0353] In some embodiments, exemplary dyes can include NIR contrast agents that fluoresce in the near-infrared region of the spectrum. Exemplary near-infrared fluorophores can include dyes and other fluorophores having emission wavelengths (e.g., peak emission wavelengths) between about 630-1000 nm, e.g., between about 630-800 nm, between about 800-900 nm, between about 900-1000 nm, between about 680-750 nm, between about 750-800 nm, between about 800-850 nm, between about 850-900 nm, between about 900-950 nm, or between about 950-1000 nm. Fluorophores having emission wavelengths (e.g., peak emission wavelengths) greater than 1000 nm can also be used in the methods described herein.
[0354] In some embodiments, exemplary fluorophores include 7-amino-4-methylcoumarin-3-acetic acid (AMCA), TEXAS RED™ (Molecular Probes, Inc.; Eugene, OR), 5(6)-carboxy-X-rhodamine, Lissamine rhodamine B, 5(6)-carboxyfluorescein, fluorescein-5-isothiocyanate (FITC), 7-diethylaminocoumarin-3-carboxylic acid, tetramethylrhodamine-5(6)-isothiocyanate, 5(6)-carboxytetramethylrhodamine, 7-hydroxycoumarin-3-carboxylic acid, 6-[fluorescein-5(6)-carboxamido]hexanoic acid, N-(4,4-difluoro-5,7-dimethyl-4-bora-3a,4a-diaza-3-indacenepropionic acid, eosin-5-isothiocyanate, erythrosine-5-isothiocyanate, and CASCADE™ blue acetyl azide (Molecular Probes, Inc., Eugene, OR) and ATTO dye.
[0355] Further labelling agents include 177-lutetium, 89-zirconium, 131-iodine, 68-gallium, 99m technetium, 225-actinium, 213-bismuth, 90-yttrium and 212-lead.
[0356] 5.3 Chelating agents Below is a list of commonly applicable chelating agents and their abbreviations, including their corresponding salts.
[0357] Acetylacetone (ACAC), ethylenediamine (EN), 2-(2-aminoethylamino)ethanol (AEEA), diethylenetriamine (DIEN), iminodiacetate (IDA), triethylenetetramine (TRIEN), triaminotriethylamine, nitrilotriacetic acid (NTA) and its salts such as Na3NTA or FeNTA, ethylenediaminotriacetate (TED), ethylenediaminetetraacetic acid (EDTA) and its salts such as Na2EDTA and CaNa2EDTA, diethylenetriaminepentaacetic acid (DTPA), 1,4,7,10-tetraazacyclododecane-1,4,7,10-tetraacetic acid (DOTA), oxalate (OX), tartrate (TART), citrate ( CIT), dimethylglyoxime (DMG), 8-hydroxyquinoline, 2,2'-bipyridine (BPY), 1,10-phenanthroline (PHEN), dimercaptosuccinic acid (DMSA), 1,2-bis(diphenylphosphino)ethane (DPPE), sodium salicylate, methoxysalicylate, British anti-Lewisite or 2,3-dimercaprol (BAL), meso-2,3-dimercaptosuccinic acid (DMSA); siderophores secreted by microorganisms, e.g., desferrioxamine or deferoxamine B, also known as Deferral (Novartis), produced by Streptomyces spp.; deferoxamine (DFO), produced by Streptomyces pilosus (Streptomyces phytochemicals such as curcuminoids and mugineic acid derivatives such as 3-hydroxymugineic acid and 2'-deoxymugineic acid; synthetic chelators such as ibuprofen; derivatives of catechol, hydroxamate, and hydroxypyridinone such as the hydroxamate desferal and the hydroxypyridinone deferiprone; deferiprone (L1 or 1,2-dimethyl-3-hydroxypyrid-4-one); β-β-dimethylcysteine or 3-mercapto-D-valine, D-penicillamine (DPA or D-PEN); tetraethylenetetraamine (TETA) or trientine and its two major metabolites N1-acetyltriethylenetetramine (MAT) and N1,N 10-diacetyltriethylenetetramine (DAT); hydroxyquinoline; clioquinol, a halogenated derivative of 8-hydroxyquinoline; and 5,7-dichloro-2-[(dimethylamino)methyl]quinolin-8-ol (PBT2).
[0358] 6 UNAA UNAA useful in the methods and kits of the present invention have been described in the prior art (see, e.g., Liu et al., Annu Rev Biochem 83:379-408, 2010; Lemke, ChemBioChem 15:1691-1694, 2014).
[0359] The UNAA can have a group (herein referred to as a "label group") that facilitates reaction with an appropriate group (herein referred to as a "docking group") on another molecule (herein referred to as a "binding partner molecule"), whereby the binding partner molecule is covalently attached to the UNAA. When a UNAA bearing a label group is translationally incorporated into a POI, the label group becomes part of the POI. Thus, a POI made according to the methods of the invention can be reacted with one or more binding partner molecules, whereby the binding partner molecule is covalently attached to (the label group of) a non-natural amino acid residue of the POI. This conjugation reaction can be used for in situ coupling of the POI within cells or tissues expressing the POI, or for site-specific conjugation of an isolated or partially isolated POI.
[0360] Particularly useful options for combinations of labeling groups and docking groups (of the binding partner molecule) are those that can react by metal-free click reactions, including strain-promoted inverse electron demand Diels-Alder cycloaddition (SPIEDAC; see, e.g., Devaraj et al., Angew Chem Int Ed Engl 2009, 48:7013) and cycloaddition reactions of strained cycloalkynyl groups or strained cycloalkynyl analogues having one or more triple bond-free ring atoms substituted with amino groups with azides, nitrile oxides, nitrones and diazocarbonyl reagents (see, e.g., Sanders et al., J Am Chem Soc 2010, 133:949; Agard et al., J Am Chem Soc 2004, 126:15046), such as strain-promoted alkyne-azide cycloaddition (SPAAC). Such click reactions allow ultrafast, bi-orthogonal, covalent, site-specific coupling of the UNAA labeling group of the POI with an appropriate group on a coupling partner molecule.
[0361] Pairs of docking groups and labeling groups that can react via the above-mentioned click reaction are known in the art. Examples of suitable UNAAs that contain docking groups include, but are not limited to, those described in WO2012 / 104422 and WO2015 / 107064.
[0362] Examples of specific suitable pairs of docking groups (in the binding partner molecule) and labeling groups (in the UNAA residue of the POI) include, but are not limited to, the following: (a) a docking group comprising (or consisting essentially of) a group selected from an azide group, a nitrile oxide functional group (i.e., a group represented by the formula, a nitrone functional group or a diazocarbonyl group) in combination with a labeling group comprising (or consisting essentially of) an optionally substituted strained alkynyl group, which groups are capable of reacting covalently in a copper-free strain-promoted alkyne-azide cycloaddition reaction (SPAAC); (b) a combination of a docking group comprising (or consisting essentially of) an optionally substituted strained alkynyl group and a labeling group comprising (or consisting essentially of) a group selected from an azide group, a nitrile oxide functional group, which groups can react covalently in a copper-free strain-promoted alkyne-azide cycloaddition reaction (SPAAC); (c) a combination of a docking group comprising (or consisting essentially of) a group selected from an optionally substituted strained alkynyl group, an optionally substituted strained alkenyl group, and a norbornenyl group, and a labeling group comprising (or consisting essentially of) an optionally substituted tetraalkynyl group, which groups are capable of reacting covalently in a copper-free strain-promoted inverse electron demand Diels-Alder cycloaddition reaction (SPIEDAC); (d) A combination of a docking group comprising (or consisting essentially of) an optionally substituted tetrazinyl group with a labeling group comprising (or consisting essentially of) a group selected from an optionally substituted strained alkynyl group, an optionally substituted strained alkenyl group, and a norbornenyl group, which groups can react covalently in a copper-free strain-promoted inverse electron demand Diels-Alder cycloaddition reaction (SPIEDAC).
[0363] Optionally substituted strained alkynyl groups include, but are not limited to, optionally substituted trans-cyclooctenyl groups (e.g., those described in WO2012 / 104422 and WO2015 / 107064). Optionally substituted strained alkenyl groups include, but are not limited to, optionally substituted cyclooctynyl groups (e.g., those described in WO2012 / 104422 and WO2015 / 107064). Optionally substituted tetrazinyl groups include, but are not limited to, those described in WO2012 / 104422 and WO2015 / 107064.
[0364] An azide group is a group of the formula -N3.
[0365] The nitrone functional group has the formula -C(Rx )=N + (R y )-O - In the formula, R x and R y are independently selected from organic residues, such as C1-C6-alkyl as described herein.
[0366] A diazocarbonyl group is a group of formula -C(O)-CH=N2.
[0367] The nitrile oxide functional group has the formula -C≡N + -O - or preferably of the formula -C=N + (R x )-O - In the formula, R x are independently selected from organic residues, such as C1-C6-alkyl as described herein.
[0368] Cyclooctynyl is an unsaturated alicyclic group having 8 carbon atoms and one triple bond in the ring structure.
[0369] "Trans-cyclooctenyl" is an unsaturated alicyclic group having eight carbon atoms in the ring structure and one double bond in the trans configuration.
[0370] "Tetradzinyl" is a six-membered monocyclic aromatic group having four nitrogen ring atoms and two carbon ring atoms.
[0371] Unless otherwise indicated, the term "substituted" means that the group is substituted with 1, 2 or 3, in particular 1 or 2, substituents. In certain embodiments, these substituents are hydrogen, halogen, C1-C4-alkyl, (R a O)2P(O)O-C1-C4-alkyl,(R bO)2P(O)-C1-C4-alkyl, CF3, CN, hydroxyl, C1-C4-alkoxy, -O-CF3, C2-C5-alkenoxy, C2-C5-alkanoyloxy, C1-C4-alkylaminocarbonyloxy or C1-C4-alkylthio, C1-C4-alkylamino, di-(C1-C4-alkyl)amino, C2-C5-alkenylamino, N-C2-C5-alkenyl-N-C1-C4-alkyl-amino and di-(C2-C5-alkenyl)amino, where R a and R b are independently selected from hydrogen or C2-C5-alkanoyloxymethyl.
[0372] The term halogen in each case denotes a fluorine, bromine, chlorine or iodine radical, in particular a fluorine radical.
[0373] C1-C4-Alkyl is a straight-chain or branched alkyl radical having 1 to 4, in particular 1 to 3, carbon atoms. Examples include C2-C4-alkyl, such as methyl and ethyl, n-propyl, isopropyl, n-butyl, 2-butyl, isobutyl and tert-butyl.
[0374] C2-C5-alkenyl is a monounsaturated hydrocarbon group having 2, 3, 4 or 5 carbon atoms. Examples include vinyl, allyl (2-propen-1-yl), 1-propen-1-yl, 2-propen-2-yl, methallyl (2-methylprop-2-en-1-yl), 1-methylprop-2-en-1-yl, 2-buten-1-yl, 3-buten-1-yl, 2-penten-1-yl, 3-penten-1-yl, 4-penten-1-yl, 1-methylbut-2-en-1-yl and 2-ethylprop-2-en-1-yl.
[0375] C1-C4-alkoxy is a group of the formula RO-, where R is a C1-C4-alkyl group as defined herein.
[0376] C2-C5-alkenoxy is a group of the formula RO-, wherein R is C2-C5-alkenyl as defined herein.
[0377] C2-C5-Alkanoyloxy is a group of the formula RC(O)-O-, where R is C1-C4-alkyl as defined herein.
[0378] C1-C4-Alkylaminocarbonyloxy is a group of the formula R-NH-C(O)-O-, wherein R is C1-C4-alkyl as defined herein.
[0379] C1-C4-Alkylthio is a group of the formula RS-, where R is C1-C4-alkyl as defined herein.
[0380] C1-C4-Alkylamino is a group of the formula R-NH-, where R is C1-C4-alkyl as defined herein.
[0381] Di-(C1-C4-alkyl)amino is a compound of the formula R x -N(R y )- group, where R x and R y are independently C1-C4-alkyl as defined herein.
[0382] C2-C5-alkenylamino is a group of the formula R-NH-, wherein R is C2-C5-alkenyl as defined herein.
[0383] N-C2-C5-alkenyl-N-C1-C4-alkylamino is a compound of the formula R x -N(R y )- group, where R x is C-C-alkenyl as defined herein; R y is C1-C4-alkyl.
[0384] Di-(C2-C5-alkenyl)amino is a compound of the formula R x -N(R y )- group, where Rx and R y are independently C2-C5-alkenyl as defined herein.
[0385] C2-C5-Alkanoyloxymethyl is a compound of the formula R x -C(O)-O-CH2- group. x is C1-C4-alkyl as defined herein.
[0386] UNAA used in the context of the present invention can be used in the form of their salts. Salts of UNAA described herein refer to acid or base addition salts, particularly addition salts with physiologically acceptable acids or bases. Physiologically acceptable acid addition salts can be formed by treating UNAA in its basic form with a suitable organic or inorganic acid. UNAA containing an acidic proton can be converted into their non-toxic metal or amine addition salt forms by treating with suitable organic and inorganic bases. UNAA and its salts described in the context of the present invention also include their hydrates and solvent addition forms, such as hydrates, alcoholates, etc.
[0387] A physiologically acceptable acid or base is one that is tolerated by the translation system used to generate the POI having UNAA residues, e.g., is substantially non-toxic to living eukaryotic cells.
[0388] UNAA and its salts useful in the context of the present invention are well known in the art and can be prepared, for example, similarly to the methods described in the various publications cited herein.
[0389] The nature of the coupling partner molecule depends on the intended use. For example, the POI may be coupled to a molecule suitable for imaging methods or may be functionalized by coupling to a biologically active molecule. For example, in addition to the docking group, the coupling partner molecule may have a group selected from, but not limited to, the following: dyes (e.g., fluorescent, luminescent or phosphorescent dyes such as dansyl, coumarin, fluorescein, acridine, rhodamine, silicon rhodamine, BODIPY or cyanine dyes); molecules capable of fluorescing upon contact with a reagent; chromophores (e.g., phytochromes, phycobilins, bilirubin, etc.); radiolabels (e.g., radioactive forms of hydrogen, fluorine, carbon, phosphorus, sulfur or iodine, e.g., tritium, 18 F, 11 C, 14 C, 32 P, 33 P, 33 S, 35 S, 11 In, 125 I, 123 I, 131 I, 212 B, 90 Y or 186Rh, etc.); MRI-sensitive spin labels; affinity tags (e.g., biotin, His tag, Flag tag, Strep tag, sugars, lipids, sterols, PEG linkers, benzylguanine, benzylcytosine, or cofactors); polyethylene glycol groups (e.g., branched PEG, linear PEG, PEG of various molecular weights, etc.); photocrosslinkers (e.g., p-azidoiodoacetanilide, etc.); NMR probes; X-ray probes; pH probes; IR probes; resins; solid supports, and bioactive compounds (e.g., synthetic drugs). Suitable bioactive compounds include, but are not limited to, cytotoxic compounds (e.g., cancer chemotherapy compounds), antiviral compounds, biological response modifiers (e.g., hormones, chemokines, cytokines, interleukins, etc.), microtubule acting agents, hormone modulating agents, and steroid compounds. Examples of useful coupling partner molecules include, but are not limited to, members of receptor / ligand pairs, members of antibody / antigen pairs, members of lectin / carbohydrate pairs, members of enzyme / substrate pairs, biotin / avidin, biotin / streptavidin, and digoxin / antidigoxin.
[0390] In particular, the ability of a particular UNAA residue (labeling group) to be covalently attached in situ to a binding partner molecule (docking group) by the click reaction described herein can be used to detect a POI bearing that UNAA residue within a eukaryotic cell or tissue expressing the POI, and to study the distribution and fate of the POI. In particular, the method of the invention for producing a POI by expression in a eukaryotic cell can be combined with super-resolution microscopy (SRM) to detect the POI within a cell or tissue of the cell. Several SRM methods are known in the art and can be adapted to utilize click chemistry to detect a POI expressed by a eukaryotic cell of the invention. Specific examples of such SRM methods include DNA-PAINT (DNA point accumulation for imaging in nanoscale topography; e.g., as described in Jungmann et al., Nat Methods 11:313-318, 2014), dSTORM (direct stochastic optical reconstruction microscopy) and STED (stimulated emission depletion) microscopy.
[0391] The following examples are illustrative only and are not intended to limit the scope of the embodiments described herein.
[0392] Numerous possible variations that will be readily apparent to those of ordinary skill in the art after reviewing the disclosure provided herein are also within the scope of the present invention. EXAMPLES
[0393] Experimental part Unless otherwise indicated, the cloning steps performed in the context of the present invention, such as restriction digestion, agarose gel electrophoresis, purification of DNA fragments, transfer of nucleic acids to nitrocellulose and nylon membranes, ligation of DNA fragments, transformation of microorganisms, culturing of microorganisms, propagation of phage, and sequence analysis of recombinant DNA, are performed by application of well-known techniques, e.g., as described in Sambrook et al. (1989), supra, unless otherwise indicated.
[0394] A. Materials and Methods chemical products TCO-E, TCO * A and SCO were purchased from SiChem (SIRIUS FINE CHEMICALS, SICHEM GMBH, Germany). N ε -tert-Butyloxycarbonyl-L-lysine (Boc) was purchased from Iris Biotech GmbH (Germany).
[0395] cell culture HEK293T cells (ATCC CRL-3216) were cultured in Dulbecco's Modified Eagle Medium (DMEM, Gibco 41965-039) supplemented with 10% FBS (Sigma-Aldrich F7524), 1% penicillin-streptomycin (Sigma-Aldrich P0781), 1% L-glutamine (Sigma-Aldrich G7513), and 1% sodium pyruvate (Life Technologies 11360). Cells were cultured at 37°C and 5% CO2 and split every 2–3 days until passage 20. HEK293T cells were seeded in 24-well plates (Nunclon Delta Surface ThermoFisher SCIENTIFIC) at a cell density of 220.000 cells / ml, 500 μl / well. Cells were seeded 16 h prior to transfection. For transfection, phenol red-free Dulbecco's modified Eagle's medium (DMEM, Gibco 11880-028) was used, and polyethyleneimine (PEI, Sigma 408727) was used as the transfection reagent at a concentration of 1 mg / ml.
[0396] Escherichia coli: ElectroMAX (trademark) DH10B F - mcrA Δ(mrr-hsdRMS-mcrBC) Φ80lacZΔM15 ΔlacX74 recA1 endA1 araD139Δ(ara,leu)7697 galU galK λ - rpsL nupG (ThermoFisher Scientific, Catalog Number: 18290015)
[0397] Culture medium 2xYT medium was prepared in-house. SOC medium was prepared in-house.
[0398] Cloning of constructs For expression in mammalian cells, we used the reporter plasmid pCI-iRFP-EGFP, shown in the upper diagram of Figure 1A, as previously published. Y39TAG -6His (sequence number 97) was used (Nikic, I. et al. Debugging Eukaryotic Genetic Code Expansion for Site-Specific Click-PAINT Super-Resolution Microscopy. Angew. Chemie-Int. Ed. 55, 16172-16176 (2016)).
[0399] Plasmid pCMV-NES-PylRS AF -U6tRNArv (SEQ ID NO: 98) is a PylRS tRNA-synthetase from Methanosarcina mazei (Mm PylRS AF ), as well as the U6 promoter signal and Methanosarcina mazei tRNA Pyl It carries a tRNA expression cassette in the reverse orientation consisting of a gene (lower panel of Fig. 1A ).
[0400] For evolution of the synthetase, the PylRS gene was cloned into the pBK plasmid, a gift from Ryan Mehl (Cooley, RB et al. Structural Basis of Improved Second-Generation 3-Nitro-tyrosine tRNA Synthetases. Biochemistry 53, 1916-1924 (2014)). First, a BglII site was introduced at bps 870–875 of the PylRS gene using the restriction sites NdeI and PstI, followed by the addition of the Methanosarcina mazei PylRSWT was cloned into the pBK plasmid to create pBK-PylRS, which contains a kanamycin resistance cassette (Kan). WT The plasmid (SEQ ID NO: 99) was obtained (FIG. 1B, upper panel). Furthermore, the BglII site was used together with PstI to insert PylRS WT By replacing this portion of the synthetase, the PylRS gene library was placed into this plasmid.
[0401] Plasmids used for evolution of the synthase: pREP-PylT (SEQ ID NO: 100) (lower diagram in FIG. 1B), pYOBB2-PylT (SEQ ID NO: 101), and pALS-sfGFP N150TAG -MbPyl-tRNA (sequence number 102) (upper and lower panels of FIG. 1C, respectively) was a gift from the Ryan Mehl lab (Cooley, RB et al. Structural Basis of Improved Second-Generation 3-Nitro-tyrosine tRNA Synthetases. Biochemistry 53, 1916-1924 (2014); Porter, JJ et al. Genetically Encoded Protein Tyrosine Nitration in Mammalian Cells. ACS Chem. Biol. 14, 1328-1336 (2019)).
[0402] The PylRS synthase mutants AF A1, AF B11, AF C11, AF G3, and AF H12 extracted by the synthase selection procedure were cloned from the pBK plasmid and digested with BglII and PstI, followed by pCMV-NES-PylRS AF -U6tRNArv. PylRS AF The corresponding nucleotides were replaced.
[0403] Site-directed mutagenesis To obtain a synthetase mutant lacking the Y306A and Y384F mutations, these amino acids were reverted to 306Y and 384Y by site-directed mutagenesis, and PylRS mutant A1 was obtained by two-step cloning. Primers PylRS A306Y fw (5'aatctttataactatatgcgcaaactggaccgtgc 3') (SEQ ID NO: 103) and PylRS A306Y rv (5'atagttataaagatttggtgctagcatagggcgc 3') (SEQ ID NO: 104) and primer pair PylRS F384Y fw (5'atggtgtatggcgacaccctggatgtcatg 3') (SEQ ID NO: 105) and PylRS F384Y rv (5'tgtcgccatacaccatacagctgtcgcccac 3') (SEQ ID NO: 106) were used.
[0404] B. Working Example Example 1: Evolution of PylRS synthase to incorporate TCO-E Methanosarcina mazei PylRS AF A synthetic NNK library based on the mutants (Y306A, Y384F) was ordered from GenScript Biotech Corp. In this library, five sites (L305, L309, C348, I405, and W417) are mutated to contain one of 20 amino acids. Two stop codons are excluded from these positions, resulting in 32 possible variants at each of the five sites, resulting in a library size of 3.3 × 10 7 The library is selected using the plasmid pBK-PylRS WT The PylRS binding pocket was replaced with the library gene, resulting in pBK-PylRS. lib The selection pREP-PylT was transformed into E. coli DH10B cells, and highly competent electrocompetent cells were freshly prepared. OD 600The cells were grown until the β-acetylglucosamine concentration was 0.5. After harvesting the cells, they were washed with 10% glycerol. To resuspend the cells after each harvesting step, the harvesting bottle should be shaken and the cells mixed gently, avoiding pipetting steps. The washing steps were repeated twice and finally the cells were taken and dispensed in as small a volume as possible (for 1 liter of expression in 1 ml of 10% glycerol).
[0405] 100ng pBK-PylRS lib was transformed into 50 μl of DH10B(pREP-PylT) cells by electroporation in a 1 mm cuvette and 800 μl of SOC medium was added directly. This step was repeated 10 times and the cells were placed in a 50 ml shake flask and incubated at 37 °C with shaking at 200 rpm for 1 h. To estimate the coverage of the library after transformation, LB-agar plates containing Tet and Kan were prepared and serial dilutions (1:10) were performed using 10 μl of the cell suspension. 2 ~1:10 7 ) was performed. 100 μl of each dilution was plated on LB-agar plates and incubated overnight in a 37°C incubator. The remaining 10 ml of transformation mixture was added to 500 ml of 2xYT medium containing Tet and Kan as antibiotics and incubated overnight at 37°C with shaking. To estimate the coverage of the library, 1:10 6 45 colonies were counted on the plate, resulting in 140-fold coverage of the library size.
[0406] The next day, the cells were diluted 1:100 in 500 ml of fresh 2xYT medium (Tet, Kan) and the OD 600 The bacteria were grown until the β-amino acid concentration reached 1 (usually this takes 2–4 hours). For the first positive selection, ten LB-agar plates (150 mm petri dishes) containing 1 mM TCO-E and 60 μg / ml chloramphenicol (Cm), Tet, and Kan were prepared. As a control, one plate without Cm was similarly prepared. The plates were prepared under sterile conditions and cooled to dryness.
[0407] To initiate the selection, 100 μl of the culture was plated on each 150 mm LB-agar plate, spread with glass beads, and dried next to a flame. The plates were grown at 37°C overnight, but not longer than 16 h. The resulting colonies were scraped off the plates with 5 ml of 2xYT medium per plate using a cell scraper. The cell suspensions of all 10 plates were pooled in a 50 ml Erlenmeyer flask and shaken at 37°C for 1 h. To isolate the library plasmids, DNA was first extracted with a Miniprep Kit (Invitrogen), followed by gel extraction, where the DNA was loaded on a 1% agarose gel and the lowest band was excised. DNA was isolated with a gel extraction kit (Invitrogen). The negative selection plasmid, pYOBB2-PylT, containing the barnase gene with two amber sites (Gln2 and Asp44), was transformed into DH10B cells and electrocompetent cells were prepared as described above. 10 ng of gel-extracted library plasmid was transformed into 50 μl of freshly prepared DH10B(pYOBB2-PylT) cells and, after electroporation, cells were recovered in 800 μl of SOC medium in 14 ml tubes at 37°C for 1 h with shaking at 200 rpm. During the incubation period, plates for negative selection were prepared. Six 150 mm dishes were cast with LB-agar medium (three of which contained 0.2% arabinose to induce barnase expression). All were incubated with 50 μg / ml Kan(pBK-PylRS lib plasmid) and 33 μg / ml Cm (pYOBB2-PylT plasmid).
[0408] 100 μl of pure cells and 100 μl of the 1:10 and 1:100 dilutions were plated on two plates each, one with arabinose and one without. The plates were dried next to a flame and then grown overnight in a 37°C incubator. Colonies were scraped off the plates and DNA was extracted as described above. Positive selection was repeated once more, transforming 10 ng of gel extracted library plasmid into 50 μl of DH10B(pREP-PylT) and recovering in 800 μl of SOC medium, shaking at 37°C for 1 h. This time, only three 15 mm LB-agar plates containing 33 μg / ml Cm(Kan, Tet) and 1 mM TCO-E and one control plate without TCO-E were used for selection. 100 μl of cell suspension was plated and grown overnight at 37°C. Surviving colonies were scraped off the plates, and library plasmids were extracted from a 1% agarose gel.
[0409] To examine the amber suppression efficiency of the remaining PylRS mutants, we performed a superfolder GFP (sfGFP)-based expression assay. The library DNA was then transfected with plasmid pALS-sfGFP, which encodes sfGFP with an amber site at position N150. N150TAG -MbPyl-tRNA and tRNA from Methanosarcina barkeri Pyl Both were transformed into DH10B cells. WT The pBK-PylRS plasmid encoding PylRS AF The plasmid encoding the GFP gene was also pALS-sfGFPN. 150TAG -MbPyl-tRNA (SEQ ID NO: 102) plasmid. Transformants were allowed to recover in SOC medium for 1 h and various amounts (50, 100 and 200 μl) were plated onto autoinducing minimal medium plates (Kan, Tet) containing 1 mM TCO-E and no ncAA. Colonies were grown at 37° C. for 24 h and, if necessary, at room temperature for an additional 24 h.
[0410] Green colonies were selected and cultured overnight in 96-well plates containing 480 μl of non-inducing minimal medium per well. WT and PylRS AF ) were also selected and cultured in this 96-well plate. The next day, two 96-well plates were prepared, with 480 μl of autoinducing minimal medium in each well, one with 1 mM TCO-E and the other without. 20 μl of each overnight culture was pipetted into the corresponding well of a new 96-well plate and cultured at 37°C for 24 h with shaking at 250 rpm. sfGFP expression was analyzed by fluorometry (BIOTEK Synergy 2 Microplate Reader). To this end, 96-well plates containing 180 μl of water per well were prepared, and 20 μl of sfGFP expression culture was pipetted into the corresponding well. The OD of each expression was 600 To correct for differences in the concentrations of 150 μl water + 50 μl expression culture, optical density scans at 600 nm were also performed (150 μl water + 50 μl expression culture).
[0411] From the 96-well plate containing non-inducing minimal medium, each well of interest can be further analyzed, for example, DNA of PylRS mutants can be extracted for sequencing and further subcloned.
[0412] In this way, the PylRS synthase mutants AF A1, AF B11, AF C11, AF G3 and AF H12 were identified.
[0413] Example 2: Estimation of the incorporation efficiency of various synthetase mutants These new novel PylRS mutants obtained in Example 1 were tested in HEK293T cells using a fluorescent reporter by fluorescence flow cytometry (FFC). The reporter contains an infrared fluorescent protein (iRFP) fused to an enhanced green fluorescent protein (EGFP), with an amber stop codon at Y39, iRFP-EGFP. Y39TAGTo analyze the integration efficiency of each new PylRS mutant, we performed transient transfections with various ncAAs and various amounts of ncAAs.
[0414] The transfected cells are able to express iRFP, which shows a signal on the vertical axis of the FFC plot. * When tRNA can be charged with TCO-A or TCO-E), a green signal due to EGFP expression is also observed on the horizontal axis.
[0415] The signals of iRFP and EGFP were measured 24 hours after transfection. Therefore, HEK293T cells were seeded in 24-well plates 16 hours before transfection. For the transfection of two plasmids, 1 μg of total DNA was used per well. This DNA was mixed with 50 μl of DMEM medium without phenol red and 3 μl of PEI was added. After vortexing for 10 seconds and a short centrifugation step, the DNA mixture was incubated in the hood for 15 minutes and then added dropwise to the wells. Master mixes were prepared for all wells at the appropriate time. After 4 hours of incubation, the medium was aspirated and fresh medium supplemented with various concentrations of ncAA was added. The concentration of ncAA was 250 μM. After 20 hours of culture, the iRFP and EGFP signals were analyzed using a flow cytometer device (LSRFortessa™, BD Biosciences). Stock solutions of all ncAAs were prepared as previously described (Nikic, I., Kang, JH, Girona, GE, Aramburu, IV & Lemke, EA Labeling proteins on live mammalian cells using click chemistry. Nat. Protoc. 10, 780-791 (2015)).
[0416] The results of these expression studies with various ncAAs are shown in Figure 2. All new variants were significantly higher than the TCOs tested here. *A (Figure 2A) and TCO-E (Figure 2B). Furthermore, when TCO-E was used, all the new mutants were able to uptake PylRS. AF Figure 2C shows the expression profile in the absence of ncAA.
[0417] Example 3: Evaluation of the importance of mutations Y306A and Y384F on the incorporation of bulky ncAAs To investigate whether the mutations Y306A and Y384F are indeed important for the uptake of bulky ncAAs, we reverted these amino acids to their original amino acid residues by site-directed mutagenesis in the PylRS AF A1 mutant, and thus the new mutant does not contain the Y306A and Y384F mutations and is called PylRS A1.
[0418] The FFC test was performed using mutants of PylRS A1, PylRS AF, PylRS AF A1, and PylRS MMA, and was performed using TCO * A, TCO-E, and also Boc. PylRS MMA has the mutations 306M 309M 348A. Figure 3 shows the data obtained by FFC in bar graphs for the various ncAAs used. The top bar graph shows the ratio of the mean GFP signal obtained by FFC divided by the mean iRFP signal, which reflects the integration efficiency. The middle bar graph shows the same data, but normalized to the ratio observed for the PylRS AF mutant with 100 μM ncAA. The bottom bar graph shows the mean GFP / mean iRFP ratio normalized to the mean GFP / mean iRFP ratio obtained at the desired ncAA concentration for the PylRS AF mutant. These bar graphs show how much higher the integration efficiency of each mutant is compared to PylRS AF.
[0419] All mutants (including PylRS A1) are able to incorporate bulky ncAAs. This finding is unexpected and new to the art, since Y306A and Y384F have always been known to be key mutations for incorporating bulky ncAAs. Figure 3A shows the incorporation of TCO-E with four different PylRS mutants. The highest incorporation efficiency is obtained with PylRS A1 and PylRS AF A1, but not with PylRS AF. Figure 3B shows the incorporation efficiency of TCO-E with four different PylRS mutants. * The FFC data using A as the ncAA shows the incorporation efficiency. In this case, PylRS A1 shows the best incorporation. Figure 3C shows the data for Boc incorporation. Again, PylRS A1 can achieve the highest incorporation efficiency.
[0420] For TCO-E, PylRS A1 was up to 30 times more effective than PylRS AF depending on the concentration used, and TCO * For A incorporation, PylRS A1 is 1.2-fold more efficient, regardless of the ncAA concentration used, and for Boc, PylRS A1 is up to 4.5-fold more efficient than PylRS AF.
[0421] In summary, the novel mutant PylRS A1 contains mutations that make it a better synthesizer for bulky ncAAs compared with the previously known PylRS AF mutant, which was unexpected from the literature.
[0422] Taken together, the novel mutant PylRS MMA exhibits up to 10-fold higher TCO-E incorporation efficiency compared to PylRS AF.
[0423] The contents of any documents cross-referenced herein are incorporated by reference.
[0424] Example 4: Evaluation of the incorporation efficiency of cyclooctyne-lysine (SCO) by the novel mutant PylRS A1 To test the incorporation efficiency of other ncAAs by the novel mutant PylRS A1, we performed an incorporation assay using HEK293T cells transfected with the reporter plasmid pCI-iRFP-EGFP. Y39TAG -6His (SEQ ID NO: 97), together with the corresponding tRNA Pyl The cells were co-transfected with a plasmid containing the NES-PylRS A1 gene. Four hours after cell transfection, various concentrations of SCO ranging from 15.625 μM to 500 μM were added to the growth medium. After 20 hours of incubation, the iRFP and EGFP signals were analyzed as described above using a flow cytometer instrument (LSRFortessa™, BD Biosciences). The corresponding FFC plots are shown in Figure 5. The new mutant A1 can efficiently incorporate SCO.
[0425] The contents of the documents cross-referenced herein above are incorporated by reference.
[0426] Sequence Listing This list is considered part of the general disclosure of the invention.
[0427] [Table 24]
[0428] [Table 25]
[0429] [Table 26]
[0430] [Table 27]
[0431] [Table 28]
[0432]
Table 29
Claims
1. 1. A modified archaeal pyrrolysyl-tRNA synthetase (PylRS), comprising: a modified archaeal PylRS comprising a combination of modified sequence motifs M1 and M3; optionally combined with at least one additional sequence motif selected from M2, M4, M5, and M6, and retaining PylRS activity; wherein M1, M2, M3, M4, M5, and M6 are arranged in the amino acid sequence of the modified PylRS in the order stated above, with M1 being closest to the N-terminus of the sequence and M6 being closest to the C-terminus of the sequence, and comprising the following sequence: M1: LRPMX 1 AX 2 X 3 L(Y / M)X 5 X 6 (M / V / C)R (SEQ ID NO: 1) M2:HLX 7 EFTMX 8 NX 9 (G / A)X 11 X 12 G (SEQ ID NO: 2) M3: VYX 13 X 14 TX 15 D (SEQ ID NO: 3) M4:SX 16 X 17 X 18 GP(R / I / N)X 20 X 21 D (SEQ ID NO: 4) M5:X 22 X 23 (I / V)X 25 X 26 PW (SEQ ID NO: 5) M6: G(A / L / I)GFGLERLL (SEQ ID NO: 6) Here, amino acid residue X 1 ~X 26 are independently selected from natural amino acid residues.
2. 2. The modified archaeal PylRS of claim 1, comprising a combination of sequence motifs M1, M3 and M2; or M1, M3, M2 and M4; or M1, M3, M2, M4 and M5 and / or M6.
3. and parent PylRSs derived from archaea of the genera Methanosarcina, Methanosarcinaceae, Methanomethylophilus, Desulfitobacterium, and Candidatus Methanoplasma, in particular Methanosarcina mazeii (SEQ ID NO: 56), Methanosarcina barkeri (SEQ ID NO: 58), Methanosarcinaceae archaeon (SEQ ID NO: 60), Methanosarcina alvus (SEQ ID NO: 62), Desulfitobacterium hafniense (SEQ ID NO: 64), and Candidatus 3. The modified archaeal PylRS of claim 1 or 2, which is derived from a parent PylRS derived from an archaeon of the species Methanoplasma termitum (SEQ ID NO: 66).
4. A modified archaeal PylRS derived from a parent PylRS having an amino acid sequence selected from SEQ ID NOs: 56, 58, 60, 62, 64, and 66, or a functional variant or fragment thereof, which retains pyrrolysyl-tRNA synthetase activity and has at least 60% sequence identity to a native pyrrolysyl-tRNA synthetase, A modified archaeal PylRS according to claim 1 or 2, comprising a combination of modified sequence motifs M1 and M3; optionally combined with at least one further sequence motif selected from M2, M4, M5 and M6, each as defined in claim 1.
5. 2. The modified archaeal PylRS of claim 1, a) the sequence motif M1 is selected from the following sequences: 【Table 1】 b) the sequence motif M2 is selected from the following sequences: 【Table 2】 c) the sequence motif M3 is selected from the following sequences: 【Table 3】 d) the sequence motif M4 is selected from the following sequences: 【Table 4】 e) the sequence motif M5 is selected from the following sequences: 【Table 5】 f) the sequence motif M6 is selected from the following sequences: 【Table 6】 Engineered archaeal PylRS.
6. 2. The modified archaeal PylRS of claim 1, a) the sequence motif M1 is selected from the following sequences: 【Table 7】 b) the sequence motif M2 is selected from the following sequences: 【Table 8】 c) the sequence motif M3 is selected from the following sequences: 【Table 9】 d) the sequence motif M4 is selected from the following sequences: 【Table 10】 e) the sequence motif M5 is selected from the following sequences: 【Table 11】 f) the sequence motif M6 is selected from the following sequences: 【Table 12】 Engineered archaeal PylRS.
7. 6. The modified archaeal PylRS of claim 5, comprising a combination of sequence motifs M1a, M2a, M3a, M4a, M5a and M6a, or the sequence motif M1a * , M2a * , M3a, M4a * 7. The modified archaeal PylRS of claim 6, comprising a combination of M5a, M6a and M7a.
8. 3. The modified archaeal PylRS of claim 1 or 2, a) PylRS A1 comprising the amino acid sequence of SEQ ID NO: 70; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 70; or a functional fragment thereof that retains PylRS activity; or b) a PylRS MMA comprising the amino acid sequence of SEQ ID NO: 72; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 72; or a functional fragment thereof that retains PylRS activity; or c) PylRS B11 comprising the amino acid sequence of SEQ ID NO: 82; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 82; or a functional fragment thereof that retains PylRS activity; or d) PylRS C11 comprising the amino acid sequence of SEQ ID NO: 84; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 84; or a functional fragment thereof that retains PylRS activity; or e) PylRS G3 comprising the amino acid sequence of SEQ ID NO: 86; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 86; or a functional fragment thereof that retains PylRS activity; or f) PylRS H12 comprising the amino acid sequence of SEQ ID NO: 88; or an amino acid sequence having at least 60% sequence identity to SEQ ID NO: 88; or a functional fragment thereof that retains PylRS activity; A modified archaeal PylRS.
9. 3. The modified archaeal PylRS of claim 1 or 2, which exhibits at least one of the following functional characteristics: a) an altered substrate profile for non-canonical amino acids (ncAA); b) Compared to mutant PylRS AF (SEQ ID NO: 68), in particular TCO-E, TCO * Improved utilization of at least one bulky ncAA selected from A and Boc; c) mutant PylRS AF A1 (SEQ ID NO: 108), respectively, especially TCO * Improved utilization of at least one bulky ncAA selected from A and Boc.
10. The modified archaeal PylRS of claim 8 a) exhibits at least one of the following functional characteristics: a) TCO-E, ... * Improved utilization of at least one bulky ncAA selected from A and Boc; b) TCO compared to mutant PylRS AF A1 (SEQ ID NO: 108), respectively * Improved utilization of at least one bulky ncAA selected from A and Boc.
11. The modified archaeal PylRS according to claim 8 b) exhibits at least the following functional characteristics: Improved utilization of the bulky ncAA, TCO-E, compared to mutant PylRS AF (SEQ ID NO: 68).
12. The modified archaeal PylRS of claim 1 or 2, comprising a nuclear export signal (NES).
13. A modified polynucleotide encoding the modified archaeal pyrrolysyl-tRNA synthetase of claim 1.
14. tRNA Pyl 14. The polynucleotide of claim 13, further encoding: The tRNA Pyl is a tRNA that can be acylated by a pyrrolysyl-tRNA synthetase encoded by the polynucleotide of claim 13.
15. At least one polynucleotide according to claim 13 and a tRNA according to claim 14. Pyl A combination of polynucleotides comprising at least one polynucleotide encoding:
16. tRNA Pyl 15. The polynucleotide of claim 14, wherein the anticodon of is the reverse complement of a codon selected from a stop codon, a four-base codon, and a rare codon.
17. The polynucleotide combination of claim 15, wherein the anticodon of the tRNA Pyl is the reverse complement of a codon selected from a stop codon, a four-base codon, and a rare codon.
18. Eukaryotic cells, particularly mammalian cells, including: (a) a polynucleotide sequence encoding the modified archaeal pyrrolysyl-tRNA synthetase of claim 1; and (b) a tRNA that can be acylated by the pyrrolysyl-tRNA synthetase encoded by the sequence of (a), or a polynucleotide sequence encoding the tRNA.
19. 1. A method for preparing a protein of interest (POI) comprising one or more unnatural amino acid residues, the method comprising the steps of: (a) providing a eukaryotic cell comprising: (i) a modified archaeal pyrrolysyl-tRNA synthetase according to claim 1 or 2; (ii) tRNA (tRNA) Pyl ); (iii) an unnatural amino acid or a salt thereof; and (iv) a polynucleotide encoding a POI, wherein any position in the POI occupied by an unnatural amino acid residue is selected from the group consisting of tRNA, tRNA, tRNA- ... Pyl a polynucleotide encoded by a codon that is the reverse complement of an anticodon contained in Here, the modified archaeal pyrrolysyl-tRNA synthetase (i) is a tRNA Pyl (ii) can be acylated with an unnatural amino acid or salt (iii); and (b) allowing translation of the polynucleotide (iv) by a eukaryotic cell, thereby producing the POI.
20. 1. A method for preparing a polypeptide conjugate, comprising the steps of: (a) preparing a POI comprising one or more unnatural amino acid residues using the method of claim 19; and (b) reacting the POI with one or more binding partner molecules such that the binding partner molecules are covalently attached to the non-natural amino acid residues of the POI.
21. 1. A kit comprising at least one unnatural amino acid or salt thereof and: (a) the polynucleotide of claim 13; or (b) a combination of polynucleotides according to claim 15; or (c) the eukaryotic cell of claim 18; Here, the archaeal pyrrolysyl-tRNA synthetase is a tRNA Pyl can be acylated with an unnatural amino acid or its salt.