Methods and compositions

JP2025023942A5Pending Publication Date: 2025-10-21UNITED KINGDOM RESEARCH AND INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024186312
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-26
Filing Date
2024-10-23
Publication Date
2025-10-21

Smart Images

  • Figure 00000114_0000
    Figure 00000114_0000
  • Figure 00000115_0000
    Figure 00000115_0000
  • Figure 00000115_0001
    Figure 00000115_0001
Patent Text Reader

Abstract

To provide a method for producing a polypeptide containing 2,3-diaminopropionic acid.SOLUTION: The invention relates to genetic incorporation of 2,3-diaminopropionic acid (DAP) into polypeptides, to unnatural amino acids comprising DAP, to a tRNA synthetase for charging tRNA with unnatural amino acids comprising DAP, and to methods of using the resulting polypeptides, for example in capturing substrates and / or intermediates in enzymatic reactions. The invention also relates to unnatural amino acids represented by the following formula, where X1 is S, Se, O, NH or the like.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to the genetic incorporation of 2,3-diaminopropionic acid (DAP) into polypeptides, unnatural amino acids that contain DAP, tRNA synthetases for charging unnatural amino acids that contain DAP onto tRNA, and methods of using the resulting polypeptides, for example, in capturing substrates and / or intermediates in enzymatic reactions. [Background technology]

[0002] Many enzymes carry out reactions that proceed through covalent intermediates attached to serine or cysteine ​​side chains in the enzyme active site (Holliday, GL, Mitchell, JBO & Thornton, JM Understanding the Functional Roles of Amino Acid Residues in Enzyme Catalysis. Journal of molecular biology 390, 560-577 (2009)). Hydroxyl or sulfhydryl groups The reaction of oxoesters with carbonyl groups in the substrate forms activated ester or thioester intermediates, which are rapidly converted to products via further reaction with selected nucleophiles. The half-lives of thioesters and esters are usually several minutes to several hours (Yang, W. & Drueckhammer, DG Understanding the relative acyl-transfer reactivity of oxoesters and Because of the lack of catalytic activity of acyl-enzyme intermediates (e.g., thioesters: computational analysis of transition state delocalization effects. Journal of the American Chemical Society 123, 11004-11009 (2001)), these key acyl-enzyme intermediates have been difficult to isolate and characterize.

[0003] Strategies to stably trap these intermediates would allow for the identification of native substrates and characterization of otherwise difficult to define intermediates and functional states. In a known approach, substrate analogs with electrophilic substitution of the carbonyl group can be used to trap analogs of the acyl-enzyme intermediates (Liu, B., Schofield, CJ & Wilmouth, RC Structural analyzes on intermediates in serine protease catalysis. Journal of Biological Chemistry 281, 24024-24035 (2006); Ngo, PD, Mansoorabadi, SO & Frey, PA Serine Protease Catalysis: A Computational Study of Tetrahedral Intermediates and Inhibitory Adducts. J Phys Chem B 120, 7353-7359 (2016); Cleary, JA, Doherty, W., Evans, P. & Malthouse, JPG Quantifying tetrahedral adduct formation and stabilization in the cysteine ​​and the serine proteases. Bba-Proteins Proteom 1854, 1382-1391 (2015)). In other known cases, mutations in the enzyme active site have been shown to This may allow stabilization of the acyl-enzyme intermediate (Scaglione, JB et al. Biochemical and structural characterization of the tautomycetin thioesterase: analysis of a stereoselective polyketide hydrolase. Angew Chem Int Ed Engl 49, 5726-5730 (2010)). Although these approaches provided valuable insights, they required the synthesis of substrate analogues, leading to conjugates with non-native active sites or non-native substrates, which are drawbacks of these techniques. In the ubiquitin field, it was possible to replace the key cysteine ​​residue of E2 with a lysine and create an amide bond at the C-terminus of ubiquitin; this provided many insights into the protein ubiquitination pathway (Cappadocia, L. & Lima, CD Ubiquitin-like Protein Conjugation: Structures, Chemistry, and Mechanism. Chem Rev 118, 889-918 (2018); Plechanovova, A., Jaffray, EG, Tatham, MH, Naismith, JH & Hay, RT Structure of a RING E3 ligase and ubiquitin-loaded E2 primed for catalysis. Nature 489, 115-U135 (2012)). However, this conjugation usually occurs at high pH This is the reason why the substitutions obtained are not isosteric. This strategy is therefore not suitable for application to most enzyme classes that proceed via an acyl-enzyme intermediate, which represents a challenge in the art.

[0004] Acyl-enzyme intermediates are formed between sulfhydryl or hydroxyl side chains of cysteine ​​or serine residues in enzymes and carbonyl groups in substrates, and are ubiquitous in a variety of biological transformations, including those mediated by non-ribosomal peptide synthetases (NRPSs) and proteases. These important thioester and ester intermediates are unstable, usually with half-lives of minutes to hours, making their characterization difficult. This is a challenge in the art.

[0005] Polypeptides containing 2,3-diaminopropionic acid (DAP) have been disclosed in the art. Specifically, one or more polypeptides having DAP have been produced by solid-phase synthesis / conjugation techniques in the prior art (see Virdee, S., Macmillan, D., & Waksman, G. (2010) Chemistry & Biology vol 17 pages 274-284 "Semisynthetic Src SH2 domains demonstrate altered phosphopeptide specificity induced by incorporation of unnatural lysine derivatives."). However, the conventional solid There are many challenges in producing polypeptides containing DAPs by solid-phase synthesis / conjugation techniques. For example, it is extremely difficult or impossible to incorporate DAPs into biologically most important domains or motifs, such as the active site of an enzyme, using prior art techniques. This is because the active site is not available for chemical conjugation and / or because proteins made by solid-phase synthesis require chemical refolding to achieve the correct conformation. This is not only laborious, but also very unreliable, unpredictable, and very often ends in failure. These are challenges in the art. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Holliday, GL, Mitchell, JBO & Thornton, JM Understanding the Functional Roles of Amino Acid Residues in Enzyme Catalysis. Journal of molecular biology 390, 560-577 (2009) [Non-Patent Document 2] Yang, W. & Drueckhammer, DG Understanding the relative acyl-transfer reactivity of oxoesters and thioesters: computational analysis of transition state delocalization effects. Journal of the American Chemical Society 123, 11004-11009 (2001) [Non-Patent Document 3] Liu, B., Schofield, CJ & Wilmouth, RC Structural analyzes on intermediates in serine protease catalysis. Journal of Biological Chemistry 281, 24024-24035 (2006) [Non-Patent Document 4] Ngo, PD, Mansoorabadi, SO & Frey, PA Serine Protease Catalysis: A Computational Study of Tetrahedral Intermediates and Inhibitory Adducts. J Phys Chem B 120, 7353-7359 (2016) [Non-Patent Document 5] Cleary, JA, Doherty, W., Evans, P. & Malthouse, JPG Quantifying tetrahedral adduct formation and stabilization in the cysteine ​​and the serine proteases. Bba-Proteins Proteom 1854, 1382-1391 (2015) [Non-Patent Document 6] Scaglione, JB et al. Biochemical and structural characterization of the tautomycetin thioesterase: analysis of a stereoselective polyketide hydrolase. Angew Chem Int Ed Engl 49, 5726-5730 (2010) [Non-Patent Document 7] Cappadocia, L. & Lima, CD Ubiquitin-like Protein Conjugation: Structures, Chemistry, and Mechanism. Chem Rev 118, 889-918 (2018) [Non-Patent Document 8] Plechanovova, A., Jaffray, EG, Tatham, MH, Naismith, JH & Hay, RT Structure of a RING E3 ligase and ubiquitin-loaded E2 primed for catalysis. Nature 489, 115-U135 (2012) [Non-Patent Document 9] Virdee, S., Macmillan, D., & Waksman, G. (2010) Chemistry & Biology vol 17 pages 274-284 "Semisynthetic Src SH2 domains demonstrate altered phosphopeptide specificity induced by incorporation of unnatural lysine derivatives" Summary of the Invention [Problem to be solved by the invention]

[0007] An important technical advance provided by the present invention is a new method for producing polypeptides containing 2,3-diaminopropionic acid (DAP), which is particularly useful for incorporating DAP into locations that were difficult or impossible to engineer using conventional techniques, such as into enzyme active sites.

[0008] These methods are based on new unnatural amino acids that are amenable to incorporation into polypeptides via the naturally occurring translation machinery (e.g., the cell's own ribosomes). These methods also involve new tRNA synthetases that can charge an orthogonal tRNA with the new unnatural amino acid. Once incorporated into a polypeptide, the new unnatural amino acid can be easily deprotected to leave the DAP in the polypeptide backbone.

[0009] Thus, the present invention allows for the incorporation of DAP into a series of proteins and / or at a series of locations within proteins that are not currently possible using prior art techniques. [Means for solving the problem]

[0010] Thus, in one aspect, the present invention provides a compound according to formula (I) or (II):

[0011] [ka]

[0012] [In the formula, R1 is H, an amino acid residue or a peptide; R2 is H, C 1-6 Alkyl, C 1-6 Haloalkyl or C 5-20 aryl, q is 1, 2 or 3; R3 or R4 is H, halo, C 1-6 Alkyl, C 1-6 Haloalkyl, C 5-20 Aryl, C 3-20 Heteroaryl, OC 1-6 Alkyl, SC 1-6 Alkyl, NH(C 1-6 Alkyl) and N(C 1-6 alkyl)2; X is X1-Y, SS-R5, Se-Se-R5, O-NH-R5, S-NH-R5, Se-NH-R5, X2-Y1, X3-Y2, N3 or NH-S(O)2-Y3, and X1 is S, Se, O, NH or N(C 1-6 alkyl), X2 is S, Se or O; X3 is NH-C(O)-O; X4 is NH-C(O)-O, O, S or NH; R5 is H, halo, C 1-6 Alkyl, C 1-6 Haloalkyl, C 5-20 Aryl, C 3-20 Heteroaryl, OC 1-6 Alkyl, NH(C 1-6 alkyl), N(C 1-6 Alkyl)2, peptide, sugar, C 3-20 is selected from heterocyclyl and nucleic acid; Y is

[0013] [ka]

[0014] is a protecting group selected from R6 is H, C 1-6 Alkyl, C 1-6 Haloalkyl, CO2H, CO2R', SO2H, SO2R', C 5-20 Aryl, C 3-20 selected from heteroaryl, NHC(O)R′ and NHR′; R7 and R8 are H, OH, O(C 1-6 alkyl), O(C 5-20 aryl) and O(C 3-20 heteroaryl), or R7 and R8 are linked together to form a O—CH2—O group; Each R' is C 1-6 Alkyl, C 1-6 Haloalkyl and C 5-20 aryl; R9 is H, C 1-6 Alkyl, C 1-6 Haloalkyl, CO2H, CO2R', SO2H, SO2R' and C 5-20 aryl; R 10 , H, C 1-6 Alkyl and C 1-6 haloalkyl; R 11 , H, C 1-6 Alkyl and C 1-6 haloalkyl; X5 is S, O, NH, NC(O)-O-R', NS(O)2H, NS(O)2R 'or NR', Y1 is,

[0015] [ka]

[0016] is a protecting group selected from Y2 is,

[0017] [ka]

[0018] , t-Bu and CH2Ph; M + Li + , Na + , K + or N(R 13 )4 + and Z is Si or Ge; R 12 is C 1-6 Alkyl or C(O)-(C 5-20 aryl), R 13 , H, C 1-6 Alkyl, aryl or C 5-20 is aryl, Y3 is a protecting group

[0019] [ka]

[0020] is] or a salt, solvate, tautomer, isomer, or mixture thereof.

[0021] In one aspect, the invention provides a polypeptide comprising an unnatural amino acid as described above, wherein the unnatural amino acid is attached to the polypeptide via a peptide bond.

[0022] In one aspect, the present invention relates to a method for preparing a polypeptide comprising DAP, comprising a step of deprotecting said polypeptide. Advantageously, said deprotecting step is carried out at 365 nm, 35 mW cm -2 This includes 1 minute of exposure at 1000 Hz.

[0023] In one aspect, the present invention relates to a PylRS tRNA synthetase comprising the mutations Y271C, N311Q, Y349F and V366C. Preferably, said PylRS tRNA synthetase is Methanosarcina barkerii PylRS (MbPylRS) tRNA synthetase comprising said mutations.

[0024] In one aspect, the invention relates to a method of producing a polypeptide comprising 2,3-diaminopropionic acid (DAP), comprising genetically incorporating an unnatural amino acid as described above into a polypeptide, and optionally deprotecting the unnatural amino acid to 2,3-diaminopropionic acid (DAP).

[0025] Preferably, the production of the polypeptide comprises (i) providing a nucleic acid encoding a polypeptide, the nucleic acid comprising an orthogonal codon encoding an unnatural amino acid according to any one of claims 1 to 10; and (ii) translating the nucleic acid in the presence of an orthogonal tRNA synthetase / tRNA pair capable of recognizing the orthogonal codon to incorporate the unnatural amino acid into a peptide chain. Includes.

[0026] Preferably, the orthogonal codon comprises an amber codon (TAG), and the tRNA is MbtRNA. CUA wherein the tRNA synthetase comprises MbPylRS synthetase having the mutations Y271C, N311Q, Y349F and V366C.

[0027] Suitably, the unnatural amino acid is

[0028] [ka]

[0029] Includes.

[0030] In one aspect, the invention relates to the aforementioned polypeptide or the aforementioned method, wherein the polypeptide is an enzyme and the unnatural amino acid is incorporated at a position corresponding to an amino acid residue in the active site of the enzyme.

[0031] In one embodiment, the present invention relates to a peptide as described above, wherein said peptide is an enzyme, said enzyme being an EC 3.4 peptidase, an EC 3.4.22.44 peptidase or an EC 2.3.2.23 E2 ubiquitin-conjugating enzyme according to the International Nomenclature and Classification of Enzymes.

[0032] In one embodiment, the present invention relates to a polypeptide as described above, comprising 1 to 20 2,3-diaminopropionic acid (DAP) groups. Advantageously, said polypeptide comprises a single 2,3-diaminopropionic acid (DAP) group.

[0033] In one aspect, the invention relates to a polypeptide or a method as described above, wherein the unnatural amino acid is incorporated at a position corresponding to a cysteine, serine or threonine residue in a wild-type polypeptide, optionally at a position corresponding to a cysteine ​​or serine residue in the wild-type polypeptide.

[0034] In one aspect, the present invention relates to the use of the aforementioned unnatural amino acids in the production of polypeptides that include 2,3-diaminopropionic acid (DAP). Production includes carrying out the methods described above.

[0035] In one aspect, the present invention provides a method for capturing a substrate for an enzyme, comprising the steps of: a) providing an enzyme that contains at least one 2,3-diaminopropionic acid (DAP) group in its active site; b) contacting the enzyme with a candidate substrate for the enzyme; c) incubating to allow the DAP group to react with the candidate substrate; The method includes: Preferably, the substrate is a metabolite. Suitably, the enzyme is a peptidase, or a ubiquitin-conjugating enzyme, or a hydrolase, or an enzyme that forms a carbon-sulfur bond. [Brief description of the drawings]

[0036] Next, embodiments of the present invention will be further described with reference to the accompanying drawings. [Figure 1] Illustrates a general mechanism for enzymes with cysteine ​​and serine nucleophiles in the active site proceeding through an acyl-enzyme intermediate. a, b, The active site serine or cysteine ​​nucleophile reacts with a carbonyl group to form a tetrahedral intermediate (not shown) that collapses to the acyl-enzyme intermediate with loss of R1-XH (where X is typically NH, O, S) . Attack of the acyl-enzyme intermediate by a nucleophile R3 (typically hydroxyl, amine, or thiol) releases the bound substrate fragment and regenerates the enzyme. c, Substitution of cysteine ​​or serine with 2,3-diaminopropionic acid (DAP) can make the enzyme proceed to the first acyl-enzyme intermediate resistant to cleavage. [Diagram 2]Diagram of valinomycin synthetase and the proposed biosynthesis of valinomycin. The subunits Vlm1 and Vlm2 of valinomycin synthetase sequentially condense D-α-hydroxyisovaleric acid (D-α-hiv), D-valine (D-val), L-lac acid (L-lac) and L-valine (L-val) to form tetradepsipeptidyl (D-hiv-D-val-L-lac-L-val) intermediates. D-α-hiv and L-lac are generated by the selection and ketoreduction of their precursor keto acids by specialized modules 1 and 3, which contain ketoreductase (KR) domains. The tetradepsipeptidyl intermediates are oligomerized to octadepsipeptidyl intermediates and then dodecadepsipeptidyl intermediates, which are cyclized by the terminal thioesterase (TE) domain to produce valinomycin. A: adenylation domain, PCP: peptidyl carrier protein domain; C: condensation domain. For the synthesis cycle of a canonical NRPS, see Supplementary Figure 1. [Figure 3-1] ~ [Figure 3-3]Figure 1. Genetic instructions for DAP incorporation into recombinant proteins. a, Structures of DAP and protected versions discussed herein. 1: 2,3-diaminopropionic acid (DAP). 2: (S)-3-(((allyloxy)carbonyl)amino)-2-aminopropanoic acid. 3: (S)-2-amino-3-((2-nitrobenzyl)amino)propanoic acid. 4: (2S)-2-amino-3-((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl)amino)propanoic acid. 5: (2S)-2-amino-3-(((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy)carbonyl)amino)propanoic acid. 6: (2S)-2-amino-3-(((2-((1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl)thio)ethoxy)carbonyl)amino)propanoic acid. b-f, Determination of intracellular concentrations of compounds 2-6 by LC-MS assay performed on extracts. Dark blue traces represent 100 μM standards for each compound. Light blue traces represent 10 μM standards for each compound. Red traces result from cells grown in the absence of compound. Brown traces result from cells grown in the absence of compound but spiked with up to 10 μM of compound. Green traces result from cells grown in the presence of 1 mM compound. g, Phenotyping of the DAPRS / tRNACUA pair. Cells containing the DAPRS / tRNACUA pair and cat(112TAG) were plated in the presence or absence of 6 at the indicated concentrations of chloramphenicol. h, Expression of sfGFP containing either 6 or BocK at position 150. After expression and purification, equal volumes of protein solutions were loaded onto SDS-PAGE gels and Coomassie stained (upper gel) or analyzed by Western blot using α-His antibody (lower gel). i, Encoded 6 was deprotected with UV light to an intermediate (red) that spontaneously fragmented to expose the DAP. j, Deprotection of 6 in sfGFP was followed by ESI-MS analysis. Green trace: purified sfGFP containing 6 at position 150: expected mass: 28096.27 Da; observed mass: 28097.21 Da.Red trace: sfGFP containing 6 after irradiation to convert 6 to an intermediate: expected mass: 27902.22 Da; observed mass: 27904.14 Da. Blue trace: sfGFP containing 6 after irradiation (to convert 6 to an intermediate) and further incubation to convert the intermediate to DAP(1): expected mass: 27798.23 Da; observed mass: 27800.88 Da. Each trace also shows the mass of the protein adduct resulting from the spontaneous loss of the N-terminal methionine. [Figure 4a] ~ [Figure 4b] Figure 4 shows stable trapping of acyl-enzyme intermediates by TEV(C151DAP). Figure 4a shows the indicated variants of TEV protease were incubated with Ub-tev-His. Use of TEV(wt) results in cleavage of the TEV cleavage sequence. Use of TEV(C151A) results in minimal cleavage. The presence of DAP at the active site of TEV results in the presence of a new band in the Coomassie gel, which represents the isopeptide-linked TEV(C151DAP)-Ub complex (left). αUb and αStrep Western blots of the reaction confirm the identity of the complex (TEV construct contains a Strep tag). Figure 4b shows tandem mass spectrometry of the isopeptide-linked TEV(C151DAP)-Ub complex. Tandem mass spectrometry clearly identifies the DAP modification at the desired site and the expected tev-Gly-Gly modification on the residues, consistent with Ub capture on the DAP. [Diagram 5]Small molecule products made by Vlm TE from tetradepsipeptidyl-SNAC clarify the oligomerization pathway. Extracted ion chromatograms (EICs) from HR-LC-ESI-MS of the reaction of tetradepsipeptidyl-SNAC 7 (1.7 mM) with Vlm TE (6.5 μM). a, TEwt produces valinomycin as the major product. The presence of octadepsipeptidyl-SNAC 11, dodecadepsipeptidyl-SNAC 15, and 16-mer depsipeptidyl-SNAC 19 confirms the oligomerization scenario in Supplementary Fig. 10b. b, TEDAP produces small amounts of octadepsipeptidyl-SNAC 11. c, Control reaction without enzyme shows small amounts of tetradepsipeptide 9 likely derived from uncatalyzed hydrolysis of the thioester in solution. For accurate mass analysis and deviation from calculated m / z for each compound, see Supplementary Table 2. [Figure 6-1] ~ [Figure 6-2]Crystal structures of complexes of TEDAP are shown. a, Representative electron density of TEwt (2mFo-DFc map contoured at 1.0σ). b, Structure of TEwt. The lid (grey) is almost perfectly ordered but has a higher B-factor than the protein core. c, Deconvoluted mass spectrum of TEDAP. d, Deconvoluted mass spectrum of TEDAP incubated with deoxy-tetradepsipeptidyl-SNAC 8. e, Deconvoluted mass spectrum of TEDAP incubated with valinomycin. f and g, Unbiased mFo-DFc electron density maps (green mesh, 2.5σ) for the depsipeptide residues of tetradepsipeptidyl-TEDAP (f) and dodecadepsipeptidyl-TEDAP (g). An amide bond links the diaminopropionic acid (DAP, brown sticks) and the depsipeptide residues (cyan sticks). h and i, Active sites of tetradepsipeptidyl-TEDAP (h) and dodecadepsipeptidyl-TEDAP (i) complexes. The carbonyl oxygen of the amide formed by DAP and val4 (h) or val12 (i) is positioned close to the oxyanion hole formed by the main chain amines of A2399 and L2464. The catalytic triad H2625 and D2490 are shown in sticks. j, The lid of tetradepsipeptidyl-TEDAP is in a similar position to that found in TEwt, while k, all crystallographically independent molecules of dodecadepsipeptidyl-TEDAP (from P1 and H3 space group structures) are in a series of similar conformations that are distinct from those found in TEwt. l, Illustrated substantial conformational changes in lid helices Lα1–Lα4 between the structures of tetradepsipeptidyl-TEDAP and dodecadepsipeptidyl-TEDAP. For clarity, mobile helices are shown in increasing wavelength colors. [Figure 7]Modeled and putative pathways of PCP and TE domain interactions. a, Superposition of the structure of dodecadepsipeptidyl-TEDAP with the EntF PCP-TE didomain (Liu, Y., Zheng, T. & Bruner, SD Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488 (2011)) shows the path of the PPE moiety into the active site. b, The lid sterically prevents the dodecadepsipeptide from linear extension, and this steric block and the highly hydrophobic nonspecific interactions of the lid with the dodecadepsipeptide favor curling back. c, Hypothetical pathways of oligomerization and cyclization starting from octadepsipeptidyl-TE. i, the position of Lα1 in the observed apo / tetradepsipeptide conformation promotes an extended peptide conformation; ii, tetradepsipeptidyl-PCP accepts the octadepsipeptide at its terminal hydroxyl, presumably using a dodecadepsipeptide-like lid conformation that could accommodate the ∼30 Å tetradepsipeptidyl-PPE bound to the PCP domain and guide it toward the active site; iii, the PCP domain presents a thioester to return to Ser2463; iv, finally, the lid conformation observed in the dodecadepsipeptide-TEDAP structure could serve to curl the dodecadepsipeptide back toward Ser2463 for cyclization. [Figure 8-1] ~ [Figure 8-2] Supplementary Figure 1. [Figure 9-1] ~ [Figure 9-2] Supplementary Figure 2. [Figure 10] Supplementary Figure 3. [Figure 11-1]~ [Figure 11-2] Supplementary Figure 4. [Figure 12] Supplementary Figure 5. [Figure 13-1] ~ [Figure 13-2] Supplementary Figure 6. [Figure 14-1] ~ [Figure 14-2] Supplementary Figure 7. [Figure 15] Supplementary Figure 8. [Figure 16] Supplementary Figure 9. [Figure 17] Supplementary Figure 10. [Figure 18] Supplementary Figure 11. [Figure 19] Supplementary Figure 12. [Figure 20] Supplementary Figure 13. [Figure 21] Supplementary Figure 14. [Figure 22] Supplementary Figure 15. [Figure 23-1] ~ [Figure 23-2] Figure 1 shows photographs. Ubiquitination studies with UBE2L3(C86DAP). A: Coomassie staining of ubiquitination reactions with either UBE2L3(wt), UBE2L3(C86A) or UBE2L3(C86DAP) in the presence and absence of Ub and β-mercaptoethanol. αHA (B) and αUBE2L3(111-125) (C) Western blots of ubiquitination reactions with either UBE2L3(wt), UBE2L3(C86A) or UBE2L3(C86DAP) in the presence or absence of Ub and β-mercaptoethanol. Ub-UBE2L3: thioester-linked (UBE2L3[wt]) or isopeptide-linked (UBE2L3[C86DAP]) E2-Ub complexes. The complex formed between UBE2L3(C86DAP) and Ub was not sensitive to the presence of β-mercaptoethanol, which differs from the complex between UBE2L3(wt) and Ub, which was reduced in the presence of β-mercaptoethanol. [Figure 24] FIG. [Diagram 25] FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0037] Here, we disclose how genetically encoded 2,3-diaminopropionic acid enables structural insight into enzymatic reactions. In particular, we illustrate this approach via an acyl-thioesterase intermediate in the biosynthesis of valinomycin.

[0038] The present invention allows for a genetically directed strategy for efficient incorporation of 2,3-diaminopropionic acid (DAP) into recombinant proteins produced in (for example) E. coli. We teach how to replace catalytic residues, e.g., cysteine ​​or serine residues, with DAP. This allows for efficient capture of acyl-enzyme complexes linked via stable amide bonds.

[0039] For example, the invention is demonstrated by elucidating a biosynthetic pathway in which the thioesterase domain of valinomycin synthetase (Vlm TE) promotes both the sequential trimerization of linear tetradepsipeptides to dodecadepsipeptides and the sequential cyclization of dodecadepsipeptides to valinomycin. By capturing the first and last acyl-TE intermediates in the catalytic cycle of Vlm TE as a DAP conjugate, the use of the invention provides structural insight into how conformational changes in the TE domain of NRPSs control the switch from oligomerization to cyclization of linear substrates. Such strategies enabled by the invention find use in facilitating the characterization of diverse acyl-enzyme intermediates. Additionally, the invention finds use in enabling the capture and identification of native substrates for enzymes of unknown function.

[0040] In one aspect, the present invention relates to a homologous recombinant polypeptide as described above. Advantageously, said polypeptide is produced by the method as described above.

[0041] In one aspect, the present invention relates to polypeptides produced according to the methods described herein. In addition to being the products of these new methods, such polypeptides have the technical feature that they contain the unnatural amino acids described above or that they contain DAPs.

[0042] Mutation has its ordinary meaning in the art and may refer to the substitution, truncation or deletion of the referred residue, motif or domain. Mutation may be made at the polypeptide level, e.g., by synthesis of a polypeptide having a mutant sequence, or at the nucleotide level, e.g., by creating a nucleic acid that encodes the mutant sequence, which can then be translated to produce the mutant polypeptide. If no replacement amino acid is specified, then preferably randomization of said site is used. Alanine (A) may be used as the default mutation. Preferably, the mutation used at a particular site is as described herein.

[0043] A fragment is preferably at least 10 amino acids in length, preferably at least 25 amino acids, preferably at least 50 amino acids, preferably at least 100 amino acids, preferably at least 200 amino acids, preferably at least 250 amino acids, preferably at least 300 amino acids, preferably at least 313 amino acids, or preferably a majority of the polypeptide of interest.

[0044] The methods of the present invention can be practiced in vivo or in vitro.

[0045] In one embodiment, preferably the method of the invention is not applied to a human or animal body. Preferably the method of the invention is an in vitro method. Preferably the method does not require the presence of a human or animal body. Preferably the method is not a method of diagnosis or surgery or therapy of a human or animal body.

[0046] The term "comprises" is used in accordance with its common usage in the art. It should be understood that the term has the general meaning, i.e. that a stated feature or group of features is included, but the term does not exclude that any other stated feature or group of features is also present.

[0047] DAP integration We hypothesized that selective replacement of catalytic cysteine ​​or serine residues with amino acids, replacing the sulfhydryl or hydroxyl groups with amino groups, would allow trapping of acyl-enzyme intermediates linked via amide bonds (Figure 1c). The conjugate acid of the lysine side chain amine has a pKa of 10.5, whereas the conjugate acid of the beta amino group of DAP has a pKa of 9.4 (Hay, RW & Morris, PJ Interaction of Dl-2,3-Diaminopropionic Acid and Its Methyl Ester with Metal Ions .1. Formation Constants. J Chem Soc A, 3562-& (1971)). Furthermore, when DAP is incorporated into a peptide, the pKa is substantially lowered by the electron-withdrawing effect of the backbone amide: a single peptide bond of DAP to a carboxylate lowers the pKa of the conjugate acid of the beta-amino group to 7.5 (Lan, Y. et al. Incorporation of 2,3-Diaminopropionic Acid into Linear Cationic Amphipathic Peptides Produces pH-Sensitive Vectors. Chembiochem 11, 1266-1272 (2010)), allowing the formation of longer peptides. A pKa of 6.3 has been reported for D peptides (Lan, Y. et al. Incorporation of 2,3-Diaminopropionic Acid into Linear Cationic Amphipathic Peptides Produces pH-Sensitive Vectors. Chembiochem 11, 1266-1272 (2010)). If AP was used to replace Cys or Ser in the active site of the enzyme, we expected that a significant portion of the DAP side chains would exist as neutral amines at physiological pH. These amines would be available to act as nucleophiles and form amide bonds with the enzyme's substrate. Since the half-life of amines in aqueous solution is approximately 500 years (Radzicka, A. & Wolfenden, R. Rates of uncatalyzed peptide bond hydrolysis in neutral solution and the transition state affinities of proteases. Journal of the American Chemical Society 118, 6105-6109 (1996)), we expected that the amide analogs of the unstable thioester and ester intermediates would be substantially stabilized, such that subsequent reactions with nucleophiles or solvents would not proceed or would be greatly attenuated (Figure 1c).

[0048] Nonribosomal peptide synthetases (NRPS) and polyketide synthases (PKS), which are large enzymes that produce secondary metabolites, play important roles in their synthesis. These molecular machines generate highly complex acyl-enzyme intermediates in cyclins. Templated biosynthetic pathways are used to assemble small acyl molecules into a wide range of biologically active natural products, including clinical anticancer, antibiotic, antifungal and immunosuppressant drugs (Supplementary Figure 1). Prior attempts to unravel their detailed molecular functions have been hampered by the challenge of characterizing their multiple acyl-enzyme intermediates at high resolution. This challenge is exemplified by thioesterase (TE) domains from the NRPS pathway, which oligomerize or cyclize linear peptidyl or depsipeptidyl substrates. These TE domains are involved in the synthesis of the antibiotic gramicidin S (Hoyer, KM, Mahlert, C. & Schneider, J. Biol. 2014;2013;2011). Marahiel, MA The iterative gramicidin s thioesterase catalyzes peptide ligation and cyclization. Chemistry & biology 14, 13-22 (2007)), the emetic toxin cereulide (Alonzo, DA, Magarvey, NA & Schmeing, TM Characterization of cereulide synthetase, a toxin-producing macromolecular machine. PloS one 10, e0128569 (2015), Magarvey, NA, Ehling-Schulz, M. & Walsh, CT Characterization of the cereulide NRPS alpha-hydroxy acid specifying modules: activation of alpha-keto acids and chiral reduction on the assembly line. Journal of the American Chemical Society 128, 10698-10699 (2006)), and the siderophores enterobactin and bacillibactin (Shaw-Reid, CA et al. Assembly line enzymology by multimodular nonribosomal peptide synthetases: the thioesterase domain of E. coli EntF catalyzes both elongation and cyclolactonization. Chemistry & biology 6, 385-400 (1999), May, JJ, Wendrich, TM & Marahiel, MA The dhb operon of Bacillus subtilis encodes the biosynthetic template for the catecholic siderophore 2,3-dihydroxybenzoate-glycine-threonine trimeric ester bacillibactin. J Biol Chem 276, 7209-7217 (2001)), the anticancer compound conglobatin (Zhou, Y. et al. Iterative Mechanism of Macrodiolide Formation in the Anticancer Compound Conglobatin. Chemistry & biology 22, 745-754 (2015)), and the DNA bisintercalator thiocoraline (Robbel, L., Hoyer, KM & Marahiel, MA TioS T-TE--a prototypical thioesterase responsible for cyclodimerization of the quinoline- and quinoxaline-type class of chromodepsipeptides. FEBS J 276, 1641-1653 (2009)) and has antibacterial, antitumor and cytotoxic properties. Characterization of the cereulide NRPS alpha-hydroxy acid specifying modules: activation of alpha-keto acids and chiral reduction on the potassium ionophore depsipeptide valinomycin (Magarvey, NA, Ehling-Schulz, M. & Walsh, CT assembly line. Journal of the American Chemical Society 128, 10698-10699 (2006), Jaitzig, J., Li, J., Sussmuth, RD & Neubauer, P. Reconstituted biosynthesis of the nonribosomal macrolactone antibiotic valinomycin in Escherichia coli. ACS Synth Biol 3, 432-438 (2014)). They were first It must oligomerize the tidyl intermediates up to, but not beyond, the copy number found in the biologically active compound, and then catalyze the final release and cyclization of the finished product. Moreover, the required oligomerization and cyclization must be fast enough so that spontaneous hydrolysis does not substantially result in the formation of free linear peptides, which are useless by-products because they cannot be reincorporated into the synthetic cycle.

[0049] High-resolution structures of acyl-TE intermediates would represent a substantial advance, providing mechanistic insight into how TEs control the fate of their substrates. A handful of high-resolution acyl-TE structures have been obtained, most notably for the TE that forms the polyketide pikromycin and for non-native substrate analogues (Akey, DL et al. Structural basis for macrolactonization by the pikromycin thioesterase. Nat Chem Biol 2, 537-542 (2006)). These helped identify the putative oxyanion hole and provided a basis for the T The interaction of the "lid" element of the E domain with the substrate was demonstrated. Structural studies of the TE domain demonstrated the association of low Kd values ​​with small molecule substrates (Samel, SA, Wagner, B., Marahiel, MA & Essen, LO The thioesterase domain of the fengycin biosynthesis cluster: a structural base for the macrocyclization of a non-ribosomal lipopeptide. Journal of molecular biology. cular biology 359, 876-889 (2006)), and the structure of the peptide chain upon binding to the TE domain. The thioesterase domain of the fengycin synthetase is characterized by multiple conformations (Tseng, CC et al. Characterization of the surfactin synthetase C-terminal thioesterase domain as a cyclic depsipeptide synthase. Biochemistry 41, 13350-13359 (2002); Bruner, SD et al. Structural basis for the cyclization of the lipopeptide antibiotic surfactin by the thioesterase domain SrfTE. Structure 10, 301-310 (2002)) and, in particular, a fast hydrolysis rate of the acyl-TE intermediate compared to the crystallographic time scale (Samel, SA, Wagner, B., Marahiel, MA & Essen, LO The thioesterase domain of the fengycin biosynthesis cluster: a structural base for the macrocyclization of a non-ribosomal lipopeptide. Journal of molecular biology 359, 876-889 (2006); Bruner, SD et al. Structural basis for the cyclization of the lipopeptide antibiotic surfactin by the thioesterase domain SrfTE. Structure 10, 301-310 (2002)). These are the challenges associated with the prior art approaches. The ability to access stable acyl-TE intermediates provided by the present invention is a significant advantage that allows one of skill in the art to characterize mechanisms of TE domain selectivity in the biosynthesis of nonribosomal peptides as well as the biosynthesis of polyketides and fatty acids.

[0050] As described in more detail below, the inventors have developed aminoacyl-tRNA synthetases / tRNAs incorporating amino acids that can be post-translationally converted to DAP under mild conditions. CUA We have evolved a pair of cereulide NRPS alpha-hydroxy acid specifying modules: activation of alpha-keto acid and amino acids to a tetradepsipeptide intermediate. We demonstrate the use of this pair for site-specific incorporation of DAP into an integration protein produced in E. coli. We demonstrate efficient capture of an acyl-enzyme intermediate with a cysteine ​​protease and an NRPS TE domain. Valinomycin synthetase (Vlm), a two-protein, four-module NRPS, has been proposed to alternatively link hydroxy acids (derived from in situ reduction of an alpha-keto acid) and amino acids to a tetradepsipeptide intermediate that the TE domain (Vlm TE) progressively trimers to a dodecadepsipeptide and then valinomycin (Magarvey, NA, Ehling-Schulz, M. & Walsh, CT Characterization of the cereulide NRPS alpha-hydroxy acid specifying modules: activation of alpha-keto acid and amino acids to a tetradepsipeptide intermediate. acids and chiral reduction on the assembly line. Journal of the American Chemical Society 128, 10698-10699 (2006), Jaitzig, J., Li, J., Sussmuth, RD & Neubauer, P. Reconstituted biosynthesis of the nonribosomal macrolactone antibiotic valinomycin in Escherichia coli. ACS Synth Biol 3, 432-438 (2014)) (Figure 2). We demonstrate the invention used to elucidate a biosynthetic pathway for converting tetradepsipeptides to valinomycin. By replacing the catalytic serine in the Vlm TE with DAP, we generate a stable deoxy-tetradepsipeptidyl-N-TE. DAPand dodecadepsipeptidyl-TE DAP We demonstrate how Vlm TEs access the conjugates that mediate the acetylation of Vlm TEs. Structural characterization of these conjugates provides insight into the first and last acyl-TE intermediates in the catalytic cycle of Vlm TEs. Thus, the present invention finds use in investigating / revealing how the fate of a substrate can be determined by conformational changes in the TE domain of NRPSs that oligomerize and cyclize linear precursors.

[0051] Suitably, the polypeptide comprises a single DAP group as defined above and / or a residue of a non-natural amino acid. This has the advantage of maintaining the specificity for any further chemical modifications that may be directed to said DAP group / non-natural amino acid and / or the specificity of capture in the case where the DAP is present at the active site of the enzyme. For example, when only a single DAP group / non-natural amino acid as defined above is present in the polypeptide of interest, potential problems of partial modification / or partial deprotection or different reaction microenvironments between alternative DAP groups in the same polypeptide (which may result in unequal reactivity between different DAP groups at different locations in the polypeptide) are advantageously avoided.

[0052] Preferably, the polypeptide comprises 2 DAP groups; preferably, the polypeptide comprises 3 DAP groups; preferably, the polypeptide comprises 4 DAP groups; preferably, the polypeptide comprises 5 DAP groups; preferably, the polypeptide comprises 10 or even more DAP groups, for example 15-20 DAP groups. Most preferably, the polypeptide comprises 1-5 DAP groups. Most preferably, the polypeptide comprises 1 DAP group.

[0053] In principle, multiple unnatural amino acids as described above (either multiple copies of the same unnatural amino acid or one or more copies of each of two or more different unnatural amino acids) can be incorporated by the same or different orthogonal codon / orthogonal tRNA pairs. Suitably, multiple unnatural amino acids are incorporated (together with the orthogonal tRNA synthetases described above) by the insertion / translation of multiple amber codons.

[0054] New chemical substances (NCE) A novel chemical entity (NCE) is a chemical entity that is not classified into two groups. As described in the specification, the formula (I) or (II)

[0055] [ka]

[0056] [In the formula, R1 is H, an amino acid residue or a peptide; R2 is H, C 1-6 Alkyl, C 1-6 Haloalkyl or C 5-20 aryl, q is 1, 2 or 3; R3 or R4 is H, halo, C 1-6 Alkyl, C 1-6 Haloalkyl, C 5-20 Aryl, C 3-20 Heteroaryl, OC 1-6 Alkyl, SC 1-6 Alkyl, NH(C 1-6 Alkyl) and N(C 1-6 alkyl)2; X is X1-Y, SS-R5, Se-Se-R5, O-NH-R5, S-NH-R5, Se-NH-R5, X2-Y1, X3-Y2, N3 or NH-S(O)2-Y3, and X1 is S, Se, O, NH or N(C 1-6 alkyl), X2 is S, Se or O; X3 is NH-C(O)-O; X4 is NH-C(O)-O, O, S or NH; R5 is H, halo, C 1-6 Alkyl, C 1-6 Haloalkyl, C 5-20 Aryl, C 3-20 Heteroaryl, OC 1-6 Alkyl, NH(C 1-6 alkyl), N(C 1-6 Alkyl)2, peptide, sugar, C 3-20 is selected from heterocyclyl and nucleic acid; Y is

[0057] [ka]

[0058] is a protecting group selected from R6 is H, C 1-6 Alkyl, C 1-6 Haloalkyl, CO2H, CO2R', SO2H, SO2R', C 5-20 Aryl, C 3-20 selected from heteroaryl, NHC(O)R′ and NHR′; R7 and R8 are H, OH, O(C 1-6 alkyl), O(C 5-20 aryl) and O(C 3-20 heteroaryl), or R7 and R8 are linked together to form a O—CH2—O group; Each R' is C 1-6 Alkyl, C 1-6 Haloalkyl and C 5-20 aryl; R9 is H, C 1-6 Alkyl, C 1-6 Haloalkyl, CO2H, CO2R', SO2H, SO2R' and C 5-20 aryl; R 10 , H, C 1-6 Alkyl and C 1-6 haloalkyl; R11 , H, C 1-6 Alkyl and C 1-6 haloalkyl; X5 is S, O, NH, NC(O)-O-R', NS(O)2H, NS(O)2R' or NR'; Y1 is,

[0059] [ka]

[0060] is a protecting group selected from Y2 is,

[0061] [ka]

[0062] , t-Bu and CH2Ph; M + Li + , Na + , K + or N(R 13 )4 + and Z is Si or Ge; R 12 is C 1-6 Alkyl or C(O)-(C 5-20 aryl), R 13 , H, C 1-6 Alkyl, aryl or C 5-20 is aryl, Y3 is a protecting group

[0063] [ka]

[0064] is] or a salt, solvate, tautomer, isomer, or mixture thereof.

[0065] The term "or salts, solvates, tautomers, isomers, or mixtures thereof" is meant to include salts, solvates, tautomers, isomeric forms of the depicted structure, and mixtures thereof means that mixtures of these forms exist, e.g., compounds of the invention may include salts of tautomers.

[0066] "Pharmaceutically acceptable" refers to materials that, within the scope of sound medical judgment, are suitable for use in contact with the tissues of a subject, are not excessively toxic, irritating, allergic, or the like, and are commensurate with a reasonable benefit-to-risk ratio and are effective for their intended use.

[0067] A "pharmaceutical composition" refers to a combination of one or more drug substances and one or more excipients.

[0068] As used herein, a "solvate" refers to a complex of variable stoichiometry formed by a solute (e.g., a compound of Formula (I)-(II) or any other compound herein or a salt thereof) and a solvent. Pharmaceutically acceptable solvates are This may be formed in the case of crystalline compounds in which solvent molecules are incorporated into the crystal lattice during crystallization. The incorporated solvent molecules may be water molecules or non-water molecules, such as, but not limited to, ethanol, isopropanol, dimethylsulfoxide, acetic acid, ethanolamine, and ethyl acetate molecules.

[0069] "Independently selected" means, for example, that "R3 or R4 is independently selected from H, halo, C 1-6 When used in the context of "independently selected from alkyl, alkyl, ...", it means that each instance of a functional group, e.g., R3, is selected from the listed options independently of every other instance of R3 or R4 in the compound. Thus, for example, the first instance of R3 in a compound may be selected as H, the next instance of R3 in a compound may be selected as methyl, and the first instance of R4 in a compound may be selected as ethyl.

[0070] In this specification, C 1-6 Alkyl is a straight-chain or branched saturated hydrocarbon group generally having 1 to 6 carbon atoms, more preferably C 1-5 Alkyl, more preferably C 1-4 Alkyl, more preferably C 1-3 Examples of alkyl groups include methyl, ethyl, n-propyl, i-propyl, n-butyl, s-butyl, i-butyl, t-butyl, pent-1-yl, pent-2-yl, pent-3-yl, 3-methylbut-1 ... Examples of the alkyl group include butylbut-2-yl, 2-methylbut-2-yl, 2,2,2-trimethyleth-1-yl, n-hexyl, and n-heptyl.

[0071] As used herein, "amino acid residue" refers to an amino acid in which either the amine or carboxylic acid terminus of the amino acid has been replaced with a peptide bond.

[0072] C 5-20 Aryl refers to fully unsaturated monocyclic, bicyclic, and polycyclic aromatic hydrocarbons having at least one aromatic ring and a specified number of carbon atoms constituting the ring members (e.g., C 6-14 Aryl refers to an aryl group having 6 to 14 carbon atoms as ring members. An aryl group can be attached to a parent group or substrate at any ring atom, and can contain one or more non-hydrogen substituents, provided that such attachment or substitution does not violate valence requirements. Examples of aryl groups include phenyl, biphenyl, cyclobutabenzenyl, naphthalenyl, benzocycloheptenyl, azulenyl, biphenylenyl, anthracenyl, phenanthrenyl, naphthacenyl, pyrenyl, groups derived from cycloheptatriene cations, and the like. Examples of aryl groups that contain fused rings, at least one of which is aromatic, include, but are not limited to, indanyl, indenyl, isoindenyl, tetralinyl, acenaphthenyl, fluorenyl, phenalenyl, acephenanthrenyl, and aceantrenyl.

[0073] "Halo" or "halogen" refers to -F, -CI, -Br or -I.

[0074] "C 1-6 "Haloalkyl" refers to a C alkyl group in which one or more hydrogen atoms are replaced by halo atoms. 1-6 It refers to a group derived from an alkyl group. 1-6 Haloalkyl is CH2F, CHF2, CF3, CH2Cl, CHCl2 or CCl3.

[0075] "C 3-20 "Heteroaryl" refers to unsaturated monocyclic, bicyclic and polycyclic aromatic groups containing 3 to 20 ring atoms, of which 1 to 10 are ring heteroatoms, whether carbon or heteroatoms. Preferably, each ring has 3 to 7 ring atoms and 1 to 4 heteroatoms. Preferably, each ring heteroatom is independently selected from nitrogen, oxygen and sulfur. Bicyclic and polycyclic may include any bicyclic or polycyclic group in which any of the monocyclic heterocycles listed above are fused to a benzene ring. A heteroaryl group may be attached to a parent group or substrate at any ring atom, and may contain one or more non-hydrogen substituents, provided such attachment or substitution does not violate valence requirements or result in a chemically unstable compound.

[0076] Examples of monocyclic heteroaryl groups include, but are not limited to, those derived from: N1: pyrrole, pyridine; O1: Franc; S1: thiophene; N1O1: oxazoles, isoxazoles, isoxazines; N2O1: oxadiazole (e.g., 1-oxa-2,3-diazolyl, 1-oxa-2,4-diazolyl, 1-oxa-2,5-diazolyl, 1-oxa-3,4-diazolyl); N3O1: oxatriazole; N1S1: thiazoles, isothiazoles; N2: imidazole, pyrazole, pyridazine, pyrimidine (e.g., cytosine, thymine, uracil), pyrazine; N3: Triazoles, triazines; and N4: tetrazole.

[0077] Examples of heteroaryls containing fused rings include, but are not limited to, those derived from: O1: benzofuran, isobenzofuran, chromene, isochromene, chroman, isochroman, dibenzofuran, xanthene; N1: indole, isoindole, indolizine, isoindoline, quinoline, isoquinoline, quinolizine, carbazole, acridine, phenanthridine; S1: benzothiofurans, dibenzothiophenes, and thioxanthenes; N1O1: benzoxazole, benzisoxazole, benzoxazine, phenoxazine; N1S1: benzothiazoles, phenothiazines; O1S1: phenoxathiin; N2: benzimidal, indazole, benzodiazine, pyridopyridine, quinoxaline, quinazoline, cinnoline, phthalazine, naphthyridine, benzodiazepine, carboline, perimidine, pyridoindole, phenazine, phenanthroline, phenazine; O2: benzodioxole, benzodioxane, oxanthrene; S2: Thianthrene; N2O1: benzofurazan; N2S1: bentothiadiazole; N3: benzotriazole; N4: Purines (e.g., adenine, guanine), pteridines.

[0078] "C 3-20 "Heterocyclyl" refers to saturated or partially unsaturated monocyclic, bicyclic and polycyclic groups having ring atoms, whether carbon or heteroatoms, composed of 3 to 20 ring atoms, of which 1 to 10 are ring heteroatoms. Preferably, each ring has 3 to 7 ring atoms and 1 to 4 heteroatoms (e.g., preferably, C 3-5Heterocyclyl refers to a heterocyclyl group having 3 to 5 ring atoms and 1 to 4 heteroatoms as ring members. The ring heteroatoms are independently selected from nitrogen, oxygen, and sulfur.

[0079] Like bicyclic cycloalkyl groups, bicyclic heterocyclyl groups can include isolated rings, spiro rings, fused rings, and bridged rings. A heterocyclyl group can be attached to a parent group or substrate at any ring atom and can contain one or more non-hydrogen substituents, unless such attachment or substitution violates valence requirements or results in a chemically unstable compound.

[0080] Examples of monocyclic heterocyclyl groups include, but are not limited to, those derived from: N1: aziridine, azetidine, pyrrolidine, pyrroline, 2H-pyrrole or 3H-pyrrole, piperidine, dihydropyridine, tetrahydropyridine, azepine; O1: oxirane, oxetane, tetrahydrofuran, dihydrofuran, tetrahydropyran, dihydropyran, pyran, oxepine; S1: thiiranes, thietanes, tetrahydrothiophenes, tetrahydrothiopyrans, thiepanes; O2: dioxoiane, dioxane and dioxepane; O3: Trioxane; N2: Imidazolidine, pyrazolidine, imidazoline, pyrazoline , piperazine; N1O1: Tetrahydrooxazole, dihydrooxazole, tetrahydroisoxazole, dihydroisoxazole, morpholine, tetrahydrooxazine, dihydrooxazine, oxazine; N1S1: thiazolines, thiazolidines, thiomorpholines; N2O1: oxadiazine; O1S1: Oxathioles and oxathianes (thioxanes); and N1O1S1: Oxathiazine.

[0081] Examples of substituted monocyclic heterocyclyl groups include those derived from cyclic forms of sugars, such as furanoses, e.g., arabinofuranose, lyxofuranose, ribofuranose and xylofuranose, and pyranoses, e.g., allopyranose, altropyranose, glucopyranose, mannopyranose, globyranose, idopyranose, galactopyranose and talopyranose.

[0082] As used herein, the term "peptide" refers to a linear molecule comprising multiple amino acid residues joined together by peptide bonds.

[0083] Protecting groups are groups that are introduced into a molecule by chemical modification of a functional group to temporarily mask the characteristic chemical properties of the functional group and prevent it from interfering with another reaction. Protecting groups that can be removed by photolysis (e.g., photolytically removable / removable / deprotected) are groups that can be removed by radiation energy, e.g., light. Protecting groups are described in Wuts, PGM and Greene, TW, Protective Groups in Organic Synthesis, 4 th Edition, Wiley-Interscience, 2007 and P. Kocienski, Protective Groups, 3rd Edition (2005).

[0084] A "sugar" substituent refers to a monosaccharide or polysaccharide in which an H from a hydroxyl group of the sugar is replaced with a bond that connects the sugar substituent to the remainder of the compound of formula (I) or (II). Preferably, the sugar is a monosaccharide. Preferably, the sugar is glucose, mannose or galactose. Preferably, the protecting group is photolabile and therefore does not show a photosensitivity at wavelengths λ of greater than 300 nm. max These are wavelengths that are not harmful to biological systems.

[0085] In one embodiment, suitably the unnatural amino acid of formula (I) or (II) is

[0086] [ka]

[0087] or a salt, solvate, tautomer, isomer, or mixture thereof. Thus, in this embodiment, the unnatural amino acid of Formula (I) or (II) is in the D form.

[0088] More preferably, the unnatural amino acid of formula (I) or (II) is

[0089] [ka]

[0090] or a salt, solvate, tautomer, isomer, or mixture thereof. Thus, in this more preferred embodiment, the unnatural amino acid of Formula (I) or (II) is in the L-form.

[0091] Suitably, the unnatural amino acid of formula (I) or (II) is

[0092] [ka]

[0093] or a salt, solvate, tautomer, isomer, or mixture thereof. Preferably, the unnatural amino acid of formula (I) or (II) is in the D-form. More preferably, the unnatural amino acid of formula (I) or (II) is in the L-form.

[0094] More preferably, the unnatural amino acid of formula (I) or (II) is

[0095] [ka]

[0096] or a salt, solvate, tautomer, isomer or mixture thereof.

[0097] Suitably, the unnatural amino acid is an unnatural amino acid of formula (I) or a salt, solvate, tautomer, isomer or mixture thereof.

[0098] Suitably, the unnatural amino acid of formula (I) has the formula

[0099] [ka]

[0100] or a salt, solvate, tautomer, isomer or mixture thereof.

[0101] Suitably, the unnatural amino acid of formula (I) has the formula

[0102] [ka]

[0103] or a salt, solvate, tautomer, isomer or mixture thereof.

[0104] Suitably, the unnatural amino acid of formula (I) has the formula

[0105] [ka]

[0106] or a salt, solvate, tautomer, isomer or mixture thereof.

[0107] Preferably, X is X1-Y, SS-R5, O-NH-R5, S-NH-R5, X2-Y1, X3-Y2, N3 or NH-S(O)2-Y3. Preferably, X is X1-Y, X2-Y1, X3-Y2, N3 or NH-S(O)2-Y3. More preferably, X is X1-Y.

[0108] Preferably, X1 is S, Se, O, NH, N(CH3) or N(CH2CH3). More preferably, X1 is S or O. More preferably, X1 is S.

[0109] Preferably, X2 is S or O. Preferably, X2 is S.

[0110] Preferably, X4 is NH-C(O)-O, O, S or NH. Preferably, X4 is NH-C(O)-O.

[0111] Preferably, X5 is S, O, NH, NC(O)-O-CH3, NC(O)-O-CH2CH3, NC(O)-O-Ph, NS(O)2CH3, NS(O)2CH2CH3, N-CH3, N-CH2CH3 or N-Ph. More preferably, X5 is S, O, NH or N-CH3.

[0112] R1 is H, an amino acid residue, or a peptide. Preferably, R1 is H or a proteinogenic amino acid residue. Most preferably, R1 is H.

[0113] Preferably, R2 is H, C 1-6 Alkyl, C 1-6 haloalkyl or phenyl More preferably, R2 is H, CH3, CH2CH3, CF3 or phenyl. More preferably, R2 is H.

[0114] Preferably, q is 1 or 2. More preferably, q is 1.

[0115] Suitably, R3 and R4 are each independently selected from H, F, Cl, Br, CH3, CH2CH3, CF3, Ph, pyridyl, pyrrolyl, imidazolyl, OCH3, OCH2CH3, SCH3, SCH2CH3, NH(CH3), NH(CH2CH3), N(CH3)2 and N(CH2CH3)2.

[0116] Preferably, each R3 is H, CH3, CH2CH3, CF3, Ph, OCH3 or OCH2CH3. More preferably, each R3 is H.

[0117] Preferably, each R4 is H, CH3, CH2CH3, CF3, Ph, OCH3 or OCH2CH3. More preferably, each R4 is H.

[0118] Suitably, R5 is H, halo, C 1-6 Alkyl, C 1-6 Haloalkyl, C 5-20 Aryl, C 3-20 Heteroaryl, OC 1-6 Alkyl, NH(C 1-6 Alkyl) and N(C 1-6 alkyl)2.

[0119] More preferably, R5 is H, F, Cl, Br, CH3, CH2CH3, CF3, Ph, pyridyl, pyrrolyl, imidazolyl, OCH3, OCH2CH3, SCH3, SCH2CH3, NH(CH3), NH(CH2CH3), N(CH3)2 or N(CH2CH3)2.

[0120] When R6 or R9 is a substituent other than H, the protecting group Y contains a stereocenter at the carbon to which R6 or R9 is attached. Suitably, Y is a racemic mixture or has the (R) or (S) configuration with respect to the stereocenter at the carbon to which R6 or R9 is attached.

[0121] In some embodiments, more preferably, Y has the (R) configuration with respect to the stereocenter at the carbon to which R6 or R9 is attached.

[0122] In some embodiments, more preferably, Y has the (S) configuration with respect to the stereocenter at the carbon to which R6 or R9 is attached.

[0123] Suitably, the compound of formula (I) or (II) or a salt, solvate, tautomer, isomer or mixture thereof comprises a Y group.

[0124] Preferably, Y is

[0125] [ka] TIFF2025023942000020.tif157154

[0126] It is.

[0127] More preferably, the unnatural amino acid is

[0128] [ka]

[0129] It includes a Y group which is

[0130] More preferably, the unnatural amino acid is

[0131] [ka]

[0132] It includes a Y group which is

[0133] Suitably, R6 is selected from H, CH3, CH2CH3, CF3, CO2H, CO2CH3, CO2CH2CH3 and Ph.

[0134] More preferably, R6 is CH3.

[0135] Suitably, R7 and R8 are independently selected from H, OH, OCH3, OCH2CH3 and O-Ph, or R7 and R8 are linked together to form an O-CH2-O group.

[0136] In some embodiments, more preferably, R7 and R8 are the same. Preferably, R7 and R8 are H, OH, OCH3 or OCH2CH3, or R7 and R8 are linked together to form an O-CH2-O group.

[0137] More preferably, R7 and R8 are linked together to form an O-CH2-O group.

[0138] Preferably, each R' is independently selected from H, CH3, CH2CH3, CF3, and Ph. can be.

[0139] Preferably, R9 is selected from H, CH3, CH2CH3, CF3, CO2H, CO2CH3, CO2CH2CH3 and Ph. Preferably, R9 is selected from H, CH3, CH2CH3 and CF3.

[0140] Preferably, R 10 is selected from H, CH3, CH2CH3, and CF3.

[0141] Preferably, R 11 is selected from H, CH3, CH2CH3, and CF3.

[0142] Suitably, in one embodiment, M + Li + , Na + Or K + It is.

[0143] In one embodiment, M + Li + In an alternative embodiment, M + Na + In an alternative embodiment, M + is N(R 13 )4 + In a more preferred embodiment, M + is K + It is.

[0144] In one embodiment, Z is Si. In an alternative embodiment, Z is Ge.

[0145] Preferably, R 12 is CH3, CH2CH3 or C(O)-(Ph).

[0146] In one embodiment, preferably, R 13 is C 1-6 Alkyl, aryl or C 5-20 In an alternative embodiment, preferably R 13 is H, CH3 or CH2CH3 or allyl. 13 is CH3 or CH2CH3.

[0147] More preferably, Y is

[0148] [ka]

[0149] It is.

[0150] In some embodiments, Y is

[0151] [ka]

[0152] It is.

[0153] In one embodiment, the compound of formula (I) or (II) is

[0154] [ka]

[0155] or a salt, solvate, tautomer, isomer or mixture thereof.

[0156] A highly preferred embodiment is

[0157] [ka]

[0158] or a salt, solvate, tautomer, isomer, or mixture thereof.

[0159] This embodiment may be referred to herein as "6" or "compound 6" or "DAP5," each of which refers to the same chemical structure shown above.

[0160] In one embodiment, more preferably the unnatural amino acid is

[0161] [ka]

[0162] or a salt, solvate, tautomer, isomer or mixture thereof.

[0163] More preferably, the unnatural amino acid is a compound of formula (III) or (IV) or a salt thereof, In one embodiment, the unnatural amino acid is a compound of formula (III) or a salt, solvate, tautomer, isomer, or mixture thereof: In another embodiment, the unnatural amino acid is a compound of formula (IV) or a salt, solvate, tautomer, isomer, or mixture thereof.

[0164] In one embodiment, one or more of compounds 2, 3, 4, or 5 of Figure 3 can be used in an in vitro translation system using the synthetases described above, or, for example, the synthetase described in (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014)), i.e., synthetase "PCC1RS", which is MbPylRS with the mutations N311M, C313Q, V366G, W382N, R85H. This PCC1RS synthetase incorporates a very similar unnatural amino acid (photocaged cysteine) and we claim that it will accept a protected version of DAP that differs by only one atom (or two atoms together with a hydrogen atom). Another technique would be to load tRNA with so-called "flexizymes" that can load tRNA with virtually any amino acid (see, for example, Morimoto et al., 2011, Acc. Chem Res. 44). These could then be used in in vitro systems. Without wishing to be bound by theory, the inventors believe that analogous compounds 2, 3, 4 or 5 in Figure 3 do not function in cells because they cannot enter the cells.

[0165] Other forms Unless otherwise specified, known ionic, salt or solvate forms of these substituents are included in the above. For example, a reference to carboxylic acid (-COOH) also refers to its anionic (carboxylate) form (-COO - ), salts or solvates. Similarly, a reference to an amino acid includes the protonated form of the amino group (-N + HR 1 R 2 ), salts or solvates, e.g., hydrochloride salts Similarly, references to a hydroxyl group also include its anionic form (-O- ), including salts or solvates.

[0166] Isomers, salts and solvates Certain compounds may exist in one or more geometric, optical, enantiomeric, diastereomeric, epimeric, atropic, stereoisomeric, tautomeric, conformational, or anomeric forms, including, but not limited to, cis and trans forms; E and Z forms; c, t and r forms; endo and exo forms; R, S and meso forms; D and L forms; d and l forms; (+) and (-) forms; keto, enol and enolate forms; syn and anti forms; synclinal and anticlinal forms; alpha and beta forms; axial and equatorial forms; boat, chair, twisted, envelope and half-chair forms; and combinations thereof, collectively referred to as "isomers" (or "isomeric forms").

[0167] As discussed below with respect to tautomeric forms, it should be noted that the term "isomer" as used herein excludes structural isomers (i.e., isomers that differ solely in the connections between the atoms rather than their positions in space). For example, a reference to a methoxy group, -OCH3, should not be construed as a reference to its structural isomer, a hydroxymethyl group, -CH2OH.

[0168] Reference to a structural class may include structural isomeric forms that fall within that class (e.g., C 1-7 Alkyl includes n-propyl and isopropyl, butyl includes n-butyl, isobutyl, sec-butyl and tert-butyl, methoxyphenyl includes ortho-methoxyphenyl, meta-methoxyphenyl and para-methoxyphenyl).

[0169] The above exclusion does not apply to tautomeric forms, such as, for example, amide / iminoalcohol -NH-C(=O)- / -N=C(-OH)-; or keto, enol and enolate forms, such as in the following tautomeric pairs: keto / enol, imine / enamine, amide / iminoalcohol, amidine / amidine, nitroso / oxime, thioketone / enethiol, N-nitroso / hydroxyazo and nitro / aci-nitro.

[0170] It should be noted that the term "isomer" specifically includes compounds with one or more isotopic substitutions. For example, H is 1 H, 2 H(D) and 3 H can be in any isotopic form, including T; 11 C. 12 C. 13 C and 14 C may be in any isotopic form; O may be in any isotopic form, 16 O and 18 For example, compound 6 can be the following:

[0171] [ka]

[0172] As shown in 13 C and 15 N isotope-labeled compounds or 18 The compound may be an O isotope-labeled compound.

[0173] Unless otherwise specified, a reference to a particular compound includes all such isomeric forms, including (wholly or partially) racemic mixtures thereof and any other mixtures. An example of a bond designated in an alternative manner is the following:

[0174] [ka]

[0175] is a CH group in DAP with the assigned stereochemistry shown in

[0176] Methods for the preparation (e.g., asymmetric synthesis) and separation (e.g., fractional crystallization and chromatographic means) of such isomeric forms are known in the art or may be readily obtained by adapting the methods taught herein or known methods in a known manner.

[0177] Unless otherwise specified, a reference to a particular compound also includes ionic, salt, solvate and protected forms thereof, for example, as discussed below.

[0178] In some embodiments, the compound of Formula (I) or (II) or a salt, solvate, tautomer, isomer, or mixture thereof includes a pharma- ceutically acceptable salt of the compound of Formula (I) or (II).

[0179] Compounds of formula (I) or (II), including those specifically named above, may include pharma- ceutically acceptable complexes, salts, solvates and hydrates, including non-toxic acid addition salts (including diacids) and base salts.

[0180] The compound has a cationic or cationic functional group (e.g., -NH2 is -NH3 +In the case of carboxylic acids such as carboxylic acids which may be carboxylic acids such as carboxylic acids, acid addition salts may be formed with suitable anions. Examples of suitable inorganic anions include, but are not limited to, those derived from the following inorganic acids: hydrochloric acid, nitric acid, nitrous acid, phosphoric acid, sulfuric acid, sulfurous acid, hydrobromic acid, hydroiodic acid, hydrofluoric acid, phosphoric acid, and phosphorous acid. Examples of suitable organic anions include, but are not limited to, those derived from the following organic acids: 2-acetyoxybenzoic acid, acetic acid, ascorbic acid, aspartic acid, benzoic acid, camphorsulfonic acid, cinnamic acid, citric acid, edetic acid, ethanedisulfonic acid, ethanesulfonic acid, fumaric acid, glucoheptonic acid, gluconic acid, glutamic acid, glycolic acid, hydroxymaleic acid, hydroxynaphthalenecarboxylic acid, isethionic acid, lactic acid, lactobionic acid, lauric acid, maleic acid, malic acid, methanesulfonic acid, mucic acid, oleic acid, oxalic acid, palmitic acid, pamoic acid, pantothenic acid, phenylacetic acid, phenylsulfonic acid, propionic acid, pyruvic acid, salicylic acid, stearic acid, succinic acid, sulfanilic acid, tartaric acid, toluenesulfonic acid, and valeric acid. Examples of suitable polymeric organic anions include, but are not limited to, those derived from the following polymeric acids: tannic acid, carboxymethylcellulose. Such salts include acetate, adipate, aspartate, benzoate, besylate, bicarbonate, carbonate, bisulfate, sulfate, borate, camsylate, citrate, cyclamate, edisylate, esylate, gypsum salts, and the like. acid salt, fumarate, gluceptate, gluconate, glucuronate, hexafluorophosphate, hybenzate, hydrochloride / chloride, hydrobromide / bromide, hydroiodide / iodide, isethionate, lactate, malate, maleate, malonate, mesylate, methylsulfonate, naphthylate, 2-napsyllate, nicotinate , nitrate, orotate, oxalate, palmitate, pamoate, phosphate, hydrogen phosphate, dihydrogen phosphate, pyroglutamate, saccharate, stearate, succinate, tannate, tartrate, tosylate, trifluoroacetate and xinofoate salts.

[0181] For example, a compound may have an anionic or potentially anionic functional group (e.g., -COOH is COO - In the case of cations such as ammonium salts, which may be ammonium salts, base salts may be formed with suitable cations. Examples of suitable inorganic cations include, but are not limited to, metal cations such as alkali metal or alkaline earth metal cations, ammonium and substituted ammonium cations, and amines. Examples of suitable metal cations include sodium (Na + ), potassium (K + ), Magnesium (Mg 2+ ), Calcium (Ca 2+ ), Zinc (Zn 2+ ) and aluminum (Al 3+ ), examples of suitable organic cations include, but are not limited to, ammonium ion (i.e., NH + ) and substituted ammonium ions (e.g., NHR + , NH2R2 + , NHR3 + , NR4 + Some suitable examples of substituted ammonium ions include ethylamine, diethylamine, dicyclohexylamine, triethylamine, butylamine, ethylenediamine, ethanolamine, diethanolamine, piperazine, benzylamine, phenylbenzylamine, choline, meglumine, and tromethamine, as well as those derived from amino acids such as lysine and arginine. An example of a common quaternary ammonium ion is N(CH3)4 +Examples of suitable amines include arginine, N,N'-dibenzylethylenediamine, chloroprocaine, choline, diethylamine, diethanolamine, dicyclohexylamine, ethylenediamine, glycine, lysine, N-methylglucamine, olamine, 2-amino-2-hydroxymethyl-propane-1,3-diol, and procaine. For a discussion of useful acid addition and base salts, see SM Berge et al., J. Pharm. Sci. (1977) 66:1-19; also see Stahl and Wermuth, Handbook of Pharmaceutical Salts: Properties, Selection, and Use (2011).

[0182] Pharmaceutically acceptable salts may be prepared using various methods. For example, a compound of formula (I) or (II) may be reacted with a suitable acid or base to produce the desired salt. A precursor of a compound of formula (I) or (II) may also be reacted with an acid or base to remove an acid-labile or base-labile group or to open a lactone or lactam group of the precursor. In addition, a salt of a compound of formula (I) or (II) may be converted to another salt by treating with a suitable acid or base or by contacting with an ion exchange resin. Following the reaction, the salt may then be isolated by filtration if the salt precipitates from the solution or by evaporation to recover the salt. The degree of ionization of the salt may vary from completely ionized to almost non-ionized.

[0183] It may be convenient or desirable to prepare, purify and / or handle the corresponding solvates of the active compounds. The term "solvate" refers to a molecular complex that contains a compound and one or more pharma- ceutically acceptable solvent molecules (e.g., EtOH). The term "hydrate" refers to a solvate in which the solvent is water. Pharmaceutically acceptable solvates include those in which the solvent may be isotopically substituted (e.g., DO, acetone-d6, DMSO-d6).

[0184] A currently accepted classification system for solvates and hydrates of organic compounds distinguishes between isolated site solvates, channel solvates and metal ion coordinated solvates and hydrates. See, for example, KR Morris (HG Brittain ed.) Polymorphism in Pharmaceutical Solids (1995). Isolated site solvates and hydrates are those in which the solvent (e.g. water) molecules are isolated from direct contact with each other by intervening molecules of the organic compound. In channel solvates, the solvent molecules lie in lattice channels where they are adjacent to other solvent molecules. In metal ion coordinated solvates, the solvent molecules are bonded to the metal ion.

[0185] If the solvent or water is tightly bound, the complex will have a well-defined stoichiometry that is independent of humidity. However, if the solvent or water is weakly bound, such as in channel solvates and hygroscopic compounds, the water or solvent content will vary with humidity and drying conditions. In such cases, non-stoichiometry will typically be observed.

[0186] Genetic integration With respect to the production of the polypeptides according to the invention by genetic integration, said genetic integration preferably uses an orthogonal or extended genetic code in which one or more specific orthogonal codons are assigned to code for an unnatural amino acid of interest such that the unnatural amino acid can be genetically incorporated by using an orthogonal tRNA synthetase / tRNA pair. The orthogonal tRNA synthetase / tRNA pair can in principle be any such pair that is capable of charging a tRNA with the unnatural amino acid and incorporating the unnatural amino acid of interest into a polypeptide chain in response to the orthogonal codon.

[0187] The orthogonal codon can be an amber, ochre, opal, or quadruplet codon of the orthogonal codon. The codon simply must correspond to the orthogonal tRNA that will be used to deliver the unnatural amino acid of interest. Most preferably, the orthogonal codon is amber.

[0188] It should be noted that many of the specific examples shown herein used amber codons and corresponding tRNA / tRNA synthetases. As mentioned above, these can vary. Alternatively, one may simply swap the anticodon region of a tRNA with the desired anticodon region for the codon of choice to use other codons without the bother of using or selecting an alternative tRNA / tRNA synthetase pair that can act on the unnatural amino acid of interest. Such swapping is entirely within the realm of the skilled worker, since the anticodon region is not involved in the charging or incorporation function of the tRNA, nor in recognition by the tRNA synthetase. Thus, in some embodiments, the anticodon region of a tRNA used in the present invention, e.g., MbtRNA CUA or MmtRNA CUA is interchangeable, i.e., chimeric tRNAs can be constructed to swap anticodons to recognize alternative codons. CUA can be used such that an unnatural amino acid of interest can be incorporated in response to different orthogonal codons discussed herein, including ochre, opal, or tetrad codons, and correspondingly, the nucleic acid encoding the polypeptide into which the unnatural amino acid of interest is to be incorporated is mutated to introduce the cognate codon at the incorporation point of the unnatural amino acid of interest. Most preferably, the orthogonal codon is amber.

[0189] Thus, alternative orthogonal tRNA synthetase / tRNA pairs may be used, if desired, provided the desired charging activity is maintained.

[0190] The Methanosarcina barkeri PylT gene encodes MbtRNA CUAtRNA (i.e., MbtRNA Pyl CUA ) is preferably encoded by tRNA CUA The sequence is as follows: tRNAcua MbPylT (from MS strain, Genbank accession number AY064401) gggaacctgatcatgtagatcgaatggactctaaatccgttcagccgggt tagattcccggggtttccgcca

[0191] There are two variants of this tRNA that can be used. One starts with ggg (see above). ), the other starts with gga (see below): tRNAcua "gga" variant MbPylT (from MS strain, Genbank accession number AY064401) ggaaacctgatcatgtagatcgaatggactctaaatccgttcagccgggt tagattcccggggtttccgcca

[0192] There is no substantial difference between these two variants.

[0193] The Methanosarcina barkeri PylS gene encodes the MbPylRS tRNA synthetase protein.

[0194] tRNA synthetase If necessary, one skilled in the art can adapt the MbPylRS tRNA synthetase protein by mutating it to optimize it for the particular unnatural amino acid used. The need for mutation (if any) depends on the particular unnatural amino acid used. One example in which it may be necessary to mutate the MbPylRS tRNA synthetase is when the particular unnatural amino acid used is not processed by the MbPylRS tRNA synthetase protein.

[0195] In the present invention, the inventors have devoted considerable intellectual effort to the creation of a novel synthetase DAPRS. See in particular Example 3. Advantageously, the synthetase of the present invention contains the following amino acids: C at position 271, Q at position 311, F at position 349 and C at position 366 (i.e., the amino acids related to MbPylRS). For example, Y271C, N311Q, Y349F and V366C), or equivalent positions when a different start synthetase sequence / backbone synthetase sequence from another species is used.

[0196] Exemplary sequences of novel synthetase DAPRS: DAPRS (mutated residues are underlined): MDKKPLDVLISATGLWMSRTGTLHKIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVP SPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINNDTELSKQIFRVDKNLCLRPMLAPTL C NYLRKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMV Q FCQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMV F GDTLDIMHGDLELSSA C VGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL

[0197] Preferably, the orthogonal synthetase / tRNA pair is Methanosarcina barkeri MS pyrrolysine tRNA synthetase (MbPylRS) and its cognate activators. Member suppressor tRNA (MbtRNA CUA ) (i.e., MbtRNA PylCUA ), wherein the MbPylRS contains mutations described herein that are important for its activity, i.e., MbPylRS contains a C at position 271, a Q at position 311, an F at position 349 and a C at position 366 (i.e., Y271C, N311Q, Y349F and V366C).

[0198] The tRNA synthetases of the present invention can vary, and although specific tRNA synthetase sequences may be used in the examples, it is not intended that the invention be limited to only those examples.

[0199] In principle, any tRNA synthetase that provides the same tRNA charging (aminoacylation) function can be used in the present invention.

[0200] For example, the tRNA synthetase can be derived from any suitable species, e.g., an archaea, e.g., Methanosarcina barkeri MS; Methanosarcina barkeri strain Fusaro; Methanosarcina mazei Go1; Methanosarcina acetivorans C2A; Methanosarcina thermophila; or Methanococcoides burtonii. Alternatively, the tRNA synthetase may be derived from a bacterium, such as Desulfitobacterium hafniense DCB-2; Desulfitobacterium hafniense Y51; Desulfitobacterium hafniense PCP1; Desulfotomaculum acetoxidans DSM 771.

[0201] Exemplary sequences from these organisms are published sequences. The following examples are provided as exemplary sequences for pyrrolysine tRNA synthetases: > M. Berkeley MS / 1~419 / Methanosarcina barkeri MS Version Q6WRH6.1 GI:74501411 MDKKPLDVLISATGLWMSRTGTLHKIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRKNDFQRLYTNDREDYLG KLERDITKFFVDRGFLEIKSPILIPAEYVERMGINNDTELSKQIFRVDKNLCLRPMLAPTLYNYLRKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMVNFCQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL > M. Berkeley F / 1~419 / Methanosarcina barkeri strain Fusaro Version YP_304395.1 GI:73668380 MDKKPLDVLISATGLWMSRTGTLHKIKHYEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTEGKTSVKVKVVSAPKVKKAMPKSVSRAPKPLENPVSAKASTDTSRSVPSPAKSTPNSPVPTSAPAPSLTRSQLDRVEALLSPEDKISLNIAKPFRELESELVTRRKNDFQRLYTNDREDYLGKLERDITKFFVDRDFLEIKSPILIPAEYVERMGINNDTELSKQIFRVDKNLCRPMLAPTLYNYLRKLDRILPDPKIFIEVGPCYRKESDGKEHLEEFTMVNFCQMGSGCTRENLESLIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELLSAVVGPVPLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL >M.マゼイ / 1~454 メレノサルシナ·マゼイ Go1 versionNP_633469.1 GI:21227547 MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSSISSISTGATASALVKGNTNPITSMSPAVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELLSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL >M.Acetivorans / 1~443 Methanosarcina acetivorans C2A Version NP_615128.2 GI:161484944 MDKKPLDTLISATGLWMSRTGMIHKIKHHEVSRSKIYIEMACGERLVVNNSRSSRTARALRHHKYRKTCRHCRVSDEDINNFLTKTSEEKTTVKVKVVSAPRVRKAMPKS VARAPKPLEATAQVPLSGSKPAPATPVSAPAQAPAPSTGSASATSASAQRMANSAAAPAAPVPTSAPALTKGQLDRLEGLLSPKDEISLDSEKPFRELESELLSRRKKDLK RIYAEERENYLGKLEREITKFFVDRGFLEIKSPILIPAEYVERMGINSDTELSKQVFRIDKNFCLRPMLAPNLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFT MLNFCQMGSGCTRENLEAIITEFLNHLGIDFEIIGDSCMVYGNTLDVMHDDLELSSAVVGPVPLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRAARSESYYNGISTNL >M.Thermophila / 1~478 Methanosarcina thermophila, version DQ017250.1 GI:67773308 MDKKPLNTLISATGLWMSRTGKLHKIRHHEVSKRKIYIEMECGERLVVNNSRSCRAARALRHHKYRKICKHCRVSDEDLNKFLTRTNEDKSNAKVTVVSAPKIRKVMPKSVARTPKPlentapvQTLPSESQPAPTTPISASTTAPASTSTTAPAPASTTAPASASTTISTSAMPASTASQGTTKFNYISGGFPRPIPVQASAPALTKSQIDRLQGLLSPKDEISLDSGTPFRKLESELLSRRRKDLKQIYAEEREHYLGKLEREITKFFVDRGFLEIKSPILIPMEYERMGIDNDKELSKQIFRVDNNFCLRPMLAPNLYNYLRKLNRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLEAIIKDFLDYLGIDFEIVGDSCMVYGDTLDVMHGDLELLSAVVGPVPMDRDWGINKPWIGAGFGLERLLKVMHNFKNIKRASRSESYYNGISTNL >M.ブルトニイ / 1~416 レトノコッコイデス·ブルトニイ DSM 6242、VERSIONYP_566710.1 GI:91774018 MEKQLLDVLVELNGVWLSRSGLLHGIRNFEITTKHIHIETDCGARFTVRNSRSSRSARSLRHNKYRKPCKRCRPADEQIDRFVKKTFKEKRQTVSVFSSPKKHVPKKPKVAVIKSFSISTPSPKEASVSNSIPTPSISVVKDEVKVPEVKYTPSQIERLKTLMSPDDKIPIQDELPEFKVLEKELIQRRRDDLKKMYEEDREDRLGKLERDITEFFVDRGFLEIKSPIMIPFEYIERMGIDKDDHLNKQIFRVDESMCLRPMLAPCLYNYLRKLDKVLPDPPIRIFEIGPCYRKESDGSSHLEEFTMVNFCQMGSGCTRENMEALIDEFLEHLGIEYEIEADNCMVYGDTIDIMHGDLELSSAVVGPIPLDREWGVNKPWMGAGFGLERLLKVRHNYTNIRRASRSELYYNGINTNL >D.Hafniens_DCB-2 / 1~279 Desulfitobacterium hafniens DCB-2 Version YP_002461289.1 GI:219670854 MSSFWTKVQYQRLKELNASGEQLEMGFSDALSRDRAFQGIEHQLMSQGKRHLEQLRTVKHRPALLELEEGLAKALHQQGFVQVVTPTIITKSALAKMTIGEDHPLFSQVFWLDGKKCLRPMLAPNLYTLWRELERLWDK PIRIFEIGTCYRKESQGAQHLNEFTMLNLTELGTPLEERHQRLEDMARWVLEAAGIREFELVTESSVVYGDTVDVMKGDLELASGAMGPHFLDEKWEIVDPWVGLGFGLERLLMIREGTQHVQSMARSLSYLDGVRLNIN >D. Hafniens_Y51 / 1~312 Desulfitobacterium hafniens Y51 Version YP_521192.1 GI:89897705 MDRIDHTDSKFVQAGETPVLPATFMFLTRRDPPLSSFWTKVQYQRLKELNASGEQLEMGFSDALSRDRAFQGIEHQLMSQGKRHLEQLRTVKHRPALLELEEGLAKALHQQGFVQVVTPTIITKSALAKMTIGEDHPLFSQVFWLDGKKCLRPPMLA PNLYTLWRELERLWDKPIRIFEIGTCYRKESQGAQHLNEFTMLNLTELGTPLEERHQRLEDMARWVLEAAGIREFELVTESSVVYGDTVDVMKGDLELASGAMGPHFLDEKWEIVDPWVGLGFGLERLLMIREGTQHVQSMARSLSYLDGVRLNIN >D. Hafniens PCP1 / 1~288 Desulfitobacterium hafniens Version AY692340.1 GI:53771772 MFLTRRDPPLSSFWTKVQYQRLKELNASGEQLEMGFSDALSRDRAFQGIEHQLMSQGKRHLEQLRTVKHRPALLELEEKLAKALHQQGFVQVVTPTIITKSALAKMTIGEDHPLFSQVFWLDGKKCLRPMLAPNLYTLWRELER LWDKPIRIFEIGTCYRKESQGAQHLNEFTMLNLTELGTPLEERHQRLEDMARWVLEAAGIREFELVTESSVVYGDTVDVMKGDLELASGAMGPHFLDEKWEIFDPWVGLGFGLERLLMIREGTQHVQSMARSLSYLDGVRLNIN >D.Acetoxydance / 1~277 Desulfotomaculum acetoxydans DSM 771 Version YP_003189614.1 GI:258513392 MSFLWTVSQQKRLSELNASEEKNMSFSSTSDREAAYKRVEMRLINESKQRLNKLRHETRPAICALENRLAAALRGAGFVQVATPVILSKKLLGKMTITDEHALFSQVFWIEENKCLRPMLAPNLYYILKDLLRLWEK PVRIFEIGSCFRKESQGSNHLNEFTMLNLVEWGLPEEQRQKRISELAKLVMDETGIDEYHLEHAESVVYGETVDVMHRDIELGSGALGPHFLDGRWGVVGPWVGIGFGLERLLMVEQGGQNVRSMGKSLTYLDGVRLNI

[0202] In cases where a particular tRNA charge (aminoacylation) function is provided by mutating a tRNA synthetase, it may not be appropriate to simply use another wild-type tRNA synthetase sequence, e.g., one selected from above. In this scenario, it becomes important to preserve the same RNA charge (aminoacylation) function. This is achieved by transferring an exemplary tRNA synthetase to an alternative tRNA synthetase backbone, e.g., one selected from above.

[0203] In this way, it should be possible to transfer selected mutations beyond the corresponding tRNA synthetase sequences, e.g., the exemplary M. barkeri and / or M. mazei sequences, to corresponding pylS sequences from other organisms.

[0204] Target tRNA synthetase proteins / backbones can be selected by alignment to known tRNA synthetases, such as the exemplary M. barkeri and / or M. mazei sequences.

[0205] This subject will now be illustrated by reference to the pylS (pyrrole lysine tRNA synthetase) sequence, although the principles apply equally to any particular tRNA synthetase of interest.

[0206] For example, an alignment of all PylS sequences can be prepared. These may have a low overall % sequence identity. It is therefore important to study the sequences, e.g., by aligning them to known tRNA synthetases (rather than simply using a low sequence identity score), to ensure that the sequence being used is in fact a tRNA synthetase.

[0207] Thus, preferably, when sequence identity is considered, it is preferably considered over the entire sequence of the exemplary tRNA synthetases. Preferably, the % identity may be defined from an alignment of the sequences.

[0208] It may be useful to focus on the catalytic region, the aim being to provide a tRNA catalytic region with a definable high % identity to capture / identify a main chain scaffold suitable to accommodate mutations translated to provide the same tRNA charge (aminoacylation) function, e.g., novel or unnatural amino acid recognition.

[0209] Therefore, preferably when sequence identity is considered it is preferably considered across the entire catalytic region. Preferably the % identity may be defined from the catalytic region.

[0210] "Moving" or "grafting" a mutation into an alternative tRNA synthetase backbone can be accomplished by site-directed mutagenesis of the nucleotide sequence encoding the tRNA synthetase backbone. This technique is well known in the art. Essentially, a backbone pylS sequence is selected (e.g., using the active site alignment discussed above) and the selected mutation is moved to (i.e., made at) the corresponding / homologous position.

[0211] When referring to specific amino acid residues using numerical addresses, unless otherwise clear, the numbering refers to the sequence encoded by the MbPylRS (Methanosarcina barkeri pyrrolysyl-tRNA synthetase) amino acid sequence (i.e., the published wild-type Methanosarcina barkeri PylS gene, accession number Q46E77) as the reference sequence): MDKKPLDVLI SATGLWMSRT GTLHKIKHYE VSRSKIYIEM ACGDHLVVNN SRSCRTARAF RHHKYRKTCK RCRVSDEDIN NFLTRSTEGK TSVKVKVVSA PKVKKAMPKS VSRAPKPLEN PVSAKASTDT SRSVPSPAKS TPNSPVPTSA PAPSLTRSQL DRVEALLSPE DKISLNIAKP FRELESELVT RRKNDFQRLY TNDREDYLGK LERDITKFFV DRDFLEIKSP ILIPAEYVER MGINNDTELS KQIFRVDKNL CLRPMLAPTL YNYLRKLDRI LPDPIKIFEV GPCYRKESDG KEHLEEFTMV NFCQMGSGCT RENLESLIKE FLDYLEIDFE IVGDSCMVYG DTLDIMHGDL ELSSAVVGPV PLDREWGIDK PWIGAGFGLE RLLKVMHGFK NIKRARSES YYNGISTNL This is done using.

[0212] This should be used as is well understood in the art to locate residues of interest. This is not necessarily a strict counting exercise - attention must be paid to the context or alignment. For example, if the proteins of interest are of slightly different lengths, locating the correct residue in that sequence that corresponds to (say) L266 may require aligned sequences and careful selection of the equivalent or corresponding residue, rather than simply taking the 266th residue of the sequence of interest. This is well within the realm of the skilled reader.

[0213] The notation of mutations used herein is standard in the art. For example, L266M means that the amino acid corresponding to L at position 266 of the wild-type sequence has been replaced with M.

[0214] In this notation, the amino acid introduced is the most important information. For example, if the "L266M" mutation is grafted into another synthetase starting sequence (backbone sequence), this alternative starting sequence may possibly not have an "L" at position 266 (or at the position corresponding to L266 in the reference sequence as explained above). However, in this case, it is important to identify the correct amino acid in the synthetase starting sequence (backbone sequence) of interest that corresponds to L266 in the reference sequence, and to change this residue (whatever it may be) to M. Thus, when grafting into an alternative backbone / starting sequence, "L266M" may be understood as "X266M".

[0215] For the exemplary DAPRS synthetases described herein, the mutations are These are designated Y271C, N311Q, Y349F and V366C compared to the reference sequence (see above), or, when grafted onto a backbone having different amino acids at those positions in the starting sequence, would be understood as X271C, X311Q, X349F and X366C.

[0216] Next, the transplantation of mutations between alternative tRNA backbones is illustrated with respect to exemplary M. barkeri and M. mazei sequences, although the same principles apply equally to transplantation into or from other backbones.

[0217] For example, Mb AcKRS is a synthetase designed for the incorporation of AcK. Parent protein / backbone: M. barkeri PylS Mutations: L266V, L270I, Y271F, L274A, C317F

[0218] Mb PCKRS: a synthetase designed for the incorporation of PCK Parent protein / backbone: M. barkeri PylS Mutations: M241F, A267S, Y271C, L274M

[0219] Synthetases with the same substrate specificity can be obtained by transferring these mutations into M. mazei PylS. Thus, by transferring mutations from the Mb backbone to the Mm tRNA backbone, the following synthetases: Mm AcKRS, which introduces the mutations L301V, L305I, Y306F, L309A, and C348F into M. mazei PylS; and Mm PCKRS introducing the mutations M276F, A302S, Y306C, and L309M into M. mazei PylS may be generated.

[0220] The full-length sequences of these exemplary grafted mutant synthetases are shown below. >Mb_PylS / 1-419 MDKKPLDVLISATGLWMSRTGTLHKIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINNDTELSKQIFRVDKNLCLRPMLAPTLYNYLRKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMVNFCQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL >Mb_AcKRS / 1-419 MDKKPLDVLISATGLWMSRTGTLHKIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSGEDINNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERMGINNDTELSKQIFRVDKNLCLRPMVAPTIFNYARKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMVNFFQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL >Mb_PCKRS / 1-419 MDKKPLDVLISATGLWMSRTGTLHKIKHHEVSRSKIYIEMACGDHLVVNNSRSCRTARAFRHHKYRKTCKRCRVSDEDINNFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKSTPNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLNMAKPFRELEPELVTRRKNDFQRLYTNDREDYLGKLERDITKFFVDRGFLEIKSPILIPAEYVERFGINNDTELSKQIFRVDKNLCLRPMLSPTLCNYMRKLDRILPGPIKIFEVGPCYRKESDGKEHLEEFTMVNFCQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHGDLELLSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL >Mm_PylS / 1-454 MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSTARARRHHKYRKTCKRCRVSDEDLN KFLTKANETQTSVKVKVVSAPTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSPAVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMLAPNLYNYLRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELLSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL >Mm_AcKRS / 1-454 MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSPAVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIFRVDKNFCLRPMVAPNIFNYARKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFFQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELLSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL >Mm_PCKRS / 1-454 MDKKPLNTLISATGLWMSRTGTIHKIKHHEVSRSKIYIEMACGDHLVVNNSRSSRTARALRHHKYRKTCKRCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAMPKSVARAPKPLENTEAAQAQPSGSKFSPAIPVSTQESVSVPASVSTSSISSISTGATASALVKGNTNPITSMSPAVQASAPALTKSQTDRLEVLLNPKDEISLNSGKPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERFGIDNDTELSKQIFRVDKNFCLRPMLSPNLCNYMRKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFCQMGSGCTRENLESIITDFLNHLGIDFKIVGDSCMVYGDTLDVMHGDLELLSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVKHDFKNIKRAARSESYYNGISTNL

[0221] The same principles apply equally to other mutations and / or other backbones.

[0222] The grafted polypeptides so produced should advantageously be tested to ensure that the desired function / substrate specificity is preserved.

[0223] In one embodiment, the tRNA can be from one species, e.g., Methanosarcina barkeri, and the tRNA synthetase can be from another species, e.g., Methanosarcina mazei. In another embodiment, the tRNA can be from a first species, e.g., Methanosarcina mazei, and the tRNA synthetase can be from a second species, e.g., Methanosarcina barkeri. When an orthogonal pair includes a tRNA and a tRNA synthetase from different species, it is necessarily with the proviso that the orthogonal pair will function together effectively, i.e., the tRNA synthetase will effectively aminoacylate the tRNA of the amino acid of interest. Most preferably, an orthogonal pair comprises a tRNA and a tRNA synthetase from the same species. The properties of tRNA synthetases and specific mutations that benefit their activity are discussed separately below.

[0224] If the charge / acylation portion of the tRNA synthetase molecule is based on or derived from a Pyl tRNA synthetase, a chimeric tRNA synthetase can be produced. That is, the anticodon portion of the tRNA molecule can vary, depending on, for example, the operator's choice to direct the tRNA in recognizing an alternative codon, such as a sense codon, a quadruplet codon, an amber codon, or another "stop" codon. However, the functional acylation / charge portion of the tRNA molecule should be conserved in order to preserve the useful charging activity with respect to the unnatural amino acids discussed above.

[0225] Both Methanosarcina barkeri and Methanosarcina mazei tRNAs are suitable. In any case, these tRNAs differ by only one nucleotide. This difference of one nucleotide does not affect their activity. Therefore, both tRNAs are equally applicable in the present invention.

[0226] The tRNA used may be varied, for example mutated. However, any such variant or mutant of a Pyl tRNA should still retain the ability to productively interact with the tRNA synthetase used to charge the tRNA with said unnatural amino acid.

[0227] tRNA synthetase Pyrrolysine-tRNA synthetases from the species Methanosarcina barkeri and Methanosarcina mazei are suitable if they contain the mutations described herein that serve to charge the tRNA with the unnatural amino acid described above.

[0228] Incorporation of unnatural amino acids via peptide bonds It will be clear that the direct product of incorporating an unnatural amino acid as described herein into a polypeptide is a polypeptide comprising a residue of said unnatural amino acid bound to the polypeptide backbone by a peptide bond. This means that in the strict sense, the polypeptide does not exactly comprise the referenced unnatural amino acid, but rather comprises its residue that has undergone a condensation reaction. This is because the amino acid group of the referenced unnatural amino acid reacts with its adjacent amino acid residue in the polypeptide chain, resulting in the incorporation of one residue of the referenced unnatural amino acid by way of a peptide bond and the liberation of one molecule of HO. References to "incorporation of unnatural amino acid X into a polypeptide chain" or "polypeptide comprising unnatural amino acid X" should be interpreted accordingly. This is entirely conventional nomenclature in the art, and for example, when an amino acid, e.g. valine, is incorporated into a polypeptide chain, the polypeptide chain is described as comprising valine, whereas in reality the polypeptide chain comprises the amino acid residue of valine bound to the polypeptide chain by reaction of its amino acid group and the formation of a peptide bond and the liberation of one molecule of HO, as described above.

[0229] In one aspect, the invention relates to polypeptides that include the unnatural amino acids described above.

[0230] In one aspect, the invention relates to a polypeptide comprising an unnatural amino acid as described above, wherein the unnatural amino acid is attached to the polypeptide via a peptide bond.

[0231] In one aspect, the invention relates to a polypeptide comprising an unnatural amino acid as described above, wherein the unnatural amino acid is attached to the polypeptide via a peptide bond.

[0232] In one aspect, the invention relates to polypeptides that include the aforementioned unnatural amino acids incorporated by or through peptide bonds.

[0233] In one aspect, the invention relates to polypeptides that include residues of the aforementioned unnatural amino acids incorporated by or through peptide bonds.

[0234] As used herein, "unnatural amino acid" refers to an amino acid that is not naturally encoded or found in the genetic code of any organism. Thus, unnatural amino acids are compounds that contain amine and carboxylic acid functional groups and side chains, but are not any of the proteinogenic amino acids used by the translation machinery to assemble proteins.

[0235] Similarly, exemplary unnatural amino acids disclosed herein may be referred to as "unnatural amino acids that contain DAP" or "amino acids that contain DAP," etc. Obviously, in the strictest sense of the IUPAC nomenclature rules, DAP has two NH groups, whereas the unnatural amino acids discussed have a single NH group, with the other NH group missing an H atom and instead bonded to the remainder of the amino acid side chain through the corresponding N atom. Notwithstanding this, "unnatural amino acids that contain DAP" may be referred to as "unnatural amino acids that contain DAP." References to "unnatural amino acids containing" or "amino acids containing DAP" are used herein as is common in the art and may refer to one or any of compounds 2, 3, 4, 5, and 6, each of which contains a HN-C-(COOH)-CN(H)-R moiety (i.e., contains "DAP").

[0236] Host cells, vectors and protein production For the aforementioned methods, the polynucleotide encoding the polypeptide of interest can be incorporated into a recombinant replicable vector. The vector can be used to replicate the nucleic acid in a compatible host. Thus, in a further embodiment, the present invention provides a method for making the polynucleotide of the present invention by introducing the polynucleotide of the present invention into a replicable vector, introducing the vector into a compatible host cell, and growing the host cell under conditions that allow replication of the vector. The vector can be recovered from the host cell. Suitable host cells include bacteria, such as E. coli.

[0237] Preferably, the polynucleotide of the invention is operably linked to a control sequence capable of enabling the expression of the coding sequence by the host cell, i.e., the vector is an expression vector. The term "operably linked" means that the described components are in a relationship permitting them to function in the intended manner. A regulatory sequence "operably linked" to a coding sequence is ligated such that expression of the coding sequence is achieved under conditions compatible with the control sequences.

[0238] The vectors of the invention can be transformed or transfected into suitable host cells as described, allowing the expression of the proteins of the invention. This process comprises the steps of culturing the host cells transformed with said expression vectors under conditions allowing the expression by the vector of the coding sequence encoding the protein, and the optional step of recovering the expressed protein.

[0239] The vector may be, for example, a plasmid or virus vector provided with an origin of replication, optionally a promoter for the expression of the polynucleotide and optionally a regulator of the promoter. The vector may contain one or more selectable marker genes, for example, the ampicillin resistance gene in the case of a bacterial plasmid. The vector may be used, for example, to transfect or transform a host cell.

[0240] Control sequences operably linked to the sequence encoding the protein of the invention include promoters / enhancers and other expression regulation signals. These control sequences may be selected to be compatible with the host cell for which the expression vector is designed to be used. The term promoter is well known in the art and encompasses nucleic acid regions of various sizes and complexity, ranging from minimal promoters to promoters and enhancers that include upstream elements.

[0241] Another aspect of the invention is a method, e.g., an in vitro method, for genetically and site-specifically incorporating an unnatural amino acid, including DAP, into a selected protein, preferably in a cell. One advantage of genetically incorporating by said method is that, in this embodiment, the protein containing DAP can be synthesized directly in the target cell, obviating the need to deliver the protein containing DAP to the cell once it is formed. The method comprises the following steps: (i) introducing an orthogonal codon, e.g., an amber codon, or replacing a particular codon with an orthogonal codon, e.g., an amber codon, at a desired site in a nucleotide sequence encoding a protein; (ii) Orthogonal tRNA synthetase / tRNA pairs, e.g., DAPRS tRNA synthetase introducing into the cell an expression system for a tRNA / tRNA pair; growing the cells in a medium containing an unnatural amino acid, including DAP, according to the invention; Includes.

[0242] Step (i) involves replacing a particular codon at a desired site in the gene sequence of the protein with an orthogonal codon, e.g., an amber codon. This can be accomplished by simply introducing a construct, e.g., a plasmid, having a nucleotide sequence encoding the protein, where the site at which it is desired to introduce or replace an unnatural amino acid, including DAP, is altered to contain an orthogonal codon, e.g., an amber codon. This is well within the capabilities of one of skill in the art, and examples are provided herein.

[0243] Step (ii) requires an orthogonal expression system that specifically incorporates the unnatural amino acid with DAP at the desired location (e.g., an amber codon). Thus, a specific orthogonal tRNA synthetase, e.g., a DAPRS-tRNA synthetase and a specific corresponding orthogonal tRNA pair, that can together charge the tRNA with the unnatural amino acid with DAP, are required. Examples of these are provided herein.

[0244] Protein expression and purification Host cells containing a polynucleotide of the present invention may be used for expressing a protein of the present invention. Suitable host cells include bacteria, such as E. coli, or certain eukaryotic cells. With regard to eukaryotic cells, many UAAs using the PylS / PylT system have been shown to function in eukaryotic cells (the most recent example is for photocaged cysteine ​​(Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society). Society 136, 2240-2243 (2014); which is incorporated herein by reference for its teaching of manipulations in eukaryotic cells.

[0245] The eukaryotic cell can be any suitable eukaryotic cell, for example, an insect cell (eg, an Sf9 insect cell), a mammalian cell, for example a mouse cell or a human cell.

[0246] Preferably, the eukaryotic cell is a mammalian cell. Preferably, the mammalian cell is a HEK293 cell, such as a HEK293T cell. In one embodiment, preferably the eukaryotic cell, mammalian cell, HEK293 cell or HEK293T cell is in vitro. In this embodiment, the cell is not an in vivo cell.

[0247] Preferably, the host cell is a bacterial cell, preferably the host cell is E. coli. Preferably, the host cell is an E. coli cell. Preferably, the E. coli cell is a BL21 DE3 E. coli cell.

[0248] In one aspect, the present invention relates to a method for producing a polypeptide comprising 2,3-diaminopropionic acid (DAP) as described above, which is carried out inside a living cell. Suitably, the method comprises the steps of genetically incorporating said unnatural amino acid into a polypeptide and the optional step of deprotecting said unnatural amino acid to 2,3-diaminopropionic acid (DAP). Suitably, the live cells comprise E. coli cells, for example BL21 DE3 E. coli cells. Suitably, the live cells comprise mammalian cells, for example HEK293T cells.

[0249] The host cells may be cultured under suitable conditions that allow expression of the proteins of the invention. Expression of the proteins of the invention may be constitutive, so that they are produced continuously, or inducible, so that a stimulus is required to initiate expression. In the case of inducible expression, protein production may be initiated, when required, for example, by the addition of an inducer substance, such as dexamethasone or IPTG, to the medium.

[0250] Human Embryonic Kidney with large T antigen (HEK293T) cells are widely available, for example, ATCC® CRL-3 216 is LGC Standards, Queens Road, Teddington, Midd. lesex, TW11 0LY, UK. HEK293T cells may be cultured as known in the art, for example, in Dulbecco's modified eagle medium (DMEM) containing appropriate supplements as required.

[0251] BL21 DE3 E. coli cells are widely available, for example, C2527I or C2527H are available from New England Biolabs, 240 County Road, New England. Ad, Ipswich, MA 01938-2732, USA. They can be cultured according to instructions from the supplier, as is well known in the art.

[0252] Proteins of the invention can be extracted from host cells by a variety of techniques known in the art, including enzymatic, chemical and / or osmotic lysis and physical disruption.

[0253] The proteins of the invention can be purified by standard techniques known in the art, such as preparative chromatography, affinity purification, or any other suitable technique.

[0254] Integration target site Preferably, the DAP and / or unnatural amino acid is incorporated at a position corresponding to a cysteine, serine, or threonine residue in the wild-type polypeptide, and may be incorporated at a position corresponding to a cysteine ​​or serine residue in the wild-type polypeptide, and is most preferably incorporated at a position corresponding to a cysteine ​​residue in the wild-type polypeptide. More preferably, the DAP and / or unnatural amino acid is incorporated at a position corresponding to a catalytic cysteine, catalytic serine or catalytic threonine residue of the wild-type polypeptide, and may be incorporated at a position corresponding to a catalytic cysteine ​​or catalytic serine residue of the wild-type polypeptide, and most preferably is incorporated at a position corresponding to a catalytic cysteine ​​residue of the wild-type polypeptide.

[0255] enzyme The present invention finds particular use in modifying the active site of an enzyme by incorporation of a DAP. Preferably, a DAP or unnatural amino acid described herein is incorporated at a position corresponding to an amino acid in the active site of a wild-type enzyme. Preferably, a DAP or unnatural amino acid described herein is incorporated at a position corresponding to a catalytic amino acid in the active site of a wild-type enzyme.

[0256] Preferably, the polypeptide is an enzyme. Preferably, it is an enzyme that generates an ester or thioester intermediate. Preferably, the polypeptide of the invention is an enzyme that follows this general mechanism; more preferably, the polypeptide may be a serine hydrolase (encoded by 1% of the genes in the human genome); preferably, the enzyme is a protease, peptidase, amidase, deubiquitinase, lipase, cholinesterase, thioesterase, phospholipase, glycan hydrolase (Di Cera, E. Serine proteases. IUBMB Life 61, 510-515 (2009); Hedstrom, L. Serine protease mechanism and specificity. Chem Rev 102, 4501-4523 (2002); Long, JZ & Cravatt, BF The Metabolic Serine Hydrolases and Their Functions in Mammalian Physiology and Disease. Chem Rev 111, 602 2-6063 (2011)), cysteine ​​proteases (e.g., caspases), or ubiquitination and / or enzymes involved in SUMOylation (e.g., E1, E2 or certain E3 families) (Verma, S., Dixit, R. & Pandey, KC Cysteine ​​Proteases: Modes of Activation and Future Prospects as Pharmacological Targets. Front Pharmacol 7 (2016); Otto, HH & Schirmeister, T. Cysteine ​​proteases and their inhibitors. Chem Rev 97, 133-171 (1997); Swatek, KN & Komander, D. Ubiquitin modifications. Cell Res 26, 399-422 (2016)).

[0257] More preferably, the DAP or unnatural amino acids described herein are incorporated by an enzyme, such as a protease.

[0258] International Nomenclature and Classification of Enzymes System (International Biochemistry Preferably, the DAP or unnatural amino acid described herein is a member of the taxonomic group:

[0259] [Table 1]

[0260] The enzyme is incorporated into one or more enzymes from.

[0261] Industrial Applications The present invention provides, inter alia, a new method for producing polypeptides containing 2,3-diaminopropionic acid (DAP), which has a variety of industrial applications, for example, exploiting its reactivity to covalently capture molecules of interest, e.g., substrates of enzymes being studied. This also allows for the dissection of metabolic pathways through a similar approach, made possible by the incorporation of DAP into the core of enzyme active sites. Applications also include the study / capture of small molecule drugs to identify how they are modified / metabolized by enzymes and / or to identify their protein targets.

[0262] Structural characterization of proteins is often the starting point for many drug discovery projects. DAP has broad applicability since it can be used for any protein-catalyzed reaction that proceeds through a covalent intermediate attached to a serine or cysteine ​​side chain in the enzyme active site. DAP can be used to structurally characterize these covalent intermediates. In addition, DAP also allows the identification of new substrates for these enzymes, which can also be the starting point for new drug development strategies.

[0263] Further uses Our technique allows the site-specific incorporation of DAP into recombinant proteins. DAP can be used for any protein-catalyzed reaction that proceeds through a covalent intermediate attached to a serine or cysteine ​​side chain in the enzyme active site. DAP can be used to structurally characterize these covalent intermediates, but also allows the identification of new substrates for these enzymes.

[0264] We have identified an aminoacyl-tRNA synthetase / tRNA that incorporates an amino acid (DAP5) that can be post-translationally converted to DAP under mild conditions. CUAWe evolved a pair that allows us to site-specifically incorporate DAP into recombinant proteins produced in E. coli, which allows us to efficiently capture the acyl-enzyme intermediate.

[0265] In one embodiment, the invention provides an amino acid (DAP5) that is incorporated into a protein using a mutant of pyrrole lysine tRNA synthetase (DAPRS) that has been evolved to load DAP5 onto a cognate tRNA, allowing the synthesis of an integrated protein with a site-specifically incorporated DAP5 that can be deprotected to DAP under mild conditions.

[0266] Preferably, the irradiation for deprotection is for less than 1 minute, preferably about 1 minute, preferably 1 minute, preferably at least 1 minute. Preferably, the irradiation for deprotection is for 1 millisecond to 120 seconds, more preferably 1 millisecond to 60 seconds. Most preferably, the irradiation for deprotection is for 1 minutes. Suitably, the protein is then incubated at any temperature above freezing after irradiation. Incubation is the second step of deprotection. The time required to complete varies depending on the protein incorporating DAP. In the examples, samples were incubated for 1-2 h after irradiation.

[0267] In one aspect, the present invention relates to a method for capturing a substrate for an enzyme.

[0268] Preferred substrates are those in which a thioester or ester intermediate bond is created by the action of an enzyme on said substrate (Note: the substrate does not contain an ester or thioester - preferably the substrate has a chemical structure such that these moieties are created during catalysis). Preferably these bonds are created following or as a result of a nucleophilic attack of the enzyme on its substrate. In the case of replacement of active site residues by DAP according to the present invention, the bonds created are aryl, ... It is a mide bond (discussed in more detail below). Preferably, the substrate is captured as a stable amide analogue. Suitably, the substrate may be an analogue of a naturally occurring substrate. Preferably, the enzyme is a cysteine ​​protease or a thioesterase. Preferably, the enzyme acts via one or more acyl-enzyme intermediates, more preferably via one or more cysteine- or serine-linked acyl-enzyme intermediates.

[0269] In one embodiment, the present invention can be used to study intermediate enzyme-substrate complexes formed upon addition of small molecules (drugs). This allows elucidation of how drugs disrupt this enzyme-substrate interaction during various complex formations. Thus, where a small molecule is a substrate for an enzyme, this is a useful application of the present invention. DAP captures any substrate, provided it replaces the nucleophilic (catalytic) residues, preferably cysteine ​​and / or serine.

[0270] The present invention finds use in exploring insight into the biosynthetic acyl-enzyme intermediates via the encoded 2,3-diaminopropionic acid.

[0271] Further particular and preferred aspects are set out in the accompanying independent and dependent claims. Features from the dependent claims may be combined with features of the independent claims as appropriate and in combinations other than those explicitly set out in the claims.

[0272] Where an arrangement of an apparatus is described as operable to provide a function, this will be understood to include an arrangement of an apparatus that provides that function or is adapted or configured to provide that function. [Example]

[0273] Genetic encoding of DAP derivatives in recombinant proteins EXAMPLES

[0274] We reasoned that the structural similarity of DAP to cysteine ​​and serine, which are constitutively present in cells, suggested that it might be difficult to find an aminoacyl-tRNA synthetase that selectively incorporates DAP. We therefore designed and synthesized four protected versions of DAP (compounds 2–5) (Figure 3a, Supplementary Figure 2). We predicted that the discovery of aminoacyl-tRNA synthetase-tRNACUA pairs for these amino acids would allow site-specific incorporation into proteins, and that post-translational deprotection of the encoded amino acid via transition metal catalysis (Li, J. et al. Palladium-triggered deprotection chemistry for protein activation in living cells. Nat Chem 6, 352-361 (2014)) (2) or light (Baker, AS & Deiters, A. Optical Control of Protein Function through Unnatural Amino Acid Mutagenesis and Other Optogenetic Approaches. Acs Chemical Biology 9, 1398-1407 (2014)) (3-5) would expose the DAP.

[0275] Compounds 3 and 4 were synthesized by the present inventors using the PCC1RS / tRNA CUA and PCC2RS / tRNA CUA These pairs include a conservative SH-NH2 substitution for a photocaged cysteine ​​derivative that was previously incorporated into the protein using the pyrrolysyl-tRNA synthetase / tRNA pair (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014)). CUAThis similarity suggested that these pairs might direct the incorporation of 3 or 4 in response to the amber codon. However, we did not inject either 3 or 4 and their When related pairs were provided, they were found to not function to suppress the amber codon in the reporter gene.

[0276] As part of an effort to discover an orthogonal aminoacyl-tRNA synthetase that incorporates amino acids 2-5, we identified MbPylRS / tRNA CUA Five paired variant libraries (Susan1, Susan2 (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014)), Susan4, PylS fwd (Nguyen, DP, Elliott, T., Holt, M., Muir, TW & Chin, JW Genetically Encoded 1,2-Aminothiols Facilitate Rapid and Site-Specific Protein Labeling via a Bio-orthogonal Cyanobenzothiazole Condensation. Journal of the American Chemical Society 133, 11418-11421 (2011)) and D3 (Supplementary Table 1)) were examined. These libraries randomize residues in the active site of the synthetase and were previously generated to allow for the incorporation of ncAAs. The Susan2 library was previously used to discover synthetases for photocaged derivatives of cysteine ​​(Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014)). We subjected each library to two consecutive rounds of positive selection in the presence of each ncAA and one round of negative selection in the absence of the ncAA (Neumann, H., Peak-Chew, SY & Chin, JW Genetically encoding N-epsilon-acetyllysine in recombinant proteins. Nature Chemical Biology 4, 232-234 (2008); Chin, JW Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu Rev Biochem 83, 379-408 (2014); Liu, CC & Schultz, PG Adding New Chemistries to the Genetic Code. Annual Review of Biochemistry, Vol 79 79, 413-444 (2010)). However, we did not find any compounds 2- No synthetase / tRNA pair for 5 was found.

[0277] Examination of the logP predicted values ​​of compounds 2-5 revealed that they were highly hydrophilic, which motivated us to examine whether they could enter cells. Using an LC-MS based amino acid uptake assay (Zhang, MS et al. Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat Methods 14, 729-736 (2017)), we found that the intracellular We did not detect substantial amounts of amino acids 2-5 in E. coli. Our data suggest that the intracellular concentrations of compounds 2-5 are substantially less than 10 μM (FIG. 3b-e). These observations suggest that compounds 2-5 are not efficiently taken up or metabolized by E. coli. This may explain why it was not possible to select synthetases for incorporation of these amino acids in vivo. EXAMPLES

[0278] Preferred Unnatural Amino Acids, Including DAP To address the challenge of encoding a protected version of DAP, we designed and synthesized amino acid 6, which we predict can be post-translationally deprotected to expose a side chain amine (Supplementary Figure 2). This amino acid has a more favorable logP prediction value than amino acids 1-5, and we found that addition of 1 mM of 6 to the cell medium resulted in an intracellular concentration of approximately 2 mM (Figure 3f). Thus, unlike amino acids 2-5, amino acid 6 can accumulate in E. coli at millimolar concentrations. EXAMPLES

[0279] Generation of DAP tRNA synthetase ("DAPRS") Given the failure of many previous approaches, the inventors have determined that the positions to be randomized are Py We devised and created an entirely new library, DAPRSlib, that was carefully selected based on a model of 6 in the active site of lRS (Supplementary Figure 3 and Supplementary Table 1). As a result of this ingenuity, five positions (Y271; N311; Y349; V366; W382) were randomized to all 20 canonical amino acids. 3.4×10 7 We determined that this would result in a theoretical library diversity of 112 different sequences. We subjected DAPRSlib to three rounds of sequential positive and negative selection (Neumann, H., Peak-Chew, SY & Chin, JW Genetically encoding N-epsilon-acetyllysine in recombinant proteins. Nature Chemical Biology 4, 232-234 (2008)) in the presence and absence of 6. Following this selection, we screened 96 clones in cells containing the cat(112TAG) reporter. We obtained a single clone that conferred high-level chloramphenicol resistance in the presence of 6 and minimal chloramphenicol resistance in the absence of 6 (FIG. 3g). The selected synthetase contains four active site mutations with respect to MbPylRS (Y271C, N311Q, Y349F, and V366C). (Note that although the library contained five randomized positions, only four mutations are present in DAPRS - the fifth position is wild type (i.e., W382) in DAPRS.) EXAMPLES

[0280] Incorporation of Unnatural Amino Acids, Including DAP, into Polypeptides To further characterize the genetically directed site-specific incorporation of 6 into proteins, a superfolder green fluorescent protein (sfGFP) containing an amber stop codon (TAG) at position 150 (sfGFP(150TAG)His6) was engineered to incorporate DAPRS / tRNA Pyl CUA and 1 mM 6 and purified by Ni-NTA affinity chromatography (Fig. 3h and Supplementary Fig. 4a). As a benchmark, the same genes were expressed in the presence of PylRS / tRNA Pyl CUA and 1 mM, which are known to provide efficient amber suppression (Virdee, S., Ye, Y., Nguyen, DP, Komander, D. & Chin, JW Engineered diubiquitin synthesis reveals Lys29-isopeptide specificity of an OTU deubiquitinase. Nature Chemical Biology 6, 750-757 (2010)). ε -tert-Butyl Expression was carried out in the presence of boronyl-lysine (BocK). The yield of GFP incorporating 6 (GFP(6)) was comparable to that of BocK incorporation, which is consistent with the DAPRS / tRNA Pyl CUA This shows the efficiency of the pair. Electrospray ionization mass spectrometry (ESI-MS) was used to identify DAPRS / tRNA Pyl CUA The pair is confirmed to direct the incorporation of 6 into the protein in response to an amber codon (Fig. 3i, j). EXAMPLES

[0281] Deprotection of DAP-containing amino acids Here, we demonstrate that deprotection of an amino acid containing DAP leaves the DAP group in the polypeptide backbone, thereby resulting in a polypeptide containing DAP.

[0282] We predicted that irradiation of a protein incorporating amino acid 6 with 365 nm light would reveal a sulfhydryl-containing intermediate that could undergo further reactions to give DAP (either by 5-exo-trig cyclization of the sulfhydryl group to the carbonyl of the carbamate and collapse of the resulting tetrahedral intermediate to release an amino group, or by forming an episulfide and carbon dioxide to release an amino group). Indeed, irradiation of GFP (6) (365 nm, 35 mW cm -2 , 1 min) resulted in the complete deprotection of 6 to the expected sulfhydryl (Fig. ​(Fig.3i,j).3i,j). Subsequent incubation of the protein at 37 °C resulted in the complete deprotection of the desired amino group, revealing amino acid 1 in GFP (Fig.3i,j). EXAMPLES

[0283] Stable trapping of cysteine ​​protease acyl-enzyme intermediates Cysteine ​​proteases, such as tobacco etch virus (TEV) protease, typically contain a Cys-His-Asp catalytic triad and generate thioester intermediates upon treatment with their cognate substrates (Verma, S., Dixit, R. & Pandey, KC Cysteine ​​Proteases: Modes of Activation and Future Prospects as Pharmacological Targets. Front Pharmacol 7 (2016); Phan, J. et al. Structural basis for the substrate specificity of tobacco etch virus protease. Journal of Biological Chemistry 277, 50564-50572 (2002)). Thus, we have demonstrated that the activity of TEV protease is a key step in the regulation of TEV signaling. The goal was to replace the site cysteine ​​with DAP, aiming to trap the acyl-enzyme intermediate. E. coli expressing His6-lipoyl-TEV(151TAG)-Strep and DAPRS / tRNA were given 0.1 mM 6. Pyl CUA Using a pair, TEV replaces the catalytic cysteine ​​in the active site with 6 Cys1516 The protease was produced. The protein was purified by tandem affinity chromatography at a yield of approximately 0.1 mg per liter of culture, and incorporation of 6 at the genetically encoded site was confirmed by ESI-MS. Photodeprotection quantitatively converted the encoded 6 to a sulfhydryl intermediate, and approximately 70% of the protein was subsequently fully deprotected and purified by TEV (TEV) as judged by ESI-MS. DAP ) amino acid 1 appeared at position 151 (Supplementary Figure 5).

[0284] To demonstrate that replacement of the catalytic cysteine ​​with DAP allows for the capture of a covalent protease-substrate intermediate, we used TEV DAP was incubated with a model substrate, Ub-tev-His6 (in which a TEV cleavage site (tev) is flanked by ubiquitin and a hexahistidine tag) and protein species were separated by SDS-PAGE. DAP We observed the formation of a new band migrating more slowly than Ub-tev-His6 and did not observe free ubiquitin that would have been generated by cleavage of the tev site (Fig. 4a and Supplementary Fig. 4). Western blots show that the new band contains both TEV and ubiquitin (Fig. 4a). Control experiments confirm that wild-type TEV cleaves Ub-tev-His6 to the more rapidly migrating Ub, and that TEV(C151A) does not cleave Ub-tev-His6 (Fig. 4a). These experiments demonstrate that the Cys151DAP substitution is essential for the formation of the slower migrating band containing TEV and Ub, and that Ub is essential for the formation of the slower migrating band containing TEV and Ub, and that the Cys151DAP substitution ... DAP Trypsin MS / MS of the slower migrating band identified an isopeptide bond between DAP and Ub, which is responsible for the release of TEV from the mutant.DAP The formation of -Ub is confirmed (Figure 4b). Thus, replacement of the catalytic cysteine ​​of TEV with DAP allows the creation of a protease that undergoes the first step of the protease cycle: nucleophilic attack on the substrate carbonyl to form a first tetrahedral intermediate. This intermediate collapses, releasing the C-terminal fragment of the substrate and leaving the N-terminal fragment of the substrate covalently bound to the protease by a stable amide bond that is not susceptible to hydrolysis. EXAMPLES

[0285] Activity and synthesis pathway of Vlm TE To gain insight into TE domain function and to prepare it for use in the DAP integration system, we cloned and expressed Vlm TE and synthesized the resulting Vlm TE (TE wt , wild-type) proteins were purified for biochemical and structural studies. The native substrate of the TE domain is a peptide intermediate linked by a thioester to phosphopantetheine-PCP (peptidyl-PCP). Gramicidin S, surfactin and fengycin synthetase (Hoyer, KM, Mahlert, C. & Marahiel, MA The iterative gramicidin s thioesterase catalyzes peptide ligation and cyclization. Chemistry & biology 14, 13-22 (2007), Samel, SA, Wagner, B., Marahiel, MA & Essen, LO The thioesterase domain of the fengycin biosynthesis cluster: a structural base for the macrocyclization of a non-ribosomal lipopeptide. Journal In NRPS TE domains, including those from Surfactin synthetase C-terminal thioesterase domain as a cyclic depsipeptide synthase. Biochemistry 41, 13350-13359 (2002)), PCP-linked substrates can be mimicked by small molecules in which a peptide intermediate is linked to N-acetylcysteine ​​by a thioester (peptidyl-SNAC). We synthesized a SNAC derivative of the native peptide, D-hiv-D-val-L-lac-L-val-SNAC (tetradepsipeptidyl-SNAC 7, Supplementary Figures 6 and 7), and demonstrated that Vlm TE wt We found that its incubation with Vlm TE resulted in the production of valinomycin (Fig. 5, Supplementary Figs. 8, 9 and Supplementary Table 2). wt demonstrate that tetradepsipeptidyl-SNAC 7 can be used to complete all steps of the catalytic cycle: oligomerization of the tetradepsipeptide intermediate to an octadepsipeptide, oligomerization of the octadepsipeptide to a dodecadepsipeptide, and cyclization of the dodecadepsipeptide to release valinomycin.

[0286] Vlm TE differentiates between two possible pathways from synthetic intermediates detected in valinomycin synthesis wtThis reveals an oligomerization pathway catalyzed by the thioesterase domain of tyrocidine synthetase (Supplementary Figure 10) (Hoyer, KM, Mahlert, C. & Marahiel, MA The iterative gramicidin s thioesterase catalyzes peptide ligation and cyclization. Chemistry & biology 14, 13-22 (2007); Trauger, JW, Kohli, RM, Mootz, HD, Marahiel, MA & Walsh, CT Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000)). Vlm TE wt is tetradepsipeptidyl-OT The D-hiv-D-val-L-lac-L-val moiety could potentially oligomerize by ester bond formation between the distal hydroxyl of D-hiv in E and the carbonyl of L-val from tetradepsipeptidyl-S-PCP ("forward transfer"), or between the distal hydroxyl of D-hiv in tetradepsipeptidyl-S-PCP and the carbonyl of L-val from tetradepsipeptidyl-O-TE ("reverse transfer", so named because the octadepsipeptide will later be transferred back to the TE domain). LC-MS of the reaction for the synthesis of valinomycin from tetradepsipeptidyl-SNAC 7 showed signals with masses corresponding to octadepsipeptidyl-SNAC 11 and dodecadepsipeptidyl-SNAC 15, intermediates produced only in the "reverse transfer" oligomerization pathway (Supplementary Fig. 7 and Supplementary Fig. 10). Consistently, experiments using a mixture of tetradepsipeptidyl-SNAC 7 and a tetradepsipeptidyl-SNAC lacking the terminal hydroxyl (deoxy-tetradepsipeptidyl-SNAC 8) showed peaks for deoxy-octadepsipeptidyl-SNAC 12 and deoxy-dodepsi-peptidyl-SNAC 16 (Supplementary Fig. 9). All of these tyrocidine synthetase enzymes use a pathway similar to that of tyrocidine S synthetase (Hoyer, KM, Mahlert, C. & Marahiel, MA The iterative gramicidin s thioesterase catalyzes peptide ligation and cyclization. Chemistry & biology 14, 13-22 (2007); Trauger, JW, Kohli, RM, Mootz, HD, Marahiel, MA & Walsh, CT Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000)). Oligomerization-cyclization of NRPS (or PKS (Zhou, Y., Prediger, P., Dias, LC, Murphy, AC & Leadlay, PF Macrodiolide formation by the thioesterase of a modular polyketide synthase. Angew Chem Int Ed Engl 54, 5232-5235 (2015))) Finally, the valinomycin synthesis assay also showed small peaks corresponding to 16-mer depsipeptidyl-SNAC 19, 20-mer depsipeptidyl-SNAC 23, and cyclic 16-mer depsipeptide 29, which suggests that Vlm The TE domain indicates that the final product is somewhat more flexible than previously thought. (Figure 5 and Supplementary Figure 8). EXAMPLES

[0287] Visualization of a key intermediate in TE domain-mediated valinomycin synthesis Next, the inventors wtWe obtained and optimized crystallization conditions for robust and repeatable growth of Vlm TE and determined its structure (Figure 6, Supplementary Figure 11, and Supplementary Table 3). The Vlm TE adopts an α / β hydrolase fold typical of type I TE domains, with the canonical Ser-His-Asp catalytic triad at Ser2463, His2625, and Asp2490 (Horsman, ME, Hari, TPA & Boddy, CN Polyketide synthase and non-ribosomal peptide synthetase thioesterase selectivity: Logic gate or a victim of fate? Natrual Products Reports (2015)) covered by the TE “lid.” The lid is a structural element known to be mobile and has been suggested to play various roles in TE domain function, ranging from substrate positioning to solvent exclusion (Samel, SA, Wagner, B., Marahiel, MA & Essen, LO The thioesterase domain of the fengycin biosynthesis cluster: a structural base for the macrocyclization of a non-ribosomal lipopeptide. Journal of molecular biology 359, 876-889 (2006); Frueh, DP et al. Dynamic thiolation-thioesterase structure of a non-ribosomal peptide synthetase. Nature 454, 903-906 (2008); Whicher, JR et al. Structure and function of the RedJ protein, a thioesterase from the prodiginine biosynthetic pathway in Streptomyces coelicolor. J Biol Chem 286, 22558-22569 (2011)).Although composition can vary substantially, a typical lid is composed of about 50 residues and about 2-4 helices. The Vlm TE lid region is about 88 residues (about 2494-2582) and is composed of an extended loop, three helices (Lα1-3) seen here as a bundle, a short five-residue helix (Lα4), a long helix (Lα5) and another short helix (Lα6) (Figure 6b). We have previously shown that the lid region of the Vlm TE is composed of about 50 residues and about 2-4 helices (Lα1-3) seen here as a bundle (Figure 6b). wt We obtained two structures of TE, which differ only in the lid region. In one structure, the lid is almost completely ordered, but the B-factor is significantly higher in the region containing Lα1–4, which therefore makes almost no contact with the rest of the domain (Supplementary Fig. 11b). wt In the structure, Lα4–5 have positions similar to those observed in the first structure, but Lα3 is rotated 10° toward the active site and Lα1–2 are too disordered to be modeled.

[0288] Incubation of Vlm TE with depsipeptidyl-SNAC molecules did not yield stable conjugates (Supplementary Fig. 12a–c and Supplementary Table 4), indicating that TE wt Several attempts to soak crystals with depsipeptidyl-SNAC molecules failed to reveal interpretable ligand electron density in the active site and conformational changes in the vicinity. Other groups have reported similar setbacks when attempting to visualize acyl-enzyme complexes from SNAC molecules (Supplementary Table 6). (Samel, SA, Wagner, B., Marahiel, MA & Essen, LO The thioesterase domain of the fengycin biosynthesis cluster: a structural basis for the macrocyclization of a non-ribosomal lipopeptide. Journal of molecular biology 359, 876-889 (2006); Bruner, SD et al. Structural basis for the cyclization of (2002)). We conclude that the acyl-intermediate in the Vlm TE-mediated synthesis of valinomycin is rapidly hydrolyzed and therefore, as expected, it is extremely difficult to visualize the biosynthetic intermediate by crystallography using wild-type Vlm TE.

[0289] To enable visualization of the Vlm TE acyl-enzyme complex, we determined that serine 2463 in the active site is a DAP (TE DAP ) was replaced by DAPRS / tRNA. CUA Expression of Vlm2TE(2463TAG) in E. coli containing the pair and supplemented with 0.1 mM 6 resulted in replacement of serine 2463 by 6. The resulting Vlm TE could be purified with a yield of about 0.1–0.5 mg per liter of culture. DAP was quantitatively produced (Fig. 6c and Supplementary Fig. 13).

[0290] To provide insight into the first acyl-TE intermediate in the catalytic cycle of Vlm TE, we synthesized a tetradepsipeptidyl-N-TE DAP The conjugate was captured. DAP When incubated with tetradepsipeptidyl-SNAC 7, stable depsipeptidyl-TE DAP The intermediate was produced in over 60% yield (Supplementary Fig. 12d and Supplementary Table 4), and we did not observe valinomycin synthesis. Notably, however, small amounts of octadepsipeptidyl-SNAC 11 were observed (Fig. 5b and Supplementary Fig. 8b). Octadepsipeptidyl-SNAC 11 was synthesized by the synthesis of tetradepsipeptidyl-N-TE DAP Hydroxyl groups of tetradepsipeptidyl-SNAC 7 against TE DAP It can be formed by catalytic attack, which is DAPhas the ability to successfully catalyze the attack of a hydroxyl on an amide. Since only small amounts of octadepsipeptidyl-SNAC 11 were formed, it is clear that this reaction is much slower than the more isoenergetic ester-ester reaction. The attack of the hydroxyl on the amide is similar to the first reaction used by related serine proteases (Ekici, OD, Paetzel, M. & Dalbey, RE Unconventional serine proteases: variations on the catalytic Ser / His / Asp triad configuration. Protein Sci 17, 2023-2037 (2008)), in which the substrate peptide backbone is cleaved from an ester-linked acyl-enzyme intermediate. However, it is surprising that this TE domain, which did not evolve to perform this reaction, was able to catalyze it.

[0291] The present inventors have discovered that tetradepsipeptidyl-N-TE DAP Conjugation of the hydroxyl groups of tetradepsipeptidyl-SNAC 7 DAP We hypothesized that catalytic attack may be a contributing factor to the observed non-quantitative yield of the conjugate. Therefore, we conjugated deoxy-tetradepsipeptidyl-SNAC 8 to TE DAP Optimized conditions for conjugating deoxy-depsipeptidyl-TE DAP The conjugate (approximately 70%) was produced (Figure 6d and Supplementary Table 4).

[0292] Deoxy-tetradepsipeptidyl-N-TE DAP To determine the structure of the conjugate, we used preformed TE DAPThe crystals were incubated with the deoxy-tetradepsipeptidyl-SNAC 8 substrate analog. The resulting electron density shows some weak but clear density for the amide bond between residues DAP2463 and L-val4 of the deoxy-tetradepsipeptide (Figure 6f,h). The carbonyl oxygen at L-val4 is close to the main chain amides of residues Ala2399 and Leu2464, a putative oxyanion hole (Tseng, CC et al. Characterization of the surfactin synthetase C-terminal thioesterase domain as a cyclic depsipeptide synthase. Biochemistry 41, 13350-13359 (2002)). There is also density for the next residue, L-lac3, but no deoxyanion bond. The si-tetradepsipeptide is arched, indicating flexibility that is insufficient to reliably model the neighboring D-val2 and D-hiv1. The deoxy-tetradepsipeptide is the first TE. wt It does not interact with the lid, which has a nearly identical conformation to the structure (Supplementary Fig. 11b).

[0293] Next, the present inventors synthesized dodecadepsipeptidyl-N-TE DAP By capturing the conjugate, we focused on gaining insight into the final acyl-TE intermediate in the catalytic cycle of Vlm TE. DAP Incubation of the 1,2-dichloro-N-tetradecane complex with 1,2-dichloro-N-tetradecane affords dodecadepsipeptidyl-N-TE by a reaction similar to the reverse of the cyclization. DAP We reasoned that this conjugate would be thermodynamically favored due to the amide bond and that under optimized conditions, Dodecadepsipeptidyl-N-TE DAP We observed that the conjugates were formed in yields of approximately 65–100% (Figure 6e and Supplementary Table 4). DAP In the crystallization test using TE wtCrystals formed in conditions similar to those described above, but had different morphologies and belonged to two distinct space groups (H3 and P1, with two and six molecules per asymmetric unit, respectively) (Supplementary Table 3).

[0294] Dodecadepsipeptidyl-TE DAP All eight crystallographically independent molecules of α-D1 show some density for the dodecadepsipeptide. Molecules P1_A-F and H3_A-B show strong density for 4, 3, 2, 2, 2, 2, 3, and 1 dodecadepsipeptide residues, respectively (Supplementary Fig. 14). Additional weaker density is present in some molecules that could accommodate up to all 12 residues (Supplementary Fig. 15), and in other molecules, weaker density suggests multiple conformations of distal residues, which could not be explicitly modeled in this density. All modeled depsipeptides follow similar trajectories away from the active site DAP. There are no consistent interactions between the depsipeptides and the TE domain that are out of reach of the L-val residues bound to the DAP (Fig. 6g, i). Rather, each depsipeptide makes different contacts with the lid. The lid forms a hemispherical pocket / steric barrier composed of helices Lα1, 3, 4, and 5 and the N-terminal strand of Lα1. Dodecadepsipeptidyl-TE DAP The lid in each crystallographically independent molecule of is in a similar, but not identical, position, and the loop between the lid helices is disordered in most molecules (Fig. 6k). This again highlights the mobility of the lid and explains why the conformation and degree of order of the dodecadepsipeptidyl-TE differs between molecules (Fig. 6k). A hemispherical barrier is distinct from the dodecadepsipeptidyl-TE with respect to the lid conformation seen in both the apo- and tetradepsipeptidyl-bound structures of Vlm TE. DAP This occurs only due to a significant rearrangement of the lid in the structure (Fig. 6l).

[0295] Comparing the position of the Vlm TE lid in the apo / tetradepsipeptide-bound structure with that in the dodecadepsipeptide-bound structure demonstrates and highlights its extreme mobility: to transition from one lid conformation to the other, helices Lα5-6 maintain their positions, Lα3-4 translocate by approximately 13 Å by rotating 45°, Lα2 translocates by approximately 25 Å, Lα1 shortens, It shifts position by approximately 13 Å and rotates by more than 90° in the opposite direction to Lα3–4 (Figure ​(Figure6l; Supplementary Animations 1, 2). This dramatic rearrangement means that the lid helices of Vlm TE fit together in ways that are significantly different in their apo / tetradepsipeptidyl- and dodecadepsipeptidyl-bound conformations.

[0296] The distinct lid conformations directly affect the possible location of the depsipeptide. In the apo / tetradepsipeptide-bound conformation of the lid, the C-terminus of helix Lα1 comes within 10 Å of Ser / DAP2463, leading the tetradepsipeptide to extend towards the TE core helix αE. In the dodecadepsipeptide-bound conformation of the lid, the loop adjacent to Lα1 blocks the location occupied by the tetradepsipeptide in the tetradepsipeptide-bound conformation. Furthermore, in the dodecadepsipeptide-bound conformation, the N-terminus of Lα1 forms part of a hemispherical pocket, which may help curl the dodecadepsipeptide back towards Ser / DAP2463 during the cyclization step.

[0297] A model of the crystal structure and structure factors have been deposited in the Protein Data Bank under accession numbers 6ECB, 6ECC, 6ECD, 6ECE and 6ECF.

[0298] method General synthetic procedure. All reagents were purchased from Sigma-Aldrich with the following exceptions: L-lactate was purchased from Fisher Scientific, EDC was purchased from Oakwood Chemicals (Estill, SC), Purchasing was at the highest purity available and used without further purification. Valinomycin was purchased from Sigma-Aldrich and BioShop Canada. All solvents were purchased from Fisher Scientific. All reactions were carried out under an argon atmosphere using dry solvents unless otherwise noted. NMR spectroscopy was performed using 1 For H spectrum, 400 MHz and 13 For the C spectrum, we used the Bruker AVANCE II operating at 100 MHz. 1 For H spectrum, 300 MHz. 13 For the C spectrum, a Bruker AVANCE 300 operating at 75 MHz was used. High-resolution mass spectroscopy (HRMS) was used to I measurements were performed using a Micromass Q-TOF I (John L. Holmes Mass Spectroscopy Facility).

[0299] Abbreviations: M = molar; conc. = concentrated; mol = mole; mmol = millimole; °C = degrees Celsius; eq. = equivalent; h = hours; min = minutes; rt = room temperature; cat. = catalyst; aq. = aqueous; Su = succinimidyl; DIPEA = diisopropylethylamine; atm = atmosphere; Boc = tert-butoxycarbonyl; t Bu = tert-butyl; Et = ethyl; Ph = phenyl; TFA = trifluoroacetic acid; THF = tetrahydrofuran; LC-MS = liquid chromatography-mass spectrometry; ELS = evaporative light scattering.

[0300] Amino acid synthesis A general synthetic scheme for the preparation of amino acids is shown below, with the reagents and conditions as follows:

[0301] [ka]

[0302] Reagents and conditions: (i) 2a (10.0 mmol), HCl (4 M in 1,4-dioxane) (8.0 eq.), EtSiH (2 eq.), rt, 1 h, 90% (2); (ii) 1b (19.2 mmol), 2-nitrobenzyl bromide (1.2 eq.), DIPEA (2.0 eq.), dry THF, rt, 10 h, 67% (3a); (iii) 3a ( 11.6 mmol), HCl (4 M in 1,4-dioxane) (8.62 eq.), Et3SiH (2.7 eq.), rt, 24 h, 99% (3·2HCl); (iv) dropwise addition of 4a (20.0 mmol) (1.563 M), conc. HNO3 (70%) in glacial CH3COOH, 0 °C, 1 h, then 40 °C, 2.5 h, 58% (4b); (v) 4b (209.0 mmol), NaBH4 (0.9 eq. .), portionwise addition (8 × 15 min), rt, CH3OH-C2H5OH-CH2Cl2 (44:29:27), then 4 h (total time = 6 h), rt, 99% (4c); (vi) 4c (75.0 mmol), PBr3 (0.4 eq., added dropwise), dry CH2Cl2, 0 °C, then dry pyridine (cat.), 0 °C, 15 min, then rt, 1.5 h, 89% (4d); (vii) Boc-L-Dap-O t Bu 1b (15.0 mmol), 13 (1.2 eq. ), DIPEA (3.0eq.), dry THF, rt, 64h, 82% (4e); (viii)4e (9.49mmol), TFA (20.64eq.), Et3SiH (6.60eq.), dry CH2Cl2, 73% (4·2CF3COOH);(ix)4c(100.0mmol), Su2O(1.5eq.), DIPEA(3.0eq.), dry CH3CN, rt, 16h, 92%(5a);(x) Method 1: Boc-L-Dap-O tBu 1b (14.05 mmol), 5a (1.2 eq.), DIPEA (3.0 eq.), dry CH2Cl2 / dry THF (2:1), r.t. 20 h, 94% (5c); Method 2: 5a (12.5 mmol), Boc-L-Dap-OH 1a (1.25 eq.), DIPEA (3.0 eq.), dry THF / dry CH3CN (9:1), r.t., 24 h, 99% (5b); (xi) 5b (9.2 mmol), TFA (21.3 eq.), dry CH2Cl2, r.t., 99% (5·CF3COOH); (xii) 4d (21.0 mmol ), 2-mercaptoethanol (1.05 eq.), 1,4-dioxane (degassed), aq. NaOH (0.5 M in degassed H2O, 1.0 eq.), r.t., 12 h, in the dark, argon atm., 95% (6a); (xiii) 6a (61.0 mmol), Su2O (1.4 eq.), dry DIPEA (4.0 eq.), dry CH3CN, r.t., in the dark, argon atm., 14 h, quantitative conversion (6b); (xiv) Method 1: 6b (18.0 mmol) , Boc-L-Dap-O tBu.HCl 1b (1.6 eq.), dry DIPEA (3.0 eq.), dry CH3CN, rt, 10 h, in the dark, argon atm., 94% (6c); Method 2: 6b (60.0 mmol), Boc-L-Dap-OH.HCl 1a (1.1 eq.), dry DIPEA (4.0 eq.), dry CH3CN, rt, 14 h, in the dark, argon atm., 96% (6d); (xv) Method 1: 6c (12.066 mmol), TFA (21.647 eq.), dry Et3SiH (10.378 eq.), dry CH2Cl2, rt, in the dark, 24 h, monitored by LC-MS instrument (reverse phase, H2O-CH3CN as mobile phase), product purified by trituration (CH3OH / Et2O) , 64% (6·TFA); Method 2: 6d (57.626 mmol), TFA (9.065 eq.), dry Et3SiH (2.173 eq.), dry CH2Cl2, rt, dark, 5 h, monitored by LC-MS instrument (reverse phase, H2O-CH3CN as mobile phase), if reaction was incomplete additional TFA (up to 2 eq.) was added and left to stir longer, product was purified by trituration (dry CH3OH / Et2O), 63% (6·TFA).

[0303] (S)-3-{[(allyloxy)carbonyl]amino}-2-aminopropanoic acid (2)

[0304] [ka]

[0305] Boc-Dap(Alloc)-OH 2a (4.325 g, 15.0 mmol, 1.0 eq., purchased from Bachem Ltd.) was placed in a dry 250 mL one-neck round-bottom flask and diluted with HCl (4 M in 1,4-dioxane, 60.0 mL, 240.0 mmol, 16.0 eq.). To this solution was added dry EtSiH (10.0 mL, 62.6075 mmol, 4.174 eq.) at rt, which immediately produced a faint white A white precipitate appeared which gradually increased in intensity with time upon stirring at rt under a nitrogen atmosphere. After 24 h, an intense white precipitate was observed. The reaction was judged complete by LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase, gradient). The mixture was then evaporated under reduced pressure and the product was dissolved in dry CH3OH (100 mL) followed by evaporation to dryness under reduced pressure. This was repeated three times to remove most of the 1,4-dioxane by azeotropic evaporation. The residue was redissolved in dry CH3OH (10 mL) and triturated with dry Et2O (500 mL) to precipitate the product. This was filtered, washed with more Et2O (2 x 125 mL) and dried overnight under high vacuum (<0.1 mbar) to give (S)-3-{[(allyloxy)carbonyl]amino}-2-aminopropanoic acid HCl salt 2 as a bright white powder (3.03 g, 90%); 1 H NMR(400.13MHz,DMSO-d6)δ3.42~3.56(m,2H), 4.48(d,J=5.3Hz,2H), 5.15 ~5.36(m,2H), 5.80~5.98(m,1H), 7.44~7.60(m,1H), 8.24~8.70(wide s,3H); 13 C NMR(100.61MHz,DMSO-d6)δ40.4(CH2), 52.4(CH), 64.7(CH2), 117.2(CH2), 133.4(CH), 156.2(C), 169.1(C); MS(ESI+) m / z(relative intensity) 189[(M+H) + ,100], 134(4), 81(9);HRMS(ESI+)m / z C7H 13 O4N2[M+H] + Calculated value: 189.0870, measured value: 189.0866 (Δ=-2.19 ppm)

[0306] tert-Butyl (S)-2-[(tert-butoxycarbonyl)amino]-3-[(2-nitrobenzyl)amino]propanoate (3a)

[0307] [ka]

[0308] Boc-L-Dap-O t Bu·HCl 1b (5.0 g, 19.206 mmol, 1.0 eq.) was placed in a dry 500 mL two-neck round-bottom flask to which was added dry THF (75 mL) followed by dry DIPEA (6.69 mL, 38.412 mmol, 2.0 eq.). The contents were allowed to stir under argon at 0 °C. 2-Nitrobenzyl bromide (4.979 g, 23.048 mmol, 1.2 eq.) was placed in another dry 250 mL one-neck round-bottom flask and dissolved in dry THF (125 mL), and the resulting solution was then added to Boc-L-Dap-O via cannula over 5 min at 0 °C under a positive pressure of argon. t The mixture was then warmed to rt and allowed to stir at rt for 10 h. The reaction mixture was then concentrated under reduced pressure and the crude mixture was extracted with EtOAc (200 mL) and washed with brine solution (3×250 mL). The organic layer was separated, dried over anhydrous Na2SO4, filtered and evaporated to dryness to give a brown viscous oil. The product was purified by flash chromatography on SiO2 (gradient; eluent: EtOAc / n-hexane=1:9→1:4) to give the desired product, tert-butyl (S)-2-[(tert-butoxycarbonyl)amino]-3-[(2-nitrobenzyl)amino]propanoate 3a as a faint yellow viscous oil (5.05 g, 67%): R f = 0.27 (SiO2 plate, EtOAc / n-hexane = 1:4); 1 H NMR (400.13 MHz, CDCl3 with TMS as internal standard) δ 1.44 (s, 9H), 1.46 (s, 9H), 2.85-3.40 (m, 2H), 4.02 (d, J = 14.5 Hz, 1H), 4.07 (d, J = 14.5 Hz, 1H), 4.22-4.35 (m, 1H), 5.20-5.55 (m, 1H), 7.41 (ddd, J = 8.4, 8.4, 2.0 Hz, 1H), 7.50-7.67 (m, 2H), 7.94 (d, J = 8.0 Hz, 1H); 13C NMR (100.61 MHz, CDCl3 with TMS as internal standard) δ 28.1 (CH3), 28.5 (CH3), 50.7 (CH2), 50.9 (CH2), 54.4 (CH), 79.9 (C), 82.3 (C), 124.9 (CH), 128.2 (CH), 131.3 (CH), 133.3 (CH), 135.5 (C), 149.2 (C), 155.7 (C), 170.9 (C).

[0309] (S)-2-Amino-3-[(2-nitrobenzyl)amino]propanoic acid 3

[0310] [ka]

[0311] A dry 100 mL single-neck round-bottom flask was charged with tert-butyl (S)-2-[(tert-butoxycarbonyl)amino]-3-[(2-nitrobenzyl)amino]propanoate 3a (4.59 g, 11.607 mmol, 1.0 eq.). HCl (25 mL, 4 M in 1,4-dioxane, 100.0 mmol, 8.615 eq.) was added followed by dry Et3SiH (5.0 mL, 31.304 mmol, 2.697 eq.). The contents were stirred in the dark under an argon atmosphere. The progress of the reaction was periodically monitored by TLC analysis (SiO2 plate, EtOAc / n-hexane=3:7). After 48 h, an intense white precipitate had formed and the reaction was judged complete by TLC and LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase, gradient). The contents were evaporated to dryness under reduced pressure. Residual 1,4-dioxane was removed by azeotropic evaporation, 3× dry CH3OH (25 mL). The contents were then redissolved in dry CH3OH (25 mL), cooled to 0 °C, triturated with dry Et2O (400 mL), and stirred vigorously at room temperature in the dark to give a dark precipitate. The precipitate was filtered, washed with more dry EtO (150 mL) followed by dry n-hexane (50 mL) and then evaporated to dryness in the dark under high vacuum (<0.1 mbar) for 14 h to give the desired product, (S)-2-amino-3-[(2-nitrobenzyl)amino]propanoic acid HCl salt 3, as an off-white powder (3.575 g, 99%): 1 H NMR(400.13MHz,CD3OD)δ3.68(dd,J=13.2,5.6Hz,1H), 3.81(dd,J=13.2,7.7Hz,1H);4.48(dd,J=7.7,5.6Hz,1H), 4.65(d,J=13.2Hz,1H), 4.69(d,J=13.2Hz,1H), 7.74~7.82(m,1H), 7.84~7.95(m,2H), 8.30(apparently d,J=8.1Hz,1H); 13C NMR(100.61MHz,CD3OD)δ47.8(CH2), 50.2(CH), 50.7(CH2), 127.1(CH), 127.3(C), 132.8 (CH), 135.3(CH), 136.0(CH), 150.3(C), 169.0(C);MS(ESI+,LC-MS)m / z(relative intensity)240[(M+H) + ,100%].

[0312] 4',5'-Methylenedioxy-2'-nitroacetophenone (4b)

[0313] [ka]

[0314] McGall et al.(McGall, GH et al. The efficiency of light-directed synthesis The procedure was slightly modified from that described by I. M. Schneider, M. D., of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997) According to the modifications, a solution of 3',4'-(methylenedioxy)acetophenone 4a (16.416 g, 0.1 mol) in glacial CH3COOH (64 mL) was added dropwise over 1 h at 0 °C to a 2 L three-necked round-bottom flask containing conc. HNO3 (136 mL, 70% strength). The reaction mixture was maintained at 0 °C with stirring under an argon atmosphere during the addition and for a further 1 h. The mixture was then warmed to 40 °C and stirred for a further 2.5 h. Finally, the mixture was cooled to rt and slowly poured onto crushed ice (1 L) in a beaker. A yellow precipitate appeared. It was stirred for 15 min and then filtered. The yellow solid was washed with water (3 x 200 mL) and dried in vacuum. The crude yellow solid was then purified by recrystallization (THF / n-hexane) followed by flash chromatography on SiO2 [eluent: CH2Cl2 / n-hexane (1:1) to 100% CH2Cl2] to afford 4',5'-methylenedioxy-2'-nitroacetophenone (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014); McGall, GH et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997); Pendrak, I., Wittrock, R. & Kingsbury, WD Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995))4b was obtained as yellow crystals (12.141 g, 58%): f =0.52(CH2Cl2);mp122.8~124.0℃((Pendrak, I., Wittrock, R. & Kingsbury, WD Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995))mp112 °C); 1 H NMR(400.13MHz,CDCl3)δ2.45(s,3H), 6.16(s,2H), 6.71(s,1H), 7.48(s,1H); 13 IR(CH2Cl2)ν max 2980, 1708, 1525, 1506, 1484, 1424, 1362, 1338, 1271, 1152, 1038, 932, 875, 819cm -1 ;MS(ESI+)m / z(relative intensity)232[(M+Na) + ,10%], 210(2), 209(7), 194(100), 171(45), 130(32), 111(9).

[0315] (R,S)-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethano Rule (4c)

[0316] [ka]

[0317] 4',5'-Methylenedioxy-2'-nitroacetophenone 4b (43.714 g, 0.209 mol, 1.0 eq.) was suspended in CHCl (400 mL), CHOH (650 mL) and anhydrous CHCHOH (425 mL) in a 2 L one-neck round-bottom flask. The mixture was sonicated at rt for 10 min to dissolve most of the yellow solid. NaBH granules (7.116 g, 0.188 mol, 0.9 eq.) were added to the yellow suspension at 15 °C in eight portions (0.890 g each) every 15 min (total time = 2 h). Note: Effervescence appeared as NaBH dissolved and the reaction mixture became a homogeneous yellow solution. After the addition was complete, the reaction mixture was stirred at rt for an additional 4 h. After this time, the reaction was judged complete by TLC analysis (SiO2, TLC eluent: 100% CH2Cl2) and was quenched by the addition of dry acetone (100 mL) with stirring at rt for an additional 2 h. The mixture was then evaporated to dryness under reduced pressure to give a yellow solid. The solid was then redissolved in CH2Cl2 (800 mL) and washed successively with saturated aq. NH4Cl solution (3 x 500 mL) and finally with saturated aq. NaCl solution (6 x 800 mL). The organic layer was separated, dried over anhydrous Na2SO4, filtered, and evaporated to dryness under high vacuum to obtain (R,S)-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethanol (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014), McGall, G. H. et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997)) 4c was obtained as a yellow solid (43.563 g, 99%): f = 0.19 (CH2Cl2);mp 76.5-77.5℃; 1 H NMR(400.13MHz,CDCl3)δ1.50(d,J=6.3Hz,3H), 2.54(d,J=3.0Hz,1H), 5.42(qd,J=6.3,3.0Hz,2H), 6.096(apparently d, 2 J HH =3.5Hz, 1H, diastereotopic OCH2O), 6.104 (apparent d, 2 J HH =3.5Hz, 1H, diastereotopic OCH2O), 7.24(s,1H), 7.42(s,1H); 13 IR(CH2Cl2)ν max 3649, 2980, 2889, 2360, 2343, 1521, 1506, 1482, 1393, 1340, 1253, 1135, 1090, 1038, 934, 819cm -1 ;MS(ESI+)m / z(relative intensity)234[(M+Na) + ,1%], 194[(M-OH) + ,100], 130(20);HRMS(ESI+)m / z C9H9NO5[M+Na] + Calculated value: 234.0373, measured value 234.0364 (Δ=-3.95 ppm).

[0318] (R,S)-1-Bromo-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethane (4d)

[0319]

change

[0320] A 1-liter three-neck round-bottom flask was dried in vacuum using a heat gun at >100°C for 15 min, purged with dry argon gas, and cooled to room temperature. To this was added (R,S)-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethanol, 4c (15.838 g, 75.0 mmol, 1.0 eq.). 4c was dissolved in dry CHCl (375 mL, sonication was required for complete dissolution) and cooled to 0°C under argon atmosphere, and the round-bottom flask was wrapped in aluminum foil to protect from light. After 20 min, PBr (2.82 mL, 30.0 mmol, 0.4 eq.) was added dropwise at 0°C using a syringe pump over 10 min, followed by dry pyridine (0.5 mL). The yellow reaction mixture was stirred at 0°C for 15 min, then brought to rt and stirred continuously for 1.5 h. The reaction was judged complete by TLC analysis (SiO2, TLC eluent: 100% CH2Cl2), cooled to 0°C, quenched by addition of dry CH3OH (15 mL), warmed to rt, and stirred under an argon atmosphere for 30 min. After the quench was complete, the reaction mixture was evaporated to dryness under reduced pressure using a rotary evaporator. The resulting yellow gum was dissolved in CH2Cl2 (300 mL) and saturated aq. NaHCO3 solution (300 mL). The contents were placed in a separatory funnel, the aqueous phase discarded, and the organic phase washed successively with additional saturated aq. NaHCO3 solution (1 x 300 mL) and saturated aq. NaCl solution (3 x 300 mL). The organic layer was separated, dried over anhydrous Na2SO4, filtered, and evaporated to dryness to give a yellow solid.The crude product was purified by flash chromatography on SiO2 [eluent: CH2Cl2 / n-hexane (1:1) then 100% CH2Cl2] to give a pure sample of (R,S)-1-bromo-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethane (Nguyen, DP et al. Genetic Encoding of Photocaged Cysteine ​​Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014)) 4d (18.330 g, 89%) as shining yellow crystals. The sample was stored in a -20 °C freezer in a dry atmosphere in the dark for several months without significant decomposition: R. f = 0.17 (CH2Cl2 / n-hexane, 1:4); mp 76.1-77.8℃; 1 H NMR(400.13MHz,CDCl3)δ2.04(d,J=6.8Hz,3H), 5.89(q,J=6.8Hz,1H), 6.13(s,2H), 7.27(s,1H), 7.35(s,1H); 13 IR(CH2Cl2)ν max 2981, 2970, 2930, 1615, 1504, 1481 , 1420, 1395, 1385, 1328, 1305, 1257, 1156, 1141, 1057, 1028, 1014, 957, 925, 872, 815, 752, 730, 719, 698cm -1 ;HRMS(ESI+)m / z C9H8 79 BrNO4[M+Na] + Calculated value: 295.9529, measured value 295.9519 (Δ=-3.45 ppm).

[0321] tert-Butyl (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]amino}propanoate (4e)

[0322] [ka]

[0323] Boc-L-Dap-O t Bu·HCl 1b (6.233 g, 21.0 mmol, 1.1 eq.) was suspended in dry THF (275 mL) in a dry 1 L three-necked round-bottom flask and dry DIPEA (9.98 mL, 57.273 mmol, 3.0 eq.) was added. The contents were stirred at rt under nitrogen atmosphere for 10 min. The flask was wrapped in aluminum foil and the contents were stored in the dark. (R,S)-1-Bromo-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethane 4d (5.232 g, 19.091 mmol, 1.0 eq.) was then added to the reaction mixture. The homogeneous yellow solution was left stirring at rt under nitrogen atmosphere in the dark for 68 h. The reaction was judged complete by TLC analysis (SiO2 plate; CH2Cl2 / n-hexane = 3:7) and was evaporated to dryness under reduced pressure to give a dark brown oil. The crude reaction oil was dissolved in CH2Cl2 (250 mL) and washed with saturated brine solution (3 x 500 mL). The organic layer was separated, dried over anhydrous Na2SO4, filtered, and evaporated to dryness to give a dark brown viscous oil. This was then purified by flash chromatography on SiO2 (gradient; eluent: 100% CH2Cl2, then CH2Cl2 / CH3OH / NEt3=94:5:1) to give the desired product, tert-butyl (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]amino}propanoate 4e as a yellow-brown sticky gum (7.49 g, 87%): R f = 0.13 (SiO2 plate, CH2Cl2); 1H NMR (400.13 MHz, CDCl3 with TMS as internal standard) δ 1.34 and 1.36 (2×d, J=3.6 and 3.6 Hz, 3H), 1.42 and 1.450 (2×s, 9H), 1.454 and 1.47 (2×s, 9H), 2.54-2.74 (m, 1H), 2.75-2.89 (m, 1H), 4.03-4.27 (m, 1H), 4.28-4.53 (m, 1H), 5.15-5.4 3 (m, 1H), 6.05-6.10 (m, 2H), 7.21 (apparently broad s, 1H), 7.345 and 7.352 (2×s, 1H); 13 C NMR (100.61 MHz, CDCl3 with TMS as internal standard) δ (mixture of diastereomers) 23.9 (CH3), 24.0 (CH3), 28.12 (CH3), 28.15 (CH3), 28.41 (CH3), 28.46 (CH3), 49.3 (CH2), 49.4 (CH2), 53.1 (CH), 53.2 (CH), 54.3 (CH), 54.5 (CH), 79.9 (C), 80.1 (C ), 82.3(C), 82.4(C), 102.8(2×CH2), 105.2(2×CH), 106.77(CH), 106.83(CH), 138.1(C), 143.30(C), 143.37( C), 146.7(C), 152.1(C), 152.2(C), 155.5(C), 155.6(C), 170.7(C), 170.8(C);MS(ESI+)m / z(relative intensity)454[(M+H) + ,86%], 301(70), 261(100), 205(7), 203(10), 186(7), 147(10);HRMS(ESI+)m / z C 21 H 32 O8N3[M+H] + Calculated value: 454.2184, measured value 454.2201 (Δ=3.65 ppm).

[0324] (2S)-2-Amino-3-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]amino}propanoic acid (4)

[0325] [ka]

[0326] tert-Butyl (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]amino}propanoate 4e (4.303 g, 9.489 mmol, 1.0 eq.) was dissolved in dry CHCl (30 mL) in a dry 250 mL round-bottom flask wrapped in aluminum foil to exclude light. Freshly distilled CFCOOH (15 mL, 195.887 mmol, 20.644 eq.) was added, turning the yellow solution brown. Dry EtSiH (10.0 mL, 62.608 mmol, 6.598 eq.) was added and the reaction mixture was stirred at rt in the dark. The reaction was periodically monitored by LC-MS analysis. After 48 h, the reaction was judged complete by LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase, gradient) and the mixture was then evaporated to dryness to give a dark brown gum. In a dry 2 L round bottom flask, this gum was dissolved in anhydrous CH3OH (10 mL) and cooled to 0 °C under argon. This was triturated by addition of dry Et2O (900 mL) at 0 °C and then stirred vigorously at rt for 1 h to give a pale yellow precipitate. The precipitate was filtered and washed with additional dry Et2O (2 x 200 mL) followed by n-hexane (150 mL). The pale yellow powder was transferred to a 100 mL round bottom flask and dried in the dark at high vacuum (<0.1 mbar) for 40 h. (2S)-2-Amino-3-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]amino}propanoic acid 4 was obtained as a free-flowing pale yellow powder (3.646 g, 73 %). The product is a approx. 1:1 mixture of the salt and epimer of CF3COOH. The product was stored in the dark under argon at -20°C: 1H NMR (400.13 MHz, DMSO-d6 with TMS as internal standard) δ 1.36 and 1.38 (2 × d, J = 3.8 and 3.8 Hz, 3H), 2.65-2.95 and 2.96-3.20 (2 × m, 1H), 3.07-3.25 (m, 1H), 3.60-3.70 (m, 1H), 3.71-3.85 and 4.18-4.44 (2 × m, 1H), 6.21 and 6.23 (2 × d, J = 3.5 and 2.9 Hz, 2H), 7.41 and 7.42 (2 × s, 1H), 7.52 and 7.53 (2 × s, 1H); 13 C NMR (100.61 MHz, DMSO-d6 with TMS as internal standard) δ (mixture of diastereomers) 22.6 (CH3), 22.7 (CH3), 38.5 (CH2), 45.8 (CH2), 51.6 (CH), 51.8 (CH), 52.3 (CH), 52.6 (CH), 103.2 (CH2), 103.3 (CH2), 104.5 (CH), 104.6 (CH), 106.4 (CH), 106.6 (CH), 117.1 (C,q, 1 J C-F =299.0Hz), 135.4(C), 135.6(C), 142.96(C), 143.01(C), 146.61(C), 146.66(C), 151.93(C), 151.96(C), 158.56(C,q, 2 J C-F =31.5Hz), 168.7(2×C), 169.6(2×C);MS(ESI+)m / z(relative intensity)298[(M+H) + ,100%], 261(10), 225(10), 211(4), 147(12), 144(9), 134(6), 105(9), 82(31);HRMS(ESI+)m / z C 12 H 16 O6N3[M+H] + Calculated value: 298.1034, measured value 298.1039 (Δ=1.74 ppm).

[0327] 2,5-Dioxopyrrolidin-1-yl (1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl)carbonate (5a)

[0328] [ka]

[0329] (R,S)-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethanol 4c (42.234 g, 200.0 mmol, 1.0 eq.) was placed in a dry 2 L three-neck round bottom flask and dissolved in dry CH3CN (1 L). To this solution was added dry DIPEA (104.5 mL, 600.0 mmol, 3.0 eq.) followed by N,N-disuccinimidyl carbonate (80.896 g of ≥95% purity, 300 mmol, 1.5 eq.). The flask was wrapped in aluminum foil and the contents were stored in the dark. The heterogeneous yellow reaction mixture was stirred at rt under an argon atmosphere in the dark. After 16 h, the reaction mixture was homogeneous and judged complete by TLC analysis (SiO2 plate, CH3CN / CH2Cl2 = 1:19). The yellow reaction mixture was then adsorbed onto Biotage® Isolute HM-N adsorbent and dried under vacuum. This was immediately subjected to flash chromatography on SiO2 in the dark [eluent: CH2Cl2, then CH3CN / CH2Cl2=1:19] to give the desired product, 2,5-dioxopyrrolidin-1-yl-(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl)carbonate 5a as yellow needles (65.051 g, 92%) [Note: flash column must be performed quickly to avoid product decomposition upon prolonged exposure to SiO2]. The product 5a was immediately used in the subsequent step. It can be stored in the dark in a freezer at -20°C: R f = 0.6 (SiO2 plate, CH3CN / CH2Cl2 = 1:19); 1 H NMR (400.13 MHz, CDCl3 with TMS as internal standard) δ 1.75 (d, J = 6.4 Hz, 3H), 2.81 (s, 4H), 6.15 (d, J = 3.0 Hz, 2H), 6.42 (q, J = 6.4 Hz, 1H), 7.11 (s, 1H), 7.51 (s, 1H); 13C NMR (100.61 MHz, CDCl3 with TMS as internal standard) δ 22.2 (CH3), 25.6 (CH2), 76.4 (CH), 103.5 (CH2), 105.5 (CH), 105.8 (CH), 133.1 (C), 141.6 (C), 148.0 (C), 150.7 (C), 153.0 (C), 168.6 (C).

[0330] (2S)-2-[(tert-butoxycarbonyl)amino]-3-({[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy]carbonyl}-amino)propanoic acid (5b)

[0331] [ka]

[0332] Boc-L-Dap-OH 1a (2.553 g, 12.5 mmol, 1.25 eq.) was suspended in dry THF (180 mL) and dry CH3CN (20 mL) in a dry 1 L one-neck round-bottom flask wrapped in aluminum foil to exclude light. To this mixture was added dry DIPEA (5.23 mL, 30.0 mmol, 3.0 eq.) and the contents were stirred at rt under argon atmosphere for 20 min, after which 2,5-dioxopyrrolidin-1-yl-(1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl)carbonate 5a (3.523 g, 10.0 mmol, 1.0 eq.) was added. The heterogeneous mixture was stirred in the dark under argon atmosphere at rt and the progress of the reaction was monitored periodically by LC-MS analysis. After several hours, the heterogeneous mixture began to become a homogeneous yellow solution. After 24 hours, the reaction was deemed complete and the contents were transferred to Biotage The eluate was adsorbed onto Isolute HM-N adsorbent and dried under reduced pressure. It was then immediately subjected to flash chromatography on SiO2 in the dark [eluent: CH2Cl2, then Upon subjecting to CH2Cl2 / CH3OH / CH3COOH=94:5:1], the desired product, (2S)-2-[(tert-butoxycarbonyl)amino]-3-({[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy]carbonyl}amino)propanoic acid 5b, was obtained as a brown-yellow gum. This was subjected to azeotropic evaporation using CH2Cl2 / cyclohexane (1:1) under reduced pressure to remove residual CH3COOH from the product 5b. The product was dried under high vacuum to give a pure sample of 5b as a yellow solid (4.360 g, 99%) and a ca. 1:1 mixture of epimers; R f = 0.41 (SiO2 plate, CH2Cl2 / CH3OH / CH3COOH = 94:5:1); MS (ESI-, LC-MS) m / z (relative intensity) 440 [(MH) - , 100%].

[0333] (2S)-2-Amino-3-({[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethoxy]carbonyl}-amino)propanoic acid (5)

[0334] [ka]

[0335] Freshly distilled CF3COOH (15 mL, 195.894 mmol, 21.293 eq.) was added to a solution of (2S)-2-[(tert-butoxycarbonyl)amino]-3-({[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)-ethoxy]carbonyl}amino)propanoic acid 5b (4.061 g, 9.2 mmol, 1.0 eq.) in dry CHCl2 (50 mL) in a dry 1 L one-neck round-bottom flask wrapped in aluminum foil. Upon addition of CF3COOH, the yellow solution turned dark brown. The reaction was stirred in the dark at rt and monitored by TLC analysis. After 2 h, the reaction was judged complete by both TLC analysis (SiO2 plate; CH2Cl2 / CH3OH / CH3COOH=94:5:1) and LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase, gradient). The reaction mixture was evaporated to dryness under reduced pressure to give a dark brown gum. This gum was dissolved in dry CH3OH (5 mL), cooled to 0 °C, and triturated with dry Et2O (0.9 L) to give a pale yellow precipitate. The mixture was allowed to stir vigorously at rt under argon atmosphere in the dark. The pale yellow precipitate was then filtered and washed with Et2O (2 x 100 mL) and dry hexane (50 mL). It was dried in vacuum (<0.1 mbar) for 2 days in the dark to give the desired (2S)-2-amino-3-({[1-(6-nitrobenzo[d][1, 3]Dioxol-5-yl)ethoxy]carbonyl}-amino)propanoic acid TFA salt 5 was obtained as a fine pale yellow powder (4.132 g, 99%) and 1 H and 13 Obtained as a ca. 1:1 mixture of epimers as observed by C NMR spectroscopy: 1 H NMR(400.13MHz,CD3OD)δ1.57(d,J=6.2Hz,3H), 2.68(s,2H), 3.41~3.58(m,1H), 3.59~3.74(m,1H ), 3.78~4.15(m,1H), 6.14(s,2H), 6.22(q,J=6.2Hz,1H), 7.12(apparently d,J=5.2Hz,1H), 7.47(s,1H); 13 C NMR(100.61MHz,CD3OD)δ22.4(CH3), 22.5(CH3), 42.1(CH2), 42.3(CH2), 55.6(CH), 55 .9(CH), 70.4(2×CH), 104.8(2×CH2), 105.7(2×CH), 106.7(CH), 106.9(CH), 118.2(C,q, 1 J C-F =292.6Hz), 136.8(C), 137.1(C), 142.8(C), 142.9(C), 148.8(C), 154.0(C), 158.5(C), 158.6(C), 163.1(C,q, 2 J C-F =34.4Hz), 170.9(2×C), 174.9(2×C);MS(ESI+)m / z(relative intensity)342[(M+H) + ,100%], 311(10), 233(5), 189(9), 130(19);HRMS(ESI+)m / z C 13 H 16 O8N3[M+H] + Calculated value: 342.0932, measured value 342.0923 (Δ=-2.63 ppm).

[0336] 2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethan-1-ol (6a)

[0337] [ka]

[0338] A freshly prepared solution of NaOH (0.5 M, 8 g, 20.0 mmol, 1 eq. in 40 mL deionized HO) was placed in a 500 mL round 3-necked round-bottomed flask and degassed by bubbling with a stream of argon gas at rt. After 30 min, mercaptoethanol (1.47 mL, 21.0 mmol, 1.05 eq.) was added to the flask and degassing was continued for another 15 min. Separately, fresh (R,S)-1-bromo-1-[4',5'-(methylenedioxy)-2'-nitrophenyl]ethane 13 (5.481 g, 20.0 mmol, 1.0 eq.) was dissolved in 1,4-dioxane (20 mL) in a 100 mL round-bottom flask wrapped in aluminum foil and degassed by bubbling a stream of argon gas through it for 15 min in the dark. The degassed solution of 13 in 1,4-dioxane was added to the flask containing the aq. NaOH and mercaptoethanol solution and degassed with argon gas at rt for a period of 90 min. The mixture was transferred dropwise using a cannula under positive pressure of 1000 ml. A yellow precipitate formed. This was then dissolved by the addition of degassed 1,4-dioxane (60 mL) followed by sonication for 30 min until a homogenous clear yellow solution was obtained. The contents were then allowed to stir at rt in the dark under an argon atmosphere for 12 h. The reaction was then determined to be complete by TLC and LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase, gradient). The mixture was then evaporated under reduced pressure to remove the volatile organic components. The yellow aqueous contents were then extracted with EtOAc (2 x 175 mL) and the combined organic phases were washed with saturated NH4Cl solution (1 x 500 mL) followed by brine solution (3 x 500 mL). The organic layer was then separated, dried over anhydrous Na2SO4, filtered and evaporated to dryness to give a yellow oil. The product was purified by flash chromatography on SiO2 in the dark (eluent: EtOAc / n-hexane = 3:7) to give 2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethan-1-ol 6a as a sticky yellow oil (5.179 g, 95%): R f = 0.33 (SiO2 plate, EtOAc / n-hexane = 3:7); 1H NMR (400.13MHz, CDCl3) δ1.55(d,J=7.0Hz,3H), 1.96(t,J=5.9Hz,1H), 2.44~2.65(m,2H), 3.5 2~3.74(m,2H), 4.78(q,J=7.0Hz,1H), 6.10(dd,J=3.8,1.0Hz,2H), 7.27(s,1H), 7.28(s,1H); 13 C NMR (100.61MHz, CDCl3) δ23.2(CH3), 34.9(CH2), 38.4(CH), 60.9(CH2), 103.1(C H2), 104.8(CH), 108.0(CH), 136.1(C), 143.3(C), 146.9(C), 152.0(C);IR(undiluted)ν max 3393, 2980, 1617, 1518, 1503, 1480, 1418, 1375, 1332, 1252, 1156, 1031, 928, 872, 817, 759;m / z(ESI-,LC-MS)270.1[(MH) - ,100%].

[0339] 2,5-Dioxopyrrolidin-1-yl-(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethyl)carbonate (6b)

[0340] [ka]

[0341] Intermediate 6b used in the subsequent reaction for the synthesis of DAP derivatives 6c and 6d was synthesized in situ starting from alcohol 6a. A 500 mL three-necked round-bottom flask was dried in vacuum using a heat gun and purged with argon gas; this procedure was repeated three times before use. A dry flask was charged with 2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethan-1-ol 6a (4.883 g, 18.0 mmol, 1.0 eq.) dissolved in dry CH3CN (90 mL). To this reaction mixture was added dry DIPEA (9.41 mL, 54.0 mmol, 3.0 eq.) followed by N,N'-disuccinimidyl carbonate (6.796 g of ≥95% purity, 25.2 mmol, 1.4 eq.) in the dark under argon atmosphere at rt. The reaction mixture turned cloudy yellow and a white precipitate began to form, after 1 h the reaction mixture became a homogeneous yellow-brown solution. The reaction was allowed to stir at rt for 12 h after which it was judged complete by TLC analysis (SiO2 plate, EtOAc / n-hexane=3:7). 2,5-dioxopyrrolidin-1-yl-(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethyl)carbonate 6b was immediately carried on to the next step without further purification: R f =0.12 (SiO2 plate, EtOAc / n-hexane = 3:7).

[0342] tert-Butyl (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoate (6c)

[0343] [ka]

[0344] Boc-L-Dap-O tBu·HCl (8.548 g, 28.8 mmol, 1.6 eq.) was added in one portion to the solution of 6b prepared above. The yellow reaction mixture became homogeneous within a few minutes and the contents were stirred in the dark under an argon atmosphere. After 10 h, TLC showed that the 6b was homogeneous. Analysis (SiO2 plate, R f The reaction was judged to be complete by both HPLC (pH 7.0, EtOAc / n-hexane = 3:7) and LC-MS analysis (C18 reverse phase column, HO-CHCN as mobile phase) confirming consumption of 6b. The reaction mixture was then purified by Biotage It was adsorbed onto Isolute® HM-N adsorbent and dried under reduced pressure, which was then subjected to flash chromatography on SiO2 in the dark [eluent: EtOAc / n-hexane=3:7] to give the desired tert-butyl (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoate 6c as a thick yellow gum (9.405 g, 94%) and a ca. 1:1 mixture of epimers: R f = 0.39 (EtOAc / n-hexane = 3:7); 1 H NMR (400.13 MHz, CDCl3) δ (mixture of epimers) 1.44 (s, 9H), 1.46 (s, 9H), 1.54 (d, J = 6.8 Hz, 3H), 2.36-2.61 (m, 2H), 3.41-3.68 (m, 2H), 4.00-4.18 (m, 2H), 4.24 (broad s, 1H), 4.85 (q, J = 6.8 Hz, 1H), 5.15 (broad s, 1H), 5.41 (broad s, 1H), 6.10 (d, J = 6.8, 2H), 7.27 (s, 1H), 7.29 (s, 1H); 13C NMR(100.61MHz,CDCl3)δ23.1(CH3), 28.1(CH3), 28.4(CH3), 30.5(CH2), 39.0(CH), 43.2(CH2), 54.6(CH), 65.2(CH2), 80.1(C), 82.9(C), 10 3.0(CH2), 104.7(CH), 108.2(CH), 136.3(C), 143.5(C), 146.9(C), 152.1(C), 155.6(C), 156.4(C), 169.7(C); m / z(ESI+,LC-MS)558.2[(M+H) + ,100%]

[0345] (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoic acid (6d)

[0346] [ka]

[0347] Boc-L-Dap-OH (13.479 g, 66.0 mmol, 1.082 eq.) was added in one portion to a solution of 6b (26.70 g, prepared above) in dry CH3CN (305 mL) under argon and stirred at rt for 12 h. The reaction was then judged complete by LC-MS (C18 reverse phase column, HO-CH3CN as mobile phase) and the contents were adsorbed onto Biotage® Isolute HM-N sorbent and dried under reduced pressure. This was then subjected to flash chromatography in the dark on spherical SiO2 [Supelco®, purchased from Sigma Aldrich Ltd., particle size 40-75 μm; gradient; eluent: Upon subjection to EtOAc / n-hexane=1:1→7:3→1:0], the desired (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoic acid 6d was obtained as a thick yellow gum (28.995 g, 95%) and a ca. 1:1 mixture of epimers: 1 H NMR [400.13 MHz, CDCl3 with 0.1% v / v TMS as internal standard] δ (mixture of epimers) 1.43 (s, 9H), 1.52 (d, J = 6.8 Hz, 3H), 2.30-2.95 (m, 2H), 3.33-3.82 (broad m, 2H), 3.86-4.18 (m, 2H), 4.20-4.48 (m, 1H), 4.64-4.97 (m, 1H), 5.34-5.58 (broad s, 1H), 5.60-5.84 (broad s, 1H), 6.20 (d, J = 8.3 Hz, 2H), 7.10-7.39 (m, 2H), 8.47 (broad s, 1H); 13 C NMR [100.61 MHz, CDCl3 with 0.1% v / v TMS as internal standard] δ 23.1 (CH3), 28.4 (3 × CH3), 30.5 (CH2), 39.0 (CH), 42.7 (CH2), 54.4 (CH), 65.3 (CH2), 80.8 (C), 103.1 (CH2), 104.7 (CH), 108.1 (CH), 136.2 (C), 143.4 (C), 146.9 (C), 152.1 (C), 156.3 (C), 157.2 (C), 173.5 (C); m / z (ESI-, LC-MS) 500.1 [(MH) - ,100%]

[0348] (2S)-2-Amino-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoic acid (6)

[0349] [ka]

[0350] Method I (prepared from 6c): A dry sample of (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoate 6c (6.728 g, 12.066 mmol, 1.0 eq.) was placed in a dry 250 mL one-neck round-bottom flask and dissolved in dry CHCl (50 mL), and the flask was wrapped in foil to exclude light. To this solution was added dry EtSiH (20 mL, 125.215 mmol, 10.378 eq.), followed by the dropwise addition of freshly distilled CFCOOH (20 mL, 261.182 mmol, 21.647 eq.) via syringe at rt over 15 min. The reaction mixture turned from yellow to brown-green in color. This was allowed to stir at rt in the dark. After 24 h, the reaction was judged complete by TLC (SiO2 plate, EtOAc / n-hexane = 3:7) and LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase). The reaction mixture was concentrated under reduced pressure in the dark to give a yellow-brown gum. This was dissolved in anhydrous CH3OH (20 mL) and evaporated to dryness under reduced pressure: this was repeated three times and dried under high vacuum (<0.1 mbar) to remove any residual CF3COOH, Et3SiH and H2O. The yellow-brown gum was then dissolved in dry CH3OH (40 mL) and transferred to a dry 2 L round bottom flask under an argon atmosphere and cooled to 0 °C. Dry Et2O (2 L) was added via cannula to this solution under a positive pressure of argon gas in the dark and the contents were stirred vigorously. A yellow precipitate formed and the contents were vigorously stirred at 0° C. for 15 min and then at rt for an additional 2 h. The pale yellow precipitate was filtered and washed with dry EtO (3×250 mL) and finally with dry n-hexane (50 mL).The product was dried in the dark in vacuum (<0.1 mbar) overnight for 14 h to give (2S)-2-amino-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}-propanoic acid TFA salt 6 as a pale yellow powder (3.940 g, 63%) and a 1:1 mixture of epimers: 1 H NMR [400.13 MHz, 1% v / v TMS as internal standard] CD3OD / CF3COOD (5:1) containing δ (mixture of epimers) 1.55 (d, J = 7.0 Hz, 3H), 2.49-2.73 (m, 2H), 3.63 (dd, J = 15.0, 6.4 Hz, 1H), 3.78 (ddd, J = 15.0, 3.6, 2.3 Hz, 1H), 4.0-4.24 (m, 3H), 4.81 (q, J = 7.0 Hz, 1H), 6.11 and 6.13 (2 × s, 1H), 6.40 (s, 1H), 7.29 and 7.33 (2 × s, 1H); 13 C NMR [100.61 MHz, CD3OD / CF3COOD (5:1) with 1% v / v TMS as internal standard] δ 23.2 (CH3), 31.4 (CH2), 40.0 (CH), 42.1 (CH2), 55.1 (CH), 65.8 (CH2), 104.8 (CH2), 105.6 (CH), 109.0 (CH), 117.0 (C, 1 J C-F =286.5Hz), 137.1(C), 144.9(C), 148.7(C), 153.7(C), 160.7(C), 160.8(C, 1 J C-F =38.1Hz), 170.2(C);MS(ESI+)m / z(relative intensity)402[(M+H) + ,100%], 386(20), 224(9), 208(11), 151(11);HRMS(ESI+)m / z C 15 H 20 N3O8S[M+H] + Calculated value: 402.0971, measured value 402.0974 (Δ=0.7 ppm).

[0351] Storage: Dried samples of DAP Amino Acid 6·TFA were stored in airtight dark glass vials in a cool, dry, dark environment and were stable for 3 years without decomposition. Handling: DAP Amino Acid 6·TFA is light sensitive and slightly hygroscopic on exposure to moist air. For this reason, samples of it in vials were always handled in a dark, dry atmosphere. Of note, vials containing 6·TFA removed from the refrigerator or freezer were always allowed to warm to rt before opening and handling.

[0352] Method II (prepared from 6d): A dry sample of (2S)-2-[(tert-butoxycarbonyl)amino]-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoic acid 6d (26.70 g, 53.2395 mmol, 1.0 eq.) was placed in a dry 1 L one-neck round-bottom flask and dissolved in dry CHCl (300 mL). The flask was wrapped in aluminum foil to exclude light, and dry EtSiH (84.69 mL, 530.24 mmol, 10.0 eq.) was added to the solution. After 5 min, freshly distilled CFCOOH (81.54 mL, 1.0648 mol, 20.0 eq.) was added dropwise to the solution over 15 min. The solution turned yellow-brown in color. It was left stirring in the dark at rt. After 5 h, the reaction was judged complete by TLC (SiO2 plate, EtOAc / CH3COOH=98:2) and LC-MS analysis (C18 reverse phase column, H2O-CH3CN as mobile phase). The solution was concentrated to dryness under reduced pressure to give a yellow-brown gum. This was dissolved in dry CH3OH (40 mL) and evaporated to dryness under reduced pressure; this was repeated three times and the product was dried under high vacuum (<0.1 mbar) to remove any residual CF3COOH, Et3SiH and H2O. The yellow-brown gum was dissolved in dry CH3OH (40 mL) and transferred to a dry 3 L round bottom flask under an argon atmosphere and cooled to 0 °C. With the contents vigorously stirred, dry Et2O (2.5 L) was added to the flask via cannula under a positive pressure of argon gas. A pale yellow precipitate formed. The contents were vigorously stirred at 0 °C for 15 min and then at rt for 2 h. The precipitate was then filtered and washed with dry Et2O (3 x 500 mL) followed by dry n-hexane (150 mL). The product was dried overnight in the dark under high vacuum (<0.1 mbar) for 14 h to give (2S)-2-amino-3-{[(2-{[1-(6-nitrobenzo[d][1,3]dioxol-5-yl)ethyl]thio}ethoxy)carbonyl]amino}propanoic acid TFA salt 6 as a pale yellow powder (20.465 g, 75%) and a ca. 1:1 mixture of epimers.

[0353] Synthesis of Vlm TE substrate

[0354] [ka]

[0355] Synthesis scheme of depsipeptidyl-SNAC compounds 7 and 8. a, Synthesis of deoxytetradepsipeptidyl-SNAC 8. a) 8c, EDC, DMAP, 72%; b) TFA, DCM, 99%; c) TBSCl, Imid., DCM; d) LiOH, THF, 78%, 2 steps; e) (COCl)2, DMF, DCM; f) TEA, DCM, 53%; g) HF, Pyr., MeCN, 84%; h) EDC, DMAP, TEA, DCM, 60%; i) LiOH, MeOH, THF, 60%; j) EDC, DMAP, DMF, 5:4dr, 92% b, Synthesis of tetradepsipeptidyl-SNAC 7. a) AllylBr, Cs2CO3, DMF, 95%; b) Boc-d-Val, EDC, DMAP, DCM, 84%; c) Pd(PPh3)4, morpholine, DCM; d) EDC, HOBt, DIPEA, DCM, 94%; e) HCl, dioxane; f) d-HIV, EDC, HOBt, DIPEA, DCM, 95%. c) Structure of deoxytetradepsipeptidyl-SNAC 8; and d Structure of tetradepsipeptidyl-SNAC 7.

[0356] (S)-S-(2-Acetamidoethyl) 2-((tert-butoxycarbonyl)amino)-3-methylbutanethioate (7a)

[0357] [ka]

[0358] Boc-L-valine (7.29 g, 33.56 mmol, 1.0 equiv.) was dissolved in CHCl. ​​N-Acetyl-cysteamine (4.00 g, 33.56 mmol, 1.0 equiv.). To this mixture was added N-(3-diaminomethylpropyl)-N'-ethylcarbodiimide hydrochloride (EDC, 7.72 g, 40.27 mmol, 1.2 equiv.) and 4-(dimethylamino)pyridine (DMAP, 410 mg, 3.36 mmol, 0.1 equiv.). The reaction was stirred at ambient temperature for 16 h. The reaction was quenched with NHCl(aq) and extracted three times with EtOAc. The organic fractions were combined, washed with brine, dried over NaSO, and concentrated. The desired product (7.69 g, 24.16 mmol, 72% yield) was purified by silica column chromatography (5% MeOH in CH2Cl2). f =0.37 (acetone:hexanes 2:3). 1 H NMR(400MHz,CDCl3)δ5.95(s,1H), 4.97(d,J=8.8Hz,1H), 4.21(dd,J=8.9,4.8Hz,1H), 3.48~3.30(m,2H), 3 .08~2.94(m,2H), 2.22(td,J=13.4,6.7Hz,1H), 1.43(s,9H), 0.96(d,J=6.9Hz,3H), 0.85(d,J=6.9Hz,3H). 13 C NMR (100MHz, CDCl3) δ201.74, 170.35, 155.66, 80.42, 65.68, 39.38, 30.77, 28.38, 28.33, 23.16, 19.40, 17.01. HRMS(ESI+) Calculated mass value (C 14 H 26 N2O4SNa) 341.1511, measured value 341.1512.

[0359] (S)-S-(2-acetamidoethyl) 2-amino-3-methylbutanethioate (7b)

[0360] [ka]

[0361] In a round-bottom flask, 7a (0.5 g, 1.57 mmol, 1.0 equiv.) was dissolved in CH2Cl2 (3 mL). The solution was cooled to 0° C. using an ice bath and trifluoroacetic acid (3 mL) was added. The reaction was allowed to proceed at ambient temperature for 45 min. The reaction mixture was concentrated and the desired product (341 mg, 1.56 mmol, >99% yield) was purified by silica column chromatography (MeOH 5% to 10% in CH2Cl2). 1 H NMR (300MHz, DMSO) δ8.45(s,157 2H), 8.10(t,J=5.5Hz,1H), 4.15(d,J=4.8Hz,1H), 3.27~3.17(m,2H), 3.13~2 .98(m,2H), 2.28~2.09(m,1H), 0.99(d,J=6.9Hz,3H), 0.95(d,J=7.0Hz,3H). 13 C NMR (75MHz, DMSO) δ196.18, 169.35, 63.48, 37.78, 30.10, 28.40, 22.50, 18.03, 17.26.

[0362] (S)-2-((tert-butyldimethylsilyl)oxy)propanoic acid (7c)

[0363] [ka]

[0364] In a round-bottom flask, L-ethyl lactate (5.08 g, 43.0 mmol, 1.0 equiv) was dissolved in CHCl (55 mL) and the solution was cooled to 0° C. using an ice bath. To this mixture, tert-butyldimethylsilyl chloride (6.48 g, 45.15 mmol, 1.05 equiv) and imidazole (3.51 g, 51.6 mmol, 1.2 equiv) were added and the reaction was then allowed to proceed at ambient temperature for 2 h. The reaction mixture was then diluted with H0 and extracted three times with CHCl. ​​The organic fractions were combined, washed with ice-cold 5% HCl (aq), washed with brine, dried over NaSO, and concentrated. The crude intermediate (S)-ethyl 2-(tert-butyldimethylsilyloxy)propanoate was dissolved in THF (215 mL). The mixture was cooled to 0 °C using an ice bath and a cold solution of LiOH (0.4 M, 215 mL) was added dropwise over 20 min. The reaction mixture was stirred at ambient temperature for 4 h. The resulting reaction mixture was concentrated to half of its original volume and the resulting aqueous solution was extracted three times with Et2O. The organic fractions were combined and extracted three times with a saturated solution of NaHCO3 (aq). The aqueous fractions were combined, acidified to pH 4 with 1 M KHSO4 (aq) and extracted three times with Et2O. The organic fractions were combined, dried over Na2SO4 and concentrated. The desired product (6.88 g, 33.7 mmol, 78% yield for two steps) was obtained, which was used without further purification. NMR data was reported from the literature (Ekici, OD, Paetzel, M. & Dalbey, RE Unconventional serine proteases: variations on the catalytic Ser / His / Asp triad The configuration was consistent with that of Protein Sci 17, 2023-2037 (2008)). 1 H NMR (300MHz, CDCl3) δ4.36(q,J=6.8Hz,1H), 1.45(d,J=6.8Hz,3H), 0.92(s,9H), 0.13(s,6H).

[0365] (S)-2-((tert-butyldimethylsilyl)oxy)propanoyl chloride (7d)

[0366] [ka]

[0367] In a round-bottom flask, 7c (3.7 g, 18 mmol, 1.0 equiv) was dissolved in DMF (45 mL) and the solution was cooled to 0° C. using an ice bath. Oxalyl chloride (13.6 mL of a 2.0 M solution in DCM, 10.0 equiv) and a catalytic amount of DMF were added. The reaction proceeded from 0° C. to ambient temperature over 2 h. The reaction mixture was concentrated. The crude oil was used in the subsequent reaction without purification.

[0368] TBSO-L-Lac-L-Val-SNAC(7e)

[0369] [ka]

[0370] In a round bottom flask, 7b (1.95 g, 9 mmol, 1.0 equiv) was dissolved in CHCl (40 mL). Crude oil 7d (18 mmol, 2.0 equiv) was dissolved in CHCl (5 mL) and added to this mixture. EtN (2.5 mL, 18 mmol, 2.0 equiv) was added and the reaction was allowed to proceed for 4 h. The reaction mixture was quenched with NHCl(aq), extracted three times with EtOAc, washed with brine, and concentrated. The desired product (1.93 g, 4.77 mmol, 53% yield) was purified from the crude mixture by silica column chromatography (50% to 90% EtOAc in hexanes). 1 H NMR (300MHz, CDCl3) δ7.22(d,J=9.3Hz,1H), 6.03(s,1H), 4.53(dd,J=9.3,4.5Hz,1H), 4.25(q,J=6 .7Hz,1H), 3.38(q,J=6.2Hz,2H), 3.07~2.98(m,2H), 2.40~2.21(m,1H), 1.93(s,3H), 1.38(d,J=159 6.7Hz,3H), 1.01~0.82(m,15H), 0.13(s,3H), 0.12(s,3H). 13C NMR (75MHz, CDCl3) δ200.34, 174.90, 170.47, 70.03, 63.48, 39.47, 31.04, 28.51, 25.82, 23.23, 22.04, 19.47, 18.00, 16.83, -4.54, -5.03.

[0371] HO-L-Lac-L-Val-SNAC(7f).

[0372] [ka]

[0373] Compound 7e (250 mg, 0.617 mmol, 1.0 equiv) was dissolved in acetonitrile (20 mL) in a 50 mL polypropylene Falcon tube. Pyridine (249 μL, 3.09 mmol, 5 equiv) and HF (48 wt.% aq. 533 μL, 30.9 mmol, 50 equiv) were added. The reaction was stirred at ambient temperature for 16 h. The reaction mixture was quenched with NH4Cl(aq), extracted three times with EtOAc, washed with brine, dried over Na2SO4, and concentrated. The desired product (150.1 mg, 0.517 mmol, 84% yield) was purified by silica column chromatography (MeOH 2% to 8% in CH2Cl2). 1 H NMR(400MHz,CDCl3)δ7.21(d,J=9.2Hz,1H), 6.20(s,1H), 4.54(dd,J=9.2,5.4Hz,1H), 4.30(q,J=6.8Hz,1H), 4.15(s,1H), 3.52~3. 32(m,2H), 3.12~2.94(m,2H), 2.36~2.21(m,1H), 1.95(s,3H), 1.44(t,J=6.3Hz,3H), 0.97(d,J=6.8Hz,3H), 0.91(d,J=6.8Hz,3H). 13 C NMR (100MHz, CDCl3) δ200.20, 175.47, 170.94, 68.66, 63.77, 39.24, 30.90, 28.71, 23.25, 21.28, 19.45, 17.27.

[0374] (S)-Allyl 2-hydroxypropanoate (7g)

[0375] [ka]

[0376] In a round bottom flask, 1 g of L-lactic acid (11.11 mmol, 1 equiv.) and 3.8 g of cesium carbonate (11.67 mmol, 1.05 equiv.) were dissolved in 13 mL of DMF. Allyl bromide (3.75 mL, 5.37 g, 44.44 mmol, 4 equiv.) was added dropwise at ambient temperature. Upon complete addition, the reaction was stirred at ambient temperature for 48 h. Upon completion, excess allyl bromide was removed by rotary evaporation and the remaining solution was diluted with water and then cooled to room temperature. Extraction with EtO three times was performed. The combined organic fractions were washed twice with water and once with brine, dried over NaSO, and concentrated to give the title compound (1.47 g, 95%) as a pale yellow oil. Characterization data were reported (Liu, Y., Zheng, T. & Bruner, SD Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488 (2011)). 1 H NMR (400MHz, CDCl3) δ 5.99-5. 82(m,1H), 5.40~5.18(m,2H), 4.71~4.59(m,2H), 4.29(q,J=6.9Hz,1H), 2.75(s,1H), 1.42(d,J=6.9Hz,3H).

[0377] (R)-(S)-1-(Allyloxy)-1-oxopropan-2-yl 2-((tert-butoxycarbonyl)amino)-3-methyl-butanoate (7h).

[0378] [ka]

[0379] In a round bottom flask, 1 g of 7g (7.69 mmol, 1 equiv.) and 1.67 g of Boc-D-Val (8.46 mmol, 1.1 equiv.) were dissolved in 39 mL of CHCl. ​​To this solution, 2.21 g of EDC (11.54 mmol, 1.5 equiv.) and 1.03 g of DMAP (8.46 mmol, 1 equiv.) were added at ambient temperature. The resulting solution was stirred at ambient temperature for 20 h. The reaction was quenched with NHCl(aq), extracted three times with CHCl, washed with NaHCO(aq), washed with brine, dried over NaSO, and concentrated. The title compound (2.12 g, 84%) was purified by silica column chromatography (20% EtOAc in hexanes). R f =0.41 (EtOAc:Hexanes 1:3) 1 H NMR (400MHz, CDCl3) δ5.95~5.81(m,1H), 5.29(dddd,J=21.3,11.7,6.6,1.3Hz,2H), 5.13(q,J=7.0Hz,1H), 4.97(d,J=8.9Hz,1H), 4.67~4.5 9(m,2H), 4.28(dd,J=8.9,4.8Hz,1H), 2.25~2.11(m,1H), 1.50(d,J=7.1Hz,3H), 1.43(s,9H), 0.97(d,J=6.9Hz,3H), 0.91(d,J=6.9Hz,3H). 13 C NMR (100MHz, CDCl3) δ171.50, 169.95, 155.56, 131.43, 118.83, 79.77, 69.17, 65.93, 58.60, 31.28, 28.32, 18.99, 17.49, 17.00. HRMS(ESI+):C 16 H 27 Calculated exact mass of NNaO6: 352.1736. Found: 352.1721.

[0380] Boc-D-Val-L-Lac-L-Val-SNAC(7i)

[0381] [ka]

[0382] In a round bottom flask, 250 mg of 7h (0.76 mmol, 1 equiv.) was dissolved in 4 mL of CHCl under nitrogen atmosphere. To this solution was added 86 μL of morpholine (87 mg, 0.99 mmol, 1.3 equiv.) and 62 mg of Pd(PPh3)4 in one portion. The reaction was stirred at ambient temperature and monitored by TLC. Upon completion, the reaction was quenched by addition of 10% aq. HCl, the organic layer was removed, and the remaining aqueous fraction was extracted three times with CHCl2. The combined organic fractions were washed with brine, dried over NaSO4, concentrated, and this intermediate, 7j, was used immediately in the subsequent reaction. To a flame-dried round bottom flask was added 194 mg of 7b (as the HCl salt, 0.76 mmol, 1 equiv.) and crude 7j (0.76 mmol, 1 equiv.) in 4 mL of CHCl2. To the resulting solution was added 400 μL of Hunig's base (295 mg, 2.28 mmol, 3 equiv), 154 mg of HOBt (1.14 mmol, 1.5 equiv) and 220 mg of EDC (1.14 mmol, 1.5 equiv). The reaction was stirred under argon at ambient temperature for 20 h. The reaction was quenched with NH4Cl(aq), extracted three times with CH2Cl2, washed with NaHCO3(aq), then with brine, dried over Na2SO4 and concentrated. The title compound (350 mg, 94% for two steps) was purified by silica column chromatography (40% acetone in hexanes). R f =0.35 (acetone:hexanes 2:3) 1 H NMR (400MHz, CDCl3) δ7.08(d,J=8.2Hz,1H), 6.07(s,1H), 5.38(q,J=6.8Hz, 1H), 5.02(d,J=7.0Hz,1H), 4.46~4.39(m,1H), 3.99(t,J=6.9Hz,1H), 3.45~ 3.30(m,2H), 3.11~2.89(m,2H), 2.30(dq,J=13.4,6.7Hz,1H), 2.11~2.01(m ,1H), 1.92(s,3H), 1.49(d,J=6.9Hz,3H), 1.39(s,9H), 1.01~0.91(m,12H).13 C NMR (100MHz, CDCl3) δ200.14, 171.72, 170.89, 170.48, 155.92, 80.45, 70.58, 64.74, 59.74, 39.30, 30.47, 30.27, 28.46, 28.26, 23.10, 19.33, 18.90, 18.49, 17.85, 17.53. HRMS(ESI+):C 22 H 39 Calculated exact mass of N3NaO7S: 512.2406. Found: 512.2391

[0383] HO-D-Hiv-D-Val-L-Lac-L-Val-SNAC(7)

[0384] [ka]

[0385] A round-bottom flask was charged with 118 mg of 7i (0.24 mmol, 1 equiv.) in a minimal amount of THF and cooled to 0° C. To this was added 1 mL of 4 M HC1 in dioxane (Sigma). l was added and the reaction was allowed to warm to ambient temperature. The reaction was monitored by TLC and upon completion, all solvent was removed by rotary evaporation. The crude intermediate 7k was used immediately in the subsequent reaction. Intermediate 7k was dissolved in 2 mL of CHCl to which was added sequentially 125 μL of Hunig's base (93 mg, 0.72 mmol, 3 equiv), 32 mg of D-α-hydroxyisovaleric acid (0.27 mmol, 1.1 equiv), 49 mg of HOBt (0.36 mmol, 1.5 equiv) and 70 mg of EDC (0.36 mmol, 1.5 equiv). The reaction was stirred at ambient temperature for 24 h and upon completion, quenched with NHCl(aq), extracted five times with CHCl, washed with NaHCO(aq) then with brine, dried over NaSO and concentrated. The title compound (111 mg, 95%) was purified by silica column chromatography (50% acetone in hexanes). 1H NMR (300MHz, CDCl3) δ7.29(s,1H), 6.20(t,J=5.7Hz,1H), 5.26(q,J=7.0Hz,1H), 4.55 (br,1H), 4.47(dd,J=9.0,6.5Hz,1H), 4.26(t,J=7.7Hz,1H), 3.99(d,J=2.9Hz,1H),3. 52~3.25(m,2H), 2.99(ddt,J=20.4,13.3,6.5Hz,2H), 2.39~2.25(m,1H), 2.20~2.06( m,2H), 1.97(s,3H), 1.54(d,J=7.0Hz,3H), 1.06~0.93(m,15H), 0.88(d,J=6.9Hz,3H). 13 C NMR(75MHz,CDCl3)δ200.05, 175.25, 171.69, 171.43(2C), 76.33, 71.14, 64.55, 58.39, 38 .95, 31.95, 30.25, 30.12, 28.64, 23.22, 19.49, 19.17, 19.13, 18.83, 18.18, 18.04, 16.15. HRMS(ESI+):C 22 H 39 Calculated exact mass of N3NaO7S: 512.2401. Found: 512.2406

[0386] (R)-Methyl 3-methyl-2-(3-methylbutanamido)butanoate (8a)

[0387] [ka]

[0388] In a round bottom flask, D-valine methyl ester hydrochloride (250 mg, 1.5 mmol, 1.0 equiv) was dissolved in CHCl (15 mL). Isovaleric acid (230 mg, 2.25 mmol, 1.5 equiv), EDC (430 mg, 2.25 mmol, 1.5 equiv), DMAP (276 mg, 2.25 mmol, 1.5 equiv) and EtN (420 μL, 3.00 mmol, 2.0 equiv) were added and the reaction was allowed to mix at ambient temperature for 16 h. The reaction was quenched with NHCl(aq), extracted three times with CHCl, washed with NaHCO(aq), washed with brine, dried over NaSO and concentrated. The desired compound (193.7 mg, 0.90 mmol, 60% yield) was purified by silica column chromatography (20-50% EtOAc in hexanes). 1 H NMR (400MHz, CDCl3) δ5.96(d,J=8.0Hz,1H), 4.57(dd,J=8.8,4.9Hz,1H), 3.71(s,3H), 2.19~2.04(m,4H), 0.97~0.86(m,12H). 13 C NMR(100MHz,CDCl3)δ172.85, 172.49, 56.91, 52.19, 46.14, 31.35, 26.29, 22.56, 22.53, 19.07, 17.93

[0389] (R)-3-Methyl-2-(3-methylbutanamido)butanoic acid (8b)

[0390] [ka]

[0391] In a round-bottom flask, 8a (180 mg, 1.2 mmol, 1.0 equiv) was dissolved in MeOH (24 mL) and THF (24 mL) and the solution was cooled to 0 °C using an ice bath. LiOH (1 M, 24 mL) was added dropwise and the solution was allowed to warm from 0 °C to ambient temperature over 4 h. The solution was concentrated to 1 / 3 volume and the resulting aqueous solution was acidified to pH 3 with 10% HCl. The solution was extracted three times with CHCl, dried over NaSO, and concentrated. The desired product (145 mg, 0.72 mmol, 60% yield) was purified by silica column chromatography (5% MeOH + 0.5% acetic acid in CHCl). 1 H NMR (300MHz, MeOD) δ4.32 (d, J=5.8Hz, 1H), 2.23~2.01 (m, 4H), 1.01~0.92 (m, 12H). 13 C NMR (75MHz, MeOD) δ175.84, 174.93, 59.00, 45.90, 31.53, 27.50, 22.76, 22.72, 19.65, 18.41.

[0392] 8(A) and 8c(B) (R)-(S)-1-(((S)-1-((2-acetamidoethyl)thio)-3-methyl-1-oxobutan-2-yl)amino)-1-oxopropan-2-yl 3-methyl-2-(3-methylbutanamido)butanoate (8)

[0393] [ka]

[0394] In a round bottom flask, alcohol 7f (25.2 mg, 0.087 mmol, 1.0 equiv) and carboxylic acid 8b (35 mg, 0.174 mmol, 2.0 equiv) were dissolved in DMF (1 mL). The solution was cooled to -20 °C using a dry ice / acetone bath and EDC (67 mg, 0.35 mmol, 4.0 equiv) and DMAP (21 mg, 0.174 mmol, 2.0 equiv) were added. The mixture was allowed to warm to ambient temperature. The reaction proceeded over 16 h. The reaction was quenched with NH4Cl(aq) and extracted three times with EtOAc. The organic fractions were combined, washed with brine, dried over Na2SO4, and concentrated. A mixture of C-2.2 diastereomers in a 5:4 ratio (A:B) (37.9 mg, 0.08 mmol, 92% yield) was purified from the crude residue by silica column chromatography (MeOH 1%-5% in CH2Cl2). The diastereomers were separated by preparative TLC. 8(A) 1 H NMR(400MHz,CDCl3)δ7.21(d,J=8.2Hz,1H), 6.15(s,1H), 5.93(d,J=6.8Hz,1H), 5.35(q,J=7.0Hz,1H), 4.44 167(dd,J=8.3,6.4Hz,1H), 4.29(t,J=7.0Hz,1H), 3.50~3.32(m,2H), 3.08~2.95(m,2H), 2. 41~2.29(m,1H), 2.19~2.00(m,4H), 1.96(s,3H), 1.53(d,J=6.9Hz,3H), 1.05~0.92(m,18H). 13 C NMR(100MHz,CDCl3)δ200.14, 173.48, 171.62, 171.04, 170.65, 71.03, 64.93, 58.71, 45.66, 39. 33, 30.37, 30.35, 28.75, 26.31, 23.27, 22.63, 22.57, 19.48, 19.07, 18.79, 18.11, 17.91.8c(B) 1H NMR (400MHz, CDCl3) δ6.95(d,J=8.7Hz,1H), 6.04(s,1H), 5.81(d,J=7.2Hz,1H), 5.25( q,J=6.8Hz,1H), 4.54(dd,J=8.8,5.9Hz,1H), 4.49(dd,J=7.3,4.7Hz,1H), 3.47~3.36(m ,2H), 3.11~2.98(m,2H), 2.38~2.26(m,2H), 2.20~2.09(m,3H), 1.95(s,3H), 1.51(d,J= 6.9Hz,3H), 1.04(d,J=6.9Hz,3H), 0.99(dd,J=6.7,2.5Hz,12H), 0.94(d,J=6.8Hz,3H). 13 C NMR(100MHz,CDCl3)δ199.74, 173.66, 170.84, 170.68, 170.49, 71.63, 64.29, 57.90, 46.06 , 39.32, 30.72, 30.55, 28.91, 26.35, 23.30, 22.64, 22.57, 19.42(2C), 18.16, 17.94, 17.70. HRMS(ESI+):C 22 H 39 Calculated exact mass of N3NaO6S: 496.2452. Found: 496.2457

[0395] Summary of Examples 1 to 8 We describe a strategy to genetically encode DAP in recombinant proteins. We show that by genetically encoding DAP in place of the catalytic cysteine ​​or serine, it is possible to trap unstable thioester or ester intermediates as their stable amide analogues. We demonstrate that cysteine ​​proteases and thioesters can be efficiently cleaved by genetic encoding. We illustrate the utility of this approach for enzymes and provide unique insight into intermediates in the synthesis of valinomycin by the Vlm TE. Our results reveal extensive rearrangements of the lid associated with the dodecapeptidyl-linked Vlm TE. Importantly, the DAP system can be used to synthesize valinomycin using both widely used reaction-competent substrates (e.g., native proteins containing protease sites), substrate analogs (in this case, SNAC), and commercially available natural products (here, valinomycin and possibly other cyclic products (Tseng, CC et al. Characterization of the surfactin synthetase C-terminal thioesterase domain as a cyclic depsipeptide synthase. Biochemistry 41, 13350-13359 (2002)) This allows the formation of near-native acyl-enzyme complexes.

[0396] These structures lack the PCP domain, which is central to the TE domain catalytic cycle, but its binding site is informative, as is the PCP-TE structure of EntF trapping a dead-end inhibitor (Liu, Y., Zheng, T. & Bruner, SD Structural basis for Inferred from phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488 (2011)) The PCP domain docks to the αE of the TE, and the PPE extends by about 15 Å, positioning the thiol close to Ser / DAP2463 (Fig. 7a, c-iii). The position of the PPE in the EntF structure is compatible with the dodecapeptide-bound conformation of the lid, but not with the apo / tetrapeptide-bound conformation in our Vlm TE structure. (The lid in the EntF structure is partially disordered.) This EntF structure showed how the PCP and TE domains can position the thioester of the depsipeptidyl-PPE close to Ser / DAP2463, but in Vlm these domains must also be able to position the terminal hydroxyl of the tetradepsipeptidyl-PPE close to Ser / DAP2463 for the oligomerization step. To do so, an additional ∼15 Å length of the tetradepsipeptide (between the terminal hydroxyl and the PPE sulfur) must be accommodated by the TE domain (compare Figures 7c-ii and 7c-iii). The lid is the same as the lid we have seen in the dodecapeptidyl-TE structure. DAP This could be facilitated, possibly using pockets similar to those observed in the structure.

[0397] Thus, the known structures can be assembled into a hypothetical pathway for oligomerization and cyclization (Figure 7). In the observed apo / tetradepsipeptide-bound conformation, Lα1 of the Vlm TE could inhibit any bound depsipeptide from curling around for cyclization (Figure 7). PCP binding could induce a TE conformation similar to that we observed for the dodecadepsipeptide-bound TE, which could accommodate and guide the tetradepsipeptidyl-PPE bound to the PCP domain of ∼30 Å toward the active site (Figure 7b). A transition to an open / mostly disordered lid (as seen in the EntF PCP-TE) could allow PCP to present a thioester back to Ser2463. Finally, a dodecadepsipeptide-TE with a hemispherical pocket could be constructed. DAP The lid conformation observed in the structure serves to curl the dodecadepsipeptide back toward Ser2463 for cyclization ( Figure 7 b, c-iv).

[0398] Dodecadepsipeptide-TE DAP The lid conformation and hemispherical pocket seen in the structure of valinomycin may be crucial during the cyclization step of the thioesterase cycle. This pocket is composed mainly of hydrophobic residues and provides steric hindrance that prevents the dodecadepsipeptide bound to Ser / DAP2463 from linearly extending (Fig. 7b). Instead, the lid conformation favors curling back the free end of the substrate toward the acyl bond between the TE and the substrate. Thus, the cyclization of the dodecadepsipeptide to valinomycin is entropically controlled by the pocket, and the dodecadepsipeptide is cyclized toward the acyl bond between the TE and the substrate. The conformation of the capeptide may be dictated by its partial confinement to the pocket and the active site of the TE domain.

[0399] Even when TE domains are covalently linked to bona fide substrates, other studies have shown that The mobility of the lid, seen in previous studies and dramatically seen here, and the lack of specific interactions between the lid and the rest of the TE domain make it unlikely that there is a single well-defined conformation at any of these steps in the synthesis cycle. The formation of a predetermined / templated conformation of the cyclization substrate has been proposed to promote cyclization in tyrocidine synthase (Trauger, JW, Kohli, RM, Mootz, HD, Marahiel, MA & Walsh, CT Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000); Trauger, JW, Kohli, RM & Walsh, CT Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001)), while specific interactions between the lid and the polyketide substrate have been proposed to achieve this in pikromycin synthase, there is no evidence for these mechanisms in Vlm TE. Indeed, specific and strong binding interactions could slow down the synthesis cycle because the tetradepsipeptide must transition back and forth between ligation to the PCP domain and ligation to the TE domain, and because the same tetradepsipeptide must adopt multiple different positions during one cycle. Rather, the conformation of the lid must fluctuate rapidly throughout the cycle, "breathing" and transiently assuming a reactive conformation. Interestingly, Mycobacterium tuberculosis polyketide synthase TE domains and novel inhibitors bind between clusters of lid helices (Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225 (2017)). Although it has been proposed to compete with substrate binding, such inhibitors could also act by preventing structural rearrangements of the lid similar to those we observed here.

[0400] Although we have focused on utilizing the encoded DAP to provide insight into the thioesterase acyl-enzyme intermediates in the synthesis of valinomycin, the DAP system has been extensively used to study a wide variety of enzymes featuring cysteine ​​or serine linked acyl-enzyme intermediates, including natural product megaenzyme domains, such as other cyclizing TE domains, transglutaminase homolog condensation domains and PKS ketosynthase domains (Holliday, GL, Mitchell, JBO & Thornton, JM Understanding the Functional Roles of Amino Acids. Residues in Enzyme Catalysis. Journal of molecular biology 390, 560-577 (2009) ) shows considerable promise for the study of DAPs. Extension of the approach reported here will facilitate structural and biochemical characterization of diverse acyl-enzyme intermediates. In addition, genetic encoding of DAPs in enzymes that proceed through acyl-enzyme intermediates (Cravatt, BF, Wright, AT & Kozarich, JW Activity-based protein profiling: from enzyme chemistry to proteomic chemistry. Annu Rev Biochem 77, 383-414 (2008)) but have unknown substrate specificities may allow for covalent capture and identification of native substrates. EXAMPLES

[0401] Expression, purification and activity test of UBE2L3-DAP GST-tag affinity purification of UBE2L3(C86DAP5) followed by cleavage of the GST-tag with TEV protease and purification by Strep-tag affinity purification was also tested. This strategy yielded a clean product. The mass of purified UBE2L3(C86DAP5)-Strep was determined by LC-ESI-MS: LC-ESI-MS of UBE2L3(C86DAP5)-Strep before UV light exposure: UBE2L3(C86DAP5)-Strep[1]: expected: 19050.57 Da, observed: 19048.79 Da; UBE2L3(C86INT);Strep[2]: expected: 18857.53 Da, observed: 18853.15 Da.

[0402] Deprotection of DAP5 occurs in two distinct steps. First, the photocaging group is removed under the action of UV light to give a semi-deprotected intermediate. The intramolecular reaction of 2 finally gives the fully deprotected DAP. After purification of UBE2L3(C86DAP5), most of the protein contains a semi-deprotected intermediate (UBE2L3[C86INT]), but some of it is in the fully photocaged form. After UV light irradiation, the protein mass was again assessed by LC-ESI-MS: LC-ESI-MS of UBE2L3(C86DAP5)-Strep after UV light irradiation: UBE2L3(C86INT)-Strep[2]: expected: 18857.53 Da, observed: 18853.15 Da.

[0403] As expected, after UV light irradiation, UBE2L3(C86DAP5)-Strep was not detectable. In fact, only UBE2L3(C86INT)-Strep could be detected. The proteins were then incubated at 37°C for 3 h and their mass was assessed by LC-ESI-MS. As expected, UBE2L3(C86DAP)-Strep was detected together with UBE2L3(C86INT)-Strep. LC-ESI-MS of UBE2L3(C86DAP5)-Strep after UV light irradiation and incubation at 37 °C for 3 h: UBE2L3(C86INT)-Strep[2]: expected: 18857.53 Da, observed: 18853.15 Da; UBE2L3(C86DAP)-Strep[3]: expected: 18753.59 Da, observed: 18753.13 Da.

[0404] Unfortunately, longer incubation times at 37°C (6h and 16h) did not lead to improved deprotection of UBE2L3(C86INT)-Strep. In fact, the proportion of proteins containing DAP did not seem to change. This corresponds to about 30% of the total protein LC-ESI-MS of UBE2L3(C86DAP5)-Strep after UV light irradiation and longer incubation times. UBE2L3(C86INT)-Strep was incubated for 6h or 16h at 37°C. No improvement in deprotection was observed and the proportion of proteins containing DAP (about 30%) remained almost unchanged.

[0405] To test whether Ub could be charged to UBE2L3(C86DAP5) following UV light irradiation and overnight incubation at 37° C., reactions were set up containing 0.2 μM E1 and HA-tagged Ub in E2 loading buffer. Reactions (positive control [wt], negative controls [C86A] and [C86DAP]) were performed with or without Ub (see FIG. 23).

[0406] As expected, a higher molecular weight band corresponding to the thioester-linked E2-Ub conjugate was observed in UBE2L3(wt). In addition, a higher molecular weight band corresponding to the isopeptide-linked E2-Ub conjugate was also detected in UBE2L3[C86DAP]. The incomplete conversion of UBE2L3 to the Ub conjugate (Figure 23) is consistent with the incomplete deprotection of DAP5 in UBE2L3. The newly formed isopeptide bond between UBE2L3(C86DAP) and Ub is redox insensitive and in the presence of β-mercaptoethanol It cannot be reduced. It is redox sensitive. This is very different from the complex formed from E2L3(wt) and Ub (Figure 23). To further characterize the identity of the distinct bands, both anti-HA and anti-UBE2L3 blots were performed (Figure 23B and C), clearly showing that the higher molecular weight band formed in the presence of both HAUb and UBE2L3(C86DAP) contains UBE2L3 and Ub.

[0407] Finally, to characterize the chemical nature of the newly formed bond, the band corresponding to the UBE2L3(C86DAP)-Ub complex was excised and analyzed by tandem mass spectrometry after trypsin digestion (performed by the Proteomics Facility, University of Bristol). Tandem mass spectrometry of isopeptide-linked UBE2L3(C86DAP)-Ub complexes. Tandem mass spectrometry clearly identifies DAP modification at the desired site and the expected Gly-Gly modification at residues consistent with Ub loading onto the DAP. This analysis clearly confirmed the formation of a stable amide bond between UBE2L3(C86DAP) and Ub. EXAMPLES

[0408] Applications in live cells In this example, we demonstrate the technique in living cells.

[0409] In this example, we demonstrate the invention in E. coli cells (BL21) and mammalian cells (HEK293T).

[0410] We refer to the TEV-GFP WB data shown in FIGS.

[0411] In particular, we refer to Figure 24 which shows Dap-mediated substrate capture in live E. coli cells. GFP (GFP has a TEV cleavage site at its C-terminus) and various variants of C-terminally Strep-tagged TEV protease (WT / Ala / TAG (with or without Dappc; (in this example, compound "DAP5" is referred to as "Dappc")) were co-expressed in E. coli BL21 cells at 20°C. After 20 h of expression, the cells were exposed to UV light (35 mW / cm 2 ) were directly irradiated for 2 min and shaken at 37°C. Equal volumes of cells were harvested at the indicated time points and analyzed by Western blot (anti-Strep against TEV and anti-GFP). Only TEV(Dap) showed UV-dependent generation of the TEV-GFP conjugate. This conjugate was detectable within 10 min after UV irradiation, and the reaction was complete within 2 h in E. coli BL21 cells.

[0412] Furthermore, we refer to Figure 25 which shows Dap-mediated substrate capture in mammalian HEK293T cells. HEK293T cells were co-transfected with GFP (GFP with a TEV cleavage site at the C-terminus) and different variants of C-terminally Strep-tagged TEV protease (WT / Ala / TAG (with or without Dappc (DAP5))). 48 h after transfection, the cells were exposed to UV light (8 mW / cm 2 ) for 2 min. Cells were then incubated at 37°C and harvested at the indicated time points. Cells were lysed and TEV was pulled down by StrepTactinXT. Pull-down results were analyzed by Western blot (anti-Strep against TEV and anti-GFP). Only TEV(Dap) showed UV-dependent generation of TEV-GFP conjugates. The conjugates began to form within 30 min after UV irradiation and were enriched with increasing incubation times in HEK293T cells.

[0413] Additional Methods of Example 10: TEV(Dap)-GFP in E. coli sub capture BL21 ( DE3) cells were induced to express the protein at 20°C. 0.1 mM Dappc (DAP5) was added to the medium to incorporate Dappc (DAP5). After 20 h, cells were transferred to a Falcon 50 mL conical centrifuge tube and exposed to UV light (365 nm, 35 mW / cm) for 2 min with gentle agitation. 2 ) were irradiated. The cells were then centrifuged at 5,000g for 5 min. The supernatant was discarded and the pellet was resuspended in fresh medium with freshly added antibiotics. The cell cultures were shaken at 37°C. At each time point indicated, 5 mL of cell culture was harvested and lysed in BugBuster (Merck). Total lysates were analyzed by WB (anti-Strep (ab76949, abcam) and anti-GFP (ab13970, abcam)).

[0414] TEV-GFP in HEK293T cells sub capture HEK293T cells were co-transfected with a plasmid containing GFP and a plasmid containing TEV. For amber suppression, 1 mM Dappc (DAP5) was added 30 min after transfection. 48 h after transfection, cells in 6-well plates were exposed to UV light (365 nm, 10 mW / cm2). 2 ) for 2 min. The medium was then replaced with fresh medium and incubated at 37°C. At each time point indicated, cells were harvested and lysed in NP lysis buffer (Cat. No. 87787, Thermo). Total lysates were used for StrepTactinXT pulldown. The eluates from the beads were stained by WB (anti-S The results were analyzed using trep (ab76949, Abcam) and anti-GFP (ab13970, Abcam).

[0415] Supplementary method List of primers used in this study. Mutated residues are shown in uppercase.

[0416] [Table 2]

[0417] Construction of DAPRSlib library by inverse PCR Plasmid pBK-pylS was used as a template (Phan, J. et al. Structural basis for the substrate specificity of tobacco etch virus protease. Journal of Biological Chemistry 277, 50564-50572 (2002)) to generate a library of 6 amino acids (DAPRSlib) by sequential five-round inverse PCR reactions using PrimeSTAR HS DNA polymerase (Takara Bio) according to the manufacturer's guidelines. The primers randomized the codons at positions Y271, N311, Y349, V366, and W382 of the pylS gene to the codons for all 20 natural amino acids ( (All primers are listed in Supplementary Table 1). The resulting PCR products were digested with BsaI-HF and DpnI and circularized with T4 DNA ligase. DNA was transformed into Eletrocompetent MegaX DH10B™ T1R Electrocomp™ E. coli cells (Invitrogen) according to the manufacturer's instructions and plasmid DNA was prepared by inoculating overnight cultures containing the appropriate antibiotic. Diversity was estimated by plating serial dilutions of the transformed rescue cultures on LB-agar plates containing the appropriate antibiotic. Ten isolated 8 The library of species transformants encompassed the theoretical diversity of the library with 97% confidence.

[0418] Selection of active aaRS by DAP derivatives Selection of synthetase mutants specific for amino acids 2 to 6 was performed as previously reported (Phan, J. et al. Structural basis for the substrate specificity of tobacco etch virus protease. Journal of Biological Chemistry 277, 50564-50572 (2002)) using the following libraries: DAPRSlib (Y271, N311, Y349, V366, W382), D3 (L270, Y271, L274, N311, C313), PylS fwd (A267, Y271, L274, C313, M315), Susan 1 (A267, Y271, Y349, V366, W382), Susan 2 (N311, C313, V366, W382, G386), and Susan 4 (A267, Y349, S364, V366, G386). The S library was subjected to five rounds of alternating positive and negative selection. Positive selection was performed using a chloramphenicol acetyltransferase reporter that has an amber codon (codon 112) at a permissive position and expresses the cognate tRNA in the presence of the desired ncAA (1 mM). Cells that survived positive selection on chloramphenicol (typically 50 μg / mL) LB agar are predicted to use either natural amino acids constitutively present in the cells or ncAAs added to the cells. The negative selection used removed synthetase variants that use natural amino acids using a barnase reporter that contains an amber codon and provides the cognate tRNA in the absence of ncAA.

[0419] Expression and purification of GFP(150TAG)His6 Superfolder green fluorescent protein (sfGFP) with 6 incorporated at position 150 was expressed from pSF-sfGFP150TAG in MegaX DH10B T1R cells containing pBK_DAPRS or pBK_PylRS vectors. mL kanamycin and 1 mM 6 or N ε The transformed cells were inoculated into LB broth supplemented with 220-tert-butyloxycarbonyl-lysine (BocK). 0.2% (w / v) L-(+)-arabinose (Sigma) was added at 37°C for 16 h. Expression was induced with shaking at rpm. The bacteria were then harvested and the protein purified by polyhistidine affinity chromatography.

[0420] Expression and purification of His6-lipoyl-TEV-Strep BL21(DE3) cells were transfected with pNHD-His6-lipoyl-TEV wt -Strep, pNHD-His6-Lipoyl-TEV Ala -Strep (gene kindly provided by Mark Allen) (Trauger, JW, Kohli, RM, Mootz, HD, Marahiel, MA & Walsh, CT Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000)) or pSF-DAPRS-PylT (Zhou, Y., Prediger, P., Dias, LC, Murphy, AC & Leadlay, PF Macrodiolide formation by the thioesterase of a modular polyketide synthase. Angew Chem Int Ed Engl 54, 5232-5235 (2015)) pNHD-His6-lipoyl-TEV Amber -Strep and grown overnight at 37°C on TB-agar plates containing 25 μg / mL tetracycline and (and 50 μg / mL kanamycin for cotransformed cells). Several transformed colonies were inoculated into TB-medium containing 12.5 μg / mL tetracycline (and 25 μg / mL kanamycin and 100 μM 6 for cotransformed cells) and incubated at 37°C; OD 600 Once the pH reached 0.5-0.7, the cultures were shifted to 20°C. After a further incubation of 30 min, the cultures were induced using 250 μM isopropyl β-D-1-thiogalactopyranoside (IPTG) and protein expression was carried out for 16 h at 20°C. Cells were harvested by centrifugation and resuspended in 50 mM tris-HCl (pH 7.5), 150 mM NaCl, 2 mM β-mercaptoethanol, 1 Roche Inhibitor Cocktail tablet / 50 mL, 0.5 mg / mL lysozyme (Sigma), 50 μg / mL The cells were resuspended in DNase (Sigma) and lysed by sonication. The mixture was clarified by centrifugation at 9,000 × g for 30 min and filtered through a 0.4 μm polyethersulfone (PES) membrane. His6-Lipoyl-TEV-Strep was purified using nickel affinity chromatography (HisTrap HP column, GE Healthcare). The protein was purified by a linear gradient of midazole (0 mM to 500 mM). The fractions containing the protein were further purified by Strep-tag affinity purification using a 5 mL StrepTrap HP column (GE Healthcare). After loading the sample, the column was The column was washed with rep binding buffer (50 mM 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid [HEPES] (pH 8.0), 150 mM NaCl, 1 mM ethylenediaminetetraacetic acid [EDTA], 5 mM dithiothreitol [DTT]). Proteins were eluted using a linear gradient of desthiobiotin (0 mM to 1.25 mM). His6-lipoyl-TEV Amber -For Strep, at the end of the purification, the protein is exposed to UV light (365 nm, 35 mW cm-2 , 1 min).

[0421] Ub tev Expression and purification BL21(DE3) cells were transformed with pNHD-Ub-tev-His6 and grown overnight at 37°C on LB agar plates containing 25 μg / mL tetracycline. LB medium containing 25 μg / mL tetracycline was inoculated with several colonies from the transformation. The culture was diluted 1:100 into fresh LB medium containing 12.5 μg / mL tetracycline; OD 600 When the NA reached 0.5, the cultures were induced with 1 mM IPTG and protein expression was carried out for 6 h at 37°C. Cells were harvested by centrifugation, resuspended in 50 mM tris-HCl (pH 7.5), 150 mM NaCl, 2 mM β-mercaptoethanol, 1 Roche Inhibitor Cocktail tablet / 50 mL, 0.5 mg / mL lysozyme (Sigma), 50 μg / mL DNase (Sigma) and lysed by sonication. Lysates were clarified by centrifugation at 39'000×g for 30 min and filtration through a 0.4 μm PES membrane. Ub was purified by cleavage of the imidazoline using nickel affinity chromatography (HisTrap HP column, GE Healthcare). The protein was purified by dialysis overnight against 10 mM tris-HCl at 4°C and purified by ion exchange chromatography (HiTrapS in 50 mM ammonium acetate (pH 4.5) using a 5 mL column (GE Healthcare) Ub was further purified using a NaCl gradient (0–1 mM). Pure fractions were pooled and then dialyzed overnight against 20 mM tris-HCl (pH 7.4). The sample was then concentrated to approximately 15 mg / mL using Amicon Ultra-15 (3 kDa MWCO) centrifugal filter devices (Millipore).

[0422] TEV and Ub tevReaction with 15 μg His6-lipoyl-TEV-Strep with 60 μg Ub tev The reaction was incubated at 30°C with 150 μL of 50 mM HEPES (pH 8.0), 150 mM NaCl, 1 mM EDTA, and 5 mM DTT overnight. 20 μL of the reaction was loaded onto a 4-12% NuPAGE Bis-Tris gel (Invitrogen) and the 2-(n-morpholino) Electrophoresis was performed for 45 min in MES buffer. Proteins were transferred to a polyvinylidene difluoride (PVSF) membrane (Roche) using 25 mM Tris (pH 8.2), 192 mM glycine, 10% (v / v) methanol. The membranes were blocked for 1 h in TBST buffer (25 mM Tris (pH 7.4), 150 mM NaCl, 0.05% [v / v] Tween 20) containing 5% (w / v) powdered milk at room temperature. Antibodies (Strep-Tactin-HRP conjugate (αStrep) [IBA Lifesciences] or P4D1 antibody (αUb) [Enzo Life Sciences]) were added in 5% TBST-milk and incubated overnight at 4°C. Secondary antibodies (for αUb antibodies) were added in 5% TBST-milk and incubated for 1 h at room temperature. Blots were developed using Amersham enhanced chemiluminescence (ECL) (GE Healthcare) and a ChemiDoc XRS+ gel imaging system (Bio-Rad). did.

[0423] Analysis of intracellular concentrations of DAP derivatives Analysis of intracellular concentrations of DAP derivatives was performed as previously described (Zhou, Y., Prediger, P., Dias, LC, Murphy, AC & Leadlay, PF Macrodiolide formation by the thioesterase of a modular polyketide synthase. Angew Chem Int Ed Engl 54, 5232-5235 (2015)). Briefly, the DAP derivatives were added to a final concentration of 1 mM in 5 mL of LB medium. A control sample was also prepared using 5 mL of unsupplemented LB medium. Each solution was inoculated with DH10B cells. The cultures were stirred at 220 rpm in the dark at 37°C for 12 h. The OD of each sample was 600 was determined and cells from each culture were harvested. The cell pellet was washed three times with 1 mL of fresh ice-cold LB medium by cycles of resuspension and centrifugation. The washed cell pellet was resuspended in a methanol:water (60:40) solution. Zirconium beads (0.1 mm) were added to each suspension. The suspension was vortexed for 12 min to lyse the cells. The lysate was centrifuged at 21000×g for 30 min at 4°C. The supernatant was carefully removed and placed in a new 1.5 mL Eppendorf tube. The solution was centrifuged again at 21000×g for 2 h at 4°C. 100 μl aliquots of the supernatant from the resulting samples were analyzed by LC-ESI-MS. A gradient of 0.5% to 95% acetonitrile in water was applied to elute the clarified lysate from a Zorbax C18 (4.6×150 mm) column. 1OD 600 8 x 10 cells per unit 8 Estimated values ​​and 0.6×10 -15 The cell volume was used to estimate the concentration.

[0424] Cloning, expression and purification of Vlm TE constructs vlm2 PCP4-TEA codon-optimized construct containing (encoding residues 2290–2655 of Vlm2 from Streptomyces tsushimaensis, GenBank: ABA59548.1) was transfected with ATUM (formerly DNA2.0) in the pJExpress411 vector, containing an N-terminal hexahistidine tag followed by a tobacco etch virus protease (TEV) cleavage recognition sequence (pJExpress411-vlm2-PCP4-TE wt ) The plasmid was synthesized as pJExpress411-vlm2-PCP4-TE. wt The nucleotides at positions 2024-2025 and The DNA contained two BamHI recognition sequences at positions 2267-2268 of the 5'-terminal end ... The plasmid pJExpress411-vlm2-TE contains the PCP4 domain sequence truncated and encodes residues 2368 to 2655 of Vlm2. wt was obtained. TE DAP To generate the expression vector for vlm2, pJExpress411-vlm2-TE wt Using primers TE_for_pNHD_fw and TE_for_pNHD_rev, wt The coding sequence was amplified by PCR. The PCR product was digested with NdeI and XhoI and ligated into the similarly digested pNHD plasmid using T4 DNA ligase to produce the plasmid pNHD-vlm2-TE wt Next, an amber stop codon was introduced in place of the codon for serine 2463 by site-directed mutagenesis using primers Vlm2_TE_Amb_Fw and Vlm2_TE_Amb_Rev to generate pNHD-vlm2-TE amber2463 was generated.

[0425] pJExpress411-vlm2-TE wt (T.E. wt ) or pNHD-Vlm2-TE amber2463 Reach and pSF-DAPRS-PylT(TE DAP The TE domain was heterologously expressed in E. coli BL21(DE3) cells cotransformed with TE wt The culture expressing -1 The strains were grown in LB medium supplemented with 100 μg of kanamycin. DAP The substances expressing the above were selected at 25 mg L -1 Kanamycin, 12.5 mg L -1 Cultures were grown in TB medium supplemented with 0.1 mM tetracycline, 0.1 mM 6 (a 100 mM stock solution of 6 was prepared in 0.4 M NaOH, added to the cultures, and neutralized using 5 M HCl). 600nm The cultures were incubated at 37°C with stirring at 220 rpm until the β-actin ratio reached 0.6, then incubated at 16°C for 30 min, and then expression was induced with 100 μM IPTG. The cultures were incubated at 16°C for a further 16 h before being harvested by centrifugation at 5000g for 20 min. The cell pellets were stored at -80°C.

[0426] For protein purification, TE wt The cell pellet was resuspended in 5 mL of buffer wt-A (50 mM TRIS (pH 7.4), 150 mM NaCl, 50 mM imidazole, 2 mM β-mercaptoethanol [βME]) + DNAse I (Bioshop) per gram of wet cells and lysed by sonication. The lysate was sonicated at 40,000 g for 20 min. The lysate was clarified by centrifugation at 1000 rpm. The clarified lysate was applied to two 5 mL HiTrap IMAC FF (GE Healthcare Life Sciences) columns connected in series on an AKTA Prime system (GE Healthcare Life Sciences). Bound proteins were eluted with Buffer wt-B (Buffer wt-A + 150 mM imidazole). TE wtFractions containing Te (determined by SDS-PAGE analysis) were pooled, incubated with TEV protease at a mass / mass ratio of 1:100 (Te:TEV) and dialyzed against buffer wt-C (50 mM TRIS (pH 7.4), 10 mM NaCl, 2 mM βME) for 16 h at 4 °C. The dialyzed sample was applied to two 5 mL HiTrap IMAC FF columns connected in series pre-equilibrated with buffer wt-A. Cleaved proteins were collected from the flow-through and transferred to two 5 mL HiTrap Q HP columns connected in series pre-equilibrated with buffer QA (50 mM TRIS (pH 7.4), 10 mM NaCl, 2 mM βME). The protein was eluted with a gradient of 0-100% of buffer QB (50 mM TRIS (pH 7.4), 500 mM NaCl, 2 mM βME) over 240 mL. wt Fractions containing TE were concentrated with a 10 kDa molecular weight cut-off Amicon® Ultra centrifugal filter (Millipore) and loaded onto a Superdex S-200 16 / 60 PG column (GE-Healthcare) pre-equilibrated with SEC buffer (25 mM HEPES (pH 7.4 or pH 8.0), 100 mM NaCl, 0.2 mM tris(2-carboxyethyl)phosphine [TCEP]). wt Fractions containing were pooled, concentrated and flash frozen.

[0427] TE DAP Cell resuspension, lysis, clarification, and Ni-IMAC purification were performed in the same manner as described above except that prolonged exposure to light was avoided. wt After elution from the Ni-IMAC column, the samples were exposed to UV light (365 nm, mW cm -2 The TEV cleavage and subsequent IMAC column were performed using TE, except that a 1:1 TE:TEV ratio was used. wt Anion exchange was performed as described for Tetrahydrofuran (Te) except that 25 mM HEPES replaced TRIS as the buffer and 0.2 mM TCEP replaced βME as the reducing agent in the mobile phase.wt The procedure was carried out as described for . The relevant fractions were concentrated and injected onto a Superdex S-7510 / 300 column pre-equilibrated with buffer T (25 mM HEPES (pH 8.0), 100 mM NaCl, 0.2 mM TCEP). Purified TE DAP Fractions containing TE were pooled, concentrated and used immediately for further experiments. wt The yield of purified TE is 30-60 mg per liter. DAP The yield was 0.1-0.5 mg per liter.

[0428] Crystallography TE wt The crystallization conditions for structure 1 were as follows: in the vapor diffusion crystallization test, a commercial screen (Qiagen) and 10 mg mL -1 Optimization of initial crystallization hits in 24-well plates yielded a protein concentration of 3.2 μL of 10 mg mL -1 TE wt The final crystallization conditions were obtained by incubating 4.0 μL of 1.65 M DL-malic acid (pH 9.5) and 0.8 μL of 17% m / v IPTG in a reservoir solution of 500 μL of 1.65 M DL-malic acid (pH 9.5). wt Structure 2 crystals, 0.5 μL of 22.4 mg mL -1 Purified TE wt and were grown under similar conditions by incubating 0.5 μL of 1.65 M DL-malic acid (pH 8.1) against a reservoir solution of 500 μL DL-malic acid (pH 8.1). Crystals appeared between 24 and 48 hours and reached maximum size in approximately one week.

[0429] Ligand-free TE DAP The crystals were grown in TE using a reservoir solution of 1.65 M DL-malic acid (pH 8.0). wt The tetradepsipeptidyl-TE DAP To obtain the complex structure, DAPOnce the crystals reached maximum size, they were incubated with deoxytetradepsipeptidyl-SNAC 8. The reservoir solution was exchanged for 2.66 M DL-malate (pH 9.5) and 32 μL of a solution of 1 mM deoxytetradepsipeptidyl-SNAC, 2.66 M DL-malate (pH 9.5), 100 mM NaCl, 25 mM HEPES (pH 9.2), 10% DMSO was added to the drop. The crystals were incubated in these conditions at room temperature for 9 days.

[0430] Dodecadepsipeptidyl-TE DAP For complex crystals, TE DAP (0.1 mg mL -1 ) to 1.1 mg mL of valinomycin in buffer T -1 The suspension was incubated at room temperature for 16 h. The sample was centrifuged at 20000 g and applied to a Superdex S-75 10 / 300 column pre-equilibrated with buffer T to remove excess valinomycin. Relevant fractions were pooled and complex formation was assessed by LC-ESI-MS (see below). Samples were diluted to 13.4 mg mL -1 Concentrate to TE wt Diffraction quality crystals with distinct morphology from the crystals were analyzed by equilibrating 1 μL of dodecadepsipeptidyl-TE against 500 μL of reservoir solution. DAP The crystals were obtained in sitting drops consisting of the complex plus 1 μL of reservoir solution (1.30–1.45 M DL-malate, pH 8.1). In an attempt to improve ligand occupancy, a subset of these crystals was further incubated with valinomycin for 24 h by the addition of 20 μL of a solution containing 555 μM valinomycin, 2 M DL-malate, pH 8.1, 11 mM HEPES, pH 8.0, 44 mM NaCl, and 0.088 mM TCEP.

[0431] TE wt and dodecadepsipeptidyl-TE DAP Add 10 μL of crystals (TE wt Structure 1 and dodecadepsipeptidyl-TE DAP ) or 20 μL (TEwt Structure 2) was cryoprotected by adding 3.6 M DL-malic acid (pH 8.1). Dodecadepsipeptidyl-Te incubated with valinomycin DAP For crystals, the drop solution was removed and replaced with 10 μL of 3.6 M DL-malic acid. The crystals were equilibrated for at least 2 min and then flash cooled in liquid nitrogen. Tetradepsipeptidyl-TE DAP The complex crystals were looped and flash-cooled directly from the incubation solution. wt Data were initially collected at the Centre for Structural Biology, McGill University, Montreal, Canada with a Rigaku RUH3R generator and R-AXIS IV++ detector. wt and depsipeptidyl-TE DAP Higher Complex Resolution data were collected using a Pilatus detector at the Canadian Light Source (CLS) 08ID-1 beamline or the Advanced Photon Source (APS) NE-CAT 24-ID-C beamline (Extended Data Table 1).

[0432] TE wt Structure determination TE wt Diffraction data from the structure 1 crystal were indexed and analyzed using iMosflm (Trauger, JW, Kohli, RM & Walsh, CT Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001)) or DIALS (Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225 (2017)) was used to integrate the ribozyme into space group P432. Intergroup determination and scaling were performed using the programs POINTLESS and SCALA (Cravatt, BF, Wright, AT & Kozarich, JW Activity-based protein profiling: from enzyme chemistry to proteomic chemistry. Annu Rev Biochem 77, 383-414 (2008)). srfA-C (PDB The structure was solved using a modified version of SCULPTOR (Pendrak, I., Wittrock, R. & Kingsbury, WD Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995)) of the TE domain of ID 2VSQ (Alonzo, DA, Magarvey, NA & Schmeing, TM Characterization of cereulide synthetase, a toxin-producing macromolecular machine. PloS one 10, e0128569, doi:10.1371 / journal.pone.0128569 (2015)) as a search model.Program Phenix(Shaw-Reid, CA et al. Assembly line enzymology by multimodular nonribosomal peptide synthetases: the thioesterase domain of E. coli EntF catalyzes both elongation and cyclolactonization. Chemistry & biology 6, 385-400, doi:10.1016 / S1074-5521(99)80050-7 (1999)) and AUTOBUILD (May, JJ, Wendrich, TM & Marahiel, MA The dhb operon of Bacillus subtilis encodes the biosynthetic template for the catecholic siderophore 2,3-dihydroxybenzoate-glycine-threonine trimeric ester bacillibactin. J Biol Chem 276, 7209-7217, The structure was iteratively refined and built using the Sigma-Aldrich (doi:10.1074 / jbc.M009140200 (2001)) and Coot (Zhou, Y. et al. Iterative Mechanism of Macrodiolide Formation in the Anticancer Compound Conglobatin. Chemistry & biology 22, 745-754, doi:10.1016 / j.chembiol.2015.05.010 (2015)) for iterative model building.TE using TopDraw (Robbel, L., Hoyer, KM & Marahiel, MA TioS T-TE--a prototypical thioesterase responsible for cyclodimerization of the quinoline- and quinoxaline-type class of chromodepsipeptides. FEBS J 276, 1641-1653, doi:10.1111 / j.1742-4658.2009.06897.x (2009)). wt Topology diagrams were generated based on results from PDBsum generate ( http: / / www.ebi.ac.uk / thornton-srv / databases / pdbsum / Generate.html ) using the structures as input.

[0433] Dodecadepsipeptidyl-TE DAP Diffraction data sets collected from the complex crystals were indexed into the P1 or H3 space group using iMosflm (Trauger, JW, Kohli, RM & Walsh, CT Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001)) or DIALS (Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225 (2017)). Most crystals from group H3 showed evidence of twinning, and only untwinned diffraction data were used for structure determination. A TE missing residues 2500-2647 was identified by molecular replacement using PHASER (McGall, GH et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997)). wt The structures were used as search models to solve structures in both the P1 and H3 space groups. The P1 structures had six molecules in the asymmetric unit, whereas the H3 structures contained two molecules in the asymmetric unit. All depsipeptidyl-TEs DAP The model was refined and mF o -F c Generate the map, then depsipet The peptide residues were constructed in the model (Fig. 4c, d and Extended Data Fig. 8). The depsipeptides were modeled using the individual monomers (DPP, VAL, 2OP The monomer libraries and the linkage constraints between the monomers (i.e., DPP->VAL, VAL->2OP, 2OP->DVA, DVA->VAD and VAD->VAL) were constructed using AceDRG (Liu, Y., Zheng, T. & Bruner, SD Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module). of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488, doi:10.1016 / j.chembiol.2011.09.018 (2011)) and merged using LIBCHECK (Trauger, JW, Kohli, RM & Walsh, CT Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001)). The resulting merged dictionary was then compared to the substrate builds in Coot (Zhou, Y. et al. Iterative Mechanism of Macrodiolide Formation in the Anticancer Compound Conglobatin. Chemistry & biology 22, 745-754, doi:10.1016 / j.chembiol.2015.05.010 (2015)) and REFMAC5 (Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225, doi:10.1016 / j.cell.2017.06.025 (2017)) and and phenix.refine (Shaw-Reid, CA et al. Assembly line enzymology by multimodular nonribosomal peptide synthetases: the thioesterase domain of E. coli EntF catalyzes both elongation and cyclolactonization. Chemistry & biology 6, 385-400, doi:10.1016 / S1074-5521(99)80050-7 (1999)). The final statistics are shown in Extended Data Table 1.

[0434] Formation of depsipeptidyl-TE complexes Final concentration 0.2mg·mL -1 TE DAP or TE wt 1.7% or 1% v / v Deoxytetradepsipeptidyl-SNAC 7 (1.7 mM) or valinomycin (50 μM) were incubated for 16 h in Buffer T containing DMSO. Reactions were concentrated in 10 kDa molecular weight cut-off Amicon® Ultra centrifugal filters (Millipore), clarified by centrifugation at 20000 g, and applied to a Superdex S-75 10 / 300 column pre-equilibrated with Buffer T to remove excess depsipeptidyl-SNAC or valinomycin before final LC-ESI-MS analysis. Deoxytetradepsipeptidyl-TE DAP To form the complex, a final concentration of 8.7 mg mL -1 TE DAP was incubated with deoxytetradepsipeptidyl-SNAC 8 (2.6 mM) in 25 mM HEPES (pH 8.6), 100 mM NaCl, 3.8% v / v DMSO for 40 h. Samples were diluted with 100 mM ammonium bicarbonate (pH 8.0) before final LC-ESI-MS analysis. All incubations were performed at room temperature.

[0435] LC-ESI-MS analysis of intact proteins For the experiments shown in Fig. 2c and Extended Data Figs. 3b and 7c, protein samples were run on a liquid chromatography (LC) system (Agilent 1200 Series) followed by a 6130 Quadrupole column. The proteins were subjected to in-line electrospray ionization mass spectrometry (ESI-MS) on a spectrometer. A Jupiter 5μ C4 300A column, 150 mm × 2.00 mm (Phenomenex) was used to run the proteins on the LC system using a gradient of 0.1% (v / v) formic acid in water (solvent A) and 0.1% (v / v) formic acid in acetonitrile (solvent B) (10% to 75% in 6 min and 75% to 95% in 1.5 min). Proteins were detected by monitoring UV absorbance at 200 and 280 nm. Protein mass was determined by M in positive ion mode using OpenLAB CDS software (Agilent Technologies). Calculated by deconvolution from the S acquisition.

[0436] For the experiments shown in Fig. 4a, b and Extended Data Figs. 6d–f, 7d, e, protein concentrations in Buffer T at 0.1 mg mL -1 and a Bruker Amazon Speed ​​ETD ion tracker. 9-well platelet counts were measured on an Agilent Technologies 1260 Infinity HPLC system coupled to a mass spectrometer. An Agilent PLRP-S (1000A 5 μM, 50 × 2.1 mm I D) 16 uL was injected onto the column. MS data was collected using ExtremeScan mass range mode. The ESI system was collected with positive ion polarity, scan range 50–3000 m / z, accumulation time 1586 μs, RF level 96%, trap drive 69.8, and PSP target mass 922 m / z, averaging five or more spectra. Calibration of the external equipment was performed with an Agilent ESI tuner. The run was performed using a 50% gradient of mobile phase B from 5% to 100% mobile phase B, followed by an isocratic step of 100% mobile phase B for 8 min. The temperature of the column compartment was set at 80 °C throughout the run. After injection, the column was washed for 5 min under the initial HPLC conditions with the sample compartment switching valve in the waste position. Then, a 5 min gradient of mobile phase B from 5% to 100% was performed, followed by an 8 min isocratic step of 100% mobile phase B. Proteins were detected by monitoring UV absorbance at 280 nm. Data were analyzed using Bruker DataAnalysis software (Bruker). Mass spectra was integrated from 10.5 to 3 min and deconvoluted using a window between 10,000 and 40,000 m / z.

[0437] Tandem MS / MS analysis Proteins were analyzed on 4–12% NuPAGE Bis-Tris gels (Invitrogen) containing MES buffer. The gel was electrophoresed on a 500 s PBS and briefly stained with InstantBlue (Expedeon). The samples were stored in 20 mM Tris (pH 7.4). Trypsin digestion and tandem MS / MS analysis were performed by Kate Heesom (Proteomics Facility, University of Bristol). It was said.

[0438] LC-ESI-MS analysis of Vlm TE reaction products 0.2 mg mL -1 (6.5 μM) purified TE wt or TE DAPwere incubated with tetradepsipeptidyl-SNAC 7 (1.7 mM) or a mixture of tetradepsipeptidyl-SNAC 7 and deoxytetradepsipeptidyl-SNAC 8 (1.7 mM each) in buffer T. Samples were incubated at room temperature for 24 h and then quenched with 1 volume of 0.1% formic acid in acetonitrile. Samples were then centrifuged at 20,000 g, flash frozen in liquid nitrogen, and stored at -80°C prior to HPLC analysis. For HPLC-MS analysis, frozen samples were thawed at room temperature, vortexed, and clarified by centrifugation at 20,000 g prior to injection. HR-LC-ESI-MS was performed in positive ESI mode at the Mass Spectroscopy Facility (Department of Chemistry, McGill University) on a Dionex Ultimate 3000U HPLC system coupled to a Bruker maXis impact QTOF mass spectrometer using an Agilent XDB-C8 (5 μm, 4.6 × 150 mm) column. Ion trap LC-ESI-MS analysis was performed on a Bruker Amazon Speed ​​ETD ion trap mass spectrometer. Positive ESI was measured on an Agilent Technologies 1260 Infinity HPLC system coupled to a The run was performed in HPLC mode. The column compartment was set at 40 °C throughout the entire run. The starting HPLC conditions were 50% mobile phase A (0.1% formic acid in HO), 50% mobile phase B (0.1% formic acid in acetonitrile). Injection (1 μL for HR-LC-ESI-MS, 5 μL for ion trap LC-ESI-MS) was followed by a gradient of 50% to 98% mobile phase B in 5 min, followed by an isocratic step of 98% mobile phase B for 20 min. For HR-LC-ESI-MS, Na +At the beginning of the first analysis using formate, an internal calibration was performed using an in-run injection, and the resulting calibration was used as the external calibration for subsequent analyses. External calibration of the ion trap was performed using an Agilent ESI tune mix. Data were analyzed using Bruker DataAnalysis software and the SmartFormula tool (Bruker).

[0439] References for Examples 1-9 1 Holliday, GL, Mitchell, JBO & Thornton, JM Understanding the Functional Roles of Amino Acid Residues in Enzyme Catalysis. Journal of molecular biology 390, 560-577 (2009) 2 Di Cera, E. Serine proteases. IUBMB Life 61, 510-515 (2009) 3 Hedstrom, L. Serine protease mechanism and specificity. Chem Rev 102, 4501-4523 (2002) 4 Long, JZ & Cravatt, BF The Metabolic Serine Hydrolases and Their Functions in Mammalian Physiology and Disease. Chem Rev 111, 6022-6063 (2011) 5 Verma, S., Dixit, R. & Pandey, KC Cysteine ​​Proteases: Modes of Activation and Future Prospects as Pharmacological Targets. Front Pharmacol 7 (2016) 6 Otto, H.H. & Schirmeister, T. Cysteine proteases and their inhibitors. Chem Rev 97, 133-171 (1997) 7 Swatek, K.N. & Komander, D. Ubiquitin modifications. Cell Res 26, 399-422 (2016) 8 Yang, W. & Drueckhammer, D.G. Understanding the relative acyl-transfer reactivity of oxoesters and thioesters: computational analysis of transition state delocalization effects. Journal of the American Chemical Society 123, 11004-11009 (2001) 9 Liu, B., Schofield, C.J. & Wilmouth, R.C. Structural analyses on intermediates in serine protease catalysis. Journal of Biological Chemistry 281, 24024-24035 (2006) 10 Ngo, P.D., Mansoorabadi, S.O. & Frey, P.A. Serine Protease Catalysis: A Computational Study of Tetrahedral Intermediates and Inhibitory Adducts. J Phys Chem B 120, 7353-7359 (2016) 11 Cleary, J.A., Doherty, W., Evans, P. & Malthouse, J.P.G. Quantifying tetrahedral adduct formation and stabilization in the cysteine and the serine proteases. Bba-Proteins Proteom 1854, 1382-1391 (2015) 12 Scaglione, J.B. et al. Biochemical and structural characterization of the tautomycetin thioesterase: analysis of a stereoselective polyketide hydrolase. Angew Chem Int Ed Engl 49, 5726-5730 (2010) 13 Cappadocia, L. & Lima, C.D. Ubiquitin-like Protein Conjugation: Structures, Chemistry, and Mechanism. Chem Rev 118, 889-918 (2018) 14 Plechanovova, A., Jaffray, E.G., Tatham, M.H., Naismith, J.H. & Hay, R.T. Structure of a RING E3 ligase and ubiquitin-loaded E2 primed for catalysis. Nature 489, 115-U135 (2012) 15 Hay, R.W. & Morris, P.J. Interaction of Dl-2,3-Diaminopropionic Acid and Its Methyl Ester with Metal Ions .1. Formation Constants. J Chem Soc A, 3562-& (1971) 16 Lan, Y. et al. Incorporation of 2,3-Diaminopropionic Acid into Linear Cationic Amphipathic Peptides Produces pH-Sensitive Vectors. Chembiochem 11, 1266-1272 (2010) 17 Radzicka, A. & Wolfenden, R. Rates of uncatalyzed peptide bond hydrolysis in neutral solution and the transition state affinities of proteases. Journal of the American Chemical Society 118, 6105-6109 (1996) 18 Hoyer, K.M., Mahlert, C. & Marahiel, M.A. The iterative gramicidin s thioesterase catalyzes peptide ligation and cyclization. Chemistry & biology 14, 13-22 (2007) 19 Alonzo, D.A., Magarvey, N.A. & Schmeing, T.M. Characterization of cereulide synthetase, a toxin-producing macromolecular machine. PloS one 10, e0128569 (2015) 20 Magarvey, N.A., Ehling-Schulz, M. & Walsh, C.T. Characterization of the cereulide NRPS alpha-hydroxy acid specifying modules: activation of alpha-keto acids and chiral reduction on the assembly line. Journal of the American Chemical Society 128, 10698-10699 (2006) 21 Shaw-Reid, C.A. et al. Assembly line enzymology by multimodular nonribosomal peptide synthetases: the thioesterase domain of E. coli EntF catalyzes both elongation and cyclolactonization. Chemistry & biology 6, 385-400 (1999) 22 May, J.J., Wendrich, T.M. & Marahiel, M.A. The dhb operon of Bacillus subtilis encodes the biosynthetic template for the catecholic siderophore 2,3-dihydroxybenzoate-glycine-threonine trimeric ester bacillibactin. J Biol Chem 276, 7209-7217 (2001) 23 Zhou, Y. et al. Iterative Mechanism of Macrodiolide Formation in the Anticancer Compound Conglobatin. Chemistry & biology 22, 745-754 (2015) 24 Robbel, L., Hoyer, K.M. & Marahiel, M.A. TioS T-TE--a prototypical thioesterase responsible for cyclodimerization of the quinoline- and quinoxaline-type class of chromodepsipeptides. FEBS J 276, 1641-1653 (2009) 25 Jaitzig, J., Li, J., Sussmuth, R.D. & Neubauer, P. Reconstituted biosynthesis of the nonribosomal macrolactone antibiotic valinomycin in Escherichia coli. ACS Synth Biol 3, 432-438 (2014) 26 Akey, D.L. et al. Structural basis for macrolactonization by the pikromycin thioesterase. Nat Chem Biol 2, 537-542 (2006) 27 Samel, S.A., Wagner, B., Marahiel, M.A. & Essen, L.O. The thioesterase domain of the fengycin biosynthesis cluster: a structural base for the macrocyclization of a non-ribosomal lipopeptide. Journal of molecular biology 359, 876-889 (2006) 28 Tseng, C.C. et al. Characterization of the surfactin synthetase C-terminal thioesterase domain as a cyclic depsipeptide synthase. Biochemistry 41, 13350-13359 (2002) 29 Bruner, S.D. et al. Structural basis for the cyclization of the lipopeptide antibiotic surfactin by the thioesterase domain SrfTE. Structure 10, 301-310 (2002) 30 Li, J. et al. Palladium-triggered deprotection chemistry for protein activation in living cells. Nat Chem 6, 352-361 (2014) 31 Baker, A.S. & Deiters, A. Optical Control of Protein Function through Unnatural Amino Acid Mutagenesis and Other Optogenetic Approaches. Acs Chemical Biology 9, 1398-1407 (2014) 32 Nguyen, D.P. et al. Genetic Encoding of Photocaged Cysteine Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014) 33 Nguyen, D.P., Elliott, T., Holt, M., Muir, T.W. & Chin, J.W. Genetically Encoded 1,2-Aminothiols Facilitate Rapid and Site-Specific Protein Labeling via a Bio-orthogonal Cyanobenzothiazole Condensation. Journal of the American Chemical Society 133, 11418-11421 (2011) 34 Neumann, H., Peak-Chew, S.Y. & Chin, J.W. Genetically encoding N-epsilon-acetyllysine in recombinant proteins. Nature Chemical Biology 4, 232-234 (2008) 35 Chin, J.W. Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu Rev Biochem 83, 379-408 (2014) 36 Liu, C.C. & Schultz, P.G. Adding New Chemistries to the Genetic Code. Annual Review of Biochemistry, Vol 79 79, 413-444 (2010) 37 Zhang, M.S. et al. Biosynthesis and genetic encoding of phosphothreonine thr ough parallel selection and deep sequencing. Nat Methods 14, 729-736 (2017) 38 Virdee, S., Ye, Y., Nguyen, D.P., Komander, D. & Chin, J.W. Engineered diubiquitin synthesis reveals Lys29-isopeptide specificity of an OTU deubiquitinase. Nature Chemical Biology 6, 750-757 (2010) 39 Phan, J. et al. Structural basis for the substrate specificity of tobacco etch virus protease. Journal of Biological Chemistry 277, 50564-50572 (2002) 40 Trauger, J.W., Kohli, R.M., Mootz, H.D., Marahiel, M.A. & Walsh, C.T. Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000) 41 Zhou, Y., Prediger, P., Dias, L.C., Murphy, A.C. & Leadlay, P.F. Macrodiolide formation by the thioesterase of a modular polyketide synthase. Angew Chem Int Ed Engl 54, 5232-5235 (2015) 42 Horsman, M.E., Hari, T.P.A. & Boddy, C.N. Polyketide synthase and non-ribosomal peptide synthetase thioesterase selectivity: Logic gate or a victim of fate? Natrual Products Reports (2015) 43 Frueh, D.P. et al. Dynamic thiolation-thioesterase structure of a non-ribosomal peptide synthetase. Nature 454, 903-906 (2008) 44 Whicher, J.R. et al. Structure and function of the RedJ protein, a thioesterase from the prodiginine biosynthetic pathway in Streptomyces coelicolor. J Biol Chem 286, 22558-22569 (2011) 45 Ekici, O.D., Paetzel, M. & Dalbey, R.E. Unconventional serine proteases: variations on the catalytic Ser / His / Asp triad configuration. Protein Sci 17, 2023-2037 (2008) 46 Liu, Y., Zheng, T. & Bruner, S.D. Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488 (2011) 47 Trauger, J.W., Kohli, R.M. & Walsh, C.T. Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001) 48 Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225 (2017) 49 Cravatt, B.F., Wright, A.T. & Kozarich, J.W. Activity-based protein profiling: from enzyme chemistry to proteomic chemistry. Annu Rev Biochem 77, 383-414 (2008) 50 McGall, G. H. et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997) 51 Pendrak, I., Wittrock, R. & Kingsbury, W. D. Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995)

[0440] References related to supplementary methods: 31 Baker, A.S. & Deiters, A. Optical Control of Protein Function through Unnatural Amino Acid Mutagenesis and Other Optogenetic Approaches. Acs Chemical Biology 9, 1398-1407 (2014) 32 Nguyen, D.P. et al. Genetic Encoding of Photocaged Cysteine Allows Photoactivation of TEV Protease in Live Mammalian Cells. Journal of the American Chemical Society 136, 2240-2243 (2014) 33 Nguyen, D.P., Elliott, T., Holt, M., Muir, T.W. & Chin, J.W. Genetically Encoded 1,2-Aminothiols Facilitate Rapid and Site-Specific Protein Labeling via a B io-orthogonal Cyanobenzothiazole Condensation. Journal of the American Chemical Society 133, 11418-11421 (2011) 34 Neumann, H., Peak-Chew, S.Y. & Chin, J.W. Genetically encoding N-epsilon-acetyllysine in recombinant proteins. Nature Chemical Biology 4, 232-234 (2008) 35 Chin, J.W. Expanding and Reprogramming the Genetic Code of Cells and Animals. Annu Rev Biochem 83, 379-408 (2014) 36 Liu, C.C. & Schultz, P.G. Adding New Chemistries to the Genetic Code. Annual Review of Biochemistry, Vol 79 79, 413-444 (2010) 37 Zhang, M.S. et al. Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat Methods 14, 729-736 (2017) 38 Virdee, S., Ye, Y., Nguyen, D.P., Komander, D. & Chin, J.W. Engineered diubiquitin synthesis reveals Lys29-isopeptide specificity of an OTU deubiquitinase. Nature Chemical Biology 6, 750-757 (2010) 39 Phan, J. et al. Structural basis for the substrate specificity of tobacco etch virus protease. Journal of Biological Chemistry 277, 50564-50572 (2002) 40 Trauger, J.W., Kohli, R.M., Mootz, H.D., Marahiel, M.A. & Walsh, C.T. Peptide cyclization catalysed by the thioesterase domain of tyrocidine synthetase. Nature 407, 215-218 (2000) 41 Zhou, Y., Prediger, P., Dias, L.C., Murphy, A.C. & Leadlay, P.F. Macrodiolide formation by the thioesterase of a modular polyketide synthase. Angew Chem Int Ed Engl 54, 5232-5235 (2015) 42 Horsman, M.E., Hari, T.P.A. & Boddy, C.N. Polyketide synthase and non-ribosomal peptide synthetase thioesterase selectivity: Logic gate or a victim of fate? Natrual Products Reports (2015) 43 Frueh, D.P. et al. Dynamic thiolation-thioesterase structure of a non-ribosomal peptide synthetase. Nature 454, 903-906 (2008) 44 Whicher, J.R. et al. Structure and function of the RedJ protein, a thioesterase from the prodiginine biosynthetic pathway in Streptomyces coelicolor. J Biol Chem 286, 22558-22569 (2011) 45 Ekici, O.D., Paetzel, M. & Dalbey, R.E. Unconventional serine proteases: variations on the catalytic Ser / His / Asp triad configuration. Protein Sci 17, 2023-2037 (2008) 46 Liu, Y., Zheng, T. & Bruner, S.D. Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488 (2011) 47 Trauger, J.W., Kohli, R.M. & Walsh, C.T. Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001) 48 Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225 (2017) 49 Cravatt, B.F., Wright, A.T. & Kozarich, J.W. Activity-based protein profiling: from enzyme chemistry to proteomic chemistry. Annu Rev Biochem 77, 383-414 (2008) 50 McGall, G. H. et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997) 51 Pendrak, I., Wittrock, R. & Kingsbury, W. D. Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995) 52 Alonzo, D. A., Magarvey, N. A. & Schmeing, T. M. Characterization of cereuli de synthetase, a toxin-producing macromolecular machine. PloS one 10, e0128569, doi:10.1371 / journal.pone.0128569 (2015) 53 Shaw-Reid, C. A. et al. Assembly line enzymology by multimodular nonribosomal peptide synthetases: the thioesterase domain of E. coli EntF catalyzes both elongation and cyclolactonization. Chemistry & biology 6, 385-400, doi:10.1016 / S1074-5521(99)80050-7 (1999) 54 May, J. J., Wendrich, T. M. & Marahiel, M. A. The dhb operon of Bacillus subtilis encodes the biosynthetic template for the catecholic siderophore 2,3-dihydroxybenzoate-glycine-threonine trimeric ester bacillibactin. J Biol Chem 276, 7209-7217, doi:10.1074 / jbc.M009140200 (2001) 55 Zhou, Y. et al. Iterative Mechanism of Macrodiolide Formation in the Anticancer Compound Conglobatin. Chemistry & biology 22, 745-754, doi:10.1016 / j.chembiol.2015.05.010 (2015) 56 Robbel, L., Hoyer, K. M. & Marahiel, M. A. TioS T-TE--a prototypical thioesterase responsible for cyclodimerization of the quinoline- and quinoxaline-type class of chromodepsipeptides. FEBS J 276, 1641-1653, doi:10.1111 / j.1742-4658.2009.06897.x (2009) 57 Liu, Y., Zheng, T. & Bruner, S. D. Structural basis for phosphopantetheinyl carrier domain interactions in the terminal module of nonribosomal peptide synthetases. Chemistry & biology 18, 1482-1488, doi:10.1016 / j.chembiol.2011.09.018 (2011) 58 Trauger, J. W., Kohli, R. M. & Walsh, C. T. Cyclization of backbone-substituted peptides catalyzed by the thioesterase domain from the tyrocidine nonribosomal peptide synthetase. Biochemistry 40, 7092-7098 (2001) 59 Aggarwal, A. et al. Development of a Novel Lead that Targets M. tuberculosis Polyketide Synthase 13. Cell 170, 249-259 e225, doi:10.1016 / j.cell.2017.06.025 (2017) 60 Neumann, H., Peak-Chew, S. Y. & Chin, J. W. Genetically encoding N(epsilon)-acetyllysine in recombinant proteins. Nat Chem Biol 4, 232-234, doi:10.1038 / nchembio.73 (2008) 61 Rogerson, D. T. et al. Efficient genetic encoding of phosphoserine and its nonhydrolyzable analog. Nat Chem Biol 11, 496-503, doi:10.1038 / nchembio.1823 (2015) 62 Zhang, M. S. et al. Biosynthesis and genetic encoding of phosphothreonine through parallel selection and deep sequencing. Nat Methods 14, 729-736, doi:10.1038 / nmeth.4302 (2017) 63 McGall, G. H. et al. The efficiency of light-directed synthesis of DNA arrays on glass substrates. Journal of the American Chemical Society 119, 5081-5090, doi:DOI 10.1021 / ja964427a (1997) 64 Nguyen, D. P. et al. Genetic encoding of photocaged cysteine allows photoactivation of TEV protease in live mammalian cells. Journal of the American Chemical Society 136, 2240-2243, doi:10.1021 / ja412191m (2014) 65 Pendrak, I., Wittrock, R. & Kingsbury, W. D. Synthesis and Anti-Hsv Activity of Methylenedioxy Mappicine Ketone Analogs. J Org Chem 60, 2912-2915, doi:DOI 10.1021 / jo00114a050 (1995) 66 Mayer, S. C., Ramanjulu, J., Vera, M. D., Pfizenmayer, A. J. & Joullie, M. M. Synthesis of New Didemnin B Analogs for Investigations of Structure / Biological Activity Relationships. J Org Chem 59, 5192-5205, doi:10.1021 / jo00097a022 (1994) 67 Faure, S. et al. Asymmetric intramolecular [2+2] photocycloadditions: alpha- and beta-hydroxy acids as chiral tether groups. J Org Chem 67, 1061-1070, doi:10.1021 / jo001631e (2002) 68 Battye, T. G., Kontogiannis, L., Johnson, O., Powell, H. R. & Leslie, A. G. iMOSFLM: a new graphical interface for diffraction-image processing with MOSFLM. Acta crystallographica. Section D, Biological crystallography 67, 271-281, doi:10.1107 / S0907444910048675 (2011) 69 Winter, G. et al. DIALS: implementation and evaluation of a new integration package. Acta crystallographica. Section D, Structural biology 74, 85-97, doi:10.1107 / S2059798317017235 (2018)

[0441] Although exemplary embodiments of the present invention have been disclosed in detail herein with reference to the accompanying drawings, the reader should note that the present invention is not limited to those exact embodiments, and that various changes and modifications may be made by those skilled in the art without departing from the scope of the present invention as defined by the appended claims and their equivalents.

Claims

1. A PylRS tRNA synthetase comprising the mutations Y271C, N311Q, Y349F and V366C.

2. The PylRS tRNA synthetase of claim 1, which is a Methanosarcina barkerii PylRS (MbPylRS) tRNA synthetase comprising the mutation.