Non-natural amino acids for intracellular synthesis for site-specific protein modification

By introducing mutant PylC and PylRS into recombinant host cells, the problem of easy degradation of protein therapeutic agents in vivo was solved, and efficient synthesis and cyclization of non-classical amino acids were achieved, improving protein stability and yield, making it suitable for large-scale production.

CN121335983APending Publication Date: 2026-01-13THE CHINESE UNIVERSITY OF HONG KONG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202480035597.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-29
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, protein-based therapeutic agents are easily degraded rapidly by proteolytic enzymes in vivo, resulting in a short serum half-life and affecting their efficacy. Furthermore, the methods for chemically synthesizing non-classical amino acids are costly and not suitable for large-scale production.

Method used

By introducing mutant PylC and PylRS into recombinant host cells, efficient synthesis and incorporation of non-classical amino acids can be achieved. This includes mutating specific residues of PylC and introducing TAG or TAA codons into the coding sequence, which combine with cyclization reactions to form covalent bonds, thereby improving protein stability and yield.

Benefits of technology

It enables the efficient incorporation and cyclization of non-classical amino acids into proteins, improving protein stability and yield in vivo, reducing production costs, and making it suitable for large-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121335983A_ABST
    Figure CN121335983A_ABST
Patent Text Reader

Abstract

Provided are host cells containing genetically modified engineered enzymes that support the synthesis of non-classical amino acids (e.g., D-Cys-Lys), as well as the incorporation of such non-classical amino acids into newly synthesized proteins. Thus, these cells are capable of synthesizing a protein containing one or more non-classical amino acids in vivo or in cells, which non-classical amino acids can be strategically placed at predetermined specific modification sites of the protein. Corresponding methods for in vivo synthesis of modified proteins and modified proteins produced by the methods are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 455,799, filed March 30, 2023, the disclosure of which is incorporated herein by reference in its entirety for all purposes. Background Technology

[0003] Protein-based therapeutics are an important alternative to small-molecule drugs for treating diseases because they generally exhibit high specificity and low toxicity. However, a major limitation of protein-based therapeutics is their rapid degradation by proteolytic enzymes in the body, resulting in a short serum half-life, which adversely affects their efficacy (Sato, Aaron K). et al. Current Opinion in Biotechnology vol. 17,6 (2006): 638-42). Protein modification (e.g., cyclization) can be a useful approach to overcome this serum stability challenge, and cyclized proteins often exhibit significantly extended in vivo lifespan and improved binding affinity due to the reduction in entropy cost of binding (Colgrave, Michelle L, and David J Craik). Biochemistry vol. 43,20 (2004): 5965-75; Ji,Yanbin et al. Journal of the American Chemical Society vol. 135,31 (2013):11623-11633; Ngo, Khac Huy et al. Chemical Communications (Cambridge,England) vol. 56,7 (2020): 1082-1084; Wilbs, Jonas et al. Nature Vommunications vol. 11,1 3890. 4 Aug. 2020; Clardy, Jon, and ChristopherWalsh. Nature vol. 432,7019 (2004): 829-37; Driggers, Edward M et al. Nature Reviews. Drug Discovery vol. 7,7 (2008): 608-24).

[0004] Previous reports have described the use of pyrrolidone (Pyl) analogues containing cysteine, D-cysteyl-N ε - -Lysine (D-Cys-ε-Lys, abbreviation: O, Figure 1BThis is used for cyclizing the arginine-glycine-aspartic acid (RGD) motif. The RGD motif is attached to mCherry proteins via inteptide-mediated natural chemical linkage (NCL) (Lee, Marianne M). et al. Chembiochem: a European Journal of Chemical Biology vol. 15,12 (2014): 1769-72; Dawson, PE et al. Science (New York, NY) vol. 266,5186 (1994): 776-9). D-Cys-ε-Lys is encoded by the amber (UAG) codon. This was achieved by *Methanococcus martensii* (…). Methanosarcina mazei )Pyrrololysyl-tRNA synthetase / tRNA pyl (PylRS / tRNA Pyl ) for the introduction of Escherichia coli ( Escherichia coli ) cells, to achieve the incorporation of D-Cys-ε-Lys into recombinant proteins (Li, Xin et al. Angewandte Chemie (International ed. in English) vol. 48,48 (2009): 9184-7). However, despite obtaining cyclized proteins with high purity, the product yield was low. Factors affecting the production yield included inefficient incorporation of D-Cys-ε-Lys into the peptide and incomplete and inefficient steps in protein cyclization.

[0005] Chemically synthesizing non-classical amino acids (ncAAs), such as D-Cys-ε-Lys, and then exogenously supplementing them into the culture medium is a common strategy for producing ncAA-containing proteins. However, this method is both expensive and time-consuming, and in some cases, ncAAs may be cell-impermeable, making them unusable for translational incorporation. Therefore, there is a need for novel methods for synthesizing Pyl analogs and other non-classical amino acids (ncAAs) that are cost-effective and sustainable for large-scale production of proteins containing ncAAs and / or Pyl analogs. As disclosed herein, the inventors have identified compositions and methods that address this need. Summary of the Invention

[0006] In a first aspect, this disclosure provides a recombinant host cell comprising a polynucleotide sequence encoding a mutant PylC capable of synthesizing non-canonical amino acids (ncAA), wherein said mutant PylC comprises the same as SEQ ID NO: 1 (wild-type Methanococcus masculinii ( M. mazeiThe recombinant host cell contains a polypeptide sequence having at least 95% sequence identity with PylC. In some embodiments, the recombinant host cell also contains (a) a polynucleotide sequence encoding wild-type PylRS or mutant PylRS, the mutant PylRS being capable of using ncAA as a substrate, and (b) a polynucleotide sequence encoding tRNA, the tRNA incorporating ncAA at the amber codon.

[0007] In some embodiments, the mutant PylC comprises a combination of mutations corresponding to residues 177, 179, 233, and 256 of the wild-type archaea PylC. In some embodiments, the combination of mutations in the mutant PylC is selected from: 1) S177N, E179P, D233S and T256V; 2) S177C, E179T, D233S, and T256V; 3) E179C and D233N; 4) S177A, E179A, D233H and T256M; 5) S177A, E179S, D233H and T256L; 6) E179C, D233H, and T256V; 7) E179V, D233N, and T256V; 8) E179V, D233H, and T256N; 9) S177T, E179P, D233H and T256M; 10) S177G, E179V, D233S and T256L; 11) E179N, D233H, and T256L; 12) S177A, E179V, D233H, and T256I; 13) S177L, E179I, D233H and T256C; 14) S177A, E179I, D233T and T256L; 15) S177H, E179L, D233S and T256V; 16) S177A, E179M, D233S, and T256L; 17) E179I, D233H, and T256I; 18) S177A, E179T, D233H and T256V; 19) S177A, E179A, D233N, and T256C; 20) S177C, E179P, D233H and T256L; 21) S177G, E179I, and D233H; 22) S177M, E179V, D233H and T256V; 23) S177V, E179V, D233H and T256V; 24) S177A, E179A, D233S and T256V; 25) S177G, E179M, D233N and T256V; 26) S177G, E179I, D233N and T256V; 27) S177G, E179T, D233S and T256L; 28) S177A, E179T, D233S and T256L; 29) S177L, E179T, D233S and T256I; 30) S177T, E179P, D233H and T256M; 31) E179L, D233L, and T256V; 32) S177A, E179V, D233Q, and T256C; 33) S177L, E179V, D233S and T256A; 34) S177A, E179C, D233L and T256C; 35) S177N, E179C, D233T and T256C; 36) S177H, E179L, D233S and T256S; 37) E179G, D233L, and T256V; 38) S177N, E179P, D233S and T256V; 39) S177M, E179G, D233H and T256Y; 40) S177M, E179L, D233S and T256C; 41) S177A, E179C, and D233H; 42) E179L, D233L, and T256I; 43) S177L, E179V, D233H and T256Y; 44) S177A, E179Q, D233N and T256C; and 45) S177L, E179A, D233H and T256V.

[0008] In some implementations, the combination of mutations in mutant PylC is S177N, E179P, D233S, and T256V.

[0009] In some implementations, the combination of mutations in mutant PylC is selected from: 1) S177G, E179V, D233N and T256V; 2) S177V, E179L, D233H and T256L; 3) S177G, E179V, D233H and T256A; 4) S177V, E179C, D233H and T256L; 5) S177M, E179V, D233H and T256Y; 6) S177C, E179P, D233H and T256L; 7) S177G, E179A, D233S and T256L; 8) S177V, E179L, D233H and T256V; 9) S177G, E179L, D233H and T256L; 10) S177I, D233H, and T256V; 11) S177L, E179V, D233T and T256I; 12) S177L, E179A, D233H and T256M; 13) S177L, E179V, D233H and T256A; 14) S177I, E179M, D233H and T256L; 15) S177A, E179I, D233N, and T256I; 16) S177I, E179C, D233H and T256L; 17) E179S and D233H; 18) S177V, E179V, D233H and T256A; 19) S177C, E179V, D233H and T256Y; 20) S177I, E179A, D233H and T256H; 21) S177L, E179C, D233H and T256V; 22) E179V and D233H; 23) S177V, E179V, D233H and T256L; 24) S177I, E179L, D233H and T256A; 25) S177L, E179A, D233H and T256Y; 26) S177G, E179V, D233L and T256L; 27) E179P, D233N, and T256V; 28) S177G, E179V, D233N and T256N; 29) S177L, E179C, D233G and T256H; 30) S177H, E179G, D233S and T256A; 31) S177A, E179C, D233S and T256L; 32) S177G, E179V, D233L and T256Y; 33) S177L, E179T, D233T and T256M; 34) S177G, E179L, D233W and T256S; 35) S177V, E179L, D233S and T256A; 36) E179V, D233L, and T256L; 37) S177L, E179P, D233S and T256N; and 38) S177V, E179V, D233T and T256L.

[0010] In some implementations, the combination of mutations in mutant PylC is S177G, E179V, D233N, and T256V.

[0011] In some implementations, mutant PylRS contains mutations corresponding to (a) G114E, C348V, and S451F of wild-type archaea PylRS, or (b) L301M, L305I, Y306L, L309A, and C348F of wild-type archaea PylRS.

[0012] In some implementations, the genome sequence encoding release factor 1 (RF1) of the recombinant host cell is at least partially missing.

[0013] In some embodiments, the recombinant host cell is a prokaryotic cell. In some embodiments, the recombinant host cell is a bacterial cell. In some embodiments, the recombinant host cell is an *E. coli* cell.

[0014] In related respects, this disclosure provides compositions comprising any recombinant host cell disclosed herein.

[0015] In related respects, this disclosure provides lysates of any recombinant host cells disclosed herein.

[0016] In another aspect, this disclosure provides a method for recombinant synthesis of a protein, comprising: (a) introducing a nucleotide sequence encoding the protein into any recombinant host cell of this disclosure, wherein the nucleotide sequence comprises (i) at least one TAG codon within the coding sequence of the nucleotide sequence, and (ii) a TAA codon or a TGA codon at the end of the coding sequence of the nucleotide sequence; and (b) culturing the recombinant host cell under conditions that allow transcription from the nucleotide sequence and protein synthesis, thereby expressing the protein. In some embodiments, the protein comprises one or more ncAAs, optionally wherein the one or more ncAAs are selected from D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2-chloroacetyl-ε-Lys.

[0017] In some embodiments, the method further includes separating the protein produced in step (b). In some embodiments, the method further includes placing the separated protein under conditions that allow the protein to cyclize, the cyclization being carried out by forming a covalent bond between (i) a first moiety on a first ncAA and (ii) a second moiety on a cysteine ​​residue, a C-terminal thioester, or a second moiety on a second ncAA.

[0018] In some embodiments, the protein comprises a peptide sequence derived from a small ubiquitin-associated modified (SUMO) protein.

[0019] In another aspect, this disclosure provides a method for recombinant synthesis of a protein, comprising contacting a nucleotide sequence encoding the protein with cell lysate of any recombinant host cell disclosed herein. In some embodiments, the protein is P16.

[0020] In related aspects, this disclosure provides a nucleic acid comprising a polynucleotide sequence encoding a mutant PylC capable of synthesizing ncAA. In some embodiments, the mutant PylC comprises a combination of mutations corresponding to residues 177, 179, 233, and 256 of the wild-type archaea PylC. In some embodiments, the combination of mutations in the mutant PylC is selected from: 1) S177N, E179P, D233S and T256V; 2) S177C, E179T, D233S, and T256V; 3) E179C and D233N; 4) S177A, E179A, D233H and T256M; 5) S177A, E179S, D233H and T256L; 6) E179C, D233H, and T256V; 7) E179V, D233N, and T256V; 8) E179V, D233H, and T256N; 9) S177T, E179P, D233H and T256M; 10) S177G, E179V, D233S and T256L; 11) E179N, D233H, and T256L; 12) S177A, E179V, D233H, and T256I; 13) S177L, E179I, D233H and T256C; 14) S177A, E179I, D233T and T256L; 15) S177H, E179L, D233S and T256V; 16) S177A, E179M, D233S, and T256L; 17) E179I, D233H, and T256I; 18) S177A, E179T, D233H and T256V; 19) S177A, E179A, D233N, and T256C; 20) S177C, E179P, D233H and T256L; 21) S177G, E179I, and D233H; 22) S177M, E179V, D233H and T256V; 23) S177V, E179V, D233H and T256V; 24) S177A, E179A, D233S and T256V; 25) S177G, E179M, D233N and T256V; 26) S177G, E179I, D233N and T256V; 27) S177G, E179T, D233S and T256L; 28) S177A, E179T, D233S and T256L; 29) S177L, E179T, D233S and T256I; 30) S177T, E179P, D233H and T256M; 31) E179L, D233L, and T256V; 32) S177A, E179V, D233Q, and T256C; 33) S177L, E179V, D233S and T256A; 34) S177A, E179C, D233L and T256C; 35) S177N, E179C, D233T and T256C; 36) S177H, E179L, D233S and T256S; 37) E179G, D233L, and T256V; 38) S177N, E179P, D233S and T256V; 39) S177M, E179G, D233H and T256Y; 40) S177M, E179L, D233S and T256C; 41) S177A, E179C, and D233H; 42) E179L, D233L, and T256I; 43) S177L, E179V, D233H and T256Y; 44) S177A, E179Q, D233N and T256C; and 45) S177L, E179A, D233H and T256V.

[0021] In some implementations, the combination of mutations in mutant PylC is S177N, E179P, D233S, and T256V.

[0022] In some implementations, the combination of mutations in mutant PylC is selected from: 1. S177G, E179V, D233N, and T256V; 2. S177V, E179L, D233H, and T256L; 3. S177G, E179V, D233H, and T256A; 4. S177V, E179C, D233H, and T256L; 5. S177M, E179V, D233H and T256Y; 6. S177C, E179P, D233H, and T256L; 7. S177G, E179A, D233S, and T256L; 8. S177V, E179L, D233H and T256V; 9. S177G, E179L, D233H, and T256L; 10. S177I, D233H, and T256V; 11. S177L, E179V, D233T, and T256I; 12. S177L, E179A, D233H, and T256M; 13. S177L, E179V, D233H and T256A; 14. S177I, E179M, D233H and T256L; 15. S177A, E179I, D233N, and T256I; 16. S177I, E179C, D233H and T256L; 17. E179S and D233H; 18. S177V, E179V, D233H and T256A; 19. S177C, E179V, D233H, and T256Y; 20. S177I, E179A, D233H and T256H; 21. S177L, E179C, D233H, and T256V; 22. E179V and D233H; 23. S177V, E179V, D233H and T256L; 24. S177I, E179L, D233H and T256A; 25. S177L, E179A, D233H and T256Y; 26. S177G, E179V, D233L and T256L; 27. E179P, D233N, and T256V; and 28. S177G, E179V, D233N and T256N.

[0023] In some implementations, the combination of mutations in mutant PylC is S177G, E179V, D233N, and T256V.

[0024] On the other hand, this disclosure provides a nucleic acid comprising a polynucleotide sequence encoding a mutant PylRS, which is capable of using ncAA as a substrate. In some embodiments, the mutant PylRS comprises mutations corresponding to (a) G14E, C348V, and S451F of wild-type archaea PylRS, or (b) L301M, L305I, Y306L, L309A, and C348F of wild-type archaea PylRS.

[0025] In related respects, this disclosure provides compositions comprising any nucleic acid disclosed herein.

[0026] In related aspects, this disclosure provides expression components or vectors that include polynucleotide sequences of any nucleic acid disclosed herein. Attached Figure Description

[0027] Figure 1A The incorporation of non-classical amino acid (ncAA) D-Cys-ε-Lys into an example protein is shown. In this example, the protein comprises a cyclized therapeutic peptide (e.g., a P16 peptide) at one end and a cyclized targeting peptide (e.g., an arginine-glycine-aspartic acid (RGD) targeting peptide) at the other end.

[0028] Figure 1B A schematic diagram of the plasmid vector used for PylRS protein evolution is shown. The top section shows the pPylST-KanR(TAG) construct, which has one selection gene (KanR(TAG)) for the first round of PylRS screening. The middle section shows pPylST-KanR(TAG)-mCh(TAG), which has two selection genes (KanR(TAG) and mCh(TAG)) for the second and subsequent rounds of PylRS screening. The bottom section shows the plasmid vector containing the encoding tRNA. M15 And the evolution of PylRS EVF The pPylST.tL-mCh(TAG) gene. pPylST.tL-mCh(TAG) was used in an optimized read-through system for efficient incorporation of D-Cys-ε-Lys.

[0029] Figures 1C-1E The optimization of the D-Cys-ε-Lys readout system is shown. Figure 1C The chemical structure of D-Cys-ε-Lys is shown. Figure 1DThis study compares the mCherry readthrough assays using the original UAG readthrough system and the optimized UAG readthrough system. The lysine codon at position 55 of the mCherry gene was mutated to TAG, and the culture medium was supplemented with 0 mL, 2 mM, or 5 mM cys-ε-lys. Data represent mean fluorescence intensity ± standard error of the mean (n = 3). WT: wild type; M15: tRNA M15 U: Unmodified pPylST; T1: Modified pPylST, in which the two T7lac promoters are replaced by Ptac and PLlacO1 respectively; R2: Rosetta 2 (DE3); C321: C321.ΔA.M9 adapted. Figure 1E The original UAG readout system and the optimized system used for incorporating D-Cys-ε-Lys into different recombinant proteins are shown. O represents D-Cys-ε-Lys. CaM is an abbreviation for calmodulin.

[0030] Figures 2A-2C Engineered PylC for intracellular synthesis of D-Cys-ε-Lys is shown. Figure 2A A schematic diagram of the reaction catalyzed by wild-type PylC and engineered PylC is shown. Figure 2B Screening for PylC mutants based on mCherry fluorescence projections is shown. These mutants recognize D-cysteine ​​and catalyze the production of D-Cys-ε-Lys. Site-directed saturation mutagenesis was performed on four residues of PylC: S177, E179, D233, and T256. Data represent the mean fluorescence intensity ± standard error of the mean (n = 3). Figure 2C This shows the production of wild-type PylC (PylC) in D-Cys-ε-Lys at different D-cysteine ​​concentrations. WT ) and the evolved PylC mutant (PylC NPSV In comparison, the D-Cys-ε-Lys was used to incorporate it into UAG-containing proteins. Readability of samples was evaluated based on the readability protein generated by exogenously supplementing 4 mM D-Cys-ε-Lys.

[0031] Figures 3A-3E Computational analysis of cyclized P16 peptides is shown. Figure 3A The amino acid sequences of the cyclic P16 peptide (cycP16p) and the linear P16p are shown. The cyclic P16 peptide shows the D-Cys-ε-Lys incorporation site (denoted by O). Black brackets indicate D-Cys-ε-Lys-mediated cyclization. Figure 3BA cycP16p model interacting with CDK6 is shown. The structure of the p16p (band-like model) and CDK6 (space-filling model) complex is derived from PDB: 1BI7. Residues interacting with CDK6 are labeled. Figure 3C This shows a comparison of molecular dynamics (MD) simulation results between linear P16p (LinP16p) (left) and cycP16p (right). Figure 3C The superposition of 10 rounds of LinP16p (left) and CycP16p (right) MD simulations at 10 ns is shown. Figure 3D RMSD was displayed, and Figure 3E The RMSF spectra of LinP16p and CycP16p are shown in a 300 ns MD simulation. The experimental results are calculated based on the main chain atoms. Error bars represent the standard deviation (SD) of three repeated simulation runs and are plotted as shaded areas in the RMSD spectra.

[0032] Figures 4A-4C The intramolecular circularization of GFP-P16p was shown. Figure 4A A schematic diagram of the protein construct GFP-O-P16p-inteptide-CBD-His7 and the mechanism of protein cyclization based on D-Cys-ε-Lys are shown. Figure 4B The SDS-PAGE gel images of reaction samples collected at different time points (hours (h)) during GFP-O-P16p circularization are shown. Experimental results are illustrated by Coomassie blue staining (top panel) and in-gel GFP fluorescence detection (bottom panel). Figure 4C The deconvolution mass spectrum of GFP-cycP16p obtained by ESI-Orbitrap mass spectrometry is shown.

[0033] Figure 5 The cyclic P16p (MBP-cycP16p) is shown to exhibit a higher CDK4 binding affinity than its linear counterpart. Binding curves from MST determinations of MBP-P16p and MBP-cycP16p with GST-CDK4 are presented. Error bars represent the standard deviation (SD) of three repeated measurements.

[0034] Figure 6 A schematic diagram of the protein construct used for P16p cyclization is shown. Top section: GFP-cycP16p. Middle section: MBP-cycP16p. Bottom section: cycRGD-mCh-cycP16p. "O" indicates D-Cys-ε-Lys.

[0035] Figures 7A-7D The effects of cycRGD-mCh-cycP16p on MCF-7 cells were demonstrated. Figure 7AIn this study, MCF-7 cells were exposed to different treatments for 24 hours, followed by cell cycle analysis using a BD flow cytometer. A 10 nM actinomycin was included as a positive control. Figure 7B MCF-7 cells treated with different peptides for 24 hours are shown. Cell numbers were normalized relative to the PBS control group. Data are expressed as mean ± SD (n = 3). Results showed statistical significance relative to the cycRGD-mCh-cycP16p group. Figure 7C The percentage of MCF-7 cells arrested in the G0 / G1 phase at different time points after exposure to different treatments is shown. Error bars represent the standard deviation (SD) of three repeated measures. P-values ​​were calculated using a one-way ANOVA test, with ns indicating no significance. p < 0.05, p < 0.001, p < 0.0001, ns, not significant. Figure 7D Western blot analysis showing the phosphorylation status of Rb in MCF-7 cells exposed to different treatments using anti-pRb antibody.

[0036] Figures 8A-8B The synthesis of D-Pra-ε-Lys using engineered PylC was demonstrated. Figure 8A (R)-2-aminopentan-4-alkynic acid was shown to be used as a substrate in combination with endogenous L-lysine for the synthesis of D-Pra-ε-Lys. Figure 8B Thirty-one PylC variants were identified following antibiotic screening. mCherry fluorescence readout assays were used to evaluate the PylC variants.

[0037] Figures 9A-9B The synthesis of D-allyl-ε-Lys using engineered PylC was demonstrated. Figure 9A The reaction of L-lysine with D-allyl-OH and its conversion to D-allyl-ε-Lys was shown. Figure 9B The mCherry fluorescence intensity from 35 selected colonies is shown.

[0038] Figure 10 The intracellular synthesis of 2-chloroacetyl-ε-Lys (2ClAcK) from 2-chloroacetic acid and lysine was demonstrated.

[0039] Figures 11A-11C The three-dimensional structure of an example cyclic SUMO1-derived peptide and Coomassie staining gel are shown. Figure 11A The structure of SUMOmini2 is shown. The mutations used for cyclization are shown in a simplified diagram. The C52S mutation prevents undesirable side reactions. Figure 11BThe SDS-PAGE gel images of the circular His-SUMOmini2-LVPRGS-SUMOCt-MBP are shown before and after thrombin cutting. Figure 11C The SDS-PAGE gel image of cyclic His-SUMOmini2 purified by RP-HPLC after thrombin cleavage is shown.

[0040] Figure 12 The results showed that linear, cyclic, and bicyclic His-SUMOmini2 could be incubated at room temperature for up to 1.5 hours in the presence of 0.1 μg / mL proteinase K.

[0041] Figure 13 The aggregation kinetics of α-synuclein (aSyn) are shown using thioflavin T (ThT) assays in the absence of two different His-SUMOmini2 constructs with varying stoichiometric ratios. The His-SUMOmini2 constructs are linear His-SUMOmini2, monocyclic (“cyclic”) His-SUMOmini2, and bicyclic His-SUMOmini2. At an α-synuclein to His-SUMOmini2 ratio of 1:1 (left panel), all SUMOmini2 peptides inhibited α-synuclein aggregation. At an α-synuclein to His-SUMOmini2 ratio of 1:0.2 (right panel), monocyclic His-SUMOmini2 was less effective than linear SUMOmini2 in inhibiting α-synuclein aggregation, while bicyclic SUMOmini2 showed enhanced inhibition of α-synuclein aggregation.

[0042] Detailed description

[0043] I. Definition

[0044] For the purpose of facilitating an understanding of the principles of this disclosure, reference will now be made to preferred embodiments, and preferred embodiments will be described using specific language. However, it should be understood that this is not intended to limit the scope of this disclosure, and changes and further modifications to this disclosure as illustrated herein are considered to be commonly apparent to those skilled in the art related to this disclosure.

[0045] As used in this article, the terms "D-Cys-ε-Lys" and "D-Cys-N" ε -Lys", "D-Cys-ε-L-Lys", "D-Cys-N ε -L-Lys", D-Cys-ε-L-lysine", D-Cys-N ε -L-lysine", D-cysteyl-ε-Lys", D-cysteyl-N ε-Lys", D-cysteyl-ε-L-Lys", D-cysteyl-N ε -L-Lys", D-cysteyl-ε-L-lysine and D-cysteyl-N ε "-L-lysine" is synonymous and interchangeable, referring to... Figure 1C The chemical structure shown is “D-Cys-ε-Lys”.

[0046] As used in this article, the terms "D-Pra-ε-Lys" and "D-Pra-N" ε -Lys", D-Pra-ε-L-lysine", D-Pra-N ε -L-lysine", D-Pra-ε-L-Lys", D-Pra-N ε -L-Lys", D-Pra-ε-lysine", D-Pra-N ε "-lysine" and "D-N6-(2-(R)-propylatedglycyl)lysine" are synonymous and interchangeable, referring to... Figure 8A The chemical structure shown is “D-Pra-ε-Lys”.

[0047] As used herein, the terms "D-allyl-ε-Lys", "D-allyl-N" ε -Lys", D-allyl-ε-L-Lys", D-allyl-N ε -L-Lys", D-allyl-ε-lysine", D-allyl-N ε -Lysine", "D-allyl-ε-L-lysine" and "D-allyl-N ε "-L-lysine" is synonymous and can be used interchangeably, referring to... Figure 9A The chemical structure shown is "D-allyl-ε-Lys".

[0048] As used herein, the terms "2-chloroacetyl-ε-Lys", "2-chloroacetyl-N" ε -Lys", "2-Chloroacetyl-ε-L-Lys", "2-Chloroacetyl-N ε -L-Lys", 2-chloroacetyl-ε-lysine", 2-chloroacetyl-N ε -Lysine", "2-Chloroacetyl-ε-L-lysine", "2-Chloroacetyl-N ε "-L-lysine" and "2ClAcK" are synonyms and can be used interchangeably, referring to... Figure 10 The chemical structure shown is "2-chloroacetyl-N ε -L-L-lysine (2ClAcK)".

[0049] As used herein, the terms “intracellular” and “in vivo” are synonymous and interchangeable, referring to processes that occur in living cells. In a non-limiting example, “intracellular synthesis of D-Cys-ε-Lys” or “in vivo synthesis of D-Cys-ε-Lys” refers to the synthesis of the chemical substance D-Cys-ε-Lys in living cells, as opposed to the in vitro chemical synthesis of D-Cys-ε-Lys using a chemical substrate in a reaction vessel (e.g., a test tube). According to this disclosure, “intracellular synthesis” or “in vivo synthesis” generally occurs in recombinant host cells.

[0050] As used herein, the terms “pyrrololysine,” “Pyl,” or “O” refer to (2S)-2-amino-6-[[(2R,3R)-3-methyl-3,4-dihydro-2H-pyrrolo-2-carboxyl]amino]hexanoic acid with PubChem CID5460671, IUPAC name (2S)-2-amino-6-[[(2R,3R)-3-methyl-3,4-dihydro-2H-pyrrolo-2-carboxyl]amino]hexanoic acid and molecular formula C 12 H 21 Compounds of N3O3. Pyl is a lysine derivative that, in some cases, can be incorporated into polypeptide chains at positions corresponding to the succinate (UAG) stop codon.

[0051] As used herein, the term "PylC" refers to... pylC The protein encoded by the gene, pylC Genes typically catalyze the ATP-dependent ligation of lysine and another substrate to produce ncAA. In some cases, wild-type PylC catalyzes the ATP-dependent ligation of (3R)-3-methyl-D-ornithine and L-lysine to produce (3R)-3-methyl-D-ornithine-N6-L-lysine. In some cases, PylC refers to 3-methyl-D-ornithine-L-lysine ligase. Although PylC mutants or variants can be derived from any naturally occurring archaea or bacterial PylC, in some non-limiting examples, the naturally occurring PylC is PylC of *Methanococcus martensii* (UniProt ID: Q8PWY3; SEQ ID NO: 1). In some cases, PylC mutants or variants can be used for the synthesis of ncAA (e.g., but not limited to D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK).

[0052] As used herein, the term "PylRS" refers to... pylS Genes that encode and typically catalyze the binding of pyrrolidone (Pyl) to tRNA to produce tRNA PylThe protein. In some cases, PylRS refers to pyrrolidone-lysyl-tRNA synthetase or pyrrolidone-lysine-tRNA ligase. Although PylRS mutants or variants can be derived from any naturally occurring archaea or bacterial PylRS, in some non-limiting examples, the naturally occurring PylRS is that of *Methanococcus martensii* (UniProt ID: Q6WRH6). In some cases, PylRS mutants or variants can be used to catalyze the ligation of ncAA (e.g., but not limited to D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK) to tRNA. Thus, PylRS mutants or variants can be used to incorporate ncAA into nascent polypeptide chains via ribosomes, for example, at the amber (TAG / UAG) codon in mRNA.

[0053] As used herein, “UAG,” “TAG,” or “amber”; “UAA,” “TAA,” or “ochre”; and “UGA,” “TGA,” or “milky” are three stop codons that interrupt translation. The Pyl and certain ncAAs disclosed herein, such as D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAck, can be encoded by stop codons. In some cases, ncAAs are encoded by the “TAG,” “UAG,” or “amber” codons.

[0054] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleotides or ribonucleotides in single-stranded or double-stranded form and their polymers. Unless specifically defined, the term includes nucleic acids containing known analogs of natural nucleotides, having similar binding properties to reference nucleic acids, and being metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specified, a particular nucleic acid sequence also implicitly includes variants of its conserved modifications (e.g., degenerate codon substitutions) and complementary sequences, as well as explicitly stated sequences. Specifically, degenerate codon substitutions can be achieved by generating a sequence in which the third position of one or more selected (or all) codons is replaced with a mixture of bases and / or deoxyinosine residues (Batzer). et al., Nucleic Acid Res. , 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. , 260:2605-2608 (1985); and Cassol et al., (1992); Rossolini et al., Mol. Cell. Probes , 8:91-98 (1994)). The terms nucleic acid and polynucleotide are used interchangeably with gene, cDNA and mRNA encoded by gene.

[0055] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein and refer to polymers of amino acid residues. This term applies to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of corresponding naturally occurring amino acids, as well as to both naturally occurring and non-naturally occurring amino acid polymers. As used herein, the term includes amino acid chains of any length, including full-length proteins (i.e., antigens), where amino acid residues are linked by covalent peptide bonds.

[0056] The term "amino acid" refers to both naturally occurring amino acids and synthetic amino acids, including amino acid analogs and amino acid mimics, which function in a manner similar to that of naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those subsequently modified, such as hydroxyproline, γ-carboxyglutamic acid, and O-phosphoserine. As used herein, "classical amino acids" refers to the 20 common amino acids used in ribosome biosynthesis: alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine. Therefore, as used herein, the terms "non-classical amino acids," "ncAA," and "non-natural amino acids" refer to amino acids not included in the aforementioned list of 20 classical amino acids. In a non-limiting sense, example ncAAs include Pyl D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2-chloroacetyl-ε-Lys. Amino acid analogs are compounds having the same basic chemical structure as naturally occurring amino acids (i.e., having an α-carbon atom bonded to a hydrogen atom, a carboxyl group, an amino group, and an R group), such as homoserine, ortholeucine, methionine sulfoxide, and methionine methylsulfonium. Such analogs have modified R groups (e.g., ortholeucine) or modified peptide backbones, but retain the same basic chemical structure as naturally occurring amino acids. An "amino acid mimic" is a compound having a structure different from the general chemical structure of an amino acid, but functioning in a manner similar to that of a naturally occurring amino acid. Amino acids are referred to herein by their commonly known three-letter symbols or by the single-letter symbols recommended by the IUPAC-IUB Committee on Biochemistry Nomenclature. Similarly, nucleotides are referred to by their commonly accepted single-letter codes.

[0057] In the context of two nucleic acids or peptides, the phrases “percent identical,” “percent identity,” or equivalents refer to a sequence that has at least a specified level of identity (e.g., at least 50% sequence identity) with a reference sequence (e.g., any peptide sequence or any nucleotide sequence included herein). Optionally, the identity percentage can be any integer from 50% to 100%. Using the procedures described herein (e.g., BLAST with standard parameters as described below), some embodiments include at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity compared to a reference sequence.

[0058] For sequence comparisons, a reference sequence is typically used, and the test sequence is compared to this reference sequence. When using a sequence comparison algorithm, the test and reference sequences are input into the computer. If necessary, the coordinates of the subsequences are specified, along with the sequence algorithm program parameters. Default program parameters can be used, or optional parameters can be specified. The sequence comparison algorithm then calculates the percentage of sequence identity between the test sequence and the reference sequence based on the program parameters.

[0059] As used herein, a “comparison window” comprises any number of segments selected from 20 to 600 (typically about 50 to 200, more typically about 100 to 150) consecutive positions within which, after optimal alignment of two sequences, the sequences are compared with a reference sequence having the same number of consecutive positions. Alignment methods for the sequences used for comparison are well known in the art. Optimal alignment of the sequences used for comparison can be performed, for example, by Smith & Waterman... Adv. Appl. Math 2:482 (1981) Local homology algorithm; by Needleman & Wunsch, J. Mol. Biol. The homology alignment algorithm of 48:443 (1970); through Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA Similarity search methods of 85:2444 (1988); by computer implementation of these algorithms (GAP, BESTFIT, FASTA and TFASTA in Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI); or by manual comparison or visual inspection.

[0060] The algorithms suitable for determining sequence identity percentage and sequence similarity are BLAST and BLAST 2.0, which are described in Altschul respectively. et al (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res 25: 3389-3402. The software used for BLAST analysis is publicly available from the National Center for Biotechnology Information (NCBI) website. The algorithm involves first identifying high-scoring sequence pairs (HSPs) by recognizing short words of length W in the query sequence. When a short word aligns with a word of the same length in a database sequence, it is either a perfect match or satisfies some positive threshold score T. T is called the neighborhood word score threshold (Altschul et al, ibid.). These initial neighborhood word matches serve as seeds to begin the search for longer HSPs containing them. The word matches are then extended in both directions along each sequence until the cumulative alignment score reaches a level that can be increased. For nucleotide sequences, the cumulative score is calculated using parameters M (reward score for matching residue pairs; always > 0) and N (penalty score for non-matching residues; always < 0). For amino acid sequences, a score matrix is ​​used to calculate the cumulative score. The extension of word matching in each direction stops when: the cumulative alignment score decreases by X from its maximum value; the cumulative score becomes zero or lower due to the accumulation of one or more negatively scored residues; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the alignment sensitivity and speed. The BLASTN program (for nucleotide sequences) uses a word length (W) of 28, an expected value (E) of 10, M = 1, N = -2, and a comparison of two strands as default values. For amino acid sequences, the BLASTP program uses a word length (W) of 3, an expected value (E) of 10, and a BLOSUM62 scoring matrix as default values ​​(see Henikoff & Henikoff). Proc. Natl. Acad. Sci. USA 89:10915 (1989)).

[0061] The BLAST algorithm also performs statistical analysis of the similarity between two sequences (see, for example, Karlin & Altschul). Proc. Natl. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which provides an indication of the probability that a match will occur by chance between two nucleotide or amino acid sequences. For example, if the minimum sum probability in a comparison of the test nucleic acid with the reference nucleic acid is less than about 0.01, more preferably less than about 10, then the probability is significantly higher. -5 The optimal value is less than approximately 10.-20 If the nucleic acid is similar to the reference sequence, then it is considered to be similar.

[0062] For sequence comparisons, typically one sequence is used as a reference sequence, and the test sequence is compared to this reference sequence. When using a sequence comparison algorithm, the test and reference sequences are input into the computer. If necessary, the coordinates of the subsequences are specified, along with the sequence algorithm program parameters. Default program parameters can be used, or optional parameters can be specified. The sequence comparison algorithm then calculates the percentage of sequence identity between the test sequence and the reference sequence based on the program parameters.

[0063] An "expression assembly" is a recombinant or synthetically produced nucleic acid construct having a series of designated nucleic acid elements that allow a specific polynucleotide sequence to be transcribed in a host cell. An expression assembly can be part of a plasmid, a viral genome, or a nucleic acid fragment. Typically, an expression assembly comprises a polynucleotide to be transcribed that is operatively linked to a promoter. In this context, "operatively linked" means that two or more genetic elements (e.g., a polynucleotide coding sequence and a promoter) are positioned relative to each other to allow the elements to properly perform their biological functions (e.g., the promoter directs the transcription of the coding sequence). Other elements that may be present in an expression assembly include those that enhance transcription (e.g., enhancers) and terminate transcription (e.g., terminators), as well as those that confer a certain binding affinity or antigenicity to the recombinant protein produced by the expression assembly.

[0064] When used to refer to, for example, cells, nucleic acids, proteins, or vectors, the terms "recombinant" or "engineered" indicate that the cell, nucleic acid, protein, or vector has been modified by introducing a heterologous nucleic acid or protein or by altering the native nucleic acid or protein, or that the cell is derived from a cell that has been so modified. Such modifications are typically accomplished by manipulating isolated nucleic acid segments and can include techniques such as genetic engineering. Therefore, recombinant or engineered cells express genes not found in the natural (non-recombinant or non-engineered) form of the cell, or native genes that are abnormally expressed, poorly expressed, or not expressed at all in recombinant or engineered cells.

[0065] As used herein, “inhibiting” or “inhibition” refers to any detectable negative impact on a target biological process, such as protein-protein specific binding or interaction, biological activity of a target protein, RNA / protein expression of a target gene, cell signal transduction, cell proliferation, presence / level of an organism (particularly microorganisms), any measurable biomarker, bioparameter, or symptom of the subject. Typically, inhibition is reflected as a reduction of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or greater in the target process (e.g., by inhibiting α-synuclein aggregation using ncAA-circulated proteins) or any of the aforementioned downstream parameters compared to a control. “Inhibition” also includes a 100% reduction (i.e., complete elimination, prevention, or abolition) of the target biological process or signal or disease / symptom. Other relative terms, such as “suppressing,” “suppression,” “reducing,” and “reduction,” are used in a similar manner in this disclosure to refer to a reduction of a target biological process or signal or disease / symptom to different levels (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or greater, compared to a control level) up to complete elimination. On the other hand, terms such as “activate,” “activating,” “activation,” “increase,” “increasing,” “promote,” “promoting,” “enhance,” “enhancing,” or “enhancement” are used in this disclosure to include positive changes at different levels in the incidence of a target process, signal, or symptom / disease (e.g., at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, or more, such as 3-fold, 5-fold, 8-fold, 10-fold, or 20-fold increase compared to a control level).

[0066] As used herein, “increase” or “decrease” refers to a detectable positive or negative change in quantity (such as the average level of α-synuclein aggregation as measured by a thioflavin T assay) compared to a comparative control (e.g., an established standard control). An increase is a positive change that is typically at least 10%, or at least 20%, or 50%, or 100% of the control value, and can be as high as at least 2, or at least 5, or even 10 times the control value. Similarly, a decrease is a negative change that is typically at least 10%, or at least 20%, 30%, or 50% of the control value, or even as high as at least 80% or 90% of the control value. Other terms indicating changes or differences in quantity relative to a comparative basis, such as “more,” “less,” “higher,” and “lower,” and terms indicating actions that cause these changes or differences, such as “increase,” “promote,” “enhance,” “decrease,” “inhibit,” and “repress,” are used in this application in the same manner as described above. In contrast, the terms "substantially the same" or "substantially no change" indicate that the quantity has changed little or no compared to the standard control, typically within ±10% of the standard control, or within ±5%, 2%, or even less.

[0067] As used herein, the singular forms “an,” “a,” and “the” include plural indicators unless the context explicitly specifies otherwise. Thus, for example, references to “ncAA” optionally include combinations of two or more such molecules, etc.

[0068] As used herein, the term "about," when modifying any quantity, refers to a variation in that quantity that is commonly encountered by those skilled in the art. For example, the term "about" refers to normal variation encountered in measurements of a given analytical technique, both within a batch or sample and between batches or samples. Therefore, the term "about" can include variations of + / - 1% to 10% of a measured value, such as variations of + / - 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of a measured value. The quantities disclosed herein include quantities equivalent to those quantities, including quantities modified or unmodified by the term "about."

[0069] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0070] II. Introduction

[0071] This article discloses recombinant cells and related polynucleotides, peptides, tRNAs, as well as methods for intracellular synthesis of non-classical amino acids (ncAAs) and proteins containing ncAAs. In some embodiments, ncAAs are D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK. Intracellular synthesis of ncAAs and proteins containing ncAAs is achieved by modifying the archaea pyrrolysine (Pyl;O) biosynthetic pathway.

[0072] Archaea are single-celled organisms that resemble bacteria in appearance but possess distinct characteristics. Archaea are prokaryotes lacking a nucleus. Archaic cells possess a series of enzymes, PylB, PylC, and PylD; regarding the biosynthesis of Pyl, the 22nd proteogenic amino acid (Meng, Kexin) appears in proteins synthesized in archaic cells and some bacterial cells. et al. Frontiers in microbiology vol. 13 1007832. 8 Sep. 2022; Wan, Wei et al. Biochimica et Biophysica Acta vol. 1844,6 (2014): 1059-70). Archaea cells also contain pyrrolysine-tRNA synthetase (PylRS) and pylT tRNA genes, which are responsible for pyrrolysine-tRNA (tRNA) synthesis. Pyl The synthesis of ) and thus the ability to incorporate Pyl into nascent polypeptide chains with corresponding amber (UAG) codons (Wan, Wei et al. Biochimica et Biophysica Acta vol. 1844,6 (2014): 1059-70). In the genome of Methanocytococcaceae, the pyl gene is... pylTSBCD It exists in the form of continuous gene clusters, in which pylT and pylS Encoding tRNA respectively Pyl and pylRS (Borrel, Guillaume) et al. Archaea (Vancouver, BC) vol.2014 374146. 27 Jan. 2014).

[0073] Amber codon suppression refers to the use of TAG / UAG (stop) codons as coding codons in translation. Complementary amber tRNA CUA Aminoacylation of the tRNA by orthogonal aminoacyl-tRNA synthetase CUAIt is specifically designed to accept only ncAAs. This results in the incorporation of one or more ncAAs into the protein. The development of amber codon suppression has been discussed in the following literature, for example, Brabham, Robin, and Martin A. Fascione. Chembiochem: a European Journal of Chemical Biology vol. 18,20 (2017): 1973-1983; and Wals, Kim, and Huib Ovaa. Frontiers in Chemistry vol. 2 15. 1 Apr. 2014.

[0074] The inventors mutated the archaea enzymes PylC and PylRS through site-directed saturation mutagenesis and directed evolution, respectively, and obtained PylC and PylRS variants capable of synthesizing ncAA (e.g., but not limited to, D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK) and their homologous tRNAs, respectively. Therefore, this disclosure provides compositions and methods for engineering enzymes that support the biosynthesis of ncAA and their corresponding tRNAs, thereby enabling the incorporation of these ncAAs into various recombinant proteins. These methods and compositions represent improvements in the art, for example, Lee... et al. Chemistry Europe The same applies to 15(12):1769-1772, 2014); U.S. Patent No. 8,921,571; and U.S. Patent Application Publication No. 2014 / 0302553.

[0075] In one aspect, this disclosure provides a PylC mutant capable of linking a first amino acid to a second amino acid, i.e., reacting a reactive group on the first amino acid with a lysine residue on the second amino acid, and at the ε-nitrogen (N) of the lysine residue. ε Covalent bonds are formed at positions 177, 179, 233, and 256 to generate ncAA. In some embodiments, ncAA is D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK. In some embodiments, mutant PylC is derived from wild-type archaea PylC and contains four point mutations at positions 177, 179, 233, and 256 corresponding to the amino acid sequence of wild-type archaea PylC (e.g., *Methanococcus martensii* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)).

[0076] On the other hand, this disclosure provides wild-type PylRS or PylRS mutants that are capable of linking ncAA (e.g., but not limited to D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK) to tRNA to produce tRNA.ncAA In some implementations, the mutant PylRS is derived from the wild-type archaea PylRS and contains three point mutations at positions 14, 348, and 451 corresponding to the amino acid sequence of the wild-type archaea PylRS.

[0077] On the other hand, this disclosure provides genetically modified host cells expressing PylC, PylRS, and tRNA, which support the synthesis of ncAAs (e.g., but not limited to D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK) and the incorporation of ncAAs into newly synthesized recombinant proteins. Thus, these cells are capable of supporting intracellular synthesis of proteins containing such ncAAs, which are strategically placed in predetermined locations to allow for protein modifications, such as cyclization, oligomerization, and PEGylation. Specifically, the recombinant host cell comprises (1) a polynucleotide sequence encoding a mutant PylC capable of synthesizing ncAAs; (2) a polynucleotide sequence encoding wild-type or mutant PylRS capable of linking ncAAs to tRNAs; and (3) a polynucleotide sequence encoding a tRNA loaded with PylRS and ncAAs, wherein the tRNA incorporates ncAAs into the nascent polypeptide chain.

[0078] In some embodiments, the recombinant host cell has been genetically modified to inactivate its endogenous releasing factor 1 (RF1), thereby enhancing readthrough at the UAG codon, where ncAA (e.g., but not limited to, D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK) is incorporated into the protein at the UAG codon. For example, the genomic sequence encoding RF1 in the recombinant host cell can be deleted, truncated, or mutated to reduce or eliminate RF1 activity. In some embodiments, the host cell is a prokaryotic cell, such as a bacterial cell, like *Escherichia coli*. In some embodiments, the host cell is a eukaryotic cell.

[0079] In some embodiments, such as as a result of PylC activity, the recombinant host cell comprises D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK. In some embodiments, such as as a result of PylRS activity, the recombinant host cell comprises D-Cys-ε-Lys-tRNA, D-Pra-ε-Lys-tRNA, D-allyl-ε-Lys-tRNA, or 2ClAcK-tRNA.

[0080] This document also provides methods for cellular synthesis of proteins containing the ncAA disclosed herein, and said methods include using recombinant host cells disclosed herein. In some embodiments, the ncAA in the protein can be used to generate cyclized, oligomerized, and / or PEGylated proteins. In some embodiments, the protein is modified intracellularly, while in other embodiments, the protein is modified in vitro. In some embodiments, proteins modified according to the methods disclosed herein (e.g., cyclized, oligomerized, or PEGylated) exhibit longer serum half-lives, increased resistance to proteolytic degradation, improved serum stability, improved binding to targets, and / or improved efficiency.

[0081] III. PylC

[0082] PylC is capable of catalyzing the ATP-dependent linkage of lysine and a second substrate to produce ncAA, for example, D-cysteyl-ε-L-lysine (D-Cys-ε-Lys), D-N6-(2-(R)-propargylglycyl)lysine (D-Pra-ε-Lys), D-allyl-N ε -L-lysine (D-allyl-ε-Lys) or 2-chloroacetyl-N ε Enzymes containing -L-lysine (2ClAcK). Many wild-type archaea and bacteria PylC, including proteins similar to archaea or bacteria PylC (e.g., proteins with at least 50% identity to archaea or bacteria PylC), can be used to produce PylC variants capable of generating ncAA. Non-limiting examples of PylC and PylC-like proteins include: *Methanococcus acetate* (… Methanosarcina acetivorans (Strain ATCC 35395 / DSM 2834 / JCM 12185 / C2A) PylC (UniProt ID: Q8TUC0), Methanogenic Micrococcus pasteurella ( Methanosarcina barke ri) (strain Fusaro / DSM 804) PylC (UniProt ID: Q46E79), Methanococcus martensii ( Methanosarcina mazei (Strain ATCC BAA-159 / DSM 3647 / Goe1 / Go1 / JCM 11833 / OCM 88) (Ferris methanogens) Methanosarcina frisia PylC (UniProt ID: Q8PWY3), methyl-utilizing methanogenic cocci ( Methanococcoides methylutens 3-Methylornithine-L-lysine ligase (UniProt ID: A0A099T2V5), *Methanococcus thermophilus* ( Methanosarcina thermophilaCHTI-55 strain pyrrolysine synthase (UniProt ID: A0A0E3HAD6), *Methanococcus methanans* WH1 pyrrolysine synthase (UniProt ID: A0A0E3L5I7), *Methanococcus thermophilus* (strain ATCC 43570 / DSM 1825 / OCM 12 / VKM B-1830 / TM-1) pyrrolysine synthase (UniProt ID: A0A0E3NDZ8), *Methanococcus methanans* WWM596 pyrrolysine synthase (UniProt ID: A0A0E3NNH7), *Methanococcus sicily* (… Methanosarcina siciliae T4 / M pyrrolidone synthase (UniProt ID: A0A0E3P0K1), *Methanococcus methanans* MTP4 pyrrolidone synthase (UniProt ID: A0A0E3P2G1), *Methanococcus sicily* HI350 pyrrolidone synthase (UniProt ID: A0A0E3P9R8), *Methanococcus sicily* C2J pyrrolidone synthase (UniProt ID: A0A0E3PJB5), *Methanococcus martensii* WWM610 pyrrolidone synthase (UniProt ID: A0A0E3Q1K8), *Malvastatinus methanans* (… Methanosarcina vacuolata Z-761 pyrrolysine synthase (UniProt ID: A0A0E3Q9K9), *Methanococcus kolksee* pyrrolysine synthase (UniProt ID: A0A0E3QFB4), *Methanococcus pasteurella* Wiesmoor strain pyrrolysine synthase (UniProt ID: A0A0E3QJ27), *Methanococcus pasteurella* MS pyrrolysine synthase (UniProt ID: A0A0E3QRW9), *Methanococcus pasteurella* 227 pyrrolysine synthase (UniProt ID: A0A0E3R2H9), *Methanococcus martensii* SarPi pyrrolysine synthase (UniProt ID: A0A0E3REK9), *Methanococcus martensii* S-6 pyrrolysine synthase (UniProt ID: A0A0E3RL10), *Methanococcus martensii* LYC pyrrolidone synthase (UniProt ID: A0A0E3RQK1), *Methanococcus martensii* C16 pyrrolidone synthase (UniProt ID: A0A0E3S363), *Methanococcus pasteurella* 3 pyrrolidone synthase (UniProt ID: A0A0E3SNN3), *Methyl-utilizing Methanococcus* MM1 pyrrolidone synthase (UniProt ID: A0A0E3SP52), *Lake-sedimentary Methanococcus* ( Methanosarcina lacustrisZ-7289 pyrrolidone synthase (UniProt ID: A0A0E3WSH3), Methanococcus pyrenoidosa ( Methanosarcina horonobensis HB-1 = JCM 15518 pyrrolysine synthase (UniProt ID: A0A0E3WWW5), Methanogenic Micrococcus pasteurella CM1 pyrrolysine biosynthesis protein PylC (UniProt ID: A0A0G3CBZ6), methanophilic bacteria ( Methanohalophilus DAL1 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A1B8WZK4), DAL1 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A1B8WZU9) of halophilic methanogens, 2-GBenrich protein containing ATP-grasp domain of halophilic methanogens (UniProt ID: A0A1D2UZX5), A14 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A1D2WWF0) of halophilic methanogens, Ant1 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A1E7GEX0) of halophilic methanogens, and *Vibrio volcanopteris* (… Methanococcoides vulcani )Pyrrolidone-lysine biosynthetic protein PylC (UniProt ID:A0A1H9Z9J2), deep-sea methanogenic bacteria ( Methanolobus profundi Pyrrolidone biosynthetic protein PylC (UniProt ID: A0A1I4T458), thermophilic methanogenic pyrrolidone biosynthetic protein PylC (UniProt ID: A0A1I6ZM81), halophilic methanogenic bacteria ( Methanohalophilus halophilus 3-Methylornithine-L-lysine ligase PylC (pyrrolelysine biosynthetic protein PylC) (UniProt ID: A0A1L3Q0L2), *Methanophilus lucida* ( Methanohalophilus portucalensis FDF-1 3-methylornithine-L-lysine ligase PylC (pyrrolelysine biosynthesis protein PylC) (UniProt ID: A0A1L9C5H7), methylmethanophiles ( Methanomethylovorans PtaU1.Bin093 3-Methylornithine-L-lysine ligase (UniProt ID: A0A1V4YHV6), halophilic bacterium eutum ( Methanohalophilus euhalobius PylC (UniProt ID: A0A285FYL0), a pyrrolidone-based biosynthetic protein, and *Dystrophus methanans* ( Methanosarcina spelaei3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A2A2HRV9), *Methanophilus lucida* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A2D3C6Q6), *Halophilus evo-malachite* 3-methylornithine-L-lysine ligase PylC (pyrrolelysine biosynthetic protein PylC) (UniProt ID: A0A315A1A6), *Methanophilus thermophilus* pyrrolelysine synthase (UniProt ID: A0A3G9CRP1), *Methanophilus lucida* RSK 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A3M9LRP0), *Methanophilus methanophilus* MSH10X1 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A2A2HRV9), A0A498GWN6), halophilic methanogen WG1-DM containing ATP-grasp domain protein (UniProt ID: A0A498H9H1), halophilic methanophores (A0A498GWN6), Methanolobus halotolerans 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A4E0Q579), *Methanococcus martensii* (*Methanococcus feris*) 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A4P8QYJ1), *Methanococcus faecalis* (*Methanococcus xanthoides*) Methanosarcina flavescens 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A660HT74), *Methanophora zinceri* ( Methanolobus zinderi 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A7D5I3G2), Methanocytic Archaea ( Methanosarcinaceae archaeon 3-Methylornithine-L-lysine ligase PylC (UniProt ID:A0A7J4PI19), Archaea of ​​the Methanocytiales order ( Methanosarcinales archaeon 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A7L4QV83), *Methanophora volcanoglycens* ( Methanolobus vulcaniPyrrololysine biosynthetic protein PylC (UniProt ID: A0A7Z7AZJ6), *Methanococcus volcanicus* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A7Z8P2G7), *Methanococcus methanans* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A832LB55), *Methanococcus acetate* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A832SGS6), *Methanococcus methanans* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A832U1W8), *Methanococcus methanans* 3-methylornithine-L-lysine ligase PylC (UniProt ID: A0A847PSL9), *Methanococcus leucosus* (… Methanococcoides seepicolus 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0A9E5D7X3), 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0AA51UHC3) for *Methanophora FTZ2*, 3-Methylornithine-L-lysine ligase PylC (UniProt ID: A0AA51UQ13) for *Methanophora FTZ6*, and *Methanophora martensii* (… Methanohalophilus mahii (Strain ATCC 35705 / DSM 5219 / SLP) contains ATP-grasp domain protein (UniProt ID: D5E9G2), and is a cryptometazobacter species. Methanohalobium evestigatum (Strain ATCC BAA-1072 / DSM3721 / NBRC 107634 / OCM 161 / Z-7303) Contains ATP-grasp domain protein D7E781, *Methanobacterium quaternaria* ( Methanosalsum zhilinae (Strain DSM 4017 / NBRC 107636 / OCM 62 / WeN5) (Halophilus quaternariae) Methanohalophilus zhilinae Contains ATP-grasp domain protein (UniProt ID: F7XLV6), and Dutch methanogen ( Methanomethylovorans hollandica (Strain DSM 15978 / NBRC 107637 / DMS1) Pyrrolysine biosynthetic protein PylC (UniProt ID: L0KW95), *Methanococcus martensii* Tuc01 pyrrolysine synthase (UniProt ID: M1Q3M7), *Bacillus burgdorferi* ( Methanococcoides burtonii(Strain DSM 6242 / NBRC 107633 / OCM 468 / ACE-M) contains ATP-binding proteins with the DUF201 domain, pyrrolidone biosynthesis (UniProt ID: Q12UB8), *Methanococcus pastoris* PylC (UniProt ID: Q8NKQ4), and *Methanococcus tindarius* (…). Methanolobus tindarius DSM 2278 PylC (UniProt ID: W9DXF9), a pyrrolidone biosynthetic protein.

[0083] In some implementations, site-directed saturation mutagenesis is a strategy used to generate PylC variants. In site-directed saturation mutagenesis, one or more sites along the protein sequence are identified as potentially harboring beneficial or desired mutations, and then randomized, i.e., amino acids at these sites are replaced with random amino acids. Site-directed saturation mutagenesis is discussed in the following literature: Nov, Yuval. Applied and environmental microbiology vol. 78,1 (2012): 258-62; and Gupta, Kritika, and Raghavan Varadarajan. Current opinion in structural biology vol. 50 (2018): 117-125.

[0084] In some embodiments, the PylC variant is derived from the wild-type archaea PylC polypeptide sequence, such as, but not limited to, PylC of *Methanococcus martensii* (UniProt ID: Q8PWY3; SEQ ID NO: 1). Non-limiting examples of PylC are discussed at Gaston, Marsha A. et al. Current Opinion in Microbiology vol. 14,3 (2011):342-9.

[0085] In some embodiments, the PylC variant contains at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylC polypeptide sequence (e.g., *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)). In some embodiments, the PylC variant also contains a polypeptide sequence in which the amino acids corresponding to positions S177, E179, D233, and / or T256 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1) are substituted with non-natural amino acids. In some embodiments, the amino acid corresponding to S177 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1) can be any amino acid other than S. In some embodiments, the amino acid corresponding to E179 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1) can be any amino acid other than E. In some embodiments, the amino acid corresponding to D233 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1) can be any amino acid other than D. In some embodiments, the amino acid corresponding to T256 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1) can be any amino acid other than T.

[0086] In some embodiments, the PylC variant is derived from the polypeptide sequence of *Methanococcus martensii* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1), as shown in SEQ ID NO: 1 below. In some embodiments, the PylC variant includes mutations at one or more of S177, E179, D233, and / or T256 in SEQ ID NO: 1.

[0087] a. A PylC variant for the synthesis of D-cysteyl-ε-L-lysine (D-Cys-ε-Lys).

[0088] In some implementations, the PylC variant is able to generate D-Cys-ε-Lys by using D-cysteine ​​(D-Cys) and L-lysine (L-Lys) as substrates. Figure 2AD-Cys-ε-Lys is a diamino acid composed of the ε-nitrogen linkage of D-Cys and L-Lys. D-Cys-ε-Lys can be used in native chemical linking reactions, i.e., reacting a C-terminal peptide thioester with an N-terminal cysteine ​​peptide to create a native peptide bond between the two fragments. Non-limiting examples of suitable reactions (e.g., protein cyclization and ubiquitination) are discussed in the following literature: Lee, Marianne M. et al., Chembiochem : a European Journal of Chemical Biology vol. 15,12 (2014): 1769-72; Li, Xin et al., Angewandte Chemie (International ed. In English) vol. 48,48 (2009): 9184-7; and Tai, Jingxuan et al., Journal of the American Chemical Society vol. 145,18(2023): 10249-10258. doi:10.1021 / jacs.3c01291.

[0089] In some embodiments, the D-Cys-ε-Lys PylC variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylC polypeptide sequence (e.g., *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)). In some embodiments, D-Cys-ε-Lys PylC also comprises mutations at positions corresponding to S177, E179, D233, and / or T256 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1), wherein one or more of these amino acids are substituted with one or more non-natural amino acids. Exemplary amino acid mutations for S177, E179, D233, and / or T256 are provided in Table 1 below. In some implementations, the PylC variant contains the mutations S177N, E179P, D233S, and T256V (PylC). NPSV mutation).

[0090]

[0091] In some implementations, the PylC variant polynucleotide sequence (coding sequence) comprises SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 22, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 28, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 44, SEQ ID NO: 46, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 59. The sequence identity is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of any of the PylC variant polynucleotide sequences shown in SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 64, SEQ ID NO: 66, SEQ ID NO: 68, SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, or SEQ ID NO: 88.In some embodiments, the PylC variant polypeptide sequence comprises the sequences corresponding to SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 25, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 39, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 45, SEQ ID NO: 47, SEQ ID NO: 49, SEQ ID NO: 51, SEQ ID NO: 53, SEQ ID NO: 55, SEQ ID NO: 57, SEQ ID NO: 59, and SEQ ID NO: 51. The sequence must possess at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity among any one of the PylC variant polypeptide sequences shown in SEQ ID NO: 61, SEQ ID NO: 63, SEQ ID NO: 65, SEQ ID NO: 67, SEQ ID NO: 69, SEQ ID NO: 71, SEQ ID NO: 73, SEQ ID NO: 75, SEQ ID NO: 77, SEQ ID NO: 79, SEQ ID NO: 81, SEQ ID NO: 83, SEQ ID NO: 85, SEQ ID NO: 87, or SEQ ID NO: 89. In some embodiments, the PylC variant polynucleotide sequence contains one or more mutations at the nucleic acids corresponding to S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant polynucleotide sequence is the sequence of SEQ ID NO: 74. In some embodiments, the PylC variant polypeptide sequence contains mutations at one or more of S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant is variant K150-19 (named PylC). NPSV (SEQ ID NO: 75).

[0092] b. PylC variants used for the synthesis of D-N6-(2-(R)-propargylglycyl)lysine (D-Pra-ε-Lys)

[0093] In some implementations, the PylC variant is able to generate D-Pra-ε-Lys by using (R)-2-aminopentan-4-ynetic acid and L-Lys as substrates. Figure 8A D-Pra-ε-Lys is a diamino acid composed of a propargyl glycine amino acid group linked to the ε-nitrogen of L-Lys. D-Pra-ε-Lys contains a modifiable alkynyl group to enable click reactions and other reactions (e.g., reactions involving addition with cysteine ​​and reactions involving thiols via thiol-alkynyl reactions), which are discussed in the following literature: Mons, Elma et al. Journal of the American Chemical Society vol.143,17 (2021): 6423-6433; Fairbanks, Benjamin D et al. Macromolecules vol.42,1 (2009): 211-217; Lowe, Andrew B. et al. Journal of Materials Chemistry vol. 20 (23): 4745; and Li, Xin et al. Chemistry, an Asian journal vol. 5,8(2010): 1765-9. doi:10.1002 / asia.201000205.

[0094] In some embodiments, the D-Pra-ε-Lys PylC variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylC polypeptide sequence (e.g., *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)). In some embodiments, D-Pra-ε-Lys PylC also comprises mutations at positions corresponding to S177, E179, D233, and / or T256 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1), wherein one or more of these amino acids are substituted with one or more non-natural amino acids. Exemplary amino acid mutations for S177, E179, D233, and / or T256 are provided in Table 2 below. In some implementations, the PylC variant contains the mutation D233H. In some implementations, the PylC variant contains the mutations S177G, E179V, D233N, and T256V (PylC GVNV mutation).

[0095]

[0096] In some implementations, the PylC variant polynucleotide sequence (coding sequence) comprises SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96, SEQ ID NO: 98, SEQ ID NO: 100, SEQ ID NO: 102, SEQ ID NO: 104, SEQ ID NO: 106, SEQ ID NO: 108, SEQ ID NO: 110, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 116, SEQ ID NO: 118, SEQ ID NO: 120, SEQ ID NO: 122, SEQ ID NO: 124, SEQ ID NO: 126, SEQ ID NO: 130, SEQ ID NO: 132, SEQ ID NO: 134, SEQ ID NO: 136, SEQ ID NO: 138, SEQ ID NO: 140 or SEQ ID NO: The sequence identity of any one of the PylC variant polynucleotide sequences shown in 142 is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%.In some embodiments, the PylC variant polypeptide sequence comprises the following SEQ ID NO: 91, SEQ ID NO: 93, SEQ ID NO: 95, SEQ ID NO: 97, SEQ ID NO: 99, SEQ ID NO: 101, SEQ ID NO: 103, SEQ ID NO: 105, SEQ ID NO: 107, SEQ ID NO: 109, SEQ ID NO: 111, SEQ ID NO: 113, SEQ ID NO: 115, SEQ ID NO: 117, SEQ ID NO: 119, SEQ ID NO: 121, SEQ ID NO: 123, SEQ ID NO: 125, SEQ ID NO: 127, SEQ ID NO: 129, SEQ ID NO: 131, SEQ ID NO: 133, SEQ ID NO: 135, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141 or SEQ ID NO: 95. Any one of the PylC variant polypeptide sequences shown in NO: 143 has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In some embodiments, the PylC variant polynucleotide sequence contains one or more mutations at the nucleic acids corresponding to S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant polynucleotide sequence is the sequence of SEQ ID NO: 142. In some embodiments, the PylC variant polypeptide sequence contains mutations at one or more of S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant is variant K200-12 (named PylC). GVNV (SEQ ID NO: 143).

[0097] c. Used in the synthesis of D-allyl-N ε PylC variant of -L-lysine (D-allyl-ε-Lys)

[0098] In some implementations, the PylC variant is able to generate D-allyl-ε-Lys by using (D)-allyl-OH and L-Lys as substrates. Figure 9AD-Allyl-ε-Lys consists of an allyl group linked to an ε-nitrogen of an L-Lys group. D-Allyl-ε-Lys contains a reactive vinyl stem that can be used in thiol-alkene reactions (also known as alkene hydrothioalkylation), which involve radical addition or Michael addition reactions between the alkene and the thiol of the vinyl stem. Thiol-alkene chemistry, thiol-Michael addition, and related click chemistry are discussed in detail in the following literature: Nolan, Mark D, and Eoin M Scanlan, Frontiers in Chemistry vol. 8 583272. 12 Nov. 2020; Hoyle, Charles E, and Christopher N Bowman. Angewandte Chemie (International ed. inEnglish) vol. 49,9 (2010): 1540-73; and Devatha P. Nair, et al., Chemistry of Materials 2014 26 (1), 724-744.

[0099] In some embodiments, the D-allyl-ε-Lys PylC variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% sequence identity with the wild-type archaea PylC polypeptide sequence (e.g., *Methanococcus martensii* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)). In some embodiments, the D-allyl-ε-Lys PylC also comprises a polypeptide sequence corresponding to *Methanococcus martensii* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1). 1) Mutations at the positions of S177, E179, D233, and / or T256, wherein one or more of these amino acids are substituted by one or more non-natural amino acids. Exemplary amino acid mutations for S177, E179, D233, and / or T256 are provided in Table 3 below. In some embodiments, the PylC variant comprises mutations S177G, E179V, D233N, and T256V (PylC GVNV mutation).

[0100]

[0101] In some embodiments, the PylC variant polynucleotide sequence (coding sequence) contains at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the PylC variant polynucleotide sequence shown in SEQ ID NO: 142. In some embodiments, the PylC variant polypeptide sequence contains at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the PylC variant polypeptide sequence shown in SEQ ID NO: 143. In some embodiments, the PylC variant polynucleotide sequence contains one or more mutations at the nucleic acids corresponding to S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant polynucleotide sequence is the sequence of SEQ ID NO: 142. In some embodiments, the PylC variant polypeptide sequence contains mutations at one or more of S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant is variant K200-12 (named PylC). GVNV (SEQ ID NO: 143).

[0102] d. PylC variants used for the synthesis of 2-chloroacetyl-ε-Lys (2ClAcK)

[0103] In some implementations, the PylC variant is able to generate 2ClAcK (using 2-chloroacetic acid and Lys as substrates) Figure 10 2ClAcK consists of a 2-chloroacetyl group linked to the ε-nitrogen of an L-Lys group. The 2-chloroacetyl group in 2ClAcK promotes protein dimerization, cyclization, and in vivo ubiquitination. In some embodiments, 2ClAcK reacts with a suitable active group (e.g., cysteine) to form a crosslink when it is very close to the active group.

[0104] In some embodiments, the 2ClAcK PylC variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylC polypeptide sequence (e.g., *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1)). In some embodiments, 2ClAcK PylC also comprises mutations at positions corresponding to S177, E179, D233, and / or T256 of *Methanococcus flavonoids* PylC (UniProt ID: Q8PWY3; SEQ ID NO: 1), wherein one or more of these amino acids are substituted with one or more non-natural amino acids. Exemplary amino acid mutations for S177, E179, D233, and / or T256 are provided in Tables 1 to 3 above. In some implementations, the PylC variant contains the mutations S177N, E179P, D233S, and T256V (PylC). NPSV (Mutations). In some implementations, PylC variants contain the mutations S177G, E179V, D233N, and T256V (PylC). GVNV mutation).

[0105] In some embodiments, the PylC variant polynucleotide sequence (coding sequence) comprises a sequence that is identical to SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 6, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 12, SEQ ID NO: 14, SEQ ID NO: 16, SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 22, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 28, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 44, SEQ ID NO: 46, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 64, SEQ ID NO: 66, SEQ ID NO: 68, SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, SEQ ID NO: 88, SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96, SEQ ID NO: 98, SEQ ID NO: 100, SEQ ID NO: 102, SEQ ID NO: 104, SEQ ID NO: 106, SEQ ID NO: 108, SEQ ID NO: 110, SEQ ID NO: 112, SEQ ID NO: 114, SEQ ID NO: 116, SEQ ID NO: 118, SEQ ID NO: 120, SEQ ID NO: 122, SEQ ID NO: 124, SEQ ID NO: 126, SEQ ID NO: 130, SEQ ID NO: 132, SEQ ID NO: 134, SEQ ID NO: 136, SEQ ID NO: 1<<END]]The sequence identity is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of any of the PylC variant polynucleotide sequences shown in SEQ ID NO: 140 or SEQ ID NO: 142. In some embodiments, the PylC variant polypeptide sequence comprises the sequences corresponding to SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 7, SEQ ID NO: 9, SEQ ID NO: 11, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, SEQ ID NO: 23, SEQ ID NO: 25, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 39, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 45, SEQ ID NO: 47, SEQ ID NO: 49, SEQ ID NO: 51, SEQ ID NO: 53, SEQ ID NO: 55, SEQ ID NO: 57, SEQ ID NO: 59. SEQ ID NO: 61, SEQ ID NO: 63, SEQ ID NO: 65, SEQ ID NO: 67, SEQ ID NO: 69, SEQ ID NO: 71, SEQ ID NO: 73, SEQ ID NO: 75, SEQ ID NO: 77, SEQ ID NO: 79, SEQ ID NO: 81, SEQ ID NO: 83, SEQ ID NO: 85. SEQ ID NO: 87, SEQ ID NO: 89, SEQ ID NO: 91, SEQ ID NO: 93, SEQ ID NO: 95, SEQ ID NO: 97, SEQ ID NO: 99, SEQ ID NO: 101, SEQ ID NO: 103, SEQ ID NO: 105, SEQ ID NO: 107, SEQ ID NO: 109. SEQ ID NO: 111. SEQ ID NO: 113. SEQ ID NO: 115. SEQ ID NO: 117. SEQ. IDNO: 119, SEQ ID NO: 121, SEQ ID NO:The sequence must possess at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity of any one of the PylC variant polypeptide sequences shown in SEQ ID NO: 123, SEQ ID NO: 125, SEQ ID NO: 127, SEQ ID NO: 129, SEQ ID NO: 131, SEQ ID NO: 133, SEQ ID NO: 135, SEQ ID NO: 137, SEQ ID NO: 139, SEQ ID NO: 141, or SEQ ID NO: 143. In some embodiments, the PylC variant polynucleotide sequence contains one or more mutations at the nucleic acids corresponding to S177, E179, D233, and / or T256 of SEQ ID NO: 1. In some embodiments, the PylC variant polynucleotide sequence is the sequence of SEQ ID NO: 74 or SEQ ID NO: 142. In some embodiments, the PylC variant polypeptide sequence contains mutations at one or more of S177, E179, D233, and / or T256 in SEQ ID NO: 1. In some embodiments, the PylC variant is variant K150-19 (named PylC). NPSV ; SEQ ID NO: 75) or variant K200-12 (named PylC) GVNV (SEQ ID NO: 143).

[0106] IV. PylRS (pyrrolidone-lysyl-tRNA synthetase)

[0107] PylRS is a pyrrolidone-lysyl-tRNA synthetase (also known as a pyrrolidone-tRNA ligase) that ligates ncAAs (e.g., pyrrolidone, D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) to transfer ribonucleic acid (tRNA). Homologous PylRS, together with succinate-repressed tRNA, plays a role in introducing ncAAs at succinate (TAG / UAG) codons during ribosomal translation of polypeptides. Figure 1A Therefore, the recombinant host cells of this disclosure typically contain the genes for both PylRS and tRNA, and simultaneously express both PylRS and tRNA. In archaea, PylRS is usually expressed by a single gene. pylSGene encoding. In some implementations, PylRS links D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK to the tRNA at the amber (TAG / UAG) codon. Non-limiting examples of PylRS and tRNA are discussed in the following literature: Wan, Wei et al. Biochimica et Biophysica Scta vol. 1844,6 (2014): 1059-70; Yuan, Jing et al. FEBS Letters vol. 584,2 (2010): 342-9; Brabham, Robin, and Martin A Fascione . Chembiochem: a European Journal of Chemical Biology vol. 18,20 (2017): 1973-1983; and Li, Wen-Tai et al. Journal of Molecular Biology vol. 385,4 (2009): 1156-64.

[0108] PylRS exhibits high substrate side chain heterogeneity, low selectivity for its substrate α-amine, and high selectivity for tRNA. Pyl The low selectivity of anticodons. These features of PylRS allow the Pyl incorporation mechanism to be engineered for the genetic incorporation of ncAA or α-hydroxy acids into proteins with amber UAG codons and the redistribution of other codons, such as ochre codons (TAA / UAA), milky codons (TGA / UGA), and tetrabasic AGGA codons, to encode ncAA (Wan, Wei et al. Biochimica et Biophysica Acta vol. 1844,6 (2014): 1059-70). Amber-repressed tRNA is discussed in detail below. In some embodiments, PylRS links Pyl, D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAck, or any combination thereof, to tRNA.

[0109] According to this disclosure, a variety of wild-type or modified archaea and bacteria PylRS can be used, such as, but not limited to, wild-type or modified PylRS from the following bacteria and archaea: *Bacillus orientalis* (…). Desulfosporosinus orientis ), Eubacterium mucilaginosa ( Eubacterium limosum ), halophilic methanogens ( Methanohalophilus mahii ), halophilic methanophiles ( Methanohalophilus halophilus ), Pasteurella multocida ( Methanosarcina barkeri ), Methanococcus martensii ( Methanosarcina mazei ), thermophilic methanococcus ( Methanosarcina thermophila), methanococcus acetate ( Methanosarcina acetivorans ), Vacuole methanococcus ( Methanosarcina vacuolata Tindarius methanophora ( Methanolobus tindarius ), Methyl-utilizing methanococci ( Methanococcoides methylutens ), desulfuric bacteria, vacuole desulfuric bacteria ( Desulfobacterium vacuolatum ), Haemophilus cryptometazoa ( Methanohalobium evestigatum ), Arabinose ( Acetohalobium arabaticum ), Methanococcus pseudomorphosus ( Methanococcoides burtonii ), Pseudomonas fibrinolyticus ( Pseudobacteroides cellulosolvens ), Biobacterium wartii ( Bilophila wadsworthia ), underground desulfurizing staphylococci ( Desulfacinum infernum ), dehalogenated and desulfurized bacteria ( Desulfitobacterium dehalogenans ), Volcano methanophora ( [[ID=4,4]]Methanolobus vulcani ), Sicilian methanococcus ( Methanosarcina siciliae ), halophilic methanogens of Portugal ( Methanohalophilus portucalensis ), *Methanobacterium quaternaria* ( Methane salt Zhilin ), Gyrocephalosporium ( Sporomusa sphaeroides ), Hafniens sulfite ( Desulfitobacterium hafniense ), methanophilic bacteria ( Methanohalophilus euhalobius ), formic acid acetic acid dehalogenated bacteria ( Dehalobacterium formicoaceticum ), Green respiratory desulfurization bacteria ( Desulfitobacterium chlororespirans ), fecal acetic acid bacteria ( Acetobacterium fimetarium ), Desulfuric Helicobacter johnsonii ( Desulfospira joergensenii ), heterotrophic iron-oxidizing carbon thermophilic bacteria ( Carboxydothermic iron-reducing ), Forest soil acetic acid rodenticide ( Sporomusa silvacetica Acetic acid oxidation desulfurization bacteria ( Desulfofarcimen acetoxidans ), Southern desulfurized spores ( Desulfosporosinus meridienii Brown thermophilic acetic bacteria ( Thermacetogenium phaeum Singapore desulfurization tubular bacteria ( Sulfuropalus Singaporean ), cockroach methanococcus ( Methanimicrococcus blatticola ), Dutch methanogen ( Methanomethylovorans hollandica ), Desulfurization bacteria ( Desulfallas gibsonii ), Gnaphalium affine ( Sporomusa acidovorans ), malonic acid rodentium ( Sporomusa malonica ), eat less parasporal bacteria ( Parasporobacterium paucivorans ), desulfurization bacteria PCE1, uncultured eubacteria, lake sediment methanogenic octopus ( Methanosarcina lacustris ), alkali-resistant desulfurizing bacteria ( Desulfitibacter alkalitolerant ), Acetobacter ferrithermos ( Thermincola ferriacetica ), uncultured *Gnaphalium affine*, and *Arsenic chloris* from the lake ( Halarsenatibacter silvermanii), Lake desulfurized spores ( Desulfosporosinus lacus ), Aminethiosulfate bacillus ( Dethiosulfatibacter aminovorans Algarviolavios nematode ( Olavius ​​of Algarve ), Delta endosymbiont, Young's desulfurized spores ( Desulfosporosinus young ), Methanococcus frenulum ( Methanosarcina horonobensis ), butyric acid amine-loving bacteria ( Aminipila buttermilk ), Megalomonas monomorpha ( Megamonas funicularis ), Deep-sea methanophytes ( Methanol deep ), Schlinke acetic acid symbiotic bacteria ( Syntrophaceticus schinkii ), Methanophora zinnei ( Methanolobus zinderi ), Campylobacter haematobium desulfurization ( Desulfosporosinus hippei ), Bacillus bacitratus ( Alkalibaculum bacchi ), Micrococcus methanophilus ( Methermicoccus shengliensis ), Haemophilus 4130, Proteus Delta 1 associated with Olavius ​​nematode, and Brassica raphe ( Scalloped cabbages ), ( Thermincola powerful ), sulfite bacteria LBE, rodenticide KB1, frozen soil methanogenic octopus ( Methanosarcina soligelidi ), cavernous methanococcus ( Methanosarcina spelei ), pyrogenic archaea ( Thermoplasmatales archaeon BRNA1, Luminidomemethane Marseilles ( Methanomassiliicoccus from Luminy PylRS1, Luminidomemethane Marseilles ( Methanomassiliicoccus luminyensis PylRS2, psychrophilic methanophores ( Methanolobus psychrophilus R15, *Rabiella ovalis* ( Sporomusa oval DSM 2662, Methanococcus pseudomorphosus AM1, Methanococcus pseudomorphosus NM1, and Dead Valley Magnetosaurus ( Desulfame plus magnetovallimortis ), Eubacterium 68-3-10, Firmicutes bacteria ( Firmicutes bacteria CAG:238, candidate intestinal methanophiles ( Methanomethylophilus alvus ), Volcanoformis ( Methanococcoides vulcani ), candidate enteric methanogenic muscarinic bacteria ( Candidate Methanomassiliicococcus intestinalis Methanogenic Micrococcus MTP4, Methanogenic Micrococcus WWM596, Methanogenic Archaea ( Methanogenic archaeon Mixed culture (ISO4-G1), *Methanococcus pyogenes* 2.HT1A.6, *Methanococcus pyogenes* 1.HA2.2, *Methanococcus pyogenes* 1.HT1A.1, *Peptococcus* bacteria ( Peptococcaceae bacteria SCADC1 2 3, Desulfurized Campylobacter HMP52, Methanogenic Archaea ISO4-H5, Peptococcus pGA-8, Desulfurized Campylobacter BICA1-9, Ratella GT1, Methanogenic Archaea DOK, Candidate Termite Methanoplasma ( Methanoplasma termite ), desulfurized Campylobacter I2, desulfurized Campylobacter BG, candidate group MSBL1 archaea SCGC-AAA382A20, Methanogenicales ( Methanomassilicoccales Archaea RumEn M1, yellow methanogenic octopus ( Methanosarcina flavescens ), desulfurizing bacteria BRH c19, candidate methanogenetic bacteria 1R26, Firmicutes bacteria ML8 F2, Timone stress bacteria ( Emergency_pilot Methanophilic bacteria T328-1, methanophores T82-4, and spirochetes ( Spirochaetes bacteria GWB1 66 5, Methanophilic bacteria PtaU1.Bin093, Archaea of ​​the Methanogenic Cyclocarya order PtaU1.Bin030, Archaea of ​​the Methanogenic Cyclocarya order PtaU1.Bin124, Archaea of ​​the Methanogenic Cyclocarya order Mx-06, Actinomycetes ADurb.BinA094, Methanophilic halophiles DAL1, Desulfurizing bacteria of the Cordyceps family ( Desulfobulbaceae bacterium S5133MH15, psychrothermic methanophores ( Methanol psychrotolerant ), Methanococcus pyogenes Ant1, Campylobacter desulfurized spores OL, Trichophyceae bacteria ( Lachnospiraceae bacteria Acetic acid bacteria MES1, candidate thermophilic methanalous archaea ( Methanohalarchaeus thermophilus PylRS1, candidate thermophilic methanogenic archaea; PylRS2, thermophilic methanogenic archaea; anaerobic methylophilic Mussa ( Methylomusa anaerophila ), Archaea of ​​the Methanocyticaceae family, Campylobacter desulfospores FKB, Bacteria of the Desulfobacterium family 4572 123, Campylobacter fructophilic desulfospores ( Desulfosporosinus fruit-eating ), candidate deep archaea ( Bathyarchaeota Archaea, Phycoccaceae ( Phycisphaeraceae Bacteria, Archaea ( Thaumarchaeota Archaea, thermophilic oleophytes ( Thermoleophilia Bacteria, methanophores SY-01, deep-sea halophilic methanogens ( Methanohalophilus profundus ), Nitrifying Cocci ( Nitrososphere Archaea, Pyrogenales ( Thermoplasmatales Archaea, desulfurized Staphylococcus, emergency bacteria 1XD21-50, methanophilic bacteria RSK, methanophilic bacteria WG1-DM, methanophilic octopus MSH10X1, amineophilic bacteria JN-18, formic acid-loving Zhao bacillus ( Zhaonella antivorans), desulfurized spore-forming Campylobacter Sb-LF, desulfurized rod-shaped bacteria ( Desulfobacillus ), Crestedomonas ( Selenomonas DSM 106892, amine-loving bacteria CBA3637, candidate Clostridium cryptotachycephala ( Cryptoclostridium obscurum ), Methanococcus pseudometaphorae SA1, candidate digestive methanogens ( Methanomethyl mesodigestus ), candidate deep-sea hydrothermal archaea ( Hydrothermarchaeum deep JdFR-18, candidate deep archaea JdFR-11, candidate primary archaea ( Korarchaeota Archaea, Methanobacteria ( Methanomicrobia Archaea ( JdFR-19 ), candidate Sif archaea ( Sifarchaeum lw 55 reseqmb2.7 and candidate *Veibenborough* archaea ( Borrarchaeum weybense For PylRS sequences from these organisms, see Guo, Li-Tao. et al. The Journal of Biological Chemistry vol. 298,11(2022): 102521.

[0110] In some embodiments, PylRS is a wild-type archaea PylRS polypeptide sequence, such as, but not limited to, PylRS of *Methanococcus martensii* (UniProt ID: Q8PWY1). In some embodiments, PylRS variants are derived from wild-type archaea PylRS polypeptide sequences, such as, but not limited to, PylRS of *Methanococcus martensii* (UniProt ID: Q8PWY1; SEQ ID NO: 144), PylRS of *Methanococcus pasteurellii* (UniProt ID: Q6WRH6), PylRS of *Methanococcus acetate* (UniProt ID: Q8TUB8), PylRS of *Havniens sulfites*, and PylRS of the methanogenic archaea ISO4-G1. In some embodiments, directed evolution is a strategy used to generate PylRS variants (Tai, Jingxuan). et al. Journal of the American Chemical Society vol. 145,18 (2023): 10249-10258).

[0111] In some embodiments, the PylRS variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylRS polypeptide sequence (e.g., *Methanococcus flavonoids* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144)). In some embodiments, the PylRS variant also comprises mutations at positions corresponding to L301, L305, Y306, L309, and / or C348 of *Methanococcus flavonoids* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144), wherein one or more of these amino acids are substituted with one or more non-natural amino acids. In some embodiments, the amino acid L301 corresponding to *PylRS* methanogenic *Dystrophus* (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than L. In some embodiments, the amino acid L305 corresponding to *PylRS* methanogenic *Dystrophus* (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than L. In some embodiments, the amino acid Y306 corresponding to *PylRS* methanogenic *Dystrophus* (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than Y. In some embodiments, the amino acid L309 corresponding to *PylRS* methanogenic *Dystrophus* (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than L. In some embodiments, the amino acid corresponding to C348 of *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than C. In some embodiments, PylRS variants comprise mutants L301M, L305I, Y306L, L309A, and / or C348F. See, for example, *Kobayashi*. et al. Journal of the American Chemical Society 2016 138 (45), 14832-14835. In some implementations, the PylRS variant includes mutations L301M, L305I, Y306L, L309A, and C348F (PylRS). mutation).

[0112] In some embodiments, the PylRS variant comprises a polypeptide sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the wild-type archaea PylRS polypeptide sequence (e.g., *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144)). In some embodiments, the PylRS variant also comprises mutations at the G14, S451, and / or C348 positions corresponding to *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144), wherein one or more of these amino acids are substituted with one or more non-natural amino acids. In some embodiments, the amino acid corresponding to G14 of *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than G. In some embodiments, the amino acid S451 corresponding to *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than S. In some embodiments, the amino acid C348 corresponding to *Methanococcus flavonoides* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144) can be any amino acid other than C. In some embodiments, the PylRS variant comprises mutant G14E, S451F, and / or C348V. In some embodiments, the PylRS variant comprises mutant G14E, S451F, and C348V (PylRS... EVF mutation).

[0113] In some embodiments, the PylRS variant polynucleotide sequence (coding sequence) contains at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the PylRS variant polynucleotide sequence shown in SEQ ID NO: 145. In some embodiments, the PylRS variant polypeptide sequence contains at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the PylRS variant polypeptide sequence shown in SEQ ID NO: 146. In some embodiments, the PylRS variant polynucleotide sequence contains one or more mutations at the G14, S451, and / or C348 nucleic acids corresponding to *Methanococcus masculinii* PylRS (UniProtID: Q8PWY1; SEQ ID NO: 144). In some embodiments, the PylRS variant polynucleotide sequence is the sequence of SEQ ID NO: 145. In some embodiments, the PylRS variant polypeptide sequence contains mutations at one or more of G14, S451, and / or C348 in *Methanococcus masculinii* PylRS (UniProt ID: Q8PWY1; SEQ ID NO: 144). In some embodiments, the PylRS polypeptide variant is PylRS EVF Variant (shown as SEQ ID NO: 146). In some embodiments, the PylRS variant is a PylRS variant containing the mutations L301M, L305I, Y306L, L309A, and C348F. Variations, such as Kobayashi et al. Journal of the American Chemical Society Discussed in 2016 138 (45), 14832-14835.

[0114] In some implementations, the PylRS capable of linking D-Cys-ε-Lys to tRNA at the amber (TAG / UAG) codon (and thus enabling the incorporation of D-Cys-ε-Lys into the protein) is the PylRS having the amino acid sequence SEQ ID NO: 144, or the PylRS having the amino acid sequence SEQ ID NO: 146. EVF ).

[0115] In some implementations, the PylRS capable of linking D-Pra-ε-Lys to tRNA at the amber (TAG / UAG) codon (and thus enabling the incorporation of D-Pra-ε-Lys into the protein) is the PylRS having the amino acid sequence SEQ ID NO: 144, or the PylRS having the amino acid sequence SEQ ID NO: 146. EVF ).

[0116] In some embodiments, the PylRS capable of linking D-allyl-ε-Lys to tRNA at the amber (TAG / UAG) codon (and thus enabling the incorporation of D-allyl-ε-Lys into the protein) is the PylRS having the amino acid sequence SEQ ID NO: 144, or the PylRS having the amino acid sequence SEQ ID NO: 146. EVF ).

[0117] In some implementations, PylRS capable of linking 2ClAcK to tRNA at the amber (TAG / UAG) codon (and thus enabling the incorporation of 2ClAcK into the protein) is PylRS having the amino acid sequence SEQ ID NO: 144, PylRS having the amino acid sequence SEQ ID NO: 146, or PylRS (PylRS... EVF ), or PylRS having the amino acid sequence of SEQ ID NO: 144, which contains mutations corresponding to L301M, L305I, Y306L, L309A and C348F of SEQ ID NO: 144 (PylRS mutation).

[0118] V. Amber inhibits transfer RNA (tRNA)

[0119] PylT Encoding wild-type tRNA Pyl Function pylT mRNA transcripts, the tRNA Pyl Contains the CUA anticodon used to identify the amber (TAG / UAG) codon. Wild-type tRNA Pyl Wild-type PylRS were loaded, and then Pyl was incorporated into the nascent polypeptide during translation. Figure 1A tRNA Pyl It exhibits orthogonality in commonly used bacterial and eukaryotic cells, and when tRNA... PylWhen coupled with PylRS in these cell lines, it plays a role in the redistribution of amber (TAG / UAG) codons. Using orthogonal pairs of PylRS and tRNA to incorporate ncAA at the amber (TAG / UAG) stop codon effectively inhibits translation termination in the presence of ncAA. Examples of amber-inhibiting and non-restrictive effects of tRNA and PylRS are discussed in the following literature: Wan, Wei et al. Biochimica et Biophysica Acta vol. 1844,6 (2014): 1059-70;Yuan, Jing et al. FEBS Letters vol. 584,2 (2010): 342-9; Brabham, Robin, andMartin A Fascione. Chembiochem: a European Journal of Chemical Biology vol.18,20 (2017): 1973-1983; Gaston, Marsha A et al. Current Opinion in Microbiology vol. 14,3 (2011): 342-9; and Li, Wen-Tai et al. Journal of Molecular Biology vol. 385,4 (2009): 1156-64.

[0120] In some implementations, the tRNA is derived from wild-type archaea tRNA. Pyl For example, but not limited to, tRNA of *Methanococcus martensii*. Pyl (SEQ ID NO: 147) or Methanococcus masculinii pylT Gene (see, e.g., GenBankLocus ID: MW879724.1). In some implementations, tRNA variants exhibit stronger tertiary interactions, e.g., the G19:C56 pair between the D-loop and T-loop in the tRNA (Serfling, Robert). et al. Nucleic Acids Research vol. 46,1 (2018): 1-10). In some embodiments, the tRNA variant also includes, in its polynucleotide sequence, a tRNA corresponding to *Methanococcus masculinii* tRNA. PylMutations at positions A11, U15, U19, U25, U29a, and / or A56 of nucleic acid (SEQ ID NO: 147). In some embodiments, the nucleic acid at A11 can be any nucleic acid other than A. In some embodiments, the nucleic acid at U15 can be any nucleic acid other than U. In some embodiments, the nucleic acid at U19 can be any nucleic acid other than U. In some embodiments, the nucleic acid at U25 can be any nucleic acid other than U. In some embodiments, the nucleic acid at U29a can be any nucleic acid other than U. In some embodiments, the nucleic acid at A56 can be any nucleic acid other than A.

[0121] In some implementations, the tRNA variant contains tRNA similar to that of *Methanococcus masculinus*. Pyl (SEQ ID NO:147) or Methanococcus masculinus pylT The tRNA possesses at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the gene (see, for example, GenBank Locus ID: MW879724.1). In some embodiments, the tRNA further comprises tRNA corresponding to *Methanococcus masculinii*. Pyl (SEQ ID NO: 147) or Methanococcus masculinii pylT One or more nucleic acid substitutions at positions A11G, U15, U19, U25, U29a, and / or A56 of the gene (see, for example, GenBank Locus ID: MW879724.1). In some embodiments, tRNA Pyl Including A11G, U15G, U19G, U25C, U29aC and A56C mutations (tRNA) M15 (Mutation). In some implementations, tRNA is as described in the examples below and in Serfling, Robert et al. Nucleic Acids Research tRNAs discussed in vol.46,1 (2018): 1-10 M15 .

[0122] VI. Polynucleotides

[0123] This article discloses the polynucleotide sequences encoding pylC, pylRS, and tRNA, which are used for the intracellular synthesis of ncAA and proteins containing ncAA. The polynucleotide sequences of PylC, PylRS, and tRNA are disclosed in the corresponding sections above.

[0124] This document also discloses polynucleotide sequences encoding target proteins, each containing one or more amber (TAG / UAG) codons for incorporating ncAA at each amber (TAG / UAG) codon in the polynucleotide sequence. In some embodiments, the polynucleotide sequence encoding the target protein also contains a stop codon for translation termination, and this stop codon is not an amber (TAG / UAG) codon, but rather an ochre (TAA / UAA) codon or a milky white (TGA / UGA) codon. A variety of proteins are suitable for incorporating ncAA into their polypeptide sequences, such as, but not limited to, therapeutic peptides, antibodies, contractile proteins, enzymes, hormone proteins, structural proteins, storage proteins, transport proteins, fibrous proteins, globular proteins, membrane proteins, and proteins conjugated with other organic and / or inorganic groups (e.g., metals, lipids, sugars, and / or phosphates). Non-limiting examples of target proteins include P16 peptides, small ubiquitin-associated modified (SUMO) proteins, and variants thereof (including mutants, truncated forms, and fusions), as discussed in the examples below. Non-limiting examples of SUMO proteins and methods of SUMOylation are disclosed in U.S. Patent Application Publication No. 2022 / 0194997.

[0125] The polynucleotide sequences disclosed herein can be introduced into host cells (e.g., *E. coli*) for expression. In some embodiments, one or more polynucleotide sequences are subcloned into one or more expression vectors. Each expression vector may contain 5' and 3' regulatory sequences operatively linked to the polynucleotide sequence. Typically, the expression vector contains a promoter for driving transcription, a transcription / translation terminator, and a ribosome-binding site for initiating translation. The expression may also contain an origin of replication, one or more selection markers, multiple restriction sites, and / or a recombination site for inserting the polynucleotide to bring it under the transcriptional regulation of the regulatory element. Suitable expression vectors and elements therein used in host cells are well known in the art and discussed in references such as Sambrook and Russell. Molecular Cloning, A Laboratory Manual (3rd ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); Ausubel et al., eds., Current Protocols in Molecular Biology (1994); and Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY.

[0126] The polynucleotide sequences disclosed herein can be modified to produce protein variants. A variety of methods readily available to those skilled in the art can be used to modify polynucleotide sequences, including, but not limited to, site-directed saturation mutagenesis, directed evolution, and site-directed mutation. See, for example, Nov, Yuval. Applied and Environmental Microbiology vol.78,1 (2012): 258-62; Gupta, Kritika, and Raghavan Varadarajan. Current Opinion in Structural Biology vol. 50 (2018): 117-125; Tai, Jingxuan et al. Journal of the American Chemical Society vol. 145,18 (2023): 10249-10258; and Bachman, Julia. Methods in Enzymology vol. 529 (2013): 241-8.

[0127] A variety of methods can be used to introduce the polynucleotides disclosed herein into host cells, including, but not limited to, electroporation, nanoparticle delivery, viral delivery, contact with nanowires or nanotubes, receptor-mediated internalization, translocation via cell-penetrating peptides, liposome-mediated translocation, DEAE dextran, lipofectamine, calcium phosphate, or any currently known or future-identified method for introducing nucleic acids into prokaryotic or eukaryotic host cells. Targeted nuclease systems (e.g., RNA-guided nucleases (e.g., CRISPR-Cas nucleases)), transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), or megaTAL (MTs) can also be used to introduce polynucleotides into the genome of host cells. See, for example, Zhang, Hong-Xia et al. Molecular Therapy: the Journal of the American Society of Gene Therapy vol. 27,4 (2019): 735-746; and González Castro, Nicolás et al. International Journal of Molecular Sciences vol. 22,19 10355. 26 Sep. 2021.

[0128] VII. Recombinant host cells

[0129] This document discloses recombinant host cells capable of expressing the disclosed PylC, PylRS, and / or tRNAs. It also discloses recombinant host cells for intracellular synthesis of ncAA and proteins containing ncAA. Non-limiting examples of ncAA include D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAck.

[0130] Recombinant host cells capable of intracellularly synthesizing ncAA and proteins containing ncAA typically contain a polynucleotide sequence encoding a mutant PylC capable of synthesizing the desired ncAA. In some embodiments, the recombinant host cell also contains a polynucleotide sequence encoding a wild-type or mutant PylRS capable of linking the desired ncAA to a tRNA. In some embodiments, the recombinant host cell further contains a polynucleotide sequence encoding a wild-type or mutant tRNA loaded with ncAA via wild-type and mutant PylRS, and the wild-type or mutant tRNA incorporates ncAA at a succinct (TAG / UAG) codon during polypeptide translation. PylC, PylRS, and tRNA, and the polynucleotide sequences encoding them, have been discussed in detail above. In some embodiments, the recombinant host cell also contains a polynucleotide sequence encoding a target protein having one or more succinct (TAG / UAG) codons for incorporating ncAA at each succinct (TAG / UAG) codon in its polypeptide sequence. In some implementations, the polynucleotide sequence encoding the target protein also contains a stop codon for terminating translation, and this stop codon is not an amber (TAG / UAG) codon, but an ochre (TAA / UAA) codon or a milky white (TCA / UGA) codon.

[0131] In some embodiments, the recombinant host cell is capable of synthesizing D-Cys-ε-Lys. In some embodiments, the recombinant host cell contains PylC NPSV In some implementations, the recombinant host cell also contains wild-type PylRS or PylRS. EVF In some implementations, the recombinant host cell also contains tRNA. Pyl or tRNA M15 In some embodiments, the recombinant host cell further comprises a polynucleotide sequence encoding the target protein having one or more amber (TAG / UAG) codons placed at predetermined positions for incorporating D-Cys-ε-Lys at each amber codon in its corresponding polypeptide sequence. In some embodiments, the polynucleotide sequence encoding the target protein includes ochre (TAA / UAA) or milky white (TCA / UGA) stop codons for translation termination.

[0132] In some embodiments, the recombinant host cell is capable of synthesizing D-Pra-ε-Lys. In some embodiments, the recombinant host cell contains PylC. GVNV In some implementations, the recombinant host cell also contains PylRS. EVF In some implementations, the recombinant host cell also contains tRNA.Pyl or tRNA M15 In some embodiments, the recombinant host cell further comprises a polynucleotide sequence encoding the target protein having one or more amber (TAG / UAG) codons placed at predetermined positions for incorporating D-Pra-ε-Lys at each amber (TAG / UAG) codon in its corresponding polypeptide sequence. In some embodiments, the polynucleotide sequence encoding the target protein includes an ochre (TAA / UAA) or milky white (TCA / UGA) stop codon for translation termination.

[0133] In some embodiments, the recombinant host cell is capable of synthesizing D-allyl-ε-Lys. In some embodiments, the recombinant host cell comprises PylC. GVNV In some implementations, the recombinant host cell also contains PylRS. EVF In some implementations, the recombinant host cell also contains tRNA. Pyl or tRNA M15 In some embodiments, the recombinant host cell further comprises a polynucleotide sequence encoding the target protein, having one or more amber (TAG / UAG) codons at predetermined positions for incorporating D-allyl-ε-Lys at each amber (TAG / UAG) codon in its corresponding polypeptide sequence. In some embodiments, the polynucleotide sequence encoding the target protein further comprises ochre (TAA / UAA) or milky white (TCA / UGA) stop codons for translation termination.

[0134] In some embodiments, the recombinant host cell is capable of synthesizing 2ClAcK. In some embodiments, the recombinant host cell comprises a PylC variant capable of synthesizing 2ClAcK. In some embodiments, the recombinant host cell also comprises PylS. or PylRS EVF In some implementations, the recombinant host cell also contains tRNA. Pyl or tRNA M15 In some embodiments, the recombinant host cell further comprises a polynucleotide sequence encoding the target protein, having one or more amber (TAG / UAG) codons at predetermined positions for incorporating 2ClAcK at each amber (TAG / UAG) codon in its corresponding polypeptide sequence. In some embodiments, the polynucleotide sequence encoding the target protein further comprises ochre (TAA / UAA) or milky white (TCA / UGA) stop codons for translation termination.

[0135] In some embodiments, the recombinant host cell is a bacterial cell. Methods for culturing bacteria and expressing proteins using bacteria (including intracellular synthesis of ncAA and proteins containing ncAA) are well known to those skilled in the art. Bacteria suitable for culturing and intracellularly synthesizing ncAA and proteins containing ncAA include Gram-negative and Gram-positive bacteria, such as Enterobacteriaceae, such as Escherichia coli spp. Escherichia ), Enterobacteriaceae ( Enterobacter Erwinia () Erwinia Klebsiella spp. Klebsiella ), Proteus spp. Proteus Salmonella ( Salmonella ), Serratia ( Serratia ), Shigella spp. Shigella ); and Bacillus spp. ( Bacilli ), Pseudomonas spp. Pseudomonas ) and Streptomyces ( Streptomyces Any strain of bacteria derived from these bacteria can be used in the methods of this disclosure. In some embodiments, the bacterial host cell is *Escherichia coli*. In some embodiments, the *E. coli* strain is strain A (K-12), B, C, or D. Suitable culture media for culturing bacteria and for protein overexpression are well known to those skilled in the art. Examples of culture medium formulations can be found in literature such as Allikian... et al., (2019). Fundamentals of Fermentation Media. In: Berenjian,A. (eds) Essentials in Fermentation Technology. Learning Materials inBiosciences. Springer, Cham .

[0136] In some embodiments, the recombinant host cell is a eukaryotic cell. Methods for culturing eukaryotic cells and for expressing proteins and tRNAs using eukaryotic cells (including intracellular synthesis of ncAA and proteins containing ncAA) are well known to those skilled in the art. Non-limiting examples of eukaryotic cells include mammalian cells, yeast, and insect cells, which are well known in the art and commercially available. Eukaryotic expression systems are available at Geisse, S. et al. Protein Expression and Purification vol. 8,3 (1996): 271-82; and Khan, Kishwar Hayat. Advanced Pharmaceutical Bulletin This was discussed in vol. 3,2 (2013): 257-63.

[0137] Standard transfection methods are used to generate bacterial, mammalian, yeast, or insect cell lines expressing the PylC, PylRS, and tRNA disclosed herein, as well as for intracellular synthesis of ncAA and proteins containing ncAA. Transformation of prokaryotic and eukaryotic cells is performed according to standard techniques (see, e.g., Morrison, ...). J. Bact. 132: 349-351 (1977); Clark-Curtiss & Curtiss, Methods in Enzymology 101: 347-362 (Wu et al., (eds., 1983). Any well-known method for introducing exogenous nucleotide sequences into host cells can be used. These methods include transfection with calcium phosphate, polyglobulin, protoplast fusion, electroporation, liposomes, microinjection, plasma vectors, viral vectors, and any other well-known method for introducing cloned genomic DNA, cDNA, synthetic DNA, or other exogenous genetic material into host cells (see, for example, Sambrook and Russell, ibid.). It is only required that the specific genetic engineering method used successfully introduces at least one gene into the host cell, thereby enabling the recombinant host cell to express the gene product (e.g., PylC, PylRS, and tRNA of this disclosure), as well as ncAA and proteins containing ncAA. The proteins of this disclosure, including proteins containing ncAA, can be purified using standard techniques (see, for example, Colley). et al., Methods in Enzymology vol. 182 (1990): 1-818; and Boneta, Laura. Nature vol. 439,7079(2006): 1018).

[0138] a. Genome modification

[0139] In some embodiments, the recombinant host cell of this disclosure further includes modifications in its genomic sequence encoding release factor 1 (RF1). RF1 is a protein that allows translation to be terminated by recognizing the UAG and UAA stop codons in the mRNA sequence. Disruption of RF1 expression levels or activity can help enhance readthrough at the amber (UAG) codon, where ncAA (e.g., but not limited to, D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK) is incorporated into the protein. For example, the genomic sequence encoding RF1 in the recombinant host cell can be deleted, truncated, or mutated to reduce or eliminate RF1 expression or activity. For example, insertions into the RF1 gene in the recombinant host cell can also reduce or eliminate RF1 expression or activity. Many genome modification systems and corresponding nucleases can be used to disrupt, reduce, or eliminate RF1 expression or activity, for example, but not limited to modifying or deleting the genomic coding sequence of RF1. Non-limiting examples of genome modification systems and nucleases include homing nuclease peptides; FokI peptides; transcription activator-like effector nucleases (TALENs); megaTAL peptides; large-scale nuclease peptides; zinc finger nucleases (ZFNs); ARCUS nucleases, etc. Large-scale nucleases can be engineered from LADLIDADG homing endonucleases (LHEs). MegaTAL peptides can contain a TALE DNA-binding domain and an engineered large-scale nuclease. See, for example, International Patent Application Publication No. WO 2004 / 067736 (Homing Endonucleases); Urnov et al. (2005) Nature 435:646 (ZFN); Mussolino et al. (2011) Nucleic Acids Research 39:9283 (TALE nuclease); Boissel et al. (2013) Nucleic Acids Research 42:2591 (MegaTAL). In some embodiments, the genome modification system is a CRISPR-Cas system comprising a CRISPR-Cas nuclease and a guide polynucleotide, such as guide RNA (gRNA). In some embodiments, for example, the CRISPR-Cas nuclease is CRISPR-Cas9. See, for example, Møller, Henrik Devitt et al. Nucleic Acids Research (2018); Burstein et al. (2017) Nature 542:237;Slaymaker et al. Science, 351(6268):84-8 (2016); and Kleinstiver et al. Nature, 529(7587):490-5 (2016).

[0140] In some embodiments, the recombinant host cell of this disclosure also includes modifications in its genome sequence, wherein one, some, or all of the natural TAG / UAG stop codons are mutated. In some embodiments, one, some, or all of the natural TAG / UAG stop codons are mutated to TAA / UAA stop codons, TGA / UGA stop codons, or a combination of both. As discussed above, the polynucleotide sequence encoding the target protein contains one or more TAG / UAG codons for incorporating ncAA at each TAG / UAG codon in the polynucleotide sequence. Therefore, in some embodiments, it is advantageous for the recombinant host cell to contain TAA / UAA stop codons and / or TGA / UGA stop codons (instead of TAG / UAG stop codons) to prevent ncAA from being incorporated into any or all polypeptides expressed by the cell other than the target protein. In some embodiments, the recombinant host cell contains TAA / UAA stop codons and / or TGA / UGA stop codons (instead of TAG / UAG stop codons) to prevent the cell's ncAA library from being used to synthesize proteins other than the target protein. In some implementations, the recombinant host cell does not contain any TAG / UAG stop codons.

[0141] b. Reagent cell lysate

[0142] In some embodiments, the recombinant host cells disclosed herein (containing the polynucleotides of this disclosure as discussed above) can be used to generate reagent cell lysates for use in cell-free expression systems. Using reagent cell lysates containing PylC, PylRS, and tRNA of this disclosure, ncAAs—D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and 2ClAcK—and / or any protein containing one or more of these ncAAs can be generated in cell-free expression systems. In some embodiments, the reagent cell lysate contains at least one PylC variant, one PylRS variant, and one tRNA variant of this disclosure.

[0143] In some embodiments, the recombinant host cells of this disclosure are cultured under suitable conditions in the presence of a suitable PylC variant substrate to produce one or more ncAAs, such as D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK. Methods for intracellular synthesis of ncAAs are discussed in detail below. Typically, the recombinant host cells are harvested from the culture and lysed to produce reagent-grade cell lysates.

[0144] In some embodiments, a polynucleotide containing the coding sequence of the desired protein is added to a reagent-grade cell lysate for cell-free expression of the desired protein. Typically, the coding sequence for the desired protein contains (1) one or more amber (TAG / UAG) codons for incorporating ncAA at each amber (TAG / UAG) codon in the protein sequence, and (2) a stop codon for terminating translation, wherein the stop codon is not an amber (TAG / UAG) codon but an ochre (TAA / UAA) codon or a milky white (TGA / UGA) codon.

[0145] In some embodiments, additional reagents are added to the reagent cell lysate for cell-free expression of the desired protein. In some embodiments, the additional reagents are one or more ncAAs (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK). In some embodiments, one or more substrates of the PylC variant in the reagent cell lysate are added to the reagent cell lysate prior to, or in addition to, cell-free expression of the desired protein to generate one or more ncAAs (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK).

[0146] Typically, reagent-grade cell lysates contain components capable of translating mRNA encoding a target protein. In some cases, reagent-grade cell lysates may contain components capable of transcribing DNA encoding the desired protein, such as, but not limited to, DNA-directed RNA polymerase (RNA polymerase), transcription activator, tRNA, aminoacyl-tRNA synthetase, 70S ribosome, N10-formyltetrahydrofolate, formylmethionine-tRNAfMet synthetase, peptidyl transferase, initiation factors (e.g., IF-1, IF-2, and IF-3), elongation factors (e.g., EF-Tu, EF-Ts, EF-G), and release factors (e.g., RF-2 and RF-3). Methods for preparing reagent-grade cell lysates may include growing the recombinant host cells of this disclosure to the exponential phase using suitable growth media and conditions. The recombinant cells are harvested and lysed to produce reagent-grade cell lysates. Methods for preparing and using reagent-grade cell lysates and cell-free expression systems are described in references such as Perez and Jessica G. et al. Cold Spring Harbor Perspectives in Biology vol. 8,12 a023853. 1 Dec. 2016; August et al. Life (Basel, Switzerland) vol. 11,12 1367. 8 Dec. 2021; Smolskaya, Sviatlana et al. International Journal of Molecular Sciences vol. 21,3 928. 31 Jan.2020; Zawada, James F. Methods in Molecular Biology (Clifton, NJ) vol. 805(2012): 31-41; Jewett et al., Molecular Systems Biology , 4, 1-10 (2008); and Shin J. and Norieaux V., J. Biol. Eng., 4:8 (2010); Zawada et al., Biotechnol Bioeng , 108:1570-1578 (2011); Yin et al., mAbs , 4(2):217-225 (2012); and Zimmerman et al., Bioconjug Chem , 25(2):351-61 (2014).

[0147] In some embodiments, the reagent cell lysate contains ncAA for producing ncAA-containing proteins. In some embodiments, reagent ncAA (e.g., ncAA from other sources) may be added to the reagent cell lysate to produce ncAA-containing proteins. In some embodiments, a substrate of the PylC variant (discussed in detail in the sections above and below regarding methods of use) is added to the reagent cell lysate for producing ncAA and / or for producing ncAA-containing proteins.

[0148] VIII. Usage Method

[0149] a. Synthesis of non-classical amino acids (ncAA) and recombinant proteins

[0150] This document also discloses methods for synthesizing ncAA intracellularly and producing recombinant proteins containing ncAA using the recombinant host cells of this disclosure. In some embodiments, the ncAA is D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, and / or 2ClAcK. Typically, the recombinant host cells of this disclosure are cultured in the presence of the substrate used in PylC for ncAA synthesis of this disclosure, and the recombinant cells are cultured under conditions suitable for ncAA synthesis and protein overexpression. In some embodiments, the recombinant host cells for synthesizing D-Cys-ε-Lys are cultured in a medium containing D-Cys and L-Lys. In some embodiments, the recombinant host cells for synthesizing D-Pra-ε-Lys are cultured in a medium containing (R)-2-aminopentan-4-ynthetic acid and L-Lys. In some embodiments, recombinant host cells for the synthesis of D-allyl-ε-Lys are cultured in a medium containing (D)-allyl-OH and L-Lys. In some embodiments, recombinant host cells for the synthesis of 2ClAcK are cultured in a medium containing chloroacetic acid and Lys. In some embodiments, the recombinant host cells are cultured in culture plates, test tubes, shake flasks, bioreactors, or fermenters. Bacterial, mammalian, and other expression systems and methods of using them are known in the art (see, for example, Ausubel). et al., 1988, Current Protocols in Molecular Biology , Wiley & Sons; Green and Sambrook, 2012, Molecular Cloning--A Laboratory Manual , 4th Ed., Cold Spring Harbor Laboratory Press, New York (2001); and Kaszubska et al., 2000, Protein Expression and Purification 18:213-20).

[0151] This document also discloses methods for generating ncAAs (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) and recombinant proteins containing ncAAs in a cell-free reaction using reagent-grade cell lysates prepared from recombinant host cells of the present disclosure. In some embodiments, the reagent-grade cell lysates contain ncAAs previously generated intracellularly. In some embodiments, the reagent-grade cell lysates are used to generate ncAAs. In some embodiments, the reagent-grade cell lysates used for the synthesis of D-Cys-ε-Lys are supplemented with D-Cys and L-Lys. In some embodiments, the reagent-grade cell lysates used for the synthesis of D-Pra-ε-Lys are supplemented with (R)-2-aminopentan-4-acetylic acid and L-Lys. In some embodiments, the reagent-grade cell lysates used for the synthesis of D-allyl-ε-Lys are supplemented with (D)-allyl-OH and L-Lys. In some embodiments, the reagent-based cell lysate used to synthesize 2ClAcK is supplemented with chloroacetic acid and Lys. Methods for preparing and using the reagent-based cell lysate and cell-free expression system are described in references such as Perez and Jessica G. et al. Cold Spring Harbor Perspectives in Biology vol. 8,12 a023853. 1 Dec. 2016; Augustet al. Life (Basel, Switzerland) vol. 11,12 1367. 8 Dec. 2021; Smolskaya,Sviatlana et al. International Journal of Molecular Sciences vol. 21,3 928.31 Jan. 2020; Zawada, James F. Methods in Molecular Biology (Clifton, NJ)vol. 805 (2012): 31-41; Jewett et al., Molecular Systems Biology , 4, 1-10(2008); and Shin J. and Norieaux V., J. Biol. Eng., 4:8 (2010); Zawada et al., Biotechnol Bioeng, 108:1570-1578 (2011); Yin et al., mAbs , 4(2):217-225(2012); and Zimmerman et al., Bioconjug Chem , 25(2):351-61 (2014).

[0152] Proteins containing ncAA can be isolated from recombinant host cells or cell-free expression systems using a variety of methods available to those skilled in the art. Standard purification methods include electrophoresis, molecular techniques, immunological techniques, and chromatographic techniques, including ion-exchange chromatography, hydrophobic chromatography, affinity chromatography, and reversed-phase HPLC. Ultrafiltration and dialysis techniques combined with protein concentration are also useful. See, for example, Scopes, 1994. Protein Purification , 3rd edition, Springer-Verlag, New York City, New York.

[0153] Methods for determining the yield or purity of purified proteins are known in the art and include, for example, Bradford assay, UV spectroscopy, biuret assay, Lowry assay, amino black assay, high-performance liquid chromatography (HPLC), mass spectrometry (MS), and gel electrophoresis methods (e.g., using protein staining, such as Coomassie blue or colloidal silver staining).

[0154] b. Site-specific modifications of proteins

[0155] Incorporating ncAA (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) into recombinant proteins as disclosed herein can be used for site-specific protein modification. Non-limiting examples of protein modifications using ncAA include protein cyclization, oligomerization, and PEGylation. In some embodiments, proteins containing ncAA and then modified with ncAA exhibit longer serum half-life, enhanced resistance to protein hydrolytic degradation, improved serum stability, improved binding to targets, and / or improved efficiency.

[0156] ncAAs (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) can be incorporated into a polypeptide sequence, wherein a cross-link (or "staple") needs to be formed between the ncAA and a second amino acid in the same polypeptide sequence. In some embodiments, a covalent bond is formed when a cross-link is formed in an intracellular or in vitro reaction, as discussed in the examples below. The cross-link is a covalent bond between a first portion on the ncAA and a second portion on the second amino acid. In some embodiments, the cross-link is formed between two ncAAs, between an ncAA and a thiol group, or between an ncAA and a thioester group (e.g., a C-terminal thioester group). In some embodiments, the formation of the cross-link causes polypeptide cyclization.

[0157] In some embodiments, cross-links form intracellularly, and the recombinant protein is separated as a cross-linked protein. In some embodiments, where cross-links on the recombinant protein formed intracellularly cause cyclization of the recombinant protein, the recombinant protein is separated as a cyclized protein.

[0158] In some embodiments, crosslinking is formed in vitro, wherein the recombinant protein is first isolated as an uncrosslinked protein and then subjected to one or more chemical reactions to produce a crosslinked protein. In some embodiments, crosslinking formed in vitro on the isolated recombinant protein causes cyclization of the isolated recombinant protein.

[0159] In some embodiments, ncAAs (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) can be used to form protein oligomers. In some embodiments, a first ncAA is incorporated into a first polypeptide sequence, and a second ncAA is incorporated into a second polypeptide sequence. In some embodiments, the method comprises two or more polypeptides, and each polypeptide contains one or more ncAAs. In some embodiments, the formation of crosslinks between two or more polypeptides causes polypeptide oligomerization. In some embodiments, the one or more ncAAs are different. In some embodiments, the one or more ncAAs are the same. In some embodiments, the two or more polypeptides are different. In some embodiments, the two or more polypeptides are the same. In some embodiments, crosslinks between ncAAs, between ncAAs and thiol groups, or between ncAAs and thioester groups (e.g., C-terminal thioesters) are formed intracellularly, and the recombinant protein separates as a crosslinked oligomeric protein. In some implementations, the cross-linking is formed in vitro, wherein the recombinant protein is first isolated as an uncross-linked protein and then subjected to one or more chemical reactions to produce a cross-linked oligomeric protein.

[0160] In some embodiments, ncAA is D-Cys-ε-Lys, and a cross-link is formed between the D-Cys moiety of D-Cys-ε-Lys and the C-terminal thioester in the same protein. In some embodiments, the protein comprises an intepid, and the C-terminal thioester is generated by incubating the protein with MESNA (sodium 2-mercaptoethanesulfonate) to catalyze the cleavage of the intepid. In these embodiments, the MESNA linked to the protein is replaced by the thiol group of D-Cys-ε-Lys.

[0161] In some embodiments, ncAA is D-Pra-ε-Lys, D-allyl-ε-Lys, or 2-chloroacetyl-ε-Lys, and a cross-link is formed between ncAA in the same protein and a cysteine ​​thiol group on another amino acid.

[0162] In some embodiments, ncAA is D-allyl-ε-Lys, and a crosslink is formed between two D-allyl-ε-Lys using a Grubbs catalyst.

[0163] In some embodiments, ncAA is 2ClAcK, and a crosslink is formed between the 2-chloroacetyl group on 2ClAcK and a suitable active group (e.g., cysteine ​​or a small molecule containing a thiol group).

[0164] This article also discloses a method for adding polyethylene glycol (PEG) groups to recombinant proteins containing ncAA. This process is called site-specific PEGylation. The PEG moiety offers numerous advantages for increasing protein stability and half-life cycling. Due to its flexibility, hydrophilicity, variable size, and low toxicity, PEGylation is used to extend protein half-life. PEG has been approved by the Food and Drug Administration (FDA) as “generally considered safe” (Gaberc-Porekar, Vladka). et al. Current Opinion in Drug Discovery & Development vol. 11,2 (2008): 242-50). Protein PEGylation is discussed in the following literature, for example, Dozier, Jonathan K, and Mark D Distefano. International Journal of Molecular Sciences vol. 16,10 25831-64. 28 Oct.2015; Liu, Yang et al. EJNMMI Research vol. 14, 1 15. 7 Feb. 2024; and Naowarojna, Nathchar et al. Synthetic and Systems Biotechnology vol. 6,1 32-49. 15 Feb. 2021.

[0165] Most therapeutic proteins rely on nonspecific PEGylation of proteins via reaction with lysine side chains and amino groups at the N-terminus. While this method facilitates the production of PEGylated conjugates, it can result in a heterogeneous mixture of PEGylated materials, where each PEGylated conjugate possesses its own activity and stability characteristics. For example, the PEGylation of INF-α2a yielded eight different PEGylated proteins via eight different lysine residues. The activity levels of these different PEGylated isomers ranged from the most active to the least active isomer by a factor of three (Foser, Stefan). et al. Protein Expression and Purification vol. 30,1 (2003): 78-87). For the PEGylated form of INF-α2b (Wang, Yu-Sen et al. Advanced Drug Delivery Reviewsvol. 54,4 (2002):547-70) and growth factor analogs (Veronese, Francesco M., ed.). Springer Science & Business Media Similar results have been observed in

[2009] . In contrast, site-specific protein PEGylation using intracellularly synthesized ncAA and ncAA-incorporated proteins disclosed herein can be used to generate predictable and homogeneous populations of proteins with the desired PEGylation.

[0166] In some embodiments, ncAA (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) can be used to inhibit proteins that have a thiol group at or near the active site of a protein. See, for example, Mons, Elma et al. Journal of the American Chemical Society vol. 143,17(2021): 6423-6433; and Yu, Yongsheng et al. Molecular Pharmaceutics vol. 14,5(2017): 1548-1557.

[0167] In some embodiments, ncAA (e.g., D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys, or 2ClAcK) can be used in engineered molecular scaffolds. See, for example, Lu, Yao et al. Bioconjugate Chemistry vol. 25,5 (2014): 989-99.

[0168] I. Reagent Kit

[0169] This document also discloses a kit comprising the live recombinant host cells, polynucleotides, PylC, PylRS, tRNA, and / or reagent cell lysates of this disclosure. The recombinant host cells can be provided in various forms, such as, but not limited to, as cryopreservation media, on agar plates, in liquid cultures, and / or in stab cultures. The polynucleotides can be provided as one or more expression vectors. PylC, PylRS, and tRNA can be provided as purified proteins or tRNA. The reagent cell lysates can be provided as frozen cell lysates prepared from the recombinant host cells of this disclosure. The kit may also include growth media, reagents, and / or instructions for use, including instructions for producing a target protein containing at least one ncAA in its polypeptide sequence.

[0170] Example

[0171] The following examples demonstrate a scalable and stable system for the intracellular synthesis of non-classical amino acids (ncAAs) coupled with the Pyl translation mechanism for protein modification.

[0172] I. Synthesis and Uses of D-Cys-ε-Lys (Overview of Examples 1 to 7)

[0173] Intracellular synthesis of ncAA D-Cys-ε-Lys is achieved by hijacking the pyrrolidone (pyl) biosynthetic pathway. This is achieved using an optimized pyrrolidone lysyl-tRNA synthetase and homologous tRNA pairs (PylRS / tRNA). pyl ) and cell lines incorporate D-Cys-ε-Lys genetics into proteins ( Figure 1A (Left figure). The system was then applied to cyclize the 23-mer therapeutic P16 peptide transplanted onto the fusion protein, resulting in near-complete cyclization of the target cyclic subunit within 3 hours. The cyclic P16 peptide fusion protein product exhibited a higher CDK4 binding affinity than its linear counterpart. Furthermore, a bifunctional bicyclic protein with a cyclic P16 peptide was produced and analyzed. This bifunctional bicyclic protein contains an arginine-glycine-aspartic acid (RGD) motif for targeting cancer cells at one end and a cyclic P16 peptide tumor suppressor (…). Figure 1A (See right figure). This bifunctional bicyclic protein has been found to be a potent cell cycle arrestor with improved serum stability.

[0174] Example 1 - Materials and Methods

[0175] Plasmid construction for PylRS directed evolution. The plasmid pPylST-KanR(TAG)-mCh(TAG) used for directed evolution of PylRS is derived from pPylST-mCh(TAG), and contains the PylRS gene from *Methanococcus martensii* and... pylT The gene, and a pETDuet-based plasmid of the mCherry gene, wherein the PylRS gene encodes PylRS (flanked by NcoI and BamHI sites), the pylT Gene encoding tRNA pyl (Flanked by XbaI sites), the lysine codon at position 55 of the mCherry gene is mutated to TAG and cloned between the NdeI and KpnI sites. 45-46 To prepare the pPylST-KanR(TAG)-mCh(TAG) plasmid, the kanamycin resistance gene encoding an aminoglycoside kinase containing a K42O mutation (i.e., KanR(TAG)) was cloned together with a gene spacer containing an additional ribosome binding site upstream of the mCh(TAG) gene. Figure 1B (The middle part). By performing two rounds of site-directed mutagenesis, such as Serfling. et al The study indicated 24 By replacing wild-type p ylTThe base generation in genes involves tRNA M15 The modified plasmids were used. The primers used were PylTm15_Fwd (SEQ ID NO: 148) and PylTm15_Rev (SEQ ID NO: 149) for the first round of mutations, and PylTstarSecondC_Fwd (SEQ ID NO: 150) and PylTstarSecondC_Rev (SEQ ID NO: 151) for the second round of mutations.

[0176] PylRS library construction. To generate a PylRS mutant library, random mutations were introduced into the genes encoding PylRS and its variants via error-prone PCR (epPCR) using the GeneMorph II random mutation kit, following the manufacturer's protocol or referring to previously reported protocols, using a Taq DNA polymerase with error-prone conditions. This resulted in a mutation frequency of approximately 2.5 nucleotides per amplicon. 47 The primers used were pylS Fwd and R4, located flanking the pylS gene. PylRS mutants containing G14E and S451F mutations were subsequently identified. Since C348 has been reported to be crucial for substrate binding, site-directed saturation mutagenesis (SSM) was performed to randomize the C348 residues of the identified PylRS mutants using primer F2-PylS-Cys348 (containing a degenerate NNN codon at C348) and R4. In each round of directed evolution, following Miyazaki's protocol, gel-purified epPCR or SSM products were subcloned into pPylSTKanR(TAG)-mCh(TAG) via large primer PCR on a complete plasmid (MEGAWHOP). 48 The PCR products obtained by treating with DpnI were then transformed into E. coli BL21 (DE3) cells for library screening.

[0177] PylRS library filtering. E. coli BL21 (DE3) cells were transformed with a mutant library via heat shock or electroporation and then revived by shaking culture at 37°C for 1 hour. The transformed cells were plated on LB agar selective plates supplemented with 50 μM IPTG, 100 μg / mL ampicillin, 2 mM D-Cys-ε-Lys, and 50 to 100 μg / mL kanamycin, and then incubated overnight at 37°C.

[0178] Positive clones selected by kanamycin were re-streaked onto LB agar containing ampicillin and incubated overnight at 37°C. Fresh clones were then inoculated into wells containing 150 μL of LB medium supplemented with 100 μg / mL ampicillin and grown at 37°C with shaking at 230 rpm for 5–6 hours. These cultures were then used for a second round of selection based on mCherry fluorescence. For this purpose, 5 μL of each culture was added to the wells of a 24-well plate containing 495 μL of LB medium or M9 glucose basal medium (containing 6 g / L Na2HPO4, 3 g / L KH2PO4, 1 g / L NH4Cl, 0.5 g / L NaCl, 3 mg / L CaCl2, 0.4% glucose, and 1 mM MgSO4 in MilliQ water) and incubated overnight at 30°C. The LB medium or M9 glucose basal medium was supplemented with 50 μM IPTG, ampicillin, and 2 mM D-Cys-ε-Lys. 150 μL of each induced culture was transferred to a black 96-well plate with a clear bottom for measurement using a Tecan Infinite M1000 PRO microplate reader. The normalized fluorescence intensity of each sample [measured fluorescence (Ex: 587 ± 5 nm; Em: 610 ± 5 nm) divided by the optical density at 700 nm] was used for selection. Cells exhibiting stronger mCherry fluorescence than the parental PylRS were selected and cultured overnight in 5 mL LB broth supplemented with ampicillin for plasmid DNA extraction.

[0179] To confirm that the selected PylRS mutant was a true positive (i.e., it did not use any endogenous amino acids as substrates and exhibited higher catalytic efficiency than the parent), mCherry fluorescence screening was repeated three times using newly transformed colonies grown in LB medium supplemented or unsupplemented with D-Cys-ε-Lys. A control without ncAA was included as a negative control.

[0180] Construction of pPylST.tL plasmid. To optimize the yield of ncAA-containing proteins, a recombinant Escherichia coli strain, C321.ΔA.M9 adapted, was used for protein overexpression. This E. coli strain was engineered to enhance the incorporation of non-classical amino acids. 29However, this *E. coli* strain is incompatible with the pETDuet-derived pPylST vector for two reasons. First, the strain is ampicillin-resistant, therefore ampicillin cannot be used as a selection marker for mutant library transformants. Second, it does not produce T7 RNA polymerase, which recognizes the T7 promoter in pETDuet-based plasmids and is essential for transcription. Therefore, in order to utilize C321.ΔA.M9 adapted cells, for cells carrying tRNA... M15 The pPylST-mCh(TAG) vector based on pETDuet and the evolutionary PylRS obtained from directed evolution studies EVF Several modifications were performed. In short, these modifications involved (1) replacing the T7lac promoter (containing tRNA) located upstream of the first multiple cloning site (MCS1) with the Ptac promoter. M15 and PylRS EVF (1) The gene was modified by replacing the T7lac promoter located upstream of the second multiple cloning site (MCS2) with the PLlacO1 promoter (containing the gene carrying the TAG codon in the frame); and (2) the ampicillin resistance gene (amR) was replaced with the streptomycin / spectinomycin resistance gene (smR). The promoter sequences are shown in SEQ ID NO: 152-155. The primers used to amplify the modified gene fragments are shown in SEQ ID NO: 156-182. Gibson assembly was used to assemble the amplicon to construct the resulting pPylST.tL-mCh(TAG) plasmid. Figure 1B (bottom part).

[0181] mCherry Readout Measurement To evaluate the efficiency of different readout systems, *E. coli* competent cells (Rosetta 2 (DE3) or C321.ΔA.M9 adapted) were transformed with the relevant pPylST-mCh(TAG) or pPylST.tL-mCh(TAG) plasmids. The transformed cells were then seeded into 50 mL of M9 medium containing appropriate antibiotics (ampicillin and / or spectinomycin) in 96-well microtiter plates. After incubation at 30°C with shaking for 5–6 hours, 250 μL of each culture was added to 48-well plates. Appropriate antibiotics, 0.5 mM IPTG, and different concentrations of D-Cys-ε-Lys were added to the medium. The plates were then incubated overnight with shaking at 25°C. The mCherry fluorescence intensity of the induced cultures was measured as described above. Three replicate measurements were performed.

[0182] A PylC mutant library was generated through site-directed saturation mutagenesis.Two sets of degenerate primers (primers: PylC-S177E179mut_Fwd, PylC-D233mut_Rev, PylC-D233mut_Fwd, and PylC-T256mut_Rev) were used to perform site-directed saturation mutagenesis on residues S177, E179, D233, and T256 of PylC fused to the N-terminal SUMO tag (for purification). The full-length SUMO-PylC fragment was extended by overlap extension PCR using two additional sets of non-mutated primers (primers: rbs-NdeI-SUMO_Fwd, PylC-S177up_Rev, PylC-T256down_Fwd, and PylC-KpnI-pDuet_Rev). The extended SUMO-PylC mutant insert was cloned into the pACYCDuet-1 vector, whose promoter was replaced with Plpp to achieve constitutive expression of PylC (primers: Plpp-rbs_Fwd). The reaction product generated a PylC mutant library, which could be screened in PylST-expressing cells.

[0183] Screening of pylC mutant libraries. For screening of the PylC mutant library, BL21(DE3) Star (Thermo Fisher) cells were transformed with the plasmid pPylST.tL-KanR(TAG)-mCh(TAG). Single colonies of transformed cells were selected to prepare competent cells containing pPylST.tLKanR(TAG)-mCh(TAG), which were subsequently used for transformation with the PylC mutant library via electroporation. Cells transformed with the mutant library were plated on LB agar plates containing spectinomycin, chloramphenicol, 50 μM IPTG, 5 mM D-cysteine, and different concentrations of kanamycin (150 μg / mL and 200 μg / mL). Each transformation yielded approximately 68,300 transformants.

[0184] Colonies grown on the positive selection plate were transferred to a 96-well microtiter plate containing 150 μL of LB medium supplemented with chloramphenicol, spectinomycin, 50 μM IPTG, and 5 mM D-cysteine ​​for quantification based on mCherry fluorescence readout efficiency as described above. The top 50 mutants producing more mCherry fluorescence than wild-type PylC were selected for plasmid extraction and subsequent DNA sequencing to identify the corresponding mutations.

[0185] Computational modeling. The modeling of PylRS C348V is based on the crystal structure of the MmPylRS C-terminal domain (CTD) bound to adenosine-modified Pyl (PDB: 2ZIM). 25The C348 residue was mutated to valine, and the ligand was modified to adenosylated D-Cys-ε-Lys in PyMOL. 49 Then, UCSF CHIMERA was used for protein structure minimization and surface hydrophobicity analysis. 28 .

[0186] use D. hafniense PylRS CTD-tRNA pyl Using the complex structure (PDB: 2ZNI)51 as a reference, MmPylRS and tRNA were manually constructed in PyMOL based on the two stored structures. pyl The complex model, wherein the two stored structures are: MmPylRS NTD-tRNA pyl Complex (PDB: 5UD5) 50 and MmPylRSCTD (PDB: 2ZIM) bound to adenosine-modified Pyl.

[0187] To generate the MmPylC model bound to D-Cys-ε-Lys, a model was first constructed in the SWISSMODEL52 server, using the crystal structure of *PylCWT* (methanococcus pasteurellii) bound to D-ornithine-ε-Lys (PDB ID: 4FFM)31 as a template. Then, D-ornithine-ε-Lys in Pymol was replaced with D-Cys-ε-Lys, and mutations (S177N, E179P, D233S, T256V) were incorporated into the PylC chain. The final structure was refined using simulated annealing / molecular dynamics from the CNS package. 32-33 Binding and surface hydrophobicity analyses were performed using UCSF CHIMERA.

[0188] Molecular dynamics (MD) simulations of P16p. MD simulations of P16p were performed using GROMACS software version 53 2021.4 with opls2001 force fields and a TIP3P water model. Topology files for D-Cys-ε-Lys were generated in GROMACS format using LigParGen server 54-56. The initial peptides were solvated through a dodecahedral water box containing approximately 4000 water molecules, and further solvated by adding Cl... -Ions were neutralized. The solvation of the system was minimized using the steepest descent method with a tolerance of 1000 KJ / mol·nm and a step size of 0.01 nm. The system was gradually heated from 0 K to 298 K over 100 ps at a pressure of 1 bar. Production runs were conducted for 300 ns in 2 fs steps. The temperature was maintained at 298 K using a modified Berendsen thermostat with a time constant of 1 ps. The pressure was maintained at 1 bar with a time constant of 2 ps using the Parrinello-Rahman scheme, and the isothermal compressibility was 4.5 × 10⁻⁶. -5 -bar -1 Long-range electrostatic interactions were calculated using the Particle Mesh Ewald (PME) method, while short-range electrostatic interactions and van der Waals interactions were calculated using a 1 nm cutoff distance. The LINCS algorithm was used to constrain all covalent bonds involving hydrogen atoms. Independent 300 ns simulations were performed on both peptides starting from the same initial structure, with each simulation repeated three times. An additional 10 rounds of 10-ns simulations were run under the same conditions used for structural comparison. RMSD and RMSF analyses were performed using algorithms in GROMACS, where least-squares fitting was calculated based on the main chain atoms.

[0189] Plasmid construction targeting D-Cys-ε-Lys-incorporated proteins. The optimized plasmid pPylST.tL was used to subclone different protein constructs for subsequent circularization studies. In short, the gene fragment encoding the target protein containing the UAG codon and an inteptide-CBD-His7 tag was inserted between the KpnI and NdeI sites via Gibson Assembly (NEB). For the construct cycRGD-mCh-cycP16p, an upstream N-terminal SUMO (small ubiquitin-like modified protein) tag was added to enhance protein expression. A linear equivalent of the D-Cys-ε-Lys-incorporated protein studied in the subcloning was used, differing in that the UAG codon in the circularized protein gene fragment was replaced with the GCG codon encoding alanine by mutating the Pfu Turbo DNA polymerase (Agilent Technologies) according to the manufacturer's instructions and primers Oto-Ala-mutant_Fwd and O-to-Ala-mutant_Rev.

[0190] Protein expression, purification, and cyclization of proteins containing D-Cys-ε-Lys. For protein expression using chemically synthesized D-Cys-ε-Lys, the pPylST.tL plasmid containing the protein construct for circularization was transformed into E. coli strain C321. Figure 6The resulting transformants were seeded into LB medium supplemented with spectinomycin and grown at 30°C / 220 rpm until the OD600 reached 0.6. The cells were centrifuged to remove the medium, and the resulting cell pellet was resuspended in 1 / 7 of the original culture volume of 2×YT medium supplemented with spectinomycin. The cells were induced with 0.5 mM IPTG and 4 mM D-Cys-ε-Lys at 25°C / 220 rpm for 16 hours.

[0191] For protein expression using intracellular or in vivo synthesized D-Cys-ε-Lys, pACYCDuet-SUMO-PylC NPSV The plasmid was co-transformed with the pPylST.tL plasmid into E. coli C321 cells and cultured in 200 mL of LB medium supplemented with spectinomycin, chloramphenicol, and 5 mM D-cysteine ​​(pH adjusted to 8.0 with 5 M Tris). When the OD600 reached 0.6, the cell culture was centrifuged and resuspended in 50 mL of LB medium and induced with 0.5 mM IPTG at 25 °C / 220 rpm for 16 hours.

[0192] At the end of the induction period, cells were collected by centrifugation and resuspended in lysis buffer (20 mM Tris-HCl pH 8.0, 500 mM NaCl, 1 mM PMSF, 1 mM benzoyl peroxide). Cells were lysed by sonication on ice for 20 min and purified by Ni2+ affinity chromatography, except for SUMO-cycRGD-O-P16p-inteptide-CBD-His7, which was purified using chitin resin to facilitate on-column cleavage by SUMO protease to remove the N-terminal SUMO tag, followed by on-column cyclization as described below. The purity of the purified protein was verified by SDS-PAGE prior to cyclization.

[0193] The cyclization of the purified protein was achieved by adding 100 mM sodium 2-mercaptoethanesulfonate (MESNA) to initiate intepitide cleavage, followed by the addition of 2 mM tris(2-carboxyethyl)phosphine (TCEP, pH adjusted to 8.0) to maintain a reducing environment. The cyclization process was carried out at room temperature for 3 hours. Afterward, the reaction mixture was incubated with chitin resin (New England Biolabs) to remove the cleaved intepitide-CBD-His7, and the cyclized protein was collected from the mobile phase. An additional cyclization step was performed to cyclize the N-terminus of the cycRGD-O-P16p construct, wherein the protein was subjected to air oxidation at 4°C with gentle agitation for 24 hours.

[0194] The linear counterpart of the cyclized protein was expressed in LB medium supplemented with 100 μg / mL ampicillin in transformed *E. coli* R2 strain cultured at 37°C / 220 rpm until an OD600 of 0.6 was reached. At this point, expression was induced for 16 hours at 25°C / 220 rpm using 0.5 mM IPTG. The linear counterpart was purified using the same purification method as the corresponding cyclized protein described above. To remove the inteptide-CBD-His7, 50 mM DTT was added to the reaction mixture, followed by chitin purification as described previously to remove the inteptide-CBDHis7.

[0195] Electrospray ionization mass spectrometry. Cut off the protein band corresponding to GFP-O-P16p and cut it into 1 mm segments. 3 The gel was decolorized by repeated washing with 50% MeOH / 10 mM NH4HCO3, then dehydrated with ACN, and subsequently vacuum dried. To reduce the protein, 25 mM DTT was added and incubated at 56°C for 1 hour, followed by washing to remove the remaining DTT and dehydration again. Trypsin was added to the dehydrated gel block and incubated overnight at 37°C for digestion. The digested peptides were extracted from the gel by sonication, and the extracted samples were then separated by HPLC and analyzed on an Orbitrap Fusion Lumos Tribrid mass spectrometer (Thermo Fisher Scientific). For mass spectrometry analysis of the intact protein, purified GFP-cycP16p was incubated with 25 mM DTT for 24 hours, then desalted using Bio-Gel P-30 size exclusion resin and denatured with 0.1% formic acid, followed by HPLC-MS analysis on an Orbitrap Fusion Lumos Tribrid mass spectrometer (Thermo Fisher Scientific). The capillary voltage was set to 3500 V. Spectra from 500 m / z to 2000 m / z were obtained at a resolution of 120,000. Mass spectra were analyzed and deconvolved using the Xtract algorithm by BioPharma Finder (Thermo Fisher Scientific).

[0196] Analytical size exclusion chromatography.To assess whether the cyclic P16p subunit on GFP-cycP16p binds to CDK4, size analysis was performed on a mixture of GST-CDK4 (0.32 mg / mL, Sino Biological) and GFP-cycP16p (0.13 mg / mL) in 25 μL SEC running buffer (20 mM Tris pH 8.0, 200 mM NaCl) using a Superdex 200 increase 5 / 150 GL analytical size exclusion column (Sigma-Aldrich) at a flow rate of 0.35 mL / min. Individual GST-CDK4 and GFP-cycP16p proteins were analyzed under the same conditions as a reference.

[0197] Combined with research. Microthermophoresis (MST) was performed to measure the biomolecular interaction between GSTCDK4 (Sino Biological) and linear or cyclic MBP-P16p. Target MBP-cycP16p or MBP-P16p was dialyzed against labeling buffer (20 mM Tris, 200 mM NaCl, pH 8.0) for 2 h, followed by labeling with Alexa Fluor 647 NHS ester dye (Thermo Fisher Scientific) at a molar ratio of 1:10 (protein:dye) for 2 h at room temperature. Free dye was removed using a gravity flow column B provided in the Monolith Protein Labeling Kit (Nanotemper). Sixteen 1:1 serial dilutions of the binding ligand GST-CDK4 were prepared using MST optimization buffer (20 mM Tris, 200 mM NaCl, 0.05% Tween 20, pH 8.0), while maintaining a constant target concentration of 70 nM. For MBP-cycP16p and MBP-P16p, the highest concentrations of GST-CDK4 used were 1.92 μM and 3.84 μM, respectively. Samples were mixed and loaded into Monolith standard capillaries (NanoTemper) by pipetting, and their binding affinity was measured using a Monolith NT.115 with Nano-RED excitation type, with the MST power set to medium. The collected data were processed using NanoTemper Affinity analysis software to determine the dissociation constant (Kd).

[0198] MCF-7 cell lysate pull-down assay. MCF-7 cells were cultured at 4.0 × 10⁻⁶ 5Cells were seeded per well in 6-well plates and cultured for 20 h in RPMI-1640 medium supplemented with 10% FBS and penicillin / streptomycin (P / S). Cells were washed twice with ice-cold PBS and lysed in NP-40 lysis buffer (50 mM Tris, 150 mM NaCl, 2 mM EDTA, 1% NP-40, 0.1% SDS, pH 7.5) supplemented with a protease inhibitor mixture (Roche), followed by incubation at low temperature with gentle shaking for 20 min. Cell lysates were clarified by centrifugation at 15000 rpm at 4°C for 15 min. The total protein concentration of the supernatant was measured using the Pierce™ BCA protein assay kit (Thermo Fisher Scientific) and diluted to 1.5 mg / mL with PBS. The bait protein solution containing CDK4 / 6 was then gently mixed with purified MBPcycP16p immobilized on amylose resin and incubated at 4°C for 3 h. The reaction mixture was washed five times with PBS, and then the resin was analyzed by Western blotting using anti-CDK4 (1:500, Biolegend) according to standard protocol. The same bait protein solution was incubated with amylose resin as a negative control only.

[0199] Cellular uptake assay. MCF-7 cells in RPMI-1640 medium supplemented with 10% FBS and P / S (Cytiva) were cultured at 2.0 × 10⁻⁶. 5 Cells were seeded at 1000 cells / mL in 35 mm glass-bottomed confocal culture dishes (MatTek) and cultured overnight at 37°C / 5% CO2 for adhesion. The next day, cells were washed with PBS to remove non-adherent cells, and adherent cells were starved for 24 h in serum-free RPMI-1640 medium to induce cell cycle synchronization. Cells were then treated for 24 h in RPMI-1640 medium supplemented with 10% FBS and P / S with 15 μM cycRGD-mCh-P16p or cycRGD-mCh-cycP16p. Prior to confocal imaging, cell nuclei and plasma membranes were counterstained with Hoechst 33342 (Thermo Fisher) and wheat germ lectin (Thermo Fisher), respectively. Images were captured using a TCS SP8 confocal microscope (Leica) with excitation wavelengths set at 488 nm, 561 nm, and 633 nm.

[0200] Cell cycle arrest assay. MCF-7 cells were seeded in 12-well plates (1.5 × 10⁻⁶ cells per well). 5Cells were cultured in RPMI-1640 medium supplemented with 10% FBS and P / S at 37°C in wells (cells / well). After 24 hours of adhesion, the cells were starved for another 24 hours by replacing the medium with serum-free RPMI-1640 medium. At the end of the starvation period, the medium was replaced with RPMI-1640 supplemented with 10% FBS. Cells were then treated for different durations (24 h, 48 h, 72 h) with 15 μM R9-P16p, 15 μM cycRGD-mCh-P16p or 15 μM cycRGD-mCh-cycP16p, PBS (negative control), and 10 nM actinomycin (positive control). At the end of the treatment period, cells were washed twice with PBS, digested with trypsin, and fixed on ice with 70% ethanol for 45 min. Fixed cells were washed with PBS and stained with PI staining solution (10 μg / mL PI in PBS supplemented with RNase) for 30 minutes in the dark at room temperature, followed by flow cytometry analysis using a FACSVerse flow cytometer (BD Biosciences). Flow cytometry data were analyzed using a ModFit LT 5.0 (BD Biosciences).

[0201] Cell proliferation assay. MCF-7 in RPMI-1640 medium supplemented with 10% FBS and P / S was incubated at 1.6 × 10⁻⁶. 4 Cells were seeded in 96-well plates to allow overnight adhesion. Cells were then washed with PBS and treated with 5 μM to 20 μM cycRGD-mCh-P16p, cycRGD-mCh-cycP16p, or R9-P16p (Pepmic) for 24 hours. The number of viable cells was determined using the CellTiter96 Cell Proliferation Assay Kit (Promega) according to the manufacturer's instructions and scanned on a TecanSpark 10M microplate reader with absorbance set to 490 nm. Data were normalized relative to values ​​from PBS-treated control cells. The experiment was repeated three times, yielding similar results.

[0202] Western blot analysis. With 4.0×10 5MCF7 cells were seeded in 6-well plates in RPMI-1640 medium supplemented with 10% FBS and P / S for overnight adhesion. The medium was then replaced with serum-free RPMI-1640 and the cells were starved for 24 hours. At the end of starvation, the serum-free medium was replaced with RPMI-1640 supplemented with 10% FBS and P / S. 15 μM of cycRGD-mCh-P16p or cycRGD-mCh-cycP16p was added to the cells, and the cells were incubated at 37°C / 5% CO2 for 24 hours. At the end of treatment, the cells were washed with ice-cold PBS and lysed in NP-40 lysis buffer (50 mM Tris, 150 mM NaCl, 2 mM EDTA, 1% NP-40, 0.1% SDS, pH 7.5) supplemented with a protease inhibitor mixture (Roche). Cell debris was removed by centrifugation at 15,000 rpm at 4°C for 10 minutes. Samples were separated by 10% SDS-PAGE, transferred to nitrocellulose membranes, and probed using pRb (Ser780) antibody (1:1000, Cell Signaling), pRb (Ser795) antibody (1:500, Cell Signaling), and β-actin antibody (1:2500, Sigma-Aldrich). Protein visualization was performed using an ECL system (Amersham).

[0203] Statistical analysis. Data from repeated experiments are expressed as mean ± standard error of the mean (SD). Statistical significance is indicated appropriately in the legend. To compare data in different groups, ordinary one-way ANOVA and Tukey's multiple comparison test were performed using GraphPad Prism 7 software (GraphPad, San Diego, USA). A p-value less than 0.05 was considered statistically significant. p < 0.05, p < 0.001, p < 0.0001, ns, not significant.

[0204] Example 2 - Optimization of the D-Cys-ε-Lys Readout System

[0205] This embodiment relates to optimizing the UAG read-through system to enhance D-Cys-ε-Lys incorporation into proteins. D-Cys-ε-Lys is not a natural substrate of wild-type pyrrololysyl-tRNA synthetase (PylRS). Improvements in the substrate specificity or activity of PylRS for D-Cys-ε-Lys can enhance the catalytic efficiency of PylRS, leading to more efficient D-Cys-ε-Lys incorporation. Therefore, directed evolution was employed to evolve PylRS mutants that specifically recognize D-Cys-ε-Lys and exhibit improved efficiency in its genetic incorporation.

[0206] To generate a PylRS mutant library, a library containing tRNA was constructed. M15 The plasmid encoding the wild-type PylRS gene. tRNA M15 It is tRNA pyl A more stable variant, which has been shown to have improved amber inhibition efficiency. 23-24 The coding sequences for the kanamycin resistance protein aphA-3 and the mCherry protein were determined. Both AphA-3 and mCherry were engineered with permitted UAG codons in their coding sequences to facilitate a two-stage selection and screening protocol for rapid detection of active PylRS mutants and their corresponding D-Cys-ε-Lys incorporation activity. First, positive selection of mutants was performed under kanamycin selection pressure, depending on D-Cys-ε-Lys incorporation into the aphA-3 kanamycin resistance peptide. Then, the selected clones were screened for mCherry fluorescence, achieved through read-through of the UAG codons in the mCherry transcript. After iterative mutagenesis and screening, triple mutant PylRS carrying G14E, C348V, and S451F mutations (PylRS) were identified. EVF — PylRS ( ) was identified as exhibiting improved readability. Figure 1D (Left image).

[0207] Computational modeling was then performed to determine how these three mutations affect substrate binding. It has been reported that, among the three mutated residues, C348 is crucial for substrate binding. 25 Previous studies on the structure of MmPylRS have revealed that C348, together with Y306, Y384, V401, and W417, forms a deep hydrophobic pocket and is one of the key residues involved in the recognition and binding of Pyl and its analogues. 26-27CHIMERA analysis confirmed that replacing the polar cysteine ​​residue with the more hydrophobic valine residue likely enhances the hydrophobicity of the binding pocket, thereby increasing its affinity for D-Cys-ε-Lys. Two other mutant residues, G14 and S451, are located at the N-terminal tRNA-binding domain and C-terminal tRNA minimal core binding surface of PylRS, respectively. Subsequent optimization of other parameters, including replacing the T7 promoter with the T7lac promoter, resulted in further yield improvements compared to using WT PylRS and tRNA. pyl Compared to readout, the net increase in mCherry fluorescence was approximately 20-fold.

[0208] To further improve the yield of proteins containing D-Cys-ε-Lys, an E. coli strain (C321.ΔA.M9 adapted) was used. 29 In this strain, release factor 1 was missing. Based on mCherry fluorescence, the use of this strain resulted in a further 5.3-fold improvement (…). Figure 1E ).like Figure 1E As shown, the improved strains produced higher levels of full-length proteins of various sizes. These results are consistent with enhanced D-Cys-ε-Lys incorporation and UAG readthrough efficiency.

[0209] Example 3: Intracellular synthesis of D-Cys-ε-Lys

[0210] Most current research on the production of ncAA-containing proteins using non-classical amino acids (ncAAs) relies on the exogenous supplementation of chemically synthesized ncAAs with the desired function. However, this approach is cost-inefficient and unsuitable for large-scale production of ncAA-containing proteins. A more economical and practical strategy is to produce the desired ncAA in vivo (e.g., in E. coli) using a basic carbon source or amino acid (which can be purchased inexpensively), and then directly incorporate these ncAAs into the protein using UAG codons.

[0211] In this regard, the elucidated Pyl biosynthetic pathway 30 This provides a blueprint for establishing the endogenous biosynthetic mechanism of ncAAs (e.g., D-Cys-ε-Lys). Pyl is generated by the enzymes PylB, PylC, and PylD. In the first step, lysine is converted to (2R,3R)-3-methylornithine by PylB. In the second step, (2R,3R)-3-methylornithine is coupled to another lysine via PylC to form -Lysine-Nε-3R-methyl-D-ornithine (Lys-Nε-3MO). Lys-Nε-3MO is a precursor of pyl. In the third step, PylD converts the terminal ornithine amide to a carbonyl group, which then reacts with an amine to form a pyrrololine group.

[0212] Previous studies have shown that Pyl can be synthesized in the absence of PylB. 30 Based on this, PylC was modified to use other amino acids (e.g., D-cysteine ​​(D-Cys)) as substrates, and -Lysine to form D-Cys-ε-Lys( Figure 2A The complex structure of PylC and D-ornithine in *Methanococcus pasteurellis* was previously determined. 31 The study revealed four active residues: S177, E179, D233, and T256, which are important for D-ornithine binding. Therefore, to evolve a PylC mutant capable of using D-Cys as a substrate, site-directed saturation mutagenesis was performed on S177, E179, D233, and T256, resulting in a mutant library. The identified PylC variants are shown in Table 1 above. This mutant library was used to transform PylRS from Example 2 above. EVF / tRNA M15 Optimized strains. Mutant PylC strains containing the S177N, E179P, D233S, and T256V mutations were identified by kanamycin selection and mCherry fluorescence screening. NPSV ( Figure 2B ).

[0213] Application of computational model to analyze MmPylC WT The substrate-specific conversion basis was determined, and MmPylC bound to D-Cys-ε-Lys was generated using the CNS program. NPSV 32-33 Based on these models, the mutations in E179P and T256V appear to transform the original hydrophilic D-ornithine binding site into a hydrophobic pocket, making the binding of the charged amino group on the D-ornithine side chain less favorable. On the other hand, based on analysis using CHIMERA, PylC NPSV The mutations T256V and S177N in D-Cys-ε-Lys lead to van der Waals interactions with the thiol group of D-Cys-ε-Lys. This combination of the weaker binding of D-ornithine-ε-Lys and the improved binding of D-Cys-ε-Lys may result in the observed substrate-specific switching.

[0214] To evaluate the readthrough efficiency of intracellularly biosynthesized D-Cys-ε-Lys and compare it with that of exogenously added chemically synthesized counterparts, different protein constructs were tested. Results showed that at D-Cys concentrations of 5 mM or higher, the optimized PylRS... EVF / tRNA M15 PylC in the strain NPSV It can achieve readability comparable to that of chemically synthesized D-Cys-ε-Lys with exogenous addition of 4 mM. Figure 2C Subsequently, LC-MS / MS analysis of the corresponding trypsin-digested proteins was performed to identify peptide fragments containing D-Cys-ε-Lys, confirming the successful intracellular biosynthesis of D-Cys-ε-Lys.

[0215] Example 4 - Cycloning of the structural effects of P16p

[0216] This embodiment relates to the development of a protein cyclization method using D-Cys-ε-Lys. Previously, slow and incomplete cyclization of the RGD motif was observed when it was placed at the C-terminus of mCherry fusion proteins. 10 At this point, it was speculated that the low yield was due to the high degree of freedom of the reactive groups. Further supporting this observation were insights from other studies, namely that the relative spatial positions of the N-terminus and C-terminus of the peptide to be cyclized significantly influence the cyclization rate and yield. 34-36 Therefore, it is hypothesized that cyclization using structurally better-positioned D-Cys-ε-Lys and C-terminal thioester fragments could lead to enhanced cyclization efficiency.

[0217] The tumor suppressor protein P16 was chosen for cyclization using D-Cys-ε-Lys because the P16 protein has a pre-formed circular motif within its native sequence; the portion of P16 that binds to CDK4 / 6 forms a helical-turn-helical structure. 19 P16 acts as a CDK4 / 6 inhibitor by preventing phosphorylation in retinoblastoma (Rb), thereby leading to G0 / G1 cell cycle arrest. The P16-binding sequence to CDK4 / 6 is a 20-demerit peptide containing residues 84-103 (P16p). 18 This embodiment involves determining whether the cyclization of P16p enhances its binding affinity and cell stability, thereby increasing its therapeutic potential.

[0218] The design of the D-Cys-ε-Lys incorporation site was guided by the structure of the P16 / CDK6 protein complex. Molecular simulations were performed to validate the design strategy. Based on structural analysis, an incorporation site was selected at the N-terminus of P16p, which was extended by one or more residues to include the 83rd histidine residue. Furthermore, an N-terminal glycine was added to spatially approximate the D-Cys-ε-Lys to the C-terminal thioester formed during cyclization. Figures 3A to 3B ).

[0219] To gain a visual understanding of the designed cyclic P16 peptide, a corresponding molecular model was built using Pymol, which showed that the D-Cys-ε-Lys-linkage occurs on opposite sides of the P16-CDK4 / 6 interface. Figure 3B Then, molecular dynamics simulations were performed on the peptide and its linear counterpart to evaluate their structural stability.

[0220] In the simulations, linear P16p (LinP16p) exhibited significantly higher structural flexibility, particularly at the ends of the two helices, while cyclic P16p (cycP16p) showed a more stable conformation with minimal variation compared to the reference structure. Figure 3C This observation was confirmed by their respective RMSD and RMSF distribution plots, which depicted that the amplitude fluctuations of linear P16p were much larger compared to cycP16p. Figures 3D to 3E These results appear to support the view that cyclization would benefit P16p, at least in the case of reduced configurational entropy, by enhancing its stability.

[0221] To experimentally verify these results, the aforementioned P16p design was incorporated into a fusion protein (GFP-O-P16p-intein-CBD-His7) containing green fluorescent protein (GFP), P16p, an inteptide for thioester production, and a CBD-His7 tag for affinity purification. Figure 4A ). Utilizing optimized PylRS EVF / tRNA M15 PylC for engineering NPSV (i.e., the PylC mutant) yielded approximately 19 mg of purified GFP-O-P16p-in-peptide-CBD-His7 protein containing D-Cys-ε-Lys from 200 mL of cell culture grown in medium supplemented with 5 mM D-Cys. This represented a 10-fold increase in yield compared to the unoptimized system. Cycloning of the fusion protein was initiated by the addition of MESNA. As indicated by the flattening of the GFP signal after 2.5 hours, cycloning was completed within 3 hours. Figure 4B Subsequent mass spectrometry (MS) analysis confirmed that the obtained cyclized product was a cyclized form of GFP-P16p with a single isotopic mass of 29,643.77 (theoretical mass of 29,642.72) (GFP-cycP16p). Figure 4C Furthermore, MS analysis showed that GFP-cycP16p was the major component in the cyclization reaction mixture. This contrasts with previous studies. 10 Compared to the cyclization of pyl, which takes >96 h and achieves only 58% of the RGD effect, the simplified cyclization of P16p within 3 h approaches quantitative levels—significantly improving efficiency and completion. Therefore, these results confirm the hypothesis that pre-formed structural motifs can facilitate simplified cyclization.

[0222] Example 5 - In vitro binding study of cyclized P16p and CDK4

[0223] To evaluate the effect of cyclization on the binding affinity of P16p to CDK4, analytical size exclusion chromatography (SEC) and microthermophoresis (MST) were performed. SEC analysis showed that GFP-cycP16p interacted with GST-CDK4 to form a complex. In the MST binding assay, the GFP tag of GFP-cycP16p was replaced with an MBP tag to avoid fluorescence interference. Protein pull-down assays confirmed that the resulting MBP-P16p could still capture CDK4 from MCF-7 cell lysates. Subsequent MST binding studies showed that the cyclized MBP-P16p protein (MBP-cycP16p) (Kd = 27.7 ± 9.4 nM; Figure 5 It exhibits similarity to its linear counterpart, MBP-P16p (Kd = 117.1 ± 41.4 nM). Figure 5 Compared to GST-CDK4, it has a binding affinity that is 4 times higher.

[0224] Example 6 - Generation and in vitro characterization of bifunctional bicyclic CycRGD-mCh-CycP16p

[0225] This embodiment relates to further improving the therapeutic potential of cyclic P16 by providing a circular RGD motif for cancer cell targeting to cyclic P16 (cycP16p). Imagine a dumbbell-shaped protein having (1) a circular RGD at one end that acts as a tumor homing device that specifically binds to the αvβ3 integrin receptor, which is typically overexpressed on cancer cells, and (2) a cyclic P16 at the other end that acts on an intracellular target to elicit an anticancer response.

[0226] To this end, a fusion protein cycRGD-mCherry-OP16p-inteptide-CBD-His7 was designed. Figure 6 The protein contains (1) an N-terminal disulfide-based cyclic RGD motif 37 (cycRGD) to enhance cell specificity and uptake via integrase binding, and (2) mCherry to facilitate improved confocal imaging. 38(3) P16p as described in the preceding examples, and (4) the inteptide-CBD-His7 tag. CycRGD-mCherry-OP16p-inteptide-CBD-His7 was expressed, purified, and cyclized according to a similar procedure to that described in the preceding examples for GFP-cycP16p and MBP-cycP16p. To achieve N-terminal RGD cyclization, the protein was air-oxidized for 24 hours to produce bicyclic cycRGD-mCh-cycP16p. Cellular uptake and distribution of cycRGD-mCherry carrying linear uncyclized P16p (cycRGD-mCh-P16p) and cyclized P16p (cycRGD-mCh-cycP16p) were then evaluated by incubating various proteins with MCF-7 breast cancer cells for 20 hours. Significant mCherry fluorescence was detected intracellularly in both treatments, as shown by cross-sectional scanning using a confocal laser scanning microscope. These results indicate that both linear and cyclic cycRGD-mCh-P16p can enter cells.

[0227] Example 7 - Effect of Cyclation on the Inhibition and Stability of p16 Peptide

[0228] This embodiment relates to determining whether P16p cyclization can reliably confer cell cycle arrest on cycRGD-mCh-cycP16p and produce a more effective inhibitory effect on cancer cell growth. MCF-7 cells were treated with different concentrations of linear or cyclic cycRGD-mCherry-P16p. Linear P16p fused to the well-known polyarginine cell-penetrating sequence (R9-P16p) was also included for comparison. Flow cytometry analysis of the treated cells from different groups showed that at a concentration of 15 μM, cycRGD-mCh-cycP16p was the most effective cell cycle arrestor (76.71%). Figure 7A (Bottom, left image), and has advantages over R9-P16p, which has moderate performance (60.71%). Figure 7A (Top, right figure), while the linear counterpart cycRGD-mCh-P16p had the least effect on G0 / G1 phase retardation (46.66%). Figure 7A (Top, middle image). Similar results were observed in cell proliferation assays, where, at the tested concentration, compared to R9-P16p (… Figure 7A (Top, right image) and cycRGD-mCh-P16p ( Figure 7A Compared to (top, middle image), cycRGD-mCh-cycP16p ( Figure 7A(Bottom, left figure) exhibits the strongest inhibition of MCF-7 growth. These enhanced properties of cycRGD-mCh-cycP16p can be partly attributed to its higher affinity binding to CDK4 in MCF-7 cells induced by the cycP16 subunit, as indicated by the binding constant obtained from MST, as discussed above in Example 5 and... Figure 5 As indicated in the document.

[0229] The stability of cycRGD-mCh-P16p was evaluated by monitoring its ability to induce cell cycle arrest over a prolonged period of up to 72 hours. This time-series study showed that cycRGD-mCh-cycP16p maintained its ability to induce G0 / G1 cell cycle arrest even after prolonged exposure to serum-supplemented medium, as evidenced by approximately 88% of cells being arrested at 72 h. Figure 7C By comparison, linear cycRGDmCh-P16p failed to induce meaningful cell cycle arrest. These results were further confirmed by querying the phosphorylation status of specific sites on retinoblastoma (Rb), which are known to be associated with the G1-S phase transition of the cell cycle in differentially treated MCF cells. Western blot analysis using phosphate-specific anti-Rb antibodies against serine 780 and serine 795 revealed significantly reduced phosphorylation levels at the specified sites in MCF-7 cells treated with cyclic cycRGD-mCh-cycP16p compared to those treated with the other two linear P16p formulations. Figure 7D Serine 780 and serine 795 are two residues whose phosphorylation activates cell cycle progression and arrests cell cycle arrest in pBR, respectively. Overall, the results indicate that cyclization enhances the potency of p16p as a CDK4 / 6 inhibitor, partly due to enhanced serum stability and binding affinity to CDK4. Notably, the use of D-Cys-ε-Lys allows for the generation of isopeptide-linked cyclic p16p subunits, which are more resistant to protein hydrolysis and stable in reduced cytosol—a crucial feature essential for peptide therapeutic development.

[0230] Summarize

[0231] These examples demonstrate an efficient and robust platform developed based on Pyl technology for the production of cyclized proteins. The system differs in that the ncAA substrate is synthesized intracellularly from low-cost and commercially available simple amino acids, significantly reducing the production cost of the target protein—a necessary attribute in the development of therapeutic products. By optimizing the readthrough system and implementing a structure-inspired macrocyclization strategy, proteins carrying ncAA can be produced in high yields and cyclization can be efficiently brought to near completion. Using this optimized system, the P16 peptide subunit of the fusion protein was successfully cyclized within 3 h under mild and reducing conditions. Compared to the linear counterpart, the cyclized P16 p subunit exhibited enhanced binding affinity to CDK4 and more effective inhibition of pRb phosphorylation, resulting in meaningful G0 / G1 arrest in MCF-7 cells. Considering the important role of cyclic proteins / peptides in drug development, our incorporation system and cyclization strategy could potentially be applied to cyclize other potential therapeutic candidates, thus providing different and complementary approaches to the preparation of proteins containing cyclic peptides.

[0232] II. Synthesis and Use of D-Pra-ε-Lys and D-Allyl-ε-Lys (Overview of Examples 8-9)

[0233] Intracellular synthesis of ncAAs D-Pra-ε-Lys and D-allyl-ε-L-Lys is achieved by hijacking the Pyl biosynthetic pathway. A PylC library was generated using site-directed saturation mutagenesis to produce suitable PylC variants capable of producing D-Pra-ε-Lys and D-allyl-ε-L-Lys in cells. D-Pra-ε-Lys contains an alkyne amino acid and is available for click chemistry reactions, etc. D-allyl-ε-L-Lys contains a reactive vinyl stalk that can be used for Michael addition reactions of thiols.

[0234] Example 8 - Intracellular synthesis of D-Pra-ε-Lys

[0235] This embodiment relates to the synthesis of D-Pra-ε-Lys using a PylC variant. The Pyl analog D-Pra-ε-Lys is biosynthesized from the starting chemical (R)-2-aminopentan-4-acetylic acid and endogenous L-lysine. Figure 8A Site-directed saturation mutagenesis was performed on *Methanococcus pasteurellii* PylC, and a variant of PylC capable of producing D-Pra-ε-Lys was selected, as discussed in Examples 1 and 3 above. After selection, 31 colonies survived, and each of these colonies was grown in 250 μL of culture and inoculated with ITPG. Their mCherry fluorescence intensity was measured, as shown in... Figure 8BAs shown in Table 2 above, 28 PylC mutants exhibiting the strongest mCherry fluorescence signal were sent for Sanger sequencing. The identified PylC variants are shown in Table 2 above. Sequence analysis revealed that the four mutated residues in PylC—S177, E179, D233, and T256—prefer hydrophobicity, with the D233H mutation being highly conserved.

[0236] Next, use the PylC variant PylC GVNV D-Pra-ε-Lys is incorporated into the protein at the UAG codon. PylC GVNV With PylS WT -PylT M15 The -GFP-UAG-P16p-in-peptide-CBD-His7 expression vector was co-expressed in the presence of 5 mM propargyl chemical supplement. Recombinant protein expression was induced overnight at 25°C. Expression was achieved via Ni... 2+ GFP-UAG-P16p protein was purified by affinity chromatography, and the GFP-positive fraction was collected. The GFP-UAG-P16p protein was treated with DTT and trypsin. D-Pra-ε-Lys was identified in the UAG-intercalated GFP protein by ESI-MS / MS analysis. The results demonstrate the success of intracellular synthesis of D-Pra-ε-Lys using the PylC variant.

[0237] Example 9 - Intracellular synthesis of D-allyl-ε-Lys

[0238] This embodiment relates to the synthesis of D-allyl-ε-Lys using a PylC variant. Pyl analogues of D-allyl-ε-Lys are biosynthesized from the starting compound D-allyl-OH and endogenous L-lysine. Figure 9A Site-directed saturation mutagenesis was performed on *Methanococcus pasteurellii* PylC, and variants of PylC capable of producing D-allyl-ε-Lys, as discussed in Examples 1, 3, and 8 above, were identified using antibiotic selection and mCherry fluorescence readout assays. Figure 9B The top 10 PylC variants with the highest readability and possibly the highest level of D-allyl-ε-Lys biosynthesis were sequenced, and site-specific mutations were identified, as shown in Table 3 above.

[0239] Next, use the PylC variant PylC GVNV D-allyl-ε-Lys was incorporated into the protein at the UAG codon. For this purpose, target cells were transformed using two vectors: constitutively expressing PylC... GVNV The carrier and PylRS WT -PylT M15-GFP-UAG-P16p-inner peptide-CBD-His7 expression vector. In the presence of D-allyl-OH, the initial growth phase allows PylC... GVNV The synthesis of D-allyl-ε-Lys was mediated, followed by induction of GFP-UAG-P16p protein expression with isopropyl-β-d-1-thiogalactopyranoside (IPTG), and cells were allowed to culture overnight at 25°C. This was achieved via Ni... 2+ GFP-UAG-P16p protein was purified by affinity chromatography, and the GFP-positive fraction was collected. The GFP-UAG-P16p protein was treated with DTT and trypsin. D-allyl-ε-Lys was identified in the UAG-intercalated GFP protein by ESI-MS / MS analysis. These results demonstrate the successful intracellular synthesis of D-allyl-ε-Lys using the PylC variant.

[0240] III. Synthesis and Use of 2-Chloroacetyl-ε-Lys (Overview of Examples 10 to 12)

[0241] Intracellular synthesis of ncAA 2-chloroacetyl-ε-Lys(2ClAcK); Figure 10 This is achieved by hijacking the Pyl biosynthetic pathway. Using directed evolution experiments, PylRS / tRNAs suitable for incorporating 2ClAcK into proteins were identified. pyl Yes. A unique feature of 2ClAcK is that the reactivity of its 2-chloroacetyl group with cysteine ​​is modulated, such that coupling requires the two moieties to be in an extended local proximity state for the reaction to occur. This requirement is advantageous because it reduces off-target reactions and allows selective coupling only after translational incorporation and when 2ClAcK is spatially located near the target cysteine ​​within the target protein.

[0242] Using tRNA (UAG) PylT ) and optimized pyrrololysyl tRNA synthetase (PylS (e.g., wild-type archaea or bacteria PylT, or PylT) M15 Mutants (such as Serfling, Robert) et al. Nucleic Acids Research vol. 46,1 (2018): 1-10 (discussed), and PylS Mutants (such as Kobayashi) et al. Journal of the American Chemical Society (Discussed in 2016 138 (45), 14832-14835)). 2ClAcK was successfully incorporated into the His-SUMOmini2 peptide model system derived from SUMO1, which can be used as an inhibitor of α-synuclein aggregation and for the treatment of Parkinson's disease (Liang). et al. Cell Chem Biol , 2021, 28, 180).

[0243] In the case of the SUMO1 peptide, the cyclic His-SUMOmini2 peptide is generated from the His-SUMOmini2-LVPRGSSUMOCt-MBP construct, which contains a core SUMO1 protein domain with an internal thrombin-cleaving sequence LVPRGS and a C-terminal maltose-binding protein (MBP) purification tag. The core SUMO1 domain acts as a scaffold, positioning the 2ClAcK and Cys residues in appropriate chemical spatial orientation. The construct is cleaved with thrombin, followed by nickel affinity chromatography, allowing the N-terminal cyclic His-SUMOmini2 peptide to separate from the remaining SUMO1 C-terminal fragment (SUMOCt) and the MBP domain. Figures 11A to 11C The formation of a 2ClAcK-Cys crosslink in the cyclic His-SUMOmini2 peptide was confirmed by mass spectrometry. The generation of the bicyclic His-SUMOmini2 peptide was also tested utilizing the spatial specificity of the 2ClAcK-Cys coupling reaction. In His-SUMOmini2, two UAG sites were engineered for 2ClAcK incorporation, and two cysteine ​​residues were positioned such that each cysteine ​​is spatially close to 2ClAcK for easy crosslinking. The benefits of the bicyclic SUMO1 peptide include enhanced resistance to protease degradation and improved inhibition of α-synuclein aggregation. Figure 12 and Figure 13 ).

[0244] Example 10 - Intracellular synthesis of 2-chloroacetyl-ε-Lys

[0245] This example relates to the synthesis using a PylC variant. The Pyl analog 2-chloroacetyl-ε-LyS (2ClAcK) is biosynthesized from 2-chloroacetic acid and endogenous lysine. Figure 10 Site-directed saturation mutagenesis was performed on PylC of *Methanococcus pasteurellii*, and PylC variants capable of producing 2ClAcK were selected, as described in Examples 1, 3, 8, and 9 above. PylC variants were analyzed by sequencing.

[0246] Next, using the identified PylC variant, 2ClAcK was incorporated into the protein at the UAG codon. The PylC variant is similar to PylS... T-mCherry(UAG)-kanR(UAG) reporter plasmid co-expression, the PylS The T-mCherry(UAG)-kanR(UAG) reporter plasmid contains an optimized pyrrolidone-lysyl-tRNA synthetase (PylS). This process allows pyrrololysyl-tRNA (UAG) (PylT) to be loaded with 2ClAcK for translational incorporation. PylC mutants capable of producing 2ClAcK were first identified by screening co-transformed cells on high concentrations of kanamycin in the presence of IPTG, lysine, and 2-chloroacetic acid. Since the expression of the KanR gene requires the presence of 2ClAcK, only cells with PylC mutants capable of producing 2ClAcK survived. After the initial round of kanamycin selection, the readthrough efficiency of the identified cells was evaluated by culturing them on a small scale (in 96-well plates) in the presence of IPTG, lysine, and 2-chloroacetic acid and measuring the relative levels of mCherry fluorescence. Following the identification of the initial mutants, further evolution and selection were conducted to identify PylC mutants with better readthrough capabilities for improved 2ClAcK biosynthesis. The identified PylC mutants were overexpressed, purified, and characterized, for example, by using protein crystallography to determine their three-dimensional structures, in order to understand how to modify the substrate specificity of PylC to catalyze reactions using lysine and 2-chloroacetic acid as substrates (for 2ClAck) instead of D-ornithine (for Pyl).

[0247] Example 11 - Generation of cyclic SUMO-derived peptides

[0248] This embodiment relates to the generation of cyclic SUMO-derived peptides by incorporating 2ClAcK. It has previously been demonstrated that the early linear form of the SUMO1 peptide (15-55) (SUMOmini2) can effectively inhibit α-synuclein aggregation (Liang) in vitro and in vivo. et al. Cell Chem Biol (2021, 28, 180). To generate both monocyclic and bicyclic forms of the His-SUMOmini2 peptide, synthetic 2ClAcK was incorporated into the His-SUMOmini2-LVPRGS-SUMOCt-MBP construct during protein translation. Alternatively, PylS... T-His-SUMOmini2-LVPRGS-SUMOCt-MBP and a plasmid were co-transformed into *E. coli*, the plasmid constitutively expressing the optimized PylC mutant identified in Example 10 above. *E. coli* were cultured in a medium supplemented with 2-chloroacetic acid to enable intracellular synthesis of 2ClAcK produced from the mutant PylC, which could then be directly incorporated into the optimized tRNA / PylS... In the nascent polypeptide of His-SUMOmini2-LVPRGS-SUMOCt-MBP protein mediated by SUMOmini2-LVPRGS-SUMOCt-MBP protein.

[0249] The purified SUMOmini2-LVPRGS-SUMOCt-MBP construct was incubated with thrombin and then purified by nickel affinity chromatography. Figures 11A to 11C It is worth noting that previous attempts to produce disulfide-bonded stable His-SUMOmini2 peptides using chemical synthesis have been unsuccessful. The ability to generate both monocyclic and bicyclic His-SUMOmini2 peptides, as discussed in this embodiment, mitigates this barrier to chemical synthesis. This is based on the enhanced resistance of the bicyclic His-SUMOmini2 peptide to proteolytic degradation. Figure 12 And the enhanced ability to inhibit α-common nucleoprotein aggregation ( Figure 13 Bicyclic His-SUMOmini2 peptide appears to be the most effective. The strategy of cyclizing SUMO-derived peptides shown in this paper can be used to inhibit α-synuclein aggregation in Parkinson's disease.

[0250] References

[0251] The references below refer to Examples 1 to 7 above.

[0252] 1. Sato, AK; Viswanathan, M.; Kent, RB; Wood, CR. Therapeutic peptides: technological advances driving peptides into development. Curr. Opin. Biotechnol. 2006, 17 (6), 638-642.

[0253] 2. Colgrave, ML; Craik, DJ, Thermal, chemical, and enzymaticstability of the cyclotide kalata B1: the importance of the cyclic cystineknot. Biochemistry 2004, 43 (20), 5965-5975.

[0254] 3. Ji, Y.; Majumder, S.; Millard, M.; Borra, R.; Bi, T.; Elnagar, AY; Neamati, N.; Shekhtman, A.; Camarero, JA, In vivoactivation of thep53 tumor suppressor pathway by an engineered cyclotide. J. Am. Chem. Soc.2013, 135 (31), 11623-11633.

[0255] 4. Ngo, K. H.; Yang, R.; Das, P.; Nguyen, G. K.; Lim, K. W.; Tam, J.P.; Wu, B.; Phan, A. T., Cyclization of a G4-specific peptide enhances itsstability and G-quadruplex binding affinity. Chem. Commun. 2020, 56 (7),1082-1084.

[0256] 5. Wilbs, J.; Kong, X.-D.; Middendorp, S. J.; Prince, R.; Cooke, A.;Demarest, C. T.; Abdelhafez, M. M.; Roberts, K.; Umei, N.; Gonschorek, P.,Cyclic peptide FXII inhibitor provides safe anticoagulation in a thrombosismodel and in artificial lungs. Nat. Commun. 2020, 11 (1), 1-13.

[0257] 6. Clardy, J.; Walsh, C., Lessons from natural molecules. Nature2004, 432 (7019), 829-837.

[0258] 7. Driggers, E. M.; Hale, S. P.; Lee, J.; Terrett, N. K., The exploration of macrocycles for drug discovery—an underexploited structural class. Nat. Rev. Drug Discovery 2008, 7 (7), 608-624.

[0259] 8. Deyle, K.; Kong, X. D.; Heinis, C., Phage selection of cyclic peptides for application in research and drug development. Acc. Chem. Res. 2017, 50 (8), 1866-1874.

[0260] 9. Tavassoli, A.; Benkovic, S. J., Split-intein mediated circular ligation used in the synthesis of cyclic peptide libraries in E. coli. Nat. Protoc. 2007, 2 (5), 1126-33.

[0261] 10. Lee, M. M.; Fekner, T.; Lu, J.; Heater, B. S.; Behrman, E. J.; Zhang, L.; Hsu, P. H.; Chan, M. K., Pyrrolysine-inspired protein cyclization. Chembiochem 2014, 15 (12), 1769-1772.

[0262] 11. Bouclier, C.; Simon, M.; Laconde, G.; Pellerano, M.; Diot, S.;Lantuejoul, S.; Busser, B.; Vanwonterghem, L.; Vollaire, J.; Josserand, V.,Stapled peptide targeting the CDK4 / cyclin D interface combined withAbemaciclib inhibits KRAS mutant lung cancer growth. Theranostics 2020, 10(5), 2008.

[0263] 12. Abdelkader, E. H.; Qianzhu, H.; George, J.; Frkic, R. L.;Jackson, C. J.; Nitsche, C.; Otting, G.; Huber, T., Genetic encoding ofcyanopyridylalanine for in-cell protein macrocyclization by thenitrileaminothiol click reaction. Angew. Chem., Int. Ed. 2022, 61 (13),e202114154.

[0264] 13. Dong, H.; Li, J.; Liu, H.; Lu, S.; Wu, J.; Zhang, Y.; Yin, Y.;Zhao, Y.; Wu, C., Design and ribosomal incorporation of noncanonicaldisulfide-directing motifs for the development of multicyclic peptidelibraries. J. Am. Chem. Soc. 2022, 144 (11), 5116-5125.

[0265] 14. Dawson, P. E.; Muir, T. W.; Clark-Lewis, I.; Kent, S. B.,Synthesis of proteins by native chemical ligation. Science 1994, 266 (5186),776-779.

[0266] 15. Li, X.; Fekner, T.; Ottesen, J. J.; Chan, M. K., A pyrrolysineanalogue for site-specific protein ubiquitination. Angew. Chem. 2009, 121(48), 9348-9351.

[0267] 16. Serrano, M.; Hannon, G. J.; Beach, D., A new regulatory motif incell-cycle control causing specific inhibition of cyclin D / CDK4. Nature 1993,366 (6456), 704-707.

[0268] 17. Finn, R. S.; Martin, M.; Rugo, H. S.; Jones, S.; Im, S. A.;Gelmon, K.; Harbeck, N.; Lipatov, O. N.; Walshe, J. M.; Moulder, S.,Palbociclib and letrozole in advanced breast cancer. N. Engl. J. Med. 2016,375 (20), 1925-1936.

[0269] 18. Fåhraeus, R.; Paramio, J. M.; Ball, K. L.; Laín, S.; Lane, D. P.,Inhibition of pRb phosphorylation and cell-cycle progression by a 20-residuepeptide from p16CDKN2 / INK4A. Curr. Biol. 1996, 6 (1), 84-91.

[0270] 19. Russo, A. A.; Tong, L.; Lee, J.-O.; Jeffrey, P. D.; Pavletich, N.P., Structural basis for inhibition of the cyclin-dependent kinase Cdk6 bythe tumour suppressor p16INK4a. Nature 1998, 395 (6699), 237-243.

[0271] 20. Fujimoto, K.; Hosotani, R.; Miyamoto, Y.; Doi, R.; Koshiba, T.;Otaka, A.; Fujii, N.; Beauchamp, R. D.; Imamura, M., Inhibition of pRbphosphorylation and cell cycle progression by an antennapediap16INK4A fusionpeptide in pancreatic cancer cells. Cancer Lett. 2000, 159 (2), 151-158.

[0272] 21. Hosotani, R.; Miyamoto, Y.; Fujimoto, K.; Doi, R.; Otaka, A.;Fujii, N.; Imamura, M., Trojan p16 peptide suppresses pancreatic cancergrowth and prolongs survival in mice. Clin. Cancer Res. 2002, 8 (4), 1271-1276.

[0273] 22. Camarero, J. A.; Pavel, J.; Muir, T. W., Chemical synthesis of acircular protein domain:evidence for folding-assisted cyclization. Angew.Chem., Int. Ed. 1998, 37 (3), 347-349.

[0274] 23. Fan, C.; Xiong, H.; Reynolds, N. M.; Soll, D., Rationallyevolving tRNA pyl for efficient incorporation of noncanonical amino acids.Nucleic Acids Res. 2015, 43 (22), e156.

[0275] 24. Serfling, R.; Lorenz, C.; Etzel, M.; Schicht, G.; Böttke, T.; Mörl, M.; Coin, I., Designer tRNAs for efficient incorporation of noncanonicalamino acids by the pyrrolysine system in mammalian cells. Nucleic Acids Res.2018, 46 (1), 1-10.

[0276] 25. Kavran, J. M.; Gundllapalli, S.; O'Donoghue, P.; Englert, M.; Söll, D.; Steitz, T. A., Structure of pyrrolysyl-tRNA synthetase, an archaealenzyme for genetic code innovation. Proc. Natl. Acad. Sci. 2007, 104 (27),11268-11273.

[0277] 26. Hohl, A.; Karan, R.; Akal, A.; Renn, D.; Liu, X.; Ghorpade, S.;Groll, M.; Rueping, M.; Eppinger, J., Engineering a polyspecific pyrrolysyl-tRNA synthetase by a high throughput FACS screen. Sci. Rep. 2019, 9 (1), 1-9.

[0278] 27. Nguyen, D. P.; Elliott, T.; Holt, M.; Muir, T. W.; Chin, J. W.,Genetically encoded 1, 2-aminothiols facilitate rapid and site-specificprotein labeling via a bio-orthogonal cyanobenzothiazole condensation. J. Am.Chem. Soc. 2011, 133 (30), 11418-11421.

[0279] 28. Pettersen, E. F.; Goddard, T. D.; Huang, C. C.; Couch, G. S.;Greenblatt, D. M.; Meng, E. C.; Ferrin, T. E., UCSF Chimera—a visualizationsystem for exploratory research and analysis. J. Comput. Chem. 2004, 25 (13),1605-1612.

[0280] 29. Wannier, T. M.; Kunjapur, A. M.; Rice, D. P.; McDonald, M. J.;Desai, M. M.; Church, G. M., Adaptive evolution of genomically recodedEscherichia coli. Proc. Natl. Acad. Sci. 2018, 115 (12), 3090-3095.

[0281] 30. Gaston, M. A.; Zhang, L.; Green-Church, K. B.; Krzycki, J. A.,The complete biosynthesis of the genetically encoded amino acid pyrrolysinefrom lysine. Nature 2011, 471 (7340), 647-650.

[0282] 31. Quitterer, F.; List, A.; Beck, P.; Bacher, A.; Groll, M.,Biosynthesis of the 22nd genetically encoded amino acid pyrrolysine:structureand reaction mechanism of PylC at 1.5 Å resolution. J. Mol. Biol. 2012, 424(5), 270-282.

[0283] 32. Brünger, A. T.; Adams, P. D.; Clore, G. M.; DeLano, W. L.; Gros,P.; Grosse-Kunstleve, R. W.; Jiang, J.-S.; Kuszewski, J.; Nilges, M.; Pannu,N. S., Crystallography & NMR system:A new software suite for macromolecularstructure determination. Acta Crystallogr. D 1998, 54 (5), 905-921.

[0284] 33. Brunger, A. T., Version 1.2 of the crystallography and NMRsystem. Nat. Protoc. 2007, 2 (11), 2728-2733.

[0285] 34. Wadhwani, P.; Afonin, S.; Ieronimo, M.; Buerck, J.; Ulrich, A.S., Optimized protocol for synthesis of cyclic gramicidin S:starting aminoacid is key to high yield. J. Org. Chem. 2006, 71 (1), 55-61.

[0286] 35. Yu, Z.; Yu, X. C.; Chu, Y.-H., MALDI-MS determination of cyclicpeptidomimetic sequences on single beads directed toward the generation oflibraries. Tetrahedron Lett. 1998, 39 (1-2), 1-4.

[0287] 36. Perlman, Z. E.; Bock, J. E.; Peterson, J. R.; Lokey, R. S.,Geometric diversity through permutation of backbone configuration in cyclicpeptide libraries. Bioorg. Med. Chem. Lett. 2005, 15 (23), 5329-5334.

[0288] 37. Sugahara, K. N.; Teesalu, T.; Karmali, P. P.; Kotamraju, V. R.;Agemy, L.; Girard, O. M.; Hanahan, D.; Mattrey, R. F.; Ruoslahti, E., Tissue-penetrating delivery of compounds and nanoparticles into tumors. Cancer Cell2009, 16 (6), 510-20.

[0289] 38. Ettinger, A.; Wittmann, T., Fluorescence live cell imaging.Methods Cell Biol. 2014, 123, 77-94.

[0290] 39. Ehrlich, M.; Gattner, M. J.; Viverge, B.; Bretzler, J.; Eisen, D.; Stadlmeier, M.; Vrabel, M.; Carell, T., Orchestrating the biosynthesis of an unnatural pyrrolysine amino acid for its direct incorporation into proteins inside living cells. Chem. Eur. J. 2015, 21 (21), 7701-7704.

[0291] 40. Exner, M. P.; Kuenzl, T.; To, T. M. T.; Ouyang, Z.; Schwagerus, S.; Hoesl, M. G.; Hackenberger, C. P.; Lensen, M. C.; Panke, S.; Budisa, N., Design of S-allylcysteine in situ production and incorporation based on a novel pyrrolysyl-tRNA synthetase variant. Chembiochem 2017, 18 (1), 85-90.

[0292] 41. Bauer, T. M.; Shaw, A. T.; Johnson, M. L.; Navarro, A.; Gainor, J. F.; Thurm, H.; Pithavala, Y. K.; Abbattista, A.; Peltz, G.; Felip, E., Brain penetration of lorlatinib: cumulative incidences of CNS and non-CNS progression with lorlatinib in patients with previously treated ALK-positive non-small-cell lung cancer. Target. Oncol. 2020, 15 (1), 55-65.

[0293] 42. Zhang, L.; Tam, J. P., Synthesis and application of unprotectedcyclic peptides as building blocks for peptide dendrimers. J. Am. Chem. Soc.1997, 119 (10), 2363-2370.

[0294] 43. Le Chevalier Isaad, A.; Papini, A. M.; Chorev, M.; Rovero, P.,Side chain-to-side chain cyclization by click reaction. J. Pept. Sci. 2009,15 (7), 451-454.

[0295] 44. Jagasia, R.; Holub, J. M.; Bollinger, M.; Kirshenbaum, K.; Finn,M., Peptide cyclization and cyclodimerization by CuI-mediated azidealkynecycloaddition. J. Org. Chem. 2009, 74 (8), 2964-2974.

[0296] 45. Li, X.; Fekner, T.; Chan, M. K., N6-(2-(R)-propargylglycyl)lysine as a clickable pyrrolysine mimic. Chem. Asian J. 2010, 5 (8), 1765-9.

[0297] 46. Fekner, T.; Li, X.; Lee, M. M.; Chan, M. K., A pyrrolysineanalogue for protein click chemistry. Angew. Chem. Int. Ed. 2009, 48 (9),1633-5.

[0298] 47. Wilson, D. S.; Keefe, A. D., Random mutagenesis by PCR. Curr.Protoc. Mol. Biol. 2000, 51 (1), 8.3. 1-8.3. 9.

[0299] 48. Miyazaki, K., MEGAWHOP cloning:a method of creating randommutagenesis libraries via megaprimer PCR of whole plasmids. Meth. Enzymol.2011, 498, 399-406.

[0300] 49. DeLano, W. L. The PyMOL molecular graphics system. http: / / www.pymol.org.

[0301] 50. Suzuki, T.; Miller, C.; Guo, L.-T.; Ho, J. M.; Bryson, D. I.;Wang, Y.-S.; Liu, D. R.; Söll, D., Crystal structures reveal an elusivefunctional domain of pyrrolysyl-tRNA synthetase. Nat. Chem. Biol. 2017, 13(12), 1261-1266.

[0302] 51. Nozawa, K.; O’Donoghue, P.; Gundllapalli, S.; Araiso, Y.;Ishitani, R.; Umehara, T.; Söll, D.; Nureki, O., Pyrrolysyl-tRNA synthetase–tRNA pyl structure reveals the molecular basis of orthogonality. Nature 2009,457 (7233), 1163-1167.

[0303] 52. Waterhouse, A.; Bertoni, M.; Bienert, S.; Studer, G.; Tauriello,G.; Gumienny, R.; Heer, F. T.; de Beer, T. A. P.; Rempfer, C.; Bordoli, L.,SWISS-MODEL:homology modelling of protein structures and complexes. NucleicAcids Res. 2018, 46 (W1), W296-W303.

[0304] 53. Abraham, M. J.; Murtola, T.; Schulz, R.; Páll, S.; Smith, J. C.;Hess, B.; Lindahl, E., GROMACS:High performance molecular simulations throughmulti-level parallelism from laptops to supercomputers. SoftwareX 2015, 1,19-25.

[0305] 54. Jorgensen, W. L.; Tirado-Rives, J., Potential energy functionsfor atomic-level simulations of water and organic and biomolecular systems.Proc. Natl. Acad. Sci. 2005, 102 (19), 6665-6670.

[0306] 55. Dodda, L. S.; Vilseck, J. Z.; Tirado-Rives, J.; Jorgensen, W. L.,1.14 CM1A-LBCC:localized bond-charge corrected CM1A charges for condensed-phase simulations. J. Phys. Chem. B 2017, 121 (15), 3864-3870.

[0307] 56. Dodda, LS; Cabeza de Vaca, I.; Tirado-Rives, J.; Jorgensen, WL, LigParGen web server: an automatic OPLS-AA parameter generator for organicligands. Nucleic Acids Res. 2017, 45 (W1), W331-W336.

[0308] All publications, granted patents and patent applications cited in this specification are incorporated herein by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated herein by reference.

[0309] It should be understood that this disclosure is not limited to the specific methods, protocols, cell lines, animal species or genera and reagents described, as they can vary. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the scope of this disclosure, which will be limited only by the appended claims.

[0310] Informal sequence list

Claims

1. A recombinant host cell comprising a polynucleotide sequence encoding a mutant PylC capable of synthesizing non-classical amino acids (ncAA), wherein the mutant PylC comprises the same as SEQ ID NO: 1 (wild-type Methanococcus masculinii ( M.mazei PylC) is a polypeptide sequence with at least 95% sequence identity.

2. The recombinant host cell of claim 1, further comprising: (a) A multinucleotide sequence encoding wild-type PylRS or mutant PylRS, wherein the mutant PylRS is capable of using ncAA as a substrate, and (b) A polynucleotide sequence encoding tRNA, wherein the ncAA is incorporated at the amber codon.

3. The recombinant host cell of claim 1 or 2, wherein the mutant PylC comprises a mutant combination of residues 177, 179, 233 and 256 corresponding to wild-type archaea PylC.

4. The recombinant host cell of claim 3, wherein the combination of mutations in the mutant PylC is selected from: 1) S177N, E179P, D233S and T256V; 2) S177C, E179T, D233S, and T256V; 3) E179C and D233N; 4) S177A, E179A, D233H and T256M; 5) S177A, E179S, D233H and T256L; 6) E179C, D233H, and T256V; 7) E179V, D233N, and T256V; 8) E179V, D233H, and T256N; 9) S177T, E179P, D233H and T256M; 10) S177G, E179V, D233S and T256L; 11) E179N, D233H, and T256L; 12) S177A, E179V, D233H, and T256I; 13) S177L, E179I, D233H and T256C; 14) S177A, E179I, D233T and T256L; 15) S177H, E179L, D233S and T256V; 16) S177A, E179M, D233S, and T256L; 17) E179I, D233H, and T256I; 18) S177A, E179T, D233H and T256V; 19) S177A, E179A, D233N, and T256C; 20) S177C, E179P, D233H and T256L; 21) S177G, E179I, and D233H; 22) S177M, E179V, D233H and T256V; 23) S177V, E179V, D233H and T256V; 24) S177A, E179A, D233S and T256V; 25) S177G, E179M, D233N and T256V; 26) S177G, E179I, D233N and T256V; 27) S177G, E179T, D233S and T256L; 28) S177A, E179T, D233S and T256L; 29) S177L, E179T, D233S and T256I; 30) S177T, E179P, D233H and T256M; 31) E179L, D233L, and T256V; 32) S177A, E179V, D233Q, and T256C; 33) S177L, E179V, D233S and T256A; 34) S177A, E179C, D233L and T256C; 35) S177N, E179C, D233T and T256C; 36) S177H, E17L9, D233S and T256S; 37) E179G, D233L, and T256V; 38) S177N, E179P, D233S and T256V; 39) S177M, E179G, D233H and T256Y; 40) S177M, E179L, D233S and T256C; 41) S177A, E179C, and D233H; 42) E179L, D233L, and T256I; 43) S177L, E179V, D233H and T256Y; 44) S177A, E179Q, D233N and T256C; and 45) S177L, E179A, D233H and T256V.

5. The recombinant host cell of claim 3 or 4, wherein the combination of mutations in the mutant PylC is S177N, E179P, D233S and T256V.

6. The recombinant host cell of claim 3, wherein the combination of mutations in the mutant PylC is selected from: 1) S177G, E179V, D233N and T256V; 2) S177V, E179L, D233H and T256L; 3) S177G, E179V, D233H and T256A; 4) S177V, E179C, D233H and T256L; 5) S177M, E179V, D233H and T256Y; 6) S177C, E179P, D233H and T256L; 7) S177G, E179A, D233S and T256L; 8) S177V, E179L, D233H and T256V; 9) S177G, E179L, D233H and T256L; 10) S177I, D233H, and T256V; 11) S177L, E179V, D233T and T256I; 12) S177L, E179A, D233H and T256M; 13) S177L, E179V, D233H and T256A; 14) S177I, E179M, D233H and T256L; 15) S177A, E179I, D233N, and T256I; 16) S177I, E179C, D233H and T256L; 17) E179S and D233H; 18) S177V, E179V, D233H and T256A; 19) S177C, E179V, D233H and T256Y; 20) S177I, E179A, D233H and T256H; 21) S177L, E179C, D233H and T256V; 22) E179V and D233H; 23) S177V, E179V, D233H and T256L; 24) S177I, E179L, D233H and T256A; 25) S177L, E179A, D233H and T256Y; 26) S177G, E179V, D233L and T256L; 27) E179P, D233N, and T256V; 28) S177G, E179V, D233N and T256N; 29) S177L, E179C, D233G and T256H; 30) S177H, E179G, D233S and T256A; 31) S177A, E179C, D233S and T256L; 32) S177G, E179V, D233L and T256Y; 33) S177L, E179T, D233T and T256M; 34) S177G, E179L, D233W and T256S; 35) S177V, E179L, D233S and T256A; 36) E179V, D233L, and T256L; 37) S177L, E179P, D233S and T256N; and 38) S177V, E179V, D233T and T256L.

7. The recombinant host cell of claim 3 or 6, wherein the combination of mutations in the mutant PylC is S177G, E179V, D233N, and T256V.

8. The recombinant host cell according to any one of claims 1 to 7, wherein the mutant PylRS comprises mutations corresponding to (a) G114E, C348V and S451F of wild-type archaea PylRS, or (b) L301M, L305I, Y306L, L309A and C348F of wild-type archaea PylRS.

9. The recombinant host cell according to any one of claims 1 to 8, wherein the genome sequence encoding release factor 1 (RF1) of the recombinant host cell is at least partially deleted.

10. The recombinant host cell according to any one of claims 1 to 9, wherein the cell is a prokaryotic cell.

11. The recombinant host cell according to any one of claims 1 to 10, wherein the cell is a bacterial cell.

12. The recombinant host cell according to any one of claims 1 to 11, wherein the cell is *Escherichia coli* (…). E. coli )cell.

13. A composition comprising the recombinant host cell of any one of claims 1 to 12.

14. The lysate of recombinant host cells according to any one of claims 1 to 12.

15. Methods for recombinant protein synthesis, including: (a) Introducing a nucleotide sequence encoding the protein into a recombinant host cell according to any one of claims 1 to 12, wherein the nucleotide sequence comprises (i) at least one TAG codon within the coding sequence of the nucleotide sequence, and (ii) a TAA codon or a TGA codon at the end of the coding sequence of the nucleotide sequence; and (b) The recombinant host cells are cultured under conditions that allow transcription from the nucleotide sequence and protein synthesis, thereby producing the protein; and wherein... The protein contains one or more ncAAs, optionally wherein the one or more ncAAs are selected from D-Cys-ε-Lys, D-Pra-ε-Lys, D-allyl-ε-Lys and 2-chloroacetyl-ε-Lys.

16. The method of claim 15, further comprising separating the protein produced in step (b).

17. The method of claim 16, further comprising placing the isolated protein under conditions that allow the protein to cyclize, the cyclization being carried out by forming a covalent bond between (i) a first portion on a first ncAA and (ii) a second portion on a cysteine ​​residue, a C-terminal thioester, or a second portion on a second ncAA.

18. The method of any one of claims 15 to 17, wherein the protein comprises a polypeptide sequence derived from a small ubiquitin-associated modified (SUMO) protein.

19. A method for recombinant synthesis of a protein, comprising contacting a nucleotide sequence encoding the protein with a cell lysate of a recombinant host cell according to any one of claims 1 to 12.

20. The method of any one of claims 15 to 17, wherein the protein is P16.

21. A nucleic acid comprising a polynucleotide sequence encoding a mutant PylC capable of synthesizing ncAA.

22. The nucleic acid of claim 21, wherein the mutant PylC comprises a combination of mutations corresponding to residues 177, 179, 233, and 256 of the wild-type archaea PylC.

23. The nucleic acid of claim 22, wherein the combination of mutations in the mutant PylC is selected from: 1) S177N, E179P, D233S and T256V; 2) S177C, E179T, D233S, and T256V; 3) E179C and D233N; 4) S177A, E179A, D233H and T256M; 5) S177A, E179S, D233H and T256L; 6) E179C, D233H, and T256V; 7) E179V, D233N, and T256V; 8) E179V, D233H, and T256N; 9) S177T, E179P, D233H and T256M; 10) S177G, E179V, D233S and T256L; 11) E179N, D233H, and T256L; 12) S177A, E179V, D233H, and T256I; 13) S177L, E179I, D233H and T256C; 14) S177A, E179I, D233T and T256L; 15) S177H, E179L, D233S and T256V; 16) S177A, E179M, D233S, and T256L; 17) E179I, D233H, and T256I; 18) S177A, E179T, D233H and T256V; 19) S177A, E179A, D233N, and T256C; 20) S177C, E179P, D233H and T256L; 21) S177G, E179I, and D233H; 22) S177M, E179V, D233H and T256V; 23) S177V, E179V, D233H and T256V; 24) S177A, E179A, D233S and T256V; 25) S177G, E179M, D233N and T256V; 26) S177G, E179I, D233N and T256V; 27) S177G, E179T, D233S and T256L; 28) S177A, E179T, D233S and T256L; 29) S177L, E179T, D233S and T256I; 30) S177T, E179P, D233H and T256M; 31) E179L, D233L, and T256V; 32) S177A, E179V, D233Q, and T256C; 33) S177L, E179V, D233S and T256A; 34) S177A, E179C, D233L and T256C; 35) S177N, E179C, D233T and T256C; 36) S177H, E17L9, D233S and T256S; 37) E179G, D233L, and T256V; 38) S177N, E179P, D233S and T256V; 39) S177M, E179G, D233H and T256Y; 40) S177M, E179L, D233S and T256C; 41) S177A, E179C, and D233H; 42) E179L, D233L, and T256I; 43) S177L, E179V, D233H and T256Y; 44) S177A, E179Q, D233N and T256C; and 45) S177L, E179A, D233H and T256V.

24. The nucleic acid of claim 22 or 23, wherein the combination of mutations in the mutant PylC is S177N, E179P, D233S and T256V.

25. The nucleic acid of claim 22, wherein the combination of mutations in the mutant PylC is selected from:

1. S177G, E179V, D233N, and T256V; 2. S177V, E179L, D233H, and T256L; 3. S177G, E179V, D233H, and T256A; 4. S177V, E179C, D233H, and T256L; 5. S177M, E179V, D233H and T256Y; 6. S177C, E179P, D233H, and T256L; 7. S177G, E179A, D233S, and T256L; 8. S177V, E179L, D233H and T256V; 9. S177G, E179L, D233H, and T256L; 10. S177I, D233H, and T256V; 11. S177L, E179V, D233T, and T256I; 12. S177L, E179A, D233H, and T256M; 13. S177L, E179V, D233H and T256A; 14. S177I, E179M, D233H and T256L; 15. S177A, E179I, D233N, and T256I; 16. S177I, E179C, D233H and T256L; 17. E179S and D233H; 18. S177V, E179V, D233H and T256A; 19. S177C, E179V, D233H, and T256Y; 20. S177I, E179A, D233H and T256H; 21. S177L, E179C, D233H, and T256V; 22. E179V and D233H; 23. S177V, E179V, D233H and T256L; 24. S177I, E179L, D233H and T256A; 25. S177L, E179A, D233H and T256Y; 26. S177G, E179V, D233L and T256L; 27. E179P, D233N, and T256V; and 28. S177G, E179V, D233N and T256N.

26. The nucleic acid of claim 22 or 25, wherein the combination of mutations in the mutant PylC is S177G, E179V, D233N and T256V.

27. The nucleic acid as claimed in claim is a polynucleotide sequence comprising encoding a mutant PylRS, wherein the mutant PylRS is capable of using ncAA as a substrate.

28. The nucleic acid of claim 27, wherein the mutant PylRS comprises mutations corresponding to (a) G14E, C348V and S451F of wild-type archaea PylRS, or (b) L301M, L305I, Y306L, L309A and C348F of wild-type archaea PylRS.

29. A composition comprising the nucleic acid of any one of claims 21 to 28.

30. An expression component or vector comprising a polynucleotide sequence of the nucleic acid according to any one of claims 21 to 28.

Citation Information

Patent Citations

  • Biosynthetically generated pyrroline-carboxy-lysine and site specific protein modifications via chemical derivatization of pyrroline-carboxy-lysine and pyrrolysine residues

    US20140302553A1

  • SUMO peptides for treating neurodegenerative diseases

    US20220194997A1

  • Pyrrolysine analogs

    US8921571B2

  • Custom-made meganuclease and use thereof

    WO2004067736A2