Compounds comprising non-canonical amino acids and uses thereof

By employing a propeptide-based strategy with the bacterial ABC transporter system to enhance ncAA uptake, the method addresses low protein yields and bioavailability issues, achieving efficient incorporation and production of proteins with diverse functionalities.

WO2026078118A1PCT designated stage Publication Date: 2026-04-16ETH ZURICH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/079084
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-09
Filing Date
2025-10-09
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

The widespread use of genetic code expansion for incorporating non-canonical amino acids (ncAAs) into proteins is limited by low protein production yields, enzymatic inefficiencies, and intracellular bioavailability issues, necessitating improved strategies for enhancing ncAA uptake and bioavailability.

Method used

Engineering a propeptide-based strategy combined with the bacterial ABC transporter system to actively transport isopeptide-linked lysine derivatives, specifically using the oligopeptide permease transporter Opp, to increase intracellular concentrations of ncAAs, and evolving the periplasmic binding protein for preferential uptake of these derivatives.

Benefits of technology

This approach achieves high incorporation efficiencies of ncAAs into proteins, rivaling wild-type expressions, providing a versatile toolbox for functionalities like bioorthogonal labeling and chemical ligation, and enabling efficient production of proteins with expanded alphabets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025079084_16042026_PF_FP_ABST
    Figure EP2025079084_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are compounds comprising an isopeptide-linked lysine or an analog thereof. Further provided herein are methods for engineering peptide-binding proteins having increased affinity and / or selectivity for the compound according to the invention, as well as peptide-binding proteins with increased selectivity for the compound according to the invention. Moreover, methods for incorporating non-canonical amino acids into proteins are provided herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPOUNDS COMPRISING NON-CANONICAL AMINO ACIDS AND USES THEREOF

[0002] BACKGROUND TO THE INVENTION

[0003] Co-translational incorporation of non-canonical amino acids (ncAAs) into proteins via genetic code expansion (GCE) has become a versatile approach to increase the chemical space of the proteome. By leveraging orthogonal aminoacyl-tRNA synthetases (aaRSs), a plethora of diverse functionalities (1) have been incorporated into proteins of interest (POIs) across all domains of life, most commonly in response to the amber nonsense codon (amber suppression). Site-specific encoding of functionalities such as post-translational modifications (PTMs), bioorthogonal handles, crosslinking moieties, spectroscopic probes, and photocaged amino acids offers a wide range of tools for studying and modifying protein functions and generating proteins with potential therapeutic and biotechnological significance (2, 3).

[0004] While GCE holds promise in a wide range of fields, its widespread use in commercial and research applications remains limited due to several challenges, most notably low protein production yields. This stems from insufficient enzymatic activities of orthogonal aaRSs towards many ncAA substrates, as well as unfavorable competition of aminoacylated tRNAs with release factors, causing premature translational termination at introduced nonsense codons (4, 5) In addition, many ncAAs require advanced expertise in chemical synthesis, or are prohibitively expensive, and are used at high millimolar concentrations in typical expression experiments, a factor that is exacerbated by low protein yields.

[0005] Significant advances have been made in recent years towards addressing these challenges. Optimization of aaRS / tRNA expression systems combined with novel selection and evolution strategies to improve both orthogonal aaRSs and tRNAs have considerably increased suppression efficiencies (6-10). Efforts to bypass competition with release factors have led to orthogonal ribosomes that read quadruplet codons (11, 12), release factor knockout strains (13, 14) and recoded strains with compressed genetic codes (15, 16), allowing also for sense codon reassignment (17).

[0006] Another less investigated factor impacting ncAA incorporation efficiencies lies in their intracellular bioavailability. In typical GCE experiments, chemically synthesized ncAAs are added to expression media and enter cells via passive diffusion, resulting in intracellular concentrations at best equal but most likely severely below those added to media. This specifically reduces aminoacylation efficiencies of aaRS / ncAA combinations that are operating below saturating conditions (4). Furthermore, the reliance on passive diffusion limits the full potential of introducing new-to-nature functionalities, as cell-permeability needs to be considered when designing new ncAAs. For a handful of ncAAs, efforts in introducing and engineering enzymatic pathways to biosynthesize corresponding ncAAs directly within cells have resulted in increased incorporation efficiencies compared to exogenous ncAA addition (18-23). Despite being an attractive prospect for GCE, introducing novel metabolic pathways into new hosts may require significant development. In addition, many ncAA functionalities lack biosynthetic strategies, necessitating substantial advancements before such approaches can be widely applied for increasing ncAA bioavailability and thus genetic encoding.

[0007] Engineering membrane transport systems as an alternative strategy for boosting intracellular ncAA uptake, is relatively unexplored, but holds the potential of being widely applicable for increasing expression yields of modified proteins. Previous research in this direction has investigated mutants of a periplasmic leucine binding protein to improve uptake and incorporation of various known ncAAs (24). Another study has used a 'trojan-horse' approach, in which a sulfonic acid moiety attached to a cargo molecule serves as recognition motif for a promiscuous bacterial sulfonate importer. Attaching an impermeant ncAA to a sulfonate carrier, followed by cleavage of the carrier via an engineered enzyme, led to increased cytosolic concentrations of this ncAA, but incorporation into proteins was not shown due to lack of a specific aaRS (25). The use of peptide transporters has also been investigated (26, 27), most notably for transporting phosphotyrosine as a lysine-linked dipeptide (28), which was believed to be a substrate for the dipeptide transporter Dpp. Peptide transporters, especially ATP-binding cassette (ABC) transporters like the Dpp-transporter, are promising candidates for engineering ncAA uptake due to their substrate promiscuity and ability to maintain high concentration gradients across the membrane (29). In addition, attaching ncAAs to peptide carriers is easily achievable via solid-phase peptide synthesis (SPPS) requiring minimal chemical expertise.

[0008] Despite these recent progresses, there is still a need in the art for more universal strategies to supply cells with sufficient amounts of non-canonical amino acids.

[0009] SUMMARY OF THE INVENTION

[0010] The present invention is characterized in the herein provided embodiments and claims. In particular, the present invention relates, inter alia, to the following embodiments:

[0011] 1. A compound comprising an isopeptide-linked lysine or an analogue thereof, the compound being a compound of formula (I): or a salt thereof, wherein:

[0012] A is selected from -NHz, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), -N+(Ci-s alkyl)(Ci-5alkyl), -NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2, or A is a peptidyl group;

[0013] B is selected from: wherein the left empty valence is connected to A and the right empty valence is connected to C, wherein:

[0014] Z is hydrogen or a side chain of an amino acid, in particular wherein Z is selected from hydrogen, Ci-8 alkyl, C2-s alkenyl, C2-s alkynyl, -(Ci-6 alkylene)-N3, -(Ci-6 alkylene)-CN, -(Ci-6 alkylene)-Hal, -(Ci-6 alkylene)-O-Rz, -(Ci-6 alkylene)-S-Rz, -(Ci-6 alkylene)-N(Rz)-Rzz, -(Ci-6 alkylene)-CO-Rz, -(Ci-6 alkylene)-COO-Rz, -(Ci-6 alkylene)-O-CO-(Ci-6 alkyl), -(Ci-6 alkylene)-CO-N(Rz)-Rz, -(Ci-6 alkylene)-N(Rz)-CO- (Ci-6 alkyl), -(Ci-6 alkylene)-CO-N(Rz)-O-Rz, -(Ci-6alkylene)-O-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-C(=N-Rz)-N(Rz)-Rz, -(Ci-6alkylene)-SO3-Rz, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z groups are each optionally substituted with one or more groups selected from -OH, -SH and Hal, further wherein one -CH2- group in said alkyl may be replaced with further wherein each Rzis independently selected from hydrogen nd wherein Rzzis selected from hydrogen, Ci-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-5 alkenyl), -CO(Ci-s alkynyl), -CO(Ci-s alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)-carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -CO-O-(Ci-5 alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-s alkynyl), -CO-0-(Co-s alkylene)- carbocyclyl, -CO-0-(Co-s alkylene)-heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs, Z1and Z2are each independently selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, -(C1-6 alkylene)-S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO- Rz, -(C1-6 alkylene)-COO-Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO- N(RZ)-RZ, -(C1-6 alkylene)-N(Rz)-CO-(Ci-6alkyl), -(Ci-6alkylene)-CO-N(Rz)-O-Rz, -(Ci-6alkylene)-O-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)- C(=N-RZ)-N(RZ)-RZ, -(Ci-6 alkylene)-SO3-Rz, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH, -SH or Hal, wherein each Rzis independently selected from hydrogen and Ci-6 alkyl, and Rzzis selected from hydrogen, Ci-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-s alkenyl), -CO(Ci-s alkynyl), -CO(Ci-s alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)- carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -CO-O-(Ci-s alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-5 alkynyl), -CO-0-(Co-s alkylene)-carbocyclyl, -CO-0-(Co-s alkylene)- heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs, and further wherein one - N = N

[0015] CH2- group in said alkyl may be replaced with provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen, or Z1and Z2are joined together to form, together with carbon atom(s) that otherwise carry Z1and Z2, a non-aromatic carbocyclic ring or non-aromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non- aromatic heterocyclic ring are each optionally substituted with one or more Rs;

[0016] -C-D- are defined as follows:

[0017] C is selected from -CO-NH-, -CO-O-, -CO-S- and -CO-N(Ci-s alkyl)-, wherein the left empty valence is connected to B and the right empty valence is connected to D, valence is connected to C and the right empty valence is connected to -CO-NH-, wherein: X is a hydrogen or a side chain of an amino acid, in particular wherein X is selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(Ci-6 alkylene)- N3, -(Ci-6 alkylene)-CN, -(Ci-6 alkylene)-Hal, -(Ci-6 alkylene)-O-Rx, -(Ci-6 alkylene)-S-Rx, -(Ci-6 alkylene)-N(Rx)-Rx, -(Ci-6 alkylene)-CO-Rx, -(Ci-6 alkylene)-COO-Rx, -(Ci-6 alkylene)-O-CO-(Ci-6 alkyl), -(Ci-6 alkylene)-CO-N(Rx)- Rx, -(Ci-6 alkylene)-N(Rx)-CO-(Ci-6alkyl), -(Ci-6alkylene)-CO-N(Rx)-O-Rx, -(Ci-6alkylene)-O-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(Ci-6 alkylene)-SO3-Rx, -(Co-6 alkylene)- carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X groups are each optionally substituted with one or more -OH, -SH or -Hal, wherein each Rxis independently selected from hydrogen and Ci-6 alkyl, and further wherein N=N one -CH2- group in said alkyl may be replaced with

[0018] X1and X2are each independently selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)- Hal, -(C1-6 alkylene)-O-Rx, -(C1-6 alkylene)-S-Rx, -(C1-6 alkylene)-N(Rx)-Rx, -(C1-6 alkylene)-CO-Rx, -(C1-6 alkylene)-COO-Rx, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), - (C1-6 alkylene)-CO-N(Rx)-Rx, -(C1-6 alkylene)-N(Rx)-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rx)-O-Rx, -(Ci-6alkylene)-O-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-CO-N(Rx)- Rx, -(Ci-6alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(Ci-6alkylene)-SO3-Rx, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)- heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X1and X2groups are each optionally substituted with one or more -OH, -SH or Hal (preferably with one or more -OH), wherein each Rxis independently selected from hydrogen and Ci-6 alkyl, further N = N wherein one -CH2- group in said alkyl may be replaced with provided that at least one of X1and X2is not hydrogen, further provided that if X1and X2are connected to the same carbon atom, then both X1and X2are not hydrogen, or X1and X2are joined together to form, together with carbon atom(s) that otherwise carry X1and X2, a non-aromatic carbocyclic ring or non-aromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non-aromatic heterocyclic ring are each optionally substituted with one or more Rs; or -C-D- taken together is -CO-(N-heterocycloalkylene)-, wherein said N- heterocycloalkylene is connected to said -CO- group of -C-D- through its nitrogen atom and wherein said N-heterocycloalkylene is optionally substituted with one or more Rs;

[0019] E is - (C1-6 alkyleneJ-CHY^2, wherein Y1is selected from hydrogen, -SH, -OH and -NH2, and wherein Y2is hydrogen, COOH or CONH2, provided that at least one of Y1and Y2is not hydrogen, further wherein said alkylene is optionally substituted with an -OH group, -SH group or Hal, and wherein one -CH2- group in said alkylene may be replaced with each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-5 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s al kyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-5 alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-5 alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH(Ci-s alkylene)-CN, -(C0-3 alkylene)-CO-N(Ci-5 alkyl)(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-(Ci-s alkylene)- CN, -(C0-3 alkylene)-N(Ci-5 alkyl)-CO-(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-5 alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s a I ky l)-SC>2-(Ci-5 alkyl), -(C0-3 alkylene)-SO2-(Ci-5alkyl), -(C0-3 alkylene)-SO-(Ci-5alkyl), -OSO2F, -B(OH)2, -PO4H, -PO3H, - SO4H, -SO3H, -L-(CO-3 alkylene)-carbocyclyl, and -L-(Co-3 alkylene)-heterocyclyl, wherein L is selected from -(C0-3 alkylene)-NH-, -(C0-3 alkylene)-O-, -(C0-3 alkylene)-CO-, -(C0-3 alkylene)-NH-CO-, -(C0-3 alkylene)-CO-NH-, -(C0-3 alkylene)-NH-CO-O-, -(C0-3 alkylene)— O-CO-NH-, -(C0-3 alkylene)-NH-CO-NH- and -(C0-3 alkylene)-NH-CO-NH-, and wherein the carbocyclyl moiety in said -L-(Co-3 alkylene)-carbocyclyl and the heterocyclyl moiety in said -L-(CO-3alkylene)-heterocyclyl are each optionally substituted with one or more groups independently selected from C1-4 alkyl, halogen, -CN, -NO2, -OH, -O-(Ci-4 alkyl), - SH, -S-(Ci-4alkyl), -NHz, -NH(CI-4alkyl), -N(CI-4alkyl)(Ci-4alkyl), -COOH, -COO(Ci-4alkyl), -CONHz, -CONH(CI-4alkyl), -CON(CI-4al kyl)(Ci-4alkyl), -NHCO(CI-4alkyl) and -N(CI-4al ky I )- CO(Ci-4 alkyl). The compound of embodiment 1, wherein A is selected from -NHz, -OH, -SH, -NH(Ci-s alkyl), -N(CI-5alkyl)(Ci-5alkyl), -NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2. The compound of embodiment 1 or 2, wherein A is selected from -NHz, -OH, -SH, -NH(Ci- 5 alkyl), -N(CI-5al kyl)(Ci-s alkyl), -COOH, and -N3. The compound of any one of embodiments 1 to 3, wherein A is -NHz, -NH(CH3), -COOH, -SH, and Ns, preferably wherein a is NHz •

[0020] Z The compound of any one of embodiments 1 to 4, wherein B is preferably

[0021] Z wherein wherein the left empty valence is connected to A and the right empty valence is connected to C. The compound of embodiment 5, wherein Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), preferably wherein Z is selected from hydrogen, -CH2CHzCH2CHz-NHz and -CH2CHzCHzCHz-NH-CO-CH3, or wherein Z is selected from any one of:

[0022]

[0023] 7. The compound of embodiment 6, wherein Z is hydrogen.

[0024] 8. The compound of any one of embodiments 1 to 7, wherein C is selected from -CO-NH- and -CO-O-, wherein the left empty valence is connected to B and the right empty valence is connected to D. 9. The compound of embodiment 8, wherein C is -CO-NH-, wherein the left empty valence is connected to B and the right empty valence is connected to D. X

[0025] 10. The compound of any one of embodiments 1 to 9, wherein D is preferably

[0026] X wherein D , wherein the left empty valence is connected to C and the right empty valence is connected to -CO-NH- group.

[0027] 11. The compound of embodiment 10, wherein X is selected from Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CI, -(C1-6 alkylene)-NH2, -(Ci-6 alkylene)-NH-CO-NH2, -(Ci-6alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, -(Ci-6alkylene)-CO-NH2, -(Ci-6alkylene)-CO-NH-OH, -(Ci-6alkyleneJ-SOsH, phenyl, -(Ci-6 alkylene)-phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)-heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)- heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH, and further wherein one -CH2- group in said alkyl may be replaced

[0028] 12. The compound of embodiment 11, wherein X is selected from methyl, ethyl, isopropyl, sec-butyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl-diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4- yl)methyl and benzyl.

[0029] 13. The compound of any one of embodiments 1 to 7, wherein -C-D- is a moiety according to formula wherein the left empty valence is connected to B, and the right empty valence is connected to -CO-NH- group.

[0030] 14. The compound of any one of embodiments 1 to 13, wherein E is selected from: - CH2CH2CH2CH2-CH(-NH2)-COOH, -CH2CH2CH2-CH(-NH2)-COOH, -CH2CH2CH2CH2CH2- COOH, -CH2CH2CH2CH2CH2-NH2, -CH2CH2CH2CH2-CH(-NH2)-CONH2, -CH2CH2CH2CH2-CH(- OH)-COOH; -CH2CH2-CH(-NH2)-COOH, and -CH2-CH(-NH2)-COOH, preferably wherein E is CH2CH2CH2CH2-CH(-NH2)-COOH. The compound of embodiment 1, wherein the compound is a compound according to or its salt, preferably wherein the compound of formula (II) is preferably a compound of or its salt, wherein Z and X are as defined in any one of embodiments 1 to 14, preferably wherein Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-e alkylene)-N(Rz)-CO-(Ci-6 alkyl), more preferably wherein Z is selected from hydrogen, - CH2CH2CH2CH2-NH2and -CH2CH2CH2CH2-NH-CO-CH3, even more preferably wherein Z is hydrogen, or wherein Z is selected from any one of:

[0031] preferably wherein X is selected from methyl, ethyl, isopropyl, sec-butyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl- diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4-yl)methyl and benzyl; or wherein the compound is a compound of formula (III):

[0032] or its salt, preferably wherein the compound of formula (III) is a compound of formula or its salt, wherein Rpis hydrogen or -OH, and wherein Z is as defined in any one of embodiments 1 to 14, preferably wherein Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), more preferably wherein Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and -CH2CH2CH2CH2-NH-CO-CH3, even more preferably wherein Z is hydrogen; or wherein Z is selected from any one of:

[0033] The compound of embodiment 1, wherein the compound is a compound selected from Table 1, or its salt. A method for identifying an engineered peptide-binding protein having increased affinity and / or specificity for the compound according to any one of embodiments 1 to 16, the method comprising the steps of: a) providing a nucleic acid encoding a peptide-binding protein comprising one or more mutations; b) expressing the nucleic acid of step (a) in a cell comprising an orthogonal translation system, wherein the orthogonal translation system comprises an orthogonal aminoacyl- tRNA synthetase / tRNA pair specific for an isopeptide-linked lysine or an analog thereof, and a reporter protein gene comprising one or more codons that have been re-allocated for the incorporation of the isopeptide-linked lysine or the analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; c) contacting the cell of step (b) with a compound according to any one of embodiments 1 to 16, wherein the compound is transported into the cell and cleaved inside the cell to release the isopeptide-linked lysine or an analog thereof; d) detecting the expression of the reporter protein gene by the cell, wherein the expression is dependent on the successful uptake of the compound in step (c); and e) identifying the engineered peptide-binding protein as having increased affinity and / or specificity for the compound according to any one of embodiments 1 to 16 based on the expression of the reporter protein gene detected in step (d). The method according to embodiment 17, wherein the peptide-binding protein is the peptide-binding protein of an oligopeptide permease (OppA). The method according to embodiment 17 or 18, wherein the peptide-binding protein is the peptide-binding protein of the E. coli oligopeptide permease (OppA; SEQ ID NO:1). The method according to embodiment 19, wherein in step (a) one or more mutations are introduced at positions V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, 5460, V482, L531 and / or N533 of SEQ ID NO:l. The method according to any one of embodiments 17 to 20, wherein the mutations introduced in step (a) are introduced using error-prone PCR and / or saturated mutagenesis and / or continuous directed evolution. The method according to any one of embodiments 17 to 21, wherein the cell in step (b) is a prokaryotic cell, in particular an E. coli cell. The method according to any one of embodiments 17 to 22, wherein the orthogonal aminoacyl-tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair. The method according to any one of embodiments 17 to 23, wherein the reporter protein gene in step (b) comprises one or more amber stop codons that have been reallocated to incorporate the isopeptide-linked lysine or the analog thereof. The method according to any one of embodiments 17 to 24, wherein in step (c), the cell is contacted with a compound according to any one of embodiments 1 to 16 in the presence of a peptide that competes for binding to the peptide-binding protein. The method according to any one of embodiments 17 to 25, wherein the reporter protein is a fluorescent protein. The method according to embodiment 26, wherein the detection of the expression of the reporter protein gene in step (d) is performed using fluorescence-activated cell sorting (FACS). The method according to any one of embodiments 17 to 25, wherein the detection of the expression of the reporter protein gene in step (d) is performed using antibiotic resistance screening and / or using growth-based selection. The method according to any one of embodiments 17 to 28, wherein the isopeptide- linked lysine or the analog thereof is obtained by intracellular cleavage of the compound of any one of embodiments 1 to 16. The method according to any one of embodiments 17 to 29, wherein identifying the engineered peptide-binding protein as having increased specificity for the compound in step (e) comprises a step of determining the sequence of the nucleic acid encoding the peptide-binding protein and / or determining specific mutations in the nucleic acid encoding the peptide-binding protein. The method according to any one of embodiments 17 to 30, wherein the isopeptide- linked lysine or the analog thereof has the structure: wherein C' is -NH2 or -OH, and wherein D and E is defined as in any one of embodiments 1 to 16, wherein in D the left empty valence is connected to C'. An engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, 5460, V482, L531 and / or N533. The engineered variant of embodiment 32, wherein the one or more mutations are located in positions: V60, 863, Y135, T173, H187, V193, D221, W222, 1303, N337, R439, W442, C443, D445, 8460, L531 and / or N533.

[0034] 34. The engineered variant of embodiment 32 or 33, wherein the one or more mutations are located in positions: V60, 863, Y135, H187, D221, W222, N337, R439, W442, C443, D445, 8460, L531 and / or N533.

[0035] 35. The engineered variant of any one of embodiments 32 to 34, wherein the one or more mutations are located in positions: V60, 863, Y135, H187, D221, W222, R439, W442, C443, D445, 8460, L531 and / or N533.

[0036] 36. The engineered variant of any one of embodiments 32 to 35, wherein the one or more mutations are located in positions: V60, 863, Y135, H187, R439, W442, C443, D445, 8460, L531 and / or N533.

[0037] 37. A nucleic acid encoding the engineered variant of any one of embodiments 32 to 36.

[0038] 38. A vector comprising the nucleic acid according to embodiment 37.

[0039] 39. A cell comprising the nucleic acid according to embodiment 37 or the vector according to embodiment 38.

[0040] 40. The cell according to embodiment 39, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

[0041] 41. A method for incorporating a non-canonical amino acid, or an analog thereof, into a protein, the method comprising the steps of: a) providing a cell comprising at least one orthogonal translation system, wherein the at least one orthogonal translation system comprises an orthogonal aminoacyl-tRNA synthetase / tRNA pair specific for a non-canonical amino acid, or an analog thereof, comprised in the compound according to any one of embodiments 1 to 16, and a nucleic acid molecule encoding the protein, wherein the nucleic acid molecule encoding the protein comprises one or more codons that have been re-allocated for the incorporation of the non-canonical amino acid or an analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; b) contacting the cell of step (a) with the compound according to any one of embodiments 1 to 16, wherein the compound is transported into the cell and cleaved inside the cell to release an isopeptide-linked lysine or an analog thereof and, optionally, another non-canonical amino acid or an analog thereof; Y1 c) incorporating (i) the isopeptide-linked lysine or an analog thereof and / or (ii) the other non-canonical amino acid, or the analog thereof, into the protein in response to the one or more re-allocated codon using the orthogonal translation system.

[0042] 42. The method according to embodiment 41, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

[0043] 43. The method according to embodiment 42 or 43, wherein the cell expresses an oligopeptide permease.

[0044] 44. The method according to embodiment 43, wherein the oligopeptide permease comprises a peptide-binding protein that has been engineered for increased specificity for the compound according to any one of embodiments 1 to 16.

[0045] 45. The method according to embodiment 44, wherein the engineered peptide-binding protein is OppA of E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445 T429, 5460, V482, L531 and / or N533.

[0046] 46. The method according to any one of embodiments 41 to 45, wherein the orthogonal aminoacyl-tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair.

[0047] 47. The method according to any one of embodiments 41 to 46, wherein the nucleic acid molecule encoding the protein comprises one or more amber stop codons that have been re-allocated to incorporate the non-canonical amino acid or an analog thereof.

[0048] 48. The method according to any one of embodiments 41 to 47, wherein cleavage of the molecule according to any one of embodiments 1 to 16 inside the cell releases a second non-canonical amino acid (ncAA), and wherein the second ncAA is incorporated into the same or a different protein in response to a re-allocated codon.

[0049] 49. The method according to any one of embodiments 41 to 48, wherein the isopeptide- linked lysine or the analog thereof has the structure: wherein C' is -NH2 or -OH, and wherein D and E is defined as in any one of embodiments 1 to 16, wherein the left empty valence in D is connected to C'.

[0050] Herein the inventors leverage a propeptide-based strategy combined with the engineering of a bacterial ABC transporter system to boost intracellular ncAA concentrations and expression yields of modified proteins. The inventors found that easily synthesized isopeptide-linked tripeptides (G-XisoK) are actively transported into E. coli cells via the oligopeptide permease transporter Opp and are processed in the cytosol to reveal dipeptidic XisoK ncAAs. These isopeptide-linked lysine derivatives allow for the incorporation of various functionalities available for genetic code expansion, with efficiencies rivaling wild type (wt) protein expressions, providing a toolbox of novel ncAAs covering approaches from bioorthogonal labeling over chemical and photocrosslinking to native chemical ligation and chemoenzymatic conjugation. To exploit efficient G-XisoK-uptake for cheap and scalable production of modified proteins, the inventors evolved the periplasmic binding protein of the Opp transporter via a fluorescence-activated cell sorting platform, facilitating the preferential transport of G-XisoK tripeptides over linear tripeptides present in expression media. By introducing the evolved Opp binding protein variant into the E. coli genome, the inventors created a novel E. coli strain that allows high XisoK incorporation efficiencies across a range of target proteins for both single as well as multiple XisoK encoding. Furthermore, the inventors adapted the tripeptide scaffold for incorporation of two different ncAAs in response to two different nonsense codons via their concomitant transport using a single tripeptide, demonstrating the significant impact that active ncAA uptake has on the efficient production of proteins with an expanded alphabet.

[0051] In one aspect, the invention relates to a compound comprising an isopeptide-linked lysine or an analogue thereof. The skilled person understands that isopeptide-linked lysine, which also may be referred to as isopeptide-bond linked lysine, is a lysine forming an amide bond with its side chain amino group. To this end, it is to be understood that preferably said amide bond is a peptide bond, i.e., amide bond formed with a peptide. Furthermore, the lysine is otherwise preferably not modified, and its both main chain functional groups, i.e., carboxylic acid group and amino group, are present. Accordingly, a lysine residue forming an isopeptide bond, as referred to herein, is a group according to formula:

[0052] It is preferred that the lysine involved in formation of the isopeptide bond, as referred to herein, is L-lysine.

[0053] Thus, preferably, the compound comprising an isopeptide-linked lysine or an analogue thereof is defined as a peptide comprising at least two amino acid residues, and wherein the C- terminal carboxyl group of said C-terminal amino acid residue is modified as -CO-NH-E.

[0054] In said -CO-NH-E group, E is -(Ci-6 alkyleneJ-CHY^2, wherein Y1is selected from hydrogen, - SH, -OH and -NH2, preferably wherein Y1is hydrogen or -NH2, and wherein Y2is hydrogen, COOH or CONH2, provided that at least one of Y1and Y2is not hydrogen.

[0055] Accordingly, in one embodiment, Y1is -NH2and Y2is COOH. In one embodiment, Y1is hydrogen and Y2is COOH. In one embodiment, Y1is -NH2and Y2is hydrogen. In one embodiment, Y1is - OH and Y2is -COOH. In one embodiment, Y1is -NH2and Y2is -CONH2.

[0056] Said alkylene in -(Ci-6 alkyleneJ-CHY^2is further optionally substituted with a group selected from -OH, -SH and Hal. Said alkylene is optionally substituted with an -OH group. Preferably however, said alkylene is not substituted with an -OH group.

[0057] N=N

[0058] Furthermore, one -CH2- group in said alkylene may be replaced with Preferably

[0059] N = N however, no -CH2- group in said alkylene is replaced with

[0060] Preferably, said -(Ci-6 alkylene)- in E is preferably -(C3-4 alkylene)-. More preferably, said - (C1-6 alkylene)- in E is preferably -(C4 alkylene)-, such as -CH2CH2CH2CH2-.

[0061] Accordingly, particularly preferred E are selected from -CH2CH2CH2CH2-CH(-NH2)COOH, - CH2CH2CH2-CH(-NH2)COOH, -CH2CH2-CH(-NH2)COOH, -CH2-CH(-NH2)COOH - CH2CH2CH2CH2CH2COOH, -CH2CH2CH2CH2COOH, CH2CH2CH2CH2CH2NH2, -CH2CH2CH2CH2- CH(-OH)COOH, and -CH2CH2CH2CH2-CH(-NH2)CONH2. More preferably, E is -CH2CH2CH2CH2- CH(-NH2)-COOH.

[0062] Preferably, in E, the configuration on carbon atom that carries -COOH group and -NH2 group is the same as in natural L-lysine.

[0063] Preferably, in the compound comprising an isopeptide-linked lysine or an analogue thereof, the C-terminal amino acid residue of said peptide is not glycine. Furthermore, it is preferred that the peptide in the compound of the present invention is a dipeptide. It is further preferred that the dipeptide, as referred to herein, is not glycylglycine.

[0064] According to the present invention, the compound comprising an isopeptide-linked lysine or an analogue thereof, is a compound of formula (I): or a salt thereof.

[0065] It is to be understood that when referring to the compound comprising an isopeptide-linked lysine or an analogue thereof, which is a compound of formula (I), a reference to a compound of formula (I) is preferably made.

[0066] In formula (I), A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s al kyl)(Ci-s alkyl), - N+(CI-5alkyl)(Ci-5alkyl)(Ci-5alkyl), -NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2, or A is a peptidyl group. It is to be understood that if A is a charged group, a suitable counterion (or suitable counter-charge in the compound of formula (I)) is to be present.

[0067] Preferably, A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s a I ky l)(Ci-s alkyl), -NH-CO- (C1-5 alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2, or A is a peptidyl group.

[0068] More preferably, A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), - NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2.

[0069] Even more preferably, A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), -NH-CO-(CI-5alkyl), -COOH, -N3, and -NO2. Even more preferably, A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), -COOH, and N3.

[0070] Again more preferably, A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), and -N(Ci-s alkyl)(Ci- 5 alkyl).

[0071] Still more preferably, A is selected from -NHz, and -OH.

[0072] Most preferably, A is -NH2.

[0073] Alternatively, A is NH(CH3), -COO

[0074] In formula (I), B is selected from: that the left empty valence as depicted herein for B is connected to A, and that the right empty valence as depicted herein for B, is connected to C.

[0075] Z

[0076] Preferably, B is The stereogenic carbon atom carrying Z in B can have both possible

[0077] Z stereoconfigurations, as encompassed by the invention. However, preferably B is In other words, preferably, the configuration of the stereogenic carbon atom carrying Z in B is the same as the stereoconfiguration of an a-carbon atom in a natural L-amino acid residue. It is to be understood that, the left empty valence is connected to A and the right empty valence is connected to C.

[0078] In formula (I), Z is hydrogen or a side chain of an amino acid. The present invention does not limit Z to natural amino acid, and accordingly Z can be a side chain of a natural amino acid or a side chain of a nonnatural amino acid. Exemplary preferred side chains of natural and nonnatural amino acid are recited in the following.

[0079] Accordingly and preferably, Z is selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, - (C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, -(C1-6 alkylene)- S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO-Rz, -(C1-6 alkylene)-COO-Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO-(Ci-g alkyl), - (C1-6 alkylene)-CO-N(Rz)-O-Rz, -(Ci-6alkylene)-O-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-CO-N(Rz)- Rz, -(C1-6 alkylene)-N(Rz)-C(=N-Rz)-N(Rz)-Rz, -(C1-6 alkylene)-SO3-Rz, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z groups are each optionally substituted with one or more groups selected from -OH, -SH and Hal (preferably with one or N=N more -OH), further wherein one -CH2- group in said alkyl may be replaced with further wherein each Rzis independently selected from hydrogen and C1-6 alkyl and wherein Rzzis selected from hydrogen, C1-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-s alkenyl), -CO(Ci-s alkynyl), -CO(Ci-s alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)-carbocyclyl, -CO(Co-5 alkylene)-heterocyclyl, -CO-O-(Ci-s alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-s alkynyl), - CO-0-(Co-5 alkylene)-carbocyclyl, -CO-0-(Co-s alkylene)-heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs(preferably selected from hydrogen, C1-6 alkyl, aryl-CH2--O-CO- (such as benzyl-O-CO-) and CH2=CH-CH2-O-CO-, more preferably selected from hydrogen, and C1-6 alkyl).

[0080] More preferably, Z is selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-Hal, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-NH2, -(Ci-6 alkylene)-NH-CO-NH2, - (C1-6 alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, -(Ci-6alkylene)-CO-NH2, -(Ci-6 alkylene)-CO-NH-OH, -(Ci-6 alkyleneJ-SOsH, phenyl, -(Ci-6 alkylene)- phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)-heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)- cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH, further wherein one -CH2- group in said alkyl may be replaced

[0081] Even more preferably, Z is selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, phenyl, -(C1-6 alkylene)-phenyl, cycloalkyl, -(C1-6 alkylene)-cycloalkyl, heteroaryl, -(C1-6 alkylene)- heteroaryl, heterocycloalkyl, and -(C1-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(C1-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(C1-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(C1-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH.

[0082] It is particularly preferred that in the compound of formula (I), Z is selected from hydrogen, - (Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl). More preferably, Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and -CH2CH2CH2CH2-NH-CO-CH3.

[0083] Further preferred Z are listed in any one of specific embodiments of the compound of formula (I).

[0084] In formula (I), Z1and Z2are each independently selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, - (C1-6 alkylene)-S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO-Rz, -(C1-6 alkylene)-COO- Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-O-Rz, -(C1-6 alkylene)-O-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO- N(RZ)-RZ, -(C1-6 alkylene)-N(Rz)-C(=N-Rz)-N(Rz)-Rz, -(Ci-6alkylene)-SO3-Rz, -(C0-6alkylene)- carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH, -SH or Hal (preferably with one or more - OH), wherein each Rzis independently selected from hydrogen and C1-6 alkyl, and Rzzis selected from hydrogen, C1-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-s alkenyl), -CO(Ci-s alkynyl), - CO(Ci-5 alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)-carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -CO-O-(Ci-s alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-s alkynyl), -CO-0-(Co- 5 alkylene)-carbocyclyl, -CO-0-(Co-s alkylene)-heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs(preferably selected from hydrogen, C1-6 alkyl, aryl-CH2--O-CO- (such as benzyl-O-CO-) and CH2=CH-CH2-O-CO-, more preferably selected from hydrogen, and C1-6 alkyl), and further

[0085] N=N wherein one -CH2- group in said alkyl may be replaced provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen, or Z1and Z2are joined together to form, together with carbon atom(s) that otherwise carry Z1and Z2, a non-aromatic carbocyclic ring or non-aromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non-aromatic heterocyclic ring are each optionally substituted with one or more Rs.

[0086] Preferably, Z1and Z2are each independently selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, - (C1-6 alkylene)-S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO-Rz, -(C1-6 alkylene)-COO- Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-O-Rz, -(C1-6 alkylene)-O-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO- N(RZ)-RZ, -(C1-6 alkylene)-N(Rz)-C(=N-Rz)-N(Rz)-Rz, -(Ci-6alkylene)-SO3-Rz, -(C0-6alkylene)- carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH, -SH or Hal (preferably with one or more - OH), wherein each Rzis independently selected from hydrogen and C1-6 alkyl, and Rzzis selected from hydrogen, C1-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-s alkenyl), -CO(Ci-s alkynyl), - CO(Ci-5 alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)-carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -CO-O-(Ci-s alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-s alkynyl), -CO-0-(Co- 5 alkylene)-carbocyclyl, -CO-0-(Co-s alkylene)-heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs(preferably selected from hydrogen, C1-6 alkyl, aryl-CH2--O-CO- (such as benzyl-O-CO-) and CH2=CH-CH2-O-CO-, more preferably selected from hydrogen, and C1-6 alkyl), and further N=N wherein one -CH2- group in said alkyl may be replaced provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen.

[0087] More preferably, Z1and Z2are each independently selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-Hal, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-NH2, - (C1-6 alkylene)-NH-CO-NH2, -(Ci-6alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, -(Ci-6 alkylene)-CO-NH2, -(Ci-6 alkylene)-CO-NH-OH, -(Ci-6 alkylene)-SO3H, phenyl, -(Ci-6 alkylene)-phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)-heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)- heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)- heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH,

[0088] N = N further wherein one -CH2- group in said alkyl may be replaced with provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen.

[0089] Even more preferably, Z1and Z2are each independently selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, phenyl, -(C1-6 alkylene)-phenyl, cycloalkyl, -(C1-6 alkylene)-cycloalkyl, heteroaryl, -(C1-6 alkylene)-heteroaryl, heterocycloalkyl, and -(C1-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(C1-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(C1-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(C1-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said - (C1-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH, provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen.

[0090] Further preferred Z1and Z2groups are listed in any one of specific embodiments of the compound of formula (I).

[0091] In formula (I), -C-D- are defined as follows:

[0092] C is selected from -CO-NH-, -CO-O-, -CO-S- and -CO-N(Ci-s alkyl)-, wherein the left empty valence is connected to B and the right empty valence is connected to D, and connected to C and the right empty valence is connected to -CO-NH- (i.e., -CO-NH- connected to -E), or -C-D- taken together is -CO-(N-heterocycloalkylene)-, wherein said N-heterocycloalkylene is connected to said -CO- group of -C-D- through its nitrogen atom and wherein said N- heterocycloalkylene is optionally substituted with one or more Rs.

[0093] Preferably, C is selected from -CO-NH-, -CO-O-, -CO-S- and -CO-N(Ci-s alkyl)-, wherein the left empty valence is connected to B and the right empty valence is connected to D, and D is selected from wherein the left empty valence is connected to C and the right empty valence is connected to -CO-NH-.

[0094] Accordingly, and preferably, C is selected from -CO-NH- and -CO-O-, wherein the left empty valence is connected to B and the right empty valence is connected to D.

[0095] More preferably, C is -CO-NH-, wherein the left empty valence is connected to B and the right empty valence is connected to D.

[0096] Alternatively, C is -CO-O-, wherein the left empty valence is connected to B and the right empty valence is connected to D. In particular, it is preferred that if A is not -NH2, then C is - CO-O-, as described herein, i.e., wherein the left empty valence is connected to B and the right empty valence is connected to D.

[0097] In formula (I), preferably D i wherein the left empty valence is connected to C and the right empty valence is connected to -CO-NH- group. It is to be understood that both configurations at the stereogenic carbon atom carrying X are possible, as encompassed by the present invention. Preferably however, D is . In other words, preferably, the configuration of the stereogenic carbon atom carrying X in D is the same as the stereoconfiguration of an a-carbon atom in a natural L-amino acid residue.

[0098] In formula (I), X is hydrogen or a side chain of an amino acid. The present invention does not limit X to natural amino acid, and accordingly C can be a side chain of a natural amino acid or a side chain of a nonnatural amino acid. Exemplary preferred side chains of natural and nonnatural amino acid are recited in the following.

[0099] Preferably, X is selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rx, -(C1-6 alkylene)-S-Rx, -(C1-6 alkylene)-N(Rx)-Rx, -(C1-6 alkylene)-CO-Rx, -(C1-6 alkylene)-COO-Rx, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rx)-Rx, -(C1-6 alkylene)-N(Rx)-CO-(Ci-g alkyl), -(C1-6 alkylene)-CO- N(RX)-O-RX, -(C1-6 alkylene)-O-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(Ci-6 alkylene)-SO3-Rx, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X groups are each optionally substituted with one or more -OH, -SH or -Hal (preferably with one or more -OH), wherein each Rxis independently selected from hydrogen and Ci-6 alkyl, and further wherein one -CH2-

[0100] N = N group in said alkyl may be replaced with

[0101] More preferably, X is selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-NH2, -(C1-6 alkylene)-NH-CO-NH2, - (C1-6 alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, -(Ci-6alkylene)-CO-NH2, -(Ci-6 alkylene)-CO-NH-OH, -(Ci-6 alkyleneJ-SOsH, phenyl, -(Ci-6 alkylene)- phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)-heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)- cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH, further wherein one -CH2- group in said alkyl may be replaced

[0102] Even more preferably, X is selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, phenyl, -(Ci-6 alkylene)-phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)- heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH.

[0103] It is particularly preferred that in the compound of formula (I), X is selected from methyl, ethyl, isopropyl, sec-butyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl-diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4- yl)methyl and benzyl.

[0104] Further preferred X are listed in any one of specific embodiments of the compound of formula (I).

[0105] In formula (I), X1and X2are each independently selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rx, - (C1-6 alkylene)-S-Rx, -(C1-6 alkylene)-N(Rx)-Rx, -(C1-6 alkylene)-CO-Rx, -(C1-6 alkylene)-COO- Rx, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rx)-Rx, -(C1-6 alkylene)-N(Rx)-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rx)-O-Rx, -(C1-6 alkylene)-O-CO-N(Rx)-Rx, -(C1-6 alkylene)-N(Rx)-CO- N(RX)- Rx, -(C1-6 alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(C1-6 alkylene)-SO3-Rx, -(Co-6 alkylene)- carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X1and X2groups are each optionally substituted with one or more -OH, -SH or Hal (preferably with one or more - OH), wherein each Rxis independently selected from hydrogen and C1-6 alkyl, further wherein

[0106] N = N one -CH2- group in said alkyl may be replaced with provided that at least one of X1and X2is not hydrogen, further provided that if X1and X2are connected to the same carbon atom, then both X1and X2are not hydrogen, or X1and X2are joined together to form, together with carbon atom(s) that otherwise carry X1and X2, a non-aromatic carbocyclic ring or nonaromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non-aromatic heterocyclic ring are each optionally substituted with one or more Rs.

[0107] More preferably, X1and X2are each independently selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-Hal, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-NH2, - (C1-6 alkylene)-NH-CO-NH2, -(Ci-6alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, -(Ci-6 alkylene)-CO-NH2, -(Ci-6 alkylene)-CO-NH-OH, -(Ci-6 alkyleneJ-SOsH, phenyl, -(Ci-6 alkylene)-phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)-heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)- heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)- heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned X1and X2groups are each optionally substituted with one or more -OH, N = N further wherein one -CH2- group in said alkyl may be replaced with provided that at least one of X1and X2is not hydrogen, further provided that if X1and X2are connected to the same carbon atom, then both X1and X2are not hydrogen.

[0108] Even more preferably, X1and X2are each independently selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, phenyl, -(C1-6 alkylene)-phenyl, cycloalkyl, -(C1-6 alkylene)-cycloalkyl, heteroaryl, -(C1-6 alkylene)-heteroaryl, heterocycloalkyl, and -(C1-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(C1-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(C1-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(C1-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said - (C1-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, and further wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned X1and X2groups are each optionally substituted with one or more -OH, provided that at least one of X1and X2is not hydrogen, further provided that if X1and X2are connected to the same carbon atom, then both X1and X2are not hydrogen.

[0109] Further preferred X1and X2groups are listed in any one of specific embodiments of the compound of formula (I).

[0110] In formula (I), E is as defined herein, including any preferred definition of E and any specific embodiment of E.

[0111] In formula (I), each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-5 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-S(Ci-5 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-5 haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-5 alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH(Ci-5 alkylene)-CN, -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-(Ci-5 alkylene)-CN, -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-s alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-SO2-(Ci-s alkyl), -(C0-3 alkylene)-SO2-(Ci-5alkyl), -(C0-3 alkylene)-SO-(Ci-5alkyl), -OSO2F, -B(OH)2, -PO4H, -PO3H, -SO4H, -SO3H, -L-(CO-3 alkylene)-carbocyclyl, and -L-(Co-3 alkylene)-heterocyclyl, wherein L is selected from -(C0-3 alkylene)-NH-, -(C0-3 alkylene)-O-, -(C0-3 alkylene)-CO-, -(C0-3 alkylene)-NH-CO-, -(C0-3 alkylene)-CO-NH-, -(C0-3 alkylene)-NH-CO-O-, -(C0-3 alkylene)— O-CO-NH-, -(C0-3 alkylene)-NH- CO-NH- and -(C0-3 alkylene)-NH-CO-NH-, and wherein the carbocyclyl moiety in said -L-(Co-3 alkylene)-carbocyclyl and the heterocyclyl moiety in said -L-(Co-3 alkylene)-heterocyclyl are each optionally substituted with one or more groups independently selected from C1-4 alkyl, halogen, -CN, -NO2, -OH, -O-(Ci-4alkyl), -SH, -S-(Ci-4alkyl), -NH2, -NH(CI-4alkyl), -N(Ci-4alkyl)(Ci- 4 alkyl), -COOH, -COO(Ci-4alkyl), -CONH2, -CON H(CI-4alkyl), -CON(CI-4alkyl)(Ci-4alkyl), - NHCO(Ci-4 alkyl) and -N(Ci-4 alkyl)-CO(Ci-4 alkyl).

[0112] Preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-5 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-S(Ci-5 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-5 haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-5 alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-s alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s al ky l)-SO2-(Ci-s alkyl), -(C0-3 alkylene)-SO2-(Ci-5 alkyl), -(C0-3 alkylene)-SO-(Ci-s alkyl), -(C0-3 alkylene)-carbocyclyl, and -(C0-3 alkylene)-heterocyclyl, wherein the carbocyclyl moiety in said -(C0-3 alkylene)-carbocyclyl and the heterocyclyl moiety in said -(C0-3 alkylene)-heterocyclyl are each optionally substituted with one or more groups independently selected from C1-4 alkyl, halogen, -CN, -NO2, -OH, -O- (C1-4 alkyl), -SH, -S-(Ci-4alkyl), -NH2, -NH(CI-4alkyl), -N(CI-4alkyl)(Ci-4alkyl), -COOH, -COO(Ci-4alkyl), -CONH2, -CONH(CI-4alkyl), -CON(CI-4alkyl)(Ci-4alkyl), -N HCO(CI-4alkyl) and -N(CI-4alkyl)-CO(Ci-4 alkyl).

[0113] More preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-5 alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-s alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-SO2-(Ci-s alkyl), -(C0-3 alkylene)-SO2-(Ci-5 alkyl), -(C0-3 alkylene)-SO-(Ci-s alkyl), -(C0-3 alkylene)-carbocyclyl, and -(C0-3 alkylene)-heterocyclyl.

[0114] More preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-5 alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-s alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-SO2-(Ci-s alkyl), -(C0-3 alkylene)-SO2-(Ci-5 alkyl), and -(C0-3 alkylene)-SO-(Ci-s alkyl).

[0115] Even more preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-5 alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-s alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-SO2-(Ci-s alkyl), -(C0-3 alkylene)-SO2-(Ci-5 alkyl), and -(C0-3 alkylene)-SO-(Ci-s alkyl).

[0116] Still more preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-5 alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-O-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-5 alkyl), and -(C0-3 alkylene)-O-CO-N(Ci-s a I ky l)-(Ci-s alkyl).

[0117] Again more preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -OH, -O(Ci-s alkyl), -O(Ci-s alkylene)-OH, -O(Ci-s alkylene)-O(Ci-s alkyl), -SH, -S(Ci-s alkyl), -S(Ci-5alkylene)-SH, -S(Ci-s alkylene)-S(Ci-s alkyl), -NH2, -NH(Ci-s alkyl), -N(Ci-s a I kyl)(Ci-s alkyl), -NH-OH, -N(Ci-s alkyl)-O(Ci-s alkyl), halogen, -(C1-5 haloalkyl), -O-(Ci-s haloalkyl), -CN, -COOH, -CO-O-(Ci-5alkyl), -O-CO-(Ci-5alkyl), -CO-NH2, -CO-NH(CI-5alkyl), -CO-N(CI-5alkyl)(Ci-5alkyl), -NH-CO-(CI-5alkyl), -N(CI-5alkyl)-CO-(Ci-5alkyl), -NH-CO- O-(Ci-5alkyl), -N(CI-5alkyl)-CO-O-(Ci-5alkyl), -O-CO-NH-(CI-5alkyl), and -O-CO-N(CI-5alkyl)-(Ci-5alkyl).

[0118] Still more preferably, each Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -OH, -O(Ci-s alkyl), -O(Ci-s alkylene)-OH, -SH, -S(Ci-s alkyl), -S(Ci-s alkylene)-SH, -NH2, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), halogen, -(C1-5 haloalkyl), -O-(Ci-s haloalkyl), -CN, -COOH, -CO-O-(Ci-5alkyl), -O-CO-(Ci-5alkyl), -CO-NH2, -CO-NH(CI-5alkyl), -CO-N(CI-5 al ky l)(Ci-s alkyl), -NH-CO-(Ci-s alkyl), and -N(Ci-s alkyl)-CO-(Ci-s alkyl). Even more preferably, each Rsis independently selected from C1-5 alkyl, -OH, -O(Ci-s alkyl), -SH, -S(Ci-s alkyl), -NH2, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), halogen, -(C1-5 haloalkyl), -O-(Ci-s haloalkyl), -CN, and -COOH.

[0119] Again more preferably, each Rsis independently selected from -OH, -SH, -NH2, halogen, -CN, and -COOH.

[0120] Most preferably, each Rsis independently selected from -OH, halogen, and -CN.

[0121] In certain embodiments, when A is N3, Z and X cannot both be hydrogen.

[0122] Preferably, the compound according to formula (I), as provided herein, is a compound according to formula (II): or its salt. In formula (II), Z and X are as defined in formula (I).

[0123] As mentioned before, it is preferred that lysine residue engaged in formation of an isopeptide bond is characterized by L configuration at its stereogenic centre. Accordingly, the compound of formula (II) is preferably a compound of formula (Ila):

[0124] or its salt. In formula (Ila), Z and X are as defined in formula (I), including any preferred definition and any specific embodiment of formula (I).

[0125] In one preferred embodiment, the compound of formula (Ila) is preferably a compound of definition and any specific embodiment of formula (I).

[0126] In one preferred embodiment, the compound of formula (Ila) is preferably a compound of formula (lie):

[0127] or its salt. In formula (lie), Z and X are as defined in formula (I), including any preferred definition and any specific embodiment of formula (I).

[0128] Preferably, in formula (II) (or formula (II*), (Ila), ( 11 * a), (lib), (lib*), (lie) or ( 11 * c)), Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl). More preferably, Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and -CH2CH2CH2CH2-NH-CO-CH3. Even more preferably, Z is hydrogen.

[0129] Alternatively, in formula (II) (or formula (II*), (Ila), (I I *a), (lib), (lib*), (lie) or (I I *c)), Z is selected from any one of:

[0130] and hydrogen.

[0131] Preferably, in formula (II) (or formula (Ila), (lib) or (lie)), X is selected from methyl, ethyl, isopropyl, sec-butyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl-diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4- yl)methyl and benzyl.

[0132] In one embodiment of the present invention, the compound of formula (I) is a compound of formula (III): or its salt. In formula (III), Rpis hydrogen or -OH, preferably Rpis hydrogen. Further in formula

[0133] (III), Z is as defined for formula (I), including any preferred definition and any specific embodiment of formula (I).

[0134] Preferably, the compound of formula (III) is a compound of formula (Illa): or its salt, or its salt. In formula (Illa), Rpis hydrogen or -OH, preferably Rpis hydrogen. Further in formula (Illa), Z is as defined for formula (I), including any preferred definition and any specific embodiment of formula (I).

[0135] Preferably, the compound of formula (Illa) is a compound of formula (lllb): or its salt. In formula (lllb), Rpis hydrogen or -OH, preferably Rpis hydrogen. Further in formula (lllb), Z is as defined for formula (I), including any preferred definition and any specific embodiment of formula (I).

[0136] Preferably, in formula (III) (including formula (Illa) and (lllb), Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl). More preferably, Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and -CH2CH2CH2CH2-NH-CO-CH3, even more preferably Z is hydrogen.

[0137] Alternatively, in formula (III) (including formula (Illa) and ( 111 b), Z is selected from any one of: and hydrogen. Particularly preferred compounds of formula (I) are listed in Table 1. Accordingly, it is preferred that the compound of formula (I) is a compound selected from the compounds listed in Table 1, or its salt. Table 1: Preferred compounds according to the invention. -X*- in any of the following compounds is -O- or -NH-. Accordingly, a formula with -X*- shown in the following Table refers specifically and individually both to a compound bearing -X*- = -O- and to a compound bearing -X*- = -NH-.

[0138]

[0139] In the following specific embodiments, preferred definitions of the amino acid side chain are defined. Accordingly, the expression "any one of Z, Z1, Z2, X, X1and X2" is to be understood as encompassing individualized reference to each and every of variables any one of Z, Z1, Z2, X, X1and X2, as referred to herein, as well as any combinations thereof. These variables are descriptive of amino acid sidechains in the compound of formula (I). Thus, as provided herein, any one of Z, Z1, Z2, X, X1and X2may refer to Z. Alternatively, any one of Z, Z1, Z2, X, X1and X2may refer to Z1, Z2or Z1and Z2. Alternatively, any one of Z, Z1, Z2, X, X1and X2may refer to X. Alternatively, any one of Z, Z1, Z2, X, X1and X2may refer to X1, X2or X1and X2. In other words, the expression "any one of Z, Z1, Z2, X, X1and X2" may also be understood as referring to an amino acid side chain, when discussing any amino acid side chain in the compound of the present invention. Accordingly and preferably, the expression "any one of Z, Z1, Z2, X, X1and X2", includes a specific and individual reference to Z, and a specific and individual reference to X.

[0140] In a first specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is Ci-6 alkyl. Particularly preferred Ci-6 alkyl groups are methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, and tert-butyl. Thus, in this embodiment any one of Z, Z1, Z2, X, X1and X2can be selected from methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, and tertbutyl. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from methyl, isopropyl, isobutyl, and sec-butyl. Even more preferably, any one of Z, Z1, Z2, X, X1and X2is isopropyl.

[0141] In a second specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is selected from phenyl and -(Ci-6 alkylene)-phenyl, preferably any one of Z, Z1, Z2, X, X1and X2is selected from phenyl, benzyl, and phenethyl. More preferably, any one of Z, Z1, Z2, X, X1and X2is benzyl.

[0142] In a third specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is a side chain of an amino acid, preferably of an a-amino acid (particularly a naturally occurring a-amino acid, including a proteinogenic a-amino acid or a non-proteinogenic a-amino acid). Examples of any one of Z, Z1, Z2, X, X1and X2being a side chain of an amino acid are illustrated in the following table.

[0143] Table 2: Examples of a side chain of an amino acid which can be any one of L, Z1, Z2, X, X1and X2.

[0144] In a fourth specific embodiment, any one of Z, Z1, Z2, X, X1and X2is selected from hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n-pentyl, n-hexyl, - CH2-CH=CH-CH3, -CHZ-C CH, phenyl, 4-hydroxyphenyl, 3,5-dihydroxyphenyl, 3,5-dichloro-4- hydroxyphenyl, -CHz-phenyl, -CH2-(4-hydroxyphenyl), -CH2-(3-hydroxyphenyl), -CH2-(3,4- dihydroxyphenyl), -CH2-(3-nitro-4-hydroxyphenyl), -CH2-(4-methoxyphenyl), -CH2-(3-hydroxy- 4-methoxy-5-methylphenyl), -CH2-(4-prenyloxy phenyl), -CH2-(4-(N,N-dimethylamino)phenyl), -CH(-OH)-phenyl, -CH(-OH)-(4-hydroxyphenyl), -CH(-OH)-(4-methoxyphenyl), -CH(-OH)-(4- aminophenyl), -CH(-OH)-(4-nitrophenyl), -CH(-CH3)-phenyl, -CH2CH2-phenyl, -CH2CH2-(4- hydroxyphenyl), -CH(-OH)-CH(-OH)-(4-hydroxyphenyl), -CH(-OH)-CH2-(4-nitrophenyl), -CH2- (lH-indol-3-yl), -CH(-CH3)-(lH-indol-3-yl), -CH(-OH)-(lH-indol-3-yl), -CH2-(5-hydroxy-lH-indol- 3-yl), -CH2-(l-acetyl-6-hydroxy-lH-indol-3-yl), -CH2-( l-(l,l-dimethy lai lyl)-lH-indol-3-yl), -CH(- CH3)-(2-oxo-indolin-3-yl), -CH2-(4-nitro-lH-indol-3-yl), -CH2-(5-nitro-lH-indol-3-yl), -CH2-(5- chloro-lH-indol-3-yl), -CH2-(6-chloro-lH-indol-3-yl), -CH2-(6-bromo-lH-indol-3-yl), -CH2-(7- fluoro-lH-indol-3-yl), -CH2-(7-chloro-lH-indol-3-yl), -CH2-(7-bromo-lH-indol-3-yl), -CH2-(6,7- dichloro-lH-indol-3-yl), -CH2-(lH-imidazol-4-yl), -CH(-OH)-(lH-imidazol-4-yl), -CH2-(5- mercapto-l-methyl-lH-imidazol-4-yl), -CH2-(3-furanyl), -CH2-(3-pyridinyl), -CH(-CH3)-CH(-OH)- (5-hydroxypyridin-2-yl), pyrrolidin-2-yl, -CH2-(2-imino-4-imidazolidinyl), 2- iminohexahydropyrimidin-4-yl, 1-methylcycloprop-l-yl, -CH2-(2-nitrocycloprop-l-yl), -CH2- NH2, -C(-OH)-(CH2)2-NH-C(=NH)-NH2, -CH2-C(-OH)-CH2-NH-C(=NH)-NH2, -C(-OH)-C(-OH)-CH2- NH-C(=NH)-NH2, -(CH2)4-NH-C(=NH)-NH2, -C(-OH)-(CH2)3-NH-C(=NH)-NH2, -CH2-C(-

[0145] OH)-CH2CH2-NH-C(=NH)-NH2, and -CH(-OH)-CH(-OH)-CH2-O-CO-NH2.

[0146] In a fifth specific embodiment, any one of L, Z1, Z2, X, X1and X2is selected from hydrogen, methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n-pentyl, n-hexyl, - CH2-CH=CH-CH3, -CH2-C CH, phenyl, 4-hydroxyphenyl, 3,5-dihydroxyphenyl, 3,5-dichloro-4- hydroxyphenyl, -CH2-phenyl, -CH2-(4-hydroxyphenyl), -CH2-(3-hydroxyphenyl), -CH2-(3,4- dihydroxyphenyl), -CH2-(3-nitro-4-hydroxyphenyl), -CH2-(4-methoxyphenyl), -CH2-(3-hydroxy- 4-methoxy-5-methylphenyl), -CH2-(4-prenyloxy phenyl), -CH2-(4-(N,N-dimethylamino)phenyl), -CH(-OH)-phenyl, -CH(-OH)-(4-hydroxyphenyl), -CH(-OH)-(4-methoxyphenyl), -CH(-OH)-(4- aminophenyl), -CH(-OH)-(4-nitrophenyl), -CH(-CH3)-phenyl, -CH2CH2-phenyl, -CH2CH2-(4- hydroxyphenyl), -CH(-OH)-CH(-OH)-(4-hydroxyphenyl), -CH(-OH)-CH2-(4-nitrophenyl), -CH2- (lH-indol-3-yl), -CH(-CH3)-(lH-indol-3-yl), -CH(-OH)-(lH-indol-3-yl), -CH2-(5-hydroxy-lH-indol- 3-yl), -CH2-(l-acetyl-6-hydroxy-lH-indol-3-yl), -CH2-(l-(l,l-dimethylallyl)-lH-indol-3-yl), -CH(- CH3)-(2-oxo-indolin-3-yl), -CH2-(4-nitro-lH-indol-3-yl), -CH2-(5-nitro-lH-indol-3-yl), -CH2-(5- chloro-lH-indol-3-yl), -CH2-(6-chloro-lH-indol-3-yl), -CH2-(6-bromo-lH-indol-3-yl), -CH2-(7- fluoro-lH-indol-3-yl), -CH2-(7-chloro-lH-indol-3-yl), -CH2-(7-bromo-lH-indol-3-yl), -CH2-(6,7- dichloro-lH-indol-3-yl), -CH2-(lH-imidazol-4-yl), -CH(-OH)-(lH-imidazol-4-yl), -CH2-(5- mercapto-l-methyl-lH-imidazol-4-yl), -CH2-(3-furanyl), -CH2-(3-pyridinyl), -CH(-CH3)-CH(-OH)- (5-hydroxypyridin-2-yl), pyrrolidin-2-yl, -CH2-(2-imino-4-imidazolidinyl), 2- iminohexahydropyrimidin-4-yl, 1-methylcycloprop-l-yl, and -CH2-(2-nitrocycloprop-l-yl).

[0147] In a sixth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is hydrogen, which corresponds to a side chain of glycine.

[0148] In a seventh specific embodiment, any one of Z, Z1, Z2, X, X1and X2is selected from Ci-8 alkyl. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, sec-butyl, tert-butyl, n-pentyl, and n-hexyl. In certain embodiments of the present invention, any one of Z, Z1, Z2, X, X1and X2is methyl, which corresponds to a side chain of alanine. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is ethyl, which corresponds to a side chain of homoalanine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is n-propyl, which corresponds to a side chain of norvaline. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is isopropyl, which corresponds to a side chain of valine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is n-butyl, which corresponds to a side chain of norleucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is isobutyl, which corresponds to a side chain of leucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is sec-butyl, which corresponds to a side chain of isoleucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is tert-butyl, which corresponds to a side chain of terleucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is n-pentyl, which corresponds to a side chain of homonorleucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is n-hexyl, which corresponds to a side chain of 2-aminocaprylic acid.

[0149] In an eighth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is selected from C2-8 alkenyl and C2-8 alkynyl. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from -CH2-CH=CH-CH3 and -CH2-C CH. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-CH=CH-CH3, which corresponds to a side chain of 2-amino-4-hexenoic acid. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-C CH, which corresponds to a side chain of propargylglycine.

[0150] In a ninth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is phenyl or -(C1-6 alkylene)-phenyl, wherein said phenyl and the phenyl group in said -(C1-6 alkylene)-phenyl are each optionally substituted with one or more groups Rs. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from phenyl, 4-hydroxyphenyl, 3,5-dihydroxyphenyl, 3,5-dichloro-4-hydroxyphenyl, -CH2-phenyl, -CH2-(4-hydroxyphenyl), -CH2-(3-hydroxyphenyl), -CH2-(3,4-dihydroxyphenyl), -CH2-(3-nitro-4-hydroxyphenyl), -CH2-(4-methoxyphenyl), -CH2- (3-hydroxy-4-methoxy-5-methylphenyl), -CH2-(4-prenyloxyphenyl), -CH2-(4-(N,N- dimethylamino)phenyl), -CH(-OH)-phenyl, -CH(-OH)-(4-hydroxyphenyl), -CH(-OH)-(4- methoxyphenyl), -CH(-OH)-(4-aminophenyl), -CH(-OH)-(4-nitrophenyl), -CH(-CH3)- phenyl, -CH2CH2-phenyl, -CH2CH2-(4-hydroxyphenyl), -CH(-OH)-CH(-OH)-(4-hydroxyphenyl), and -CH(-OH)-CH2-(4-nitrophenyl).

[0151] In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is phenyl, which corresponds to a side chain of phenylglycine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is 4-hydroxyphenyl, which corresponds to a side chain of 4-hydroxyphenylglycine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is 3,5-dihydroxyphenyl, which corresponds to a side chain of 3,5-dihydroxyphenylglycine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-phenyl, which corresponds to a side chain of phenylalanine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(4-hydroxyphenyl), which corresponds to a side chain of tyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is - CH2-(3,4-dihydroxyphenyl), which corresponds to a side chain of DOPA. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(3-nitro-4- hydroxyphenyl), which corresponds to a side chain of 3-nitrotyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(4- methoxyphenyl) which corresponds to a side chain of O-methyltyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(3-hydroxy-4- methoxy-5-methylphenyl) which corresponds to a side chain of 3-OH-5-Me-OMe-tyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(4-prenyloxyphenyl) which corresponds to a side chain of O-prenyltyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2- phenyl which corresponds to a side chain of homophenylalanine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2-(4- hydroxyphenyl) which corresponds to a side chain of homotyrosine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-OH)-CH2-(4- nitrophenyl) which corresponds to a side chain of P-OH-p-NO2-homophenylalanine.

[0152] In a tenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(Co-6 alkylene)-heterocyclyl, wherein said heterocyclyl is optionally substituted with one or more groups Rs. Preferably, the heterocyclyl in said -(Co-6 alkylene)-heterocyclyl is lH-indol-3- yl. Further preferably, any one of Z, Z1, Z2, X, X1and X2is selected from -CH2-(lH-indol-3-yl), - CH(-CH3)-(lH-indol-3-yl), -CH(-OH)-(lH-indol-3-yl), -CH2-(5-hydroxy-lH-indol-3-yl), -CH2-(1- acetyl-6-hydroxy-lH-indol-3-yl), -CH2-(l-(l,l-dimethylallyl)-lH-indol-3-yl), -CH(-CH3)-(2-oxo- indolin-3-yl), -CH2-(4-nitro-lH-indol-3-yl), -CH2-(5-nitro-lH-indol-3-yl), -CH2-(5-chloro-lH- indol-3-yl), -CH2-(6-chloro-lH-indol-3-yl), -CH2-(6-bromo-lH-indol-3-yl), -CH2-(7-fluoro-lH- indol-3-yl), -CH2-(7-chloro-lH-indol-3-yl), -CH2-(7-bromo-lH-indol-3-yl), and -CH2-(6,7- dichloro-lH-indol-3-yl). Thus, in a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(lH-indol-3-yl) which corresponds to a side chain of tryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-CH3)-(lH-indol-3-yl) which corresponds to a side chain of (3-methyl-tryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is - CH2-(5-hydroxy-lH-indol-3-yl) which corresponds to a side chain of 5-hydroxytryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is - CH2-(l-acetyl-6-hydroxy-lH-indol-3-yl) which corresponds to a side chain of 6-hydroxy- N-acetyltryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-( l-(l,l-dimethy la I ly l)-lH-indol-3-yl) which corresponds to a side chain of N-dimethylallyltryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-CH3)-(2-oxo-indolin-3-yl) which corresponds to a side chain of oxo-(3-methyl-tryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(4-nitro-lH-indol-3-yl) or -CH2-(5-nitro-lH- indol-3-yl) which corresponds to a side chain of nitrated tryptophan. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(5-chloro-lH- indol-3-yl), -CH2-(6-chloro-lH-indol-3-yl), -CH2-(6-bromo-lH-indol-3-yl), -CH2-(7-fluoro-lH- indol-3-yl), -CH2-(7-chloro-lH-indol-3-yl), or -CH2-(7-bromo-lH-indol-3-yl), which correspond to a side chain of halogenated tryptophan.

[0153] In an eleventh specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(Co-6 alkylene)-heterocyclyl, wherein said heterocyclyl is optionally substituted with one or more groups Rs, any one of Z, Z1, Z2, X, X1and X2is preferably selected from -CH2-(lH-imidazol- 4-yl), -CH(-OH)-(lH-imidazol-4-yl), -CH2-(5-mercapto-l-methyl-lH-imidazol-4-yl), -CH2-(3- furanyl), -CH2-(3-pyridinyl), -CH(-CH3)-CH(-OH)-(5-hydroxypyridin-2-yl), pyrrolidin-2-yl, -CH2- (2-imino-4-imidazolidinyl), and 2-iminohexahydropyrimidin-4-yl. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(lH-imidazol-4-yl) which corresponds to a side chain of histidine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(5-mercapto-l-methyl-lH-imidazol-4-yl) which corresponds to a side chain of ovothiol. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-(3-furanyl) which corresponds to a side chain of 3-(3-furyl)-alanine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is pyrrolidin-2-yl which corresponds to a side chain of tambroline.

[0154] In a twelfth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is cycloalkyl or -(C1-6 alkylene)-cycloalkyl, wherein said cycloalkyl and the cycloalkyl group in said -(C1-6 alkylene)-cycloalkyl are each optionally substituted with one or more groups Rs. Preferably, any one of Z, Z1, Z2, X, X1and X2is cycloalkyl, preferably a cyclopropyl, which may be optionally substituted with one or more groups Rs. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from 1-methylcycloprop-l-yl, and -CH2-(2-nitrocycloprop-l-yl).

[0155] In a thirteenth specific embodiment, any one of Z, Z1, Z2, X, X1and X2is selected from C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-OH and -(C1-6 alkylene)-O-( C1-6 alkyl), wherein said alkyl, said alkenyl, said alkynyl, and said alkylene group are each optionally substituted with one or more -OH. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from -CH2- OH, -CH2CH2-OH, -CH2-O-CH3, -CH2CH2-O-CH3, -CH2CH2-O-CH2CH3, -CH(-CH3)-OH, -C(- CH3)(-OH)-CH3, -CH(-OH)-CH(-CH3)-CH3, -CH(-OH)-CH(-CH3)-CH2-CH=CH-CH3, and -CH(-OH)- C CH. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -C(-CH3)(-OH)-CH3, which corresponds to a side chain of (3-hydroxyvaline. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-OH)- CH(-CH3)-CH3, which corresponds to a side chain of (3-hydroxyleucine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-OH)-C CH, which corresponds to a side chain of (3-ethynylserine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-OH which corresponds to a side chain of serine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2-OH which corresponds to a side chain of homoserine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2- O-CH3 which corresponds to a side chain of O-methyl-homoserine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2-O-CH2CH3 which corresponds to a side chain of O-ethyl-homoserine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH(-CH3)-OH which corresponds to a side chain of threonine.

[0156] In a fourteenth specific embodiment, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-SH or -(C1-6 alkylene)-S-(Ci-6 alkyl). Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from -CH2- SH, -CH2CH2-SH, -CH2CH2-S-CH3, and -CH2CH2-S-CH2CH3. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-SH which corresponds to a side chain of cysteine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2-SH which corresponds to a side chain of homocysteine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2- S-CH3 which corresponds to a side chain of methionine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2CH2-S-CH2CH3 which corresponds to a side chain of ethionine.

[0157] In a fifteenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkyleneJ-SOsH or -(C1-6 alkylene)-SO3-(Ci-6 alkyl). Preferably, any one of Z, Z1, Z2, X, X1and X2is -CH2-SO3H. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-SO3H which corresponds to a side chain of cysteate.

[0158] In a sixteenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-COOH or -(C1-6 alkylene)-COO-(Ci-6 alkyl), wherein said alkyl, and said alkylene group are each optionally substituted with one or more -OH. Preferably, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-COOH. Preferably, in certain embodiments any one of Z, Z1, Z2, X, X1and X2is selected from -CH2-COOH, -CH2CH2-COOH, -CH(-OH)-COOH, -CH(-OH)- CH2-COOH, -CH(-CH3)-COOH, -CH(-CH3)-CH2-COOH. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -CH2-COOH, which corresponds to a side chain of aspartate.

[0159] In a seventeenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-CO-NH2, -(C1-6 alkylene)-CO-N(Ci-6 alkyl)H or -(C1-6 alkylene)-CO-N(Ci-6 alkyl)-(Ci-e alkyl), wherein said alkyl, and said alkylene group are each optionally substituted with one or more -OH. Preferably, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-CO-NH2. Preferably, in certain embodiments any one of Z, Z1, Z2, X, X1and X2is selected from -CH2-CO- NH2, -CH2CH2-CO-NH2, -CH(-OH)-CO-NH2, and -CH2-NH-CO-NH2.

[0160] In an eighteenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is selected from -(CH2)2-NH-CO-NH2, -(CH2)3-NH-CO-NH2, and -(CH2)2-CO-NH-OH.

[0161] In a nineteenth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-NH2, -(C1-6 alkylene)-N(Ci-6 al kyl)-H or -(C1-6 alkylene)-N(Ci-6 a I kyl)-(Ci-6 alkyl), and wherein said alkylene group is optionally substituted with one or more -OH. Preferably, any one of Z, Z1, Z2, X, X1and X2is -(C1-6 alkylene)-NH2. Preferably, any one of Z, Z1, Z2, X, X1and X2is selected from -CH2-NH2, -(CH2)2-NH2, -(CH2)3-NH2, -(CH2)4-NH2, -CH(- CH3)-NH2, -CH(-CH3)-(CH2)2-NH2, -C(-OH)-(CH2)3-NH2, -CH2-C(-OH)-(CH2)2-NH2, and -(CH2)2-C(- OH)-C(-CH2-OH)-NH2. In a particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(CH2)3-NH2, which corresponds to the side chain of ornithine. In a further particular embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is -(CH2)4- NH2, which corresponds to the side chain of lysine.

[0162] In a twentieth specific embodiment of the present invention, any one of Z, Z1, Z2, X, X1and X2is selected from -(CH2)3-NH-C(=NH)-NH2, -CH(-CH3)-(CH2)2-NH-C(=NH)-NH2, -C(-OH)-(CH2)2-NH- C(=NH)-NH2, -CH2-C(-OH)-CH2-NH-C(=NH)-NH2, -C(-OH)-C(-OH)-CH2-NH-C(=NH)-NH2, -(CH2)4- NH-C(=NH)-NH2, -C(-OH)-(CH2)3-NH-C(=NH)-NH2, and -CH2-C(-OH)-CH2CH2-NH-C(=NH)-NH2. In a particular embodiment, any one of Z, Z1, Z2, X, X1and X2is -(CH2)3-NH-C(=NH)-NH2, which corresponds to the side chain of arginine.

[0163] In a twenty-first specific embodiment, any one of Z, Z1, Z2, X, X1and X2is -CH(-OH)-CH(-OH)- CH2-O-CO-NH2.

[0164] In a twenty-second specific embodiment, any one of Z, Z1, Z2, X, X1and X2is C1-8 alkyl, N = N wherein one -CH2- group in said alkyl is replaced with Preferably, in this specific embodiment, any one of Z, Z1, Z2, X, X1and X2is (3-methyl-diazirin-3-yl)methyl. In a twenty-third specific embodiment, X, which can be a side chain of a natural or a nonnatural amino acid, is selected from any one of the following groups:

[0165]

[0166] As understood herein, the wavy bond indicates the point of attachment of the monovalent group.

[0167] In a twenty-fourth specific embodiment, Z, which can be a side chain of a natural or a non- natural amino acid, is selected from any one of the following groups:

[0168]

[0169] In this twenty fourth specific embodiment, Z may also, alternatively, be selected from :

[0170] 5 and hydrogen In a twenty-fifth specific embodiment of the compound of formula (I), B is selected from: wherein the left empty valence is connected to A and the right empty valence is connected to C.

[0171] In a twenty sixth specific embodiment of the compound of formula (I), -C-D- taken together is -CO-(N-heterocycloalkylene)-, wherein said N-heterocycloalkylene is connected to said -CO- group of -C-D- through its nitrogen atom and wherein said N-heterocycloalkylene is optionally substituted with one or more Rs. Preferably, in this specific embodiment, -C-D- is a moiety according to formula wherein the left empty valence is connected to B, and the right empty valence is connected to -CO-NH- group.

[0172] In a twenty seventh specific embodiment of the compound of formula (I), X is -(C1-6 alkylene)- Hal, in particular -(C1-6 alkylene)-Br or -(C1-6 alkylene)-CI.

[0173] The following definitions apply throughout the present specification and the claims, unless specifically indicated otherwise.

[0174] The term "hydrocarbon group" refers to a group consisting of carbon atoms and hydrogen atoms.

[0175] As used herein, the term "alkyl" refers to a monovalent saturated acyclic (i.e., non-cyclic) hydrocarbon group which may be linear or branched. Accordingly, an "alkyl" group does not comprise any carbon-to-carbon double bond or any carbon-to-carbon triple bond. A "Ci- 6 alkyl" denotes an alkyl group having 1 to 6 carbon atoms. Preferred exemplary alkyl groups are methyl, ethyl, propyl (e.g., n-propyl or isopropyl), or butyl (e.g., n-butyl, isobutyl, sec-butyl, or tert-butyl). Unless defined otherwise, the term "alkyl" preferably refers to C1-4 alkyl, more preferably to methyl or ethyl, and even more preferably to methyl.

[0176] As used herein, the term "alkenyl" refers to a monovalent unsaturated acyclic hydrocarbon group which may be linear or branched and comprises one or more (e.g., one or two) carbon- to-carbon double bonds while it does not comprise any carbon-to-carbon triple bond. The term "C2-6 alkenyl" denotes an alkenyl group having 2 to 6 carbon atoms. Preferred exemplary alkenyl groups are ethenyl, propenyl (e.g., prop-l-en-l-yl, prop-l-en-2-yl, or prop-2-en-l-yl), butenyl, butadienyl (e.g., buta-l,3-dien-l-yl or buta-l,3-dien-2-yl), pentenyl, or pentadienyl (e.g., isoprenyl). Unless defined otherwise, the term "alkenyl" preferably refers to C2-4 alkenyl.

[0177] As used herein, the term "alkynyl" refers to a monovalent unsaturated acyclic hydrocarbon group which may be linear or branched and comprises one or more (e.g., one or two) carbon- to-carbon triple bonds and optionally one or more (e.g., one or two) carbon-to-carbon double bonds. The term "C2-6 alkynyl" denotes an alkynyl group having 2 to 6 carbon atoms. Preferred exemplary alkynyl groups are ethynyl, propynyl (e.g., propargyl), or butynyl. Unless defined otherwise, the term "alkynyl" preferably refers to C2-4 alkynyl.

[0178] As used herein, the term "alkylene" refers to an alkanediyl group, i.e. a divalent saturated acyclic hydrocarbon group which may be linear or branched. A "C1-6 alkylene" denotes an alkylene group having 1 to 6 carbon atoms, and the term "Co-6 alkylene" indicates that a covalent bond (corresponding to the option "Co alkylene") or a C1-6 alkylene is present. Preferred exemplary alkylene groups are methylene (-CH2-), ethylene (e.g., -CH2-CH2- or -CH(-CH3)-), propylene (e.g., -CH2-CH2-CH2-, -CH(-CH2-CH3)-, -CH2-CH(-CH3)-, or -CH(-CH3)- CH2-), or butylene (e.g., -CH2-CH2-CH2-CH2-). Unless defined otherwise, the term "alkylene" preferably refers to C1-4 alkylene (including, in particular, linear C1-4 alkylene), more preferably to methylene or ethylene, and even more preferably to methylene.

[0179] As used herein, the term "carbocyclyl" refers to a hydrocarbon ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings), wherein said ring group may be saturated, partially unsaturated (i.e., unsaturated but not aromatic) or aromatic. Unless defined otherwise, "carbocyclyl" preferably refers to aryl, cycloalkenyl or cycloalkyl.

[0180] As used herein, the term "heterocyclyl" refers to a ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings), wherein said ring group comprises one or more (such as, e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, and the remaining ring atoms are carbon atoms, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) may optionally be oxidized, wherein one or more carbon ring atoms may optionally be oxidized (i.e., to form an oxo group), and further wherein said ring group may be saturated, partially unsaturated (i.e., unsaturated but not aromatic) or aromatic. For example, each heteroatom-containing ring comprised in said ring group may contain one or two O atoms and / or one or two S atoms (which may optionally be oxidized) and / or one, two, three or four N atoms (which may optionally be oxidized), provided that the total number of heteroatoms in the corresponding heteroatom-containing ring is 1 to 4 and that there is at least one carbon ring atom (which may optionally be oxidized) in the corresponding heteroatom-containing ring. Unless defined otherwise, "heterocyclyl" preferably refers to heteroaryl, heterocycloalkenyl or heterocycloalkyl.

[0181] As used herein, the term "aryl" refers to an aromatic hydrocarbon ring group, including monocyclic aromatic rings as well as bridged ring and / orfused ring systems containing at least one aromatic ring (e.g., ring systems composed of two or three fused rings, wherein at least one of these fused rings is aromatic; or bridged ring systems composed of two or three rings, wherein at least one of these bridged rings is aromatic). "Aryl" may, e.g., refer to phenyl, naphthyl, dialinyl (i.e., 1,2-dihydronaphthyl), tetralinyl (i.e., 1,2,3,4-tetrahydronaphthyl), indanyl, indenyl (e.g., lH-indenyl), anthracenyl, phenanthrenyl, 9H-fluorenyl, or azulenyl. Unless defined otherwise, an "aryl" preferably has 6 to 14 ring atoms, more preferably 6 to 10 ring atoms, even more preferably refers to phenyl or naphthyl, and most preferably refers to phenyl.

[0182] As used herein, the term "heteroaryl" refers to an aromatic ring group, including monocyclic aromatic rings as well as bridged ring and / or fused ring systems containing at least one aromatic ring (e.g., ring systems composed of two or three fused rings, wherein at least one of these fused rings is aromatic; or bridged ring systems composed of two or three rings, wherein at least one of these bridged rings is aromatic), wherein said aromatic ring group comprises one or more (such as, e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, and the remaining ring atoms are carbon atoms, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) may optionally be oxidized, and further wherein one or more carbon ring atoms may optionally be oxidized (i.e., to form an oxo group). For example, each heteroatom-containing ring comprised in said aromatic ring group may contain one or two O atoms and / or one or two S atoms (which may optionally be oxidized) and / or one, two, three orfour N atoms (which may optionally be oxidized), provided that the total number of heteroatoms in the corresponding heteroatom-containing ring is 1 to 4 and that there is at least one carbon ring atom (which may optionally be oxidized) in the corresponding heteroatom-containing ring. "Heteroaryl" may, e.g., refer to thienyl (i.e., thiophenyl), benzo[b]thienyl, naphtho[2,3- b]thienyl, thianthrenyl, furyl (i.e., furanyl), benzofuranyl, isobenzofuranyl, chromanyl, chromenyl (e.g., 2H-l-benzopyranyl or 4H-l-benzopyranyl), isochromenyl (e.g., 1H-2- benzopyranyl), chromonyl, xanthenyl, phenoxathiinyl, pyrrolyl (e.g., lH-pyrrolyl), imidazolyl, pyrazolyl, pyridyl (i.e., pyridinyl; e.g., 2-pyridyl, 3-pyridyl, or 4-pyridyl), pyrazinyl, pyrimidinyl, pyridazinyl, indolyl (e.g., 3H-indolyl), isoindolyl, indazolyl, indolizinyl, purinyl, quinolyl, isoquinolyl, phthalazinyl, naphthyridinyl, quinoxalinyl, cinnolinyl, pteridinyl, carbazolyl, (3-carbolinyl, phenanthridinyl, acridinyl, perimidinyl, phenanthrolinyl (e.g., [l,10]phenanthrolinyl, [l,7]phenanthrolinyl, or [4,7]phenanthrolinyl), phenazinyl, thiazolyl, isothiazolyl, phenothiazinyl, oxazolyl, isoxazolyl, oxadiazolyl (e.g., 1,2,4-oxadiazolyl, 1,2,5- oxadiazolyl (i.e., furazanyl), or 1,3,4-oxadiazolyl), thiadiazolyl (e.g., 1,2,4-thiadiazolyl, 1,2,5- thiadiazolyl, or 1,3,4-thiadiazolyl), phenoxazinyl, pyrazolo[l,5-a]pyrimidinyl (e.g., pyrazolo[l,5-a]pyrimidin-3-yl), l,2-benzoisoxazol-3-yl, benzothiazolyl, benzothiadiazolyl, benzoxazolyl, benzisoxazolyl, benzimidazolyl, benzo[b]thiophenyl (i.e., benzothienyl), triazolyl (e.g., lH-l,2,3-triazolyl, 2H-l,2,3-triazolyl, lH-l,2,4-triazolyl, or 4H-l,2,4-triazolyl), benzotriazolyl, lH-tetrazolyl, 2H-tetrazolyl, triazinyl (e.g., 1,2,3-triazinyl, 1,2,4-triazinyl, or 1,3,5-triazinyl), furo[2,3-c]pyridinyl, dihydrofuropyridinyl (e.g., 2,3-dihydrofuro[2,3-c]pyridinyl or l,3-dihydrofuro[3,4-c]pyridinyl), imidazopyridinyl (e.g., imidazo[l,2-a]pyridinyl or imidazo[3,2-a]pyridinyl), quinazolinyl, thienopyridinyl, tetrahydrothienopyridinyl (e.g., 4,5,6,7-tetrahydrothieno[3,2-c]pyridinyl), dibenzofuranyl, 1,3-benzodioxolyl, benzodioxanyl (e.g., 1,3-benzodioxanyl or 1,4-benzodioxanyl), or coumarinyl. Unless defined otherwise, the term "heteroaryl" preferably refers to a 5 to 14 membered (more preferably 5 to 10 membered) monocyclic ring orfused ring system comprising one or more (e.g., one, two, three orfour) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized; even more preferably, a "heteroaryl" refers to a 5 or 6 membered monocyclic ring comprising one or more (e.g., one, two or three) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized.

[0183] As used herein, the term "cycloalkyl" refers to a saturated hydrocarbon ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings; such as, e.g., a fused ring system composed of two or three fused rings). "Cycloalkyl" may, e.g., refer to cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, decalinyl (i.e., decahydronaphthyl), or adamantyl. Unless defined otherwise, "cycloalkyl" preferably refers to a C3-11 cycloalkyl, and more preferably refers to a C3-7 cycloalkyl. A particularly preferred "cycloalkyl" is a monocyclic saturated hydrocarbon ring having 3 to 7 ring members (e.g., cyclopropyl or cyclohexyl).

[0184] As used herein, the term "heterocycloalkyl" refers to a saturated ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings; such as, e.g., a fused ring system composed of two or three fused rings), wherein said ring group contains one or more (such as, e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, and the remaining ring atoms are carbon atoms, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) may optionally be oxidized, and further wherein one or more carbon ring atoms may optionally be oxidized (i.e., to form an oxo group). For example, each heteroatom-containing ring comprised in said saturated ring group may contain one or two O atoms and / or one or two S atoms (which may optionally be oxidized) and / or one, two, three or four N atoms (which may optionally be oxidized), provided that the total number of heteroatoms in the corresponding heteroatom-containing ring is 1 to 4 and that there is at least one carbon ring atom (which may optionally be oxidized) in the corresponding heteroatom-containing ring. "Heterocycloalkyl" may, e.g., refer to aziridinyl, azetidinyl, pyrrolidinyl, imidazolidinyl, pyrazolidinyl, piperidinyl, piperazinyl, azepanyl, diazepanyl (e.g., 1,4-diazepanyl), oxazolidinyl, isoxazolidinyl, thiazolidinyl, isothiazolidinyl, morpholinyl (e.g., morpholin-4-yl), thiomorpholinyl (e.g., thiomorpholin-4-yl), oxazepanyl, oxiranyl, oxetanyl, tetra hydrofuranyl, 1,3-dioxolanyl, tetrahydropyranyl, 1,4-dioxanyl, oxepanyl, thiiranyl, thietanyl, tetrahydrothiophenyl (i.e., thiolanyl), 1,3-dithiolanyl, thianyl, 1,1-dioxothianyl, thiepanyl, decahydroquinolinyl, decahydroisoquinolinyl, or 2-oxa-5-aza-bicyclo[2.2.1]hept-5- yl. Unless defined otherwise, "heterocycloalkyl" preferably refers to a 3 to 11 membered saturated ring group, which is a monocyclic ring or a fused ring system (e.g., a fused ring system composed of two fused rings), wherein said ring group contains one or more (e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized; more preferably, "heterocycloalkyl" refers to a 5 to 7 membered saturated monocyclic ring group containing one or more (e.g., one, two, or three) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized.

[0185] As used herein, the term "cycloalkenyl" refers to an unsaturated alicyclic (non-aromatic) hydrocarbon ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings; such as, e.g., a fused ring system composed of two or three fused rings), wherein said hydrocarbon ring group comprises one or more (e.g., one or two) carbon-to-carbon double bonds and does not comprise any carbon-to-carbon triple bond. "Cycloalkenyl" may, e.g., refer to cyclopropenyl, cyclobutenyl, cyclopentenyl, cyclohexenyl, cyclohexadienyl, cycloheptenyl, or cycloheptadienyl. Unless defined otherwise, "cycloalkenyl" preferably refers to a C3-11 cycloalkenyl, and more preferably refers to a C3-7 cycloalkenyl. A particularly preferred "cycloalkenyl" is a monocyclic unsaturated alicyclic hydrocarbon ring having 3 to 7 ring members and containing one or more (e.g., one or two; preferably one) carbon-to-carbon double bonds.

[0186] As used herein, the term "heterocycloalkenyl" refers to an unsaturated alicyclic (non-aromatic) ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings; such as, e.g., a fused ring system composed of two or three fused rings), wherein said ring group contains one or more (such as, e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, and the remaining ring atoms are carbon atoms, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) may optionally be oxidized, wherein one or more carbon ring atoms may optionally be oxidized (i.e., to form an oxo group), and further wherein said ring group comprises at least one double bond between adjacent ring atoms and does not comprise any triple bond between adjacent ring atoms. For example, each heteroatom-containing ring comprised in said unsaturated alicyclic ring group may contain one or two O atoms and / or one or two S atoms (which may optionally be oxidized) and / or one, two, three or four N atoms (which may optionally be oxidized), provided that the total number of heteroatoms in the corresponding heteroatom-containing ring is 1 to 4 and that there is at least one carbon ring atom (which may optionally be oxidized) in the corresponding heteroatom-containing ring. "Heterocycloalkenyl" may, e.g., refer to imidazolinyl (e.g., 2-imidazolinyl (i.e., 4,5-dihydro-lH-imidazolyl), 3-imidazolinyl, or 4-imidazolinyl), tetrahydropyridinyl (e.g., 1,2,3,6-tetrahydropyridinyl), dihydropyridinyl (e.g., 1,2- dihydropyridinyl or 2,3-dihyd ropyridinyl), pyranyl (e.g., 2H-pyranyl or4H-pyranyl), thiopyranyl (e.g., 2H-thiopyranyl or 4H-thiopyranyl), dihydropyranyl, dihydrofuranyl, dihydropyrazolyl, dihydropyrazinyl, dihydroisoindolyl, octahydroquinolinyl (e.g., 1, 2, 3, 4, 4a, 5,6,7- octahydroquinolinyl), or octahydroisoquinolinyl (e.g., 1,2,3,4,5,6,7,8-octahydroisoquinolinyl). Unless defined otherwise, "heterocycloalkenyl" preferably refers to a 3 to 11 membered unsaturated alicyclic ring group, which is a monocyclic ring or a fused ring system (e.g., a fused ring system composed of two fused rings), wherein said ring group contains one or more (e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, wherein one or more carbon ring atoms are optionally oxidized, and wherein said ring group comprises at least one double bond between adjacent ring atoms and does not comprise any triple bond between adjacent ring atoms; more preferably, "heterocycloalkenyl" refers to a 5 to 7 membered monocyclic unsaturated non-aromatic ring group containing one or more (e.g., one, two, or three) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, wherein one or more carbon ring atoms are optionally oxidized, and wherein said ring group comprises at least one double bond between adjacent ring atoms and does not comprise any triple bond between adjacent ring atoms.

[0187] As understood herein, the term "heterocycloalkylene" refers to a heterocycloalkyl group, as defined herein above, but having two points of attachment, i.e. a divalent saturated ring group, including monocyclic rings as well as bridged ring, spiro ring and / or fused ring systems (which may be composed, e.g., of two or three rings; such as, e.g., a fused ring system composed of two or three fused rings), wherein said ring group contains one or more (such as, e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, and the remaining ring atoms are carbon atoms, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) may optionally be oxidized, and further wherein one or more carbon ring atoms may optionally be oxidized (i.e., to form an oxo group). For example, each heteroatom-containing ring comprised in said saturated ring group may contain one or two O atoms and / or one or two S atoms (which may optionally be oxidized) and / or one, two, three or four N atoms (which may optionally be oxidized), provided that the total number of heteroatoms in the corresponding heteroatom-containing ring is 1 to 4 and that there is at least one carbon ring atom (which may optionally be oxidized) in the corresponding heteroatom-containing ring. "Heterocycloalkylene" may, e.g., refer to aziridinylene, azetidinylene, pyrrolidinylene, imidazolidinylene, pyrazolidinylene, piperidinylene, piperazinylene, azepanylene, diazepanylene (e.g., 1,4-diazepanylene), oxazolidinylene, isoxazolidinylene, thiazolidinylene, isothiazolidinylene, morpholinylene, thiomorpholinylene, oxazepanylene, oxiranylene, oxetanylene, tetrahydrofuranylene, 1,3-dioxolanylene, tetrahydropyranylene, 1,4-dioxanylene, oxepanylene, thiiranylene, thietanylene, tetrahydrothiophenylene (i.e., thiolanylene), 1,3-dithiolanylene, thianylene, 1,1-dioxothianylene, thiepanylene, decahydroquinolinylene, decahydroisoquinolinylene, or 2- oxa-5-aza-bicyclo[2.2.1]hept-5-ylene. Unless defined otherwise, "heterocycloalkylene" preferably refers to a divalent 3 to 11 membered saturated ring group, which is a monocyclic ring or a fused ring system (e.g., a fused ring system composed of two fused rings), wherein said ring group contains one or more (e.g., one, two, three, or four) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized; more preferably, "heterocycloalkylene" refers to a divalent 5 to 7 membered saturated monocyclic ring group containing one or more (e.g., one, two, or three) ring heteroatoms independently selected from O, S and N, wherein one or more S ring atoms (if present) and / or one or more N ring atoms (if present) are optionally oxidized, and wherein one or more carbon ring atoms are optionally oxidized.

[0188] As used herein, the term " / V-heterocycloalkylene" refers to the heterocycloalkylene group as defined hereinabove wherein said heterocycloalkylene includes at least one nitrogen atom which serves as an attachment point of said heterocycloalkylene.

[0189] As used herein, the term "halogen" refers to fluoro (-F), chloro (-CI), bromo (-Br), or iodo (-1). The terms "Hal" and "halogen" can be used interchangeably.

[0190] As used herein, the term "nitro" refers to a group -NO2. The terms "bond" and "covalent bond" are used herein synonymously, unless explicitly indicated otherwise or contradicted by context.

[0191] As used herein, the terms "optional", "optionally" and "may" denote that the indicated feature may be present but can also be absent. Whenever the term "optional", "optionally" or "may" is used, the present invention specifically relates to both possibilities, i.e., that the corresponding feature is present or, alternatively, that the corresponding feature is absent. For example, the expression "X is optionally substituted with Y" (or "X may be substituted with Y") means that X is either substituted with Y or is unsubstituted. Likewise, if a component of a composition is indicated to be "optional", the invention specifically relates to both possibilities, i.e., that the corresponding component is present (contained in the composition) or that the corresponding component is absent from the composition.

[0192] Various groups are referred to as being "optionally substituted" in this specification. Generally, these groups may carry one or more substituents, such as, e.g., one, two, three or four substituents. It will be understood that the maximum number of substituents is limited by the number of attachment sites available on the substituted moiety. Unless defined otherwise, the "optionally substituted" groups referred to in this specification carry preferably not more than two substituents and may, in particular, carry only one substituent. Moreover, unless defined otherwise, it is preferred that the optional substituents are absent, i.e. that the corresponding groups are unsubstituted.

[0193] The term "peptide" refers to a polymer of two or more amino acids linked via amide bonds that are formed between an amino group of one amino acid and a carboxylic acid group of another amino acid. The amino acids comprised in the peptide or protein, which are also referred to as amino acid residues, may be selected from the 20 standard proteinogenic a- amino acids (i.e., Ala, Arg, Asn, Asp, Cys, Glu, Gin, Gly, His, lie, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, and Vai) but also from non-proteinogenic and / or non-standard a-amino acids (such as, e.g., ornithine, citrulline, homolysine, pyrrolysine, 4-hydroxyproline, a-methylalanine (i.e., 2-aminoisobutyric acid), norvaline, norleucine, terleucine (tert-leucine), labionin, or an alanine or glycine that is substituted at the side chain with a cyclic group such as, e.g., cyclopentylalanine, cyclohexylalanine, phenylalanine, naphthylalanine, pyridylalanine, thienylalanine, cyclohexylglycine, or phenylglycine) as well as (3-amino acids (e.g., (3-alanine), y-amino acids (e.g., y-aminobutyric acid, isoglutamine, or statine) and 6-amino acids. Preferably, the amino acid residues comprised in the peptide or protein are selected from a- amino acids, more preferably from the 20 standard proteinogenic a-amino acids (which can be present as the L-isomer or the D-isomer, and are preferably all present as the L-isomer). The peptide may be unmodified or may be modified, e.g., at its N-terminus, at its C-terminus and / or at a functional group in the side chain of any of its amino acid residues (particularly at the side chain functional group of one or more Lys, His, Ser, Thr, Tyr, Cys, Asp, Glu, and / or Arg residues). Such modifications may include, e.g., the attachment of any of the protecting groups described for the corresponding functional groups in: Wuts PG & Greene TW, Greene's protective groups in organic synthesis, John Wiley & Sons, 2006. Such modifications may also include the covalent attachment of one or more polyethylene glycol (PEG) chains (forming a PEGylated peptide), the glycosylation and / or the acylation with one or more fatty acids (e.g., one or more C8-30 alkanoic or alkenoic acids; forming a fatty acid acylated peptide or protein). Moreover, such modified peptides or proteins may also include peptidomimetics, provided that they contain at least two amino acids that are linked via an amide bond (formed between an amino group of one amino acid and a carboxyl group of another amino acid). The amino acid residues comprised in the peptide or protein may, e.g., be present as a linear molecular chain (forming a linear peptide) or may form one or more rings (corresponding to a cyclic peptide). The peptide may also form oligomers consisting of two or more identical or different molecules. The term "peptide" may also refer to a monovalent radical derived from a "peptide" as described herein above, wherein preferably an amino group or a carboxylic acid group, in particular the N-terminal amino group or the C-terminal carboxylic acid group, serves as point of attachment, preferably through an amide bond. Thus, the term "peptide" may refer to a peptidyl moiety attached to the rest of the molecule through a -CO- group (formed from a carboxylic acid group, e.g., from its C-terminal carboxylic acid group).

[0194] A skilled person will appreciate that the substituent groups comprised in the compounds of the present invention may be attached to the remainder of the respective compound via a number of different positions of the corresponding specific substituent group. Unless defined otherwise, the preferred attachment positions for the various specific substituent groups are as illustrated in the examples.

[0195] As used herein, unless explicitly indicated otherwise or contradicted by context, the terms "a", "an" and "the" are used interchangeably with "one or more" and "at least one". Thus, for example, a composition comprising "a" compound of formula (I) can be interpreted as referring to a composition comprising "one or more" compounds of formula (I).

[0196] It is to be understood that wherever numerical ranges are provided / disclosed herein, all values and subranges encompassed by the respective numerical range are meant to be encompassed within the scope of the invention. Accordingly, the present invention specifically and individually relates to each value that falls within a numerical range disclosed herein, as well as each subrange encompassed by a numerical range disclosed herein.

[0197] As used herein, the term "about" preferably refers to ±10% of the indicated numerical value, more preferably to ±5% of the indicated numerical value, and in particular to the exact numerical value indicated. If the term "about" is used in connection with the endpoints of a range, it preferably refers to the range from the lower endpoint -10% of its indicated numerical value to the upper endpoint +10% of its indicated numerical value, more preferably to the range from of the lower endpoint -5% to the upper endpoint +5%, and even more preferably to the range defined by the exact numerical values of the lower endpoint and the upper endpoint.

[0198] As used herein, the term "comprising" (or "comprise", "comprises", "contain", "contains", or "containing"), unless explicitly indicated otherwise or contradicted by context, has the meaning of "containing, inter alia", i.e., "containing, among further optional elements, ..." . In addition thereto, this term also includes the narrower meanings of "consisting essentially of" and "consisting of". For example, the term "A comprising B and C" has the meaning of "A containing, inter alia, B and C", wherein A may contain further optional elements (e.g., "A containing B, C and D" would also be encompassed), but this term also includes the meaning of "A consisting essentially of B and C" and the meaning of "A consisting of B and C" (i.e., no other components than B and C are comprised in A).

[0199] The scope of the invention embraces all salts, in particular all pharmaceutically acceptable salt forms of the compounds of formula (I) which may be formed, e.g., by protonation of an atom carrying an electron lone pair which is susceptible to protonation, such as an amino group, with an inorganic or organic acid, or as a salt of an acid group (such as a carboxylic acid group) with a physiologically acceptable cation. Exemplary base addition salts comprise, for example: alkali metal salts such as sodium or potassium salts; alkaline earth metal salts such as calcium or magnesium salts; zinc salts; ammonium salts; aliphatic amine salts such as trimethylamine, triethylamine, dicyclohexylamine, ethanolamine, diethanolamine, triethanolamine, procaine salts, meglumine salts, ethylenediamine salts, or choline salts; aralkyl amine salts such as N,N- dibenzylethylenediamine salts, benzathine salts, benethamine salts; heterocyclic aromatic amine salts such as pyridine salts, picoline salts, quinoline salts or isoquinoline salts; quaternary ammonium salts such as tetramethylammonium salts, tetraethylammonium salts, benzyltrimethylammonium salts, benzyltriethylammonium salts, benzyltributylammonium salts, methyltrioctylammonium salts or tetra butylammonium salts; and basic amino acid salts such as arginine salts, lysine salts, or histidine salts. Exemplary acid addition salts comprise, for example: mineral acid salts such as hydrochloride, hydrobromide, hydroiodide, sulfate salts (such as, e.g., sulfate or hydrogensulfate salts), nitrate salts, phosphate salts (such as, e.g., phosphate, hydrogenphosphate, or dihydrogenphosphate salts), carbonate salts, hydrogencarbonate salts, perchlorate salts, borate salts, or thiocyanate salts; organic acid salts such as acetate, propionate, butyrate, pentanoate, hexanoate, heptanoate, octanoate, cyclopentanepropionate, decanoate, undecanoate, oleate, stearate, lactate, maleate, oxalate, fumarate, tartrate, malate, citrate, succinate, adipate, gluconate, glycolate, nicotinate, benzoate, salicylate, ascorbate, pamoate (embonate), camphorate, glucoheptanoate, or pivalate salts; sulfonate salts such as methanesulfonate (mesylate), ethanesulfonate (esylate), 2-hydroxyethanesulfonate (isethionate), benzenesulfonate (besylate), p-toluenesulfonate (tosylate), 2-naphthalenesulfonate (napsylate), 3-phenylsulfonate, or camphorsulfonate salts; glycerophosphate salts; and acidic amino acid salts such as aspartate or glutamate salts. Preferred pharmaceutically acceptable salts of the compounds of formula (I) include a hydrochloride salt, a hydrobromide salt, a mesylate salt, a sulfate salt, a tartrate salt, a fumarate salt, an acetate salt, a citrate salt, and a phosphate salt. A particularly preferred pharmaceutically acceptable salt of the compound of formula (I) is a hydrochloride salt. Accordingly, it is preferred that the compound of formula (I), including any one of the specific compounds of formula (I) described herein, is in the form of a fumarate salt, a maleate salt, an oxalate salt, a malate salt, a tartrate salt, and a mesylate salt.

[0200] The present invention also specifically relates to the compound of formula (I), including any one of the specific compounds of formula (I) described herein, in non-salt form.

[0201] Moreover, the scope of the invention embraces the compounds of formula (I) in any solvated form, including, e.g., solvates with water (i.e., as a hydrate) or solvates with organic solvents such as, e.g., methanol, ethanol, isopropanol, acetic acid, ethyl acetate, ethanolamine, DMSO, or acetonitrile. All physical forms, including any amorphous or crystalline forms (i.e., polymorphs), of the compounds of formula (I) are also encompassed within the scope of the invention. It is to be understood that such solvates and physical forms of pharmaceutically acceptable salts of the compounds of the formula (I) are likewise embraced by the invention.

[0202] Furthermore, the compounds of formula (I) may exist in the form of different isomers, in particular stereoisomers (including, e.g., geometric isomers (or cis / trans isomers), enantiomers and diastereomers) or tautomers (including, in particular, prototropic tautomers, such as keto / enol tautomers or thione / thiol tautomers). All such isomers of the compounds of formula (I) are contemplated as being part of the present invention, either in admixture or in pure or substantially pure form. As for stereoisomers, the invention embraces the isolated optical isomers of the compounds according to the invention as well as any mixtures thereof (including, in particular, racemic mixtures / racemates and non-racemic mixtures). The racemates can be resolved by physical methods, such as, e.g., fractional crystallization, separation or crystallization of diastereomeric derivatives, or separation by chiral column chromatography. The individual optical isomers can also be obtained from the racemates via salt formation with an optically active acid followed by crystallization. The present invention further encompasses any tautomers of the compounds of formula (I). It will be understood that some compounds may exhibit tautomerism. In such cases, the formulae provided herein expressly depict only one of the possible tautomeric forms. The formulae and chemical names as provided herein are intended to encompass any tautomeric form of the corresponding compound and not to be limited merely to the specific tautomeric form depicted by the drawing or identified by the name of the compound.

[0203] The scope of the invention also embraces compounds of formula (I), in which one or more atoms are replaced by a specific isotope of the corresponding atom. For example, the invention encompasses compounds of formula (I), in which one or more hydrogen atoms (or, e.g., all hydrogen atoms) are replaced by deuterium atoms (i.e.,2H; also referred to as "D"). Accordingly, the invention also embraces compounds of formula (I) which are enriched in deuterium. Naturally occurring hydrogen is an isotopic mixture comprising about 99.98 mol-% hydrogen-1 ^H) and about 0.0156 mol-% deuterium (2H or D). The content of deuterium in one or more hydrogen positions in the compounds of formula (I) can be increased using deuteration techniques known in the art. For example, a compound of formula (I) or a reactant or precursor to be used in the synthesis of the compound of formula (I) can be subjected to an H / D exchange reaction using, e.g., heavy water (D2O). Further suitable deuteration techniques are described in: Atzrodt J et al., Bioorg Med Chem, 20(18), 5658-5667, 2012; William JS et al., Journal of Labelled Compounds and Radiopharmaceuticals, 53(11-12), 635- 644, 2010; Modvig A et al., J Org Chem, 79, 5861-5868, 2014. The content of deuterium can be determined, e.g., using mass spectrometry or NMR spectroscopy. Unless specifically indicated otherwise, it is preferred that the compound of formula (I) is not enriched in deuterium. Accordingly, the presence of naturally occurring hydrogen atoms or1H hydrogen atoms in the compounds of formula (I) is preferred.

[0204] The present invention also embraces compounds of formula (I), in which one or more atoms are replaced by a positron-emitting isotope of the corresponding atom, such as, e.g.,18F,nC,13N,150,76Br,77Br,120l and / or124L Such compounds can be used as tracers, trackers or imaging probes in positron emission tomography (PET). The invention thus includes (i) compounds of formula (I), in which one or more fluorine atoms (or, e.g., all fluorine atoms) are replaced by18F atoms, (ii) compounds of formula (I), in which one or more carbon atoms (or, e.g., all carbon atoms) are replaced bynC atoms, (iii) compounds of formula (I), in which one or more nitrogen atoms (or, e.g., all nitrogen atoms) are replaced by13N atoms, (iv) compounds of formula (I), in which one or more oxygen atoms (or, e.g., all oxygen atoms) are replaced by15O atoms, (v) compounds of formula (I), in which one or more bromine atoms (or, e.g., all bromine atoms) are replaced by76Br atoms, (vi) compounds of formula (I), in which one or more bromine atoms (or, e.g., all bromine atoms) are replaced by77Br atoms, (vii) compounds of formula (I), in which one or more iodine atoms (or, e.g., all iodine atoms) are replaced by120l atoms, and (viii) compounds of formula (I), in which one or more iodine atoms (or, e.g., all iodine atoms) are replaced by124l atoms. In general, it is preferred that none of the atoms in the compounds of formula (I) are replaced by specific isotopes. Throughout the present application, reference is made to an isopeptide-linked lysine or an analog thereof. This isopeptide-linked lysine or analog thereof is released when the covalent bond C of the compound of formula (I) according to the invention is hydrolyzed or cleaved inside the cell. Accordingly, the isopeptide-linked lysine or analog thereof has the following structure according to formula (IV): wherein C' is -NH2 or -OH,

[0205] D is as defined for formula (I) (with the difference that its left empty valence is connected to

[0206] C' instead of C), in particular D is selected from the left empty valence is connected to C' and the right empty valence is connected to -CO-NH- , wherein X, X1and X2are as defined for formula (I), including any preferred definition and any specific embodiment of the compound of formula (I), and wherein E is as defined for formula (I), in particular wherein E is -(C1-6 alkyleneJ-CHY^2, wherein Y1is -NH2 and Y2is COOH, preferably wherein the configuration on the carbon atom that carries Y1and Y2is the same as in L-lysine; or wherein E is selected from: -CH2CH2CH2CH2- CH(-NH2)-COOH, -CH2CH2CH2-CH(-NH2)-COOH, -CH2CH2CH2CH2CH2-COOH, CH2CH2CH2CH2CH2-NH2, -CH2CH2CH2CH2-CH(-NH2)-CONH2, -CH2CH2CH2CH2-CH(-OH)-COOH; - CH2CH2-CH(-NH2)-COOH, and -CH2-CH(-NH2)-COOH, preferably wherein E is CH2CH2CH2CH2- CH(-NH2)-COOH.

[0207] In another aspect, the invention relates to a method for identifying an engineered peptide- binding protein that promotes uptake of the compound according to the present invention, the method comprising the steps of: a) providing a nucleic acid encoding a peptide-binding protein comprising one or more mutations; b) expressing the nucleic acid of step (a) in a cell comprising an orthogonal translation system, wherein the orthogonal translation system comprises an orthogonal aminoacyl- tRNA synthetase / tRNA pair specific for an isopeptide-linked lysine or an analog thereof, and a reporter protein gene comprising one or more codons that have been re-allocated for the incorporation of the isopeptide-linked lysine or the analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; c) contacting the cell of step (b) with a compound according to the invention, wherein the compound is transported into the cell and hydrolyzed or cleaved inside the cell to release the isopeptide-linked lysine or the analog thereof; d) detecting the expression of the reporter protein gene by the cell, wherein the expression is dependent on the successful uptake of the compound in step (c); and e) identifying the engineered peptide-binding protein that promotes uptake of the compound according to the invention based on the expression of the reporter protein gene detected in step (d).

[0208] That is, another aspect of the invention relates to a method for engineering a peptide-binding protein to promote the uptake of the compound according to the invention. It has been shown in the appended experimental Examples that engineering of the peptide-binding protein OppA from E. coli resulted in higher affinity and selectivity for the compounds according to the invention, i.e., compounds comprising an isopeptide-linked lysine residue or an analog thereof, which resulted in increased cellular uptake of these compounds; see Example 4.

[0209] Herein, an engineered peptide-binding protein is defined as promoting the uptake of the compound of the invention if, when expressed by a cell, it leads to increased intracellular accumulation of the compound compared to a cell expressing the non-engineered counterpart. Intracellular levels of the compound may be measured, for example, by HPLC as described herein.

[0210] Alternatively or in addition, the method of the invention may be used to identify an engineered peptide-binding protein having increased affinity and / or selectivity for the compound according to the present invention.

[0211] Thus, in certain embodiments, the invention relates to a method for identifying an engineered peptide-binding protein having increased affinity and / or selectivity for the compound according to the present invention, the method comprising the steps of: a) providing a nucleic acid encoding a peptide-binding protein comprising one or more mutations; b) expressing the nucleic acid of step (a) in a cell comprising an orthogonal translation system, wherein the orthogonal translation system comprises an orthogonal aminoacyl- tRNA synthetase / tRNA pair specific for an isopeptide-linked lysine or an analog thereof, and a reporter protein gene comprising one or more codons that have been re-allocated for the incorporation of the isopeptide-linked lysine or the analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; c) contacting the cell of step (b) with a compound according to any one of claims 1 to 16, wherein the compound is transported into the cell and hydrolyzed or cleaved inside the cell to release the isopeptide-linked lysine or the analog thereof; d) detecting the expression of the reporter protein gene by the cell, wherein the expression is dependent on the successful uptake of the compound in step (c); and e) identifying the engineered peptide-binding protein as having increased affinity and / or selectivity for the compound according to any one of claims 1 to 16 based on the expression of the reporter protein gene detected in step (d).

[0212] A peptide-binding protein is defined to have increased affinity for the compound of the invention if it demonstrates a stronger binding interaction with the compound compared to its non-engineered counterpart. This increased affinity can be quantified by measuring the dissociation constant (Kd) for the compound of the invention. A lower Kd value indicates a higher affinity, as it reflects a tighter binding interaction between the peptide-binding protein and the compound.

[0213] A peptide-binding protein is defined to have increased selectivity for the compound of the invention if its preference for the compound compared to linear peptides is greater than that of a non-engineered counterpart. This increased selectivity can be quantified by comparing the dissociation constants (Kd) forthe compound of the invention and the linear peptides, and calculating the selectivity ratio (SR). The selectivity ratio is defined as the ratio of the Kd for the linear peptides to the Kd for the compound of the invention (SR = Kd (linear peptide) / Kd (compound according to the invention)). A higher selectivity ratio indicates greater selectivity for the compound of the invention, demonstrating that the engineered protein has a stronger preference for the compound compared to its non-engineered counterpart.

[0214] A higher selectivity ratio indicates greater selectivity for the compound of the invention. For instance, if the affinity for the compound of the invention remains the same while the affinity for linear peptides decreases, the selectivity ratio will increase, demonstrating that the engineered protein has a stronger preference for the compound of the invention.

[0215] That is, a higher selectivity ratio forthe compound according to the invention may be achieved by engineering the peptide-binding protein such that its affinity for the compound of the invention is maintained or increased, while its affinity for linear peptides is decreased. This approach ensures that the protein preferentially binds to the compound of the invention over other similar peptides, thereby enhancing its specificity. Alternatively, a higher selectivity ratio for the compound according to the invention may also be achieved by engineering the peptide-binding protein to increase its affinity for the compound of the invention while maintaining or decreasing its affinity for linear peptides. This strategy similarly enhances the protein's preference for the target compound, ensuring that it binds more effectively and selectively to the compound of the invention compared to linear peptides.

[0216] In certain embodiments, an engineered peptide-binding protein is determined to have a higher affinity for the compound according to the invention, if the affinity for the compound according to the invention is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400% or 500% higher than that of the non-engineered counterpart.

[0217] In certain embodiments, an engineered peptide-binding protein is determined to have a higher selectivity for the compound according to the invention, if the selectivity ratio for the compound according to the invention is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400% or 500% higher than that of the non-engineered counterpart.

[0218] To standardize the determination of the affinity and the selectivity ratio, the same compounds as in Example 4 may be used. That is, as an example for the compound according to the invention, the compound G-SlsoK may be used and as an example for a linear peptide, the peptide GSK may be used. These compounds differ only in that in G-SlsoK, the lysine residue is attached to the serine residue via an isopeptide bond formed between the carboxyl group of the serine and the amino group in the side chain of the lysine.

[0219] The skilled person is aware of methods to determine the affinity of a protein for a ligand. For example, microscale thermophoresis may be used as described in appended Example 4. Alternative methods known to the person skilled in the art include, without limitation, isothermal titration calorimetry, surface plasmon resonance, fluorescence anisotropy, and equilibrium dialysis. These techniques provide reliable measurements of binding constants (Kd), allowing for the accurate assessment of protein-ligand interactions and the determination of binding affinities.

[0220] The starting point for the method of the invention is a nucleic acid encoding a peptide-binding protein. A peptide-binding protein is a type of protein that has the ability to specifically recognize and bind to peptide molecules. These proteins typically contain binding sites that interact with the peptide's amino acid residues through various non-covalent interactions, such as hydrogen bonds, ionic interactions, and hydrophobic effects.

[0221] The peptide-binding protein may be any peptide-binding protein from any organism. However, it is preferred herein, that the peptide-binding protein is the substrate-binding protein of an ABC transporter, in particular a prokaryotic ABC transporter. An ATP-binding cassette (ABC) transporter is a type of membrane protein complex that utilizes the energy derived from the hydrolysis of adenosine triphosphate (ATP) to transport various molecules across cellular membranes. ABC transporters are characterized by their ATP-binding domains (nucleotide-binding domains or NBDs) and transmembrane domains (TMDs) that form the pathway for substrate transport. A critical component of many ABC transporters, particularly in prokaryotes, is the substrate-binding protein (SBP). The SBP specifically binds to the substrate (e.g., a peptides) with high affinity and delivers it to the transmembrane components of the transporter. This interaction ensures the specificity and efficiency of the transport process, allowing the substrate to be translocated into the cell in an ATP-dependent manner. In bacteria, the substrate-binding protein is typically a periplasmic protein (in Gramnegative bacteria) or an extracellular protein that is tethered to the outer membrane (in Gram-positive bacteria).

[0222] That is, in certain embodiments, the peptide-binding protein is preferably a peptide-binding protein of a prokaryotic peptide transporter. In a particular embodiment, the invention relates to the method according to the invention, wherein the peptide-binding protein is the peptide- binding protein of a dipeptide permease (DppA),an oligopeptide permease (OppA) or a murein peptide-binding protein (MppA).

[0223] Dipeptide permease is a type of ABC transporter system found in bacteria, such as E. coli, that is specialized in the uptake and transport of dipeptides across the cell membrane. The system comprises several components, including a periplasmic dipeptide-binding protein (DppA), which specifically binds dipeptides in the periplasmic space. Once bound, the dipeptide is delivered to the transmembrane components of the system (DppBCDF), which facilitate its transport into the cytoplasm using the energy derived from ATP hydrolysis. The Dpp system plays a crucial role in nutrient acquisition, allowing the bacterium to utilize dipeptides as sources of amino acids and nitrogen for growth and metabolism.

[0224] Dipeptide permease (Dpp) and oligopeptide permease (Opp) are types of ATP-binding cassette (ABC) transporter systems found in bacteria, such as E. coli, that specialize in the uptake and transport of peptides across the cell membrane. The Dpp system is responsible for the transport of dipeptides, while the Opp system handles oligopeptides (short chains of amino acids). Both systems consist of several key components, including a periplasmic peptide- binding protein (DppA for dipeptides and OppA for oligopeptides), which specifically binds the respective peptides in the periplasmic space. Once bound, the peptides are delivered to the transmembrane components of the systems (DppBCDF for dipeptides and OppBCDF for oligopeptides), which facilitate their transport into the cytoplasm using the energy derived from ATP hydrolysis. These permease systems are essential for nutrient acquisition, enabling the bacterium to utilize peptides as sources of amino acids and nitrogen, thereby supporting growth and metabolic processes.

[0225] MppA is a periplasmic substrate-binding protein that is part of the murein tripeptide permease system in bacteria. MppA specifically binds to murein tripeptides, which are degradation products of peptidoglycan, in the periplasmic space. Once bound, MppA delivers the murein tripeptides to the transmembrane components of the Opp system (OppBCDF), which then transport the peptides into the cytoplasm using the energy derived from ATP hydrolysis. The Mpp system, utilizing the OppBCDF components, plays a crucial role in recycling cell wall components and in nutrient acquisition, allowing the bacterium to utilize murein tripeptides as sources of amino acids and nitrogen for growth and metabolism.

[0226] In a preferred embodiment, the starting point for the method according to the invention is the peptide-binding protein OppA. It has been demonstrated in Example 4 that OppA can be engineered to have a higher specificity for the compounds according to the invention. The OppA may be derived from any organism. More preferably, the OppA is a prokaryotic OppA. In a most preferred embodiment, the OppA is derived from E. coli. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the peptide-binding protein is the peptide-binding protein of the E. coli oligopeptide permease (OppA; SEQ ID NO:1).

[0227] In a first step of the method of the invention, a nucleic acid encoding a peptide-binding protein comprising one or more mutations is provided. The one or more mutations may be introduced into the nucleic acid encoding the peptide-binding protein by any suitable method known in the art. The skilled person is aware of such methods.

[0228] Mutations may be introduced into the nucleic acid in a random or site-directed manner. Random mutagenesis techniques, such as error-prone PCR, can be used to generate a diverse library of variants by introducing mutations throughout the entire gene. Alternatively, site- directed mutagenesis allows for the precise introduction of specific mutations at predetermined positions within the gene, enabling the study of the effects of particular amino acid changes on the protein's function.

[0229] In a particular embodiment, the invention relates to the method according to the invention, wherein the mutations introduced in step (a) are introduced using error-prone PCR and / or saturated mutagenesis and / or continuous directed evolution.

[0230] Error-prone PCR is a molecular biology technique used to introduce random mutations into a specific DNA sequence during the amplification process. Unlike standard PCR, which aims to correctly replicate the DNA template, error-prone PCR employs modified reaction conditions, such as the use of a low-fidelity DNA polymerase, imbalanced nucleotide concentrations, or the inclusion of mutagenic agents, to increase the error rate of DNA synthesis. This results in the generation of a diverse library of mutant DNA sequences, each containing various point mutations. Error-prone PCR is particularly useful in directed evolution experiments, where the goal is to create a wide array of genetic variants that can be screened for desirable traits. Error- prone PCT may be used to identify positions in a protein that have an impact on the functionality of the protein. For example, error-prone PCR may be used to identify positions in a peptide-binding protein that are involved in binding the compound of the invention.

[0231] Saturated mutagenesis is a molecular biology technique used to generate a comprehensive library of genetic variants by introducing all possible nucleotide substitutions at specific positions within a target DNA sequence. This method involves systematically mutating one or more codons to include every possible amino acid residue, thereby creating a diverse set of protein variants. Saturated mutagenesis is typically performed using synthetic oligonucleotides that contain degenerate codons at the desired mutation sites, which are then incorporated into the target gene through techniques such as PCR or site-directed mutagenesis. This approach allows researchers to explore the functional impact of every possible amino acid change at specific positions, making it a valuable tool for studying structure-function relationships, optimizing protein properties, and engineering proteins with enhanced or novel functionalities. Saturated mutagenesis may be used to mutagenize positions in a peptide-binding protein that have previously been identified by error-prone PCR to have an impact on peptide-binding.

[0232] Continuous directed evolution is a powerful technique used to evolve proteins or nucleic acids in a continuous, iterative manner to achieve desired traits or functions. Unlike traditional directed evolution methods that rely on discrete rounds of mutagenesis and selection, continuous directed evolution integrates these processes into a single, ongoing cycle. This is typically achieved by coupling the evolutionary process to the replication or survival of the host organism, often using systems such as phage-assisted continuous evolution (PACE) or yeast surface display. In these systems, genetic variants that exhibit improved or desired characteristics are preferentially replicated or survive better under selective conditions, thereby enriching the population for beneficial mutations over time. Continuous directed evolution allows for the rapid and efficient optimization of biomolecules, enabling the development of proteins with enhanced activities, specificities, or stabilities, as well as the discovery of novel functions.

[0233] To identify peptide-binding proteins with high specificity for the compound of the invention, an initial round of mutagenesis may be conducted using error-prone PCR. This technique introduces random mutations throughout the peptide-binding protein, allowing for the identification of positions potentially involved in peptide binding. Following this, a second round of mutagenesis may be performed using saturated mutagenesis at the positions identified by error-prone PCR. This approach systematically introduces all possible amino acid substitutions at these specific sites, enabling the verification of their involvement in peptide binding and the identification of preferred amino acid residues. Additionally, multi-site saturation mutagenesis can be employed to explore the combinatorial effects of mutations at multiple positions, further refining the binding specificity and affinity of the peptide-binding protein for the compound of the invention. This iterative process of mutagenesis and selection facilitates the development of highly specific and optimized peptide-binding proteins.

[0234] In Example 4, the inventors have identified positions in E. coli OppA (SEQ ID NO:1) by error- prone PCR that potentially play a role in binding of the compound of the invention. These positions involve: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, 5460, V482, L531 and / or N533 of SEQ ID NO:1.

[0235] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein in step (a) one or more mutations are introduced at positions V60, S63, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, C443, W442, T429, D445, S460, V482, L531 and / or N533 of SEQ ID NO:1.

[0236] Positions of particular interest in E. coli OppA (SEQ ID NO:1) for binding the compound of the invention are R439 and S460. Mutations in these positions appeared more than once when error-prone PCR was used to identify mutations that can confer higher specificity to the compound of the invention. Similarly, mutations in position D221 and W222 were identified in the error-prone PCR screen, indicating that these neighboring positions are involved in binding of the compound of the invention.

[0237] Thus, in a preferred embodiment, the invention relates to the method according to the invention, wherein in step (a) one or more mutations are introduced at positions D221, W222, R439, and / or S460 of SEQ ID NO:1.

[0238] In certain embodiments, the nucleic acid encoding the peptide-binding protein comprising one or more mutations is comprised in a plasmid, in particular an expression plasmid comprising a promoter for expression of the peptide-binding protein comprising one or more mutations. The term "plasmid" as used herein refers to a small, circular, double-stranded DNA molecule that is distinct from a cell's chromosomal DNA. Plasmids are capable of autonomous replication within a host cell and are commonly found in bacteria, although they can also be present in archaea and eukaryotic organisms. In the context of genetic engineering and biotechnology, plasmids are often used as vectors to introduce foreign genetic material into a host cell. They typically contain an origin of replication, selectable marker genes, and multiple cloning sites, which facilitate the insertion and expression of target genes.

[0239] In certain embodiments, the nucleic acid encoding the peptide-binding protein comprising one or more mutations is comprised in a library of nucleic acids, preferably a plasmid library, wherein each member of the library preferably encodes a peptide-binding protein comprising one or more mutation.

[0240] The nucleic acid encoding the peptide-binding protein and comprising one or more mutations may then be expressed in a cell. To achieve this, the nucleic acid, or a plurality of nucleic acids, first needs to be introduced into the cells. The skilled person is aware of methods for introducing a nucleic acid into a cell, including, but not limited to, transformation, transfection, electroporation, and viral transduction. These techniques facilitate the uptake and integration of the nucleic acid into the host cell, enabling the subsequent expression of the engineered peptide-binding protein for further analysis and characterization.

[0241] The cell may be any type of cell. Preferably, the cell is a prokaryotic cell, more preferably An E. coli cell. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the cell in step (b) is a prokaryotic cell, in particular an E. coli cell.

[0242] In a preferred embodiment, the nucleic acid encoding the peptide-binding protein comprising one or more mutations is a plasmid that is introduced into a cell, preferably a prokaryotic cell, using transformation and / or electroporation.

[0243] The term "transformation" refers to the process by which a cell takes up and expresses foreign DNA, such as a plasmid, from its surrounding environment. Transformation can be facilitated by various methods, including chemical treatment to make the cell membrane more permeable or by applying an electric field in a process known as electroporation, which creates temporary pores in the cell membrane through which the plasmid DNA can enter. These techniques enable the efficient introduction of the plasmid into the host cell, allowing for the expression of the engineered peptide-binding protein and subsequent analysis of its properties.

[0244] The skilled person is aware of the conditions necessary to achieve the expression of the peptide-binding protein comprising one or more mutations. Specifically, this involves the induction of the promoter that controls the expression of the mutated peptide-binding protein. These conditions may include the appropriate temperature, pH, and the presence of specific inducers or cofactors that activate the promoter. Additionally, the skilled person understands how to optimize these conditions to maximize protein expression, taking into account factors such as the choice of host cell, the strength of the promoter, and the efficiency of the translation and transcription machinery.

[0245] In certain embodiments, expression of the peptide-binding protein comprising one or more mutations is driven by an inducible or a constitutive promoter. Inducible promoters require specific conditions or the presence of certain molecules to initiate transcription. For example, the pBAD promoter is inducible by the sugar arabinose, allowing for controlled expression of the target gene. In contrast, constitutive promoters drive continuous expression of the gene regardless of environmental conditions. Other non-limiting examples of inducible promoters include the lac promoter, which is induced by the presence of lactose or its analog IPTG (isopropyl (3-D-l-thiogalactopyranoside), and the tac promoter, which is a hybrid of the trp and lac promoters and is also inducible by IPTG.

[0246] In contrast, constitutive promoters drive continuous expression of the gene regardless of environmental conditions. These promoters are always active and do not require any external inducers. Examples of constitutive promoters that can be used in bacteria include the lacllV5 promoter, which is a mutant form of the lac promoter with higher activity, the T7 promoter, which is recognized by T7 RNA polymerase for high-level expression, and the rpsL promoter, which is derived from the ribosomal protein 512 gene and provides strong, consistent expression. By selecting the appropriate promoter and optimizing the induction conditions, the desired levels of expression of the peptide-binding protein can be reliably achieved.

[0247] The cell used for identifying peptide-binding proteins with increased affinity and / or selectivity for the compound of the invention comprises an orthogonal translation system that enables the incorporation of an isopeptide-linked lysine or an analog thereof into a protein of interest. This incorporation serves as a readout for the efficiency with which the compound of the invention is transported into the cell by peptide transporters. By quantifying the incorporation of the isopeptide-linked lysine or its analog into the protein of interest, the performance of peptide-binding proteins, and their mutants, in mediating the uptake of the target compound can be assessed, thereby identifying those with enhanced affinity and / or selectivity and transport efficiency.

[0248] Cells expressing peptide-binding proteins with increased affinity and / or selectivity for the compound of the invention will accumulate higher concentrations of the compound inside the cell. Once inside, the compound is hydrolyzed or cleaved in any other suitable way to release an isopeptide-linked lysine or an analog thereof. The availability of the isopeptide-linked lysine or its analog within the cell for the orthogonal translation system directly correlates with the specificity of the peptide-binding protein for the compound of the invention. Therefore, the incorporation of the isopeptide-linked lysine or its analog into the protein of interest serve as effective indicators of the peptide-binding protein's affinity and / or selectivity for the compound.

[0249] The term "orthogonal translation system" as used herein refers to a specialized translation machinery within a cell that operates independently of the cell's native translation system. This orthogonal system typically includes an orthogonal aminoacyl-tRNA synthetase / tRNA pair that is specific for a non-canonical amino acid, such as an isopeptide-linked lysine or an analog thereof. The orthogonal aminoacyl-tRNA synthetase charges the orthogonal tRNA with the non-canonical amino acid, which is then incorporated into a protein of interest at designated codon positions that have been re-allocated for this purpose. This system allows for precise control over the incorporation of non-standard amino acids into proteins, enabling the study and engineering of proteins with novel properties and functions. In the context of the invention, the orthogonal translation system is used to monitor the uptake and incorporation of the compound of the invention, providing a readout for the efficiency and specificity of peptide-binding proteins in transporting the target compound into the cell.

[0250] The orthogonal aminoacyl-tRNA synthetase may be any aminoacyl-tRNA synthetase that is capable of attaching an isopeptide-linked lysine or an analog thereof to an orthogonal tRNA.

[0251] In certain embodiments, the amino acid residue at position Z of the peptide may be incorporated into a protein as a readout for efficient peptide uptake and subsequent cleavage. This approach is especially relevant in cases where a lysine analog at position E (XisoK) lacks the functional group required for aminoacylation onto tRNA and thus cannot be incorporated into a protein, serving solely as a carrier for Z. In such embodiments, the ncAA at position Z can be selectively incorporated into a protein via a cognate aminoacyl-tRNA synthetase (aaRS) / tRNA system, as demonstrated in Figure 5e (where only Z is incorporated) and Figure 5f (where both Z and XisoK are incorporated using two mutually orthogonal aaRS / tRNA systems). Thus, the incorporation of Z into a protein provides a direct and functional readout of the uptake and cleavage of the Z-XisoK compound, even when XisoK itself cannot be incorporated due to the absence of a lysine "stem" or suitable functional group.

[0252] However, it is preferred herein that the aminoacyl-tRNA synthetase is a pyrrolysyl-tRNA synthetase (PylRS). Thus, in a particular embodiment, the invention relates to the method according the invention, wherein the orthogonal aminoacyl-tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair.

[0253] Py rrolysy 1-tRN A synthetases are specific for their cognate tRNA and the non-canonical amino acid pyrrolysine. However, it has been demonstrated in the past that pyrrolysyl-tRNA synthetases also recognize other lysine-derivatives than pyrrolysine. For example, it was demonstrated in Example 3, that a PylRS from Methanosarcina barkeri (MbPyIRS; SEQ ID NO:2) can be used to incorporate various isopeptide linked lysins (XIsoK) into proteins. In addition, engineered variants of MbPyIRS were shown in Example 3 to have increased specificity for isopeptide linked lysins.

[0254] An overview of orthogonal PylRS / tRNAPylpairs are provided by Koch et al. (Int J Mol Sci. 2021 Oct 17;22(20):11194), Koch and Budisa (Chem Rev. 2024 Aug 28;124(16):9580-9608), and more recently by Dunkelmann and Chin (https: / / doi.org / 10.1021 / acs.chemrev.4c00243), which are incorporated herein in their entirety.

[0255] In certain embodiments, the PylRS is an archaeal PylRS. In certain embodiments, the PylRS is derived from a Methanosarcina species, such as Methanosarcina barkeri or Methanosarcina mazei. In certain embodiments, the PylRS is derived from a Methanomethylophilus species, such as Methanomethylophilus alvus.

[0256] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (Mb PylRS; SEQ ID NO:2) or a variant thereof. In certain embodiments, the MbPyIRS comprises one or more mutations at the following positions: Y271, and / or C313.

[0257] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (MbPyIRS) as set forth in SEQ ID NO:2.

[0258] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (MbPyIRS) as set forth in SEQ ID NO:3 comprising the mutation C313V.

[0259] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (MbPyIRS) as set forth in SEQ ID NO:4 comprising the mutations Y271L and C313T. In a particular embodiment, the invention relates to a protein comprising the amino acid set forth in SEQ ID NO:4.

[0260] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (MbPyIRS) as set forth in SEQ ID NO:79 comprising the mutations Y271A and L274M. In a particular embodiment, the invention relates to a protein comprising the amino acid set forth in SEQ ID NO:79. In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanomethylophilus alvus (MaPyIRS) as set forth in SEQ ID NO:6 comprising the mutations H227I and Y228P.

[0261] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina mazei (MmPyIRS) as set forth in SEQ ID NO:78 comprising the mutations Y306A and Y384F. In a particular embodiment, the invention relates to a protein comprising the amino acid set forth in SEQ ID NO:78.

[0262] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina mazei (MmPyIRS) as set forth in SEQ ID NQ:80 comprising the mutations L309A, N346Q, C348S. In a particular embodiment, the invention relates to a protein comprising the amino acid set forth in SEQ ID NQ:80.

[0263] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina barkeri (MbPylRS) as set forth in SEQ ID NO:2 comprising one or more mutations at positions: M241, M365, L266, A267, L270, Y271, L274, 1287, M309, N311, C313, M315, V348, Y349, S364, V366, V370, L372, 1378, W382, and / or G386. These positions of MbPylRS have been reported to be involved in substrate binding.

[0264] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanosarcina mazei (MmPyIRS) as set forth in SEQ ID NO:7 comprising one or more mutations at positions: M276, M300, L301, A302, L305, Y306, L309, 1322, M344, N346, C348, M350, V383, Y384, S399, V401, 1405, L407, 1413, W417, and / or G421. These positions of MmPyIRS have been reported to be involved in substrate binding.

[0265] In certain embodiments, the pyrrolysyl-tRNA synthetase is the pyrrolysyl-tRNA synthetase from Methanomethylophilus alvus (MaPyIRS) as set forth in SEQ ID NO:8 comprising one or more mutations at positions: M96, L121, A122, L125, Y126, M129, 1142, N166, V168, M170, V205, Y206, S221, A223, H227, L229, V235, W239, and / or G243. These positions of MaPylRS have been reported to be involved in substrate binding.

[0266] In certain embodiments, the amino acyl-tRNA synthetase is a tyrosyl-tRNA synthetase from Methanocaldococcus janaschii (MyTyrRS) as set forth in SEQ ID NO:9 comprising one or more mutations at positions: Y32, L65, H70, F108, Q109, D158, 1159, and / or L162. These positions of MyTyrRS have been reported to be involved in substrate binding.

[0267] In certain embodiments, the amino acyl-tRNA synthetase is a tyrosyl-tRNA synthetase from Archaeoglobus fulgidus (A / TyrRS) as set forth in SEQ ID NQ:10 comprising one or more mutations at positions: Y36, L69, H74, Q116, D165, and / or 1166. These positions of A / TyrRS have been reported to be involved in substrate binding.

[0268] In certain embodiments, the amino acyl-tRNA synthetase is a phosphoseryl-tRNA synthetase from Methanococcus maripaludis (MmSepRS) as set forth in SEQ ID NO:11 comprising one or more mutations at positions: E412, E414, P495, 1496, and / or F529. These positions of MmSepRS have been reported to be involved in substrate binding.

[0269] With the teaching provided herein, the skilled person would have no difficulty in identifying a pyrrolysyl-tRNA synthetase (PylRS) that is suitable for attaching an isopeptide-linked lysine, or an analog thereof, to its cognate tRNA. Furthermore, based on the teaching provided herein, identifying additional PylRS variants capable of attaching an isopeptide-linked lysine, or an analog thereof, to their cognate tRNA would not pose an undue burden on the skilled person. In this regard, it is worth noting that the process of identifying or evolving a new PylRS variant for a certain isopeptide-linked lysine amino acid is greatly facilitated by the higher concentration of the amino acid inside the cell due to the increased uptake of the compound of the invention. The methods and criteria for selecting and optimizing such PylRS variants are well within the expertise of those skilled in the art, ensuring efficient and accurate incorporation of the desired non-canonical amino acids into proteins.

[0270] The orthogonal translation system further comprises an orthogonal tRNA that is recognized by the aminoacyl-tRNA synthetase and can be loaded with the isopeptide-linked lysine or the analog thereof. The skilled person is aware of cognate aminoacyl-tRNA synthetase / tRNA pairs.

[0271] The orthogonal tRNA preferably comprises an engineered anticodon loop to enable incorporation of the isopeptide-linked lysine or the analog thereof in response to a reallocated codon. For example, where the re-allocated codon is the amber stop codon TAG, the anticodon loop is engineered to comprise the anticodon CUA.

[0272] The orthogonal translation system is used to incorporate one or more isopeptide-linked lysins or the analogs thereof into a reporter protein in response to one or more re-allocated codons comprised in the gene encoding the reporter protein.

[0273] The term "reporter protein" as used herein refers to a protein that is used as a marker to monitor and measure biological processes, gene expression, or cellular events. Reporter proteins produce a detectable signal, such as fluorescence, luminescence, or enzymatic activity, which can be easily quantified using various analytical techniques. Preferably, the reporter protein is a fluorescent protein which allows for real-time visualization and quantification of protein expression. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the reporter protein is a fluorescent protein.

[0274] The term "fluorescent protein" as used herein refers to a type of protein that exhibits fluorescence, meaning it can absorb light at a specific wavelength and subsequently emit light at a longer wavelength. This property makes fluorescent proteins valuable tools in molecular and cellular biology for visualizing and tracking biological processes in real-time. Fluorescent proteins are often used as reporter proteins to monitor gene expression, protein localization, and interactions within living cells. The most well-known fluorescent protein is green fluorescent protein (GFP), originally derived from the jellyfish Aequorea victoria. Variants of GFP and other naturally occurring fluorescent proteins have been engineered to emit light in different colors, such as blue (BFP), cyan (CFP), yellow (YFP), and red (RFP), expanding their utility in multicolor imaging and multiplexed assays.

[0275] However, the reporter protein can also be an antibiotic resistance marker or an auxotrophic selection marker.

[0276] Antibiotic resistance markers, such as P-lactamase or neomycin phosphotransferase, confer resistance to specific antibiotics, enabling the selection of successfully modified cells by allowing them to grow in the presence of the antibiotic, while cells that do not express the antibiotic resistance marker are inhibited or killed.

[0277] Auxotrophic selection markers, on the other hand, involve genes that complement a nutritional deficiency in the host cell. For example, a cell that is auxotrophic for a particular amino acid or nucleotide can only grow if the medium is supplemented with that nutrient. By introducing a gene that restores the ability to synthesize the missing nutrient, only the cells that successfully express the gene can grow in minimal media lacking the supplemented nutrient.

[0278] In the context of this invention, reporter proteins are utilized to indicate the successful incorporation of non-canonical amino acids, such as isopeptide-linked lysine or its analogs, into a protein of interest. The expression and activity of the reporter protein serve as a readout for the efficiency and specificity of the underlying biological process, facilitating the identification and characterization of peptide-binding proteins with enhanced affinity and / or selectivity for the compound of the invention.

[0279] To enable the expression of the reporter protein gene, the gene comprises one or more codons that have been re-allocated for the incorporation of the isopeptide-linked lysine or an analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair. Several codons have been described in the art that can be re-allocated for the incorporation of non-canonical amino acids. The most prominent example is the amber stop codon (TAG / UAG). In its native context, the amber stop codon signals the termination of protein synthesis. However, in an orthogonal translation system, the amber stop codon can be repurposed to incorporate non- canonical amino acids, such as isopeptide-linked lysine, into the protein of interest.

[0280] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the reporter protein gene in step (b) comprises one or more amber stop codons that have been re-allocated to incorporate the isopeptide-linked lysine or the analog thereof.

[0281] In embodiments where the re-allocated codon is the amber stop codon, the cell preferably lacks release factor 1 (RF1) to improve incorporation of the non-canonical amino acid in response to the amber stop codon.

[0282] Besides the amber stop codon other codons may be re-allocated for the incorporation of the isopeptide-linked lysine or an analog thereof. For example, the re-allocated codon may be the ochre stop codon (TAA / UAA), a quadruplet codon (as described for example by Guo and Niu (J Mol Biol. 2022 Apr 30; 434(8): 167346)) or a liberated sense codon (as described by Robertson et al., (Science 372, 1057-1062(2021). DOI:10.1126 / science.abg3029)).

[0283] It is to be noted that the cell may also comprise two or more mutually orthogonal translation systems and the reporter protein gene may comprise two or more codons re-allocated for the incorporation of non-canonical amino acids. This is demonstrated in Example 5, where a first orthogonal translation system is used to incorporate a non-canonical amino acid in response to a first re-allocated codon (TAA) and a second orthogonal translation system is used to incorporate an isopeptide-linked lysine in response to a second re-allocated codon (TAG).

[0284] In a next step, cells comprising an orthogonal translation system and expressing a peptide- binding protein comprising one or more mutations are contacted with the compound according to the invention.

[0285] Whether and how efficiently the compound according to the invention can be taken up by the cell depends on the affinity and / or selectivity of the peptide-binding protein for the compound according to the invention. That is, if the compound according to the invention is selectively bound by the peptide-binding protein, it will be efficiently transported into the cell. Inside the cell, the compound according to the invention will then by hydrolyzed or cleaved to release the isopeptide-linked lysine or the analog thereof, which can subsequently by incorporated into the reporter protein.

[0286] The term "contact" as used herein refers to the process of bringing two or more substances or entities into close proximity or direct interaction with each other. In the context of this invention, "contact" typically involves exposing a cell to a compound, such as the compound of the invention, under conditions that facilitate the interaction between the cell and the compound. This may include incubating the cell with the compound in a suitable medium, under appropriate temperature, pH, and other environmental conditions that support the uptake, binding, or reaction of the compound with cellular components. The term "contact" encompasses both direct physical interaction and indirect interactions facilitated by the surrounding environment, ensuring that the desired biological or chemical processes can occur effectively.

[0287] In certain embodiments, cells are contacted with the compound according to the invention in a growth medium. In certain embodiments, the growth medium is a defined growth medium. In certain embodiments, the defined growth medium contain no or little amounts of peptides that can compete with the compound according to the invention for binding to the peptide- binding protein.

[0288] However, if the aim of the method is to further increase the specificity of the peptide-binding protein for the compound according to the invention, the method may be repeated several times and the concentration of competing peptides in the medium may be increased in every reiteration. As such, peptide-binding proteins may be obtained that are highly specific for compounds comprising isopeptide-linked lysins or analogs thereof, but have only limited affinity to linear peptides.

[0289] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein in step (c), the cell is contacted with a compound according to the invention in the presence of a peptide that competes for binding to the peptide-binding protein.

[0290] As mentioned above, the compound according to the invention is hydrolyzed or cleaved inside the cell to release the isopeptide-linked lysine or an analog thereof. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the isopeptide-linked lysine or the analog thereof is obtained by intracellular cleavage of the compound of the invention.

[0291] If the isopeptide-linked lysine or its analog is linked to an amino acid residue via a peptide bond, the release can be catalyzed by intracellular peptidases. As demonstrated in the appended Examples, E. coli possesses various endogenous peptidases capable of catalyzing this reaction. Alternatively, the cell may be engineered to express a recombinant peptidase known for its promiscuous activity, thereby enhancing the efficiency of hydrolysis or cleavage and ensuring the effective release of the isopeptide-linked lysine or its analog.

[0292] In certain embodiments, the cell may recombinantly express one or more peptidase selected from: pepA (SEQ ID NO:55), pepB (SEQ ID NO:56), pepT (SEQ ID NO:57), pepN (SEQ ID NO:58), ddpX (SEQ ID NO:59), ydpE (SEQ ID NQ:60), ypdF (SEQ ID NO:61), frvX (SEQ ID NO:62), or sgcX (SEQ ID NO:63).

[0293] Of note, certain peptidases require the presence of a free N-terminal amino group to hydrolyze the peptide bond. Thus, it is preferred herein that when the bond C of the compound according to the invention is a peptide bond (CO-NH), then the group A is an amino group (-NH2).

[0294] In certain embodiments, the isopeptide-linked lysine or its analog may be linked to an amino acid residue via an ester bond. In such embodiments, the release can be catalyzed by intracellular esterases. Alternatively, the cell may be engineered to express a recombinant esterase known for its promiscuous activity, thereby enhancing the efficiency of hydrolysis and ensuring the effective release of the isopeptide-linked lysine or its analog.

[0295] In certain embodiments, the cell may recombinantly express an esterase from Bacillus subtilis (SEQ ID NO:64).

[0296] It is preferred herein that when the group A of the compound according to the invention is not NH2, the bond C is preferably an ester bond (i.e., C is -CO-O-, as defined herein), since esterases do not require the presence of an N-terminal amino group to recognize their targets.

[0297] In a next step, expression of the reporter protein gene is detected and, preferably, quantified in the cell. As explained herein above, the expression level of the reporter protein gene correlates with the efficiency with which the isopeptide-linked lysine, or the analog thereof, is incorporated into the reporter protein, which in turn corresponds with the efficiency with which the compound according to the invention is taken up by the cell, mediated by the engineered peptide-binding protein. Thus, the affinity and / or specificity of the peptide- binding protein for the compound according to the invention correlates with the expression of the reporter protein gene.

[0298] How the expression of the reporter protein gene can be detected and / or quantified depends on the type of reporter protein. If the reporter gene is a fluorescent protein, as described herein, expression of the reporter protein gene is preferably detected and quantified by flow cytometry. Flow cytometry is a powerful analytical technique used to measure and analyze the physical and chemical characteristics of individual cells or particles as they flow in a fluid stream through a beam of light, typically a laser. This technique allows for the rapid quantification of various parameters, such as cell size, granularity, and fluorescence intensity, which can be indicative of specific cellular properties or the presence of particular biomolecules. Flow cytometry is widely used in research and clinical settings for applications such as cell sorting, biomarker detection, and the assessment of cellular health and function.

[0299] A preferred embodiment of flow cytometry is fluorescence-activated cell sorting (FACS). FACS is a specialized type of flow cytometry that not only analyzes but also sorts cells based on their fluorescence characteristics. In FACS, cells are labeled with fluorescent markers. As the cells pass through the laser beam, their fluorescence is detected and measured. Based on the fluorescence intensity and other parameters, the cells are then sorted into different populations using an electrostatic deflection system. FACS allows for the precise isolation of specific cell types from heterogeneous mixtures.

[0300] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the detection of the expression of the reporter protein gene is performed using fluorescence-activated cell sorting (FACS).

[0301] The skilled person is capable of isolating cells expressing high levels of the reporter protein gene by flow cytometry. In particular, the skilled person is capable of applying gating strategies to select cells that exhibit high fluorescence levels, effectively isolating those with elevated expression of the reporter protein gene.

[0302] Alternatively, selection strategies may be applied to detect and / or quantify the expression of the reporter protein gene. For example, the reporter protein gene may be an antibiotic resistance gene, and successful incorporation of the isopeptide-linked lysine or an analog thereof will result in the expression of this gene. In such embodiments, the level of expression of the antibiotic resistance gene directly correlates with the cell's ability to tolerate higher concentrations of the corresponding antibiotic. Therefore, cells expressing higher levels of the reporter protein gene will survive and proliferate in the presence of increasing antibiotic concentrations, allowing for the effective detection and quantification of gene expression.

[0303] When using antibiotic resistance markers as the reporter protein gene, multiple rounds of mutagenesis may be performed, with increasing concentrations of the antibiotic applied in each round. This iterative selection process ensures the enrichment of cells that exhibit enhanced incorporation of the isopeptide-linked lysine or its analog. By progressively challenging the cells with higher antibiotic concentrations, the selection pressure drives the evolution of peptide-binding proteins with improved specificity and efficiency for the compound of the invention.

[0304] Alternatively, growth-based selection strategies may be applied to detect and / or quantify the expression of the reporter protein gene. For example, the reporter protein gene may be an auxotrophic marker gene that complements a nutritional deficiency in the host cell. Successful incorporation of the isopeptide-linked lysine or an analog thereof will result in the expression of the auxotrophic marker gene. In such embodiments, cells expressing the reporter protein gene will be able to grow in minimal media lacking the supplemented nutrient that the auxotrophic marker gene restores. The higher the expression of the auxotrophic marker gene, the better the growth and survival of the cells under these selective conditions.

[0305] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the detection of the expression of the reporter protein gene is performed using antibiotic resistance screening and / or using growth-based selection.

[0306] In the final step, cells that have been detected and isolated based on their ability to express the reporter protein gene are determined to express a peptide-binding protein having increased affinity and / or selectivity for the compound of the invention. These isolated cells may be further analyzed to verify that the enhanced expression of the reporter protein gene correlates with the improved binding characteristics of the peptide-binding protein. This evaluation may include additional biochemical assays, binding studies, or functional tests to ensure that the peptide-binding protein exhibits the desired specificity and affinity for the target compound.

[0307] Preferably, identifying an engineered peptide-binding protein as having increased affinity and / or selectivity for the compound according to the invention involves a step of determining the nucleic acid sequence of the gene encoding the engineered peptide-binding protein, and / or identifying specific mutations in the gene encoding the engineered peptide-binding protein.

[0308] That is, cells expressing high levels of the reporter protein gene may be isolated and the gene encoding the engineered peptide-binding protein may be sequenced to identify the nucleic acid sequence encoding an engineered peptide-binding protein having increased affinity and / or selectivity for the compound according to the invention. Sanger sequencing, as known in the art, may be used to identify the nucleic acid sequence encoding an engineered peptide- binding protein. Alternatively, a plurality of cells that have been identified to express engineered peptide- binding proteins with increased affinity and / or selectivity for the compound of the invention may be subjected to deep sequencing. This approach allows for the comprehensive analysis of the genetic diversity within the population, enabling the identification of specific positions and / or mutations in the peptide-binding protein that are critical for binding the compound. By examining the frequency and enrichment of particular mutations, the genetic changes that contribute to the enhanced binding characteristics can be pinpointed. The skilled person is aware of methods to perform deep sequencing and subsequent bioinformatic analyses to identify such positions and / or mutations. This detailed genetic information provides valuable insights into the molecular basis of the improved binding properties, facilitating the rational design and further optimization of peptide-binding proteins for the compound of the invention.

[0309] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein identifying the engineered peptide-binding protein as having increased specificity for the compound comprises a step of determining the sequence of the nucleic acid encoding the peptide-binding protein and / or determining specific positions and / or mutations in the nucleic acid encoding the peptide-binding protein.

[0310] The method described herein above was successfully applied by the inventors to identify engineered variants of the peptide-binding protein OppA from E. coli having increased affinity and selectivity for the compound according to the invention. Thus, another aspect of the invention relates to an engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, 5460, V482, L531 and / or N533.

[0311] That is, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V60. Improved variants of OppA were identified using the method of the invention comprising the mutations V60D, V60S, V60Q, and V60E. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V60, preferably wherein the valine at position 60 is replaced by aspartate, glutamate, serine or glutamine.

[0312] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position S63. Improved variants of OppA were identified using the method of the invention comprising the mutations S63G, S63A, S63C and S63M. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position 563, preferably wherein the serine at position 63 is replaced by glycine, alanine, cysteine or methionine.

[0313] That is, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position L78. An improved variant of OppA was identified using the method of the invention comprising the mutation L78L Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position L78, preferably wherein the leucine at position 78 is replaced by an isoleucine or another hydrophobic amino acid residue such as methionine, alanine or valine.

[0314] That is, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position Y135. An improved variant of OppA was identified using the method of the invention comprising the mutation Y135L. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position Y135, preferably wherein the tyrosine at position 135 is replaced by leucine or another hydrophobic amino acid residue such as alanine, valine or isoleucine.

[0315] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position T173. An improved variant of OppA was identified using the method of the invention comprising the mutation T173N. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position T173, preferably wherein the threonine at position 173 is replaced by an asparagine or another polar amino acid residue such as serine or glutamine.

[0316] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position H187. Improved variants of OppA were identified using the method of the invention comprising the mutations H187N, H187G and H187L. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position H187, preferably wherein the histidine at position 187 is replaced by asparagine, glycine or leucine.

[0317] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V193. An improved variant of OppA was identified using the method of the invention comprising the mutation V193A. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V193, preferably wherein the valine at position 193 is replaced by an alanine or another hydrophobic amino acid residue such as methionine, leucine, or isoleucine.

[0318] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position D221. Improved variants of OppA were identified using the method of the invention comprising the mutation D221G, D221L or D221M. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position D221, preferably wherein the aspartate at position 221 is replaced by a glycine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position D221, preferably wherein the aspartate at position 221 is replaced by a leucine or methionine or another hydrophobic amino acid residue such as alanine, valine or isoleucine.

[0319] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position W222. Improved variants of OppA were identified using the method of the invention comprising the mutations W222R and W222A. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position W222, preferably wherein the tryptophan at position 222 is replaced by an arginine or another positively charged amino acid residue such as histidine or lysine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position W222, preferably wherein the tryptophan at position 222 is replaced by an alanine or another hydrophobic amino acid residue such as valine, methionine, leucine or isoleucine.

[0320] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position 1303. An improved variant of OppA was identified using the method of the invention comprising the mutation 1303V. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position 1303, preferably wherein the isoleucine at position 303 is replaced by a valine or another hydrophobic amino acid residue such as methionine, leucine, or alanine.

[0321] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position K307. An improved variant of OppA was identified using the method of the invention comprising the mutation K307R. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position K307, preferably wherein the lysine at position 307 is replaced by an arginine or another positively charged amino acid residue such as histidine.

[0322] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position K333. An improved variant of OppA was identified using the method of the invention comprising the mutation K333N. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position K333, preferably wherein the lysine at position 333 is replaced by an asparagine or another polar amino acid residue such as serine, threonine or glutamine.

[0323] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position N337. An improved variant of OppA was identified using the method of the invention comprising the mutation N337D. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position N337, preferably wherein the asparagine at position 337 is replaced by an aspartate or another negatively charged amino acid residue such as glutamate.

[0324] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position K371. An improved variant of OppA was identified using the method of the invention comprising the mutation K371E. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position K371, preferably wherein the lysine at position 371 is replaced by a glutamate or another negatively charged amino acid residue such as aspartate.

[0325] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppAfrom E. coli (SEQ ID NO:1) comprises a mutation at position R439. Improved variants of OppA were identified using the method of the invention comprising the mutations R439H, R439Q, R439L, R439A, or R439D. Thus, in certain embodiments, the engineered variant of an oligopeptide- binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position R439, preferably wherein the arginine at position 439 is replaced by a histidine or another positively charged amino acid residue such as lysine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position R439, preferably wherein the arginine at position 439 is replaced by a glutamine or another polar amino acid residue such as serine, threonine or asparagine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position R439, preferably wherein the arginine at position 439 is replaced by a leucine or alanine or another hydrophobic amino acid residue such as methionine, isoleucine or valine. In another embodiment, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position R439, wherein the arginine at position 439 is replaced by aspartate or another negatively charged amino acid residue such as glutamate.

[0326] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position C443. Improved variants of OppA were identified using the method of the invention comprising the mutations C443L and C443G. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position C443, preferably wherein the cysteine at position 443 is replaced by leucine or glycine.

[0327] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position D445. Improved variants of OppA were identified using the method of the invention comprising the mutation D445L. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position D445, preferably wherein the aspartate at position 445 is replaced by leucine or another hydrophobic amino acid residue such as alanine, valine or isoleucine.

[0328] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position T429. An improved variant of OppA was identified using the method of the invention comprising the mutation T429A. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position T429, preferably wherein the threonine at position 429 is replaced by an alanine or another hydrophobic amino acid residue such as methionine, leucine, isoleucine or valine.

[0329] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position S460. Improved variants of OppA were identified using the method of the invention comprising the mutations S460N, S460H or S460C. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position S460, preferably wherein the serine at position 460 is replaced by an asparagine or another polar amino acid residue such as threonine or glutamine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position S460, preferably wherein the serine at position 460 is replaced by a histidine or another positively charged amino acid residue such as lysine or arginine. In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position 5460, preferably wherein the serine at position 460 is replaced by a cysteine or another polar amino acid residue such as serine, threonine, asparagine or glutamine.

[0330] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V482. An improved variant of OppA was identified using the method of the invention comprising the mutation V482A. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position V482, preferably wherein the valine at position 482 is replaced by an alanine or another hydrophobic amino acid residue such as methionine, leucine, or isoleucine.

[0331] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position L531. Improved variants of OppA were identified using the method of the invention comprising the mutations L531G and L531C. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position L531, preferably wherein the leucine at position 531 is replaced by glycine or cysteine.

[0332] In certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position N533. Improved variants of OppA were identified using the method of the invention comprising the mutations N533S, N533G and N533A. Thus, in certain embodiments, the engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprises a mutation at position N533, preferably wherein the asparagine at position 533 is replaced by serine, glycine or alanine.

[0333] In certain embodiments, the engineered OppA variant according to the invention comprises at least one mutation in a position selected from the group consisting of: T173, V193, D221, W222, 1303, N337, R439, and S460. Thus, in a particular embodiment, the invention relates to the engineered variant of the invention, wherein the one or more mutations are located in positions: T173, V193, D221, W222, 1303, N337, R439, and / or S460.

[0334] In certain embodiments, the engineered OppA variant according to the invention comprises at least one mutation in a position selected from the group consisting of: D221, W222, N337, R439, and 5460. Thus, in a particular embodiment, the invention relates to the engineered variant of the invention, wherein the one or more mutations are located in positions: D221, W222, N337, R439, and / or 5460.

[0335] In certain embodiments, the engineered OppA variant according to the invention comprises at least one mutation in a position selected from the group consisting of: D221, W222, R439, and 5460. Thus, in a particular embodiment, the invention relates to the engineered variant of the invention, wherein the one or more mutations are located in positions: D221, W222, R439, and / or 5460.

[0336] In certain embodiments, the engineered OppA variant according to the invention comprises at least one mutation in a position selected from the group consisting of: R439, and 5460. Thus, in a particular embodiment, the invention relates to the engineered variant of the invention, wherein the one or more mutations are located in positions: R439, and / or 5460.

[0337] In certain embodiments, the engineered OppA variant according to the invention comprises at least one mutation in a position selected from the group consisting of: Y135, H187, W442, and D445. Thus, in a particular embodiment, the invention relates to the engineered variant of the invention, wherein the one or more mutations are located in positions: Y135, H187, W442, and / or D445.

[0338] The mutations are preferably any of the mutations listed herein above. That is, in certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: L78I, Y135L, T173N, H187L, V193A, D221G, D221L, D221M, W222R, W222A, 1303V, K307R, K333N, N337D, K371E, R439H, R439Q, R439L, R439A, R439D, D445L, T429A, S460N, S460H, S460C, and / or V482A.

[0339] In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: T173N, V193A, D221G, D221L, D221M, W222R, W222A, 1303V, N337D, R439H, R439Q, R439L, R439A, R439D, S460N, S460H, and / or S460C.

[0340] In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: D221G, D221L, D221M, W222R, W222A, N337D, R439H, R439Q, R439L, R439A, R439D, S460N, S460H, and / or S460C.

[0341] In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: D221G, D221L, D221M, W222R, W222A, R439H, R439Q, R439L, R439A, R439D, S460N, S460H, and / or S460C.

[0342] In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: R439H, R439Q, R439L, R439A, R439D, S460N, S460H, and / or S460C.

[0343] In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations: R439Q and / or S460C. In certain embodiments, the invention relates to an engineered OppA variants comprising one or more of mutations at: position V60, preferably wherein the mutation is V60E, V60D, V60S, or V60Q; position 563, preferably wherein the mutation is S63G, S63A, S63C, or S63M; position H187, preferably wherein the mutation is H187N or H187G; position R439, preferably wherein the mutation is R439H, or R439S; position C443, preferably wherein the mutation is C443L, or C443G; position L531, preferably wherein the mutation is L531G, or L531C, and / or position N533, preferably wherein the mutation is N533S, N533A, or N533G.

[0344] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: V60E, S63G, L531G, N533S. In a preferred embodiment, the invention relates to an engineered OppA variant comprising an amino acid sequence as set forth in SEQ ID NO:69.

[0345] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: S63A, R439H, L531G, N533A. In a preferred embodiment, the invention relates to an engineered OppA variant comprising an amino acid sequence as set forth in SEQ ID NO:70.

[0346] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: V60D, S63C, L531C, N533G.

[0347] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: V60S, S63G, R439S, L531G, N533A.

[0348] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: S63G, H187N, C443L.

[0349] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: V60Q, S63M, H187G, C443G.

[0350] In one embodiment, the invention relates to an engineered OppA variant comprising one or more of the following mutations: Y135L, H187L, D445L.

[0351] In one embodiment, the invention relates to an engineered OppA variant comprising the following mutations: R439D. Furtherencompassed are variants of the E. coli OppA protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95 sequence identity to SEQ ID NO:1, provided that they comprise one or more of the mutations disclosed herein above and show increased affinity and / or selectivity for the compound according to the invention compared to wild type OppA (SEQ ID NO:1).

[0352] The present invention further relates to a nucleic acid encoding the engineered OppA variants disclosed herein. Thus, in a particular embodiment, the invention relates to a nucleic acid encoding the engineered variant according to the invention.

[0353] The skilled person will recognize that these nucleic acids can be synthesized using standard molecular biology techniques and can be introduced into suitable host cells for expression. These nucleic acids may include regulatory elements such as promoters, enhancers, and terminators to ensure efficient transcription and translation of the engineered OppA variants.

[0354] Additionally, the nucleic acids can be incorporated into various vectors, including plasmids, for ease of manipulation and transfer into host cells. Thus, in a particular embodiment, the invention relates to a vector comprising the nucleic acid according to the invention.

[0355] By encoding the specific mutations identified in the engineered OppA variants, these nucleic acids enable the production of peptide-binding proteins with enhanced affinity and / or selectivity for the compound of the invention.

[0356] The nucleic acid and / or vector encoding the engineered OppA variants may be introduced into a host cell. Thus, in a particular embodiment, the invention relates to a cell comprising the nucleic acid according to the invention or the vector according to the invention.

[0357] That is, the cell may comprise a vector or plasmid encoding the engineered OppA variant. The skilled person is aware of methods to introduce a vector or plasmid into a cell and maintain it inside the cell, such as transformation, transfection and electroporation. These techniques ensure the stable presence and expression of the engineered OppA variant within the host cell.

[0358] Alternatively, the nucleic acid encoding the engineered OppA variant may be integrated into the genome ofthe host cell. This can be achieved using methods such as the Datsenko-Wanner method, which allows for precise genomic integration through homologous recombination (Proc Natl Acad Sci U S A. 2000 Jun 6;97(12):6640-5) or viral transduction. By integrating the nucleic acid into the host genome, stable and long-term expression of the engineered OppA variant can be achieved, providing a reliable system for studying and utilizing the enhanced peptide-binding properties of the protein.

[0359] The host cell may be any type of cell. It is, however, preferred herein that the host cell is a prokaryotic cell. The term "prokaryotic cell" as used herein refers to a type of cell that lacks a membrane-bound nucleus and other membrane-bound organelles. Prokaryotic cells are characterized by their relatively simple structure, with genetic material organized in a single, circular chromosome located in a region called the nucleoid. Additionally, prokaryotic cells may contain plasmids, which are small, circular DNA molecules that replicate independently of the chromosomal DNA. Prokaryotic cells include organisms from the domains Bacteria and Archaea. Preferably, the cell according to the invention is a bacterial cell. Even more preferably, the cell according to the invention is an E. coli cell. Thus, in a particular embodiment, the invention relates to the cell according to the invention, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

[0360] Another aspect of the invention relates to a method for incorporating a non-canonical amino acid into a protein, wherein the non-canonical amino acid is derived from the compound according to the invention.

[0361] Thus, in a particular embodiment, the invention relates to a method for incorporating a non- canonical amino acid, or an analog thereof, into a protein, the method comprising the steps of: a) providing a cell comprising at least one orthogonal translation system, wherein the at least one orthogonal translation system comprises an orthogonal aminoacyl-tRNA synthetase / tRNA pair specific for a non-canonical amino acid comprised in the compound according to the invention, and a nucleic acid molecule encoding the protein, wherein the nucleic acid molecule encoding the protein comprises one or more codons that have been re-allocated for the incorporation of the non-canonical amino acid by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; b) contacting the cell of step (a) with the compound according to the invention, wherein the compound is transported into the cell and hydrolyzed or cleaved inside the cell to release an isopeptide-linked lysine or an analog thereof and, optionally, another non- canonical amino acid or an analog thereof; c) incorporating (i) the isopeptide-linked lysine or an analog thereof and / or (ii) the other non-canonical amino acid or an analog thereof into the protein in response to the one or more re-allocated codon using the orthogonal translation system.

[0362] It has been convincingly demonstrated herein that the compound according to the invention, which comprises an isopeptide-linked lysine or an analog thereof, can be efficiently taken up by a cell through endogenous peptide transporters or engineered variants thereof. Notably, the import of the compound according to the invention has been shown to be more efficient than the uptake of isopeptide-linked lysine, or its analog, alone. This enhanced uptake efficiency underscores the effectiveness of the compound in leveraging the cellular transport mechanisms, thereby facilitating its intracellular delivery and subsequent utilization.

[0363] The method for incorporating a non-canonical amino acid into a protein starts with a cell comprising at least one orthogonal translation system. The orthogonal translation system may comprise any of the aminoacyl-tRNA synthetase / tRNA pairs that have been disclosed elsewhere herein for the incorporation of the isopeptide-linked lysine or an analog thereof. That is, the orthogonal translation system may be used for incorporating an isopeptide-linked lysine or an analog thereof, obtained by intracellular cleavage of the compound according to the invention, into a protein of interest.

[0364] In such embodiments, the cell preferably comprises a pyrrolysyl-tRNA synthetase / tRNA pair for the incorporation of the isopeptide-linked lysine or an analog thereof. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the orthogonal aminoacyl-tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair.

[0365] When the isopeptide-linked lysine or an analog thereof is incorporated into a protein using an orthogonal translation system, it is to be understood that the isopeptide-linked lysine or analog thereof comprises an alpha amino group and an alpha carboxyl-group.

[0366] Alternatively or in addition, the cell may comprise another orthogonal translation system that does not recognize the isopeptide-linked lysine or an analog thereof, but another amino acid comprised in the compound according to the invention. That is, the compound according to the invention may not just be used to transport isopeptide-linked lysines or analogs thereof into a cell, but may also be used to transport other amino acids or analogs thereof that would otherwise not be efficiently taken up by the cell. This other amino acid or analog thereof is preferably linked to the isopeptide-linked lysine or analog thereof by a bond that can be hydrolysed inside the cell.

[0367] In certain embodiments, the analog of the other amino acid is an alpha-hydroxy acid where the alpha-amino group is replaced with a hydroxy group. Incorporation of such amino acid analogs is described, without limitation, by Owczarek et al. (Biochemistry, 2008; 47( 1) :301-7).

[0368] In certain embodiments, the analog of the other amino acid may be a mercapto acid where the alpha-amino group is replaced with a thiol group.

[0369] To stick with the nomenclature of the present invention, the other or second (non-canonical) amino acid would comprise the moiety B comprising the side chain Z (or Z1and Z2). More preferably, the other amino acid would have the structure NH2-B-COOH once it is released from the compound according to the invention by a hydrolyzing enzyme. The other amino acid analog would preferably have the structure HO-B-COOH or HS-B-COOH once it is released from the compound according to the invention by a hydrolyzing enzyme, preferably an esterase.

[0370] In embodiments where the compound of the invention is a peptide having the structure Z- XisoK, the other amino acid would be residue Z.

[0371] Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein hydrolysis of the molecule according to the invention inside the cell releases a second non-canonical amino acid (ncAA), and wherein the second ncAA is incorporated into the same or a different protein in response to a re-allocated codon.

[0372] There is no limitation regarding the other amino acid or its analog. That is, the other amino acid (Z) may, inter alia, be any non-canonical amino acid. The skilled person is capable of selecting an amino acid or analog thereof for which an orthogonal translation system is known in the art.

[0373] The cell further comprises a nucleic acid encoding a protein of interest, specifically the protein into which one or more non-canonical amino acids are to be incorporated. This nucleic acid sequence is designed to include codons that have been re-allocated for the incorporation of non-canonical amino acids by the orthogonal translation system(s). The presence of this nucleic acid ensures that the protein of interest is expressed within the cell, allowing for the precise incorporation of the desired non-canonical amino acids at specific positions.

[0374] Of note, the nucleic acid encoding the protein of interest may comprise one or more codons that have been re-allocated for the incorporation of non-canonical amino acids by the orthogonal aminoacyl-tRNA synthetase / tRNA pair(s).

[0375] That is, in certain embodiments, the nucleic acid may comprise one or more codons that have been re-allocated for the incorporation of an isopeptide-linked lysine or an analog thereof by an orthogonal aminoacyl-tRNA synthetase / tRNA pair.

[0376] In certain embodiments, the nucleic acid may comprise one or more codons that have been re-allocated for the incorporation of another amino acid or analog thereof by an orthogonal aminoacyl-tRNA synthetase / tRNA pair, wherein the other amino acid or analog thereof is also comprised in the compound according to the invention. In certain embodiments, the nucleic acid may comprise one or more codons that have been re-allocated for the incorporation of an isopeptide-linked lysine or an analog thereof by a first orthogonal aminoacyl-tRNA synthetase / tRNA pair and one or more codons that have been reallocated for the incorporation of another amino acid or analog thereof by a second orthogonal aminoacyl-tRNA synthetase / tRNA pair, wherein the other amino acid or analog thereof is also comprised in the compound according to the invention.

[0377] In a particular embodiment, the invention relates to the method according to the invention, wherein the nucleic acid molecule encoding the protein comprises one or more amber stop codons that have been re-allocated to incorporate the non-canonical amino acid.

[0378] As described herein above, the amber stop codon is well established for the incorporation of non-canonical amino acids. However, also other codons may be re-allocated for the incorporation of non-canonical amino acids, including the ochre stop codon, quadruplet codons or liberated sense codons, as described in more detail elsewhere herein.

[0379] In embodiments where the cell comprises two orthogonal translation systems, the nucleic acid encoding the protein of interest may comprise two re-allocated codons for the incorporation of non-canonical amino acids. This allows for the simultaneous incorporation of different non- canonical amino acids into a single protein, thereby enabling the creation of proteins with novel properties and functionalities.

[0380] Alternatively, the cell may comprise two separate nucleic acids encoding different proteins of interest. In this case, the first nucleic acid may comprise one or more codons that are reallocated for the incorporation of a non-canonical amino acid by a first orthogonal translation system, while the second nucleic acid may comprise one or more codons that are re-allocated for the incorporation of a non-canonical amino acid by a second orthogonal translation system. This configuration allows forthe independent expression and incorporation of distinct non-canonical amino acids into different proteins within the same cell, thereby expanding the range of possible protein modifications and applications.

[0381] The protein encoded by the nucleic acid may be any protein, including: an antibody or antibody fragment (including nanobodies), a fluorescent protein (such as sfGFP), Ubiquitin, SUMO1, SUMO2, a histone (such as Histone H3), a Proliferating cell nuclear antigen (PCNA), Calmodulin, (3-lactamase, Hsp82, a cytokine (such as interleukin-2, IFN-(3 or IFN- y), or a hormone (such as Human growth hormone (hGH)).

[0382] Antibodies may include full length antibodies such as, without limitation Trastuzumab. An antibody fragment is a portion of an antibody that retains the ability to bind to an antigen, and is preferably a single-chain fragment. Single-chain fragments, such as single-chain variable fragments (scFvs) and nanobodies, are engineered to consist of the variable regions of the heavy (VH) and light (VL) chains connected by a short flexible linker, allowing them to retain the antigen-binding specificity of the original antibody while being smaller and more stable. Examples of single-chain fragments include scFvs, which are fusion proteins of the VH and VL regions, and nanobodies, also known as VHHs or single-domain antibodies, which are derived from the unique heavy-chain-only antibodies found in camelids (such as camels, llamas, and alpacas). Nanobodies are particularly noted for their small size, high stability, and strong binding affinity, making them valuable in therapeutic treatments, diagnostic assays, and research applications. Antibody fragments may bind to any target including, without limitation, eGFP or HER2. In certain embodiments, the antibody fragment is a Fab fragment. The Fab fragment may be further PEGylated.

[0383] The cell may be any type of cell, as explained for the other methods disclosed herein. However, it is preferred herein that the cell is a prokaryotic cell, more preferably a bacterial cell, most preferably an E. coli cell. Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

[0384] It has been demonstrated herein that endogenous peptide transporters can import the compound according to the invention. Thus, the method of the invention is not restricted to cells comprising an engineered peptide-binding protein, such as any one of the engineered peptide-binding proteins disclosed herein.

[0385] The most promising endogenous peptide transporters for the uptake of the compound according to the invention are the dipepitde permease (Dpp) and, in particular, the oligopeptide permease (Opp). Thus, in a particular embodiment, the invention relates to the method according to the invention, wherein the cell expresses a dipeptide permease and / or an oligopeptide permease

[0386] However, the inventors have observed that the compound according to the invention competes with linear peptides for endogenous peptide transporters. Therefore, it may be desirable to utilize peptide transporters that have been engineered for enhanced affinity and / or selectivity for the compound according to the invention.

[0387] Accordingly, in a particular embodiment, the invention relates to the method according to the invention, wherein the dipeptide permease and / or the oligopeptide permease comprises a peptide-binding protein that has been engineered for increased affinity and / or selectivity for the compound according to invention.

[0388] Methods for obtaining engineered peptide-binding proteins with increased affinity and / or selectivity for the compound of the invention are disclosed herein.

[0389] In particular, the cell may comprise any one of the engineered OppA variants disclosed herein. That is, in a particular embodiment, the invention relates to a method according to the invention, wherein the engineered peptide-binding protein is OppA of E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, C443, T429, W442, D445, 5460, V482, L531 and / or N533.

[0390] With the teaching provided herein, the skilled person would have no difficulty engineering an OppA variant suitable for a given peptide. The application identifies specific positions within OppA that impact substrate specificity, enabling targeted mutagenesis to alter binding preferences. By focusing on these key residues, the skilled person can efficiently generate OppA variants optimized for the transport of desired peptides, going beyond the specific combinations of mutations disclosed herein.

[0391] To avoid lengthy repetition, it is to be understood that all combinations of positions and / or specific mutations in OppA disclosed elsewhere herein can be used for the method of incorporating non-canonical amino acids into a protein.

[0392] In a particularly preferred embodiment, the cell is a cell expressing an oligopeptide permease comprising a wild-type or engineered peptide-binding protein OppA.

[0393] In an even more preferred embodiment, the cell is a cell expressing an oligopeptide permease comprising a wild-type or engineered peptide-binding protein OppA from E. coli.

[0394] Of note, the cell may comprise a heterologous OppA variant, meaning an OppA variant derived from a different organism. For example, a wild-type or engineered E. coli OppA variant may be heterologously expressed in another organism to enhance the uptake of the compound of the invention. In other embodiments, the entire oligopeptide permease system of E. coli, including a wild-type or engineered E. coli OppA variant, may be heterologously expressed in another organism to further improve the uptake of the compound of the invention.

[0395] Alternatively, a DppA or OppA variant from another organism that has been engineered using the method disclosed herein herein may be heterologously expressed in E. coli cells to facilitate the uptake of the compound according to the invention. The cell is then contacted with the compound according to the invention as described elsewhere herein. Contacting the cell with the compound according to the invention will result in uptake of the compound according to the invention into the cell and the subsequent hydrolysis of the compound according to the invention to release an isopeptide-linked lysine or an analog thereof and, optionally, another non-canonical amino acid.

[0396] Once released inside the cell, the isopeptide-linked lysine or analog thereof and / or the other non-canonical amino acid may be incorporated into one or more proteins of interest using orthogonal translation systems.

[0397] Successful incorporation of non-canonical amino acids into proteins may be confirmed by methods known in the art, including mass spectrometry.

[0398] BRIEF DESCRIPTION OF THE DRAWINGS

[0399] Figure 1. Isopeptide-linked tripeptides are privileged scaffolds for efficient E. coli uptake a. Chemical structures for G-XisoK, XisoK and BocK. X is alanine for G-AisoK and AisoK. b. SDS- PAGE analysis of sfGFP-N150TAG expression using wt-MbPylRS / PylT in the presence of 2 mM BocK, AisoK or G-AisoK. (trunc. denotes truncated protein), c. LC-MS analysis of sfGFP- N150AisoK. d. Time course measurements of sfGFP fluorescence from cultures expressing sfGFP-N150TAG and wt-MbPylRS / PylT in the presence of G-AisoK, AisoK or BocK, or grown without ncAAs. e. Extracted ion chromatograms for determining intracellular concentrations of G-AisoK (teal) and AisoK (pink) by an LC-MS assay, performed on K12 cell extracts. Intracellular G-AisoK concentrations in K12 cells grown in the presence of 2 mM G-AisoK are negligible (dark grey, left). Intracellular AisoK concentrations in K12 cells grown in the presence of 2 mM G-AisoK (dark grey, right) are 5-10-fold higher than when grown in the presence of 2 mM AisoK (light grey, right), f. Proposed model for increased AisoK incorporation in the presence of G-AisoK. The tripeptide G-AisoK is actively taken up by K12 cells via a transporter. Within the cytoplasm G-AisoK is processed to AisoK, which is a substrate for wt-MbPyIRS / PyIT and is incorporated site-specifically into a protein of interest (POI).

[0400] Figure 2. The Opp transporter is responsible for efficient G-AisoK uptake, a. SDS-PAGE analysis and time course fluorescence measurements of sfGFP-N150TAG expression in the presence of BocK, AisoK or G-AisoK in K12 cells and in DoppA, DoppB or DoppD knockouts indicating that the Opp transporter is responsible for G-AisoK uptake. Results for other knockouts can be found in Fig. S3a. b. AlphaFold2(40) predicted structure of the Opp transporter consisting of the periplasmic peptide binding protein OppA, two transmembrane domains (TMDs) OppB,C and two nucleotide binding domains (NBDs) OppD,F. c. Proposed mechanism of Opp- mediated uptake. G-AisoK binds to OppA in the periplasm and is shuttled to membrane-bound OppB,C. The tripeptide is actively transported into the cytosol in an ATP dependent manner where it is cleaved by endogenous peptidases to AisoK. OppA in its apo-form is released from the TMDs to allow for binding of new G-AisoK. d. Extracted ion chromatograms of E. coli K12 lysates for determining intracellular AisoK concentrations in K12 versus oppA K12 cells. Genomic deletion of oppA results in undetectable AisoK concentrations when growing cells in the presence of 2 mM G-AisoK (grey, dashed). Consistent results were obtained over three distinct replicate experiments, e. SDS-PAGE analysis and time course fluorescence measurements of sfGFP-N150TAG expression in AoppA cells, constitutively expressing OppA variants (left: wt, right: D445A). Plasmid-based constitutive expression of wt-OppA rescues full-length sfGFP expression in presence of G-AisoK, while this is not the case for OppA-D445 expression, indicating that D445 is essential for G-AisoK binding and uptake. BocK dependent sfGFP expression remains unchanged in both cases. Arrow indicates full-length sfGFP, asterisk indicates truncated protein), f. SDS-PAGE analysis of sfGFP-N150TAG expression with BocK or G-AisoK in single peptidase knockouts ApepN and ApepA as well as the double knockout ApepN / pepA. G-AisoK-dependent full-length sfGFP expression is significantly reduced in ApepN / pepA, indicating that pepA and pepN are the main peptidases responsible for cleavage of the N-terminal glycine. Consistent results were obtained over three distinct replicate experiments. Arrow indicates full-length sfGFP, asterisks indicates truncated sfGFP.

[0401] Figure 3. A versatile G-XisoK toolbox, a. Structure of a generalized G-XisoK tripeptide, b. X-ray structure of OppA bound to G-SisoK (PDB ID: 9RD1). G-SisoK forms extensive interactions with OppA residues via its N- and C-termini as well as via its backbone amide groups, c. The OppA:G-SisoK complex around the serine side chain reveals a large cavity capable of accommodating bulky side chains, d. All functional groups incorporated via the G-XisoK scaffold, e. SDS-PAGE analysis of sfGFP-N150TAG expression in the presence of either 2 mM XisoK or G-XisoK. All G-XisoK derivatives show higher levels of full-length sfGFP expression using the corresponding Pyl RS / Py IT pairs in comparison to cells grown with the corresponding XisoK. Arrow indicates full-length sfGFP, asterisk indicates truncated sfGFP. f. LC-MS analysis of tyrosinase-mediated labelling of 3C-Ub bearing PisoK at K63 with p-cresol demonstrates quantitative conversion, g. SDS-PAGE analysis of CuAAC labeling of purified eGFPNb- R75PrgisoK with an Atto647-Azide fluorophore. No labeling is observed for eGFPNb-R75BocK. h. Western blot analysis of GST-dimer crosslinking for GST-E51pLisoK upon UV365 ™ illumination. No crosslink is observed for GST-wt. i. Western blot analysis of proximity-induced chemical crosslinking between Rablb-R79CIAisoK and its interactor DrrA-D512C339-522- Cells expressing both binding partners in the presence of G-CIAisoK display a higher molecular weight band corresponding to the crosslinked complex in both anti-H6 and anti-Strep blots, e- i) Consistent results were obtained over three distinct replicate experiments. Figure 4. Scalable XisoK incorporation through OppA evolution, a. SDS-PAGE analysis of sfGFP- N150TAG expression in the presence of SisoK or G-SisoK in Al media (left) or 2-YT (middle, right). In K12 cells full-length sfGFP expression yields are significantly reduced when using tryptone containing 2-YT (middle) media due to competition between tryptic peptides and G- SisoK for OppA binding. sfGFP expression in 2-YT is recovered when using the engineered lsoK12 strain (right), b. Scheme for OppA evolution for better G-SisoK uptake in tryptone containing media. OppA libraries were screened for increased G-SisoK uptake under increasing tryptone concentrations via amber suppression of sfGFP-N150TAG by sorting for fluorescent cells. Initial screening of an error-prone library yielded 4 variants which were used as basis for creating a cassette mutagenesis library, whose screening yielded OppA-ev2. c. Extracted ion chromatograms to determine intracellular SisoK concentrations. lsoK12 cells (purple) grown in 2-YT show 7-10-fold higher intracellular SisoK concentrations in comparison to K12 cells (grey) when adding 2 mM G-SisoK to 2-YT media, d. Affinity measurements of G-SisoK and a linear GSK peptide (GSK(lin)) towards OppA using microscale thermophoresis. e. SDS-PAGE analysis of sfGFP-N150TAG expression in wt-K12 cells versus lsoK12 cells grown in 2-YT media in presence of G-XisoK derivatives with X = P, Prg, CIA, C, pL. Full-length sfGFP expression is severely increased in lsoK12 cells for all X functionalities. Full gels as well as expression gels for other G-XisoK derivatives can be found in Fig. S12. f. SDS-PAGE analysis of sfGFP-N150TAG expression in 2xYT, comparing K12 with lsoK12 cells at different G-SisoK concentrations, g. SDS-PAGE analysis of Histone H3 with single (K122TAG), double (K79TAG, K122TAG) and triple (K27TAG, K79TAG, K122TAG) amber suppression in presence of 1 mM G-SisoK comparing K12 with lsoK12 cells grown in 2-YT media.

[0402] Figure 5. Generalized ncAA uptake using Z-AisoK tri peptides a. Structure of tested Z-residues within the Z-AisoK scaffold, b. SDS-PAGE analysis of sfGFP-N150TAG expression in the presence of 2 mM of BocK or Z-AisoK tripeptides with wt-MbPylRS / PylT. Expression levels of full-length sfGFP indicate successful transport of the Z-AisoK tripeptides, subsequent cleavage to Z and AisoK and AisoK incorporation. Top: expression in wt-K12. Bottom: Expression in evolved strain K12-Z2. Consistent results were obtained over three distinct replicate experiments. Arrow indicates full-length sfGFP, asterisk indicates truncated sfGFP. c. Scheme for OppA evolution to accommodate novel Z-AisoK substrates. Successful transport by OppA library variants was evaluated by incorporation of AisoK into sfGFP-N150TAG. Variants with high sfGFP fluorescence were enriched via three rounds of FACS. d. X-ray crystal structure of the OppA:G-SisoK complex highlighting four residues surrounding the N-terminal glycine of G-SisoK that were targeted for site-saturation mutagenesis to enable recognition of novel Z-AisoK substrates, e. Left: SDS-PAGE analysis of sfGFP-N150TAG expression with 0.25 mM LipK or the corresponding Z-AisoK tripeptide 13 with a LipK-specific MmPyIRS / PyIT variant (Y306A, Y384F). Expression levels are highest with peptide 13 in the K12-Z1 strain, which expresses OppA-Zl, evolved specifically for the transport of 13. Right: LC-MS analysis of sfGFP purified from K12-Z1 cultures grown with 13. Observed mass confirms incorporation of LipK. Consistent results were obtained over three distinct replicate experiments, f. Dual stop codon suppression using a single isopeptide-linked tripeptide. Left: Chemical structure of the tri peptide AcK-pLisoK (16). Middle: SDS-PAGE analysis of sfGFP-N40TAA-N150TAG expression in the presence of either AcK and G-pLisoK (or pLisoK) separately added to media or in the presence of tripeptide AcK-pLisoK (16) in lsoK12. Right: LC-MS analysis of purified sfGFP confirms dual ncAA incorporation (AcK and pLisoK) upon addition of tripeptide AcK-pLisoK (16). Consistent results were obtained over three distinct replicate experiments.

[0403] Figure 6. Substrates of OppA with different sidechains at 'Z' position (Position B): compounds of type Z-AisoK with Z being compounds 17-23 are transported by Opp transporter using wt- OppA or evolved variants of OppA as specified. Successful transport is confirmed by incorporation of AisoK into sfGFP via genetic code expansion.

[0404] Figure 7. Substrates of OppA with different linkages between Z and X (Position C): compounds of type Z-AisoK with an amide linkage between 'Z' and AisoK (Z-AisoK) or an ester linkage between 'Z' and AisoK (Z-LacK) with Z being compound 24-32 (amides) and 33-41 (esters) are transported by wt Opp. Successful transport is confirmed by incorporation of AisoK or LacK into sfGFP via genetic code expansion, measured by sfGFP fluorescence.

[0405] Figure 8. Transport of tripeptides that contain lysine analogues by oppA: Isopeptide linked compounds with analouges of lysine at the isopeptide linkage (position E) are also transported via the opp transporter. In peptides of type Lys(Ac)-Ala-stem where the stem is Lysine (42), ornithine (43), 6-aminohexanoic acid (44), 1,5 diaminopentane (45), Lysineamide (46) or 2,3- Diaminopropionic acid (48) are transported into E. coli in a oppA dependent manner. Succesful transport by opp is confirmed via the incorporation of Lys(Ac) into sfGFP via genetic code expansion in response to a amber stop codon. The absence of sfGFP expression or reduced sfGFP expression in E. coli opp . cells confirmed that transport of the compound is oppA dependant. For the tripeptide containing molecule 47 as ,stem' we did not observe any difference between K12 and opp . cells, indicating that transport of this tripeptide is not dependant on the Opp transporter.

[0406] Figure 9. Incorporation of hydroxy-acid lysine analogs. Isopeptide linked compounds with a hydroxy acid derivative of lysine at the isopeptide linkage (position E) can be transported by the opp transporter. Compounds of type G-XisoK(OH) are transported and cleaved in the cell to reveal XisoK(OH) which is incorporated into sfGFP via genetic code expansion. Successful transport is confirmed by the expression of full length sfGFP. An absence / reduction of sfGFP expression in E. coli opp . cells confirms that the transport is oppA dependent. Incorporation of hydroxy acids into proteins leads to an ester bond in the protein backbone which is partially cleaved during lysate preparation.

[0407] Figure 10. Incorporation of AzGGisoK using the MF18 Methanosarcina mazei PylRS harboring the mutations:L309A, N346Q, C348S. Uptake of AzGGisoK was improved by using an engineered OppA variant comprising mutations Y135L, H187L and D445L.

[0408] EXAMPLES

[0409] Example 1: Isopeptide-linked tripeptides are privileged scaffolds for efficient E. coli uptake

[0410] Previous work in our group has pioneered the use of transpeptidases in combination with genetic code expansion to generate defined protein-protein conjugates. Through site-specific incorporation of an ncAA bearing an azide-caged dipeptidic nucleophile acceptor (AzGGisoK) and subsequent Staudinger reduction, GGisoK-modified proteins can engage in transpeptidation with donor proteins modified at their C-terminus with an appropriate recognition sequence. We have leveraged these approaches for successfully generating ubiquitin (Ub)- and Ub-like modifier (Ubl)-POI conjugates using sortase or OaAEPl as transpeptidases (30-32). In our quest to diversify the linker sequence in the generated protein conjugates, we explored the site-specific incorporation of ncAAs resembling a general G-XisoK scaffold (Fig. la). While AzGGisoK required intensive PylRS engineering, we were not able to identify a PylRS-variant for direct GGisoK incorporation(30). In contrast, addition of the alanine-bearing G-XisoK tripeptide (G-AisoK, Fig. la) to growing E. coli cells (K12 strain), transformed with plasmids coding for the wild-type Methanosarcina barkeri Pyrrolysyl-tRNA synthetase / tRNA pair (wt-MbPyIRS / PyIT) and super-folder green fluorescence protein bearing an amber codon at position 150 (sfGFP-N150TAG), led to highly efficient sfGFP expression, comparable to wt-sfGFP production and amber suppression yields using the gold-standard ncAA BocK (Fig. la,b). Interestingly, MS-analysis of purified sfGFP expressed in presence of G- AisoK revealed site-specific incorporation of AisoK (Fig lc). This suggests that the N-terminal glycine is cleaved off intracellularly, either on the free ncAA, or co / post-translationally. On the other hand, supplementing K12 with AisoK instead of G-AisoK showed almost no visible amber suppression and sfGFP production (Fig. lb). This is corroborated by live-cell sfGFP- fluorescence measurements in the presence of different ncAAs. Cells grown in the presence of AisoK showed minimal fluorescence, while fluorescence is observed at earlier timepoints and reaches higher values with G-AisoK when compared to BocK (Fig. Id). Similarly, G-AisoK- mediated AisoK incorporation was observed for other amber codon-bearing target proteins. Intrigued by this observation, we set out to determine intracellular levels of the corresponding ncAAs, using LC-MS based uptake assays (18, 33). Upon G-AisoK addition, we could not observe any intracellular G-AisoK, but AisoK accumulated at 5-10-fold higher concentrations compared to supplementing K12 cells with AisoK itself (Fig le).

[0411] This observation led us to formulate the hypothesis that a specific transport mechanism may actively pump G-AisoK into cells. Within the cytosol G-AisoK is then enzymatically processed to AisoK, which accumulates in high concentrations and serves as a substrate for wt-PyIRS, leading to efficient AisoK encoding (Fig. If).

[0412] Example 2: A bacterial ABC-transporter facilitates tripeptide uptake

[0413] In gram-negative organisms such as E. coli, small peptidic substrates gain access to the periplasm via diffusion through outer membrane pore-forming proteins known as porins (34). Depending on the mode of action there are two major classes of active peptide transporters located in the inner membrane that shuffle specific peptides from the periplasm to the cytosol: (1) proton-dependent oligopeptide transporters (POT) that leverage a proton gradient for peptide import and (2) A BC-tra ns porters that use ATP hydrolysis as energy source for peptide import. Bacterial POTs are monomeric, multi-transmembrane (TM) helix-containing proteins that typically import various di- or tripeptides (35). Bacterial ABC-transporters associated with peptide uptake (e.g. Dpp-, Ddp- and Opp-transporters) require a periplasmic binding protein that delivers the captured peptidic substrate to a multi-subunit transmembrane complex for ATP-dependent promiscuous di- and oligopeptide transport.

[0414] In order to identify a potential uptake system for G-AisoK, we used single gene knockouts (36), which have individual transporter domains deleted, for amber suppression of sfGFP-N150TAG in the presence of the wt-PyIRS / PyIT pair and G-AisoK. We hypothesized that amber suppression of sfGFP would be abolished or diminished, if a potentially involved transporter system was absent. Deletion of the POT-family members and components of the ddp or dpp tranporters did not have any influence on efficient sfGFP production. In contrast, individual knockouts of genes constituting the opp operon led to complete abolishment of sfGFP expression in the presence of G-AisoK (Fig. 2a). The Opp transporter shares its general domain organization with other ABC transporters (37): it consists of the periplasmic binding protein (OppA), two transmembrane domains (TMDs, OppB and OppC) that span the periplasmic membrane and two cytoplasmic nucleotide-binding domains (NBDs, OppD and OppF, Fig. 2b) that bind and hydrolyze ATP. Docking of peptide-bound OppA to the TMDs triggers ATP binding and dimerization of the NBDs, leading to uptake of the substrate into the translocation channel. Hydrolysis of ATP results in dissociation of the NBD-dimerization interface and flips the TMDs to trigger release of the substrate into the cytosol. ADP-ATP exchange re-induces NBD-dimerization and the TMDs can again bind to OppA (Fig. 2c). Ill

[0415] Individual deletions of either the binding protein OppA, or any of the two TMDs or NBDs led to complete loss of amber suppression and sfGFP fluorescence in the presence of G-AisoK, while no differences in sfGFP-N150BocK expression yields were observed, indicating that the Opp transporter may be specifically involved in G-AisoK uptake (Fig. 2a). Confirming these expression results, uptake assays showed unchanged intracellular BocK concentrations between K12 cells and DoppA-K12 cells, while AisoK, which accumulated in millimolar concentrations in K12 cells upon addition of G-AisoK, was not detectable when the oppA gene was deleted (Fig. 2d).

[0416] Remarkably, plasmid-based constitutive expression of OppA in DoppA-K12 cells completely rescued amber suppression in the presence of G-AisoK (Fig. 2e).

[0417] OppA is known to promiscuously bind to 2-5 amino acid long peptides and shows some preference for tri- and tetrapeptides containing positively charged side chains. Previously reported crystal structures of the apo and the liganded form of OppA have elucidated its binding mechanism: OppA shows a three-domain architecture and upon ligand binding, the N- and C-terminal domains rearrange in a Venus fly-trap mechanism to adopt a closed conformation, in which the bound peptide is completely engulfed inside the protein (38). Furthermore, available structural information of a OppA:tripeptide complex indicates that the OppA binding cavity provides large hydrated pockets for various side chains of the tripeptide. Direct interactions between the peptidic substrate and OppA target the backbone and its termini. The N-terminus of the bound tripeptide makes multiple interactions with OppA residues and is kept in place by an extensive network of hydrogen bonds and salt bridges. Specifically, the protonated a-amine of the bound tripeptide forms a salt bridge with OppA residue D445, which is kept in its deprotonated form by hydrogen-bonding to Y135 and H187. The C-terminus of the tripeptide forms a salt bridge with R439. To experimentally verify if these interactions are also necessary for recognition of G-AisoK, we constitutively coexpressed OppA-D445A or OppA-R439A in DoppA-K12 cells. Expression of the OppA-R439A variant did not impact G-AisoK dependent amber suppression of sfGFP, implying that R439 may not be involved in interacting with the C-terminus of isopeptide-linked G-AisoK. In stark contrast, overexpression of the OppA-D445A variant did not lead to any sfGFP expression in the presence of G-AisoK, indicating- that the N-terminal amine of G in G-AisoK is interacting with OppA in a similar way as the one of a linear tripeptide (Fig. 2e).

[0418] To identify a potential peptidase responsible for G-AisoK to AisoK processing within the E. coli cytosol, we performed amber suppression experiments in single-gene knockouts (36) that had individual peptidases or proteases deleted in the presence of G-AisoK. We reasoned that lack of N-terminal glycine processing should lead to diminished protein expression yields. We could however not identify one specific peptidase solely responsible for G-AisoK to AisoK processing and speculate that G-AisoK may be a substrate for several promiscuous aminopeptidases (39).

[0419] We therefore generated multi-peptidase knockouts using a CRISPR-Casl2a-based genome editing platform for E. col\ (64). Notably, only cells with both pepN and pepA deleted ( pepN / pepA), showed a drastic reduction in sfGFP expression with G-AisoK, while amber suppression yields with BocK remained unaffected (Fig. 2f). Complementation with either pepA or pepN restored sfGFP expression with G-AisoK, indicating that either peptidase is sufficient for G-AisoK processing.

[0420] Together, these findings support a model in which G-AisoK is actively imported via the Opp transporter into the E. coli cytosol, where it is processed by endogenous peptidases, releasing AisoK for efficient amber suppression.

[0421] Example 3: A versatile toolbox for efficient encoding of diverse chemical functionalities into POIs

[0422] Next, we tested whether amino acids in a general G-XisoK (Fig. 3a) scaffold behaved similarly. Indeed, SisoK, bearing serine instead of alanine was similarly incorporated in a tripeptide (G-SisoK)-dependent manner. OppA is known to promiscuously bind 2-5 amino acid long peptides, favoring positively charged side chains. To elucidate how OppA distinguishes G-XisoK from XisoK, we solved the crystal structure of OppA bound to G-SisoK (PDB ID: 9RD1). The structure overlays well with previous ligand-bound OppA conformations (65) and adopts the closed state, with G-SisoK enclosed in the binding pocket (Fig. 3b). G-SisoK engages in extensive interactions with OppA through its backbone and termini. The N-terminal glycine forms key hydrogen bonds and electrostatic contacts: its protonated a-amine interacts with D445, while the C-terminal carboxylate is stabilized by hydrogen bonds involving the side chains of R439, H397, and N392 (Fig. 3b). To validate these interactions, we expressed OppA variants in AoppA. Expression of wt-OppA fully restored sfGFP expression with G-SisoK, while the D445A variant, disrupting the interaction with the N-terminal a-amine of G-SisoK, failed to rescue expression. Mutations targeting the hydrogen bonding network at the C-terminus of G-SisoK (e.g. R439A), had less pronounced effects, suggesting that the OppA binding site possesses some structural flexibility. These results highlight the essential role of the interaction between the a-amine of glycine and D445 for effective OppA binding and transport.

[0423] The OppA:G-SisoK crystal structure revealed no specific interactions with the serine side chain, which is accommodated in a spacious pocket (Fig. 3c). This suggests that OppA binding and uptake rely primarily on recognition of the tripeptide backbone and termini rather than side chain identity. Accordingly, this mechanism may represent a more general concept applicable to a variety of ncAAs presented within a general G-XisoK scaffold. Based on this, we expanded our propeptide strategy to efficiently incorporate XisoK derivatives bearing functionalities commonly used in GCE, including moieties for site-specific protein conjugation and crosslinking (Fig. 3d). Supplementing E. coli K12 with G-XisoK derivatives - where X represents various side chains - enabled efficient suppression of sfGFP-N150TAG and Ub-K63TAG using either wt-MbPyIRS / PyIT or suitable synthetase variants identified from an in cellulo screen (Fig. 3e). MS-analysis confirmed site-specific incorporation of the respective XisoK dipeptides, while supplementation with free XisoK derivatives led to minimal protein expression (Fig. 3e).

[0424] Efficient encoding of such XisoK derivatives is interesting, as lysine aminoacylation, in which the e-amino group of lysine is covalently linked to the a-carboxyl of any of the 20 naturally occurring amino acids, was recently identified as a reversible PTM (41, 42). As previous attempts at directly encoding S / T / P / CisoK derivatives via GCE proved very inefficient(19, 41, 43-45), functional studies on these PTMs are elusive.

[0425] Furthermore, CisoK-modified proteins are ideally suited for native chemical ligation approaches, as has been shown for generating Ub-conjugates (46, 47). Comparing CisoK- incorporation efficiencies obtained via our propeptide strategy with published CisoK- incorporation yields using specifically evolved PylRS-variants, impressively shows the advantage of actively pumping G-CisoK into bacterial cells (19). This indicates that intracellular ncAA concentration may be more crucial for efficient ncAA incorporation than finely tuning and evolving PylRS-variants.

[0426] Site-specific incorporation of XisoK derivatives, in which X is a natural amino acid, offers also the unique possibility to introduce a specific natural amino acid with an N-terminal a-amine within a POI sequence, equipping a protein with a second, artificial N-terminus. This allows internal protein labeling by applying various bioconjugation strategies that have been developed for targeting specific N-termini. (48)

[0427] By efficient encoding of PisoK (via uptake of G-PisoK, Fig. 3e), we introduce proline bearing a free a-amine at an internal site of a target protein and show that this allows tyrosinase- mediated labeling with phenol derivatives. (49) Tyrosinase oxidizes p-cresol to the corresponding o-quinone intermediate that can oxidatively couple to the free a-amine of proline in PisoK. We show specific and quantitative labeling of PisoK-modified Ub, establishing a solid chemoenzymatic labeling approach that is applicable to any user-chosen site in various POIs (Fig. 3f). When screening for PylRS-va Hants for incorporation of XisoK derivatives bearing aromatic side chains at position X, such as in HisoK (Fig. 3d), we could not identify any positive hits upon supplementing cells with G-HisoK. To verify if this stems from poor G-HisoK recognition by OppA or lack of appropriate PylRS-va riants for HisoK, we created a custom-designed MbPylRS library with five randomized positions in the active site and subjected this library to alternating rounds of positive and negative selection in K12 cells combined with fluorescence readout in the presence of G-HisoK. Gratifyingly, a novel Mb PylRS-va riant supported HisoK incorporation in the presence of G-HisoK, but not when supplementing cells with HisoK (Fig. 3e), indicating that OppA also binds and delivers G-XisoK amino acids with bulky and aromatic X-side chains. Efficient incorporation of synthetically easily accessible histidine-containing ncAAs may prove useful for expanding the range of genetically encoded metal coordination environments and may advance the generation and engineering of metalloenzymes with optimized properties and novel activities. (50, 51)

[0428] Importantly, also non-canonical side chains can be incorporated at position X to further expand the G-XisoK-toolbox for efficient incorporation of diverse functionalities. We explored the uptake of G-XisoK derivatives, where X contained functionalities for bioorthogonal labeling (52), as well as for light-induced and proximity-induced chemical crosslinking (53, 54). As a considerable advantage of our propeptide strategy, G-XisoK derivatives can be easily synthesized at large scales via SPPS with many non-canonical side chains commercially available as amino acid building blocks. In standard GCE approaches, apart from limiting incorporation yields, the synthetic accessibility of ncAAs often presents a considerable challenge. We first explored incorporation of functionalities that would be amenable for bioorthogonal protein labeling via Cu(l)-catalyzed azide alkyne cycloaddition (CuAAC). A G-XisoK derivative, bearing propargyl-glycine at position X, allowed very efficient installation of PrgisoK into diverse target proteins (Fig. 3e), including an eGFP-specific nanobody (eGFPNb- R75TAG), which was successfully labeled with a fluorophore-conjugated azide moiety via CuAAC (Fig. 3g). As seen for all other tested derivatives, PrgisoK incorporation was successful only when cells were supplemented with G-PrgisoK, but not when supplemented with PrgisoK (Fig. 3e). Similarly, PrgisoK incorporation yields using the corresponding G-bearing tripeptide, compare very favorably to recently reported PrgisoK incorporation efficiencies using a specific PylRS-va riant identified for PrgisoK incorporation (41).

[0429] To expand the G-XisoK-toolbox to ncAAs useful for mapping and trapping PPIs, we explored efficient incorporation of different crosslinker moieties using our propeptide strategy. We integrated commercially available photoleucine (pL) as residue X in the G-XisoK scaffold. Supplementing K12 cells with this tripeptide, enabled efficient encoding of a diazirine moiety into a POI using an Methanomethylophilus alvus (Ma) PylRS-derived variant (Fig. 3e) and UV- induced crosslinking of PPIs, as exemplified by crosslinking of glutathione-S-transferase (GST) dimer, as well as sfGFP dimer (Fig. 3h). As pL is commercially available, synthesis of the needed tri eptide via SPPS is much easier, more efficient, and scalable than synthetic access to other diazirine-bearing ncAAs reported for GCE and photo-crosslinking (53-55), highlighting an important advantage of our strategy.

[0430] For trapping transient PPIs in a proximity-induced manner, we designed a G-XisoK scaffold, bearing the finely tuned electrophile chloroalanine (CIA) at position X. Incubation of K12 cells with G-CIAisoK led to efficient incorporation of CIAisoK into different POIs (Fig. 3e). When CIAisoK is placed in proximity to a nucleophilic amino acid in an interacting protein, an SN2 nucleophilic reaction can take place, covalently stabilizing the protein complex by a stable thioether bridge (Fig. 3i). By pairwise incorporation of CIAisoK and cysteine residues at protein-protein interfaces, we covalently stabilized a variety of low-affinity protein complexes with KDS in the micromolar to low millimolar range in living cells, such as a sfGFP homodimer, the interaction between affibody and protein Z (56) as well as the ternary GDP-bound complex between small G-protein Rablb and the guanine nucleotide exchange factor (GEF) domain of DrrA (57) (Fig. 3i). Distances of 8-12 A between the corresponding Ca atoms of cysteine and the lysine of CIAisoK could be efficiently crosslinked. Apart from cysteine, also histidine and glutamate showed successful covalent homodimer formation with CIAisoK-modified sfGFP. In vitro crosslinking using purified affibody- and protein-Z-variants confirmed that crosslinking is specific for CIAisoK, while no crosslinking was observed for AisoK-bearing affibody, as expected.

[0431] Example 4: Scalable and cheap XisoK incorporation through OppA evolution

[0432] For all tested G-XisoK tripeptides, we observed highly efficient protein production in chemically-defined auto induction (Al) media (58). Amber suppression in nutrient-rich media, such as lysogeny broth (LB) or 2-YT, did, however, not give the desired yields of protein expression (Fig. 4a). We hypothesized that G-XisoK tripeptides may compete for OppA-binding with short linear peptides that are abundant in nutrient-rich expression media based on peptone or tryptone. This was confirmed by uptake assays determining intracellular SisoK concentrations, which upon G-SisoK supplementation, were sixfold lower in 2-YT compared to Al media, indicating that OppA-mediated active transport is impaired in such conditions. This drastically impacts the implementation of our versatile G-XisoK toolbox for cheap and scalable protein production, as chemically-defined Al media is much more expensive and cumbersome to prepare than regular nutrient-rich growth media and does not support as high cell-density cultures leading to decreased biomass and lower overall expression yields of ncAA-modified proteins. We therefore aimed at evolving an engineered OppA-variant that preferentially binds G-SisoK over short linear peptides as present in peptone or tryptone-based media. We developed a fluorescence-activated cell sorting (FACS)-based screening platform to alter the binding preference and selectivity of OppA using directed evolution. Our screening system couples uptake of the G-SisoK tripeptide to sfGFP fluorescence by suppressing the amber codon in sfGFP-N150TAG (Fig. 4b). We co-transformed an error-prone OppA-library together with the wt-MbPylRS / PylT pair and sfGFP-N150TAG into DoppA-K12 cells in Al media and subjected these cells to subsequent rounds of enrichment from low (lg / L) to high (16g / L) tryptone concentrations. We identified four converging OppA-variants each harboring 4-5 mutations distributed all over the OppA-fold. Mutations that were present in more than one variant or were structurally close to each other, were identified as hotspot positions and selected for saturation mutagenesis on all four OppA-variants and wt-OppA. The obtained library was again subjected to multiple FACS-based enrichment steps in LB and 2-YT media. The final selected OppA-variant (OppA-ev2) contains a total of seven mutations, distributed over the entire OppA-fold, with only one mutation (R439Q) being closer than 4 A to the binding site. We introduced these mutations into the K12 genome via lambda red-mediated homologous recombination to create a K12-derived engineered E. coli strain for scalable and cheap production of XisoK-bearing proteins in peptide containing media, which we dubbed lsoK12. Doubling times for lsoK12 and its parent strain were comparable both in Al and 2-YT media. With the lsoK12 strain in hand, we first measured intracellular SisoK concentrations in presence of G-SisoK in 2-YT media. Gratifyingly, intracellular SisoK concentrations substantially increased (7-10-fold) when compared to K12 cells (Fig. 4c), while BocK concentrations did not differ between the two E. coli strains, indicating that OppA-ev2 indeed preferentially binds and transports G-SisoK over linear peptides present in 2-YT (Fig. 4c). Importantly, the observed elevated SisoK concentrations in lsoK12 cells led to efficient amber suppression of TAG- containing constructs in 2-YT media, exhibiting similar efficiencies as observed for K12 cells grown in defined Al media (Fig. 4a).

[0433] To confirm OppA-ev2 binding preference, we measured KDS of wt-OppA and OppA-ev2 towards G-SisoK, a linear GSK tripeptide (mimicking linear tryptone peptides) and SisoK via microscale thermophoresis (Fig. 4d). While both wt-OppA and OppA-ev2 showed similarly low binding affinities towards SisoK (~300 pM), affinity of G-SisoK towards OppA-ev2 is slightly increased compared to wt-OppA (37 pM versus 50 pM). Interestingly, binding of linear GSK was four-fold decreased for OppA-ev2 versus wt-OppA (275 pM versus 71 pM), corroborating the in cellulo data that G-SisoK is able to compete with linear tryptone-derived peptides in the presence of OppA-ev2. In an attempt to structurally rationalize this binding behavior, we mapped the seven mutations found in OppA-ev2 onto a published OppA structure liganded to a linear tripeptide (PDB ID: 3TCF)(38). Most of the mutations are not directly contacting the bound tripeptide, but interestingly residue R439 that is important for keeping the C-terminus of the linear tripeptide in place is mutated to glutamine in 0ppA-ev2. The R439Q mutation may therefore be responsible for the observed lowered affinity for linear GSK, while the binding affinity of G-SisoK, which does not display a carboxy group at this position, is not affected by this mutation.

[0434] As the OppA-ev2-variant was evolved in the presence of G-SisoK, we wondered if preferential binding and uptake was extendable to other G-XisoK derivatives. Indeed, all of the tested G-XisoK derivatives (X=A, S, T, C, V, L, P, H, Prg, pL and CIA) led to very efficient amber suppression of sfGFP-N150TAG, with hardly any amounts of truncated side-product when expressed in lsoK12 cells grown in 2-YT media, resembling protein yields obtained in K12 cells grown in chemically-defined and peptide-free Al media (Fig. 4e). In fact, tripeptide uptake was so efficient in lsoK12 cells that G-SisoK concentrations as low as 50-100 pM led to similar incorporation yields as observed in K12 cells, supplemented with 1 mM G-SisoK, thereby decreasing necessary ncAA concentrations by 10-fold (Fig. 4f). We show that our propeptide strategy combined with engineered lsoK12 enables high-yielding XisoK-incorporation at various positions in diverse target proteins spanning sizes from 7 to 85 kDa. Amber suppression of PCNA, R-lactamase, SUMO2, Calmodulin, eGFPNb, and Hsp82 attest to the wide applicability of the G-XisoK / lsoK12 combination resulting in highly efficient production of XisoK-modified proteins matching or even exceeding wt expression yields. Interestingly, in many cases G-XisoK uptake was also improved when lsoK12 cells were grown in Al media. Preparative large-scale production and purification of a GFP nanobody with an ncAA bearing a propargyl moiety showed that the G-PrgisoK / lsoK12 combination resulted in similar purified protein yields (44 mg / L) as obtained for wt expression (41 mg / L) and exceeded yields from a previously reported optimized alkyne-bearing ncAA / PyIRS combination. Increasing intracellular ncAA concentrations via efficient tripeptide uptake facilitated also multi-site amber suppression within one target protein. We introduced TAG-codons at up to three positions into Histone H3 (K27, K79, K122) and expressed the corresponding variants in the presence of G-SisoK. lsoK12 outperformed K12 in single, double and triple amber suppression with expression yields rivaling wt-H3 production, while only minute amounts of doubly- and triply-suppressed protein were obtained with the gold-standard ncAA BocK (Fig. 4g).

[0435] Evolution of OppA gave rise to the following variants:

[0436] Example 5: Generalized ncAA uptake using Z-AisoKs

[0437] Efficient uptake and processing of G-XisoK results in high intracellular concentrations of both the XisoK dipeptide and cleaved N-terminal glycine. Crystallographic analysis of the OppA:G-SisoK complex revealed a spacious cavity extending from the serine side chain to the N-terminal glycine. Given OppA's promiscuity, we hypothesized that other side chains, including those of ncAAs, could also be accommodated at this position, enabling efficient transport of diverse non-canonical tripeptides into the E. coli cytosol. If different N-terminal residues (Z) in a Z-AisoK scaffold (Fig. 5a) are also efficiently cleaved, the strategy could broadly enable intracellular delivery of various amino acids.

[0438] We synthesized a panel of 14 Z-AisoK tripeptides bearing either natural amino acids or ncAAs with diverse side chains, including bulky or negatively charged groups with poor or negligible cell permeability as N-terminal Z-residues (Fig. 5a). To monitor uptake and cleavage of Z, we assessed AisoK incorporation into sfGFP-N150TAG using the wt-MbPylRS / tRNA pair. Excitingly, for eight out of 14 Z-residues, supplementation of K12 with corresponding Z-AisoK tripeptides, yielded amber suppression efficiencies comparable to G-AisoK (Fig. 5b), with LC-MS confirming AisoK incorporation in all cases. No sfGFP expression was observed in DoppA cells, confirming dependence on OppA-mediated uptake.

[0439] In contrast, tripeptides bearing bulkier or negatively charged Z-side chains (compounds 5-6 and 12-15, Fig. 5a), resulted in reduced (5-6) or completely abolished (12-15) amber suppression (Fig. 5b), suggesting inefficient uptake and / or cleavage. To address this, we leveraged our OppA engineering platform for evolving new OppA variants capable of accommodating tripeptides with larger or negatively charged Z-residues (Fig. 5c). Guided by our OppA:G-SisoK structure, we selected four residues around glycine of G-SisoK for sitesaturation mutagenesis (Fig. 5d) and subjected a corresponding library to multiple FACS-based enrichments in the presence of tripeptides 13 and 15. From these screens, we isolated two OppA variants with enhanced uptake: OppA-Zl (evolved with 13) and OppA-Z2 (evolved with 15). Both variants feature small side chains at the targeted positions likely enlarging the binding pocket to accommodate larger substrates, and in the case of OppA-Z2, a R439H mutation. These findings provide further evidence that the R439 variant can enhance binding of isopeptide-linked scaffolds.

[0440] We genomically introduced these OppA-variants, generating K12-Z1 and K12-Z2 strains, respectively. Supplementing K12-Z1 with bulky Z-tripeptides (5, 6, 12, 13) enabled efficient AisoK incorporation, but uptake of tripeptides with negatively charged Z-residue (14, 15) was not supported by this engineered strain. In contrast, K12-Z2 enabled efficient uptake of all 14 Z-AisoK tripeptides, including those bearing negatively charged ncAAs such as SucK and GluK as Z-residues (Fig. 5b).

[0441] To demonstrate that tripeptide-based uptake is superior to direct ncAA supplementation, we focused on selected Z-residues known for low suppression efficiencies, where limited uptake was suspected to be the main bottleneck. We compared amber suppression yields upon supplementing cells either Z directly or the Z-AisoK tripeptide. Compound 3 (with AcK as Z-residue) yielded higher AcK incorporation than AcK alone using the Mb Py IRS-variant AcKRS3 (60), highlighting the benefit of active delivery. Notably, compound 13 (bearing LipK as Z-residue) or LipK alone led to negligible full-length sfGFP-expression in presence of the LipKRS / tRNA pair (66). In contrast, K12-Z1 supplemented with 13 enabled efficient sfGFP production with LipK incorporation confirmed by LC-MS (Fig. 5e). Similar improvements were observed for tripeptides 5 and 6 in K12-Z2, confirming that evolved OppA variants enable effective delivery of tri peptides impermeable to wt-strains.

[0442] Further unnatural amino acids in position Z were explored and successful import into the cell and incorporation into sfGFP could be demonstrated (see Fig.6).

[0443] Example 6: Efficient encoding of two distinct ncAAs

[0444] Given OppA's adaptability in substrate recognition, we envisioned a broadly applicable strategy to co-deliver two ncAAs using a single, easily synthesized Z-XisoK tripeptide. To leverage such a mechanism for site-specific dual ncAA incorporation into a single protein, we designed tri peptide 16, bearing AcK (a PTM) as Z and pLisoK (a photocrosslinker) as XisoK (Fig. 5f). Such a setup would represent an ideal tool to investigate PTM-specific protein interactors or to chemically stabilize transient POI-reader complexes (59).

[0445] Importantly, the PylRS / PyIT pairs for AcK and pLisoK are mutually orthogonal (61). Upon OppA-mediated uptake and cleavage of 16, both AcK and pLisoK accumulate intracellularly, allowing dual suppression of TAA and TAG codons within the same target protein (sfGFP- N40AcK-N150pLisoK, Fig. 5f), using the respective PylRS / PyIT pairs. Notably, protein yields were significantly higher when cells were supplemented with tripeptide 16 compared to addition of free AcK and G-pLisoK, underscoring the enhanced efficiency of transporter- mediated delivery. LC-MS of full-length sfGFP confirmed dual incorporation of AcK and pLisoK (Fig. 5f). Mutual orthogonality and accurate decoding of both ncAAs was furthermore shown using a Ub-SUMO2 fusion construct containing a TEV protease site (Ub-K48TAG-TEV-SUMO-K11TAA). LC-MS analysis after TEV-cleavage, revealed specific pLisoK incorporation at TAG48 of Ub and AcK encoding in response to TAA11 of SUMO2.

[0446] Example 7: Exploring different linkages between Z and X (position C) and different terminal groups (position A)

[0447] A number of constructs (24-41) comprising either an amide linkage (24-32) or an ester linkage (33-41) between residues Z and X (corresponding to position C in the claims) and comprising different terminal groups (corresponding to position A in the claims) were tested for their potential to be taken up into cells and processed intracellularly. In case of successful uptake and processing incorporation of AisoK or LacK into a protein, in this case sfGFP, determined by fluorescence, is observed. The results are shown in Fig.7.

[0448] Example 8: Exploring different lysine derivatives or analogues in position E

[0449] Successful incorporation of isopeptide-linked XisoK ncAAs into proteins, wherein the XisoK was incorporated into the protein via the a-amino and a-carboxyl groups of the lysine moiety, was demonstrated herein above. It was further explored whether the lysine in position E can be replaced with lysine derivatives or analogues without affecting the import of the tripeptides into the cell, the intracellular processing of the tripeptides and the incorporation of the Z aminoacids (being Lys(Ac)) into proteins. For that tripeptide of the type Lys(Ac)-Ala- stem where the stem is lysine (42), ornithine (43), 6-aminohexanoic acid (44), 1,5 diaminopentane (45), lysineamide (46) or 2,3-diaminopropionic acid (48) were tested. The results are shown in Figure 8. Moreover, incorpration of isopeptide-linked ncAAs comprising a hydroxy-acid derivative of lysine (K(OH)) is shown in Figure 9.

[0450] Example 9: Improving uptake of peptides comprising an azide group in position A

[0451] It was demonstrated that peptides comprising an azide group in position A instead of an amino group, i.e., the peptide AzGGisoK, can be taken up and incorporated into proteins with the MF18 Methanosarcina mazei PylRS harboring the mutations:L309A, N346Q, C348S (Fig.10), although relatively high concentrations were needed. By screening a OppA library having mutations at positions Y135, H187, W442, and D445, an OppA variant (Y135L, H187L, D445L; OppA_Z7) was identified that resulted in a more efficient uptake of such peptides. Conclusion and discussion

[0452] Low protein yields and the limited chemical accessibility of many ncAAs remain major obstacles to the routine application of GCE for the generation of proteins with therapeutic or biotechnological potential. Here, we overcome these drawbacks by combining a modular propeptide strategy with the programmed hijacking of the Opp transporter, enabling active and tailored import of a broad range of ncAAs, including those that have historically been refractory to efficient uptake and encoding. Isopeptide-linked Z-XisoK tripeptides act as trojan horses that can be readily synthesized from commercially available building blocks via SPPS, and function as privileged ligands for the periplasmic binding protein OppA, enabling their ATP-driven transport into the cytosol. Once inside the cell, Z-XisoK tripeptides are enzymatically processed, leading to intracellular accumulation of both isopeptide-linked XisoK ncAAs, as well as Z-residues. Through a FACS-based directed evolution approach, we reprogrammed OppA to selectively discriminate against linear peptides present in complex media. The resulting engineered E. coli strain, lsoK12, enables cost-effective, high-yield production of modified proteins in nutrient-rich conditions using minimal tripeptide concentrations. This platform allows robust incorporation of eleven previously inaccessible XisoK ncAAs, significantly expanding the chemical space available through GCE. Among these functionalities are ncAAs bearing bioorthogonal handles, novel PTMs, chemical and photocrosslinkers, as well as functional groups for chemoenzymatic ligations, all of which are demonstrated in proof-of-principle applications. The positively charged side chains of XisoK ncAAs make them ideal moieties for applications such as protein labeling, where site-directed incorporation at surface-exposed positions is essential. This strategy avoids aggregation and misfolding often associated with bulky, hydrophobic ncAAs commonly used for labeling purposes (62). With these advantages, we anticipate that the XisoK-toolbox can be further expanded through aaRS engineering to include functionalities for inverse-electron-demand Diels-Alder cycloadditions or spectroscopic probes.

[0453] Importantly, uptake and processing of Z-XisoK tripeptides can also be leveraged for customized delivery and intracellular accumulation of challenging Z-residues that typically lack cell-permeability. By synthesizing a panel of 14 different Z-AisoK tri peptides, we demonstrate that AisoK incorporation into sfGFP serves as a straightforward readout for efficient tripeptide uptake and processing, thereby eliminating the need for intensive MS-based uptake assays.21Our platform provides a foundation for developing OppA variants with specific binding sites for tripeptides carrying otherwise cell-impermeable Z-residues, such as bulky and negatively charged groups. Thus, our innovation enables the intracellular delivery of these challenging ncAAs. Site-specific incorporation of Z-residues from Z-XisoK significantly outperforms direct Z supplementation. As OppA evolution for a specific Z-residue is decoupled from availability of a Z-specific orthogonal aaRS, our system presents a practical and modular strategy to dissect ncAA uptake from its co-translational incorporation, complementing recent efforts in decoupling and optimizing individual GCE steps towards synthesis of non-canonical biopolymers in E. coli (63, 13). While the current system relies on orthogonality of Z incorporation to AisoK-selective PylRS variants, future efforts will focus on relaxing this constraint by diversifying the XisoK scaffold to further expand the system's versatility and applicability. Notably, the ability to co-import two distinct ncAAs via a single Z-XisoK scaffold enables efficient dual ncAA incorporation. We envision extending our concept toward multi- ncAA encoding in synergy with advances in non-canonical polymer biosynthesis (63, 13), and E. coli strains with compressed genomes (14, 16, 17).

[0454] Methods

[0455] General methods: Plasmids and Reagents

[0456] Genes encoding GST-E51TAG, tyrosinase and Histone H3 mutants were ordered as DNA string from Twist biosciences and cloned into the pBAD vector via standard restriction cloning with enzymes from New England Biolabs. Single point mutants of OppA as well as insertion of a 3C site on UbK63 were done via Site-directed, Ligase-Independent Mutagenesis (SLIM) (J. Chiu, P. E. March, R. Lee, D. Tillett, Site-directed, Ligase-Independent Mutagenesis (SLIM): a singletube methodology approaching 100% efficiency in 4 h. Nucleic Acids Res 32, el74 (2004)). Recombination plasmid pSIJ8 was purchased from Addgene (Addgene ID: 68122). Oligonucleotide primers were ordered from Microsynth AG unless otherwise specified. General solvents and chemical reagents were purchased from Sigma Aldrich, cslabs, Fisher Scientific, Carbolution or Acros Organics. Fmoc building blocks of photoleucine, propagylglycine and chloroalanine were purchased from Iris Biotech. All reagents were used without further purification.

[0457] SDS-PAGE gels (Bolt™ Bis-Tris Plus Mini Protein Gels, 4-12%, Invitrogen) were run on a Bolt™ Mini Gel Tank (Invitrogen) system (165 V for 40 minutes) and stained with Quick Coomassie Stain (Generon). As a marker, PageRuler Prestained Plus Protein Ladder 10-250 kDa (ThermoFisher) was used. Proteins were transferred onto a nitrocellulose membrane using a Bio-Rad Trans-blot Turbo Transfer System. After transfer, the membrane was blocked with 5% skim milk in TBS-T buffer for 1 h at RT and incubated with the appropriate HRP-coupled antibody. Imaging of western blots was performed on an Amersham ImageQuant™ 800 using Immobilon® Forte Western HRP Substrate (Millipore).

[0458] Peptide and protein LC-MS was done using an Agilent Technologies 1260 Infinity LC-MS system with a 6310 Quadrupole spectrometer. Proteins were measured on a Phenomenex Jupiter C4 300 A LC Column (150 x 2 mm, 5 pm) and peptides, on a Luna Omega PS C18 (100 x 2.1 mm, 3 pm). The solvent system consisted of 0.1% formic acid in water (solvent A) and 0.1% formic acid in acetonitrile (solvent B). Both proteins and peptides were analyzed in positive mode.

[0459] Table 3: Plasmids used in this study Table 4: Primers used in this study

[0460] Table 5: Primers used for SLIM cloning Protein Sequences sfGFP N 150TAG H6 (SEQ ID NO: 30):

[0461] MPSKGEELFTGVVPILVELDGDVNGH KFSVRGEGEGDATNGKLTLKFICTTGKLPVPWPTLVTTLTYGVQ

[0462] CFSRYPDHMKRH DFFKSAMPEGYVQERTISFKDDGTYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGH K

[0463] LEYNFNSH *VYITADKQKNGIKANFKIRH NVEDGSVQLADHYQQNTPIGDGPVLLPDN HYLSTQSVLSKD PN EKRDH MVLLEFVTAAGITHGM DELYKGSH HHHH H

[0464] UbK63TAG H6 (SEQ ID NO:31):

[0465] MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLSDYN IQ* ESTLHLVLRL

[0466] RGGHHH HHH

[0467] MDKKPLDVLISATGLWMSRTGTLH KIKHHEVSRSKIYIEMACGDHLVVN NSRSCRTARAFRHHKYRKTCK

[0468] RCRVSDEDINN FLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKST

[0469] PNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLN MAKPFRELEPELVTRRKNDFQRLYTN DREDYLGKLER

[0470] DITKFFVDRGFLEIKSPILIPAEYVERMGIN NDTELSKQIFRVDKNLCLRPMLAPTLYNYLRKLDRILPGPIKIF

[0471] EVGPCYRKESDGKEH LEEFTMVNFCQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHG

[0472] DLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTNL

[0473] MaPylRS H227I Y228P (SEQ ID NO:6):

[0474] MTVKYTDAQIQRLREYGNGTYEQKVFEDLASRDAAFSKEMSVASTDNEKKIKGMIANPSRHGLTQLMN

[0475] DIADALVAEGFIEVRTPIFISKDALARMTITEDKPLFKQVFWIDEKRALRPMLAPNLYSVMRDLRDHTDGP

[0476] VKIFEMGSCFRKESHSGM HLEEFTMLNLVDMGPRGDATEVLKNYISVVMKAAGLPDYDLVQEESDVYKE

[0477] TIDVEINGQEVCSAAVGPIPLDAAH DVH EPWSGAGFGLERLLTIREKYSTVKKGGASISYLNGAKIN

[0478] DiazKRS (SEQ ID NO:5):

[0479] MDKKPLNTLISATGLWMSRTGTIHKIKHH EVSRSKIYIEMACGDH LVVN NSRSSRTARALRHHKYRKTCK

[0480] RCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAM PKSVARAPKPLENTEAAQAQPSGSKFSPAIPV

[0481] STQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLNSG

[0482] KPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIF

[0483] RVDKNFCLRPMLAPNLMNYARKLDRALPDPIKIFEIGPCYRKESDGKEHLEEFTMLNFAQMGSGCTRENL

[0484] ESIITDFLNH LGIDFKIVGDSCMVYGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKV

[0485] KH DFKNIKRAARSESYYNGISTNL eGFP-N B wt (SEQ ID NO:32):

[0486] MKYLLPTAAAGLLLLAAQPAMAQVQLVESGGALVQPGGSLRLSCAASGFPVNRYSMRWYRQAPGKERE

[0487] WVAGMSSAGDRSSYEDSVKGRFTISRDDARNTVYLQM NSLKPEDTAVYYCNVNVGFEYWGQGTQVTV SSKKKH HHHH H

[0488] WVAGMSSAGDRSSYEDSVKGRFTISRDDARNTVYLQMNSLKPEDTAVYYCNVNVGFEYWGQGTQVTV

[0489] SSKKKHHHHHH eGFP-NB R75TAG H6 (SEQ ID NO:34):

[0490] MKYLLPTAAAGLLLLAAQPAMAQVQLVESGGALVQPGGSLRLSCAASGFPVNRYSMRWYRQAPGKERE

[0491] WVAGMSSAGDRSSYEDSVKGRFTISRDDA*NTVYLQMNSLKPEDTAVYYCNVNVGFEYWGQGTQVTV SSKKKHHHHHH

[0492] GST E51TAG H6 (SEQ ID NO:35):

[0493] MSPILGYWKIKGLVQPTRLLLEYLEEKYEEHLYERDEGDKWRNKKFELGL*FPNLPYYIDGDVKLTQSMAII

[0494] RYIADKHNMLGGCPKERAEISMLEGAVLDIRYGVSRIAYSKDFETLKVDFLSKLPEMLKMFEDRLCHKTYL

[0495] NGDHVTHPDFMLYDALDVVLYMDPMCLDAFPKLVCFKKRIEAIPQIDKYLKSSKYIAWPLQGWQATFGG

[0496] GDHPPKHHHHHH

[0497] 3C-Ub-K63TAG H6 (SEQ ID NO:36):

[0498] MENLLEVLFQGPGGGGSMQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAGKQLEDGR

[0499] TLSDYNIQ*ESTLHLVLRLRGGHHHHHH

[0500] OppA H6 (SEQ ID NO:37):

[0501] MGTNITKRSLVAAGVLAALMAGNVALAADVPAGVTLAEKQTLVRNNGSEVQSLDPHKIEGVPESNISRD

[0502] LFEGLLVSDLDGHPAPGVAESWDNKDAKVWTFHLRKDAKWSDGTPVTAQDFVYSWQRSVDPNTASPY

[0503] ASYLQYGHIAGIDEILEGKKPITDLGVKAIDDHTLEVTLSEPVPYFYKLLVHPSTSPVPKAAIEKFGEKWTQP

[0504] GNIVTNGAYTLKDWVVNERIVLERSPTYWNNAKTVINQVTYLPIASEVTDVNRYRSGEIDMTNNSMPIEL

[0505] FQKLKKEIPDEVHVDPYLCTYYYEINNQKPPFNDVRVRTALKLGMDRDIIVNKVKAQGNMPAYGYTPPYT

[0506] DGAKLTQPEWFGWSQEKRNEEAKKLLAEAGYTADKPLTINLLYNTSDLHKKLAIAASSLWKKNIGVNVKL

[0507] VNQEWKTFLDTRHQGTFDVARAGWCADYNEPTSFLNTMLSNSSMNTAHYKSPAFDSIMAETLKVTDE

[0508] AQRTALYTKAEQQLDKDSAIVPVYYYVNARLVKPWVGGYTGKDPLDNTYTRNMYIVKHHHHHH

[0509] OppA ev2 H6 (SEQ ID NO:38):

[0510] MGTNITKRSLVAAGVLAALMAGNVALAADVPAGVTLAEKQTLVRNNGSEVQSLDPHKIEGVPESNISRD

[0511] LFEGLLVSDLDGHPAPGVAESWDNKDAKVWTFHLRKDAKWSDGTPVTAQDFVYSWQRSVDPNTASPY

[0512] ASYLQYGHIAGIDEILEGKKPITDLGVKAIDDHTLEVTLSEPVPYFYKLLVHPSTSPAPKAAIEKFGEKWTQP

[0513] GNIVTNGAYTLKLAVVNERIVLERSPTYWNNAKTVINQVTYLPIASEVTDVNRYRSGEIDMTNNSMPIELF

[0514] QKLKKEIPDEVHVDPYLCTYYYEVNNQKPPFNDVRVRTALKLGMDRDIIVNKVKAQGDMPAYGYTPPYT

[0515] DGAKLTQPEWFGWSQEKRNEEAKKLLAEAGYTADKPLTINLLYNTSDLHKKLAIAASSLWKKNIGVNVKL VNQEWKTFLDTRHQGTFDVAQAGWCADYNEPTSFLNTMLSNHSMNTAHYKSPAFDSIMAETLKVTDE

[0516] AQRTALYTKAEQQLDKDSAIVPVYYYVNARLVKPWVGGYTGKDPLDNTYTRNMYIVKHHHHHH

[0517] OppA Z1 (SEQ ID NO:69)

[0518] MTNITKRSLVAAGVLAALMAGNVALAADVPAGVTLAEKQTLVRNNGSEVQSLDPHKIEGEPEGNISRDL

[0519] FEGLLVSDLDGHPAPGVAESWDNKDAKVWTFHLRKDAKWSDGTPVTAQDFVYSWQRSVDPNTASPY

[0520] ASYLQYGHIAGIDEILEGKKPITDLGVKAIDDHTLEVTLSEPVPYFYKLLVHPSTSPVPKAAIEKFGEKWTQP

[0521] GNIVTNGAYTLKDWVVNERIVLERSPTYWNNAKTVINQVTYLPIASEVTDVNRYRSGEIDMTNNSMPIEL

[0522] FQKLKKEIPDEVHVDPYLCTYYYEINNQKPPFNDVRVRTALKLGMDRDIIVNKVKAQGNMPAYGYTPPYT

[0523] DGAKLTQPEWFGWSQEKRNEEAKKLLAEAGYTADKPLTINLLYNTSDLHKKLAIAASSLWKKNIGVNVKL

[0524] VNQEWKTFLDTRHQGTFDVARAGWCADYNEPTSFLNTMLSNSSMNTAHYKSPAFDSIMAETLKVTDE

[0525] AQRTALYTKAEQQLDKDSAIVPVYYYVNARLVKPWVGGYTGKDPGDSTYTRNMYIVKH

[0526] OppA Z2 (SEQ ID NO:70)

[0527] MTNITKRSLVAAGVLAALMAGNVALAADVPAGVTLAEKQTLVRNNGSEVQSLDPHKIEGVPEANISRDL

[0528] FEGLLVSDLDGHPAPGVAESWDNKDAKVWTFHLRKDAKWSDGTPVTAQDFVYSWQRSVDPNTASPY

[0529] ASYLQYGHIAGIDEILEGKKPITDLGVKAIDDHTLEVTLSEPVPYFYKLLVHPSTSPVPKAAIEKFGEKWTQP

[0530] GNIVTNGAYTLKDWVVNERIVLERSPTYWNNAKTVINQVTYLPIASEVTDVNRYRSGEIDMTNNSMPIEL

[0531] FQKLKKEIPDEVHVDPYLCTYYYEINNQKPPFNDVRVRTALKLGMDRDIIVNKVKAQGNMPAYGYTPPYT

[0532] DGAKLTQPEWFGWSQEKRNEEAKKLLAEAGYTADKPLTINLLYNTSDLHKKLAIAASSLWKKNIGVNVKL

[0533] VNQEWKTFLDTRHQGTFDVAHAGWCADYNEPTSFLNTMLSNSSMNTAHYKSPAFDSIMAETLKVTDE

[0534] AQRTALYTKAEQQLDKDSAIVPVYYYVNARLVKPWVGGYTGKDPGDATYTRNMYIVKH

[0535] 3-lactamase wt H6 (SEQ ID NO:39):

[0536] MGADLADRFAELERRYDARLGVYVPATGTTAAIEYRADERFAFCSTFKAPLVAAVLHQNPLTHLDKLITYT SDDIRSISPVAQQHVQTGMTIGQLCDAAIRYSDGTAANLLLADLGGPGGGTAAFTGYLRSLGDTVSRLDA EEPELNRDPPGDERDTTTPHAIALVLQQLVLGNALPPDKRALLTDWMARNTTGAKRIRAGFPADWKVID KTGTGDYGRANDIAVVWSPTGVPYVVAVMSDRAGGGYDAEPREALLAEAATCVAGVLAGSGGSGHHH

[0537] HHH

[0538] 3-lactamase K221TAG H6 (SEQ ID NQ:40):

[0539] MGADLADRFAELERRYDARLGVYVPATGTTAAIEYRADERFAFCSTFKAPLVAAVLHQNPLTHLDKLITYT

[0540] SDDIRSISPVAQQHVQTGMTIGQLCDAAIRYSDGTAANLLLADLGGPGGGTAAFTGYLRSLGDTVSRLDA

[0541] EEPELNRDPPGDERDTTTPHAIALVLQQLVLGNALPPDKRALLTDWMARNTTGA*RIRAGFPADWKVID

[0542] KTGTGDYGRANDIAVVWSPTGVPYVVAVMSDRAGGGYDAEPREALLAEAATCVAGVLAGSGGSGHHH

[0543] HHH

[0544] SUM02 wt H6 (SEQ ID N0:41): MADEKPKEGVKTENNDHINLKVAGQDGSVVQFKIKRHTPLSKLMKAYCERQGLSMRQIRFRFDGQPINE TDTP AQLE M E D E DTI D VFQQQTG G H H H H H H

[0545] SUMO2 K45TAG H6 (SEQ ID NO:42):

[0546] MADEKPKEGVKTENNDHINLKVAGQDGSVVQFKIKRHTPLSKLM*AYCERQGLSMRQIRFRFDGQPINE

[0547] TDTP AQLE M E D E DTI D VFQQQTG G H H H H H H

[0548] Calmodulin wt strep (SEQ ID NO:43):

[0549] MADQLTEEQIAEFKEAFSLFDKDGDGTITTKELGTVMRSLGQNPTEAELQDMINEVDADGNGTIDFPEFL

[0550] TMMARKMKDTDSEEEIREAFRVFDKDGNGYISAAELRHVMTNLGEKLTDEEVDEMIREADIDGDGQVN YEEFVQMMTAKSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK

[0551] Calmodulin G40TAG strep (SEQ ID NO:44):

[0552] MADQLTEEQIAEFKEAFSLFDKDGDGTITTKELGTVMRSL*QNPTEAELQDMINEVDADGNGTIDFPEFL

[0553] TMMARKMKDTDSEEEIREAFRVFDKDGNGYISAAELRHVMTNLGEKLTDEEVDEMIREADIDGDGQVN YEEFVQMMTAKSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK

[0554] Histone H3 wt H6 (SEQ ID NO:45):

[0555] MARTKQTARKSTGGKAPRKQLATKAARKSAPATGGVKKPHRYRPGTVALREIRRYQKSTELLIRKLPFQRL

[0556] VREIAQDFKTDLRFQSSAVMALQEASEAYLVALFEDTNLCAIHAKRVTIMPKDIQLARRIRGERARSHHHH HH

[0557] Histone H3 K112TAG (SEQ ID NO:46):

[0558] MARTKQTARKSTGGKAPRKQLATKAARKSAPATGGVKKPHRYRPGTVALREIRRYQKSTELLIRKLPFQRL

[0559] VREIAQDFKTDLRFQSSAVMALQEASEAYLVALFEDTNLCAIHAKRVTIMP*DIQLARRIRGERARSHHHH HH

[0560] Histone H3 K79 K112TAG H6 (SEQ ID NO:47):

[0561] MARTKQTARKSTGGKAPRKQLATKAARKSAPATGGVKKPHRYRPGTVALREIRRYQKSTELLIRKLPFQRL

[0562] VREIAQDF*TDLRFQSSAVMALQEASEAYLVALFEDTNLCAIHAKRVTIMP*DIQLARRIRGERARSHHHH HH

[0563] Histone H3 K27 K79 K112TAG H6 (SEQ ID NO:48):

[0564] MARTKQTARKSTGGKAPRKQLATKAAR*SAPATGGVKKPHRYRPGTVALREIRRYQKSTELLIRKLPFQRL

[0565] VREIAQDF*TDLRFQSSAVMALQEASEAYLVALFEDTNLCAIHAKRVTIMP*DIQLARRIRGERARSHHHH HH inase H6 ID NO: MSNKYRVRKNVLH LTDTEKRDFVRTVLILKEKGIYDRYIAWHGAAGKFHTPPGSDRNAAHMSSAFLPWH

[0566] REYLLRFERDLQSIN PEVTLPYWEWETDAQMQDPSQSQIWSADFMGGNGNPIKDFIVDTGPFAAGRW

[0567] TTIDEQGNPSGGLKRNFGATKEAPTLPTRDDVLNALKITQYDTPPWDMTSQNSFRNQLEGFINGPQLH N

[0568] RVHRWVGGQMGVVPTAPN DPVFFLHHANVDRIWAVWQIIH RNQNYQPMKNGPFGQN FRDPMYP WNTTPEDVMNHRKLGYVYDIELRKSKRSSLEH HH HHH

[0569] Hsp82 H7 (SEQ ID NO:50):

[0570] MASETFEFQAEITQLMSLIINTVYSN KEIFLRELISNASDALDKIRYKSLSDPKQLETEPDLFIRITPKPEQKVLE

[0571] IRDSGIGMTKAELIN NLGTIAKSGTKAFMEALSAGADVSMIGQFGVGFYSLFLVADRVQVISKSNDDEQYI

[0572] WESNAGGSFTVTLDEVNERIGRGTILRLFLKDDQLEYLEEKRIKEVIKRHSEFVAYPIQLVVTKEVEKEVPIPE

[0573] EEKKDEEKKDEEKKDEDDKKPKLEEVDEEEEKKPKTKKVKEEVQEIEELNKTKPLWTRNPSDITQEEYNAFY

[0574] KSISN DWEDPLYVKHFSVEGQLEFRAILFIPKRAPFDLFESKKKKNNIKLYVRRVFITDEAEDLIPEWLSFVKG

[0575] VVDSEDLPLN LSREMLQQNKIMKVIRKNIVKKLIEAFN EIAEDSEQFEKFYSAFSKNIKLGVHEDTQNRAAL

[0576] AKLLRYNSTKSVDELTSLTDYVTRMPEHQKNIYYITGESLKAVEKSPFLDALKAKNFEVLFLTDPIDEYAFTQL

[0577] KEFEGKTLVDITKDFELEETDEEKAEREKEIKEYEPLTKALKEILGDQVEKVVVSYKLLDAPAAIRTGQFGWS

[0578] ANM ERIMKAQALRDSSMSSYMSSKKTFEISPKSPIIKELKKRVDEGGAQDKTVKDLTKLLYETALLTSGFSL

[0579] DEPTSFASRIN RLISLGLNIDEDEETETAPEASTAAPVEEVPADTEMEEVDPGEQKCEEWKRRYEKEKEKN

[0580] ARLKGKVEKLEIELARWRPGSAWSHHH HHH H

[0581] Hsp82 D452TAG H7 (SEQ ID NO:51):

[0582] MASETFEFQAEITQLMSLIINTVYSN KEIFLRELISNASDALDKIRYKSLSDPKQLETEPCLFIRITPKPEQKVLE

[0583] IRDSGIGMTKAELIN NLGTIAKSGTKAFMEALSAGADVSMIGQFGVGFYSLFLVADRVQVISKSNDDEQYI

[0584] WESNAGGSFTVTLDEVNERIGRGTILRLFLKDDQLEYLEEKRIKEVIKRHSEFVAYPIQLVVTKEVEKEVPIPE

[0585] EEKKDEEKKDEEKKDEDDKKPKLEEVDEEEEKKPKTKKVKEEVQEIEELNKTKPLWTRNPSDITQEEYNAFY

[0586] KSISN DWEDPLYVKHFSVEGQLEFRAILFIPKRAPFDLFESKKKKNNIKLYVRRVFITDEAEDLIPEWLSFVKG

[0587] VVDSEDLPLN LSREMLQQNKIMKVIRKNIVKKLIEAFN EIAEDSEQFEKFYSAFSKNIKLGVHEDTQNRAAL

[0588] AKLLRYNSTKSV* ELTSLTDYVTRMPEHQKN IYYITGESLKAVEKSPFLDALKAKNFEVLFLTDPIDEYAFTQL

[0589] KEFEGKTLVDITKDFELEETDEEKAEREKEIKEYEPLTKALKEILGDQVEKVVVSYKLLDAPAAIRTGQFGWS

[0590] ANM ERIMKAQALRDSSMSSYMSSKKTFEISPKSPIIKELKKRVDEGGAQDKTVKDLTKLLYETALLTSGFSL

[0591] DEPTSFASRIN RLISLGLNIDEDEETETAPEASTAAPVEEVPADTEMEEVDPGEQKCEEWKRRYEKEKEKN

[0592] ARLKGKVEKLEIELARWRPGSA WSH HHH HHH

[0593] MAPTSSSTKKTQLQLEHLLLDLQM ILNGIN NYKNPKLTRMLTFKFYMPKKATELKHLQCLEEELKPLEEVLL

[0594] AQSKNFHLRPRDLISN INVIVLELKGSETTFMCEYADETATIVEFLNRWITFCQSIISTLTGSGSHHH HHH

[0595] IL-2 R38TAG (SEQ ID NO:72)

[0596] MAPTSSSTKKTQLQLEHLLLDLQM ILNGIN NYKNPKLT* MLTFKFYMPKKATELKHLQCLEEELKPLEEVLN

[0597] LAQSKN FH LRPRDLISNINVIVLELKGSETTFMCEYADETATIVEFLNRWITFCQSIISTLTGSGSHHH HHH TRX H6 hGH wt (SEQ ID NO:73)

[0598] MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPK

[0599] YGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDAN LAGSGSGHHHH HHHGSGSGENLYFQGFPTIPLSR LFDNAMLRADRLNQLAFDTYQEFEEAYIPKEQKYSFLQNPQTSLCFSESIPTPSNREETQQKSNLELLRISLL LIQSWLEPVQFLRSVFANSLVYGASDSNVYDLLKDLEEKIQTLMGRLEDGSPRTGQIFKQTYSKFDTNSH N DDALLKNYGLLYCFNADMSRVSTFLRTVQCRSVEGSCGF

[0600] TRX H6 hGH Y35TAG (SEQ ID NO:74)

[0601] MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPK

[0602] YGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDAN LAGSGSGHHHH HHHGSGSGENLYFQGFPTIPLSR LFDNAMLRADRLNQLAFDTYQEFEEA*IPKEQKYSFLQN PQTSLCFSESIPTPSNREETQQKSNLELLRISLL LIQSWLEPVQFLRSVFANSLVYGASDSNVYDLLKDLEEKIQTLMGRLEDGSPRTGQIFKQTYSKFDTNSH N DDALLKNYGLLYCFNADMSRVSTFLRTVQCRSVEGSCGF

[0603] TRX H6 hGH K38TAG (SEQ ID NO:75)

[0604] MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPK

[0605] YGIRGIPTLLLFKNGEVAATKVGALSKGQLKEFLDAN LAGSGSGHHHH HHHGSGSGENLYFQGFPTIPLSR LFDNAMLRADRLNQLAFDTYQEFEEAYIP* EQKYSFLQNPQTSLCFSESIPTPSNREETQQKSNLELLRISLL LIQSWLEPVQFLRSVFANSLVYGASDSNVYDLLKDLEEKIQTLMGRLEDGSPRTGQIFKQTYSKFDTNSHN DDALLKNYGLLYCFNADMSRVSTFLRTVQCRSVEGSCGF

[0606] GST TEV RanGAPl H6 wt (SEQ ID NO:76)

[0607] MSPILGYWKIKGLVQPTRLLLEYLEEKYEEH LYERDEGDKWRN KKFELGLEFPNLPYYIDGDVKLTQSMAII

[0608] RYIADKHNM LGGCPKERAEISMLEGAVLDIRYGVSRIAYSKDFETLKVDFLSKLPEMLKMFEDRLCHKTYL

[0609] NGDHVTHPDFMLYDALDVVLYM DPMCLDAFPKLVCFKKRIEAIPQIDKYLKSSKYIAWPLQGWQATFGG

[0610] GDH PPKGIEENLYFQGNTGEPAPVLSSPPPADVSTFLAFPSPEKLLRLGPKSSVLIAQQTDTSDPEKVVSAF

[0611] LKVSSVFKDEATVRMAVQDAVDALMQKAFNSSSFNSNTFLTRLLVHMGLLKSEDKVKAIAN LYGPLMAL

[0612] NHMVQQDYFPKALAPLLLAFVTKPNSALESCSFARHSLLQTLYKVH HHH HH

[0613] GST TEV RanGAPl H6 K524TAG (SEQ ID NO:77)

[0614] MSPILGYWKIKGLVQPTRLLLEYLEEKYEEH LYERDEGDKWRN KKFELGLEFPNLPYYIDGDVKLTQSMAII

[0615] RYIADKHNM LGGCPKERAEISMLEGAVLDIRYGVSRIAYSKDFETLKVDFLSKLPEMLKMFEDRLCHKTYL

[0616] NGDHVTHPDFMLYDALDVVLYM DPMCLDAFPKLVCFKKRIEAIPQIDKYLKSSKYIAWPLQGWQATFGG

[0617] GDH PPKGIEENLYFQGNTGEPAPVLSSPPPADVSTFLAFPSPEKLLRLGPKSSVLIAQQTDTSDPEKVVSAF

[0618] LKVSSVFKDEATVRMAVQDAVDALMQKAFNSSSFNSNTFLTRLLVHMGLL*SEDKVKAIANLYGPLMAL

[0619] NHMVQQDYFPKALAPLLLAFVTKPNSALESCSFARHSLLQTLYKVH HHH HH

[0620] AcKRS3 (SEQ ID NO:52): MDKKPLDVLISATGLWMSRTGTLH KIKHHEVSRSKIYIEMACGDHLVVN NSRSCRTARAFRHHKYRKTCK

[0621] RCRVSGEDIN NFLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVGAKASTNTSRSVPSPAKST

[0622] PNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLN MAKPFRELEPELVTRRKNDFQRLYTN DREDYLGKLER

[0623] DITKFFVDRGFLEIKSPIUPAEYVERMGIN NDTELSKQIFRVDKNLCLRPMMAPTIFNYARKLDRILPGPIKIF

[0624] EVGPCYRKESDGKEH LEEFTMVNFFQMGSGCTRENLEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMHG

[0625] DLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTN L

[0626] MmPyIRS Y306A Y384A (SEQ ID NO:78)

[0627] MDKKPLNTLISATGLWMSRTGTIHKIKHH EVSRSKIYIEMACGDH LVVN NSRSSRTARALRHHKYRKTCK

[0628] RCRVSDEDLNKFLTKANEDQTSVKVKVVSAPTRTKKAM PKSVARAPKPLENTEAAQAQPSGSKFSPAIPV

[0629] STQESVSVPASVSTSISSISTGATASALVKGNTNPITSMSAPVQASAPALTKSQTDRLEVLLNPKDEISLNSG

[0630] KPFRELESELLSRRKKDLQQIYAEERENYLGKLEREITRFFVDRGFLEIKSPILIPLEYIERMGIDNDTELSKQIF

[0631] RVDKN FCLRPMLAPNLANYLRKLDRALPDPIKIFEIGPCYRKESDGKEH LEEFTMLNFCQMGSGCTREN LE

[0632] SIITDFLN HLGIDFKIVGDSCMVFGDTLDVMHGDLELSSAVVGPIPLDREWGIDKPWIGAGFGLERLLKVK

[0633] HDFKNIKRAARSESYYNGISTNL

[0634] MbPyIRS Y271A L274M (SEQ ID NO:79)

[0635] MDKKPLDVLISATGLWMSRTGTLH KIKHHEVSRSKIYIEMACGDHLVVN NSRSCRTARAFRHHKYRKTCK

[0636] RCRVSDEDINN FLTRSTESKNSVKVRVVSAPKVKKAMPKSVSRAPKPLENSVSAKASTNTSRSVPSPAKST

[0637] PNSSVPASAPAPSLTRSQLDRVEALLSPEDKISLN MAKPFRELEPELVTRRKNDFQRLYTN DREDYLGKLER

[0638] DITKFFVDRGFLEIKSPILIPAEYVERMGIN NDTELSKQIFRVDKNLCLRPMLAPTLANYM RKLDRILPGPIKI

[0639] FEVGPCYRKESDGKEHLEEFTMVNFCQMGSGCTREN LEALIKEFLDYLEIDFEIVGDSCMVYGDTLDIMH

[0640] GDLELSSAVVGPVSLDREWGIDKPWIGAGFGLERLLKVMHGFKNIKRASRSESYYNGISTN L sfGFP N40TAA N150TAG H6 (SEQ ID NO:53):

[0641] MPSKGEELFTGVVPILVELDGDVNGH KFSVRGEGEGDAT*GKLTLKFICTTGKLPVPWPTLVTTLTYGVQC

[0642] FSRYPDHM KRHDFFKSAMPEGYVQERTISFKDDGTYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGH KL

[0643] EYNFNSH*VYITADKQKNGIKAN FKIRHNVEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSVLSKDP NEKRDH MVLLEFVTAAGITHGM DELYKGSHH HHHH

[0644] Ub K48TAA TEV SUMO2K11TAG H6 (SEQ ID NO:54):

[0645] MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQQRLIFAG*QLEDGRTLSDYN IQKESTLHLVLRL

[0646] RGGEDLYFQSMADEKPKEGVKTENN DH IN L*VAGQDGSVVQFKIKRHTPLSKLMKAYCERQGLSMRQIR

[0647] FRFDGQPINETDTPAQLEMEDEDTIDVFQQQQNGLH HH HHH H

[0648] Synthesis of peptides via solid phase peptide synt...

Claims

CLAIMS1. A compound comprising an isopeptide-linked lysine or an analogue thereof, the compound being a compound of formula (I):or a salt thereof, wherein:A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(Ci-s alkyl)(Ci-s alkyl), -N+(Ci-s alkyl)(Ci-5alkyl), -NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2, or A is a peptidyl group;B is selected from:wherein the left empty valence is connected to A and the right empty valence is connected to C, wherein:Z is hydrogen or a side chain of an amino acid, in particular wherein Z is selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, -(C1-6 alkylene)-S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO-Rz, -(C1-6 alkylene)-COO-Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-Rz, -(C1-6 alkylene)-N(Rz)-CO- (C1-6 alkyl), -(C1-6 alkylene)-CO-N(Rz)-O-Rz, -(Ci-6alkylene)-O-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-C(=N-Rz)-N(Rz)-Rz, -(Ci-6alkylene)-SO3-Rz, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z groups are each optionally substituted with one or more groups selected from -OH, -SH and Hal, further wherein one -CH2- group in said alkyl may be replaced with further wherein each Rzis independently selected from hydrogennd wherein Rzzis selected from hydrogen, Ci-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-5 alkenyl), -CO(Ci-s alkynyl), -CO(Ci-s alkylene)-COOH, -CO(Ci-salkenylene)-COOH, -CO(Co-s alkylene)-carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -C0-0-(Ci-5 alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-s alkynyl), -CO-0-(Co-s alkylene)- carbocyclyl, -CO-0-(Co-s alkylene)-heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs,Z1and Z2are each independently selected from hydrogen, Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)-Hal, -(C1-6 alkylene)-O-Rz, -(C1-6 alkylene)-S-Rz, -(C1-6 alkylene)-N(Rz)-Rzz, -(C1-6 alkylene)-CO- Rz, -(C1-6 alkylene)-COO-Rz, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO- N(RZ)-RZ, -(C1-6 alkylene)-N(Rz)-CO-(Ci-6alkyl), -(Ci-6alkylene)-CO-N(Rz)-O-Rz, -(Ci-6alkylene)-O-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)-CO-N(Rz)-Rz, -(Ci-6alkylene)-N(Rz)- C(=N-RZ)-N(RZ)-RZ, -(Ci-6 alkylene)-SO3-Rz, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned Z1and Z2groups are each optionally substituted with one or more -OH, -SH or Hal, wherein each Rzis independently selected from hydrogen and Ci-6 alkyl, and Rzzis selected from hydrogen, Ci-6 alkyl, -CHO, -CO(Ci-s alkyl), -CO(Ci-s alkenyl), -CO(Ci-s alkynyl), -CO(Ci-s alkylene)-COOH, -CO(Ci-s alkenylene)-COOH, -CO(Co-s alkylene)- carbocyclyl, -CO(Co-s alkylene)-heterocyclyl, -CO-O-(Ci-s alkyl), -CO-O-(Ci-s alkenyl), -CO-O-(Ci-5 alkynyl), -CO-0-(Co-s alkylene)-carbocyclyl, -CO-0-(Co-s alkylene)- heterocyclyl, wherein said carbocyclyl and said heterocyclyl are each optionally substituted with one or more groups selected from Rs, and further wherein one - N = NCH2- group in said alkyl may be replaced with, provided that at least one of Z1and Z2is not hydrogen, further provided that if Z1and Z2are connected to the same carbon atom, then both Z1and Z2are not hydrogen, or Z1and Z2are joined together to form, together with carbon atom(s) that otherwise carry Z1and Z2, a non-aromatic carbocyclic ring or non-aromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non- aromatic heterocyclic ring are each optionally substituted with one or more Rs;-C-D- are defined as follows:C is selected from -CO-NH-, -CO-O-, -CO-S- and -CO-N(Ci-s alkyl)-, wherein the left empty valence is connected to B and the right empty valence is connected to D, andD is selected fromwherein the left empty valence is connected to C and the right empty valence is connected to -CO-NH-, wherein:X is a hydrogen or a side chain of an amino acid, in particular wherein X is selected from hydrogen, Ci-s alkyl, C2-8 alkenyl, C2-8 alkynyl, -(Ci-6 alkylene)- N3, -(Ci-6 alkylene)-CN, -(Ci-6 alkylene)-Hal, -(Ci-6 alkylene)-O-Rx, -(Ci-6 alkylene)-S-Rx, -(Ci-6 alkylene)-N(Rx)-Rx, -(Ci-6 alkylene)-CO-Rx, -(Ci-6 alkylene)-COO-Rx, -(Ci-6 alkylene)-O-CO-(Ci-6 alkyl), -(Ci-6 alkylene)-CO-N(Rx)- Rx, -(Ci-6 alkylene)-N(Rx)-CO-(Ci-6alkyl), -(Ci-6alkylene)-CO-N(Rx)-O-Rx, -(Ci-6alkylene)-O-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(Ci-6 alkylene)-SO3-Rx, -(Co-6 alkylene)- carbocyclyl, and -(Co-6 alkylene)-heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)-carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X groups are each optionally substituted with one or more -OH, -SH or -Hal, wherein each Rxis independently selected from hydrogen and Ci-6 alkyl, and further whereinN=N one -CH2- group in said alkyl may be replaced withX1and X2are each independently selected from hydrogen, C1-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-N3, -(C1-6 alkylene)-CN, -(C1-6 alkylene)- Hal, -(C1-6 alkylene)-O-Rx, -(C1-6 alkylene)-S-Rx, -(C1-6 alkylene)-N(Rx)-Rx, -(C1-6 alkylene)-CO-Rx, -(C1-6 alkylene)-COO-Rx, -(C1-6 alkylene)-O-CO-(Ci-6 alkyl), - (C1-6 alkylene)-CO-N(Rx)-Rx, -(C1-6 alkylene)-N(Rx)-CO-(Ci-6 alkyl), -(C1-6 alkylene)-CO-N(Rx)-O-Rx, -(Ci-6alkylene)-O-CO-N(Rx)-Rx, -(Ci-6alkylene)-N(Rx)-CO-N(Rx)- Rx, -(Ci-6alkylene)-N(Rx)-C(=N-Rx)-N(Rx)-Rx, -(Ci-6alkylene)-SO3-Rx, -(Co-6 alkylene)-carbocyclyl, and -(Co-6 alkylene)- heterocyclyl, wherein the carbocyclyl group in said -(Co-6 alkylene)- carbocyclyl and the heterocyclyl group in said -(Co-6 alkylene)-heterocyclyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group comprised in any of the aforementioned X1and X2groups are each optionally substituted with one or more -OH, -SH or Hal (preferably with one or more -OH), wherein each Rxis independently selected from hydrogen and Ci-6 alkyl, furtherN = N wherein one -CH2- group in said alkyl may be replaced withprovided that at least one of X1and X2is not hydrogen, further provided that if X1and X2are connected to the same carbon atom, then both X1and X2are not hydrogen, or X1and X2are joined together to form, together with carbon atom(s) that otherwise carry X1and X2, a non-aromatic carbocyclic ring or non-aromatic heterocyclic ring, wherein said non-aromatic carbocyclic ring and said non-aromatic heterocyclic ring are each optionally substituted with one or more Rs; or -C-D- taken together is -CO-(N-heterocycloalkylene)-, wherein said N- heterocycloalkylene is connected to said -CO- group of -C-D- through its nitrogen atom and wherein said N-heterocycloalkylene is optionally substituted with one or more Rs;E is - (C1-6 alkyleneJ-CHY^2, wherein Y1is selected from hydrogen, -SH, -OH and -NH2, and wherein Y2is hydrogen, COOH or CONH2, provided that at least one of Y1and Y2is not hydrogen, further wherein said alkylene is optionally substituted with an -OH group, -SH group or Hal, and wherein one -CH2- group in said alkylene may be replaced witheach Rsis independently selected from C1-5 alkyl, C2-5 alkenyl, C2-5 alkynyl, -(C0-3 alkylene)-OH, -(C0-3 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-O(Ci-s alkylene)-OH, -(C0-3 alkylene)-O(Ci-5 alkylene)-O(Ci-s alkyl), -(C0-3 alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-S(Ci-s alkylene)-SH, -(C0-3 alkylene)-S(Ci-s alkylene)-S(Ci-s alkyl), -(C0-3 alkylene)-NH2, -(C0-3 alkylene)-NH(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s al kyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-OH, -(C0-3 alkylene)-N(Ci-s alkyl)-O(Ci-s alkyl), -(C0-3 alkylene)-halogen, -(C0-3 alkylene)-(Ci-s haloalkyl), -(C0-3 alkylene)-O-(Ci-s haloalkyl), -(C0-3 alkylene)-CN, -(C0-3 alkylene)-NO2, -(C0-3 alkylene)-CHO, -(C0-3 alkylene)-CO-(Ci-s alkyl), -(C0-3 alkylene)-COOH, -(C0-3 alkylene)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH2, -(C0-3 alkylene)-CO-NH(Ci-s alkyl), -(C0-3 alkylene)-CO-N(Ci-5 alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-CO-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-5 alkyl)-CO-(Ci-s alkyl), -(C0-3 alkylene)-CO-NH(Ci-s alkylene)-CN, -(C0-3 alkylene)-CO-N(Ci-5 alkyl)(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-(Ci-s alkylene)- CN, -(C0-3 alkylene)-N(Ci-5 alkyl)-CO-(Ci-s alkylene)-CN, -(C0-3 alkylene)-NH-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-N(Ci-s alkyl)-CO-O-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-NH-(Ci-s alkyl), -(C0-3 alkylene)-O-CO-N(Ci-s alkyl)-(Ci-s alkyl), -(C0-3 alkylene)-SO2-NH2, -(C0-3 alkylene)-SO2-NH(Ci-5 alkyl), -(C0-3 alkylene)-SO2-N(Ci-s alkyl)(Ci-s alkyl), -(C0-3 alkylene)-NH-SO2-(Ci-5 alkyl), -(C0-3 alkylene)-N(Ci-s a I ky l)-SC>2-(Ci-5 alkyl), -(C0-3 alkylene)-SO2-(Ci-5alkyl), -(C0-3 alkylene)-SO-(Ci-5alkyl), -OSO2F, -B(OH)2, -PO4H, -PO3H, -SO4H, -SO3H, -L-(CO-3 alkylenej-carbocyclyl, and -L-(Co-3 alkylene)-heterocyclyl, wherein L is selected from -(C0-3 alkylene)-NH-, -(C0-3 alkylene)-O-, -(C0-3 alkylene)-CO-, -(C0-3 alkylene)-NH-CO-, -(C0-3 alkylene)-CO-NH-, -(C0-3 alkylene)-NH-CO-O-, -(C0-3 alkylene)— O-CO-NH-, -(C0-3 alkylene)-NH-CO-NH- and -(C0-3 alkylene)-NH-CO-NH-, and wherein the carbocyclyl moiety in said -L-(Co-3 alkylene)-carbocyclyl and the heterocyclyl moiety in said -L-(CO-3alkylene)-heterocyclyl are each optionally substituted with one or more groups independently selected from C1-4 alkyl, halogen, -CN, -NO2, -OH, -O-(Ci-4 alkyl), - SH, -S-(Ci-4alkyl), -NH2, -NH(CI-4alkyl), -N(CI-4alkyl)(Ci-4alkyl), -COOH, -COO(Ci-4alkyl), -CONH2, -CON H (CI-4alkyl), -CON(CI-4al kyl)(Ci-4alkyl), -N HCO(CI-4alkyl) and -N(CI-4al ky I )- CO(Ci-4 alkyl).

2. The compound of claim 1, wherein A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), - N(CI-5 alkyl)(Ci-5 alkyl), -NH-CO-(CI-5alkyl), -COOH, -SO3H, -SO2H, -N3, and -NO2.

3. The compound of claim 1 or 2, wherein A is selected from -NH2, -OH, -SH, -NH(Ci-s alkyl), -N(CI-5 al kyl)(Ci-5 alkyl), -COOH, and -N3.

4. The compound of any one of claims 1 to 3, wherein A is -NH2, -NHfCHs), -COOH, -SH, and N3, preferably wherein a is NH2.Z5. The compound of any one of claims 1 to 4, wherein B ispreferably wherein BZ iswherein the left empty valence is connected to A and the right empty valence is connected to C.

6. The compound of claim 5, wherein Z is selected from hydrogen, -(C1-6 alkylene)-N(Rz)- Rzz, and -(C1-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), preferably wherein Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and -CH2CH2CH2CH2-NH-CO-CH3; or wherein Z is selected from any one of:

7. The compound of claim 6, wherein Z is hydrogen.

8. The compound of any one of claims 1 to 7, wherein C is selected from -CO-NH- and -CO- O-, wherein the left empty valence is connected to B and the right empty valence is connected to D.

9. The compound of claim 8, wherein C is -CO-NH-, wherein the left empty valence is connected to B and the right empty valence is connected to D.X10. The compound of any one of claims 1 to 9, whereinpreferably wherein DX is' wherein the left empty valence is connected to C and the right empty valence is connected to -CO-NH- group.

11. The compound of claim 10, wherein X is selected from Ci-8 alkyl, C2-8 alkenyl, C2-8 alkynyl, -(C1-6 alkylene)-OH, -(C1-6 alkylene)-O(Ci-6 alkyl), -(C1-6 alkylene)-SH, -(C1-6 alkylene)-S(Ci-6 alkyl), -(C1-6 alkylene)-N3, -(Ci-6 alkylene)-CI, -(C1-6 alkylene)-NH2, -(C1-6 alkylene)-NH-CO- NH2, -(C1-6 alkylene)-NH-C(=NH)-NH2, -(Ci-6alkylene)-O-CO-NH2, -(Ci-6alkylene)-COOH, - (C1-6 alkylene)-CO-NH2, -(Ci-6 alkylene)-CO-NH-OH, -(Ci-6 alkyleneJ-SOsH, phenyl, -(Ci-6 alkylene)-phenyl, cycloalkyl, -(Ci-6 alkylene)-cycloalkyl, heteroaryl, -(Ci-6 alkylene)- heteroaryl, heterocycloalkyl, and -(Ci-6 alkylene)-heterocycloalkyl, wherein said phenyl, the phenyl group in said -(Ci-6 alkylene)-phenyl, said cycloalkyl, the cycloalkyl group in said -(Ci-6 alkylene)-cycloalkyl, said heteroaryl, the heteroaryl group in said -(Ci-6 alkylene)-heteroaryl, said heterocycloalkyl, and the heterocycloalkyl group in said -(Ci-6 alkylene)-heterocycloalkyl are each optionally substituted with one or more groups Rs, wherein said alkyl, said alkenyl, said alkynyl, and any alkylene group in any of the aforementioned groups are each optionally substituted with one or more -OH, andN = N further wherein one -CH2- group in said alkyl may be replaced with12. The compound of claim 11, wherein X is selected from methyl, ethyl, isopropyl, secbutyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl-diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4-yl)methyl and benzyl.

13. The compound of any one of claims 1 to 7, wherein -C-D- is a moiety according to formulawherein the left empty valence is connected to B, and the right empty valence is connected to -CO-NH- group.

14. The compound of any one of claims 1 to 13, wherein E is selected from: -CH2CH2CH2CH2- CH(-NH2)-COOH, -CH2CH2CH2-CH(-NH2)-COOH, -CH2CH2CH2CH2CH2-COOH, CH2CH2CH2CH2CH2-NH2, -CH2CH2CH2CH2-CH(-NH2)-CONH2, -CH2CH2CH2CH2-CH(-OH)-COOH; -CH2CH2-CH(-NH2)-COOH, and -CH2-CH(-NH2)-COOH, preferably wherein E is CH2CH2CH2CH2-CH(-NH2)-COOH.

15. The compound of claim 1, wherein the compound is a compound according to formulaor its salt, preferably wherein the compound of formula (II) is preferably a compound of formula (lib):or its salt, wherein Z and X are as defined in any one of claims 1 to 14, preferably wherein Z is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-e alkylene)-N(Rz)-CO-(Ci-6 alkyl), more preferably wherein Z is selected from hydrogen, -CH2CH2CH2CH2-NH2and -CH2CH2CH2CH2-NH-CO-CH3, even more preferably wherein Z is hydrogen; or wherein Z is selected from any one of:preferably wherein X is selected from methyl, ethyl, isopropyl, sec-butyl, isobutyl, hydroxymethyl, 2-hydroxypropyl, thiomethyl, propargyl, chloromethyl, (3-methyl- diazirin-3-yl)methyl, azidomethyl, aminomethyl, (imidazol-4-yl)methyl and benzyl; or wherein the compound is a compound of formula (III):or its salt, preferably wherein the compound of formula (III) is a compound of formulaor its salt, wherein Rpis hydrogen or -OH, and wherein Z is as defined in any one of claims 1 to 14, preferably wherein is selected from hydrogen, -(Ci-6 alkylene)-N(Rz)-Rzz, and -(Ci-6 alkylene)-N(Rz)-CO-(Ci-6 alkyl), more preferably wherein Z is selected from hydrogen, -CH2CH2CH2CH2-NH2 and - CH2CH2CH2CH2-NH-CO-CH3, even more preferably wherein Z is hydrogen; or wherein Z is selected from any one of:

16. The compound of claim 1, wherein the compound is a compound selected from Table 1, or its salt.

17. A method for identifying an engineered peptide-binding protein having increased affinity and / or specificity for the compound according to any one of claims 1 to 16, the method comprising the steps of: a) providing a nucleic acid encoding a peptide-binding protein comprising one or more mutations; b) expressing the nucleic acid of step (a) in a cell comprising an orthogonal translation system, wherein the orthogonal translation system comprises an orthogonal aminoacyl- tRNA synthetase / tRNA pair specific for a non-canonical amino acid, preferably an isopeptide-linked lysine, or an analog thereof, and a reporter protein gene comprisingone or more codons that have been re-allocated for the incorporation of the isopeptide- linked lysine or the analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; c) contacting the cell of step (b) with a compound according to any one of claims 1 to 16, wherein the compound is transported into the cell and cleaved inside the cell to release the isopeptide-linked lysine or an analog thereof; d) detecting the expression of the reporter protein gene by the cell, wherein the expression is dependent on the successful uptake of the compound in step (c); and e) identifying the engineered peptide-binding protein as having increased affinity and / or specificity for the compound according to any one of claims 1 to 16 based on the expression of the reporter protein gene detected in step (d).

18. The method according to claim 17, wherein the peptide-binding protein is the peptide- binding protein of an oligopeptide permease (OppA).

19. The method according to claim 17 or 18, wherein the peptide-binding protein is the peptide-binding protein of the E. coli oligopeptide permease (OppA; SEQ ID NO:1).

20. The method according to claim 19, wherein in step (a) one or more mutations are introduced at positions V60, S63, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, C443, T429, W442, D445, S460, V482, L531 and / or N533 of SEQ ID NO:1.

21. The method according to any one of claims 17 to 20, wherein the mutations introduced in step (a) are introduced using error-prone PCR and / or saturated mutagenesis and / or continuous directed evolution.

22. The method according to any one of claims 17 to 21, wherein the cell in step (b) is a prokaryotic cell, in particular an E. coli cell.

23. The method according to any one of claims 17 to 22, wherein the orthogonal aminoacyl- tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair.

24. The method according to any one of claims 17 to 23, wherein the reporter protein gene in step (b) comprises one or more amber stop codons that have been re-allocated to incorporate the isopeptide-linked lysine or the analog thereof.

25. The method according to any one of claims 17 to 24, wherein in step (c), the cell is contacted with a compound according to any one of claims 1 to 16 in the presence of apeptide that competes for binding to the peptide-binding protein.

26. The method according to any one of claims 17 to 25, wherein the reporter protein is a fluorescent protein.

27. The method according to claim 26, wherein the detection of the expression of the reporter protein gene in step (d) is performed using fluorescence-activated cell sorting (FACS).

28. The method according to any one of claims 17 to 25, wherein the detection of the expression of the reporter protein gene in step (d) is performed using antibiotic resistance screening and / or using growth-based selection.

29. The method according to any one of claims 17 to 28, wherein the isopeptide-linked lysine or the analog thereof is obtained by intracellular cleavage of the compound of any one of claims 1 to 16.

30. The method according to any one of claims 17 to 29, wherein identifying the engineered peptide-binding protein as having increased specificity for the compound in step (e) comprises a step of determining the sequence of the nucleic acid encoding the peptide- binding protein and / or determining specific mutations in the nucleic acid encoding the peptide-binding protein.

31. The method according to any one of claims 17 to 30, wherein the isopeptide-linked lysine or the analog thereof has the structure:wherein C' is -NH2 or -OH, and wherein D and E is defined as in any one of claims 1 to 16, wherein in D the left empty valence is connected to C'.

32. An engineered variant of an oligopeptide-binding protein OppA from E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, 563, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, 5460, V482, L531 and / or N533.

33. The engineered variant of claim 32, wherein the one or more mutations are located inpositions: V60, 863, Y135, T173, H187, V193, D221, W222, 1303, N337, R439, W442, C443, D445, 8460, L531 and / or N533.

34. The engineered variant of claim 32 or 33, wherein the one or more mutations are located in positions: V60, 863, Y135, H187, D221, W222, N337, R439, C443, D445, 8460, L531 and / or N533.

35. The engineered variant of any one of claims 32 to 34, wherein the one or more mutations are located in positions: V60, 863, Y135, H187, D221, W222, R439, C443, D445, 8460, L531 and / or N533.

36. The engineered variant of any one of claims 32 to 35, wherein the one or more mutations are located in positions: V60, 563, Y135, H187, R439, C443, D445, 5460, L531 and / or N533.

37. A nucleic acid encoding the engineered variant of any one of claims 32 to 36.

38. A vector comprising the nucleic acid according to claim 37.

39. A cell comprising the nucleic acid according to claim 37 or the vector according to claim 38.

40. The cell according to claim 39, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

41. A method for incorporating a non-canonical amino acid, or an analog thereof, into a protein, the method comprising the steps of: a) providing a cell comprising at least one orthogonal translation system, wherein the at least one orthogonal translation system comprises an orthogonal aminoacyl-tRNA synthetase / tRNA pair specific for a non-canonical amino acid, or an analog thereof, comprised in the compound according to any one of claims 1 to 16, and a nucleic acid molecule encoding the protein, wherein the nucleic acid molecule encoding the protein comprises one or more codons that have been re-allocated for the incorporation of the non-canonical amino acid or an analog thereof by the orthogonal aminoacyl-tRNA synthetase / tRNA pair; b) contacting the cell of step (a) with the compound according to any one of claims 1 to 16, wherein the compound is transported into the cell and cleaved inside the cell to release an isopeptide-linked lysine or an analog thereof and, optionally, another non- canonical amino acid or an analog thereof;c) incorporating (i) the isopeptide-linked lysine or an analog thereof and / or (ii) the other non-canonical amino acid, or the analog thereof, into the protein in response to the one or more re-allocated codons using the orthogonal translation system.

42. The method according to claim 41, wherein the cell is a prokaryotic cell, in particular an E. coli cell.

43. The method according to claim 42 or 43, wherein the cell expresses an oligopeptide permease.

44. The method according to claim 43, wherein the oligopeptide permease comprises a peptide-binding protein that has been engineered for increased specificity for the compound according to any one of claims 1 to 16.

45. The method according to claim 44, wherein the engineered peptide-binding protein is OppA of E. coli (SEQ ID NO:1) comprising mutations in one or more of the following positions: V60, S63, L78, Y135, T173, H187, V193, D221, W222, 1303, K307, K333, N337, K371, R439, W442, C443, D445, T429, S460, V482, L531 and / or N533.

46. The method according to any one of claims 41 to 45, wherein the orthogonal aminoacyl- tRNA synthetase / tRNA pair is a Pyrrolysyl-tRNA synthetase / tRNA pair.

47. The method according to any one of claims 41 to 46, wherein the nucleic acid molecule encoding the protein comprises one or more amber stop codons that have been reallocated to incorporate the non-canonical amino acid or an analog thereof.

48. The method according to any one of claims 41 to 47, wherein cleavage of the molecule according to any one of claims 1 to 16 inside the cell releases a second non-canonical amino acid (ncAA), and wherein the second ncAA is incorporated into the same or a different protein in response to a re-allocated codon.

49. The method according to any one of claims 41 to 48, wherein the isopeptide-linked lysine or the analog thereof has the structure:wherein C' is -NH2 or -OH, and wherein D and E is defined as in any one of claims 1 to 16, wherein the left empty valence in D is connected to C'.

Citation Information

Patent Citations

  • Fluorine- substituted 2 ', 6 ' -dimethyl-l-tyrosine-1, 2,3, 4-tetrahydr0is0QUIN0line-3-CARB0xylic acid peptides (DMT-tic) for use as MYU- and delta-opioid receptor probes in

    WO2009032840A1

  • Means and methods for site-specific protein modification using transpeptidases

    WO2020007899A1