Adaptor systems for nonribosomal peptide synthetases and polyketide synthases
By employing SYNZIP coiled coils for protein-protein interactions, the assembly of NRPS/PKS enzyme complexes is facilitated, addressing the challenge of efficient peptide production in NRPS/PKS expression systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2026-03-10
AI Technical Summary
Existing systems for expressing nonribosomal peptide synthetases (NRPSs) and polyketide synthases (PKSs) face challenges in efficiently producing peptide products due to difficulties in assembling these large multidomain proteins, particularly when longer NRPs are involved.
The use of protein fragments with intentionally introduced protein-protein interactions, facilitated by SYNZIP coiled coils, allows for post-translationally assembling functional NRPS/PKS enzyme complexes through vector systems.
This approach enables flexible and efficient recombinant production of NRPS modules, overcoming assembly limitations and enhancing peptide production capabilities.
Smart Images

Figure 0007827305000002 
Figure 0007827305000003 
Figure 0007827305000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to systems for expressing nonribosomal peptide synthetases (NRPSs), polyketide synthases (PKSs), or NRPS / PKS hybrid synthases (synthetases). NRPSs, PKSs, or hybrids thereof are large multidomain proteins or complexes, the expression of which often makes peptide production difficult. Therefore, the present invention relates to systems for expressing enzyme fragments that can be post-translationally assembled through intentionally introduced protein-protein interactions to form functional multienzyme complexes. The present invention discloses such protein fragments and the nucleic acids encoding them. Also disclosed are vector systems for the protein fragments of the present invention and their use for producing functional NRPS / PKS enzyme complexes. [Background technology]
[0002] Nonribosomal peptides (NRPs) are peptides produced by nonribosomal peptide synthetases (NRPSs) with a wide structural diversity, particularly those characterized by cyclic, branched, or other complex primary structures (NPL 1, NPL 2). Due to their structural complexity, many of these molecules exhibit therapeutic properties, including antibiotic, immunosuppressive, or anticarcinogenic mechanisms of action (NPL 3, NPL 4). As a result, they are used not only as a source of new medicinal uses but also as building blocks for pharmaceutical reagents (NPL 5). For this purpose, natural substances are usually chemically modified by semisynthesis, a process that combines biosynthesis and organic synthesis (NPL 6). Alternatively, novel NRPs can be produced by reprogramming the NRPS responsible for their synthesis. For this purpose, modular synthetases are often utilized and genetically engineered. Several examples of successful reprogramming of NRPSs are known in the literature (NPL 7, NPL 8). Nevertheless, the productivity of these synthetases is usually severely limited (NPL 9).
[0003] The structural and functional diversity of NRPs is the result of the introduction of D-amino acids (AS), heterocyclic elements, or N-methylated side chains, as well as the addition of lipids, sugars, and halogens (Non-Patent Document 10). Examples of NRPs with such structural specificity are, for example, bacitracin and vibriobactin, which have heterocyclic rings, or cyclosporin A and tyrocidine A, which are characterized by the introduction of D-AS. On the other hand, daptomycin is an acetylated peptide and has a potent antibacterial effect due to its fatty acid, while balhimycin and syringomycin are examples of halogenated NRPs with antibacterial and antifungal properties.
[0004] SYNZIP is a heterospecific synthetic coiled coil that allows for the control of protein interactions for use in synthetic biology. Coiled coils generally consist of two, three, or four amphipathic 20-50 AS-long α-helices that form an intertwined left-handed supercoil. Coiled coils are a structural motif in many proteins, and are present, for example, in the leucine zipper region of human bZIP transcription factors. Coiled coils follow a heptad pattern (abcdefg) with hydrophobic ASs at positions a and d, but usually electrostatic ASs at positions e and g. n In the secondary structure of the α-helix, these hydrophobic ASs interact with each other to form a narrow hydrophobic interface (Non-Patent Document 11).
[0005] SYNZIP was originally developed for heterospecific interactions with the leucine zipper domains of human bZIP transcription factors (TFs). For this purpose, 48 artificial peptides were constructed computationally and then investigated for their interactions with these peptides (Non-Patent Document 12). In further studies, interactions of peptides with each other were also examined. To this end, Reinke et al. performed a protein microarray assay, testing all 48 artificial coiled-coils, as well as the coiled-coils of seven additional human bZIPs, with each other (Figure 1). From the results of the assay, 27 pairs, 23 synthetics (i.e., SYNZIP1-SYNZIP23), and three human bZIP structures were selected that exhibited strong heterospecific interactions but low homospecific interactions. As shown in Figure 1A, these peptides participate in at least one to up to seven interactions, potentially forming various networks (Figure 1B). Examples of these networks include linear, cyclic, branched, and orthogonal networks (Non-Patent Document 13). Furthermore, it was concluded that most pairs must be parallel heterodimers based on the Asn-Asn pairing at the aa' positions (Non-Patent Document 14).
[0006] Patent Document 1 describes a system for the assembly and modification of NRPSs. This system uses novel, well-defined building blocks (units) containing condensed subdomains. This strategy allows for the efficient combination of assembly elements, called exchange units (XU2.0), regardless of the naturally occurring specificity for the subsequent NRPS adenylation domain. The system of Patent Document 1 allows for the simple assembly of NRPSs with the activity to synthesize peptides of any amino acid sequence, without the limitations imposed by naturally occurring NRPS units. This system also allows for the exchange of natural NRPS components with the inventive XU2.0, resulting in the production of modified peptides. Although this system allows for the simple combination of XU2.0 units, it requires the expression of the assembled NRPS protein within an open reading frame (ORF). This poses problems, especially for longer NRPs. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] International Publication No. 2019 / 138117 [Non-patent literature]
[0008] [Non-Patent Document 1] Caradec et al., 2014 [Non-patent document 2] Caboche et al., 2010 [Non-patent document 3] Finking and Marahiel, 2004 [Non-patent document 4] Felnagle et al., 2008 [Non-Patent Document 5] Cane et al., 1998 [Non-patent document 6] Kirschning and Hahn, 2012 [Non-Patent Document 7] Schneider et al., 1998 [Non-patent document 8] Chiocchini et al., 2006 [Non-Patent Document 9] Suo, 2005 [Non-Patent Document 10] Hur et al., 2012 [Non-Patent Document 11] Lumb et al., 1994 [Non-Patent Document 12] Grigoryan et al., 2009 [Non-Patent Document 13] Thompson et al., 2012 [Non-Patent Document 14] Reinke et al., 2010 Summary of the Invention
[0009] It is therefore an object of the present invention to efficiently and flexibly recombinantly produce NRPS modules or submodules (i.e., domains), such as, for example, the XU2.0 unit from WO 02 / 04790.
[0010] Generally, and by way of brief description, the main aspects of the present invention can be described as follows:
[0011] In a first aspect, the present invention relates to a protein or protein fragment comprising at least a first domain or partial domain (first PKS-NRPS domain) of a non-ribosomal peptide synthetase (NRPS), a polyketide synthase (PKS), or an NRPS / PKS hybrid synthase (synthetase), wherein the protein or protein fragment has an N-terminus or C-terminus that comprises a first binding domain, and this first binding domain preferably represents the N-terminus or C-terminus, respectively, of the protein or protein fragment, and wherein the first binding domain is characterized by the property that it is capable of specific protein-protein binding with at least one corresponding second binding domain.
[0012] In a second aspect, the present invention relates to an isolated nucleic acid construct comprising a first coding region having a nucleic acid sequence encoding a protein or protein fragment of the first aspect.
[0013] In a third aspect, the present invention relates to a vector system for producing a functional NRPS or PKS, the vector system comprising at least one nucleic acid construct according to the second aspect, and the at least one nucleic acid construct is suitable for expressing at least two proteins or protein fragments according to the first aspect, and the at least two proteins or protein fragments are different and together form a functional NRPS, PKS, or NRPS / PKS hybrid.
[0014] In a fourth aspect, the present invention relates to a method for producing a functional (complete) NRPS or PKS, comprising contacting at least a first protein or protein fragment according to the first aspect with a second protein or protein fragment according to the first aspect, wherein the first protein or protein fragment has a terminal first binding domain and the second protein or protein fragment has a terminal second binding domain in place of the terminal first binding domain.
[0015] Detailed Description of the Invention Elements of the present invention are described below. These elements are described in specific embodiments. However, it should be understood that elements of the present invention can be combined with each other in any manner and in any number to yield further embodiments. The various described examples and preferred embodiments should not be construed as limiting the invention to only the explicitly described embodiments or examples. The present disclosure should be understood to describe and include embodiments in which two or more explicitly described embodiments or elements are combined with each other, or in which one or more explicitly described embodiments are combined with any number of disclosed and / or preferred elements. Furthermore, all permutations and combinations of all elements described in this application should be considered to be disclosed by the detailed description of this application, unless otherwise indicated or permitted by context or technical context.
[0016] The term "partial domain" or "partial C or C / E domain" or "partial domain," or similar terms, refers to a nucleic acid sequence encoding an incomplete (non-full-length) NRPS-PKS domain or its protein sequence. In this context, this term should be understood to mean that, compared to the full-length domain, the partial domain has at least 20%, preferably 30% or 40% or more of the contiguous sequence of the full-length domain. Thus, the partial domain has a very high degree of sequence identity (90% or more) with respect to a consistent portion (at least 20%, preferably 30% or 40% or more) of the sequence of the full-length domain. For example, this expression describes a C domain sequence or a C / E domain sequence that does not include both the donor and acceptor sites of the NRPS-C domain or C / E domain. An NRPS partial domain has a sequence length of, for example, 100 or more, 150 or more, or about 200 amino acids.
[0017] "Arrangement" refers to several domains. A multi-NRPS / PKS assembly includes a complete NRPS-PKS. One or more polypeptides may contain modules. A combination of modules then catalyzes a longer peptide. In one example, a module may contain a C domain (condensation domain), an A domain (adenylation domain), and a peptidyl carrier protein domain.
[0018] Further structural information on A domains, C domains, didomains, domain-domain interfaces, and complete modules can be found in Conti et al. (1997), Sundlov et al. (2013), Samel et al. (2007), Tanovic et al. (2008), Strieker and Marahiel (2010), Mitchell et al. (2012), and Tan et al. (2015).
[0019] "Initiator module" refers to an N-terminal module that can transfer the first monomer to another module (e.g., an extender module or a terminal module). In some cases, the additional module is not the second module, but one of the modules following the C-terminus (e.g., in the case of the nocardicin NRPS). In the case of an NRPS, the initiator module contains, for example, an A (adenylation) domain and a PCP (peptidyl carrier protein) domain or a T (thiolation) domain. The initiator module may also contain the first C domain and / or an E domain (epimerization domain). In the case of a PKS, a possible initiator module consists of an AT domain (acetyltransferase) and an ACP domain (acyl carrier protein). The initiator module is preferably located at the amino terminus of the polypeptide of the first module in the "assembly series." Each assembly series preferably contains an initiator module.
[0020] The term "extension module" or "extension module" refers to a module that adds a donor monomer to an acceptor monomer or multimer, thereby extending a peptide chain. An extender module may comprise a C (condensation) domain, a Cy (heterocyclization) domain, an E domain, a C / E domain, an MT (methyltransferase) domain, an A-MT (adenylation and methylation domain combination) domain, an Ox (oxidase) domain, or a Re (reductase) domain, an A domain, or a T domain. An extender domain may further comprise an additional E domain, a Re domain, a DH (dehydration) domain, an MT domain, an NMet (N-methylation) domain, an AMT (aminotransferase) domain, or a Cy domain. Furthermore, the extension module may be of PKS origin and contain the respective domains (ketosynthase (KS), acyltransferase (AT), ketoreductase (KR), dehydratase (DH), enoylreductase (ER, thiolation (T)).
[0021] A "termination module" refers to a module that releases or cleaves a molecule (e.g., an NRP, a PK, or a combination thereof) from an assembly train. The molecule can be released, for example, by hydrolysis or cyclization. A termination module can include a TE (thioesterase) domain, a Cterm domain (terminal C domain), or an Re domain. A termination module is preferably located at the carboxy terminus of an NRPS or PKS polypeptide. A termination module can further comprise additional enzymatic activities (e.g., oligomerase activity).
[0022] By "domain" is meant a polypeptide sequence or fragment of a larger polypeptide sequence that has one or more specific enzymatic activities (i.e., a C / E domain has the functions of C and E in one domain) or another conserved function (i.e., as a binding function for an ACP domain or a T domain). Thus, a single polypeptide may contain multiple domains. Multiple domains may form a module. Examples of domains are C (condensation), Cy (heterocyclization), A (adenylation), T (thiolation), TE (thioesterase), E (epimerization), C / E (condensation / epimerization), MT (methyltransferase), Ox (oxidase), Re (reductase), KS (ketosynthase), AT (acyltransferase), KR (ketoreductase), DH (dehydratase), and ER (enoylreductase).
[0023] "Non-ribosomally synthesized peptide," "non-ribosomal peptide," or "NRP" refers to any polypeptide that is not produced by a ribosome. NRPs can be linear, cyclic, or branched and can contain proteinogenic amino acids, natural amino acids, or unnatural amino acids, or any combination thereof. NRPs include peptides produced in a sort of assembly line or train (i.e., the modular nature of an enzyme system in which building blocks can be added stepwise to form a final product).
[0024] "Polyketide" refers to a compound containing multiple ketone units.
[0025] "Nonribosomal peptide synthetase" or "nonribosomal peptide synthase" or "NRPS" refers to a polypeptide or series of interacting polypeptides that produce nonribosomal peptides and thus can catalyze peptide bond formation without a ribosomal component. "Polyketide synthase" (PKS) refers to a polypeptide or series of polypeptides that produce polyketides without a ribosomal component.
[0026] "Non-ribosomal peptide synthetase / polyketide synthase hybrid" or "hybrid of non-ribosomal peptide synthetase and polyketide synthase" or "NRPS / PKS hybrid" or "NRPS and PKS hybrid" or "PKS and NRPS hybrid" and other corresponding expressions refer to an enzyme system containing any domain or module of an NRPS and a PKS. Such hybrids catalyze the synthesis of natural hybrid substances.
[0027] By "structural change" is meant any change in chemical (eg, covalent or non-covalent) bonds compared to a reference structure.
[0028] "Mutation" refers to a change in a nucleic acid sequence such that the amino acid sequence encoded by the nucleic acid sequence has at least one amino acid change compared to the naturally occurring sequence. The mutation may be, but is not limited to, an insertion mutation, a deletion mutation, a frameshift mutation, or a missense mutation. The term also refers to the protein encoded by the mutated nucleic acid sequence.
[0029] A "variant" is a polypeptide or polynucleotide having at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence identity to a reference sequence. Sequence identity is typically measured using sequence analysis software (e.g., sequence analysis software packages from the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wisconsin 53705, USA, programs: BLAST, BESTFIT, gap, or PILEUP / PRETTYBOX). This type of software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications (substitution / scoring matrices: e.g., PAM, Blosum, GONET, JTT).
[0030] In a first aspect, the present invention relates to a protein or protein fragment comprising at least a first domain or partial domain (first PKS-NRPS domain) of a non-ribosomal peptide synthetase (NRPS), a polyketide synthase (PKS), or an NRPS / PKS hybrid synthase (synthetase), wherein the protein or protein fragment has an N-terminus or C-terminus that comprises a first binding domain, and this first binding domain preferably represents the N-terminus or C-terminus, respectively, of the protein or protein fragment, and wherein the first binding domain is characterized by the property that it is capable of specific protein-protein binding with at least one corresponding second binding domain.
[0031] A protein or protein fragment in the sense of the present disclosure is preferably a polypeptide comprising an amino acid sequence having high sequence identity to a contiguous portion of an NRPS / PKS, and one or more modules thereof, and protein portions thereof.
[0032] In the context of the present application, the term "binding domain" is intended to denote a polypeptide element, domain, or sequence capable of forming a specific or non-specific covalent or non-covalent bond with another polypeptide sequence. In a particular embodiment, the second binding domain is, for example, an endogenous NRPS / PKS sequence, which preferably allows interaction with SYNZIP according to the present invention. Binding domains that form specific binding interactions with other corresponding binding domains are preferred for the purposes of the present invention. In the context of the present invention, these are also called protein interaction domains (PIDs). For example, these may be polypeptide domains responsible for the formation of protein homodimers or heterodimers. These types of domains are, for example, coiled-coil domains, CH3 domains, and leucine zipper domains. So-called SYNZIP domains are particularly preferred.
[0033] Coiled-coil protein interaction domains are known in the art. Some non-limiting examples of computer programs that generate such PIDs include SOCKET (described, e.g., in Walshaw & Woolfson, J. Mol. Gen. Biol, 2001; 307 (5), 1427-1450, and available at the Woolfson group website at the University of Bristol), COILS (described, e.g., in Lupas et al., Science. 1991; 252: 1162-1164, incorporated herein by reference, and available at the ch.EMBnet.org website), PAIRCOIL (described, e.g., in Berger et al., Proc Natl. Acad. Sci. UNITED STATES OF AMERICA. 1995; 92, 8259-8263 and available from the group at csail.mit.edu / cb / paircoil / cgi-bin / paircoil.cgi), and MULTICOIL (described, for example, by Wolf et al., Protein Sci. 1997; 6: 1179-1189 and available from the group at csail.mit.edu / cb / multicoil / cgi-bin / multicoil.cgi).
[0034] In some embodiments, the coiled-coil-forming PIDs are those listed in Table I by Mueller et al., Methods Enzymol. 2000; 328, 261, which is incorporated herein by reference in its entirety. For example, coiled-coil-forming PIDs include leucine zippers (e.g., in the proteins GCN4, Fos, Jun, C / EBP, and variants or mutants thereof), the peptide "Velcro" (e.g., as described by O'Shea et al., Curr Biol. 1993; 3(10):658-67), E-Coil / K-Coil (e.g., as described by Tripet et al., Protein Eng. 1996; 9, 1029), and WinZip-A2 and WinZip-B1 (e.g., as described by Arndt et al., Structure. 2002; (9): 1235-48).
[0035] In some embodiments, the coiled-coil-forming PID is a heterospecific synthetic coiled-coil peptide, known as SYNZIP, such as SYNZIP 1 to SYNZIP 22. Detailed information about SYNZIP 1 to SYNZIP 22 is disclosed in Thompson KE et al., "SYNZIP protein interaction toolbox: in vitro and in vivo specifications of heterospecific coiled-coil interaction domains" (ACS Synth Biol. 2012 Apr 20;1(4):118-29), which is incorporated herein by reference in its entirety. In some embodiments, the PID fused either C- or N-terminally to the NRPS-PKS domain or subdomain is SYNZIP17 (NEKEELKSKKKAELRNRIEQLKQKREQLKQKIANLRKEIEAYK, SEQ ID NO: 1) and / or SYNZIP 18 (SIAATLENDLARLENARLEKDIANLAKLEREEAYEAYEAYEF, SEQ ID NO: 2). Other combinations of SYNZIPs that can be used as binding domain pairs in the context of the present invention are listed in the matrix of FIG.
[0036] In some embodiments, PIDs contemplated by the present disclosure include those disclosed on the website of Dr. Tony Pawson at Mount Sinai Hospital in Toronto. For example, PIDs include a 14-3-3 domain, an ADF domain, an ANK repeat, an ARM repeat, an amphiphysin bar domain, a BEACH domain, a Bcl-2 homology domain (BH) (e.g., BH1, BH2, BH3, BH4), a BIR domain, a BRCT domain, a bromodomain, a BTB / POZ domain, a CI domain, a C2 domain, a caspase recruitment domain (CARD), a lymphoid myeloid (CALM) domain associated with clathrin assembly, a calponin homology (CH) domain, a chromatin organization modifier (CHROMO / Chr) domain, a CUE domain, a death (DD) domain, a death effector (DED) domain, a DEP domain, a Dbl homology (DH) domain, an EF domain, a IFN-γ ... hand (EFh) domain, Epsl5 homology (EH) domain, epsin NH2-terminal homology (ENTH) domain, Ena / Vasp homology domain 1 (EVH1 domain), Fox-Box domain, FERM domain, FF domain, formin homology domain 2 (FH2), forkhead-associated domain (FH), FYVE (Fab-1, YGL023, Vps27, and EEA1) domain, GAT (GGA and Toml) domain, gelsolin / severin / villin homology (GEL) domain, GLUE (gram-like ubiquitin binding in EAP45) domain, GRAM (glucosyltransferases, Rab-like GTPase activators, and myotubularin) domain myotubularin (HEAT) domain, GRIP domain, glycine-tyrosine-phenylalanine (GYF) domain, HEAT (Huntington, elongation factor 3, PR65 / A, TOR) domain, HECT (homologous to the E6-AP carboxyl terminus)carboxyl-terminus domain, IQ domain, LIM domain, leucine-rich repeat domain (LRR domain), malignant brain tumor domain (MBT domain), Mad homology 1 domain (MH1 domain), MH2 domain, MIU (motif interacting with ubiquitin) domain, NZF domain (Npl4 zinc finger domain), PAS domain (Per-ARNT Sim domain), Phox and Beml domain (PB1 domain), PDZ (postsynaptic density 95 (PSD-95), disc large (Dig), zona tight junction-1 (ZO-1) occludens-1, ZO-1) domain, pleckstrin homology domain (PH domain), PoloBox domain, phosphotyrosine binding domain (PTB domain), Pumilio domain (Puf domain), PWWP domain, Phox homology domain (PX domain), RGS (regulator of G protein signaling) domain, RING finger domain, SAM (sterile alphamotive) domain, shadechromo domain (CSD domain or SC domain), Src homology 2 domain (SH2 domain), Src homology 3 domain (SH3 domain), SOCS (cytokin signaling pathway suppressors) domain, SPRY domain, START (steroidogenic acute regulatory protein (StAR) related lipid transporter) domain transfer domain, SWIRM domain, Toll / 11-1 receptor (TIR) domain, tetratricopeptide repeat (TPR) motif domain, TRAF domain, SNARE (soluble NSF binding protein receptor)These include SMAP receptor (SMAP receptor) domains (e.g., T-SNARE), Tubby domain, Tudor domain, ubiquitin-associated (UBA) domain, UEV (ubiquitin E2 variant) domain, ubiquitin-interacting motif (UIM) domain, von Hippel-Lindau tumor suppressor protein (VHLP) beta domain, VHS (Vps27p, Hrs, and STAM) domain, WD40 repeat domain, and WW domain.
[0037] The PID can be linked to the C-terminus or N-terminus of a protein or protein fragment according to the invention, with or without a linker. Of course, any PID and any linker can be compatible with aspects of the invention. In some embodiments, the linker is flexible. The linker can be composed of amino acids. In some embodiments, the linker consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more than 50 amino acids. In some embodiments, the linker consists of 5 to 7 amino acids. In some embodiments, the linker is, for example, a Gly-Ser linker.
[0038] Proteins or protein fragments according to the present invention preferably contain SYNZIP as the first and / or second binding domain, where SYNZIP is selected from SYNZIP 1 to SYNZIP 23, preferably SYNZIP 1, SYNZIP 2, SYNZIP 17, SYNZIP 18, or SYNZIP 19. In some embodiments, proteins or protein fragments are preferred that are characterized by the property that the end opposite the first binding domain contains a third binding domain, and that the third binding domain is capable of specific protein-protein binding with at least one corresponding fourth binding domain. Preferably, the first and second binding domains are incapable of binding to the third and fourth binding domains. According to the present invention, NRPS / PKS proteins consisting of three or more individual polypeptides (proteins or protein fragments according to the present invention) can be produced with such a structure.
[0039] In a preferred embodiment, the SYNZIP sequence from Thompson KE, et al. may have a truncation at the C-terminus and / or N-terminus compared to the original sequence. This truncation is ideally located at the N-terminus and / or C-terminus. That is, amino acids are removed at the N-terminus or C-terminus compared to the original sequence, with the remaining sequence remaining unchanged. In a preferred embodiment, at least 7, preferably at least 10, amino acids of the SYNZIP always remain. In the case of a SYNZIP pair used in the context of the present invention, either both or even only one SYNZIP may be present in a truncated form according to the present invention.
[0040] The truncation preferably involves 1 to 15 amino acids and can occur at the C-terminus and / or N-terminus. The shorter the SYNZIP, the closer the assembled NRPS is to the native structure, since the connection point under normal circumstances involves significantly fewer amino acids. Furthermore, the truncation is preferably 1 to 10 amino acids in length. Thus, for example, an N-terminal truncation of the SYNZIP sequence of an NRPS can involve 9 amino acids, while a C-terminal truncation of the SYNZIP sequence of an NRPS involves 2 amino acids. This is shown in Figure 16 for the SYNZIP pairs SZ1 and SZ2 or SZ19 and SZ2. However, truncation of the SYNZIP sequence must not result in the loss of the pairing properties of the SYNZIP pair.
[0041] Preferably, the first binding domain and the second binding domain can specifically bind to the third binding domain or the fourth binding domain, so that a mixture of peptides can be formed.
[0042] The term "terminus" in reference to a protein or protein fragment refers to the respective termini of an amino acid polymer. The free amino terminus is referred to as the "N-terminus" and the free carboxy terminus as the "C-terminus." When a feature, such as a sequence or domain or other element, is located at the N-terminus or C-terminus, this means that the corresponding feature constitutes the final N- or C-terminal part of the entire protein and thus represents the N- or C-terminus.
[0043] In certain embodiments, the protein or protein fragment according to the invention further comprises at least one, preferably two, three, or more, additional PKS and / or NRPS domains, where the additional PKS and / or NRPS domain(s) are arranged in direct functional configuration next to the first PKS-NRPS domain, meaning that the domains are spatially capable of carrying out NRPS / PKS synthesis of the peptide or peptide / polyketide.
[0044] In a preferred embodiment, the protein or protein fragment does not contain at least one corresponding second binding domain. In this embodiment, the second binding domain is located in a second protein or protein fragment according to the invention, and the first and second, preferably also other, proteins or protein fragments according to the invention form a group of proteins or protein fragments according to the invention. This group of proteins or protein fragments according to the invention is designed so that a functional NRPS or PKS, or hybrid thereof, can be post-translationally assembled by specific or non-specific binding via the binding domain.
[0045] In a preferred embodiment of the invention, the first binding domain is positioned at a terminus such that specific or nonspecific mediated interaction / binding to a corresponding second binding domain is possible under normal conditions. In this case, the second binding domain will be found in a second protein or protein fragment of the invention. In this case, the NRPS / PKS domain composition will be different in the first and second proteins or protein fragments of the invention.
[0046] The first PKS-NRPS domain or partial domain according to the present invention is selected from any NRPS domain and / or PKS domain known to those skilled in the art, preferably an A domain, a C domain, a C / E domain, an E domain, a C ... start In a preferred embodiment, the protein or protein fragment of the present invention comprises at least one of an A domain, a C domain, and a T domain, and preferably the protein or protein fragment has at least one NRPS-PKS, initiation module, elongation module, or termination module.
[0047] 14. A protein or protein fragment according to claim 12 or 13, wherein the third binding domain is attached to the protein or protein fragment via a linker sequence.
[0048] 10. A protein or protein fragment according to any one of the preceding claims, wherein the bond between the first binding domain and the second binding domain is a non-covalent bond.
[0049] In a second aspect, the present invention relates to an isolated nucleic acid construct comprising a first coding region having a nucleic acid sequence encoding a protein or protein fragment of the first aspect.
[0050] The term "nucleic acid" refers to naturally occurring, semi-synthetic, or fully synthetic nucleic acid molecules, as well as modified nucleic acid molecules, composed of deoxyribonucleotides and / or ribonucleotides, and / or modified nucleotides, such as "peptide nucleic acids" (PNAs), "locked nucleic acids" (LNAs), or "phosphorothioates." Other modifications of the internucleotide phosphate and ribose or sugar moieties may also be present.
[0051] The so-called "coding region" refers to the sequence elements within the nucleic acid construct of the present invention which, according to the genetic code, encode an expressible protein.
[0052] In one embodiment of the nucleic acid construct of the present invention, the first coding region is operably linked to an expression promoter. An expression promoter refers to a nucleic acid element that is necessary, and preferably sufficient, to initiate RNA transcription. These types of elements are known to those skilled in the art. These types of promoters can be selected depending on the expression system.
[0053] In a further embodiment, the nucleic acid construct of the invention comprises a second coding region having a nucleic acid sequence encoding a protein or protein fragment according to the invention, wherein the first coding region and the second coding region encode non-identical proteins or protein fragments according to the invention.
[0054] In further embodiments, the nucleic acid construct of the invention may comprise one or more additional elements for recombinantly expressing the protein or protein fragment or for controlling the intensity or duration of its expression.
[0055] In a third aspect, the present invention relates to a vector system for producing a functional NRPS or PKS, or an NRPS-PKS hybrid, the vector system comprising at least one nucleic acid construct according to the second aspect, the at least one nucleic acid construct being suitable for expressing at least two proteins or protein fragments according to the first aspect, the at least two proteins or protein fragments being different and together forming a functional NRPS, PKS, or NRPS / PKS hybrid. Thus, the at least two proteins or protein fragments can be expressed via one or two nucleic acid constructs, or, to the extent that there are three or more proteins or protein fragments according to the present invention, they can be expressed by one, two, or three nucleic acid constructs. The vector system of the present invention need only provide sufficient coding regions in their entirety to express the desired number of proteins or protein fragments according to the present invention.
[0056] A preferred embodiment of the vector system of the present invention relates to a vector system in which at least two proteins or protein fragments that can be expressed via a nucleic acid construct form a functional NRPS, PKS, or NRPS / PKS hybrid via binding between a first binding domain and / or a second binding domain, meaning that the vector system according to the present invention comprises, at least in part, expressible proteins or protein fragments that have the ability to form a functional NRPS / PKS via the binding domains.
[0057] Preferably, the functional NRPS, PKS, or NRPS / PKS hybrid is capable of synthesizing a linear peptide, a cyclic peptide, a linear polyketide, a cyclic polyketide, a linear peptide-polyketide, or a cyclic peptide-polyketide.
[0058] In a further embodiment, it is preferred that the vector system comprises a nucleic acid construct suitable for expressing at least three or more proteins or protein fragments according to the present invention, or that at least two of the three or more proteins or protein fragments together form a functional NRPS, PKS, or NRPS / PKS hybrid. Furthermore, at least three proteins or protein fragments can together form a functional NRPS, PKS, or NRPS / PKS hybrid, where the functional NRPS or PKS is formed by linking the proteins or protein fragments together via a first binding domain to a second binding domain and a third binding domain to a fourth binding domain. In this case, it is more preferred that the first protein or protein fragment has a terminal first binding domain and a terminal third binding domain, the second protein or protein fragment has a terminal second binding domain, and the third protein or protein fragment has a terminal fourth binding domain.
[0059] A preferred vector system of the present invention is designed so that the binding domains that assemble the NRPS / PKS of the present invention are positioned between the NRPS / PKS domains. Preferred positions of the binding domains, particularly SYNZIP, can be found in the Examples, and only the disclosed positions where SYNZIP is integrated are indirectly generalized.
[0060] In a fourth aspect, the present invention relates to a method for producing a functional (complete) NRPS or PKS, or an NRPS / PKS hybrid, comprising contacting at least a first protein or protein fragment according to the first aspect with a second protein or protein fragment according to the first aspect, wherein the first protein or protein fragment has a terminal first binding domain and the second protein or protein fragment has a terminal second binding domain in place of the terminal first binding domain. As a further step, the method may comprise recombinant expression of the first and / or second protein or protein fragment by at least one nucleic acid construct of the invention.
[0061] The terms "the [present] invention," "according to the invention," and similar terms used herein are intended to refer to all aspects, elements, and embodiments of the described and / or claimed invention.
[0062] As used in this disclosure, the term "comprising" is intended to be interpreted as including both "including" and "consisting of," where both meanings are specifically intended and therefore represent separately disclosed embodiments according to the present invention. The term "and / or" is understood as a specific disclosure of the two indicated features or components, each of which may or may not include the other. For example, "A and / or B" should be understood as a specific disclosure of (i) A, (ii) B, and (iii) each of A and B, as if each were separately disclosed herein. In the context of the present invention, the terms "roughly" and "approximately" indicate a range of accuracy that a person skilled in the art should understand within the scope that still ensures the technical effect of the feature in question. Where an indefinite or specific article such as "one" or "the" is used when referring to a singular noun, such use includes the plural of that noun unless otherwise stated.
[0063] It will be appreciated that application of the teachings of the present invention may be adapted to particular problems or environments, and that the inclusion of variations or additional features of the present invention (e.g., further aspects and embodiments) is within the capabilities of one of average skill in the art in light of the teachings contained herein.
[0064] Unless the context requires otherwise, the above feature descriptions and definitions are not limited to any particular aspect or embodiment of the invention, but apply equally to all aspects and embodiments described.
[0065] All references, patents, and publications cited herein are hereby incorporated by reference in their entirety.
[0066] In view of the above, it should be noted that the present invention also relates to the following detailed numbered subject matter:
[0067] Subject 1: A protein or protein fragment comprising at least a first domain or partial domain (first PKS-NRPS domain) of a non-ribosomal peptide synthetase (NRPS), a polyketide synthase (PKS), or an NRPS / PKS hybrid synthase (synthetase), wherein the protein or protein fragment has an N-terminus or C-terminus comprising a first binding domain, and this first binding domain preferably represents the N-terminus or C-terminus of the protein or protein fragment, respectively, and is characterized by the property that the first binding domain is capable of specific protein-protein binding with at least one corresponding second binding domain.
[0068] Subject 2: The protein or protein fragment of Subject 1, further comprising at least one, preferably two, three, or four or more additional PKS-NRPS domains, wherein the additional PKS-NRPS domain(s) are positioned in direct functional arrangement next to the first PKS-NRPS domain.
[0069] Aspect 3: The protein or protein fragment of Aspect 1 or 2, wherein the protein or protein fragment does not comprise at least one corresponding second binding domain.
[0070] Subject 4: A protein or protein fragment described in any one of subjects 1 to 3, wherein the first binding domain is positioned at its terminus so that specific or non-specific mediated interaction / binding to the corresponding second binding domain is possible under normal conditions.
[0071] Theme 5: The first PKS-NRPS domain or partial domain is an A domain, an A-MT domain, a C domain, a C / E domain, an E domain, a C start 5. The protein or protein fragment according to any one of claims 1 to 4, wherein the protein or protein fragment is selected from the group consisting of a FT domain, a FT domain, and a T domain.
[0072] Aspect 6: A protein or protein fragment according to any one of Aspects 1 to 5, comprising at least one A domain or A-MT domain, a C domain and / or an E domain or a C / E domain or a Cy domain, and a T domain, wherein the protein or protein fragment preferably has at least one NRPS-PKS extension module.
[0073] Aspect 7: The protein or protein fragment of any one of aspects 1 to 6, wherein the binding domain is a protein sequence, preferably a protein domain, that mediates specific protein-protein binding.
[0074] Aspect 8: The protein or protein fragment of any one of Aspects 1 to 7, wherein the binding domain comprises a coiled-coil domain.
[0075] Aspect 9: The protein or protein fragment of aspect 8, wherein the binding domain comprises a synthetic coiled-coil domain (SYNZIP).
[0076] Aspect 10: The protein or protein fragment according to aspect 9, wherein SYNZIP is selected from SYNZIP 1 to SYNZIP 23, preferably SYNZIP 1, SYNZIP 2, SYNZIP 17, SYNZIP 18, or SYNZIP 19. The protein or protein fragment is characterized by the property that it comprises a third binding domain opposite the first binding domain, wherein the third binding domain is capable of specific protein-protein binding with at least one corresponding fourth binding domain.
[0077] Aspect 11: The protein or protein fragment of aspect 10, wherein the first binding domain and the second binding domain are incapable of binding to the third binding domain and the fourth binding domain.
[0078] Aspect 12: The protein or protein fragment of aspect 11, wherein the first binding domain and the second binding domain can specifically bind to the third binding domain or the fourth binding domain to form a mixture of peptides.
[0079] Aspect 13: The protein or protein fragment of any one of aspects 1 to 12, wherein the first binding domain is linked to the first PKS-NRPS domain by a linker sequence.
[0080] Aspect 14: The protein or protein fragment of aspect 12 or 13, wherein the third binding domain is attached to the protein or protein fragment via a linker sequence.
[0081] Aspect 15: A protein or protein fragment according to any one of the preceding aspects, wherein the bond between the first binding domain and the second binding domain is a non-covalent bond.
[0082] Subject 16: An isolated nucleic acid construct comprising a first coding region having a nucleic acid sequence encoding the protein or protein fragment of any one of subjects 1 to 15.
[0083] Subject 17: The isolated nucleic acid construct of subject 16, wherein the first coding region is operably linked to an expression promoter.
[0084] Subject matter 18: The isolated nucleic acid construct of Subject matter 16 or 17, further comprising a second coding region having a nucleic acid sequence encoding the protein or protein fragment of any one of Subject matters 1-15, wherein the first coding region and the second coding region encode non-identical proteins or protein fragments.
[0085] Subject matter 19: The isolated nucleic acid construct of any one of subjects 1 to 18, further comprising elements for recombinant expression of a protein or protein fragment.
[0086] Subject matter 20: A vector system for producing a functional NRPS or PKS, the vector system comprising at least one nucleic acid construct according to any one of subjects 16-19, and the at least one nucleic acid construct is suitable for expressing at least two proteins or protein fragments according to any one of subjects 1-15, and the at least two proteins or protein fragments are different and together form a functional NRPS, PKS, or NRPS / PKS hybrid.
[0087] Aspect 21: The vector system of aspect 20, wherein at least two proteins or protein fragments form a functional NRPS, PKS, or NRPS / PKS hybrid by binding of the first binding domain and the second binding domain.
[0088] Aspect 22: The vector system of Aspect 20 or 21, wherein the functional NRPS, PKS, or NRPS / PKS hybrid is capable of synthesizing a linear peptide, a cyclic peptide, a linear polyketide, a cyclic polyketide, a linear peptide polyketide, or a cyclic peptide polyketide.
[0089] Aspect 23: The vector system of any one of Aspects 20-22, wherein the vector system comprises a nucleic acid construct suitable for expression of at least three or more proteins or protein fragments of any one of Aspects 1-15, or at least two of the three or more proteins or protein fragments together form a functional NRPS, PKS, or NRPS / PKS hybrid.
[0090] Aspect 24: The vector system of aspect 23, wherein at least three proteins or protein fragments can be combined to form a functional NRPS, PKS, or NRPS / PKS hybrid, wherein the functional NRPS or PKS is formed by linking the proteins or protein fragments to each other by binding of the first binding domain to the second binding domain and the third binding domain to the fourth binding domain.
[0091] Aspect 25: The vector system of Aspect 23 or 24, wherein the first protein or protein fragment has a first binding domain at its end and a third binding domain at its opposite end, the second protein or protein fragment has a second binding domain at its end, and the third protein or protein fragment has a fourth binding domain at its end.
[0092] Aspect 26: A method for producing a functional (complete) NRPS or PKS, comprising contacting at least a first protein or protein fragment described in any one of Aspects 1 to 15 with a second protein or protein fragment described in any one of Aspects 1 to 15, wherein the first protein or protein fragment has a terminal first binding domain, and the second protein or protein fragment has a terminal second binding domain in place of the terminal first binding domain.
[0093] Aspect 27: The method of Aspect 26, wherein said contacting comprises recombinant expression of said first and / or second protein or protein fragment by at least one nucleic acid construct of any one of Aspects 16-19. [Brief explanation of the drawings]
[0094] [Figure 1]Figure 1. SYNZIP interaction partners and possible networks. A) Protein microarray assay results for 26 peptides that form specific interaction pairs. Peptides immobilized on the microarray surface are shown in columns. Peptides fluorescently labeled in solution are listed in rows. According to the array score (shown on the right), black spots indicate strong fluorescent signals (0-0.2), while white spots indicate weak fluorescent signals (>1.0). The absence of homospecific interactions is indicated by a red diagonal line. Interactions with an array score of 0.2 or less are highlighted in green. The number of strong interacting partners is shown at the bottom (Non-Patent Document 14). B) Possible SYNZIP interaction networks with corresponding SYNZIP numbers are shown: 1. Linear network, 2. Circular network, 3. Branched network, and 4. Orthogonal network. Dashed lines indicate weak interactions, and solid lines indicate strong interactions. The asterisk highlights the antiparallel interaction between SYNZIP 17 and SYNZIP 18 (Non-Patent Document 13). [Figure 2] Figure 1 shows the construction of AmbS hybrids that produce novel peptides. A: Schematic of NRPS hybrids (NRPS-3a and NRPS-3b) from XU of AmbS (black) and GxpS (red). The relative peptide production of peptide 7, peptide 8, and peptide 9 from triplicate measurements is shown in %. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18. B: Structures of the produced peptides. [Figure 3]Figure 1 shows the construction of SzeS hybrids that produce novel peptides. A: Schematic of NRPS hybrids (NRPS-4a and NRPS-4b) and covalently linked hybrids (NRPS-4c) from XU of SzeS (green) and GxpS (red). The relative peptide production of peptide 10 and peptide 11 from triplicate measurements is shown in %. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain or FT domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZ: orange, SZ17; green, SZ18. B: Structures of produced peptides. [Figure 4] Construction of XldS hybrids producing novel peptides. A: Schematic of NRPS hybrids (NRPS-5a and NRPS-5b) from XU of XldS (turquoise) and GxpS (red). The relative peptide production of peptide 12, peptide 13, peptide 14, and peptide 15 from triplicate measurements is shown in %. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18. B: Structures of the produced peptides. [Figure 5] Figure 1 shows proof-of-concept of various interfaces and SZ oligomerization contexts based on XtpS. Schematic diagram of XtpS (light green) split at TC (NRPS-13), AT (NRPS-14), and CA (NRPS-15 and NRPS-16), as well as constructs with different SZ oligomerization contexts (NRPS-15 and NRPS-16). WT-XtpS (NRPS-1) was used as a reference. Relative production of peptide 1 and peptide 2 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; yellow, SZ19. [Figure 6]Figure 1 shows the effect of SZ on the production of AT-cleaved XtpS. Schematic diagram of three control experiments: constructs without an N-terminal SZ (NRPS-14b), without an N-terminal SZ (NRPS-14c), and without both SZs (NRPS-14d), as well as a diagram of a construct with both SZs (NRPS-14a). WT-XtpS (NRPS-1) was used as a reference. Relative production of peptide 1 and peptide 2 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18. [Figure 7] Figure 1 shows the effect of SZ on the production of XtpS split at CA (SZ19 / 18). Schematic diagram of three control experiments: one without the N-terminal SZ (NRPS-16b), one without the C-terminal SZ (NRPS-16c), and one without both SZs (NRPS-16d), as well as a diagram of the construct with both SZs (NRPS-16a). WT-XtpS (NRPS-1) was used as a reference. Relative production of peptide 1 and peptide 2 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: yellow, SZ19; green, SZ18. [Figure 8]Figure 1 shows the effect of GS linkers on the production of XtpS split at CA (SZ17 / 18). Schematic diagrams of constructs without a GS linker (NRPS-15a) and with a 10-AS-long GS linker (NRPS-15b), an 8-AS-long GS linker (NRPS-15c), and a 4-AS-long GS linker (NRPS-15d) introduced between the C-terminus of the first XtpS moiety and SZ17. WT-XtpS (NRPS-1) was used as a reference. Relative production of peptide 1 and peptide 2 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18. [Figure 9] Figure 1. Productivity of tripartite XtpS. Schematic diagram of XtpS split at the T-linker (NRPS-17a) and the A-T-linker (NRPS-18a), as well as the corresponding negative control (NRPS-18b). WT-XtpS (NRPS-1) was used as a reference. Relative production of peptide 1 and peptide 2 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZ: orange, SZ17; green, SZ18; dark blue: SZ1; light blue: SZ2. [Figure 10] Figure 1 shows the productivity of tripartite GxpS. A: Schematic of GxpS (NRPS-20) split at the AT linker. WT-GxpS (NRPS-2) was used as a reference. The relative production of peptides 3, 4, 5, and 6 from triplicate measurements is shown as % of WT levels. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; dark blue: SZ1; light blue: SZ2. B: Structures of the produced peptides. [Figure 11]Figure 1. Reprogramming of XtpS for the production of novel peptides. A: Schematic of hybrid NRPS-23b and NRPS-23c, generated by replacing the XtpS (light green) tridomain with the GxpS (red) and SzeS (green) tridomains. The relative production of peptides 16a / b, 17a, 18, and 19a / b from triplicate assays is shown as a normalized percentage compared to WT (NRPS-1). The triplicate XtpS (NRPS-18) is also shown. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; dark blue, SZ1; light blue, SZ2. B: Structures of the produced peptides. [Figure 12] Figure 1. Reprogramming of GxpS for the production of novel peptides. A: Schematic of hybrid NRPS-24a and NRPS-24c, generated by replacing the GxpS (red) tridomain with the XtpS (light green) and SzeS (green) tridomains. The relative production of peptides 20, 21a / b, 22, 23a / b, 24, 25, 3, and 4 from triplicate assays is shown as a percentage compared to WT (NRPS-2). A triplicate GxpS (NRPS-20) is also shown. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; diamond, C / E domain; small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; dark blue, SZ1; light blue, SZ2. B: Structures of the produced peptides. [Figure 13]Design of XtpS hybrids for the production of novel peptides. A: Schematic of hybrids NRPS-26 and NRPS-27, made from portions of GxpS (red) and SzeS (green) and XtpS (light green), respectively. The relative percentages of peptide 20, peptide 22, and peptide 26 from triplicate measurements are shown. Also shown is a triplicate of XtpS (NRPS-18). Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C or FT domain; diamond, C / E domain; and small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; dark blue, SZ1; light blue, SZ2. B: Structures of the produced peptides. [Figure 14] Design of GxpS hybrids for the production of novel peptides. A: Schematic of hybrids NRPS-28 and NRPS-29, respectively, made from portions of XtpS (light green), SzeS (green), and GxpS (red). The relative peptide production of peptides 16a / b, 27, 17a / b, 28a / b, 18, 19, 29, 30, 31a / b, and 32 from triplicate measurements is shown in %. The triplicate GxpS (NRPS-20) is also shown. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain or FT domain; diamond, C / E domain; and small circle at the C-terminus, TE domain. Helices represent SZs: orange, SZ17; green, SZ18; dark blue, SZ1; light blue, SZ2. B: Structures of the produced peptides. [Figure 15] Figure 1 shows the division of a hybrid of NRPS and PKS modules. The substance produced by this hybrid is glidobactin A (see structure). The NRPS GlbD was divided by SZs between the A and T domains. Symbols represent domains: circle, A domain; rectangle, T domain; triangle, C domain; PKS domains in GlbB are named according to their function. Helices represent SZs: light gray, SZ17; dark gray, SZ18. [Figure 16]Figure 1 shows preferred embodiments in which the sequences of SYNZIP variants SZ1 and SZ2 (A) or SZ2 and SZ19 (B) are truncated at the N-terminus, respectively. The truncated but still fully functional SYNZIP results in improved peptide production in NRPSs. DETAILED DESCRIPTION OF THE INVENTION
[0095] The sequence shows: SEQ ID NO: 1 and SEQ ID NO: 2: SYNZIP sequences SEQ ID NO: 3 and SEQ ID NO: 5: Preferred sequence motifs for inserting binding domains according to the invention SEQ ID NO: 6 to SEQ ID NO: 30: Peptide sequences of NRPS peptides produced in this application [Example]
[0096] Certain aspects and embodiments of the present invention will now be described, by way of example, with reference to the detailed description, figures, and tables set forth herein. Such examples of materials, methods, uses, and other aspects of the invention are representative only and should not be understood as limiting.
[0097] Examples include:
[0098] Example 1: De novo design of NRPS via SYNZIP based on the XU concept This paper first addressed the de novo construction of NRPSs based on the XU concept and the production of novel peptides. By introducing the SYNZIP pair 17 / 18 into the conserved WNATE motif of the CA linker, a hybrid NRP was constructed from the two systems. The antiparallel SZ pair should act as a noncovalent intermediary between the various synthetases. With a dissociation constant (Kd) of less than 10 nM, SZ17 and SZ18 have strong affinity for each other (Thompson et al., 2017), presenting almost all the characteristics of a covalent linkage. An NRPS hybrid was created by combining the first two XUs of AmbS, SzeS, and XldS with the last three XUs of GxpS. SZ17 was added to the C-terminus of AmbS, SzeS, and XldS, and SZ18 was added to the N-terminus of GxpS. Regarding the rule established by Bozhueyuek et al. that requires consideration of the specificity of the C domain, specificity was observed for the first two hybrids (AmbS-GxpS and Sze-GxpS), but not for the last hybrid (XldS-GxpS).
[0099] Example 2: Plasmid construction and heterologous expression of GxpS hybrid in E. coli DH10B::mtaA Using gDNA from X. miraniensis DSM 17902, X. szentirmaii DSM 16338, and X. indica DSM 17382, we first amplified the first two XUs, AmbS (A1-C3), SzeS (C1-C3), and XldS (C1-C3). For this purpose, we used the primers listed in Table 1. These contained overhangs compatible with the pACYC_ara_araE vector, which already contained the sequence of SZ17. After vector linearization, plasmids pJW91 (ambS_A1-C3_SZ17), pJW92 (szeS_C1-C3_SZ17), and pJW93 (xldS_C1-C3_SZ17) were cloned from the plasmid backbone and insert by hot fusion (see 2.3.7). After screening and verification of the plasmids, these plasmids were transformed into E. coli DH10B::mtaA together with additional plasmids pJW76 (SZ18_gxpS_A3-TE) or pJW83 (gxpS_A3-TE), respectively. In this case, pJW76 contained the last three XU sequences of GxpS and the SZ18 sequence. In contrast, transformation of pJW83, which lacked the SZ18 sequence, served as a negative control. Protein production was carried out in triplicate at 22°C for 72 hours by induction with L(+)-arabinose.
[0100] In the subsequent analysis by HPLC-MS (see 2.5.2), the masses of the peptides obtained from the hybrid system were investigated. Thus, in the case of hybrid NRPS-3a composed of parts of AmbS and GxpS, the mass of 7 (linear peptide) was 607.23 [M+H + ] and 589.33 [M+H for 8 (cyclic peptide) + These masses could be calculated from the peptide sequence (sQflL). Due to the promiscuity of the A domain of the third GxpS, which can incorporate phenylalanine in addition to leucine, the m / z value of 573.36 [M+H +] (linear peptide) and 555.35 [M+H + The m / z values of 9 (cyclic peptide) were also searched for. In this case, the mass was derived from the sequence (sQllL). Peptides 7, 8, and 9, which eluted at retention times of 6, 7.1, and 7 minutes, could be identified based on their masses. On the other hand, the linear peptide with the sequence (sQllL) could not be detected. Because standards were not available during data acquisition, the results were quantified relatively and calculated from the average peak area (Figure 2). This indicated that 8 was the most frequently detected peptide. 7 and 9 were produced at 8.1% and 21.2% of 8, respectively. All peptides were identified by MS analysis. 2 This could be verified based on the spectra (Appendix Figure 2). Furthermore, measurement data from the negative control (NRPS-3b) showed that production of 7, 8, and 9 was possible even without the N-terminal SZ (Figure 2). However, peptide production was significantly lower than in the control system (NRPS-3a) with both SZs. Thus, only about 50% of 7 and 8 were produced, and only about 18% of peptide 9 was produced.
[0101] In the case of the second hybrid NRPS-4a, which was composed of the first two XUs of SzeS and the last three XUs of GxpS, the HPLC-MS data recorded showed a 634.38 [M+H] peak for the phenylalanine derivative 10 (formyl-1TflL). + ] and 601.36 [M+H for leucine derivative 11 (formyl-1TllL) +The m / z values of [peptides 10 and 11] were searched for. A construct lacking the N-terminal SZ (NRPS-4b) and a covalently linked system (NRPS-4c) were used as comparison systems. In the extracted ion chromatograms (EICs) of the measured data, both the masses of peptides 10 and 11, eluting at retention times of 8 and 7.8 min, respectively, were detectable. Peptide 10 was the most frequently measured peptide, while peptide 11 was measured at a relative value of 8% (Figure 3). Furthermore, the measured data for the negative control (NRPS-4b) showed a significant reduction of 10 and 11 by approximately 80%. The covalently linked construct (NRPS-4c) also produced significantly lower amounts of these peptides. Thus, compared to NRPS-4a, only approximately 60% of 10 and approximately 30% of 11 were detected.
[0102] Finally, HPLC-MS data for the third hybrid (NRPS-5a) composed of portions of XldS and GxpS showed the production of the most promising derivative (Figure 4). The C1 domain of XldS allows for the incorporation of C13, C14, or C15 FS at the N-terminus of the peptide, while the promiscuity of the A domain of the third GxpS resulted in a peak of 830.54 [M+H]. + ] 12(13:0-qNflL), 844.55 [M+H + ] 13(14:0-qNflL), 858.12 [M+H + ] 14(15:0-qNflL), 796.55 [M+H + ](13:0-qNllL), 810.57 [M+H + ] 15(14:0-qNllL), and 824.59[M+H +Six possible derivatives with m / z values of [15:0-qNllL] were obtained. Four of these were detected. The retention times of the C13, C14, and C15 derivatives were 11.3, 11.8, and 12.3 min, respectively. 13 was the most frequently produced peptide, whereas the remaining peptides were detected at relative amounts ranging from 2.3% (12) to 14.3% (15) (Figure 4). Furthermore, the low signal intensity of the EICs indicated low overall production. Furthermore, the negative control (NRPS-5b) did not show significant differences in peptide production from the NRPS-5a construct (Figure 4).
[0103] Example 3: Strategy for reconstructing NRPS via SYNZIP In the above, the conserved WNATE motif (SEQ ID NO: 3) in the CA linker was selected as a preferred cleavage site based on the XU concept. This cleavage site was hypothesized by Bozhueyuek et al. from sequence alignments of NRPS linker regions from Photorhabdus and Xenorhabdus, as well as from published NRPS structural data from other organisms, as an ideal fusion point. It was also noted that the AT and TC linker regions are less suitable for NRPS reprogramming because they have less conserved sequences compared to the CA linker. Nevertheless, if the SZ pair 17 / 18 is introduced into the TC and TA linker regions, comparative studies should be performed to determine whether these insertion sites are also suitable fusion points. This hypothesis was confirmed using the XtpS model system. To this end, we aligned XtpS with the structural data of bacillibactin synthetase from Bacillus subtilis published in 2017 (Tarry et al., 2017) to obtain conclusions about its possible secondary structure. Subsequently, based on this, we defined cleavage sites in the TC and AT linker regions. These cleavage sites were ultimately related to the sequence motifs RV|LP (SEQ ID NO: 4) in the TC linker and VY|AAP (SEQ ID NO: 5; vertical lines indicate the cleavage sites) in the AT linker.
[0104] Furthermore, two different SYNZIP oligomerization states should be compared, meaning that the XtpS subunits bound to the SZ are either spatially close or further apart. In principle, both parallel and antiparallel SZ orientations can be implemented. However, only the conformation in which the proteins are further apart is useful. This is because, after splitting the NRPS system into two, SZ can only be bound to two possible ends of the protein (the C-terminus of the first NRPS part and the N-terminus of the second NRPS part) rather than four (the N- and C-terminus of the first protein part and the N- and C-terminus of the second protein part). However, by introducing SZ19 and the functional reverse form of SZ18, a similar orientation seems possible. Nevertheless, this pair was not characterized by Thompson et al. (SZ19 and the forward form of SZ18), and therefore data on Kd values, interaction partners, etc. are lacking. Finally, the antiparallel SZ pair 17 / 18 was used for the wide conformation and the parallel SZ pair 19 / 18 for the close conformation.
[0105] In the case of XtpS fragments split at the TC and AT linkers (NRPS-13 and NRPS-14, see 3.2 for cleavage sites), two plasmids were assembled in each case: pNA2 (xtpS_A1-T2_SZ17) and pNA3 (SZ18_xtpS_C3-TE), and pNA4 (xtpS_A1-A2_SZ17) and pNA5 (SZ18_xtpS_T2-TE), and these were transformed together into E. coli DH10B::mtaA. In contrast, to represent the cleavage site at the CA linker, plasmids pJW61 (xtpS_A1-C3_SZ17) and pJW62 (SZ18_xtpS_A3-TE) were used. Furthermore, in the case of NRPS-16 split at the CA linker, SZ17 at the C-terminus was replaced by SZ19, while SZ18 at the reverse N-terminus was left unchanged. Production by wild-type XtpS (NRPS-1) was used as a reference. Triplicate production cultures of all constructs were prepared simultaneously, and the synthesis of 1 and 2 was examined by HPLC-MS.
[0106] Because absolute quantification was not possible due to the lack of standards, a relative evaluation of peak areas was performed (Figure 5). Unlike the relative values in Figure 5, linear peptide 1 is not formed in substantially the same amount as cyclic peptide 2, based on absolute peptide yield, but only at approximately 0.1%. This result is simply due to better ionization of the linear peptide, and this must be taken into account when considering relative values. Measurements revealed that for all constructs, 411.31 [M+H + ] and in most cases also 429.31 [M+H +It was also demonstrated that production of linear peptide 1 with an m / z value of [ ] could be demonstrated (Figure 5 ). 2 was best produced by the constructs split at the TC linker (NRPS-13) and the AT linker (NRPS-14) at approximately 80% relative to WT, whereas the two constructs split at the CA, NRPS-15 and NRPS-16, showed significantly lower production at 27% (NRPS-15) and 13% (NRPS-16). NRPS-14 instead of NRPS-13 produced negligible amounts of linear peptide 1, and NRPS-16 showed no production of 1.
[0107] Example 4: Effect of SYNZIP on the production of AT-split XtpS The effect of SZ on the production of 1 and 2 was examined for the AT-split construct NRPS-14a (Figure 6). Control experiments were performed in which, after assembly of pNA11(xtpS_A1-A2) and pNA12(xtpS_T2-TE), the N-terminal SZ (NRPS-14b) was omitted in the first case, the C-terminal SZ (NRPS-14c) in the second case, and both SZs (NRPS-14c) in the third case (Figure 6). For this purpose, we cloned plasmids pNA12(xtpS_T2-TE) and pNA11(XtpS_A1-A2) lacking the sequences for SZ17 and SZ18, respectively. The HPLC-MS data are shown in Figure 6. Negative control AT-split constructs (NRPS-14b, NRPS-14c, and NRPS-14d) showed very low production of cyclic product 2 at approximately 3%–10% of WT levels, and no linear peptide 1 at all. Overall, the controls showed a greater than 90% reduction in productivity compared to NRPS-14a (Figure 6). Furthermore, NRPS-14a produced peptide 2 and peptide 1 at 104% and 81% of WT levels, respectively.
[0108] Example 5: Effect of SYNZIP on the production of XtpS split in AC (SZ19 / 18) The same control experiments were also performed for the AC-split construct with the SZ pair 19 / 18. Because this construct already showed low production of both SZs (Figure 5), the three control experiments performed showed very low or no detectable production of 2 and 1 (Figure 7). Relative analysis of the HPLC-MS data showed no peptide production for the control experiments NRPS-16b and NRPS-16d, and very low production of 2, at 3.2% of WT levels, for NRPS-16c. Compared to the construct with both SZs (NRPS-16a), the control NRPS-16c showed an 80% reduction in production of cyclic peptide 2.
[0109] Example 6: Effect of GS linker on the production of XtpS split in AC (SZ17 / 18) The XtpS construct split at CA (NRPS-15, Figure 5) showed significantly lower productivity compared to the XtpS constructs split at TC (NRPS-13, Figure 5) and AT (NRPS-14, Figure 5). Therefore, to improve productivity, we introduced GS linkers of various lengths between the C-terminus of the first XtpS fragment and SZ17. According to the XU concept published by Bozhueyuek et al., 10 ASs of the conserved WNATE motif were deleted for the construction of reprogrammed NRPSs. The same procedure was performed for the CA construct (NRPS-15) shown in Figure 8. Therefore, we set out to introduce a GS linker with a length of 10 ASs. This was achieved by assembling the plasmid pNA8(xtpS_A1-C3_GS(10)_SZ17), a plasmid derived from pJW61(A1-C3_SZ17). Furthermore, two additional plasmids, pNA9 (xtpS_A1-C3_GS(8)_SZ17) and pNA10 (xtpS_A1-C3_GS(4)_SZ17), encoding GS linkers with lengths of eight and four AS, respectively, were constructed. Evaluation of HPLC-MS data indicated that all constructs with GS linkers (NRPS-15b, NRPS-15c, and NRPS-15d) produced better results than constructs without GS linkers. Overall, the introduction of the linkers resulted in an average productivity increase of approximately 37% for cyclic 2 and approximately 26% for linear peptide 1. Furthermore, all constructs with GS linkers produced cyclic product 2 at near-WT levels.
[0110] Example 7: Proof of concept: Productivity of the three-part XtpS system Having successfully split XtpS into two at each of the above positions, the next step was to split the system into two positions. To this end, we introduced an additional SYNZIP pair, SZ1 and SZ2, which are not linked to SZ17 and SZ18 and thus form a so-called orthogonal network. With Kd values of less than 10 nM, SZ1 and SZ2, as well as SZ17 and SZ18, show very strong affinity for each other. Furthermore, we selected only the AT and TC linker regions as positions for the three-way split. These proved to be the most favorable positions with the best production by NRPS-13 and NRPS-14 (Figure 5). Therefore, we created two split constructs, NRPS-17a and NRPS-18a, by introducing two SZ pairs into the linker regions TC and AT, respectively. Specifically, SZ was introduced into the second and third TC linkers or the second and third AT linkers. A total of four additional plasmids were cloned: pNA17 (SZ18_xtpS_C3-T3_SZ1) and pNA18 (SZ2_xtpS_C4-TE) for TC division, and pNA15 (SZ18_xtpS_T2-A3_SZ1) and pNA16 (xtpS_SZ2_T3-TE) for AT division. Additionally, plasmids lacking SZ were assembled as negative controls. This resulted in pNA19 (xtpS_T2-A3) and pNA20 (xtpS_T3-TE) for AT division (NRPS-18b, Figure 9).
[0111] In the case of both the NRPS-17a and NRPS-18a tripartite constructs, production of the linear 1 and cyclic 2 peptides was confirmed (Figure 9). Compared to NRPS-18a, NRPS-17a produced peptides at twice the amount (2) and more than three times the amount (1). Thus, 2 was confirmed at 71.7% and 32.2%, respectively, and 1 at 25.6% and 7.3%, respectively. Furthermore, the negative control (NRPS-18b) of the AT-split system showed no peptide production.
[0112] Example 8: Construction of hybrid NRPs by SYNZIP-mediated tridomain exchange In addition to XtpS, GxpS (NRPS-20) was also split into three parts (Figure 10) to create the tridomain parts of XldS and SzeS. The resulting tridomain parts of the systems should then be combined with each other in further experiments to produce novel peptides. Sequence alignment of the AT linker and TC linker regions of all four systems revealed that the cleavage site of the AT linker (see 3.2, slight variations in sequence motifs within and between systems) was the more favorable cleavage site (the more conserved cleavage site sequence motif). Therefore, only the interface at the AT linker was used in all other constructs shown. This means that the substrate specificity of the upstream C domain should be taken into account, rather than the substrate specificity of the downstream C domain as in the XU concept. In addition to XtpS, NRPS-23 also consists of parts of GxpS and SzeS (Figure 11). After confirming the productivity of the three parts NRPS-18, NRPS-23b, and NRPS-23c, the tridomains were swapped.
[0113] Example 9: Productivity of a further system divided into three parts To construct the three-part GxpS (NRPS-18) and SzeS (NRPS-19) systems, a set of three plasmids was cloned in each case, yielding plasmids pNA26 (gxpS_A1-A2_SZ17), pNA27 (SZ18_gxpS_T2-A3_SZ1), and pNA28 (SZ2_gxpS_T3-TE), as well as plasmids pNA29 (szeS_C1-A2_SZ17), pNA30 (szeS_T2-A3_SZ1), and pNA31 (SZ2_szeS_T3-TE), which were then transformed together into E. coli DH10B::mtaA.
[0114] In the case of NRPS-20 (3 parts GxpS), the values were 586.40 [M+H +], 600.41 [M+H + ], 552.41 [M+H + ], and 566.43 [M+H + All four derivatives, 3, 4, 5, and 6, with m / z values of [ ], could be measured (Figure 10). However, compared with WT-NRPS (NRPS-2), the productivity was significantly reduced. For example, only 5.2% of 3, 10.8% of 4, and approximately 14% of 5 and 6 were produced (Figure 10).
[0115] Example 10: Production of novel peptides by tridomain swapping The division of the above NRPSs into three plasmids simplifies the operation of the system. In experimental implementation, one or two plasmids are substituted for the original plasmid and transformed together into the respective expression strains in the new configuration. The post-translational linkage between the various NRPSs is then mediated by an artificial leucine zipper. Thus, for example, a new hybrid system could be constructed by replacing the second plasmid of the XtpS set with the second plasmid of the GxpS set. Overall, the plasmids constructed in this paper allow the construction of 50 hybrid synthetases, eight of which are described in the following examples.
[0116] The first tridomain exchange was aimed at allowing the substitution of the second valine in Xtp with a phenylalanine. In an experimental implementation, instead of the pNA15 plasmid, the plasmids pNA27 and pNA30 were transformed together with pNA4 and pNA16 into E. coli DH10B::mtaA, thereby enabling the generation of hybrid NRPS-23b and hybrid NRPS-23c (Figure 11).
[0117] Comparative evaluation of HPLC-MS analysis revealed the expected peptides for NRPS-23b and NRPS-23c (Figure 11). Thus, NRPS-23b detected both phenylalanine and leucine derivatives in linear forms 16 and 17 and in cyclic forms 18 and 19. The EICs of peptides 16 and 19 also showed double peaks with different retention times, but identical fragmentation in the MS2 spectra. From this, it was concluded that these peptides existed as stereoisomers and eluted at different times accordingly. Because a non-native protein-protein interface exists only for the last C / E domain of hybrid NRPS-23b, the upstream AS was presumed to exist in two conformations. In the cases of 16 and 19, these were the phenylalanine and leucine ASs, respectively. However, which AS is actually affected remains unexplored and therefore unresolved. Furthermore, peptides 16 and 18 were produced by NRPS-23c, as were peptides already produced by NRPS-23b. Because a double peak was again present for peptide 16, stereoisomers 16a and 16b were assumed. Peak areas were added for relative assessment of the isomers.
[0118] The most frequently detected peptides were linear peptides 16a / b (in both hybrids) and 17, all of which were produced in essentially equal amounts, ranging from 94.4% to 100% (Figure 11). The effect of ionization on the frequency of linear peptides is discussed in Section 4. Overall, both NRPS-23b and NRPS-23c hybrids showed similar production of peptides 16a / b and 18, with 16a / b produced at 94.4% and 100%, respectively, and 18 produced at 19.1% and 24.8%, respectively. Furthermore, very similar values were measured for the phenylalanine derivatives 16a / b and 18 and the leucine derivatives 17a and 19a / b produced by NRPS-23b.
[0119] In the second tridomain exchange, substitution of phenylalanine in Gxps with valine (from Xtp) or phenylalanine (from Sze) should be performed. Hybrid NRPS-24a and hybrid NRPS-24c were generated by replacing pNA27 with pNA15 and pNA30, respectively (Figure 12).
[0120] Relative analysis of the measured data demonstrated that peptide production could be measured for NRPS-22a and NRPS-22c. Both hybrids produced valine derivatives 20, 22, 24, and 3 and leucine derivatives 21a / b, 23, 25, and 4 in both linear and cyclic forms (Figure 12). The most commonly produced peptide for hybrid NRPS-24a was the cyclic valine derivative 22. The linear form 20 was detected at 34.7% relative to 22. Furthermore, the leucine derivative was produced in cyclic form 23 at 62.9%, less than 22. The linear peptide 21 was present at a relative value of 20.9% and was detected as the stereoisomers 21a and 21b. NRPS-24c showed overall lower production compared to NRPS-24a. Thus, the cyclic valine derivative 3 was produced at 66.1% of the level of 22, whereas the cyclic leucine derivative 4 was detected at only 16.3% of the level. Furthermore, the linear peptides 24 and 25 were detected in NRPS-22c at 13.4% and 2.7%, respectively, lower than the cyclic peptides. The structures of these peptides are shown in Figure 12B. Further combinations of the corresponding plasmids yielded correspondingly additional novel peptides, which are shown in Figures 13 and 14.
[0121] Furthermore, a hybrid of the PKS and NRPS modules shown in Figure 15 was constructed, resulting in the complete synthesis of the glidobactin peptide. To this end, GlbD was split between the ATs by SZ17 / SZ18. In negative controls lacking one or both SZs in each case, virtually no production of glidobactin A was observed.
[0122] In the context of the examples, the following plasmids were used: [Table 1]
[0123] Drawing translation Figure 1 Interaction-pair per Peptide Arrayscore Array score linear network circular network branched network orthogonal network Figure 2 Peptide [#] Peptide (number) Peptide Relative Production % cyclo Figure 3 Peptide [#] Peptide (number) Peptide Relative Production % formyl- Fig. 4 Figure 4 Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) Figure 5 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 6 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 7 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 8 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 9 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 10 Peptide [#] Peptide (number) Peptide % of WT level cyclo Figure 11 Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) cyclo Figure 12 Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) cyclo Figure 13 Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) cyclo Production Figure 14 Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) cyclo Figure 15 Peptide % of WT level Glidobactin A Figure 16A Peptide Production % of WT level Figure 16B Peptide [#] Peptide (number) Peptide Relative Production % Relative Production (%) cyclo Deletion
Claims
1. A protein or protein fragment comprising at least the first domain of a non-ribosomal peptide synthetase (NRPS), the protein or protein fragment has an N-terminus or C-terminus comprising a first binding domain, which first binding domain represents the N-terminus or the C-terminus of the protein or protein fragment, respectively, and which is characterized by the property that the first binding domain is capable of undergoing specific protein-protein binding with at least one corresponding second binding domain, wherein the first binding domain and the corresponding second binding domain each comprise a synthetic coiled-coil domain (SYNZIP); and The protein or protein fragment does not contain a second binding domain and binds to another protein or protein fragment containing an NRPS that has a second binding domain to form a functional NRPS.
2. 2. The protein or protein fragment of claim 1, further comprising at least one, two, three, or four or more additional NRPS domains, wherein the additional NRPS domain(s) are positioned in direct functional arrangement next to the first NRPS domain.
3. The first NRPS domain comprises an A domain, a C domain, a C / E domain, an E domain, a C start 3. The protein or protein fragment according to claim 1 or 2, which is selected from the group consisting of a FT domain, a T domain, and a PG domain.
4. 4. The protein or protein fragment of claim 1, comprising at least an A domain, a C domain, and a T domain, wherein the protein or protein fragment has at least one PKS extender module, initiation module, or termination module.
5. A protein or protein fragment described in any one of claims 1 to 4, characterized in that the end of the protein or protein fragment opposite the first binding domain contains a third binding domain, and the third binding domain is capable of specific protein-protein binding with at least one corresponding fourth binding domain of another protein or protein fragment.
6. 6. The protein or protein fragment of claim 5, wherein the first binding domain and the second binding domain can selectively bind to the third binding domain or the fourth binding domain.
7. The protein or protein fragment of any one of claims 1 to 6, wherein the first binding domain is linked to the first NRPS domain by a linker.
8. An isolated nucleic acid construct comprising a first coding region having a nucleic acid sequence encoding the protein or protein fragment of any one of claims 1 to 7.
9. A vector system for producing a functional NRPS, the vector system comprising at least one nucleic acid construct described in claim 8, and the at least one nucleic acid construct is suitable for expressing at least two proteins or protein fragments described in any one of claims 1 to 7, and the at least two proteins or protein fragments are different and together form a functional NRPS.
10. The vector system of claim 9 , wherein the at least two proteins or protein fragments form the functional NRPS by binding of the first binding domain and the second binding domain.
11. 11. The vector system of claim 9 or 10, comprising a nucleic acid construct suitable for expression of at least three or more proteins or protein fragments according to any one of claims 1 to 7, or at least two of the three or more proteins or protein fragments together form a functional NRPS.
12. The vector system of claim 11, wherein the protein or protein fragment is characterized by the property that the end opposite the first binding domain contains a third binding domain, and the third binding domain is capable of specific protein-protein binding with at least one corresponding fourth binding domain of another protein or protein fragment, and at least three proteins or protein fragments can together form a functional NRPS, wherein the functional NRPS is formed by binding the proteins or protein fragments to each other through binding between the first binding domain and the second binding domain, and binding between the third binding domain and the fourth binding domain.
13. A method for producing a functional (complete) NRPS, comprising contacting at least a first protein or protein fragment described in any one of claims 1 to 7 with a second protein or protein fragment described in any one of claims 1 to 7, wherein the first protein or protein fragment has a terminal first binding domain, and the second protein or protein fragment has a terminal second binding domain in place of the terminal first binding domain.
Citation Information
Patent Citations
System for the assembly and modification of non-ribosomal peptide synthases
WO2019138117A1