Fusion polypeptides for preparing cyclic peptides
Through the design of the fusion polypeptide, combining the self-protease and purification domain, the expansion and stability of the production of cyclic peptides in the prior art are solved, and the efficient large-scale production of a variety of cyclic peptides is achieved.
Patent Information
- Application Number
- CN202380080960.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to produce a variety of different cyclic peptides on a large scale, especially those without N-terminal cysteine, and existing methods cannot effectively transfer from one cyclic peptide to another, and existing tools have shortcomings in scalability and protein stability.
The fusion polypeptide method is adopted, which includes a cleavage C-inclusion peptide domain, a self-proteinase domain, a target polypeptide sequence, a cleavage N-inclusion peptide domain and a purification domain. Through the activation of the proteinase domain and the binding of the purified domain, the efficient production of the cyclic peptide is achieved.
The efficient and large-scale production of a variety of cyclic peptides is achieved, including cyclic peptides without N-terminal cysteine, which improves the scalability and protein stability of cyclic peptides and simplifies the production process.
Smart Images

Figure BDA0005414802770000151 
Figure BDA0005414802770000161 
Figure BDA0005414802770000162
Abstract
Description
[0001] The present invention relates to fusion polypeptides for preparing cyclic peptides, nucleic acid molecules encoding such fusion polypeptides, and genetically modified cells comprising such nucleic acid molecules. Additionally, the present invention relates to methods for preparing cyclic target peptides and mixtures of target peptides. Other aspects of the present invention become apparent when studying the appended patent claims and the specification (including the examples).
[0002] In recent years, new classes of cyclic peptides with beneficial therapeutic and interesting properties have gained attention. To date, there are no naturally occurring cyclic peptides (e.g., cyclopeptides) on the market because the assembly of the required multi-step organic synthesis procedures needs to be customized for any new cyclic peptide. Cyclic peptides can be present in several classes of antibiotics. For example, vancomycin and streptomycin are antibiotics inspired by naturally occurring antibiotics such as penicillin. These substances are prepared in optimized industrial processes and allow for the cost-effective production of certain cyclic peptides. However, this modular process cannot be transferred from one cyclic peptide to another. Biosynthetic methods for expressing cyclic peptides recombinantly do exist, but they are not suitable for scale-up.
[0003] In the early 21st century, methods for screening cyclic peptides for their ability to inhibit the activity of certain polypeptides in a host were developed. It is called SCICLOPPS (split intein mediated circular ligation of peptides and polypeptide). This method has been used to genetically encode variants of dozens of circular peptides and test their inhibitory effects on polypeptides co-expressed with the polypeptide to be tested. Cyclic peptides are released in nature through the split intein mechanism. An intein is a domain in a polypeptide that is encoded such that it releases itself from the polypeptide.
[0004] This split intein mechanism can be used to produce cyclic peptides in vivo. Scott et al. (Production of cyclic peptides and polypeptides in vivo, Proceedings of the National Academy of Sciences in the USA, 96(24)(13638–13643), 1999) reported using split intein domains to catalyze peptide backbone cyclization intracellularly. The "cyclization" used herein means to put a linear (poly)peptide in a cyclic (circular) conformation. This cyclization is carried out to produce peptides and polypeptides that are stable against cellular catabolism in Escherichia coli (E. coli).
[0005] In addition, some other non-naturally cyclic peptides are interesting targets for cyclization. An example is the Green Fluorescent Polypeptide (GFP), which was reported to be cyclized in vivo in the article by Iwai et al. (Cyclic Green Fluorescent Polypeptide Produced in Vivo Using an artificially split PI-PfuI Intein from Pyrococcus furiosis, The Journal of Biological Chemistry, 276(19)(16548 - 16554), 2001). The amount of cyclized GFP reported was in the milligram range.
[0006] All cyclic peptides reported in the literature were produced only in small amounts, and no reliable method for scale-up has been reported. In addition, the split intein mechanism was reported to function as an in vivo method for cyclizing peptides and polypeptides, but it has not been reported to function in vitro.
[0007] Furthermore, the known split intein method does not allow cyclization of sequences that do not contain an N-terminal cysteine. The product of the split intein can only be released and cyclized when the linear precursor of the cyclic product located between two split intein moieties carries an N-terminal cysteine. This cysteine becomes part of the sequence of the product. Cyclic peptides containing an N-terminal amino acid other than cysteine are produced by cyclizing their backbone using sortase. Since different cyclic peptides require different sortases, the sortase-based process cannot be easily transferred from one cyclic peptide to another. Thus, there is currently no general tool available for producing cyclic peptides with variable N-termini.
[0008] International Patent Application WO2019138125 discloses fusion polypeptides having a self-protease domain from N pro for preparing linear target polypeptides. These fusion polypeptides are not useful for producing cyclic target polypeptides and exhibit certain disadvantages in terms of scalability and protein stability.
[0009] Therefore, the main task of the present invention is to provide tools for the large-scale production of a variety of different cyclic peptides and also for the production of cyclized peptides and polypeptides having a native linear configuration.
[0010] This main task is solved by providing a fusion polypeptide that comprises or consists of, in the direction from the N-terminus to the C-terminus: (i) a split C-intein domain, (ii) a self-protease domain, (iii) target polypeptide sequence, (iv) split N - intein domain, and (v) purification domain, which is embedded in the split N - intein domain after aspartic acid at position 71 of the split N - intein domain, wherein the purification domain (v) binds to a carbohydrate matrix.
[0011] The split intein domain is a domain that can be a split C - intein domain or a split N - intein domain, referring to the corresponding position (C - terminus or N - terminus) of the target peptide sequence to be cyclized. These two split intein domains flank the self - protease domain and the target peptide sequence, and the purification domain is embedded in the split N - intein domain after aspartic acid at position 71 of the split N - intein domain. Unexpectedly, it was found that even when the sequence of the N - intein is interrupted, this domain still has activity.
[0012] C - intein and N - intein cyclize the fusion polypeptide as a whole at high pH values (preferably above pH 9.0). In this conformation, the purification domain is still able to bind to the carbohydrate matrix. The N - intein is located at the C - terminus of the target peptide or target polypeptide and catalyzes the release of the C - terminus of the target peptide or target polypeptide and enables the target sequence to be cyclized.
[0013] The self - protease domain exhibits the function of a self - proteolytic cleavage site, which separates the target peptide from the C - intein and the self - protease domain. This domain is activated at certain pH values. The self - protease domain according to the present invention is necessary for releasing the N - terminus of the target polypeptide to prepare it for cyclization ( Figure 1 ).
[0014] The purification domain confers binding of the fusion polypeptide to the carbohydrate matrix. The purification domain can be active in an alkaline environment. This ability is affected by the surrounding domains. The inclusion body signal and the self - protease domain can affect the deposition of the fusion polypeptide in the inclusion body.
[0015] The target peptide domain to be cyclized contains or consists of the amino acid sequence of the target peptide or target polypeptide to be produced. This domain can consist of any amino acid sequence having from 6 to more than 1000 amino acids. Preferably, the target peptide consists of an amino acid sequence of 6 to 1000 amino acids, preferably 6 to 500 amino acids, more preferably 6 to 100 amino acids, and particularly preferably 6 to 50 amino acids. In one embodiment of the invention, based on the total number of amino acids, the amount of hydrophobic amino acids in the target peptide can be ≥10%, more preferably ≥20%, particularly preferably ≥30%, and even more preferably ≥40%. In another embodiment, also based on the total number of amino acids, the amount of hydrophilic amino acids in the target peptide can be ≥10%, preferably ≥20%, particularly preferably ≥30%, and even more preferably ≥40%. In another embodiment, based on the total number of amino acids, the amount of hydrophobic and hydrophilic amino acids in the target peptide can be ≥10%, more preferably ≥20%, particularly preferably ≥30%, and even more preferably ≥40%.
[0016] The split C-intein domain and the split N-intein domain form a peptide bond at the N-terminus of the split C-intein domain and the C-terminus of the split N-intein domain to produce a cyclic polypeptide. Subsequently, the C-terminus of the target polypeptide sequence is released through the autocatalytic site of the split N-intein domain and the self-protease domain. The N-terminus of the target polypeptide sequence is released by forming a peptide bond between the N-terminus and the C-terminus of the target polypeptide. The cyclized target polypeptide is released and the linear remaining fusion polypeptide remains at the carbohydrate matrix.
[0017] Unmodified split inteins are restricted with respect to the N-terminal amino acid. Native split inteins only allow threonine, serine or cysteine to be located at the N-terminus. Replacing the unit G of the C-split intein with the self-protease domain according to the invention allows all amino acids except proline to be the N-terminal amino acid, which greatly increases the number of possible cyclic sequences.
[0018] Unexpectedly, it was found that replacing the C-terminal part of the C-split intein with different autocatalytic domains produced a wide variety of different cyclic sequences, which can be produced and are not limited to the cyclic sequences formed by native split inteins. Native C-splitting requires the N-terminal amino acid of the cyclic product to have cysteine at the N-terminus. This N-terminal cysteine is subsequently released into the product. Thus, native split inteins will only produce cyclic peptides having an internal cysteine at the previous (before cyclization) N-terminal position. This is remedied by inserting a new autocatalytic domain to replace the unit C of the C-split intein. This sequence modification affects the release mechanism, and the autocatalytic cysteine will remain as part of the fusion polypeptide after the cyclic peptide product is released. Thus, the present invention is capable of obtaining a wide variety of cyclic peptides that cannot be produced by methods according to the prior art (such as SCICLOPPS).
[0019] For the present invention, preferably the cyclic peptides produced do not require an N-terminal cysteine at the N-terminal position of the corresponding linear sequence.
[0020] The fusion polypeptides according to the present invention can produce a variety of different cyclic target polypeptides, where these target polypeptides do not need to be cyclic polypeptides in nature. In addition, target polypeptides with a linear native conformation can be cyclized using the fusion polypeptides according to the present invention.
[0021] In addition, the fusion polypeptides according to the present invention can be used to produce cyclic target polypeptides economically and on a large scale.
[0022] In a preferred embodiment, the fusion polypeptide according to the present invention has a purification domain, which comprises or consists of: the amino acid sequences according to SEQ ID No.: 1 to SEQ ID No.: 3 or amino acid sequences having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 1 to SEQ ID No.: 3.
[0023] The purification domain can consist of different carbohydrate binding modules (CBMs) from different organisms. The binding strength of the purification domain to the carbohydrate matrix can be enhanced by combining single building units of the carbohydrate binding module (CBM) with each other. The binding can be stabilized under specific reaction conditions (such as high ionic strength), and the size of the purification domain can be altered to fit the desired target peptide domain. The purification domain according to SEQ ID No.: 1 to SEQ ID No.: 3 or amino acid sequences having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 1 to SEQ ID No.: 3 shows preferred properties in terms of binding strength and optimized interaction with other domains of the fusion polypeptide.
[0024] Whenever the present disclosure refers to a percentage of sequence identity between nucleic acid or amino acid sequences, these values are defined as those obtained by using the EMBOSS Water pairwise sequence alignment (nucleotide) program for nucleic acids or the EMBOSS Water pairwise sequence alignment (polypeptide) program for amino acid sequences. An alignment or sequence comparison as used herein refers to an alignment of the entire full length of two sequences being compared to each other. Those tools for local sequence alignment provided by the European Molecular Biology Laboratory (EMBL) European Bioinformatics Institute (EBI) use a modified Smith-Waterman algorithm. When performing an alignment, the default parameters defined by the EMBL-EBI are used. Those parameters are (i) for amino acid sequences: matrix = BLOSUM62, gap open penalty = 10 and gap extension penalty = 0.5 or (ii) for nucleic acid sequences: matrix = DNAfull, gap open penalty = 10 and gap extension penalty = 0.5. A person skilled in the art will be aware of the fact that, for example, if a sequence encoding a polypeptide is to be used in another organism as compared to the original organism from which the molecule is derived, the corresponding sequence may be "codon optimized".
[0025] Another preferred embodiment of the invention relates to a fusion polypeptide according to the invention, wherein the self-protease domain (ii) comprises or consists of the following: the amino acid sequences according to SEQ ID No.: 8 to SEQ ID No.: 12 or amino acid sequences having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 8 to SEQ ID No.: 12.
[0026] Preferably (and advantageously, especially in combination with some of the preferred embodiments described herein), the self-protease domain is activated at a pH value of 6.8 or higher (i.e., not activated at a pH below 6.8), more preferably activated at a pH value of 6.8 to 7.2. The self-protease is subject to the pestivirus self-protease N proinspired and modified such that the pH of the environment (rather than the chaotropic agent concentration) is the activation trigger. The autoprotease domain of the amino acid sequence according to SEQ ID No.: 8 to SEQ ID No.: 12 or having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 8 to SEQ ID No.: 12 exhibits preferred properties in terms of targeting and activating autoproteolytic activity under the specific pH conditions defined above.
[0027] The autoproteolytic activity of the autoprotease domain is based on the catalytic diade of histidine and cysteine in the enzyme active site of the autoprotease. These enzymes are the basis of the autoprotease domain of the present invention. Such a basis for the autoprotease domain can be the autoprotease N from pestivirus pro or the autoprotease from potyvirus, picornavirus or any other viral autoprotease. By targeted recombination or redesign of these sequences, autoprotease domain building blocks can be designed that, alone or in combination, exhibit several advantages over their native counterparts. In one aspect, the pH sensitivity of the autoprotease can be precisely adjusted. This shows the advantage that the activity of the autoprotease can be controlled to suit the desired reaction conditions. The autoprotease can be precisely activated at the desired pH in the case of a very narrow pH value range and premature release of the target peptide can be avoided, or it can be carried out at extreme pH values where the native autoprotease is no longer stable.
[0028] Another preferred embodiment of the present invention relates to a fusion polypeptide according to the present invention, wherein the fusion polypeptide further comprises an inclusion body sequence selected from SEQ ID No.: 13 to SEQ ID No.: 15 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 13 to SEQ ID No.: 15.
[0029] Preferably, the inclusion body domain is embedded in the C-split intein domain and has the following sequence from the N-terminus to the C-terminus: the first part of the C-split intein, which preferably has an amino acid sequence selected from SEQ ID No.: 4 or SEQ ID No.: 6; the inclusion body sequence, which preferably has an amino acid sequence according to any one of SEQ ID No.: 13 to SEQ ID No.: 15; and the second part of the C-split intein, which preferably has an amino acid sequence selected from SEQ ID No. 5 or SEQ ID No.: 7.
[0030] Unexpectedly, it has been found that these sequences not only control the promotion of inclusion bodies, but also direct the refolding process in a strong alkaline environment.
[0031] Signal sequences are always selected based on their effect on the reprocessing of the target peptide. In one embodiment, an inclusion body signal sequence that directs the target polypeptide to inclusion bodies is used.
[0032] A preferred embodiment of the invention relates to a fusion polypeptide according to the invention, wherein the split C-intein domain (i) and the split N-intein domain (ii) are derived from a naturally occurring DnaE split intein sequence, preferably from the organism Synechocystes sp PCC6803 or Nostoc punctiforme.
[0033] Another preferred embodiment relates to a fusion polypeptide according to the invention, wherein the split C-intein domain (i) and the split N-intein domain (ii) comprise or consist of: amino acid sequence pairs selected from SEQ ID No.: 17 and SEQ ID No.: 18, SEQ ID No.: 19 and SEQ ID No.: 20, SEQ ID No.: 21 and SEQ ID No.: 22, SEQ ID No.: 23 and SEQ ID No.: 24, SEQ ID No.: 25 and SEQ ID No.: 26, SEQ ID No.: 27 and SEQ ID No.: 28, or amino acid sequence pairs having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID No.: 17 and SEQ ID No.: 18, SEQ ID No.: 19 and SEQ ID No.: 20, SEQ ID No.: 21 and SEQ ID No.: 22, SEQ ID No.: 23 and SEQ ID No.: 24, SEQ ID No.: 25 and SEQ ID No.: 26, SEQ ID No.: 27 and SEQ ID No.: 28, wherein the first SEQ ID No. relates to the C-intein domain (i) and the second SEQ ID No. relates to the N-intein domain (ii).
[0034] In a preferred embodiment, the present invention relates to a fusion polypeptide having an amino acid sequence according to any one of SEQ ID No.: 29 to 44 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 29 to 44, wherein the target polypeptide is inserted after the N-terminal cysteine of the self-protease domain.
[0035] Another aspect of the present invention relates to a nucleic acid molecule encoding a fusion polypeptide according to the present invention.
[0036] In a preferred embodiment, the present invention relates to a nucleic acid molecule encoding a fusion polypeptide having a sequence according to SEQ ID No.: 77 to 92 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 77 to 92, wherein the nucleic acid sequence encoding the target polypeptide is inserted at the following positions: after the tgc triplet at position 588 in nucleic acid sequences SEQ ID No.: 77 and 85, after the tgc or tgt triplet at position 573 in nucleic acid sequences SEQ ID No.: 78 and 86, after the tgc or tgt triplet at position 588 in nucleic acid sequences SEQ ID No.: 79, 80, 82, 87, 88 and 90, after the tgc or tgt triplet at position 1014 in nucleic acid sequences SEQ ID No.: 81 and 89, after the tgc or tgt triplet at position 582 in nucleic acid sequences SEQ ID No.: 83 and 91, after the tgc or tgt triplet at position 336 in nucleic acid sequence SEQ ID No.: 84, and after the tgc or tgt triplet at position 678 in nucleic acid sequence SEQ ID No.: 92.
[0037] Another preferred embodiment of the present invention relates to a fusion polypeptide having an amino acid according to the sequence SEQ ID No.: 37 to 44 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 37 to 44, wherein the fusion polypeptide carries a target peptide which is inserted after the N-terminal cysteine of the self-protease domain at position 196 of SEQ ID NO.: 37 to 44.
[0038] Another aspect of the present invention relates to a genetically modified cell comprising a recombinant nucleic acid molecule according to the present invention, wherein the cell is capable of expressing a fusion polypeptide according to the present invention.
[0039] In a preferred embodiment of the genetically modified cell according to the invention, the cell is selected from: Escherichia coli, Vibrio natrigens, Saccheromyces cerevisiae, Aspergillus niger, green algae, microalgae, HEK T293 and Chinese hamster ovary cell (CHO).
[0040] Another aspect of the invention relates to a method for preparing a target peptide, which comprises the following steps: (a) providing a genetically modified cell as defined above, (b) culturing the cell under conditions suitable for expressing the fusion polypeptide according to the invention, (c) obtaining the fusion polypeptide and optionally, unfolding the obtained fusion polypeptide and refolding the fusion polypeptide directionally, (d) contacting the fusion polypeptide obtained in step (c) with a carbohydrate matrix, (e) cyclizing the fusion polypeptide by splitting the C-intein domain (i) and the split N-intein domain (ii), (f) cleaving the fusion polypeptide by activating the self-protease domain of the fusion polypeptide to obtain a cyclic target peptide, (g) collecting the mixture containing the cyclic target peptide.
[0041] Step (a) includes providing a genetically modified cell that expresses the fusion polypeptide. Such cells can be obtained by introducing a nucleic acid molecule (preferably in the form of a vector) containing a sequence encoding the fusion polypeptide into the cell by known methods (e.g., by transfection or transformation).
[0042] In step (b), the cell is cultured under conditions suitable for expressing the fusion polypeptide according to the invention, preferably in high-density culture. The culture conditions and in particular the conditions for achieving high-density culture and the corresponding culture media are well known to those skilled in the art. In one embodiment of the invention, the expression of the fusion polypeptide is achieved and subsequently transported to inclusion bodies using a suitable signal sequence.
[0043] If the fusion polypeptide is present in inclusion bodies, step (c) includes obtaining the fusion polypeptide from the culture broth and optionally unfolding the obtained fusion polypeptide and refolding the fusion polypeptide directionally. Dissolution conditions for treating inclusion bodies and conditions for directional refolding are well known in the art. Preferably, the inclusion bodies are dissolved by using 6 M guanidine chloride, 8 M urea, or 2 wt% sodium dodecyl sulfate and refolded under neutral or weakly basic conditions. Directional refolding is initiated when the concentration of chaotropic agent or detergent in the composition is below 1 wt%.
[0044] In step (d), the dissolved fusion polypeptide is contacted with a carbohydrate matrix such that the fusion polypeptide binds to the matrix via its purification domain (v). Cyclization of the fusion polypeptide backbone occurs under environmental conditions present in detergent and chaotropic agent conditions. The purification domain (v) is not negatively affected by this cyclization. The cyclic fusion polypeptide binds to the starch mixture.
[0045] This step is carried out under conditions in which the autocatalytic domain (ii) is inactive, preferably by controlling the pH value rather than the concentration of chaotropic agent or denaturing agent, to avoid premature cleavage of the target peptide domain (iii) on the one hand and to induce the activity of the purification domain (v) on the other hand. In parallel, the split inteins ((i), (iv)) are initiated under environmental conditions that keep the autocatalytic enzyme inactive on the one hand but allow the purification domain (v) to be active on the other hand. Under these conditions, based on the total amount of the fusion polypeptide, the amount of the cleaved fusion polypeptide is preferably < 10%, more preferably < 5%, particularly preferably < 3%, or even more preferably < 1%.
[0046] The potential mechanisms for the inactivity of the autocatalytic domain can be described by the following two scenarios: (1) Conditions where the autocatalytic domain is constitutively inactive and is activated only by changing environmental conditions (e.g., by adjusting temperature, pH, and / or ionic strength, preferably by adjusting pH); (2) Conditions where the autocatalytic domain is constitutively active, however, its activity is not sufficient to achieve premature cleavage of the target peptide domain during the time period required for carrying out method step (d), i.e., it is kinetically inactive, preferably for a maximum of 10 minutes, more preferably for a maximum of 20 minutes, and particularly preferably for a maximum of 30 minutes.
[0047] In one embodiment, step (d) is carried out under native conditions, i.e., conditions in which the autocatalytic enzyme is constitutively active. Unexpectedly, it was found that even when the fusion polypeptide is present in its native state, the autocatalytic domain is still sufficiently inactive during step (d). This effect is particularly present when using an autocatalytic enzyme having the amino acid sequence according to SEQ ID NO.: 7 to 11.
[0048] Preferably, an insoluble carbohydrate matrix is used in step (d), which is conducive to the separation of impurities. Controlling the release of the N-terminus in step (d) mediated by the catalytic site of the pH-sensitive self-protease is crucial for the subsequent cyclization process of the target sequence. Before cyclization and release of the target sequence, the cyclic fusion polypeptide binds to the carbohydrate matrix.
[0049] In step (e), the fusion polypeptide is cleaved by the catalytic sites of the self-protease domain (ii) and the split N-intein (iv) and the cyclized target polypeptide (iii) is released. Cleavage of the fusion polypeptide can be induced by the addition of a self-proteolysis buffer, i.e., a buffer that provides conditions (i.e., acidic conditions) under which the self-protease is active.
[0050] Step (f) results in the obtaining of a mixture by eluting the cleaved target peptide from the column. Preferably, the elution is carried out using a buffer selected from: HEPES, PBS, and TrisHCl, at a concentration of 1 to 100 mM and a pH of 6.5 to 7.5. Additionally, the preferred buffer can be supplemented with arginine at a concentration of 10 to 100 mM or sucrose at a concentration of 2 to 20 mM.
[0051] A preferred embodiment of the method according to the invention relates to a method in which the carbohydrate matrix in step (d) consists of or comprises a substance selected from: starch, lignin, carbohydrate polymers, copolymers having α-1,4-glycosidic bonds and α-1,6-glycosidic bonds of glucose or other sugars, and mixtures thereof, and the carbohydrate matrix preferably exists in a packed column as a packed matrix or as starch granules composed of amylose and amylopectin.
[0052] Starch is a complex mixture of carbohydrates from different sugar polymers. Plant cells collect the sugars they produce in storage organelles called vacuoles. When the cells and organelles are mechanically disrupted, the starch granules are released. Raw starch varies according to the plant species. Starch can have different particle sizes with diameters ranging from less than 25 μm to greater than 100 μm. The higher the proportion with a diameter exceeding 75 μm, the higher the probability of non-specific adsorption and thus the higher the probability of retaining impurities in the product after starch purification. Additionally, there are starch granules that are porous and can adsorb amylase in their internal channels (e.g., wheat). Starch consists of the components amylose and amylopectin. In contrast to amylopectin, amylose is water-soluble. The swelling behavior of each starch in water also depends on the proportion of these two types. Thus, unpurified corn starch in water obtains a cement-like consistency, while table potato starch remains water-permeable. All carbohydrate-binding enzymes have a high affinity for their substrates, which also exists under extreme conditions. Preferably, the starch granules are insoluble in water. Additionally, it is preferred if the soluble amylose fraction and polypeptides have been removed from the starch.
[0053] Another preferred embodiment of the method according to the invention relates to a method in which the cyclization of the fusion polypeptide by cleavage of the C-intein domain (i) and cleavage of the N-intein domain (ii) in step (e) is carried out at a pH above 7.5, preferably at a pH above 9.0.
[0054] Another preferred embodiment of the method according to the invention relates to a method in which the activation of the autocatalytic domain in step (e) is carried out at a pH of 6 to 8, preferably at a pH of 6.5 to 7.5, particularly preferably at a pH of 7 to 7.4.
[0055] One aspect of the invention relates to a recombinant nucleic acid molecule encoding a fusion polypeptide according to the invention and a cloning site for incorporating the recombinant nucleic acid molecule according to the invention, which is optionally operably linked to an expression control sequence.
[0056] Another aspect of the invention relates to a mixture comprising or consisting of: a cyclic target peptide, preferably a synthetic cyclic target peptide, and a total of 0.001 wt% to 1 wt% of sodium and / or potassium, based on the total weight of sodium (if present), potassium (if present) and the target peptide combined, wherein the mixture is obtained or can be obtained by the method according to the invention.
[0057] Preferably, the cyclic target peptide comprised in the mixture according to the invention does not require an N-terminal cysteine at the N-terminal position of the corresponding linear sequence.
[0058] The present invention is further characterized by the following illustrative, non-limiting examples. Brief description of the sequence
[0059] SEQ ID No.: 1 to SEQ ID No.: 3 are artificial amino acid sequences encoding purification domains.
[0060] SEQ ID No.: 4 to SEQ ID No.: 7 are artificial amino acid sequences encoding C-split intein sequence variants.
[0061] SEQ ID No.: 8 to SEQ ID No.: 12 are artificial amino acid sequences encoding self-protease domains.
[0062] SEQ ID No.: 13 to SEQ ID No.: 15 are artificial amino acid signal sequences for intracellular targeting of fusion polypeptides according to the present invention.
[0063] SEQ ID No.: 16 to SEQ ID No.: 28 are artificial and non-artificial amino acid sequences encoding pairs of split C-inteins and split N-inteins. Sequences with SEQ ID No. 17, 19, 21, 23, 25 and 27 encode split C-inteins and sequences with SEQ ID No. 16, 18, 20, 22, 24, 26 and 28 encode split N-inteins.
[0064] SEQ ID No.: 29 to SEQ ID No.: 44 are artificial amino acid sequences encoding preferred fusion polypeptides, wherein the target polypeptide sequence can be inserted after the N-terminal cysteine of the self-protease domain.
[0065] SEQ ID No.: 45 to SEQ ID No.: 48 are amino acid sequences of target peptides to be cyclized.
[0066] SEQ ID No.: 49 to SEQ ID No.: 51 are artificial nucleic acid sequences encoding purification domains.
[0067] SEQ ID NO.: 52 to SEQ ID No.: 55 are artificial nucleic acid sequences encoding C-split intein sequence variants.
[0068] SEQ ID No.: 56 to SEQ ID No.: 60 are artificial nucleic acid sequences encoding self-protease domains.
[0069] SEQ ID No.: 61 to SEQ ID No.: 63 are nucleic acid sequences encoding signal sequences for intracellular targeting of fusion polypeptides according to the present invention.
[0070] SEQ ID No.: 64 to SEQ ID No.: 76 are artificial and non-artificial nucleic acid sequences encoding split C-intein and split N-intein pairs. The sequences with SEQ ID No.: 65, 67, 69, 71, 73, and 75 encode split C-inteins and the sequences with SEQ ID No.: 64, 66, 68, 70, 72, 74, and 76 encode split N-inteins.
[0071] SEQ ID No.: 77 to SEQ ID No.: 92 are artificial nucleic acid sequences encoding preferred fusion polypeptides, wherein the target polypeptide sequence can be inserted after the N-terminal cysteine of the self-protease domain.
[0072] SEQ ID No.: 93 to SEQ ID No.: 96 are nucleic acid sequences encoding target peptides to be cyclized. Description of the Drawings
[0073] Figure 1 A schematic diagram of the cyclization reaction that produces the cyclized product is shown. The reaction begins with the formation of a peptide bond at the N-terminus of the split C-intein and the C-terminus of the split N-intein, producing a circular polypeptide. The following steps occur in an orderly manner. The C-terminus of the product sequence is released through the self-catalytic site of the split N-intein, and the N-terminus of the product sequence is released by the formation of a peptide bond between the N-terminus and the C-terminus of the target peptide. In this way, the cyclic product is released and the linear sequence remains at the starch matrix.
[0074] Figure 2 A shows the results of the kinetic evaluation of the release from the fusion polypeptide according to SEQ ID No.: 39 at 421 nm and 397 nm at pH 7.0. The upper curve shows the absorption at 421 nm and the lower curve shows the absorption at 397 nm. GFP has a maximum absorption at 397 nm in its linear form and a maximum absorption at 421 nm in its cyclic form. It is shown that most of the product released from the fusion peptide is cyclic GFP, with a lower amount of linear GFP released. Figure 2 B is an unpurified sample of the fusion protein according to SEQ ID No.: 39 carrying GFP as the target peptide to be cyclized after activation. The size of the product cyclic GFP is 28.5 kDa. Figure 2Panel C shows the purified sample, in which the fusion polypeptide according to SEQ ID No.:39 carrying GFP as the target peptide was subjected to a starch matrix and incubated (binding). The self-protease was activated and the target peptide circular GFP was eluted. The sample was then analyzed on a tricine gel. It can be seen that the sample contains almost no impurities. The bands at 75 and 25 kDa in the first lane of Panel C are the bands of the marker.
[0075] Figure 3 Panel A shows the results of the kinetic evaluation of the release from the fusion polypeptide according to SEQ ID No.:40 at pH 7.0 at 421 nm and 397 nm. The upper curve shows the absorption at 421 nm and the lower curve shows the absorption at 397 nm. GFP has a maximum absorption at 397 nm in its linear form and at 421 nm in its circular form. It is shown that most of the released product from the fusion peptide is circular GFP, with a lower amount of linear GFP released. Figure 3 Panel B is an unpurified sample of the fusion protein according to SEQ ID No.:40 carrying GFP as the target peptide to be cyclized after activation. The size of the product circular GFP is 28.5 kDa.
[0076] Figure 4 Panel A shows the results of the kinetic evaluation of the release from the fusion polypeptide according to SEQ ID No.:41 at pH 7.0 at 421 nm and 397 nm. The upper curve shows the absorption at 421 nm and the lower curve shows the absorption at 397 nm. GFP has a maximum absorption at 397 nm in its linear form and at 421 nm in its circular form. It is shown that most of the released product from the fusion peptide is circular GFP, with a lower amount of linear GFP released. Figure 4 Panel B is an unpurified sample of the fusion protein according to SEQ ID No.:41 carrying GFP as the target peptide to be cyclized after activation. The size of the product circular GFP is 28.5 kDa.
[0077] Figure 5 Panel A shows the results of the kinetic evaluation of the release from the fusion polypeptide according to SEQ ID No.:42 at pH 7.0 at 421 nm and 397 nm. The upper curve shows the absorption at 421 nm and the lower curve shows the absorption at 397 nm. GFP has a maximum absorption at 397 nm in its linear form and at 421 nm in its circular form. It is shown that most of the released product from the fusion peptide is circular GFP, with a lower amount of linear GFP released. Figure 5B is an unpurified sample of the fusion protein according to SEQ ID No.:42 carrying GFP as the target peptide to be cyclized after activation. The size of the product cyclic GFP is 28.5 kDa.
[0078] Figure 6 A shows the results of the kinetic evaluation of the release from the fusion polypeptide according to SEQ ID No.:44 at pH 7.0 at 421 nm and 397 nm. The upper curve shows the absorption at 421 nm and the lower curve shows the absorption at 397 nm. GFP has a maximum absorption at 397 nm in its linear form and at 421 nm in its cyclic form. It is shown that most of the product released from the fusion peptide is cyclic GFP, with a lower amount of linear GFP released. Figure 6 B is an unpurified sample of the fusion protein according to SEQ ID No.:44 carrying GFP as the target peptide to be cyclized after activation. The size of the product cyclic GFP is 28.5 kDa. The second lane shows the marker with bands at 25 and 75 kDa.
[0079] Figure 7 A tricine gel showing an E. coli culture sample loaded with a construct expressing the fusion polypeptide according to SEQ ID NO.:39, which carries the following different target peptides: GFP according to SEQ ID No.:45 (SEQID No.:39 - PcGFP), MCoTI-II (trypsin inhibitor) snake venom (SEQ ID No.:39 - 46), and Cycloviolacin O14 (SEQ ID No.:39 - 47). The bands are as follows: Std. represents the marker, a) is the uninduced control, b) is the culture after induction, c) is the supernatant of the culture after separation of the lysate supernatant from the inclusion body precipitate, d) is the inclusion body precipitate after denaturation, and e) is sample d) mixed with starch and the buffer according to Table 10. The size of the fusion polypeptide according to SEQ ID No.:39 without any product is shown on the gel as 56.62 kDa and can also be detected if the cyclization reaction has occurred and thus the product is no longer contained in the fusion protein. The size of the fusion polypeptide containing cyclic GFP is 85.2 kDa. The size of the fusion polypeptide with MCoTI-II is 59.8 kDa and the size of the fusion polypeptide with Cycloviolacin O14 is 59.8 kDa. Thus, all three target peptides can be produced using the polypeptide according to the invention.
[0080] Figure 8A tricine gel of an E. coli culture sample loaded with a construct expressing a fusion polypeptide according to SEQ ID NO.: 40 is shown. The fusion polypeptide carries the following different target peptides: GFP according to SEQ ID No.: 45 (SEQID No.: 40-PcGFP), the MCoTI-II trypsin inhibitor from snake venom (SEQ ID No.: 40-46), and Cycloviolacin O14 (SEQ ID No.: 40-47). The bands are as follows: Std. represents the marker, a) is the uninduced control, b) is the culture after induction, c) is the supernatant of the culture after separation of the lysate supernatant from the inclusion body precipitate, d) is the inclusion body precipitate after denaturation, and e) is sample d) mixed with starch and the buffer according to Table 10. The size of the fusion polypeptide according to SEQ ID No.: 40 without any product is shown on the gel as 56.63 kDa. The size of the fusion polypeptide containing circular GFP is 85.2 kDa. The size of the fusion polypeptide with MCoTI-II is 59.8 kDa and the size of the fusion polypeptide with Cycloviolacin O14 is 59.8 kDa. Thus, all three target peptides can be produced using the polypeptide according to the invention.
[0081] Figure 9 A tricine gel of an E. coli culture sample loaded with a construct expressing a fusion polypeptide according to SEQ ID NO.: 42 is shown. The fusion polypeptide carries the following different target peptides: GFP according to SEQ ID No.: 45 (SEQID No.: 42-PcGFP), the MCoTI-II trypsin inhibitor from snake venom (SEQ ID No.: 42-46), Cycloviolacin O14 (SEQ ID No.: 42-47), and sunflower trypsin inhibitor 1 (SEQ ID No.: 42-48). The bands are as follows: Std. represents the marker, a) is the uninduced control, b) is the culture after induction, c) is the supernatant of the culture after separation of the lysate supernatant from the inclusion body precipitate, d) is the inclusion body precipitate after denaturation, and e) is sample d) mixed with starch and the buffer according to Table 10. The size of the fusion polypeptide according to SEQ ID No.: 42 without any product is shown on the gel as 56.87 kDa. The size of the fusion polypeptide containing circular GFP is 85.42 kDa. The size of the fusion polypeptide with MCoTI-II is 60.01 kDa, the size of the fusion polypeptide with Cycloviolacin O14 is 60.03 kDa, and the size of the fusion polypeptide with sunflower trypsin inhibitor 1 is 58.63 kDa. Thus, all four target peptides can be produced using the polypeptide according to the invention. Examples Example 1: Generation of circular GFP Construct design
[0082] The feasibility of the method according to the invention using the fusion polypeptide according to the invention was tested with three different peptides, said three different peptides having an N-terminal glycine and different numbers of intramolecular cysteine bridges and not being cyclizable using methods available in the prior art. GFP (SEQ ID No.: 45) was tested as a proof of principle. The peptides are listed in Table 1. Table 1: Target sequences tested SEQ ID No. Product Mass k[Da] 45 GFP 28.56 46 MCoTI-II (trypsin inhibitor)) snake venom 32.07 47 Cycloviolacin 014 31.68 48 Sunflower trypsin inhibitor 1 14.98
[0083] During the cyclization reaction, different fragments were generated to demonstrate a) the presence of the fusion polypeptide and its different parts b) different mechanisms of product release and different chemical environments.
[0084] The sizes of these fragments are listed in Table 2. The columns of the fragments show the differences in mass between different constructs. The constructs can be distinguished by single mutations or exchanges, additions or deletions of motifs. The changes or distinguishing features between different constructs can be found in the catalytic domains of both the N-terminal and C-terminal domains of the fusion protein. Table 2: Components of the fusion polypeptide located at the N-terminal and C-terminal of the product
[0085] Table 3 below shows the composition and mass of different fusion polypeptides carrying the target peptides GFP (SEQ ID No.: 45), MCo-TI-II (SEQ ID No.: 46), Cycloviolacin O14 (SEQ ID No.: 47), Sun Flower Trypsin Inhibitor I (SFT-I) (SEQ ID No.: 48) to be cyclized: Table 4: Composition of different fusion polypeptides carrying the product peptides to be cyclized. Explanation of the reaction
[0086] The fusion polypeptide tested to generate cyclized polypeptide sequences consists of a combination of N-terminal and C-terminal split inteins, an autocatalytic protease domain, a cyclized product sequence, and a purification domain. The N-terminus of the product sequence is flanked by the split C-intein and the catalytic domain. The C-terminus consists of the split N-intein domain and the starch-binding purification domain. The purification domain enhances the purification and yield of the cyclized product. When the fusion polypeptide is activated, the N-terminus of the C-split intein and the C-terminus of the N-split intein form a peptide bond, generating a circular polypeptide. When the product sequence is released through the autocatalytic protease domain and the split N-intein domain, the product sequence is cyclized. The release of the circular polypeptide produces two different products: the unloaded and linearized fusion polypeptide retained in the purification matrix and the cyclized product polypeptide also released from the column. Replacing the catalytic domain of the split C-intein with another autocatalytic protease domain will enhance both the controllability of product release and the product profile, as more amino acids other than cysteine, threonine, and serine will be allowed at the N-terminus of the cyclized product. The size of the fusion polypeptide relative to the cyclized product has an impact on the reaction, especially on the pH value of the cyclization reaction. The cyclization reaction can be adjusted for different target pH values or pH ranges to achieve better controllability of product release.
[0087] The cyclization reaction is depicted in Figure 1 and schematically shows the cyclization reaction of the product. The given pH value is an example and can vary depending on the product and the fusion polypeptide construct used. Example 2: Generation of fusion polypeptides
[0088] The loaded fusion polypeptide is produced in E. coli. The gene encoding the fusion polypeptide for circular peptide release is contained on a pET vector, which allows control and enhancement of the expression of the fusion polypeptide from the plasmid in an appropriate cell strain.
[0089] The vector containing the fusion polypeptide for circular peptide release is used as a vector to transform an appropriate host cell (e.g., E. coli) and the transformed host cells are stored on a selection agar plate. The colonies are cultured to express the recombinant fusion polypeptide.
[0090] Pick colonies into 10 mL of Luria Bertolli medium containing 30 mg / L kanamycin and 30 mg / L chloramphenicol. Stir the culture overnight at 37 °C. Transfer the overnight culture to 180 mL of Luria Bertolli medium and 20 mL of potassium phosphate solution (Table 6). Incubate the bacteria at 170 rpm in 200 mL of medium at a temperature between 23 °C and 37 °C for three to six hours. Harvest the cells and resuspend them in 900 ml of Luria Bertolli medium containing 30 mg / L kanamycin and 30 mg / L chloramphenicol and 100 mL of phosphate solution (Table 6). Before introducing the cells, the pH of the medium is slightly higher than pH 7.2, which is close to the activation point of the autocatalytic domain. For this reason, pH stabilization measures are taken. After a 2-hour incubation, add 25 mL of alkaline solution (Table 5) to one liter of medium. After another 2-hour incubation, add another 25 mL of alkaline solution (Table 5) and then induce. Allow the culture to grow for another three to six hours depending on the desired optical density of the culture. Then induce the expression culture by applying 25 mL of feed salt solution (Table 7) over a period of one hour. The feed contains isopropyl β-D-1-thiogalactopyranoside (IPTG) or IPTG is applied independently. The final concentration is 2 mM. Incubate the expression culture at 170 rpm at a temperature between 23 °C and 37 °C for three to twelve hours. Table 5: Alkaline solution Table 6: Potassium phosphate solution Table 7: Salt solution for feeding Downstream processing
[0091] Harvest the cells by centrifugation at 8,000 rpm and 4 °C for 10 minutes. Discard the medium and store the cells at -80 °C overnight. Keep the pH above 7.4.
[0092] Thaw the cells in lysis and wash buffer (Table 8) at a ratio of 35 mL per 8 g of cells (approx. 1:4 (w / v) ratio). Resuspend the cells in the buffer and immediately homogenize in a disperser at 2,800 rpm for 5 minutes 30 seconds without additional cooling. Then collect the lysate and store it at 4 °C for 20 minutes. Centrifuge the lysate at 7,000 rpm at 4 °C for five to eight minutes. Discard the supernatant and resuspend the pellet in the same volume of 4 °C cold lysis and wash buffer (Table 8) by vortexing for 60 seconds. Then store the suspension at 4 °C again and centrifuge at 7,000 rpm at 4 °C for five to eight minutes once more. Discard the supernatant and replace it with 35 mL of deionized water. Resuspend the lysate again by vortexing for 60 seconds and store it at 4 °C for at least twenty minutes. Centrifuge the lysate in water a third time at 6,000 rpm at 4 °C for five to eight minutes. Discard the supernatant containing DNA-rich contaminants. The washed inclusion bodies can be stored at -80 °C or used immediately for the next step.
[0093] Immerse 2 g of inclusion bodies in 8 mL to 10 mL of denaturing buffer (Table 7). A ratio of 1:4 to 1:10 (w / v) is suitable. Vortex the suspension for 1 minute 40 seconds and then incubate at room temperature for 8 minutes. Repeat this sequence six times. After one hour, the suspension should be clear. If not, incubate the suspension at room temperature for an additional 60 minutes. When clear, store the suspension at 4 °C for two to twelve hours and centrifuge at 8,000 to 10,000 rpm for six to eight minutes. Collect the supernatant and store it at 4 °C for 30 minutes to two hours. If more SDS polypeptide aggregates precipitate, centrifuge the suspension again. The detergent-free sample can be stored at 4 °C or room temperature for several weeks. The pH will be higher than 11.0.
[0094] Now bring the suspension into contact with a starch mixture consisting of a 2:3 mixture of wheat and potato starch washed in starch washing medium (Table 10). Sieve the starch before use. The starch has a particle size of 25 to 32 μm. Mix the starch in the buffer according to Table 6 and wash it twice therein. Centrifuge the starch at 5,000 rpm at 23 °C for 5 minutes. Load the starch in the form of a slurry into a column or centrifuge beaker. After loading the fusion polypeptide sample, remove the buffer by elution or centrifugation. Other starches, such as corn starch, rice starch, and starches from other plant, fungal, or animal sources, can be used.
[0095] After loading the sample, the buffer pH on the column was gradually decreased by adding activation buffers of different pH values (Table 11) at different rates and eluting them at different rates. The eluate or centrifuged supernatant containing the desired cyclized product was lyophilized. The supernatant was collected by centrifuging at 6000 to 8000 rpm for 6 to 8 minutes at 4 °C. Elution was carried out at a flow rate of 0.1 ml / min to 1 ml / min at atmospheric pressure. Table 8: Lysis and Wash Buffer Table 9: Denaturing Buffer Table 10: Starch Wash Buffer Table 11: Activation Buffer Example 3: Photometric and gel electrophoresis studies of the cyclization process
[0096] The fusion protein performance was tested at different pH values from 4.0 to 8.0. In this case, "performance" refers to the successful production of cyclic product peptides or polypeptides.
[0097] The test sample was a mixture of a fusion protein in a detergent-free denaturing buffer (Table 9) and nine parts of an aqueous solution with a pH value from 5.0 to 9.0. The best results regarding the release and purification efficacy of the cyclic product were found in a 1:10 (one part protein solution and nine parts buffer) mixture of the fusion polypeptide in a detergent-free denaturing buffer and an activation buffer with a pH value of 7.2 (Table 11).
[0098] The analysis and results are discussed in more detail below. Study of kinetics
[0099] A kinetic study of the fusion protein was performed photometrically using an Implen n120 nanophotometer and GFP as the product to be cyclized, which had been activated by dissolving it in the activation buffer at a ratio of 1:10. (1) Transfer 10 μl of a clarified fusion polypeptide (according to SEQ ID No.: 39, 40, 41, 42, and 44) solution in a detergent-free denaturing buffer (Table 7) to 90 μl of an activation buffer with a pH of 7.0, which allows activation of the catalytic domain of the fusion polypeptide and release of cyclic GFP. (2) Carefully mix the two solutions by slowly pipetting and slowly stirring the mixture for 30 seconds. Incubate the solution at room temperature for an additional 30 seconds and then spin down. (3) Transfer 2 μl of the solution to the sample area of the nanophotometer. The sample is referenced against the activation buffer. The sample is then measured over a 20-minute time period and at wavelengths from 250 nm to 600 nm. (4) Compare the results of the same fusion polypeptide carrying cyclic GFP as the product with the results of the same fusion polypeptide with other cyclic products at 280 nm, 397 nm, and 421 nm.
[0100] Figures 2 to 6 Release of cyclic GFP from the fusion polypeptide is shown. If absorption is detected at 421 nm, cyclic GFP is released and linear GFP has an absorption maximum at 397 nm. The figure indicates that GFP is released in its cyclic form and only a small amount is released in its linear form. Study of productivity and purification
[0101] Test product cyclization in small-scale cultures expressed and processed as described in Example 2 using fusion polypeptides according to SEQ ID No.: 39, SEQ ID No.: 40, and SEQ ID No.: 42 with GFP, MCoTI-II, Cycloviolacin O14, or sunflower trypsin inhibitor-I. The construction of these sequences is described in Example 1.
[0102] Use the following pH values and product codes for the release of cyclic products:
[0103] Collect samples of fractions from the entire production process and analyze them by tricine gel electrophoresis, as Figures 7 to 9 shown. Samples of all intermediate supernatants are studied by gel electrophoresis. Samples a) to e) shown on the gel are the following fractions: a) Sample of uninduced E. coli culture (no inclusion body production) b) Sample of induced E. coli culture (inclusion body production) c) Supernatant of E. coli culture after separation of the lysate supernatant and inclusion body precipitate. d) Inclusion body precipitate in denaturing buffer e) Protein solution after refolding from inclusion bodies and binding to the purification domain
[0104] Figure 7 Shows products PcGFP, Pc-1, and Pc2 in the fusion polypeptide according to SEQ ID No.: 39.
[0105] Figure 8Shows the products PcGFP, Pc-1, and Pc2 in the fusion polypeptide according to SEQ ID No.: 40.
[0106] Figure 9 Shows the products PcGFP, Pc-1, Pc2, and Pc3 in the fusion polypeptide according to SEQ ID No.: 42.
[0107] As can be seen from the sizes of the proteins detected in the tricine gel, all fusion polypeptides carrying the desired target protein can be produced as inclusion bodies and can thus be used in the purification method according to the invention, where the product is released in its circular form. Example 3 - Binding experiment
[0108] For samples containing the fusion polypeptides according to SEQ ID No.: 29 and 36 with circular GFP as the target protein, the samples for experiments in the presence of starch were carried out at pH 6.5 to pH 9.0. Similar experiments were carried out using the fusion polypeptides according to SEQ ID No.: 39 to 43 carrying circular GFP (SEQ ID NO.: 45) as the product.
[0109] The starch was taken from a stock mixture of 40% corn starch and 60% potato starch. The starch stock was resuspended in a buffer of 50 mM Tris, 100 mM NaCl, and 1 mM EDTA. The starch mixture was vortexed and 200 μL of the mixture was pipetted into a 2 mL reaction vessel. The vessel was centrifuged in a microcentrifuge at 13,000 rpm and room temperature for 2 minutes.
[0110] The supernatant was removed and replaced with 100 μL of the respective fusion polypeptide sample without detergent at pH 11.5. The sample vessel was vortexed and stirred at room temperature (T = 24.5 °C) for ten minutes. Then, 900 μL of buffer (Table 11, pH 7.2) was added to the respective reaction vessel containing starch and the fusion polypeptide. The mixture was now stirred as described above. Then, all samples were centrifuged in a microcentrifuge at 13,000 rpm and room temperature for 2 minutes. For each sample, the supernatant was precipitated in 1 mL of ethanol, vortexed, and the precipitate was pelleted by centrifugation. The precipitate was resuspended in 50 μL of Laemmli SDS sample loading buffer and incubated at 66 °C for 10 minutes. The starch precipitate was extracted with 50 μL of Laemmli SDS sample loading buffer and the starch precipitate was incubated at 66 °C for 10 minutes. Then the samples were loaded onto a 4% to 20% acrylamide gradient TGX stain-free polyacrylamide gel and compared with 1 μL of the Dual Xtra polypeptide standard (Marker) from Bio-Rad.
Claims
1. A fusion polypeptide comprising or consisting of, in the direction from the N-terminus to the C-terminus: (i) a split C-intein domain, (ii) an autoprotease domain, (iii) a target polypeptide sequence, (iv) a split N-intein domain, and (v) a purification domain embedded in the split N-intein domain after aspartic acid at position 71 of the split N-intein domain, wherein the purification domain (v) binds to a carbohydrate matrix.
2. The fusion polypeptide according to claim 1, wherein the purification domain comprises or consists of the following: An amino acid sequence according to SEQ ID No.: 1 to SEQ ID No.: 3 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 1 to SEQ ID No.:
3.
3. The fusion polypeptide according to claim 1 or 2, wherein the self-protease domain (ii) comprises or consists of the following: An amino acid sequence according to SEQ ID No.: 8 to SEQ ID No.: 12 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 8 to SEQ ID No.:
12.
4. The fusion polypeptide according to any one of the preceding claims, wherein the fusion polypeptide further comprises an inclusion body sequence selected from SEQ ID No.: 13 to SEQ ID No.: 15 or an amino acid sequence having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 13 to SEQ ID No.:
15.
5. The fusion polypeptide according to any one of the preceding claims, wherein the split C-intein domain (i) and the split N-intein domain (ii) are derived from a naturally occurring DnaE split intein sequence, preferably from the organism Synechocystes sp. PCC6803 or Nostoc punctiforme.
6. The fusion polypeptide according to any one of the preceding claims, wherein the split C-intein domain (i) and the split N-intein domain (ii) comprise or consist of the following: A pair of amino acid sequences selected from SEQ ID No.: 17 and SEQ ID No.: 18, SEQ ID No.: 19 and SEQ ID No.: 20, SEQ ID No.: 21 and SEQ ID No.: 22, SEQ ID No.: 23 and SEQ ID No.: 24, SEQ ID No.: 25 and SEQ ID No.: 26, SEQ ID No.: 27 and SEQ ID No.: 28, or a pair of amino acid sequences having 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with SEQ ID No.: 17 and SEQ ID No.: 18, SEQ ID No.: 19 and SEQ ID No.: 20, SEQ ID No.: 21 and SEQ ID No.: 22, SEQ ID No.: 23 and SEQ ID No.: 24, SEQ ID No.: 25 and SEQ ID No.: 26, SEQ ID No.: 27 and SEQ ID No.: 28, wherein the first SEQ ID No. relates to the C-intein domain (i) and the second SEQ ID No. relates to the N-intein domain (ii).
7. A recombinant nucleic acid molecule encoding the fusion polypeptide according to any one of the preceding claims.
8. A genetically modified cell comprising the recombinant nucleic acid molecule according to claim 7, wherein the cell is capable of expressing the fusion polypeptide according to any one of claims 1 to 6.
9. The genetically modified cell according to claim 8, wherein the cell is selected from: Escherichia coli, Vibrio natrigens, Saccheromyces cerevisiae, Aspergillus niger, green algae, microalgae, HEK T293 and Chinese hamster ovary cells (CHO).
10. A method for preparing a target peptide, comprising the steps of: (a) providing the genetically modified cell according to claim 8 or 9, (b) culturing the cell under conditions suitable for expressing the fusion polypeptide according to any one of claims 1 to 6, (c) obtaining the fusion polypeptide and optionally, unfolding the obtained fusion polypeptide and refolding the fusion polypeptide directionally, (d) contacting the fusion polypeptide obtained in step (c) with a carbohydrate matrix, (e) cyclizing the fusion polypeptide through the split C-intein domain (i) and the split N-intein domain (ii), (f) cleaving the fusion polypeptide by activating the autocatalytic domain of the fusion polypeptide to obtain a cyclic target peptide, (g) collecting the mixture containing the cyclic target peptide.
11. The method according to claim 10, wherein the carbohydrate matrix in step (d) consists of or comprises a substance selected from the group consisting of: starch, lignin, carbohydrate polymers, copolymers having α-1,4-glycosidic bonds and α-1,6-glycosidic bonds of glucose or other sugars, and mixtures thereof, and the carbohydrate matrix is preferably present in the packed column as a packed matrix or as starch granules composed of amylose and amylopectin.
12. The method according to claim 10 or 11, wherein the cyclization of the fusion polypeptide by the split C-intein domain (i) and the split N-intein domain (ii) in step (e) is carried out at a pH higher than 7.5, preferably at a pH higher than 9.
0.
13. The method according to any one of claims 10 to 12, wherein the activation of the self-protease domain in step (e) is carried out at a pH of 6 to 8, preferably at a pH of 6.5 to 7.5, and particularly preferably at a pH of 7 to 7.
4.
14. A recombinant nucleic acid molecule encoding a fusion polypeptide according to any one of claims 1 to 6 and a cloning site for incorporating the recombinant nucleic acid molecule according to claim 7, optionally operably linked to an expression control sequence.
15. A mixture comprising or consisting of the following: A cyclic target peptide, preferably a synthetic cyclic target peptide, and a total amount of 0.001 wt% to 1 wt% of sodium and / or potassium, based on the total weight of sodium (if present), potassium (if present), and the target peptide, wherein the mixture is obtained by or capable of being obtained by the method according to claims 10 to 13.
Citation Information
Patent Citations
Biological synthesis of amino acid chains for preparation of peptides and proteins
WO2019138125A1