Methods and compositions related to enhanced protein purification systems
The method stabilizes intein complexes with mutated N-terminal segments and Cognate Binding Partners for controlled cleavage, addressing inconsistent cleavage rates in intein-based purification systems, enhancing protein purification efficiency and yield.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OHIO STATE INNOVATION FOUND
- Filing Date
- 2026-01-15
- Publication Date
- 2026-07-23
AI Technical Summary
Existing intein-based protein purification systems struggle with inconsistent cleavage rates for different proteins, leading to unacceptably long process times for some targets, and lack a reliable method for selective purification across a broader range of proteins.
A method involving a stable intein system with mutated N-terminal intein segments covalently immobilized to a solid support, using a Cognate Binding Partner to stabilize the intein complex, allowing controlled cleavage under specific conditions.
Enhances cleavage rate and stability, enabling efficient purification of a wider variety of proteins with controlled cleavage kinetics, optimizing process times and yields.
Smart Images

Figure IMGF000037_0001 
Figure IMGF000038_0001 
Figure IMGF000040_0001
Abstract
Description
Attorney Docket No: 103362-056WO1METHODS AND COMPOSITIONS RELATED TO ENHANCED PROTEIN PURIFICATION SYSTEMSREFERENCE TO SEQUENCE LISTING
[0001] A Sequence Listing conforming to the rules of WIPO Standard ST.26 is hereby incorporated by reference. Said Sequence Listing has been filed as an electronicdocument via PatentCenter encoded as XML in UTF-8 text. The electronic document, created on January 8, 2026, is entitled “103362-056WO1 ST26.xml”, and is 56,006 bytes in size.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims benefit of U.S. Provisional Application No. 63 / 745,494, filed January 15, 2025, incorporated herein by reference in its entirety.BACKGROUND
[0003] Inteins are naturally occurring, self-splicing protein subdomains that are capable of excising out their own protein subdomain from a larger protein structure while simultaneously joining the two formerly flanking peptide regions (“exteins”) together to form a mature host protein.
[0004] The ability of inteins to rearrange flanking peptide bonds and retain activity when in fusion to proteins other than their native exteins, has led to a number of intein-based applications in various fields of biotechnology. These include several types of protein ligation and activation applications, as well as protein labeling and tracing applications. An important application of inteins is in the production of purified recombinant proteins. In particular, inteins have the ability to impart self-cleaving activity to a number of conventional affinity and purification tags, and thus provide a major advance in the production of recombinant protein products for research, medical and other commercial applications.
[0005] Conventional purification tags provide a simple and robust means for purifying any- tagged target protein and are commonly added to desired target proteins through simple genetic fusions. These tags are now ubiquitous in research and have formed a major platform for research and manufacturing of these important products. Once the tagged target protein is expressed in an appropriate host cell and purified via the tag, however, the presence of the tag on the purified target can lead to compromised activity, and potentially unwanted immunogenicityAttorney Docket No: 103362-056WO1 in the case of therapeutic proteins. For these reasons, the ability to remove the affinity tag after purification is of critical importance in many applications, which is conventionally done through the addition of highly specific endopeptidase enzymes. Although these enzymes are generally effective, they are too expensive to scale up for manufacturing, and their use requires an additional step for their removal.
[0006] Thus, the ability of inteins to impart self-cleaving activity to conventional tags is a significant advance, and early implementations of intein-based self-cleaving affinity tag systems have been published in several patents and hundreds of journal papers in the biological sciences. Despite their strength, however, several substantial weaknesses remain that inhibit the full implementation of intein methods. In particular, the ability to tightly control the cleaving reaction in a variety of highly relevant contexts has been elusive. In order to be useful, the intein self-cleaving reaction must be tightly suppressed during protein expression and purification, but very rapid once the tagged target protein is pure. Of the two initially available classes of conventional inteins, one is highly controllable and is triggered to cleave by the addition of thiol compounds, while the other is more loosely controlled and is triggered by small changes in pH and temperature. To beter control the activity of pH sensitive inteins, split intein systems have been developed where the intein is initially expressed as two separate segments that can only achieve cleaving capability after assembly. In these systems, one segment of the intein is affixed to a solid support, while the other acts as a tag in fusion to the target protein. In at least one of these systems, the assembled intein exhibits cleaving that is highly sensitive to pH, allowing assembly of the intein and purification of the target, followed by controlled release of the target protein from the affinity resin (Prabhala 2022).
[0007] Split intein systems have been effective in controlling premature cleavage and maximizing yields of their target proteins relative to contiguous intein systems, but some challenges remain. A significant challenge is that some target proteins cleave more quickly or slowly than others, requiring a degree of optimization when first applying an intein purification system to a specific target. Although it is possible to roughly predict the cleavage rate of different proteins, usually based on the identities of their first two amino acids (Prabhala 2022), some proteins cleave at an inconveniently slow rate. In these cases, process times can become unacceptably long, and the split intein system cannot be used for that target.
[0008] Therefore, what is needed is a more reliable method for the selective purification of a broader range of proteins using a stable, transformative intein system, wherein the intein system comprises mutants which can enhance the process, such as an increase in cleavage rate (tag self-Attorney Docket No: 103362-056WO1 removal) for currently difficult proteins, while still remaining strongly sensitive under permissive conditions, such as pH.SUMMARY
[0009] In accordance with the purpose(s) of the invention, as embodied and broadly described herein, the invention, in one aspect, relates to a method of stabilizing an N-Intein Ligand during expression and purification, purifying the N-Intein Ligand, and immobilizing the N-Intein Ligand to a solid support. In particular, disclosed is a method comprising: forming a soluble and stable intein complex via assembly of the N-Intein Ligand with a Cognate Binding Partner (e.g., a corresponding C-terminal intein segment; alone or in fusion to a cleavable or non-cleavable fusion partner); purifying the intein complex; and immobilizing the intein complex to a solid support. The intein complex can then be subjected to conditions that disrupt association between the N-Intein Ligand and the cognate binding partner; and the solid support washed to remove non-bound Cognate Binding Partner; and conditions provided that allow the N-Intein Ligand to fold into an active state.
[0010] The Cognate Binding Partner can comprise a C-terminal intein (INTc) segment that binds an N-Intein Ligand to induce a structured, soluble intein complex. The N-Intein Ligand and the Cognate Binding Partner can be co-expressed either in vivo in a single cell from a single plasmid or two-plasmid system, or in trans (expressed in separate cells) and mixed before or during the purification process. Such immobilization can take place onto a solid support, such as chromatographic media, a membrane, or a magnetic bead. In one example, the chromatographic media can be a solid chromatographic resin backbone.
[0011] Utilizing a Cognate Binding Partner to stabilize the N-Intein Ligand renders the N- Intein Ligand incapable of binding any other INTc segment. Therefore, following immobilization, the N-terminal intein segment (INTN) must be denatured or otherwise dissociated from the Cognate Binding Partner, allowing the Cognate Binding Partner to be removed, washed, or “stripped” away from the N-Intein Ligand. Once the Cognate Binding Partner is removed, the immobilized N-Intein Ligand must be reverted to an active state (capable of binding new partner), thereby forming a functional affinity capture medium.
[0012] Disclosed is a method for manufacturing an affinity medium comprising an N-Intein Ligand covalently bound to a convenient substrate, as well as compositions related to the manufacturing process. The N-Intein Ligand can comprise an internal INTN along with operably linked fusion partners. The INTN segment within the N-Intein Ligand can been derived from a native intein such as the Npu DnaE intein. The INTN segment may further be modified toAttorney Docket No: 103362-056WO1 increase its utility, as outlined herein. For example, a tag can be attached to the INTN segment within a region following the C-terminal residue of the INTN segment so as to aid in purification, detection, and / or enhancement of soluble expression of the N-Intein Ligand. The N-Intein Ligand can also comprise amino acids within a region following the C-terminal residue of the INTN segment, which allows for covalent immobilization of the N-Intein Ligand onto a substrate. The N-Intein Ligand can further comprise a sensitivity-enhancing motif, which renders its cleaving activity highly sensitive to extrinsic conditions. The sensitivity-enhancing motif can be in fusion to the N-terminus of the INTN segment. The extrinsic condition can be pH, temperature, zinc ion concentration, or a combination of these.
[0013] Also disclosed is an expression vector comprising exogenous nucleic acid, wherein the exogenous nucleic acid encodes an N-Intein Ligand and a Cognate Binding Partner, wherein the N-Intein Ligand can be encoded to be expressed with a purification tag, and wherein the Cognate Binding Partner may not be encoded for expression in fusion with a desired protein of interest. Also disclosed is a two-plasmid system wherein the N-Intein Ligand and Cognate Binding Partner are encoded on two distinct compatible plasmids housed within a single cell. Also disclosed is a cell comprising the expression vector(s). The Cognate Binding Partner can be encoded to be expressed in fusion to a protein or peptide that is not a desired protein of interest, such as an affinity tag.
[0014] Specifically, disclosed herein is a protein purification system, wherein the system comprises a split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment; wherein the N-terminal intein segment is capable of covalent immobilization onto a solid support, and wherein the C-terminal intein segment is capable of attachment to a protein of interest; and further wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q or C59L; wherein the mutation increases yield of the protein of interest, maintains or increases cleavage rate, increases stability or binding kinetics of the protein purification system, or provides increased sensitivity to extrinsic conditions compared to a native intein.
[0015] Also disclosed is an amino acid sequence comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L.
[0016] Further disclosed is a method of purifying a protein of interest, the method comprising: utilizing a split intein comprising two separate peptides: an N-terminal inteinAttorney Docket No: 103362-056WO1 segment and a C -terminal intein segment, wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L; covalently immobilizing the N-terminal intein segment to a solid support; attaching a protein of interest to the C -terminal intein segment, wherein cleaving of the C -terminal intein segment is more sensitive to extrinsic conditions compared to a native intein; exposing the N-terminal intein segment and the C-terminal intein segment to each other so that they associate; washing the solid support to remove non-bound material; placing the associated N-terminal intein segment and the C-terminal intein segment under conditions that allow for the intein to self-cleave; and isolating the protein of interest.
[0017] While aspects of the present invention can be described and claimed in a particular statutory class, such as the system statutory class, this is for convenience only and one of skill in the art will understand that each aspect of the present invention can be described and claimed in any statutory class. Unless otherwise expressly stated, it is in no way intended that any method or aspect set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not specifically state in the claims or descriptions that the steps are to be limited to a specific order, it is no way intended that an order be inferred, in any respect. This holds for any possible non-express basis for interpretation, including matters of logic with respect to arrangement of steps or operational flow, plain meaning derived from grammatical organization or punctuation, or the number or type of aspects described in the specification.BRIEF DESCRIPTION OF THE FIGURES
[0018] The accompanying figures, which are incorporated in and constitute a part of this specification, illustrate several aspects and together with the description serve to explain the principles of the invention.
[0019] Figure 1 shows position 28 mutated to all amino acids except proline and cysteine, where cysteine is the native amino acid appearing in this position. Intracellular cleaving during cytoplasmic expression in E. coli was then measured relative to the S28 control. In this case, both intein segments were expressed in the same cell, where the N-terminal segment is mutated, and the C-terminal segment is fused to FFN-sfGFP (sfGFP with the first three amino acids converted to FFN) as a surrogate target protein. Cell lysates then show the extent of cleaving during expression at 37°C for 4 hours with 0.5mM IPTG.Attorney Docket No: 103362-056WO1
[0020] Figure 2A-C shows analysis of the statistical significance of each mutation on cleavage rate, where the gels shown in Figure 1 were scanned and cleaving was quantified using ImageJ. Statistical analyses and results are indicated in the figure.
[0021] Figure 3 shows position 59 mutated to all amino acids except proline and cysteine, where cysteine is the native amino acid appearing in this position. Intracellular cleaving during cytoplasmic expression in E. coli was then measured relative to the S59 control. In this case, both intein segments were expressed in the same cell, where the N-terminal segment is mutated, and the C -terminal segment is fused to FFN-sfGFP (sfGFP with the first three amino acids converted to FFN) as a surrogate target protein. Cell lysates then show the extent of cleaving during expression at 37°C for 4 hours and 0.5mM IPTG.
[0022] Figure 4A-C shows analysis of the statistical significance of each mutation on cleavage rate, where the gels shown in Figure 3 were scanned and cleaving was quantified using ImageJ. Statistical analyses and results are indicated in the figure.
[0023] Figure 5 shows a general schematic for split inteins (SEQ ID NOS: 32-34 are shown from top to bottom).
[0024] Figure 6 shows a schematic of column cleaving test as described in Example 8.
[0025] Figure 7A-B shows (A) Representative data of cleaving performance and pH sensitivity for mutants with glycine in the 28 position and variable amino acids at the 59 position. The gel labeled SS at bottom indicates the time points on the other gel images. (B) Band intensities were measured by scanning densitometry and fractional cleavage at each time point was plotted against time on a semi-log plot to directly calculate cleavage kinetic constant for each mutant under each pH condition.
[0026] Figure 8A-B shows (A) Representative data of cleaving performance and pH sensitivity for mutants with histidine in the 28 position and variable amino acids at the 59 position. The gel labeled SS at bottom indicates the time points on the other gel images. (B) Band intensities were measured by scanning densitometry and fractional cleavage at each time point was plotted against time on a semi-log plot to directly calculate cleavage kinetic constant for each mutant under each pH condition.
[0027] Figure 9A-B shows (A) Representative data of cleaving performance and pH sensitivity for mutants with isoleucine in the 28 position and variable amino acids at the 59 position. The gel labeled SS at bottom indicates the time points on the other gel images. (B) Band intensities were measured by scanning densitometry and fractional cleavage at each time point was plotted against time on a semi-log plot to directly calculate cleavage kinetic constant for each mutant under each pH condition.Attorney Docket No: 103362-056WO1
[0028] Figure 10A-B shows ( A) Rpresentative data of cleaving performance and pH sensitivity for mutants with glutamic acid in the 28 position and variable amino acids at the 59 position. The gel labeled SS at bottom indicates the time points on the other gel images. (B) Band intensities were measured by scanning densitometry and fractional cleavage at each time point was plotted against time on a semi-log plot to directly calculate cleavage kinetic constant for each mutant under each pH condition.
[0029] Figure 11A-B shows quantifying cleavage kinetics at pH 6.2 on an AKTA FPLC. (A) The tail-end of the elution peak, boxed in red, follows an exponential decay curve characteristic of first-order kinetics. (B) A plot of ln(mAU) vs. time yields a linear graph in which slope = -A where k is the cleavage rate constant in min’1.DESCRIPTION
[0030] The present invention can be understood more readily by reference to the following detailed description of the invention and the Examples included therein.
[0031] Before the present compounds, compositions, articles, systems, devices, and / or methods are disclosed and described, it is to be understood that they are not limited to specific synthetic methods unless otherwise specified, or to particular reagents unless otherwise specified, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, example methods and materials are now described.
[0032] All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided herein can be different from the actual publication dates, which can require independent confirmation.A, DEFINITIONS
[0033] As used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. Thus, forAttorney Docket No: 103362-056WO1 example, reference to “a functional group,” “an alkyl,” or “a residue” includes mixtures of two or more such functional groups, alkyls, or residues, and the like.
[0034] Ranges can be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms a further aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. For example, if the value “10” is disclosed, then “about 10” is also disclosed. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
[0035] A weight percent (wt. %) of a component, unless specifically stated to the contrary, is based on the total weight of the formulation or composition in which the component is included.
[0036] As used herein, the terms “optional” or “optionally” means that the subsequently described event or circumstance can or can not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.
[0037] The term “contacting” as used herein refers to bringing two biological entities together in such a manner that the compound can affect the activity of the target, either directly; i.e., by interacting with the target itself, or indirectly; i.e., by interacting with another molecule, co-factor, factor, or protein on which the activity of the target is dependent. “Contacting” can also mean facilitating the interaction of two biological entities, such as peptides, to bond covalently or otherwise.
[0038] As used herein, “kit” means a collection of at least two components constituting the kit. Together, the components constitute a functional unit for a given purpose. Individual member components may be physically packaged together or separately. For example, a kit comprising an instruction for using the kit may or may not physically include the instruction with other individual member components. Instead, the instruction can be supplied as a separate member component, either in a paper form or an electronic form which may be supplied on computer readable memory device or downloaded from an internet website, or as recorded presentation.
[0039] As used herein, “instruction(s)” means documents describing relevant materials or methodologies pertaining to a kit. These materials may include any combination of theAttorney Docket No: 103362-056WO1 following: background information, list of components and their availability information (purchase information, etc.), brief or detailed protocols for using the kit, troubleshooting, references, technical support, and any other related documents. Instructions can be supplied with the kit or as a separate member component, either as a paper form or an electronic form which may be supplied on computer readable memory device or downloaded from an internet website, or as recorded presentation. Instructions can comprise one or multiple documents and are meant to include future updates.
[0040] As used herein, the terms “target protein,” “protein of interest” and “therapeutic agent” include any synthetic or naturally occurring protein or peptide. In the context of this invention, a “protein of interest” is a protein that is to be purified using split intein purification technology by an end user in a laboratory or manufacturing setting, as opposed to any context related to the manufacture of the purification medium itself. This definition would apply to any protein or peptide requiring purification for study or other research applications. The term additionally encompasses those compounds traditionally regarded as drugs, vaccines, and biopharmaceuticals including molecules such as proteins, peptides, and the like. Examples of therapeutic agents are described in well-known literature references such as the Merck Index (14th edition), the Physicians' Desk Reference (64th edition), and The Pharmacological Basis of Therapeutics (1st edition), and they include, without limitation, medicaments; substances used for the treatment, prevention, diagnosis, cure or mitigation of a disease or illness; substances that affect the structure or function of the body, or pro-drugs, which become biologically active or more active after they have been placed in a physiological environment.
[0041] As used herein, “variant” refers to a molecule that retains a functional activity that is the same or substantially like that of the original sequence. The variant may be from the same or distinct species or be a synthetic sequence based on a natural or prior molecule. Moreover, as used herein, “variant” refers to a molecule having a structure attained from the structure of a parent molecule (e.g., a protein or peptide disclosed herein) and whose structure or sequence is sufficiently similar to those disclosed herein that based upon that similarity, would be expected by one skilled in the art to exhibit the same or similar activities and utilities compared to the parent molecule. For example, substituting specific amino acids in a given peptide can yield a variant peptide with similar activity to the parent.
[0042] As used herein, “mutant” refers to a molecule that has been modified in a way which confers a different structural and / or functional characteristic to the molecule. This can be a modification to the nucleic acid which affects the amino acid sequence, or, in the case of protein synthesis, can be a direct modification to an amino acid sequence. This modification can resultAttorney Docket No: 103362-056WO1 in a difference in function between the mutant and the original, or naturally occurring, molecule. This modification can confer benefits to the mutant which are not seen in the original, or naturally occurring, molecule.
[0043] As used herein, the term “amino acid sequence” refers to a list of abbreviations, letters, characters, or words representing amino acid residues. The amino acid abbreviations used herein are conventional one letter codes for the amino acids and are expressed as follows: A, alanine; C, cysteine; D aspartic acid; E, glutamic acid; F, phenylalanine; G, glycine; H histidine; I isoleucine; K, lysine; L, leucine; M, methionine; N, asparagine; P, proline; Q, glutamine; R, arginine; S, serine; T, threonine; V, valine; W, tryptophan; Y, tyrosine.
[0044] “Peptide” as used herein refers to any peptide, oligopeptide, polypeptide, gene product, expression product, or protein. A peptide is comprised of consecutive amino acids. The term “peptide” encompasses naturally occurring or synthetic molecules.
[0045] In addition, as used herein, the term “peptide” refers to amino acids joined to each other by peptide bonds or modified peptide bonds, e.g., peptide isosteres, etc. and may contain modified amino acids other than the twenty gene-encoded amino acids. The peptides can be modified by either natural processes, such as post-translational processing, or by chemical modification techniques which are well known in the art. Modifications can occur anywhere in the peptide, including the peptide backbone, the amino acid side-chains and the amino or carboxyl termini. The same type of modification can be present in the same or varying degrees at several sites in a given polypeptide. Also, a given peptide can have many types of modifications. Modifications include, without limitation, linkage of distinct domains or motifs, acetylation, acylation, ADP-ribosylation, amidation, covalent cross-linking or cyclization, covalent attachment of flavin, covalent attachment of a heme moiety, covalent attachment of a nucleotide or nucleotide derivative, covalent attachment of a lipid or lipid derivative, covalent attachment of a phosphytidylinositol, disulfide bond formation, demethylation, formation of cysteine or pyroglutamate, formylation, gamma-carboxylation, glycosylation, GPI anchor formation, hydroxylation, iodination, methylation, myristolyation, oxidation, pergylation, proteolytic processing, phosphorylation, prenylation, racemization, selenoylation, sulfation, and transfer-RNA mediated addition of amino acids to protein such as arginylation. (See Proteins — Structure and Molecular Properties 2nd Ed., T. E. Creighton, W. H. Freeman and Company, New York (1993); Posttranslational Covalent Modification of Proteins, B. C. Johnson, Ed., Academic Press, New York, pp. 1-12 (1983)).
[0046] As used herein, “isolated peptide” or “purified peptide” is meant to mean a peptide (or a fragment thereof) that is substantially free from the materials with which the peptide isAttorney Docket No: 103362-056WO1 normally associated in nature, or from the materials with which the peptide is associated in an artificial expression or production system, including but not limited to an expression host cell lysate, growth medium components, buffer components, cell culture supernatant, or components of a synthetic in vitro translation system. The peptides disclosed herein, or fragments thereof, can be obtained, for example, by extraction from a natural source (for example, a mammalian cell), by expression of a recombinant nucleic acid encoding the peptide (for example, in a cell or in a cell-free translation system), or by chemically synthesizing the peptide. In addition, peptide fragments may be obtained by any of these methods, or by cleaving full length proteins and / or peptides.
[0047] The word “or” as used herein means any one member of a particular list and also includes any combination of members of that list.
[0048] The phrase “nucleic acid” as used herein refers to a naturally occurring or synthetic oligonucleotide or polynucleotide, whether DNA or RNA or DNA-RNA hybrid, single-stranded or double-stranded, sense or antisense, which is capable of hybridization to a complementary nucleic acid by Watson-Crick base-pairing. Nucleic acids of the invention can also include nucleotide analogs (e.g., BrdU), and non-phosphodi ester internucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids can include, without limitation, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA or any combination thereof.
[0049] As used herein, “isolated nucleic acid” or “purified nucleic acid” is meant to mean DNA that is free of the genes that, in the naturally occurring genome of the organism from which the DNA of the invention is derived, flank the gene. The term therefore includes, for example, a recombinant DNA which is incorporated into a vector, such as an autonomously replicating plasmid or virus; or incorporated into the genomic DNA of a prokaryote or eukaryote (e.g., a transgene); or which exists as a separate molecule (for example, a cDNA or a genomic or cDNA fragment produced by PCR, restriction endonuclease digestion, or chemical or in vitro synthesis). It also includes a recombinant DNA which is part of a hybrid gene encoding additional polypeptide sequences. The term “isolated nucleic acid” also refers to RNA, e.g., an mRNA molecule that is encoded by an isolated DNA molecule, or that is chemically synthesized, or that is separated or substantially free from at least some cellular components, for example, other types of RNA molecules or peptide molecules.
[0050] “Intern” refers to an in-frame intervening sequence in a protein as described by Perler (Perler, Davis et al. 1994). An intein can catalyze its own excision from the protein through a post-translational protein splicing process to yield the free intein and a mature protein. An intein can also catalyze the cleavage of the intein-extein bond at either the intein N-terminus, or theAttorney Docket No: 103362-056WO1 intein C-terminus, or both of the intein-extein termini. As used herein, “intein” encompasses mini-inteins, modified or mutated inteins, and split inteins.
[0051] A "split intein" is an intein that is comprised of two or more separate components not fused to one another. Split inteins can occur naturally or can be engineered by splitting contiguous inteins.
[0052] As used herein, the term "splice" or "splices" means to excise a central portion of a polypeptide to form two or more smaller polypeptide molecules. In some cases, splicing also includes the step of fusing together two or more of the smaller polypeptides to form a new polypeptide. Splicing can also refer to the joining of two polypeptides encoded on two separate gene products through the action of a split intein.
[0053] As used herein, the terms "cleave", "cleaves", ‘"cleavage” and “a cleaving event” refer to a chemical reaction in which a peptide bond within a polypeptide is broken, thereby dividing a single polypeptide to form two or more smaller polypeptide molecules. In some cases, cleavage is mediated by the addition of an extrinsic endopeptidase, which is often referred to as “proteolytic cleavage.” In other cases, cleaving can be mediated by the intrinsic activity of one or both of the cleaved peptide sequences, which is often referred to as “self-cleavage ” Cleavage can be controlled by extrinsic conditions (such as buffer pH), as in the action of the split intein system described herein.
[0054] By the term "fused" or “in fusion with” is meant covalently bonded to. For example, a first peptide is fused to a second peptide when the two peptides are covalently bonded to each other (e.g., via a peptide bond). Peptides and / or protein domains conjoined by peptide bonds may also be referred to as “fusion partners.”
[0055] As used herein an "isolated" or "substantially pure" substance is one that has been separated from components which naturally accompany it. Typically, a polypeptide is substantially pure when it is at least 50% (e.g., 60%, 70%, 80%, 90%, 95%, and 99%) by weight free from the other proteins and naturally occurring organic molecules with which it is naturally associated.
[0056] Herein, "bind", "binds", “binding” or “binding event” means that one molecule recognizes and adheres to another molecule in a sample but does not substantially recognize or adhere to other molecules in the sample. The terms “bind,” “binds”, “binding” and “binding event” also imply the interaction between two molecules is non-covalent and reversible. One molecule "specifically binds" another molecule if it has a binding affinity greater than about 105to 10” liters / mole for the other molecule. These terms are used interchangeably with “associate with,” “associates with,” or “associating with.”Attorney Docket No: 103362-056WO1
[0057] Nucleic acids, nucleotide sequences, proteins or amino acid sequences referred to herein can be isolated, purified, synthesized chemically, or produced through recombinant DNA technology. All of these methods are well known in the art.
[0058] As used herein, the terms “modified” or “mutated,” as in “modified intein” or “mutated intein,” refer to one or more modifications in either the nucleic acid or amino acid sequence being referred to, such as an intein, when compared to the native, or naturally occurring structure. Such modification can be a substitution, addition, or deletion. The modification can occur in one or more amino acid residues or one or more nucleotides of the structure being referred to, such as an intein.
[0059] As used herein, “operably linked” refers to the association of two or more biomolecules in a configuration relative to one another such that the normal function of the biomolecules can be performed. In relation to nucleotide sequences, “operably linked” refers to the association of two or more nucleic acid sequences, by means of enzymatic ligation or otherwise, in a configuration relative to one another such that the normal function of the sequences can be performed. For example, the nucleotide sequence encoding a pre-sequence or secretory leader is operably linked to a nucleotide sequence for a polypeptide if it is expressed as a pre-protein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the coding sequence; and a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation of the sequence.
[0060] “Sequence homology” can refer to the situation where nucleic acid or protein sequences are similar because they have a common evolutionary origin. “Sequence homology” can indicate that sequences are very similar. Sequence similarity is observable; homology can be based on the observation. “Very similar” can mean at least 70% identity, homology or similarity; at least 75% identity, homology or similarity; at least 80% identity, homology or similarity; at least 85% identity, homology or similarity; at least 90% identity, homology or similarity, such as at least 93% or at least 95% or even at least 97% identity, homology or similarity. The nucleotide sequence similarity or homology or identity can be determined using the “Align” program of Myers et al. (1988) CABIOS 4:11-17 and available at NCBI Additionally, or alternatively, amino acid sequence similarity or identity or homology can be determined using the BlastP program (Altschul et al. Nucl. Acids Res. 25:3389-3402), and available at NCBI. Alternatively or additionally, the terms “similarity” or “identity” or “homology,” for instance, with respect to a nucleotide sequence, are intended to indicate a quantitative measure of homology between two sequences.Attorney Docket No: 103362-056WO1
[0061] Alternatively or additionally, “similarity” with respect to sequences refers to the number of positions with identical nucleotides divided by the number of nucleotides in the shorter of the two sequences wherein alignment of the two sequences can be determined in accordance with the Wilbur and Lipman algorithm. (1983) Proc. Natl. Acad. Sci. USA 80:726. For example, using a window size of twenty nucleotides, a word length of 4 nucleotides, and a gap penalty of 4, and computer-assisted analysis and interpretation of the sequence data including alignment can be conveniently performed using commercially available programs (e.g., Intelligenetics™ Suite, Intelligenetics Inc. CA). When RNA sequences are said to be similar or have a degree of sequence identity with DNA sequences, thymidine (T) in the DNA sequence is considered equal to uracil (U) in the RNA sequence. The following references also provide algorithms for comparing the relative identity or homology or similarity of amino acid residues of two proteins, and additionally or alternatively with respect to the foregoing, the references can be used for determining percent homology, identity, or similarity. Needleman et al. (1970) J. Mol. Biol. 48:444-453; Smith et al. (1983) Advances App. Math. 2:482-489; Smith et al. (1981) Nuc. Acids Res. 11:2205-2220; Feng et al. (1987) J. Molec. Evol. 25:351-360; Higgins et al. (1989) CABIOS 5:151-153; Thompson et al. (1994) Nuc. Acids Res. 22:4673- 480; and Devereux et al. (1984) 12:387-395. “Stringent hybridization conditions” is a term which is well known in the art; see, for example, Sambrook, “Molecular Cloning, A Laboratory Manual” second ed., CSH Press, Cold Spring Harbor, 1989; “Nucleic Acid Hybridization, A Practical Approach”, Hames and Higgins eds., IRL Press, Oxford, 1985; see also FIG. 2 and description thereof herein wherein there is a sequence comparison.
[0062] The terms “plasmid” and “vector” and “cassette” refer to an extrachromosomal element often carrying genes which are not part of the central metabolism of the cell and usually in the form of circular double-stranded DNA molecules. Such elements may be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear or circular, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequence into a cell. Typically, a “vector” is a modified plasmid that contains additional multiple insertion sites for cloning and an “expression cassette” that contains a DNA sequence for a selected gene product (i.e., a transgene) for expression in the host cell. This “expression cassette” typically includes a 5' promoter region, the transgene ORF, and a 3' terminator region, with all necessary' regulatory sequences requiredAttorney Docket No: 103362-056WO1 for transcription and translation of the ORF. Thus, integration of the expression cassette into the host permits expression of the transgene ORF in the cassette.
[0063] The term "buffer" or "buffered solution" refers to solutions which resist changes in pH by the action of its conjugate acid-base range.
[0064] The term "loading buffer" or "binding buffer" refers to the buffer containing the salt or salts which is mixed with the protein preparation for loading the protein preparation onto a column. This buffer is also used to equilibrate the column before loading, and to wash to column after loading the protein.
[0065] The term "wash buffer" is used herein to refer to the buffer that is passed over a column (for example) following loading of a protein of interest (such as one coupled to a C-terminal intein fragment, for example) and prior to elution of the protein of interest. The wash buffer may serve to remove one or more contaminants without substantial elution of the desired protein.
[0066] The term "elution buffer" refers to the buffer used to elute the desired protein from the column. As used herein, the term "solution" refers to either a buffered or a non-buffered solution, including water.
[0067] The term "washing" means passing an appropriate buffer through or over a solid support, such as a chromatographic resin.
[0068] The term "eluting" a molecule (e.g., a desired protein or contaminant) from a solid support means removing the molecule from such material.
[0069] The term "contaminant" or "impurity" refers to any foreign or objectionable molecule, particularly a biological macromolecule such as a DNA, an RNA, or a protein, other than the protein being purified, that is present in a sample of a protein being purified.Contaminants include, for example, other proteins from cells that express and / or secrete the protein being purified.
[0070] The term "separate" or "isolate" as used in connection with protein purification refers to the separation of a desired protein from a second protein or other contaminant or mixture of impurities in a mixture comprising both the desired protein and a second protein or other contaminant or impurity mixture, such that at least the majority of the molecules of the desired protein are removed from that portion of the mixture that comprises at least the majority of the molecules of the second protein or other contaminant or mixture of impurities.
[0071] The term "purify" or "purifying" a desired protein from a composition or solution comprising the desired protein and one or more contaminants means increasing the degree ofAttorney Docket No: 103362-056WO1 purity of the desired protein in the composition or solution by removing (completely or partially ) at least one contaminant from the composition or solution.
[0072] Disclosed are the components to be used to prepare the compositions of the invention as well as the compositions themselves to be used within the methods disclosed herein. These and other materials are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these materials are disclosed that while specific reference of each various individual and collective combinations and permutation of these compounds can not be explicitly disclosed, each is specifically contemplated and described herein. For example, if a particular compound is disclosed and discussed and a number of modifications that can be made to a number of molecules including the compounds are discussed, specifically contemplated is each and every combination and permutation of the compound and the modifications that are possible unless specifically indicated to the contrary. Thus, if a class of molecules A, B, and C are disclosed as well as a class of molecules D, E, and F and an example of a combination molecule, A-D is disclosed, then even if each is not individually recited each is individually and collectively contemplated meaning combinations, A-E, A-F, B-D, B-E, B-F, C-D, C-E, and C-F are considered disclosed. Likewise, any subset or combination of these is also disclosed. Thus, for example, the sub-group of A-E, B-F, and C-E would be considered disclosed. This concept applies to all aspects of this application including, but not limited to, steps in methods of making and using the compositions of the invention. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific embodiment or combination of embodiments of the methods of the invention.
[0073] It is understood that the compositions disclosed herein have certain functions.Disclosed herein are certain structural requirements for performing the disclosed functions, and it is understood that there are a variety of structures that can perform the same function that are related to the disclosed structures, and that these structures will typically achieve the same result. For example, compounds used to control pH in the examples shown can be substituted with other buffering compounds to control pH, since pH is the critical variable to be controlled, and the specific buffering compounds can vary.B. PROTEIN PURIFICATION SYSTEMS AND METHODS THEREOFSPLIT INTEINS AND MODIFICATIONS THEREOFAttorney Docket No: 103362-056WO1
[0074] Disclosed herein are protein purification systems, wherein the system comprises an engineered split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment, wherein the N-terminal intein segment can be linked to a solid support, and wherein the C -terminal intein segment DNA is genetically fused to a desired target protein DNA such that the expressed target protein is tagged with the C -terminal intein segment, and wherein cleaving of the C -terminal intein segment is suppressed in the absence of the N-terminal intein segment, and wherein the N-terminal intein segment and C -terminal intein segment associate strongly to form an immobilized complex when contacted to each other on the solid support, and wherein the C-terminal cleaving activity of the C-terminal intein segment in the assembled N-terminal and C-terminal intein segment complex is highly sensitive to extrinsic conditions when compared to a native intein complex, and wherein the immobilized intein complex with the fused target protein can be purified through washing away of unimmobilized contaminants, and wherein the C-terminal intein segment can be induced to cleave and thereby release the untagged and substantially purified target protein from the immobilized intein complex through a controlled change in extrinsic conditions. The general schematic for how split inteins function can be found in Figure 5,
[0075] Specifically, disclosed herein are protein purification systems, wherein the system comprises a split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment; wherein the N-terminal intein segment is capable of covalent immobilization onto a solid support, and wherein the C-terminal intein segment is capable of attachment to a protein of interest; and further wherein the N-terminal intein segment comprises one or more mutations. These mutations can increase yield of the protein of interest, increase cleavage rate, increase stability of the protein purification system, or can render the intein more sensitive to extrinsic conditions compared to a native intein. Specific examples of these mutations include, but are not limited to, variants of SEQ ID NO: 1, such as C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L (or a combination thereof). These specific mutations can be found in SEQ ID NOS: 2-17 (see Table 2).
[0076] As used herein, the " N-terminal intein segment" refers to any intein sequence that comprises an N- terminal amino acid sequence that is functional for splicing and / or cleaving reactions when combined with a corresponding C-terminal intein segment. An N-terminal intein segment thus also comprises a sequence that is spliced out when splicing occurs. An N-terminal intein segment can comprise a sequence that is a modification of the N-terminal portion of a naturally occurring (native) intein sequence. For example, an N-terminal intein segment can comprise additional amino acid residues and / or mutated residues so long as the inclusion of suchAttorney Docket No: 103362-056WO1 additional and / or mutated residues does not render the intein non-functional for splicing or cleaving. Such modifications are discussed in more detail below. Preferably, the inclusion of the additional and / or mutated residues improves or enhances the splicing and / or cleaving activity and / or controllability of the intein. Non-intein residues can also be genetically fused to intein segments to provide additional functionality, such as the ability to be affinity purified or to be covalently immobilized. SEQ ID NOS: 1-17 and 20-27 represent N-terminal intein segments (see Table 2). SEQ ID NO: 1 is a native N terminal intein segment, and SEQ ID NOS: 2-17 and 20-27 are modifieds version of SEQ ID NO: 1. These modifications can provide unexpectedly better characteristics for the protein purification system when compared to a native intein. These unexpected characteristics are discussed in more detail below.
[0077] As used herein, the " C-terminal intein segment” refers to any intein sequence that comprises a C-terminal amino acid sequence that is functional for splicing or cleaving reactions when combined with a corresponding N-terminal intein segment. In one aspect, the C-terminal intein segment comprises a sequence that is spliced out when splicing occurs. In another aspect, the C-terminal intein segment is cleaved from a peptide sequence fused to its C -terminus. The sequence which is cleaved from the C-terminal intein’s C -terminus is referred to herein as a “protein of interest” or “target protein” and is discussed in more detail below. A C-terminal intein segment can comprise a sequence that is a modification of the C-terminal portion of a naturally occurring (native) intein sequence. For example, a C-terminal intein segment can comprise additional amino acid residues and / or mutated residues so long as the inclusion of such additional and / or mutated residues does not render the C-terminal intein segment non-functional for splicing or cleaving. Preferably, the inclusion of the additional and / or mutated residues improves or enhances the splicing and / or cleaving activity of the intein. SEQ ID NOS: 18 and 19 represent C-intein segments. SEQ ID NO: 18 is a native C intein segment, and SEQ ID NO: 19 is a modified version, wherein the alanine (“A”) at position 34 is substituted for a threonine (“T”). It is noted that in some instances, this mutation is referred to as “A33T”. It is noted that this occurs when methionine (“M”) is not considered in the first position. So “A33T” is considered the same as “A34T” when described herein, with “A33T” referring to SEQ ID NO: 19 with no methionine present.
[0078] The intein can be derived, for example, from an Npu DnaE intein. Npux refers to the N-terminal intein segment, while Npuc refers to the C-terminal intein segment. “TP” refers to the target protein (also referred to herein as the protein of interest), which can be coupled to Npuc. In Figure 3, a scheme is shown in which NpuN is coupled to a solid support, such as a commercialized resin. Npuc, while attached to a target protein, is exposed to immobilized Npux.Attorney Docket No: 103362-056WO1 The NpuN and Npuc then associate, allowing capture and purification of the fused target protein. When the proper conditions are present, a cleaving reaction can occur, which allows for cleaving (and subsequent elution) of the target protein from the immobilized intein segments.
[0079] The modifications to the inteins disclosed herein can exert a variety of beneficial effects. These benefits include, but are not limited to, an increased yield of the protein of interest, a maintained or increased cleavage rate, an increase in stability or binding kinetics of the protein purification system, or the mutations can render the intein more sensitive to extrinsic conditions compared to a native intein. These modifications and their benefits are demonstrated and discussed in detail in Examples 1-3 and Figures 1-4.
[0080] By “increased yield of the protein of interest” is meant that the amount of protein yielded when using the protein purification system with the modifications disclosed herein is increased by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, 110%, 120%, 130%, 140%, 150%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 fold or more when compared to the non-mutated, or native, version of the intein.
[0081] By “increased cleavage rate” is meant that the rate at which protein is yielded when using the protein purification system with the modifications disclosed herein is increased by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, 110%, 120%, 130%, 140%, 150%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 fold or more when compared to the non-mutated, or native, version of the intein.
[0082] By “increased in stability of the protein” is meant that the stability of the protein is increased when using the protein purification system with the modifications disclosed herein by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%,Attorney Docket No: 103362-056WO1 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, 110%, 120%, 130%, 140%, 150%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 fold or more when compared to the non-mutated, or native, version of the intein.
[0083] By ‘’increase in binding kinetics” is meant that the binding kinetics of the protein are increased when using the protein purification system with the modifications disclosed herein by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, 110%, 120%, 130%, 140%, 150%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 fold or more when compared to the non-mutated, or native, version of the intein.
[0084] By “more sensitive to extrinsic conditions” is meant that the sensitivity to extrinsic conditions (e g., temperature, salt concentration, redox potential, pH, etc.) of the protein is increased when using the protein purification system with the modifications disclosed herein by 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, 110%, 120%, 130%, 140%, 150%, or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 fold or more when compared to the non-mutated, or native, version of the intein.USES OF SPLIT INTEINS
[0085] Generally speaking, split inteins have many practical uses including the production of recombinant proteins from fragments, the circularization of recombinant proteins, and the fixation of proteins on protein chips (Scott et al., (1999); Xu et al. (2001); Evans et al. (2000);Attorney Docket No: 103362-056WO1 Kwon et al., (2006)). The advantages of intein-based protein cleavage methods, compared to others such as protease-based methods, have been noted previously (Xu et al. (2001)). The intein-based N- and C-cleavage methods can also be used together on a single target protein to produce precise and tag-free ends at both the N- and the C-termini, or to achieve cyclization of the target protein (ligation of the N- and C-termini) using the expressed protein ligation approach.
[0086] Some examples of other modifications to the N-terminal intein segment include those which alter the amino acid sequence so that it does not comprise any internal cysteine residues. This is desirous so as to eliminate side reactions associated with immobilization of the NpuN intein segment onto a solid support. For example, the residues can be mutated to serine. An example of NpuN in which the cysteine residues have been mutated can be found in SEQ ID NOS: 1-17. It is noted that the first cysteine residue which is replaced (the first amino acid on the intein N-terminus) can be replaced with either alanine or glycine so as to eliminate intein splicing in the assembled intein complex. The Npux intein can also be modified by the addition of an internal affinity tag to facilitate its purification, and the addition of specific peptides to its C-terminus to facilitate covalent immobilization onto a solid support. For example, an included His-tag was used to purify the NpuN-segment described herein, while a Cys residue was appended to the C-terminus of NpuN to facilitate covalent chemical immobilization of the NpuN-segment. The modified Npux segment that incorporates both the His-tag and Cys residue is shown by way of example in SEQ ID NO: 23.
[0087] Importantly, any tag can be used to purify the N-segment (not just His-tag), and several different immobilization chemistries can be used to immobilize the N-segment. “Tags” are more generally referred to herein as purification tags. In one example, the N-segment could be purified without a tag, using conventional chromatography. Furthermore, it is also possible to add a His-tag mutation in the C-terminal intein segment, and a “sensitivity enhancing domain” to the N-terminal intein segment.
[0088] The N-terminal intein segment can also comprise a purification, or affinity tag, attached to its C-terminus. This can include an affinity resin reagent. The purification tag can comprise, for example, one or more histidine residues. The purification tag can comprise, for example, a chitin binding domain protein with highly specific affinity for chitin. The purification tag can further comprise, for example, a reversibly precipitating elastin-like peptide tag, which can be induced to selectively precipitate under known conditions of buffer composition and temperature. Affinity tags are discussed in more detail below. The N-terminal intein segmentAttorney Docket No: 103362-056WO1 can also comprise amino acids at its C-terminus which allow for covalent immobilization. For example, one or more amino acids at the C-terminus can be cysteine residues.
[0089] The N-terminal intein segment can be immobilized onto a solid support. A variety of supports can be used. For example, the solid support can a polymer or substance that allows for immobilization of the N-terminal intein fragment, which can occur covalently or via an affinity tag with or without an appropriate linker. When a linker is used, the linker can be additional amino acid residues engineered to the C-terminus of the N-terminal intein segment or can be other known linkers for attachment of a peptide to a support.
[0090] The N-terminal intein segment disclosed herein can be attached to an affinity tag through a linker sequence. The linker sequence can be designed to create distance between the intein and affinity tag, while providing minimal steric interference to the intein cleaving active site. It is generally accepted that linkers involve a relatively unstructured amino acid sequence, and the design and use of linkers are common in the art of designing fusion peptides. There is a variety of protein linker databases which one of skill in the art will recognize. This includes those found in Argos et al. J Mol Biol 1990 Feb 20; 211(4) 943-58; Crasto et al. Protein Eng 2000 May; 13(5) 309-12; George et al. Protein Eng 2002 Nov; 15(11) 871-9; Arai et al. Protein Eng 2001 Aug; 14(8) 529-32; and Robinson et al. PNAS May 26, 1998, vol. 95 no. 11 5929-5934, hereby incorporated by reference in their entirety for their teaching of examples of linkers.
[0091] Examples of linkers which can be used with the present invention include, but are not limited to: (1) poly -asparagine linker consisting of 4 to 15 asparagine residues, and (2) glycine-serine linker, consisting of various combinations and lengths of polypeptides consisting of glycine and serine. One of skill in the art can easily identify and use any linker that will successfully link the CIPS with an affinity tag.
[0092] As mentioned above, the solid support can be a solid chromatographic resin backbone, such as a crosslinked agarose. The term "solid support matrix" or "solid matrix" refers to the solid backbone material of the resin which material contains reactive functionality permitting the covalent attachment of ligand (such as N-terminal intein segments) thereto. The backbone material can be inorganic (e.g., silica) or organic. When the backbone material is organic, it is preferably a solid polymer and suitable organic polymers are well known in the art. Solid support matrices suitable for use in the resins described herein include, by way of example, cellulose, regenerated cellulose, agarose, silica, coated silica, dextran, polymers (such as polyacrylates, polystyrene, polyacrylamide, polymethacrylamide including commercially available polymers such as Fractogel, Enzacryl, and Azlactone), copolymers (such as copolymers of styrene and divinyl- benzene), mixtures thereof and the like. Also, co-, ter- andAttorney Docket No: 103362-056WO1 higher polymers can be used provided that at least one of the monomers contains or can be derivatized to contain a reactive functionality in the resulting polymer. In an additional embodiment, the solid support matrix can contain ionizable functionality incorporated into the backbone thereof.
[0093] Reactive functionalities of the solid support matrix permitting covalent attachment of the N-terminal intein segments are well known in the art. Such groups include epoxide, cyanogen bromide, iodoacetyl, hydroxyl (e.g., Si-OH), carboxyl, thiol, amino, and the like. Conventional chemistry permits use of these functional groups to covalently attach ligands, such as N-terminal intein segments, thereto. Additionally, conventional chemistry permits the inclusion of such groups on the solid support matrix. For example, carboxy groups can be incorporated directly by employing acrylic acid or an ester thereof in the polymerization process. Upon polymerization, carboxyl groups are present if acrylic acid is employed or the polymer can be derivatized to contain carboxyl groups if an acrylate ester is employed.
[0094] Affinity tags can be peptide or protein sequences cloned in frame with protein coding sequences that change the protein's behavior. Affinity tags can be appended to the N- or C-terminus of proteins which can be used in methods of purifying a protein from cells. Cells expressing a peptide comprising an affinity tag can be pelleted, lysed, and the cell lysate applied to a column, resin or other solid support that displays a ligand to the affinity tags. The affinity tag and any fused peptides are bound to the solid support, which can also be washed several times with buffer to eliminate unbound (contaminant) proteins. A protein of interest, if attached to an affinity tag, can be eluted from the solid support via a buffer that causes the affinity tag to dissociate from the ligand resulting in a purified protein, or can be cleaved from the bound affinity tag using a soluble protease. As disclosed herein, the affinity tag is cleaved through the self-cleaving action of the NpuC intein segment in the active intein complex.
[0095] Examples of affinity tags can be found in Kimple et al. Curr Protoc Protein Sci 2004 Sep; Arnau et al. Protein Expr Purif 2006 Jul; 48(1) 1-13; Azarkan et al. J Chromatogr B Analyt Technol Biomed Life Sci 2007 Apr 15; 849(1-2) 81-90; and Waugh et al. Trends Biotechnol 2005 Jun; 23(6) 316-20, all hereby incorporated by reference in their entirety for their teaching of examples of affinity tags.
[0096] Examples of affinity include, but are not limited to, maltose binding protein, which can bind to immobilized maltose to facilitate purification of the fused target protein; Chitin binding protein, which can bind to immobilized chitin; Glutathione S transferase, which can bind to immobilized chitin; Poly-histidine, which can bind to immobilized chelated metals;FLAG octapeptide, which can bind to immobilized anti-FLAG antibodies.Attorney Docket No: 103362-056WO1
[0097] Affinity tags can also be used to facilitate the purification of a protein of interest using the disclosed modified peptides through a variety of methods, including, but not limited to, selective precipitation, ion exchange chromatography, binding to precipitation-capable ligands, dialysis (by changing the size and / or charge of the target protein) and other highly selective separation methods.
[0098] In some aspects, affinity tags can be used that do not actually bind to a ligand, but instead either selectively precipitate or act as ligands for immobilized corresponding binding domains. In these instances, the tags are more generally referred to as purification tags. For example, the ELP tag selectively precipitates under specific salt and temperature conditions, allowing fused peptides to be purified by centrifugation. Another example is the antibody Fc domain, which serves as a ligand for immobilized protein A or Protein G-binding domains.
[0099] As disclosed herein, a protein of interest (POI), or target protein can attach to the C-terminal intein segment at its C -terminus. The C-terminal segment can be genetically fused to the protein of interest, for example. Methods of recombinant protein production are known to those of skill in the art. The C-terminal intein segment can comprise modifications when compared to a native Npu DnaE C-terminal intein segment.
[0100] The N-terminal intein segment can further comprise a sensitivity-enhancing motif (SEM), which renders the splicing or cleaving activity of the assembled intein complex highly sensitive to extrinsic conditions. This sensitivity-enhancing motif can render the split intein, when assembled (meaning the C-terminal intein segment comprising the protein of interest and the N-terminal intein segment are non-covalently associated, or covalently linked), more likely to cleave under certain conditions. Therefore, the sensitivity-enhancing motif can render the split intein more sensitive to extrinsic conditions when compared to a native, or naturally occurring, intein. For example, the split intein disclosed herein can be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% or more sensitive to extrinsic conditions when compared to a native intein, where the sensitivity is defined as the percent cleavage under non-permissive conditions subtracted from the percent cleavage under permissive conditions. For example, if an intein shows no cleavage at pH 8.5 and 5 hours incubation, but 100% cleavage at pH 6.2 and 5 hours incubation, then the intein would show 100% sensitivity to this pH change under these conditions. Specifically, the Npu DnaE mutated intein disclosed herein can be more sensitive, and thus more likely to cleave the protein of interest, when certain conditions are present. These extrinsic conditions can be, for example, pH, temperature, or exposure to a certain compound or element, such as a chelating agent. Cleaving can occur at a greater rate, for example, at a pH of 6.2, when the sensitivity enhancing motif is present on the N-terminal inteinAttorney Docket No: 103362-056WO1 segment, as compared to either an N -terminal intein segment which doesn’t comprise the sensitivity enhancing motif, or a native or naturally occurring N-terminal intein segment. In another example, sensitivity can occur at a temperature of 0°C to 45°C, when the sensitivity enhancing motif is present on the N-terminal intein segment, as compared to either an N-terminal intein segment which doesn’t comprise the sensitivity enhancing motif, or a native or naturally occurring N-terminal intein segment.
[0101] The sensitivity enhancing motif (SEM) can be on the N-terminus of the N-terminal intein segment, for example. The SEM can be reversible. By "reversible" it is meant that the cleaving behavior of the split intein can be altered under a given extrinsic condition. This behavior of the intein, however, is reversible when the extrinsic condition is removed. For example, an intein may cleave or splice when at a pH of 6.2 but may not cleave or splice at a pH of 8.5. The SEMs disclosed herein can be designed such that when appended to, for example, the N-terminus of the N-terminal intein segment, they introduce a sensitivity enhancing element that allows for more precise control of cleavage of the protein of interest. This sensitivity has enabled the successful use of the inteins under conditions that are compatible with commercially relevant cell culture expression platforms.
[0102] In one example, a SEM can comprise the structure: aal - aa2 - aa3 - aa4, wherein aal is a non-polar amino acid, aa2 is a negatively-charged amino acid, aa3 is a non-polar amino acid and aa4 is a positively charged amino acid. For example, aal can be glycine, alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, or valine; aa2 can be aspartic acid or glutamic acid; aa3 can be glycine, alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, or valine; and aa4 can be arginine or lysine.
[0103] Also disclosed herein is a SEM, wherein the SEM comprises the structure: aal - aa2 - aa3 - aa4 - aa5, wherein aal is a non-polar amino acid, aa2 is a negatively-charged amino acid, aa3 is a non-polar amino acid, aa4 is a positively charged amino acid and aa5 is a non-polar amino acid. For example, aal can be glycine, alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, or valine; aa2 can be aspartic acid or glutamic acid or cysteine; aa3 can be glycine, alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, or valine; aa4 can be arginine or lysine; and aa5 can be glycine, alanine, isoleucine, leucine, methionine, phenylalanine, proline, tryptophan, or valine.
[0104] Disclosed herein are SEMS, wherein the reversible SEM comprises the sequence G-E-G-H-H (SEQ ID NO: 28), G-E-G-H-G (SEQ ID NO: 29), G-D-G-H-H (SEQ ID NO: 30), or G-D-G-H-G(SEQ ID NO: 31).Attorney Docket No: 103362-056WO1
[0105] Disclosed herein is a method of purifying a protein of interest, the method comprising: utilizing a split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment; immobilizing the N-terminal intein segment to a solid support; attaching a protein of interest to the C-terminal intein segment, wherein cleaving of the C-terminal intein segment is highly sensitive to extrinsic conditions when compared to a native intein; exposing the N-terminal intein segment and the C-terminal intein segment to each other so that they associate; washing the solid support to remove non-bound material; placing the associated the N-terminal intein segment and the C-terminal intein segment under conditions that allow for the intein to self-cleave; and isolating the protein of interest.SPECIFIC MUTA TIONS
[0106] As discussed above, the present invention relates to a protein purification system, wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1. These mutations can comprise, for example, C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q or C59L. Table 2 below discloses specific sequences with these mutations, as well as others, and they are described in more detail below.
[0107] Specifically, herein disclosed is an N-terminal intein segment which varies from SEQ ID NO: 1 by one or more amino acids. SEQ ID NOS: 2-12 disclosed specific single mutants in C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q and C59L (respectively). Also described are SEQ ID NOS: 15-17, which is a generic sequence where there is an “X” at position 28 (SEQ ID NO: 15), an “X” at position 59 (SEQ ID NO: 16), or an “X” at both positions 28 and 59 (SEQ ID NO: 17), where “X” can be any of the above-specified mutations, or any other mutation which confers benefits to the protein purification system disclosed herein.
[0108] Put another way, described herein is SEQ ID NO: 1 which comprises 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, or any amount in between, above, or below this amount, with respect to SEQ ID NO: 1. Specific examples of mutations within this range can be found in Table 2.
[0109] Also described herein is a C-terminal intein segment comprising SEQ ID NO: 18. SEQ ID NO: 18 can comprise an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more modifications that also provide desired properties which are preferable to the native SEQ ID NO: 18. An example of an additional modification can be found in SEQ ID NO: 19 (A34T). Put another way, described herein is a sequence which comprises 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, or any amount in between, above, or below this amount, identity with respect to SEQ ID NO: 18.Attorney Docket No: 103362-056WO1
[0110] It is noted that “at least one mutation” to SEQ ID NO: 1 means that more than one mutation can exist in the same peptide sequence. For example, SEQ ID NO: 1 can comprise 2, 3, 4, or 5 or more modifications. These can mean that multiple of these mutations are in the same peptide sequence: C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q and C59L. It can also mean that 1, 2, 3, 4, 5 or more of these specifi c mutations are found within the peptide sequence, along with at least one other additional mutation. This additional mutation can provide additional beneficial characteristics (these characteristics are described in detail above).
[0111] For example, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, or C28D compared to SEQ ID NO: 1; the N-terminal intein segment also comprises a mutation of C59S.
[0112] In another example, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L in the N-terminal intein, the C- terminal intein segment of the protein purification system can also comprise a mutation. This can include, for example, a mutation at A33T.
[0113] Table 1 shows a list of amino acids, their abbreviations, polarity, and charge.Table 1: Amino AcidsAmino Acid 3-Letter 1-Letter Polarity ChargeCode CodeAlanine Ala A nonpolar neutralBasicArginine Arg R positivepolarAsparagine Asn N polar neutralacidicAspartic acid Asp D negativepolarCysteine Cys C nonpolar neutralacidicGlutamic acid Glu E negativepolarGlutamine Gin Q polar neutralGlycine Gly G nonpolar neutralBasic Positive (10%) Histidine His HpolarNeutral (90%)Attorney Docket No: 103362-056WO1 Isoleucine Ile I nonpolar neutral Leucine Leu L nonpolar neutral BasicLysine Lys K positivepolarMethionine Met M nonpolar neutral Phenylalanine Phe F nonpolar neutral Proline Pro P nonpolar neutral Serine Ser S polar neutral Threonine Thr T polar neutral Tryptophan Trp W nonpolar neutral Tyrosine Tyr Y polar neutral Valine Val V nonpolar neutralTABLE 2: SEQUENCES AND MUTATIONS THEREOFSEQ ID NO CONTRUCT CONTRUCT SEQUENCE NAME DESCRIPTION SEQ ID NO: 1 NPUNWTWild-type Npu CLSYETEILTVEY DNAE (INTN GLLPIGKIVEKRIE segment) capable of CTVYSVDNNGNI splicing events YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 2 NpuNC28GVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC28G GTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 3 NpUNC281Variant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE of C28I ITVYSVDNNGNI YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFERELDLMRVDNLPNAttorney Docket No: 103362-056WO1 SEQ ID NO: 4 NPUNC28HVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC28H HTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 5 NpUNC28EVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE of C28E ETVYSVDNNGNI YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 6 NPUNC28DVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC28D DTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 7 NPUNC59GVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC59G CTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYGLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 8 NPUNC591Variant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC59I CTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYILED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 9 NPUNC59HVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE of C59H CTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYHLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 10 NPUNC59EVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIEofC59E CTVYSVDNNGNIAttorney Docket No: 103362-056WO1YTQPVAQWHDR GEQEVFEYELED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 11 NPUNC59QVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC59Q CTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYQLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 12 NpUNC59LVariant of SEQ ID CLSYETEILTVEY NO: 1 with mutation GLLPIGKIVEKRIE ofC59L CTVYSVDNNGNI YTQPVAQWHDR GEQEVFEYLLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 13 NpuNC28SWild-type Npu CLSYETEILTVEY DNAE ( JlN TN GLLPIGKIVEKRIE segment) capable of STVYSVDNNGNI splicing events YTQPVAQWHDR GEQEVFEYCLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 14 NPUNC59SWild-type Npu CLSYETEILTVEY DNAE (INTN GLLPIGKIVEKRIE segment) capable of CTVYSVDNNGNI splicing events YTQPVAQWHDR GEQEVFEYSLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 15 NPUNC28XWild-type Npu CLSYETEILTVEY DNAE (INTN GLLPIGKIVEKRIE segment) with XTVYSVDNNGNI variation at position YTQPVAQWHDR 28, where X can be GEQEVFEYCLED any amino acid GSLIRATKDHKFM substitution TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 16 NPUNC59XWild-type Npu CLSYETEILTVEY DNAE (INTN GLLPIGKIVEKRIE segment) with CTVYSVDNNGNI variation at position YTQPVAQWHDR 59, where X can be GEQEVFEYXLEDGSLIRATKDHKFMAttorney Docket No: 103362-056WO1any amino acid TVDGQMLPIDEIFE substitution RELDLMRVDNLPN SEQ IDNO: 17 NpuNC28X, C59XWild-type Npu CLSYETEILTVEY DNAE (INTN GLLPIGKIVEKRIE segment) with XTVYSVDNNGNI variation at positions YTQPVAQWHDR 28 and 59, where X GEQEVFEYXLED can be any amino GSLIRATKDHKFM acid substitution TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ IDNO: 18 NpucWTWild-type Npu MIKIATRKYLGKQ DNAE (INTc NVYDIGVERDH segment) capable of NFALKNGFIASN splicing eventsSEQ ID NO: 19 NpucA34TVariant of SEQ ID MIKIATRKYLGKQ NO: 18 withA34T NVYDIGVERDH mutation NFALKNGFITSNSEQ IDNO: 20 NpUNClx’C28X’C59XCleaving variant of XLSYETEILTVEY SEQ ID NO: 1; GLLPIGKIVEKRIE cleaving phenotype XTVYSVDNNGNI resulting from a YTQPVAQWHDR C1X mutation, GEQEVFEYXLED where “X”::::A or G, GSLIRATKDHKFM and C28 and / or C59 TVDGQMLPIDEIFE can be any of the RELDLMRVDNLPN mutations specifiedaboveSEQ ID NO: 22 So4b- NPUNC1X’C28X’ Variant of SEQ ID MGDGHGC59XNO: 20 modified XLSYETEILTVEY with Sensitivity- GLLPIGKIVEKRIE Enhancing Motif XTVYSVDNNGNI expressed as a YTQPVAQWHDR fusion partner at the GEQEVFEYXLED N-terminus GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN SEQ ID NO: 23 So4b- NPUNC1X’C28X’ Variant derived from MGDGHGC59XG4S-His6-Cys SEQ ID NO: 22; XLSYETEILTVEY constructed by GLLPIGKIVEKRIE adding linker, tag, XTVYSVDNNGNI and immobilization YTQPVAQWHDR moiety fusion GEQEVFEYXLED partners at the C- GSLIRATKDHKFMTVDGQMLPIDEIFEAttorney Docket No: 103362-056WO1terminus of the N- RELDLMRVDNLPN Intein Ligand. GGGGSHHHHHHCSEQ IDNO: 24 So4b- NpuNC,x’C28X’ Variant of SEQ ID MGDGHGC59XG4S-Cys-His6 NO: 23; derived XLSYETEILTVEY from an alternate GLLPIGKIVEKRIE arrangement of the XTVYSVDNNGNI I-L-T fusion partners YTQPVAQWHDR at the C -terminus. GEQEVFEYXLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN GGGGSHHHHHHC SEQ ID NO: 25 So4b- NpuNClx’C28X’ Variant of SEQ ID MGDGHGC59XG4S-Cys-His6 NO: 24 created by XLSYETEILTVEY adding an additional GLLPIGKIVEKRIE linker between I-L-T XTVYSVDNNGNI moieties at the C- YTQPVAQWHDR terminus. GEQEVFEYXLED GSLIRATKDHKFM TVDGQMLPIDEIFE RELDLMRVDNLPN GGGGSCGGGGS HHHHHH SEQ IDNO: 26 S04b-NpuNC1X,C- Variant of SEQ ID MGDGHGS’F74W_(G4S)2-Cys_ NO: 24 created by XLSYETEILTVEY G4S-His6 adding an additional GLLPIGKIVEKRIE linker between LL-T XTVYSVDNNGNI moieties at the C- YTQPVAQWHDR terminus. GEQEVFEYXLED GSLIRATKDHKFM’TVDGQMLPIDEIFE RELDLMRVDNLPN GGGGSGGGGS CGGGGSHHHHHH SEQ IDNO: 27 So4b-NpuNClx’c‘ Variant of SEQ ID MGDGHGS’F74W_(G4S)2-Cys_NO: 24 created by XLSYETEILTVEY G4S removing the His6 GLLPIGKIVEKRIE purification tag XTVYSVDNNGNI moiety from the C- YTQPVAQWHDR terminus. GEQEVFEYXLEDGSLIRATKDHKFMTVDGQMLPIDEIFEAttorney Docket No: 103362-056WO1RELDLMRVDNLPN GGGGSGGGGSCGGGGS SEQ ID NO: 28 sensitivity GEGHHenhancing motifSEQ IDNO: 29 sensitivity GEGHGenhancing motifSEQ ID NO: 30 sensitivity GDGHHenhancing motifSEQ IDNO: 31 sensitivity GDGHGenhancing motifSEQ ID NO: 32 NpuNl Figure 5 MGDGHGALSYTEIL TVEYGLPIGRIVEKR IESTVYSVDNNGNI YTQPVAVQWHDR SEQ ID NO: 33 NpuN2 Figure 5 GEQEVFEYSLEDG SLIRATKDHKFMTVD GQMLPIDEIFERELDL MRVDNPLNGSHG SEQ ID NO: 34 NcuC Figure 5 MIKIATRKYLGKQNVYGIGVERDHNFALKGNFIAHNNUCLEIC ACIDS, VECTORS, AND CELL LINES
[0114] Disclosed herein are vectors comprising nucleic acids encoding the C -terminal intein segment as disclosed herein, as well as cell lines comprising said vectors. As used herein, plasmid or viral vectors are agents that transport the disclosed nucleic acids, such as those encoding a C-terminal intein segment and a peptide of interest, into a cell without degradation and include a promoter yielding expression of the gene in the cells into which they can be delivered. In one example, a C-terminal intein segment and peptide of interest are derived from either a virus or a retrovirus. Retroviral vectors are able to carry a larger genetic payload, i.e., a transgene or marker gene, than other viral vectors and for this reason are a commonly used vector. However, they are not as useful in non-proliferating cells. Adenovirus vectors are relatively stable and easy to work with, have high titers, and can be delivered in aerosol formulation, and can transfect non-dividing cells. Pox viral vectors are large and have several sites for inserting genes; they are thermostable and can be stored at room temperature. Disclosed herein is a viral vector which has been engineered so as to suppress the immune response of a host organism, elicited by the viral antigens. Vectors of this type can carry coding regions for Interleukin 8 or 10.Attorney Docket No: 103362-056WO1
[0115] Viral vectors can have higher transfection (ability to introduce genes) abilities than chemical or physical methods to introduce genes into cells. Typically, viral vectors contain nonstructural early genes, structural late genes, an RNA polymerase III transcript, inverted terminal repeats necessary' for replication and encapsidation, and promoters to control the transcription and replication of the viral genome. When engineered as vectors, viruses typically have one or more of the early genes removed and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral DNA. Constructs of this type can carry up to about 8 kb of foreign genetic material. The necessary functions of the removed early genes are typically supplied by cell lines which have been engineered to express the gene products of the early genes in trans.
[0116] The fusion DNA encoding a modified peptide can be inserted into an appropriate expression vector, i.e., a vector which contains the necessary elements for the transcription and translation of the inserted protein-coding sequence. Depending on the host-vector system utilized, any one of a number of suitable transcription and translation elements can be used. For instance, when expressing a modified eukaryotic protein, it can be advantageous to use appropriate eukaryotic vectors and host cells. Expression of the fusion DNA results in the production of a modified protein.
[0117] Also disclosed herein are cell lines comprising the vectors or peptides disclosed herein. A variety of cells can be used with the vectors and plasmids disclosed herein. Non¬ limiting examples of such cells include somatic cells such as blood cells (erythrocytes and leukocytes), endothelial cells, epithelial cells, neuronal cells (from the central or peripheral nervous systems), muscle cells (including myocytes and myoblasts from skeletal, smooth or cardiac muscle), connective tissue cells (including fibroblasts, adipocytes, chondrocytes, chondroblasts, osteocytes and osteoblasts) and other stromal cells (e g., macrophages, dendritic cells, thymic nurse cells, Schwann cells, etc.). Eukaryotic germ cells (spermatocytes and oocytes) can also be used, as can the progenitors, precursors and stem cells that give rise to the above-described somatic and germ cells. These cells, tissues and organs can be normal, or they can be pathological such as those involved in diseases or physical disorders, including, but not limited to, infectious diseases (caused by bacteria, fungi yeast, viruses (including HIV, or parasites); in genetic or biochemical pathologies (e.g., cystic fibrosis, hemophilia, Alzheimer's disease, schizophrenia, muscular dystrophy, multiple sclerosis, etc.); or in carcinogenesis and other cancer-related processes.
[0118] The eukaryotic cell lines disclosed herein can be animal cells, plant cells (monocot or dicot plants) or fungal cells, such as yeast. Animal cells include those of vertebrate orAttorney Docket No: 103362-056WO1 invertebrate origin. Vertebrate cells, especially mammalian cells (including, but not limited to, cells obtained or derived from human, simian or other non-human primate, mouse, rat, avian, bovine, porcine, ovine, canine, feline and the like), avian cells, fish cells (including zebrafish cells), insect cells (including, but not limited to, cells obtained or derived from Drosophila species, from Spodoptera species (e.g., Sf9 obtained or derived from S. frugiperida, or HIGH FIVE™ ceils) or from Trichoplusa species (e.g., MG1, derived from T. ni)), worm cells (e.g., those obtained or derived from C. elegants), and the like. It will be appreciated by one of skill in the art, however, that cells from any species besides those specifically disclosed herein can be advantageously used in accordance with the vectors, plasmids, and methods disclosed herein, without the need for undue experimentation.
[0119] Examples of useful cell lines include, but are not limited to, HT1080 cells (ATCC CCL 121), HeLa cells and derivatives of HeLa cells (ATCC CCL 2, 2.1 and 2.2), MCF-7 breast cancer cells (ATCC BTH 22), K-562 leukemia cells (ATCC CCL 243), KB carcinoma cells (ATCC CCL 17), 2780AD ovarian carcinoma cells (see Van der Buick, A. M. et al.. Cancer Res.48:5927-5932 (1988), Raji cells (ATCC CCL 86), Jurkat ceils (ATCC TIB 152), Namalwa ceils (ATCC CRL 1432), HL-60 cells (ATCC CCL 240), Daudi cells (ATCC CCL 213), RPMI 8226 cells (ATCC CCL 155), U-937 ceils (ATCC CRL 1593), Bowes Melanoma cells (ATCC CRL 9607), WI-38VA13 subline 2R4 cells (ATCC CLL 75.1), and MOLT-4 cells (ATCC CRL 1582), as well as heterohybridoma cells produced by fusion of human cells and cells of another species. Secondary human fibroblast strains, such as WI-38 (ATCC CCL 75) and MRC-5 (ATCC CCL 171 can also be used. Other mammalian cells and cell lines can be used in accordance with the present invention, including, but not limited to CHO cells, COS cells, VERO cells, 293 cells, PER-C6 cells, Ml cells, NS-1 cells, COS-7 cells, MDBK cells, MDCK cells, MRC-5 cells, WI-38 ceils, WEHI ceils, SP2 / 0 ceils, BHK cells (including BHK-21 cells); these and other cells and cell lines are available commercially, for example from the American Type Culture Collection (P. O. Box 1549, Manassas, Va. 20108 USA). Many other cell lines are known in the art and will be familiar to the ordinarily skilled artisan; such cell lines therefore can be used equally well.
[0120] Once obtained, the proteins of interest can be separated and purified by appropri ate combination of known techniques. These methods include, for example, methods utilizing solubility such as salt precipitation and solvent precipitation; methods utilizing the difference in molecular weight such as dialysis, ultra-filtration, gel-filtration, and SDS-polyacrylamide gel electrophoresis; methods utilizing a difference in electrical charge such as ion-exchange column chromatography; methods utilizing specific affinity such as affinity chromatography; methodsAttorney Docket No: 103362-056WO1 utilizing a difference in hydrophobicity such as reverse-phase high performance liquid chromatography; and methods utilizing a difference in isoelectric point, such as isoelectric focusing electrophoresis. These are discussed in more detail below.METHODS OF USING PROTEIN PURIFICATION SYSTEMS AND KITS
[0121] Disclosed is a method of purifying a protein of interest, the method comprising: utilizing a split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment, wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L; covalently immobilizing the N-terminal intein segment to a solid support; attaching a protein of interest to the C-terminal intein segment, wherein cleaving of the C-terminal intein segment is more sensitive to extrinsic conditions compared to a native intein; exposing the N-terminal intein segment and the C-terminal intein segment to each other so that they associate; washing the solid support to remove non-bound material; placing the associated N-terminal intein segment and the C-terminal intein segment under conditions that allow for the intein to self-cleave; and isolating the protein of interest.
[0122] Also disclosed herein are kits. A kit, for example, can include the split intein disclosed herein (an N-terminal intein segment and a C-terminal intein segment) as well as a protein of interest, optionally. The kit can also include instructions for use.C. EXPERIMENTAL
[0123] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how the compounds, compositions, articles, devices, and / or methods cl imed herein are made and evaluated, and are intended to be purely exemplary of the invention and are not intended to limit the scope of what the inventors regard as their invention. Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperatures, etc.), but some errors and deviations should be accounted for.EXAMPLE 1: Site-directed mutagenesis production of single mutant constructs
[0124] The template DNA used for these experiments was pET / RBS-Zn-NpuN(C 1 A, CF)- G4S-His6-Cys-Stop-RBS-NpuC-GFP. This plasmid produces a bicistronic mRNA that encodes both the N-terminal (NpuN) and C-terminal (NpuC) intein segments as separate products. In this case, NpuC segment is fused to the gene of interest, which for this experiment is a greenAttorney Docket No: 103362-056WO1 fluorescent protein (GFP). The first three amino acids following the C-terminal intein fragment (NpuC) used were F, F, N in the +1, +2, and +3 amino acid residues, respectively. These amino acids were selected because they are known to facilitate fast cleaving, allowing a quick evaluation of the impacts of the proposed mutants on intein function. Standard recombinant DNA techniques were employed to perform site-directed mutagenesis at either the 28thor 59thposition of the NpuN intein fragment (Figures 1 and 3). All 20 possible amino acids were introduced to each position for full evaluation of the library. Mutagenesis was carried out via PCR, and the size of the PCR product was confirmed using DNA gel electrophoresis. The amplified product was re-circularized and transformed into E coli strain DH5-a and sequenced to verify each mutated sequence. Positive clones were subsequently transformed into E. coli strain BLR for protein expression.EXAMPLE 2: Protein expression and in-solution cleavage kinetic testing
[0125] Single isolated colonies of each mutant were selected from ampicillin-resistant agar plates and used to inoculate 5 mL primary cultures of selective IX Luria broth (LB) media (10 g NaCl, 5 g yeast extract, 10 g tryptone per liter). These primary cultures were grown at 37°C, 220 rpm in a water bath for 18 (± 0.5) hours. After 18 hours, the ODeoo of each culture was measured. Typical ODeoo readings were around 4.5. The primary cultures were then normalized to an ODeoo of 3.5 using Equation 1.
[0126] VP’-^-y to add=(Equation 1)OD600
[0127] Secondary cultures (5 mL) were inoculated with 3 % v / v of primary culture. The secondary cultures were grown under identical conditions through the exponential growth phase. At an OD600of 0.6-0.8, protein expression was induced with 1 mM IPTG (5 pL of a 1 M stock solution). The cultures were returned to the 37°C, 220 rpm water bath shaker for 6 (± 0.5) hours to allow for adequate protein expression. The final ODeoo of each culture was measured after 6 hours. In a 2 mL Eppendorf tube, 2 mL of culture were harvested via centrifugation at 8,000 x g for 5 minutes, 20°C. The supernatant was discarded, and an additional 2 mL of culture were harvested using the same conditions, for a total of 4 mL of cell culture pelleted. Each cell pellet was frozen at -20°C for at least 12 hours. This process was completed in triplicate for each NpuN mutant.Attorney Docket No: 103362-056WO1
[0128] After thawing, the cells were resuspended in lysis buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 8.5) normalized to an ODeoo of 10 using Equation 2:Vresuspension(μL) = 4000 μL*OD600Secondary
[0129] Vresuspension(μL) = 4000 μL*OD600Secondary / 10 (Equation 2)
[0130] Cell lysis was performed on ice, using three 30-second cycles of sonication at 7 W. Lysates were clarified via centrifugation at 21,130 x g for 15 minutes, 20°C. Samples of each clarified cell lysate (CCL) were combined in a 50:50 ratio with 2*SDS-PAGE loading dye containing beta-mercaptoethanol immediately following clarification. The samples were boiled for 7 minutes. CCL samples were analyzed using SDS-PAGE gel electrophoresis. For each mutant, 5 pL of each CCL sample was loaded into 12% acrylamide gels and ran at 200 V for 45 minutes. The gels were stained using Coomassie blue for analysis of intein cleaving function.EXAMPLE 3: Image J Analysis
[0131] The SDS-PAGE gels were scanned as photograph files (600 dpi) and uploaded into ImageJ for analysis (Figures 2 and 4). Each gel lane was selected individually near protein bands of interest and band intensity profiles were generated using the plot lanes command. The output protein band intensity peaks were manually separated using the straight-line tool and integrated using the wand tool in the ImageJ toolbox. The percentage of cleaved product was determined by taking the area under the second curve (cleaved product) and dividing by the total area under both curves (total protein of interest product) for each sample. This process was repeated for each sample and the resulting percent cleaved was recorded in a JMP table.Additionally, the percent difference from the control (C28S and C59S) was calculated and recorded.EXAMPLE 4: JMP statistical analysis
[0132] The percent cleavage for each mutant was analyzed using an ANOVA test, treating the day on which the triplicate was expressed as a blocking variable. Separate ANOVA tests were conducted for mutants in the 28thand 59thposition. For mutants in the 28thposition, the ANOVA test revealed that at least one mutant was statistically different from the others. An all¬ means Tukey-Kramer test was then performed to assess pairwise differences in means for mutations in the 28thposition. The ANOVA test for mutants in the 59thposition indicated noAttorney Docket No: 103362-056WO1 statistically significant differences between the mutants. The percent difference in cleavage rate compared to the control (S28, S59) was plotted to visualize the data trends. The significance level of all statistical tests was set to a = 0.05.EXAMPLE 5: Site-directed mutagenesis production of double mutant constructs
[0133] The template DNA used for these experiments was pET / RBS-Zn-NpuN(Cl A, CF)-G4S-His6-Cys-Stop-RBS-NpuC-Stop, where the GFP at the C -terminus of NpuC in example 1 was removed. Standard recombinant DNA techniques were employed to perform site-directed mutagenesis at either the 28thor 59thposition of the NpuN intein fragment depending on double mutant identity. Mutagenesis was earned out via PCR, and the size of the PCR product was confirmed using DNA gel electrophoresis. The amplified product was re-circularized and transformed into E. coli strain DH5-a and sequenced to verify each mutated sequence. Positive clones were subsequently transformed into E. coli strain BLR for protein expression. Amino acids glycine (G), histidine (H), leucine (L), isoleucine (I) and glutamic acid (E) were substituted into the 28thand / or 59thposition. Each residue was selected by assigning a wellperforming single mutant to represent a group of amino acid sidechain characteristics. Glycine represents small nonpolar residues. Isoleucine and leucine represent bulky nonpolar residues. Histidine represents polar, positive charged residues. Glutamic acid represents polar negative charged residues. The following 16 double mutant clones were constructed: S28G+S59G, S28G+S59H, S28G+S59L, S28G+S59E, S28H+S59G, S28H+S59H, S28H+S59L, S28H+S59E, S28I+S59G, S28I+S59H, S28I+S59L, S28I+S59E, S28E+S59G, S28E+S59H, S28E+S59L, S28E+S59E. These are summarized in the following table:
[0134] 28thpositionS28G S28H S28I S28ES59G GG HG IG EGS59H GH HH 1H EH59thpositionS59L GL HL IL ELS59E GE HE IE EEEXAMPLE 6: Protein expression and purification for the 16 double mutants and controlAttorney Docket No: 103362-056WO1
[0135] Single isolated colonies of each mutant and the S28 S59 control (17 total constructs) were selected from ampicillin-resistant agar plates and used to inoculate 5 niL primary cultures of selective IX Luria broth (LB) media. These primary cultures were grown at 37°C, 220 rpm in a water bath for 18 (± 0.5) hours. After 18 hours, the ODeoo of each culture was measured, and each primary culture was normalized to an ODeoo of 3.5 (equation 1 example 2). Secondary cultures (50 niL) were inoculated with 3 % v / v of normalized primary culture.The secondary cultures were grown at 37°C, 325 rpm in a temperature-controlled air shaker through the exponential growth phase (ODeoo of 1-2). Protein expression was induced with 1 mM IPTG, and then the temperature was dropped to 16°C and left shaking at 325 rpm for 16 (± 0.5) hours. The final ODeoo of each culture was measured at the end of expression. Each culture was harvested into a 50 mL centrifuge tube at 8,000 x g for 10 minutes, 4°C. The supernatant was discarded and each pellet was frozen at ~20°C for at least 12 hours. After thawing, the cells were resuspended in lysis buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 8.5) normalized to an OD600 of 10 using Equation 3:Vresuspension(mL) = 50 mL*OD600Secondary / 10 (Equation 3)
[0136] Cell lysis was performed on ice, using ten 30-second cycles of sonication at 7 W. Lysates were clarified via centrifugation at 21,130 x g for 30 minutes, 4°C. Each clarified cell lysate (CCL) was sampled for SDS-PAGE as described in Example 2. Each CCL was then purified via IMAC with 1 mL of bulk resin (per construct) using standard protocol. The elution fraction for each ligand was sampled for SDS-PAGE at this time.EXAMPLE 7: Sulfolink coupling for all 17 ligands
[0137] Each purified ligand was buffer exchanged into coupling buffer (200 M boric acid 50 mM sodium tetraborate 150 mM NaCl pH 8.80) over a 3 kDa dead-end centrifugal filter. The final buffer-exchanged material was concentrated to 20 mg / mL over the same filter and the precoupling mass was determined by multiplying the concentration by the concentrated sample volume. At this time, Sulfolink™ resin was added to a 20 mL gravity column in a ratio of 1 mL resin per every mL of protein at 20 mg / mL, such that 20 mg of ligand was available to 1 mL of resin for coupling. All resin-handling steps from this point passivation were performed with light-limiting precautions including wrapping the columns in tin foil due to the sensitive nature of the uncoupled Sulfolink™ resin. The storage solution was drained from the resin and then equilibrated in 10 CVs of coupling buffer. TCEP was added to a final concentration of 25 mM in the concentrated sample to reduce each ligand’s C-terminal cysteine residue. The reduced andAttorney Docket No: 103362-056WO1 concentrated sample was then applied to the capped gravity column containing an appropriate volume of equilibrated resin. The column was secured and then set on a rotating wheel placed in a 30°C incubator for 16 (± 0.5) hours. After incubation, the column was drained and the mass of protein in the flowthrough was measured. This mass was subtracted from the pre-couple mass and then divided by the resin volume to determine the ligand density in mg ligand per mL of resin. Each resin bed was then blocked with 1 CV of passivating solution (50 mM L-cystine) for one hour at 30C on the rotating wheel. Light-sensitive precautions were removed at this time as the resin coupling reaction was complete. The resin was stored in 20% ethanol at 4°C until ready for use.EXAMPLE 8: On-column cleavage test
[0138] A standard intein cleavage test was performed to determine the on-column performance of each double mutant ligand. Each Sulfolink™-coupled resin bed was made into a 50% slurry, resuspended fully, and then 400 uL of slurry was moved into a small (10 mL) gravity column and allowed to drain. The 0.2 mL CV resin bed was then stripped with 50 CVs of 150 mM H₃PO₄(aq) over the course of an hour. The acid was drained and the resin was then equilibrated in 10 CVs of intein columns buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 8.5). To test the relative cleaving rate of each mutant ligand, an intein-tagged GFP test protein was prepared, where the intein tag had a mutation of alanine to threonine at position 33, and the initial three amino acids of the GFP target protein were VSK (valine, serine, lysine). The tagged target was pre-purified via IMAC, using a His tag appended to its C -terminus. For each test, 3 ml of 1 mg / ml tagged GFP test protein was loaded onto the column, and the column was sealed and allowed to rotate gently at room temperature for one hour to maximize binding. The flowthrough was drained after binding and then the resin was then washed with high-salt buffer (20 mM AMPD, 20 mM PIPES, 500 mM NaCl, pH 8.5). Following the wash, the pH was dropped to 6.2 by passing 10 CVs of elution buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 6.2) over the resin. The column was then sealed, and one CV of elution buffer was added to the resin to create a 50% slurry. The mixture was resuspended fully and then sampled directly from the slurry to capture the first resin (“R0”) sample. Resin samples were repeated at 0.5, 1, 2.5, 5, and 24 hours for each ligand. After 24 hours, each column was drained and the elution containing cleaved protein was collected and sampled. The resin was then stored in 20% ethanol at 4°C for future use. For analysis, ImageJ was performed on all the SDS-PAGE gels for all 16 double mutants and the S28 S59 control as described in Example 3.Attorney Docket No: 103362-056WO1 EXAMPLE 9: Quantifying cleavage kinetics on an AKTA FPLC
[0139] As with Example 8, NpuC-sfGFP-His6 constructs with the N-terminus of GFP appended with the three amino acids DIG (aspartic acid, isoleucine, glycine) were initially pre¬ purified using a standard IMAC chromatography protocol. The purified samples were then diluted to 1 mg / mL in column buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 8.5). For each test, 20 mL of these GFP samples were loaded onto double-mutant FPLC columns prepared as described in Examples 6 and 7, at a flow rate of 1 mL / min. The samples were passed through the column twice, for a total of 40 mL. The column was washed with the same buffer, followed by a wash with cleavage buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 6.2). The cleavage buffer was then slowly pumped through the column at 0.01 mL / min for 6.5 hours, totaling 4 mL. Protein elution was monitored via A280 in mAU using a UV-Vis detector mounted directly downstream of the AKTA column.
[0140] GFP elution from the column is proportional to the cleavage rate, which was assumed to follow first-order kinetics according to Equation 4:Rate — k[Npu: GFP] (Equation 4)where k is the rate constant and Npu: GFP is the bound NpuN / NpuC intein with the GFP still attached. The tail end of the elution peak exhibited exponential decay behavior consistent with first-order kinetics. Based on these kinetics, the AKTA’s output was used to create a plot of ln(A280) versus time to align with the integrated rate law, which is shown in Equation 5:ln(mAU) = —kt (Equation 5)
[0141] This plot fits a linear equation with a slope of -k. These rate constants were also used to determine the half-life of each tested double-mutant using Equation 6.In 2^•1 / 2 “ (Equation 6)EXAMPLE 10: Quantifying pH sensitivity
[0142] As described in Example 9, NpuC-sfGFP-His6 constructs were pre-purified and applied to double-mutant FPLC columns. After washing away unbound protein with the same column buffer at pH 8.5, the buffer was slowly pumped through the column at 0.01 mL / min for 15 hours, totaling 9 mL. Protein elution was monitored via A280 using a UV-Vis detector mounted directly downstream of the AKTA column and collected in a fraction collector.
[0143] Following this 15-hour step at pH 8.5, the column was washed with cleavage buffer (20 mM AMPD, 20 mM PIPES, 200 mM NaCl, pH 6.2) until the pH stabilized, and a secondAttorney Docket No: 103362-056WO1 step was performed in which the lower pH buffer was slowly pumped through the column at 0.01 mL / min for 10 hours. This elution was also collected in a fraction collector.
[0144] The GFP that eluted at pH 8.5 and the GFP that eluted at pH 6.2 were quantified following a standard BCA assay. The pH sensitivity of the double mutants could then be compared by the fraction of GFP that eluted at pH 8.5 over 15 hours compared to the total GFP eluted at both pH 8.5 and pH 6.2.
Claims
Attorney Docket No: 103362-056WO1CLAIMSWhat is claimed is:
1. A protein purification system, wherein the system comprises a split intein comprising two separate peptides: an N-terminal intein segment and a C -terminal intein segment; wherein the N-terminal intein segment is capable of covalent immobilization onto a solid support, and wherein the C -terminal intein segment is capable of attachment to a protein of interest; and further wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q or C59L; wherein the mutation increases yield of the protein of interest, increases cleavage rate, increases stability of the protein purification system, or is more sensitive to extrinsic conditions compared to a native intein.
2. The protein purification system of claim 1, wherein, in addition to the at least one mutation comprising C59G, C59I, C59H, C59E, or C59L compared to SEQ ID NO: 1; the N-terminal intein segment also comprises a mutation of C28S.
3. The protein purification system of claim 1, wherein, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, or C28D compared to SEQ ID NO: 1; the N-terminal intein segment also comprises a mutation of C59S.
4. The protein purification system of any one of claims 1-3, wherein, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L, the C -terminal intein segment also comprises a mutation at A33T.
5. The protein purification system of any one of claims 1-4, wherein, in addition to the mutation of at least one of C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L of SEQ ID NO: 1, the N-terminal intein segment also comprises an additional mutation which further increases yield of the protein of interest, increases cleavage rate, increases stability of the protein purification system, or is more sensitive to extrinsic conditions compared to a native intein.
6. The protein purification system of any one of claims 1-5, wherein a purification tag is attached to the N-terminal intein segment at its C-terminus.
7. The protein purification system of claim 6, wherein the purification tag comprises one or more histidine residues.Attorney Docket No: 103362-056WO1 8. The protein purification system of any one of claims 1-7, wherein the N-terminal intein segment is immobilized on a solid chromatographic resin backbone, or another chromatographic / non-chromatographic purification system.
9. The protein purification system of any one of claims 1-8, wherein the N-terminal intein segment further comprises a sensitivity-enhancing motif, which renders it highly more sensitive to extrinsic conditions.
10. The protein purification system of claim 9, wherein the sensitivity-enhancing motif is on the N-terminus of the N-terminal intein segment.
11. The protein purification system of claim 9 or 10, wherein the extrinsic condition is pH, temperature, or both.
12. The protein purification system of any one of claims 1-11, wherein the N-terminal intein segment comprises any one of SEQ ID NO: 2-17 or 20-27.
13. A method of purifying a protein of interest, the method comprising:a. utilizing a split intein comprising two separate peptides: an N-terminal intein segment and a C-terminal intein segment, wherein the N-terminal intein segment comprises at least one mutation comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L;b. covalently immobilizing the N-terminal intein segment to a solid support;c. attaching a protein of interest to the C-terminal intein segment, wherein cleaving of the C-terminal intein segment is more sensitive to extrinsic conditions compared to a native intein; d. exposing the N-terminal intein segment and the C-terminal intein segment to each other so that they associate;e. washing the solid support to remove non-bound material;f. placing the associated N-terminal intein segment and the C-terminal intein segment under conditions that allow for the intein to self-cleave;g. isolating the protein of interest.
14. The method of claim 13, wherein, in addition to the at least one mutation comprising C59G, C59I, C59H, C59E, C59Q, or C59L compared to SEQ ID NO: 1; the N-terminal intein segment also comprises a mutation of C28S.
15. The method of claim 13, wherein, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, C59Q, or C28D compared to SEQ ID NO: 1; the N-terminal intein segment also comprises a mutation of C59S.Attorney Docket No: 103362-056WO1 16. The method of any one of claims 13-15, wherein, in addition to the at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L, the C -terminal intein segment also comprises a mutation at A33T.
17. The method of any of claims 13-16, wherein, in addition to the mutation of at least one of C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L of SEQ ID NO: 1, the N-terminal intein segment also comprises an additional mutation which further increases yield of the protein of interest, increases cleavage rate, increases stability of the protein purification system, or is more sensitive to extrinsic conditions compared to a native intein.
18. The method of any of claims 13-17, wherein a purification tag is attached to the N- terminal intein segment at its C -terminus.
19. The method of claim 18, wherein the purification tag comprises one or more histidine residues.
20. The method of any one of claims 13-19, wherein the solid support comprises a solid chromatographic resin backbone or another chromatographic / non-chromatographic purification system..
21. The method of any one of claims 13-20, wherein the N-terminal intein segment further comprises a sensitivity-enhancing motif, which renders it highly more sensitive to extrinsic conditions.
22. The method of claim 21, wherein the sensitivity-enhancing motif is on the N-terminus of the N-terminal intein segment.
23. The method of claim 22, wherein the extrinsic condition is pH, temperature, or both.
24. An amino acid sequence comprising 90% or more identity to SEQ ID NO: 1, wherein SEQ ID NO: 1 comprises at least one mutation comprising C28G, C28I, C28H, C28E, C28D, C59G, C59I, C59H, C59E, C59Q, or C59L.
25. The amino acid sequence of claim 24, wherein the amino acid sequence comprises at least one additional mutation which provides at least one benefit compared to a native intein when the amino acid sequence is used in a protein purification system.
26. The amino acid sequence of claim 25, wherein the benefit comprises increased yield of a protein of interest, increases cleavage rate, increases stability of the protein purification system, or increased sensitivity to extrinsic conditions compared to a native intein.
27. A kit comprising the amino acid sequence of any one of claims 24-26.