Protein purification using the split intein system
The modified N-intein system addresses the limitations of existing split intein systems by enabling rapid and efficient production of untagged proteins with native sequences through improved stability and cleavage kinetics, suitable for large-scale purification.
Patent Information
- Application Number
- JP2022526270
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-22
- Filing Date
- 2020-11-20
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2040-11-20
AI Technical Summary
Existing split intein systems for protein purification are limited by the requirement of a specific amino acid at the splice junction, slow cleavage kinetics, solubility issues, and lack of suitability for large-scale untagged protein purification, necessitating improvements for efficient and high-yield production of native proteins.
Development of N-intein protein variants modified to eliminate asparagine residues and allow for alkaline stability, enabling rapid and efficient cleavage of affinity tags to produce untagged proteins using a split intein system in a single chromatography step.
The modified N-intein system facilitates high-yield, rapid, and efficient production of untagged proteins with native sequences by overcoming the limitations of existing systems, providing improved stability and versatility for large-scale purification.
Smart Images

Figure 0007744026000043 
Figure 0007744026000044 
Figure 0007744026000045
Abstract
Description
[Technical Field]
[0001] The present invention relates primarily to protein purification in the field of chromatography, and more specifically to affinity chromatography using an improved split intein system with a C-intein tag and an N-intein ligand, whereby the target protein can be purified as an untagged final product with a native N-terminus. [Background technology]
[0002] Inteins are protein elements expressed as in-frame inserts that interrupt enzyme sequences and catalyze their own excision and the ligation of two adjacent polypeptides to generate active proteins. Genetically, inteins are encoded in two different ways: as intact inteins that interrupt two adjacent extein sequences, or as split inteins, in which each extein and part of the intein is encoded by two different genes. Although they hold great promise as biotechnology and protein purification tools, naturally occurring split inteins have rapid kinetic properties that depend on specific amino acids at the intein-extein junction, severely limiting the proteins that can be fused to inteins for affinity purification and recovery of native protein sequences. In particular, a typical split intein DNAE from Nostoc punctiforme exhibits favorable kinetic properties for protein purification applications. However, its activity depends on a phenylalanine at the +2 position of the C-extein. This dependency severely limits and impairs its general applicability.
[0003] Inteins have been engineered to fulfill several important functions in biotechnology, including applications as self-cleaving proteins for recombinant protein purification. Split inteins are particularly promising in this regard because they can simultaneously provide affinity ligand and self-cleaving properties. In protein purification, either extein can be substituted for the target protein being purified. To date, split inteins of the DNAE family have shown the greatest promise for C-terminally truncated protein purification approaches.
[0004] WO2014 / 004336 describes proteins fused to the N fragment of a split intein and the C fragment of a split intein that can be bound to a support. The solid support can be a particle, bead, resin, or slide.
[0005] WO2014 / 110393 describes a protein of interest fused to a split intein N fragment and a split intein C fragment that is contacted with a purification tag. The N fragment can be bound to a solid phase via the purification tag, and methods of affinity purification are discussed.
[0006] U.S. Provisional Patent Application No. 10,066,027 describes a protein purification system and a method for using the system. A split intein is disclosed, which includes an N-terminal intein segment that can be immobilized and a C-terminal intein segment that has self-cleaving properties and can bind to a protein of interest. The N-terminal intein segment is provided with a sensitivity-enhancing motif that makes it more sensitive to exogenous conditions.
[0007] US Provisional Patent Application No. 10,308,679 describes fusion proteins comprising an N-intein polypeptide and an N-intein solubilization partner, as well as affinity matrices comprising such fusion proteins.
[0008] WO 2018 / 091424 describes a method for producing an affinity chromatography resin containing an amino-terminal (N-terminal) split intein fragment as an affinity ligand, the method comprising the following steps: a) expressing the N-terminal split intein fragment protein as an insoluble protein in inclusion bodies of bacterial cells, preferably Escherichia coli (E. coli), b) harvesting the inclusion bodies; c) solubilizing the inclusion bodies to release the expressed protein; d) binding the protein to a solid support; e) refolding the protein; f) releasing the protein from the solid support; and g) immobilizing the protein as a ligand on a chromatography resin to form an affinity chromatography resin. This procedure allows for immobilization of a ligand density of 2 to 10 mg / ml of resin.
[0009] As mentioned above, split inteins have been used for protein purification using multiple affinity tags and tag cleavage mechanisms. However, the utility of such systems is limited by several factors. First, there is an amino acid requirement at the splice junction of the desired product, i.e., a Phe at position +2 of the C-extein, to effect cleavage and achieve purification of the untagged protein. The production of recombinant proteins without extraneous amino acids at the N-terminus is highly desirable. Second, cleavage to release the protein must be sufficiently fast and provide acceptable yields. Third, there is a solubility requirement for the N- or C-fragment of the split intein for its attachment to a solid support. Fourth, there are currently no available split intein systems suitable for large-scale purification of untagged proteins. [Prior art documents] [Patent documents]
[0010] [Patent Document 1] WO2014 / 004336 [Patent Document 2] WO2014 / 110393
Patent document 3
Patent document 4
Patent document 5
Non-licensed literature
[0011] [Non-licensed document 1] Proteins-Structure and Molecular Properties 2nd edition, TE Creighton, WH Freeman and Company, New York (1993) [Non-licensed document 2] Posttranslational Covalent Modification of Proteins, edited by BC Johnson, Academic Press, New York, pages 1~12 (1983) [Non-licensed document 3] Merck Index (14th edition), Physicians' Desk Reference (64th edition)
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed literature 9
[0012] The present invention overcomes the drawbacks of the prior art and allows for the general purification of untagged / native proteins in just one rapid affinity chromatography step using a split intein system.
[0013] The present invention provides N-intein protein variant sequences of naturally occurring split inteins or a consensus sequence derived from naturally occurring inteins and split inteins, where the N-intein variants are modified to eliminate all asparagine (N) amino acid residues present in the sequence compared to the native or consensus sequence. Preferably, all such N-intein variant sequences are further modified to replace cysteine (C) at position 1 with any other amino acid that is not cysteine.
[0014] The present invention provides N-intein protein variants of a consensus sequence derived from a naturally occurring split intein or intein / split intein, wherein the N-intein protein variant does not contain an asparagine (N) at position 36 of the variant sequence. This position is calculated by conventional cluster alignment with naturally occurring split inteins starting from the first catalytic cysteine, numbered 1. While this position is conserved in prior art and naturally occurring N-intein sequences at N, the inventors have found that this position can be mutated to other amino acids less susceptible to deamidation, such as histidine (H or His) or glutamine (Q or Gln), thereby achieving increased alkaline stability, which is important because it confers resistance to increasing pH values, for example, during chromatography procedures. While at least the N at position 36 must be mutated, it is also contemplated that more Ns in an N-intein sequence can be mutated, preferably to H or Q.
[0015] The present invention also provides N- and C-inteins that overcome the absolute requirement for a phenylalanine at position +2 of a target protein of interest (POI). The N- and C-inteins of the present invention can be used for the production of any recombinant protein. By using the N- and C-inteins of the present invention, tag cleavage occurs at the precise junction of the tag intein and the POI, meaning that the POI is expressed in its native form, free of extraneous amino acids encoded by the affinity tag. Furthermore, the intein sequences of the present invention allow the POI to be produced in high yield and with fast cleavage kinetics. The N-intein is coupled to a solid phase that can be renatured under alkaline conditions.
[0016] The present invention provides N-intein, C-intein, split intein systems and methods of using same as defined in the appended claims. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a graph showing the relative binding capacity of N-intein ligands (A40, A41, and A48) according to the present invention coupled to an SPR biosensor chip. [Figure 2] 1 is a staple diagram showing the relative binding ability of N-intein ligands according to the present invention (B72, B22, A48) and a comparative ligand (A53) coupled to an SPR sensor chip. [Figure 3] The static binding capacity of the N-intein ligands of the present invention is shown. Amino acid analysis (AAA) is performed by conventional methods. The A48 prototype is coupled to porous agarose particles via epoxy chemistry. [Figure 4A] Chromatogram of the purification results of Experiment 6. [Figure 4B] FIG. 1 shows SDS PAGE results from experiment 6. [Figure 5]1 is a graph showing the relative binding capacity of N-intein ligands (A40 and A48) according to the present invention coupled to an SPR biosensor chip. DETAILED DESCRIPTION OF THE INVENTION
[0018] definition As used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural unless the content clearly dictates otherwise. Thus, for example, reference to a "group," "alkyl," or "residue" includes mixtures of two or more such groups, alkyls, or residues, etc.
[0019] Ranges can be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, a further aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it is understood that the particular value forms a further aspect. It is further understood that each endpoint of a range is significant relative to the other endpoint, and independently of the other endpoint. It is also understood that there are several values disclosed herein, and that each value is also disclosed herein as "about" that particular value in addition to the value itself. For example, if the value "10" is disclosed, then "about 10" is also disclosed. It is also understood that each unit between two particular units is disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
[0020] Unless specifically stated to the contrary, the weight percent (wt%) of an ingredient is based on the total weight of the formulation or composition in which it is included.
[0021] As used herein, the term "optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes instances in which said event or circumstance occurs and instances in which it does not occur.
[0022] The term "contacting" as used herein refers to bringing two biological entities together so that the compound can affect the activity of the target either directly, i.e., by interacting with the target itself, or indirectly, i.e., by interacting with another molecule, cofactor, factor, or protein on which the activity of the target depends. "Contacting" can also mean facilitating the interaction of two biological entities, such as peptides, so as to bind covalently or otherwise.
[0023] As used herein, a "kit" refers to a collection of at least two components that make up the kit. Together, these components constitute a functional unit for a given purpose. The individual components can be physically packaged together or separately. For example, a kit that includes instructions for using the kit may or may not physically include the instructions with the other individual components. Alternatively, the instructions can be provided as separate components in paper form, or in electronic form that can be provided on a computer-readable memory device or downloaded from an internet website, or as a recorded presentation.
[0024] As used herein, "instructions" means documents that describe the relevant materials or methods associated with the kit. These materials can include any combination of the following: background information, a list of ingredients and their availability (such as purchasing information), a brief or detailed protocol for using the kit, troubleshooting, references, technical support, and any other related documentation. The instructions can be provided with the kit, or as a separate component in paper form, or in electronic form that can be provided on a computer-readable memory device or downloaded from an internet website, or as a recorded presentation. The instructions can include one or more documents and should include future updates.
[0025] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein and include proteins and fragments thereof. Polypeptides are disclosed herein as amino acid residue sequences. These sequences are written from left to right in the amino to carboxy terminus direction. According to standard nomenclature, amino acid residue sequences are designated by a three-letter or single-letter code as follows: alanine (Ala, A), arginine (Arg, R), asparagine (Asn, N), aspartic acid (Asp, D), cysteine (Cys, C), glutamine (Gln, Q), glutamic acid (Glu, E), glycine (Gly, G), histidine (His, H), isoleucine (Ile, I), leucine (Leu, L), lysine (Lys, K), methionine (Met, M), phenylalanine (Phe, F), proline (Pro, P), serine (Ser, S), threonine (Thr, T), tryptophan (Trp, W), tyrosine (Tyr, Y), and valine (Val, V). Peptides include any oligopeptide, polypeptide, gene product, expression product, or protein. Peptides consist of a sequence of amino acids and include naturally occurring or synthetic molecules.
[0026] Furthermore, as used herein, the term "peptide" refers to amino acids linked to each other by peptide bonds or modified peptide bonds, such as peptide isosteres, and can contain modified amino acids other than the 20 gene-encoded amino acids. Peptides can be modified by natural processes, such as post-translational processing, or by chemical modification techniques that are well known in the art. Modifications can occur anywhere in a peptide, including the peptide backbone, the amino acid side-chains, and the amino or carboxyl termini. The same type of modification can be present in the same or varying degrees at several sites in a given polypeptide. Furthermore, a given peptide can have many types of modifications. Modifications include, but are not limited to, linking different domains or motifs, acetylation, acylation, ADP-ribosylation, amidation, covalent cross-linking or cyclization, covalent attachment of flavin, covalent attachment of a heme moiety, covalent attachment of a nucleotide or nucleotide derivative, covalent attachment of a lipid or lipid derivative, covalent attachment of phosphatidylinositol, disulfide bond formation, demethylation, formation of cysteine or pyroglutamate, formylation, gamma-carboxylation, glycosylation, GPI anchor formation, hydroxylation, iodination, methylation, myristoylation, oxidation, PEGylation, proteolytic processing, phosphorylation, prenylation, racemization, selenation, sulfation, and transfer-RNA-mediated addition of amino acids to proteins such as arginylation. (See Proteins - Structure and Molecular Properties, 2nd ed., TE Creighton, WH Freeman and Company, New York (1993); Posttranslational Covalent Modification of Proteins, BC Johnson (ed.), Academic Press, New York, pp. 1-12 (1983)).
[0027] As used herein, "variant" refers to a molecule that retains the same or substantially similar biological activity as that of the original sequence. Variants can be from the same or different species, or can be synthetic sequences based on natural or previous molecules. Furthermore, as used herein, "variant" refers to a molecule that has a structure derived from the structure of a parent molecule (e.g., a protein or peptide disclosed herein) and whose structure or sequence is sufficiently similar to those disclosed herein that one of skill in the art would expect it to exhibit the same or similar activity and utility compared to the parent molecule based on that similarity. For example, substituting specific amino acids in a given peptide can result in a variant peptide that has similar activity as the parent.
[0028] In the context of the present invention, substitutions in mutant proteins are designated as [original amino acid / position in sequence / substituted amino acid]. For example, an asparagine (N) at position 36 of an amino acid sequence mutated to a histidine (H) is interchangeably designated as "N36H" or "N36 to H."
[0029] As used herein, the term "protein of interest (POI)" includes any synthetic or naturally occurring protein or peptide. Thus, the term encompasses compounds traditionally considered to be drugs, vaccines, and biopharmaceuticals, including molecules such as proteins and peptides. Examples of therapeutic agents are described in well-known references, such as the Merck Index (14th ed.), the Physicians' Desk Reference (64th ed.), and The Pharmacological Basis of Therapeutics (1st ed.), and include, but are not limited to, medicines; substances used to treat, prevent, diagnose, cure, or alleviate a disease or condition; substances that affect the structure or function of the body; or prodrugs that become biologically active or more active after being placed in a physiological environment.
[0030] As used herein, "isolated peptide" or "purified peptide" refers to a peptide (or fragment thereof) that is substantially free from substances with which the peptide is normally associated in nature or with which the peptide is associated in an artificial expression or production system, including, but not limited to, expression host cell lysate, growth medium components, buffer components, cell culture supernatant, or components of a synthetic in vitro translation system. The peptides disclosed herein or fragments thereof can be obtained, for example, by extraction from a natural source (e.g., mammalian cells), by expression of a recombinant nucleic acid encoding the peptide (e.g., in a cell or in a cell-free translation system), or by chemically synthesizing the peptide. Furthermore, peptide fragments can be obtained by any of these methods or by cleaving the full-length protein and / or peptide.
[0031] As used herein, the word "or" means any one member of a particular list and also includes any combination of members of that list.
[0032] As used herein, the phrase "nucleic acid" refers to a DNA or RNA or DNA-RNA hybrid, single- or double-stranded, sense or antisense, naturally occurring or synthetic oligonucleotide or polynucleotide capable of hybridizing to a complementary nucleic acid via Watson-Crick base pairing. Nucleic acids of the present invention can also contain nucleotide analogs (e.g., BrdU) and non-phosphodiester internucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids can include, but are not limited to, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, or any combination thereof.
[0033] As used herein, "isolated nucleic acid" or "purified nucleic acid" refers to DNA that is free of the genes that flank it in the naturally occurring genome of the organism from which the DNA of the invention is derived. Thus, the term includes recombinant DNA that is incorporated into a vector, such as a self-replicating plasmid or virus; or that is incorporated into the genomic DNA of a prokaryotic or eukaryotic organism (e.g., a transgene); or that exists as a separate molecule (e.g., cDNA or genomic or cDNA fragments produced by PCR, restriction enzyme digestion, or chemical or in vitro synthesis). It also includes recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence. The term "isolated nucleic acid" also refers to an RNA, e.g., mRNA molecule, that is encoded by an isolated DNA molecule, that is chemically synthesized, or that is separated from or substantially free of at least some cellular components, e.g., other types of RNA molecules or peptide molecules.
[0034] As used herein, "extein" refers to a portion of an intein-modified protein that is not part of the intein and that can be spliced or cleaved after excision of the intein.
[0035] "Intein" refers to an in-frame intervening sequence in a protein. An intein can catalyze its excision from itself by the post-translational protein splicing process, yielding a free intein and a mature protein. An intein can also catalyze cleavage of an intein-extein bond at the intein N-terminus or intein C-terminus, or both at the intein-extein terminus. As used herein, "intein" encompasses mini-inteins, modified or mutated inteins, and split inteins.
[0036] As used herein, the term "split intein" refers to any intein in which there are one or more peptide bond cleavages between the N-terminal and C-terminal intein segments, such that the N- and C-terminal intein segments are separate molecules that can be non-covalently recombined or reconstituted into an intein that is functional for a splicing or cleavage reaction. Any catalytically active intein or fragment thereof can be used to derive a split intein for use in the systems and methods disclosed herein. For example, in one embodiment, the split intein can be derived from a eukaryotic intein. In another embodiment, the split intein can be derived from a bacterial intein. In another embodiment, the split intein can be derived from an archaeal intein. Preferably, the split intein so derived has only the amino acid sequence essential for catalyzing a splicing reaction.
[0037] As used herein, "N-terminal intein segment" or "N-intein" refers to any intein sequence that includes an N-terminal amino acid sequence that, when combined with a corresponding C-terminal intein segment, is functional for splicing and / or cleavage reactions. Thus, an N-terminal intein segment also includes sequences that are spliced out when splicing occurs. An N-terminal intein segment can include sequences that are modified versions of the N-terminal portion of a naturally occurring (native) intein sequence. Non-intein residues can also be genetically fused to an intein segment to provide additional functionality, such as the ability to be affinity purified or covalently immobilized.
[0038] As used herein, a "C-terminal intein segment" or "C-intein" refers to any intein sequence that contains a C-terminal amino acid sequence that is functional for a splicing or cleavage reaction when combined with a corresponding N-terminal intein segment. In one embodiment, the C-terminal intein segment contains a sequence that is spliced out when splicing occurs. In another embodiment, the C-terminal intein segment is cleaved from the peptide sequence fused to its C-terminus. The sequence cleaved from the C-terminus of a C-terminal intein is referred to herein as the "protein of interest," which is discussed in further detail below. A C-terminal intein segment can contain a sequence that is a modified version of the C-terminal portion of a naturally occurring (native) intein sequence. For example, a C-terminal intein segment can contain additional and / or mutated amino acid residues, as long as the incorporation of such residues does not render the C-terminal intein segment non-functional for splicing or cleavage.
[0039] A consensus sequence is a DNA, RNA, or protein sequence that represents aligned, related sequences. The consensus sequence of related sequences can be defined in different ways, but is usually defined by the most common nucleotide or amino acid residue at each position. An example of a consensus sequence of the present invention is the N-intein consensus sequence of SEQ ID NO:6.
[0040] As used herein, the term "splicing" or "to splice" refers to the cutting out of a central portion of a polypeptide to form two or more smaller polypeptide molecules. In some cases, splicing also includes the fusing of two or more smaller polypeptides to form a new polypeptide. Splicing can also refer to the joining of two polypeptides encoded on two separate gene products through the action of split inteins.
[0041] As used herein, the term "cleavage" or "cleaving" refers to the division of a single polypeptide to form two or more smaller polypeptide molecules. In some cases, cleavage is mediated by the addition of an exogenous endopeptidase, which is often referred to as "proteolytic cleavage." In other cases, cleavage may be mediated by the endogenous activity of one or both of the peptide sequences being cleaved, which is often referred to as "autocleavage." Cleavage can also refer to the autocleavage of two polypeptides induced by the addition of a non-proteolytic third peptide, such as in the case of the action of the split intein system described herein.
[0042] The term "fused" means covalently linked. For example, a first peptide is fused to a second peptide when the two peptides are covalently linked to one another (e.g., via a peptide bond).
[0043] As used herein, an "isolated" or "substantially pure" material is one that has been separated from components that naturally accompany it. Generally, a polypeptide is substantially pure when it is at least 50%, by weight (e.g., 60%, 70%, 80%, 90%, 95%, and 99%) free from other proteins and naturally-occurring organic molecules with which it is naturally associated.
[0044] As used herein, "binding" or "binding" means that one molecule recognizes and attaches to another molecule in a sample, but does not substantially recognize or attach to other molecules in the sample. A molecule binds to another molecule at a binding affinity of about 10 5 ~10 6 A molecule "specifically binds" to another molecule if it has a binding affinity greater than liter / mole.
[0045] The nucleic acids, nucleotide sequences, proteins, or amino acid sequences referred to herein can be isolated, purified, chemically synthesized, or produced by recombinant DNA techniques, all of which methods are well known in the art.
[0046] As used herein, the terms "modified" or "mutated," as in "modified intein" or "mutated intein," refer to one or more modifications in a referenced nucleic acid or amino acid sequence, such as an intein, as compared to the native or naturally occurring structure. Such modifications may be substitutions, additions, or deletions. The modifications may be in one or more amino acid residues or one or more nucleotides of the referenced structure, such as an intein.
[0047] As used herein, the terms "modified peptide," "modified protein," or "modified protein of interest" or "modified target protein" refer to a modified protein.
[0048] As used herein, "operably linked" refers to the attachment of two or more biological molecules in a composition to one another in a manner that allows the biological molecules to perform their normal functions. With respect to nucleotide sequences, "operably linked" refers to the attachment of two or more nucleic acid sequences in a composition to one another, by enzymatic ligation or otherwise, in a manner that allows the sequences to perform their normal functions. For example, a nucleotide sequence encoding a presequence or secretory leader is operably linked to a nucleotide sequence of a polypeptide if it is expressed as a preprotein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the coding sequence; and a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation of the coding sequence.
[0049] "Sequence homology" can refer to the situation where nucleic acid or protein sequences are similar because they share a common evolutionary origin. "Sequence homology" can indicate that sequences are highly similar. Sequence similarity can be observable; homology can be based on observation. "Highly similar" can mean at least 70% identity, homology, or similarity; at least 75% identity, homology, or similarity; at least 80% identity, homology, or similarity; at least 85% identity, homology, or similarity; at least 90% identity, homology, or similarity; e.g., at least 93%, or at least 95%, or even at least 97% identity, homology, or similarity. Nucleotide sequence similarity, homology, or identity can be determined using the "Align" program of Myers et al. (1988) CABIOS 4:11-17, available at NCBI. Additionally, or alternatively, amino acid sequence similarity or identity or homology can be determined using the BlastP program (Altschul et al. Nucl. Acids Res. 25:3389-3402), available at NCBI. Alternatively, or additionally, the terms "similarity" or "identity" or "homology," for example, with respect to nucleotide sequences, indicate a quantitative measure of homology between two sequences.
[0050] Alternatively, or in addition, "similarity" with respect to sequences refers to the number of positions with identical nucleotides divided by the number of nucleotides in the shorter of the two sequences, where alignment of two sequences can be determined by the Wilbur and Lipman algorithm. (1983) Proc. Natl. Acad. Sci. USA 80:726. Computer-assisted analysis and interpretation of sequence data, including alignments, can be conveniently performed using commercially available programs (e.g., Intelligenetics™ Suite, Intelligenetics, Inc., CA) using, for example, a window size of 20 nucleotides, a word length of 4 nucleotides, and a gap penalty of 4. When RNA sequences are said to be similar or have a certain degree of sequence identity with a DNA sequence, thymidine (T) in the DNA sequence is considered equivalent to uracil (U) in the RNA sequence. The following references also provide algorithms for comparing the relative identity or homology or similarity of amino acid residues of two proteins, which can be used in addition to or instead of those described above to determine percent homology or identity or similarity. Needleman et al. (1970) J. Mol. Biol. 48:444~453; Smith et al. (1983) Advances App. Math. 2:482~489; Smith et al. (1981) Nuc. Acids Res. 11:2205~2220; Feng et al. (1987) J. Molec. Evol. 25:351~360; Higgins et al. (1989) CABIOS 5:151-153; Thompson et al. (1994) Nuc. Acids Res. 22:4673-480; and Devereux et al. (1984) 12:387-395."Stringent hybridization conditions" is a term well known in the art; see, e.g., Sambrook, Molecular Cloning, A Laboratory Manual, 2nd ed., CSH Press, Cold Spring Harbor, 1989; Nucleic Acid Hybridization, A Practical Approach, Hames and Higgins (eds.), IRL Press, Oxford, 1985; see also Figure 2, which contains the sequence comparison, and its description herein.
[0051] The terms "plasmid," "vector," and "cassette" refer to extrachromosomal elements that often carry genes that are not part of the cell's central metabolism and are usually in the form of circular double-stranded DNA molecules. Such elements can be linear or circular, self-replicating sequences of single- or double-stranded DNA or RNA from any source, genome-integrating sequences, phage, or nucleotide sequences in which several nucleotide sequences have been linked or recombined into a specific construct that allows the introduction of promoter fragments and DNA sequences for selected gene products, along with appropriate 3' untranslated sequences, into cells. Generally, a "vector" is a modified plasmid containing an "expression cassette" that contains multiple additional insertion sites for cloning and DNA sequences for selected gene products (i.e., transgenes) for expression in a host cell. This "expression cassette" generally includes a 5' promoter region, a transgene ORF, and a 3' terminator region, along with all necessary regulatory sequences required for the transcription and translation of the ORF. Thus, introduction of the expression cassette into a host allows expression of the transgene ORF in the cassette.
[0052] The term "buffer" or "buffered solution" refers to a solution that resists changes in pH by function of its conjugate acid-base range.
[0053] The term "loading buffer" or "equilibration buffer" refers to a buffer containing a salt or salts that is mixed with a protein preparation to load it onto a column. This buffer is also used to equilibrate the column before loading and to wash the column after loading the protein.
[0054] The term "wash buffer" is used herein to refer to a buffer that is run over a column after (for example) loading a protein of interest (e.g., coupled to a C-terminal intein fragment) and prior to elution of the protein of interest. The wash buffer can serve to remove one or more contaminants without substantially eluting the desired protein.
[0055] The term "elution buffer" refers to a buffer used to elute a desired protein from a column. As used herein, the term "solution" refers to a buffered or unbuffered solution that includes water.
[0056] The term "washing" means passing an appropriate buffer through or over a solid support, such as a chromatography resin.
[0057] The term "eluting" a molecule (eg, a protein of interest or a contaminant) from a solid support means removing the molecule from such material.
[0058] The term "contaminant" or "impurity" refers to any foreign or unwanted molecule other than the protein being purified, particularly a biological macromolecule such as DNA, RNA, or protein, present in a sample of the protein being purified. Contaminants include, for example, other proteins from the cells that express and / or secrete the protein being purified.
[0059] The terms "separating" or "isolating" as used in the context of protein purification refer to the separation of a desired protein from a second protein or a mixture of other contaminants or impurities in a mixture comprising the desired protein and the second protein or a mixture of other contaminants or impurities, such that at least the majority of the molecules of the desired protein are removed from that portion of the mixture comprising at least the majority of the molecules of the second protein or a mixture of other contaminants or impurities.
[0060] The terms "purifying" or "purifying" a desired protein from a composition or solution containing the desired protein and one or more contaminants means increasing the purity of the desired protein in the composition or solution by removing (fully or partially) at least one contaminant from the composition or solution.
[0061] N-intein protein mutants The present invention relates to a single-step affinity chromatography and affinity tag cleavage mechanism using the split intein system of the present invention, which cleaves with a wide range of amino acids to generate an untagged protein of interest (POI) as the final product. The two halves of the intein are the affinity ligand (N-intein) and the affinity tag (C-intein), which rapidly bind together. Immobilization of one half (N-intein) on a chromatography resin allows capture of the other half (C-intein) coupled to the POI from solution. Zn 2+ In the presence of ions, the cleavage reaction is inhibited, allowing stable complexes to form while impurities are washed away. After the impurities are removed, a chelator or reducing agent is added, allowing the cleavage reaction to proceed, allowing the POI to be collected while the intein tag remains noncovalently bound to its cognate intein, which is tethered to a chromatography resin.
[0062] Preferably, the present invention provides N-intein protein variant sequences of a naturally occurring split intein or a consensus sequence derived from a naturally occurring intein and a split intein, where the N-intein variant is modified to eliminate all asparagine (N) amino acid residues present in the sequence compared to the native sequence or consensus sequence. Preferably, not all such sequences include a cysteine (C) at position 1 of the N-intein variant sequence.
[0063] Preferably, the present invention provides N-intein protein variant sequences that do not contain an asparagine (N) at position 36 of the variant sequence. This position is calculated by conventional cluster alignment with naturally occurring split inteins starting from the first catalytic cysteine, numbered 1. While this position is conserved in the prior art and in naturally occurring N-intein sequences, the present inventors have found that this position can be mutated to an amino acid that provides increased alkaline stability compared to the naturally occurring N-intein protein sequence, which is important because it confers resistance to increasing pH values, for example, during chromatography procedures. Preferably, the amino acid that provides increased alkaline stability is histidine (H or His) or glutamine (Q or Gln).
[0064] Naturally occurring inteins are known in the art. A list of inteins is found in Table 1 below. All inteins can be made into split inteins, although some inteins naturally exist in split form. All of the inteins found in the table can exist as split inteins or can be made into split inteins modified according to the present invention at position 36 so that the conserved N is replaced with another amino acid that confers alkaline stability, such as H or Q.
[0065] [Table 1A]
[0066] [Table 1B]
[0067] Table 1C
[0068] Table 1D
[0069]
Table 1E
[0070]
Table 1F
[0071] Table 1G
[0072]
Table 1H
[0073]
Table 1I
[0074] Table 1J
[0075] Table 1K
[0076]
Table 1L
[0077] Table 1M
[0078]
Table 1N
[0079]
Table 1O
[0080]
Table 1P
[0081]
Table 1Q
[0082]
Table 1R
[0083]
Table 1S
[0084]
Table 1T
[0085]
Table 1U
[0086]
Table 1V
[0087]
Table 1W
[0088]
Table 1X
[0089]
Table 1Y
[0090]
Table 1Z
[0091]
Table 1AA
[0092]
Table 1BB
[0093]
Table 1CC
[0094] [Table 1DD]
[0095]
Table 1EE
[0096]
Table 1FF
[0097] Table 1GG
[0098] [Table 1HH]
[0099] [Table 1II]
[0100] [Table 1JJ]
[0101] [Table 1KK]
[0102] Split inteins that can be used in the disclosed compositions or methods can be modified or mutated inteins. Modified inteins can include modifications to the N-terminal intein segment, the C-terminal intein segment, or both. Modifications can include additional amino acids at the N-terminus, C-terminus, or within either portion of the split intein. Table 2 provides a list of amino acids, their abbreviations, polarity, and charge.
[0103] [Table 2]
[0104] Preferably, the present invention provides an N-intein protein variant of the native N-intein domain of Nostoc punctiforme (Npu), wherein the native N-intein domain has the following sequence: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRV (SEQ ID NO: 1) Here, the protein variant with an amino acid that increases the alkaline stability of the N-intein protein variant compared to the alkaline stability of the native N-intein of SEQ ID NO: 1 comprises an amino acid substitution of asparagine (N) at position 36 of SEQ ID NO: 1.
[0105] Preferably, the present invention provides an N-intein protein variant of SEQ ID NO: 1, wherein the protein variant comprises an amino acid substitution of asparagine (N) at position 36 of SEQ ID NO: 1 with an amino acid that increases the alkaline stability of the N-intein protein variant compared to the alkaline stability of the native N-intein of SEQ ID NO: 1, as well as an amino acid substitution of cysteine (C) at position 1 of SEQ ID NO: 1 with any other amino acid that is not cysteine.
[0106] The present invention also provides N-intein protein variants of a reference protein, wherein the reference protein has at least about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:1, preferably, the reference protein has at least 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:1, wherein the N-intein protein variants of the invention comprise an amino acid substitution of asparagine (N) at position 36 of the reference protein with an amino acid that increases the alkaline stability of the N-intein protein variant compared to the alkaline stability of the native N-intein of SEQ ID NO:1.
[0107] In another embodiment, the N-intein comprises the amino acid sequence of SEQ ID NO: 2, a sequence derived from the N-intein consensus. The N-intein variant sequence based on SEQ ID NO: 2 also contains an amino acid other than N at position 36, which increases the alkaline stability of the N-intein protein variant compared to the alkaline stability of the native N-intein of SEQ ID NO: 1. Preferably, the amino acid that increases alkaline stability is an amino acid that is less susceptible to deamidation compared to asparagine (N). The amino acid sequence of SEQ ID NO: 2 is as follows: ALSYDTEILTVEYGFLPIGXIVEEXIEXTVYSVDXXGFVYTQPIAQWHNRGEQEVFEYXLEDGSIIRATXDHXFMTTDGXMLPIDEIFEXGLDLXQV (SEQ ID NO: 2) During the ceremony X at positions 20, 35, 70, 73, and 95 are each independently selected from K, R, or A; X at position 28 is C, A or S; X at position 36 is N, H or Q; X at position 25 is N or R; X at position 59 is D or C; X at position 80 is E or Q; X at position 90 is Q, R or K.
[0108] Preferred embodiments of N-inteins according to the present invention are selected from the group of N-intein mutants referred to herein as A48, B22, B72 and A41, wherein: A48 has the sequence of SEQ ID NO:2, wherein: X at positions 20, 35, 70, 73 and 95 is R; The X in the 28th position is an A; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q; B22 has the sequence of SEQ ID NO:2, wherein: X at positions 20, 35, 70, 73 and 95 is A; The X in the 28th position is an A; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q; B72 has the sequence of SEQ ID NO:2, wherein: X at positions 20, 35, 70, 73 and 95 is K; X in 28th position is C; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q; A40 has the sequence of SEQ ID NO:2, wherein: X at positions 20, 35, 70, 73 and 95 is R; The X in the 28th position is an A; The X in the 36th position is N; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th place X is Q. A41 has the sequence of SEQ ID NO:2, wherein: X at positions 20, 35, 70, 73 and 95 is K; The X in the 28th position is an A; The X in the 36th position is N; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q; Comparative ligand A53 has the sequence of SEQ ID NO: 2, wherein: X at positions 20, 35, 70, 73 and 95 is K; X in 28th position is C; The X in the 36th position is N; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th place X is Q.
[0109] The N-inteins of the present invention can be coupled to a solid phase, such as a membrane, fiber, particle, bead, or chip. The solid phase can be a chromatography resin of natural or synthetic origin, such as a natural or synthetic resin, preferably a polysaccharide such as agarose. The solid phase, such as a chromatography resin, can be provided with embedded magnetic particles. In another embodiment, the solid phase is a non-diffusion-limiting resin / fiber material.
[0110] In this case, the solid phase can be formed from one or more polymer nanofiber substrates, such as electrospun polymer nanofibers. Polymer nanofibers for use in the present invention typically have an average diameter of 10 nm to 1000 nm. The length of the polymer nanofiber is not particularly limited. The polymer nanofiber may preferably be a monofilament nanofiber and may have, for example, a circular, ellipsoidal, or substantially circular / ellipsoidal cross section. Typically, one or more polymer nanofibers are provided in the form of one or more nonwoven sheets, each containing one or more polymer nanofibers. A nonwoven sheet containing one or more polymer nanofibers is a mat of said one or more polymer nanofibers, with each nanofiber essentially randomly oriented, i.e., it is not fabricated so that the nanofiber or nanofibers adopt a specific pattern. The nonwoven sheet typically has an areal density of 1 to 40 g / m2. The nonwoven sheet typically has a thickness of 5 to 120 μm. The polymer should be a polymer suitable as a chromatography medium, i.e., an adsorbent, in a chromatography method. Suitable polymers include polyamides such as nylon, polyacrylic acid, polymethacrylic acid, polyacrylonitrile, polystyrene, polysulfones such as polyethersulfone (PES), polycaprolactone, collagen, chitosan, polyethylene oxide, agarose, agarose acetate, cellulose, cellulose acetate, and combinations thereof.
[0111] The N-inteins according to the present invention can be immobilized to a very high degree on solid supports, with 0.2-2 μmol / ml N-intein coupled per ml of resin (swollen gel).
[0112] N-inteins according to the invention can be coupled to a solid phase via a Lys tail containing one or more Lys, e.g., at least two, at the C-terminus. Alternatively, the N-intein is coupled to a solid phase via a Cys tail on the C-terminus.
[0113] C-intein protein mutants Preferably, the present invention provides the following SEQ ID NO: 3: VKIVSRKSLGVQNVYDIGVEKDHNFLLANGLIASN (SEQ ID NO: 3) Alternatively, C-inteins comprising a sequence having at least 50%, 60%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, and preferably a sequence having at least 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, are also provided.
[0114] It is understood that the N-intein and C-intein selection can be from the same wild-type split intein (e.g., both from Npu) or mutants of the N- or C-intein, or can be selected from different wild-type split intein or consensus split intein sequences, as it has been discovered that the affinity of the N-fragment for different C-fragments (e.g., the Ssp C-fragment or mutant thereof with the Npu N-fragment or mutant thereof) still maintains sufficient binding affinity for use in the disclosed methods.
[0115] Vectors containing the intein mutants of the present invention In a third aspect, the present invention relates to a vector comprising the above-described C-intein of SEQ ID NO: 3 and a gene encoding a protein of interest (POI). Also disclosed herein are vectors comprising nucleic acids encoding a C-terminal intein segment, as well as cell lines comprising said vectors. As used herein, a plasmid or viral vector is an agent that transports the disclosed nucleic acids, such as those encoding a C-terminal intein segment and a peptide of interest, into cells without degradation and includes a promoter that drives gene expression in cells to which they can be delivered. In one example, the C-terminal intein segment and peptide of interest are derived from a virus or retrovirus. Retroviral vectors can carry larger gene payloads, i.e., transgenes or marker genes, than other viral vectors, and for this reason are commonly used vectors. However, they are not useful in non-proliferating cells. Adenoviral vectors are relatively stable, easy to handle, have high titers, can be delivered in aerosol formulations, and can transfect non-dividing cells. Poxviral vectors are large and have several sites for inserting genes; they are heat-stable and can be stored at room temperature.
[0116] Split intein system Preferably, the present invention provides a split intein system for affinity purification of a protein of interest (POI), comprising an N-intein and a C-intein as described above.
[0117] Preferably, the N-intein contains the N36H mutation for increased alkaline stability.
[0118] Preferably, the N-intein is bound to a solid phase, and the C-intein is co-expressed with the POI and used as a tag for affinity purification of the POI. The reverse is also possible, i.e., the C-intein is bound to a solid phase and the N-intein is used as a tag, although the former is preferred.
[0119] The alkaline stability of the N-intein ligand in the split intein system according to the present invention allows for regeneration after cleavage of the POI from the solid phase under alkaline conditions, such as 0.05-0.5 M NaOH. The solid phase can be regenerated up to 100 times.
[0120] In one embodiment, the C-intein and an additional tag are co-expressed with the POI. The additional tag may be any conventional chromatography tag, such as an IEX tag or an affinity tag.
[0121] How to purify a protein of interest (POI) In a fifth aspect, the present invention relates to a method for purifying a protein of interest (POI) using a split intein system according to the present invention, comprising binding of a C-intein and an N-intein at a neutral pH, such as 6-8, and in the presence of divalent cations (which impair spontaneous cleavage); washing the solid phase in the presence of divalent cations; adding a chelator that allows spontaneous cleavage between the C-intein and the POI; collecting the tag-free POI; and regenerating the solid phase under alkaline conditions, such as 0.5 M NaOH.
[0122] This protocol is suitable for proteins that are insensitive to Zn. The advantage is that it allows for a long contact time with the resin and a large sample volume. The sample loading can be carried out for a long time, for example, up to 1.5 hours.
[0123] According to the present invention, a yield of greater than 30%, preferably greater than 50%, and most preferably greater than 80% of the POI is achieved in less than 4 hours of cleavage.
[0124] The present invention allows for high ligand density when the N-intein is immobilized on a solid phase. Preferably, the N-intein is bound to a chromatography resin, such as agarose, or any other suitable resin for protein purification. The present invention makes it possible to achieve static binding capacities of 0.2-2 μmol / ml of C-intein-bound POI per ml of stationary resin.
[0125] Affinity tags The present invention also relates to a method for purifying a protein of interest (POI), comprising the steps of coexpressing the POI with a C-intein according to the present invention and an additional tag; binding the additional tag to its binding partner on a solid phase; cleaving the POI and C-intein; binding the C-intein to an N-intein bound to a solid phase at neutral pH and cleaving the bound C-intein and N-intein from the POI; and regenerating the solid phase under alkaline conditions, such as 0.5 M NaOH. The purpose of this twin tag is: increased purity (enabling dual affinity purification), solubility, and detectability.
[0126] An affinity tag can be a peptide or protein sequence cloned in-frame with a protein-coding sequence that alters the behavior of the protein. Affinity tags can be added to the N- or C-terminus of a protein, which can be used in methods for purifying proteins from cells. Cells expressing a peptide containing an affinity tag can be expressed with a signal sequence in the supernatant / cell culture medium. Cells expressing a peptide containing an affinity tag can be pelleted or lysed, and the cell lysate can be applied to a column, resin, or other solid support that presents a ligand to the affinity tag. The affinity tag and any fusion peptide are bound to the solid support, which can also be washed several times with buffer to remove unbound (contaminating) protein. If the protein of interest is bound to the affinity tag, it can be eluted from the solid support with a buffer that dissociates the affinity tag from the ligand and results in a purified protein, or it can be cleaved from the bound affinity tag using a soluble protease. As disclosed herein, the affinity tag is cleaved by the self-cleavage mechanism of the C-intein segment in the active intein complex.
[0127] Examples of affinities include, but are not limited to, maltose-binding proteins that can bind to immobilized maltose to facilitate purification of the fused target protein; chitin-binding proteins that can bind to immobilized chitin; glutathione S-transferases that can bind to immobilized glutathione; polyhistidines that can bind to immobilized chelated metals; and FLAG octapeptides that can bind to immobilized anti-FLAG antibodies.
[0128] Affinity tags can also be used to facilitate purification of proteins of interest using the disclosed modified peptides by a variety of methods, including, but not limited to, selective precipitation, ion exchange chromatography, binding to precipitable ligands, dialysis (by altering the size and / or charge of the target protein), and other highly selective separation methods.
[0129] In some embodiments, affinity tags can be used that do not actually bind to a ligand, but instead selectively precipitate or act as a ligand for an immobilized corresponding binding domain. In these cases, the tag is more commonly referred to as a purification tag. For example, the ELP tag selectively precipitates under specific salt and temperature conditions, allowing purification of the fusion peptide by centrifugation. Another example is an antibody Fc domain, which serves as a ligand for an immobilized Protein A or Protein G binding domain.
[0130] Protein of interest Target proteins for all protocols are: any recombinant protein, especially proteins requiring a native or near-native N-terminal sequence, e.g., therapeutic protein candidates, biopharmaceuticals, antibody fragments, antibody mimetics, protein scaffolds, enzymes, recombinant proteins or peptides, e.g., growth factors, cytokines, chemokines, hormones, antigen (viral, bacterial, yeast, mammalian) production, vaccine production, cell surface receptors, fusion proteins.
[0131] The invention will now be described in more detail with reference to some non-limiting examples and the accompanying drawings. [Example]
[0132] (Experiment 1: Alkaline stability of N-intein ligands of the present invention) N-intein ligands A40, A41, and A48 according to the present invention were immobilized on a Biacore™ CM5 sensor chip (Cytiva, Sweden) in sufficient amounts to provide an immobilization level of approximately 450 response units (RU) or greater. To monitor the relative binding capacity of C-intein-tagged POIs to the immobilized surface, 20 μg / ml of C-intein (SEQ ID NO: 3)-tagged green fluorescent protein (GFP) was flowed over the chip for 1 minute, and the signal intensity was noted. The surface was then cleaned in place (CIP), i.e., rinsed with 100 mM NaOH, 4 M guanidine-HCl at room temperature (22 ± 3°C) for 10 minutes. This was repeated for 50 cycles, and the alkaline stability of the immobilized ligand was monitored after each cycle as the relative loss of C-intein-tagged GFP binding capacity (signal intensity).
[0133] The results are shown in Figure 1 and demonstrate that ligand A48 (with the N36H mutation) has improved alkaline stability compared to ligands A41 and A40. The alkaline stability was further improved compared to the native sequence. Furthermore, the N36H mutation significantly improved alkaline stability compared to the wild-type Npu N-intein sequence (A52 with a C1A mutation compared to SEQ ID NO: 1). The relative residual binding capacity (%) after 50 CIP cycles was 55% for A40 and A41, while it was 69% for A48. The alkaline stability using 0.5 M NaOH is shown in Figure 5. Figure 5 shows the results for A40 and A48 over 20 cycles. Relative Residual Binding Capacity (%) CIP: 100 mM NaOH, 4 M Gdn-HCl for 2 min, followed by 0.5 M NaOH for 2 min.
[0134] (Experiment 2: Alkaline stability of the N-intein ligands of the present invention) Purified N-intein ligands A53, B72, B22, and A48 were immobilized on a Biacore™ CM5 sensor chip (Cytiva, Sweden) in sufficient amounts to provide an immobilization level of approximately 450 response units (RU) or greater. To monitor the relative binding capacity of the non-cleavable C-intein-tagged POI to the immobilized surface, 20 μg / ml of non-cleavable C-intein (SEQ ID NO: 3)-tagged IL-1b was flowed over the chip for 1 minute, and the signal intensity was noted. The surface was then cleaned in place (CIP), i.e., rinsed with 100 mM NaOH, 4 M guanidine-HCl at room temperature (22 ± 3°C) for 10 minutes. This was repeated for 50 cycles, and the alkaline stability of the immobilized ligand was monitored after each cycle as the relative loss of non-cleavable C-intein-tagged IL-1b binding capacity (signal intensity).
[0135] The results are shown in Figure 2 and show that all three ligands (A48, B22, and B72) with the N36H mutation have improved alkaline stability compared to ligand A53. The relative residual binding capacity (%) after 50 CIP cycles for A53 was only 20%, while for B72 it was 28%, for B22 it was 30%, and for A48 it was 35%.
[0136] (Experiment 3: Immobilization of N-intein ligand A48 on agarose gel resin) Five milliliters of epoxy-activated crosslinked gel resin was added to a polypropylene test tube. 2.7 milliliters, corresponding to 135 milligrams of N-intein ligand A48 with a C-terminal Lys tail in phosphate buffer, was added to the tube. Subsequently, 1.3 milliliters of phosphate buffer (pH 12.1) was added to adjust the agarose resin slurry to approximately 50%, followed by 2 grams of sodium sulfate. The pH of the resulting reaction mixture was adjusted to 11.5. The reaction mixture was heated to 33°C in a shaker and continued shaking at 33°C for 4 hours. The slurry was then transferred to a glass filter and washed three times with 10 milliliters of distilled water. After washing, the gel was transferred to a three-neck round-bottom flask (RBF), and 5 milliliters of Tris buffer (pH 8.6) and 375 microliters of thioglycerol were added. The reaction mixture was placed on a shaker at 45°C for 2 hours. After the reaction, the slurry was transferred to a glass filter. The gel was washed three times with 5 milliliters of basic wash buffer, then three times with 5 milliliters of acidic wash buffer. This base / acid wash was repeated two more times for a total of 18 washes. The gel resin was then washed 10 times with 5 milliliters of distilled water. Prior to analysis, the washed and drained gel was kept in 20% ethanol in a refrigerator.
[0137] The dry mass of the gel resin was determined by measuring the mass of 1 milliliter of gel. For sample preparation, 2 grams of drained gel resin was thoroughly mixed with 2 grams of water to give an approximately 50% resin slurry, and the slurry was then added to a 1 mL Teflon cube. A vacuum was then applied to drain the gel in the cube, thus obtaining 1 mL of gel. The gel was transferred onto a dry mass balance. The drying temperature was set to 105°C, and the mass was determined after 35 minutes.
[0138] After dry mass determination, amino acid analysis was performed. The corresponding dry mass of the protein and information on its size and primary amino acid sequence allowed the ligand density to be derived in mg per mL of gel resin. The result for the coupled agarose resin was a dry mass of 90.6 mg / mL and a ligand content of 18.4 mg / mL, corresponding to 1.38 μmol / mL.
[0139] (Experiment 4: Static binding capacity in relation to ligand density) The proposed capacity method presented herein allows for measuring the binding capacity of a resin in a test tube.
[0140] Reaction setup Briefly, prototype resins with immobilized A48 ligands with various ligand densities and dual-labeled test protein A43 (SEQ ID NO: 5) were separately diluted in assay buffer (2x PBS) to 2.5% resin slurry and 0.4 mg / mL, respectively. 50 μL of the 2.5% resin slurry was added to an ILLUSTRA™ microspin column, followed by 150 μL of diluted A43 (SEQ ID NO: 5). The reaction was incubated at 22°C for a fixed time of 2 hours with shaking at 1450 rpm, followed by centrifugation at 3000 rcf for 1 minute.
[0141] SDS-PAGE The centrifuged sample (containing cleaved and unbound, uncleaved protein) was mixed 1:1 with 2x SDS-PAGE reducing sample buffer, boiled at 95°C for 5 minutes, and subjected to SDS-PAGE (18 μL loaded). A C-intein-labeled test protein, A43 (SEQ ID NO: 5), was added as a standard (usually a five-point standard between 18.75 and 300 μg / mL) so that concentrations could be calculated from the densitometric volumes. The gel was Coomassie stained for 60 minutes (approximately 100 mL / gel) and subsequently destained for 120 to 180 minutes at room temperature under gentle agitation (until the background was completely clear). Densitometric quantification of uncleaved / unbound and cleaved test protein was performed using IQ TL software. The raw densitometric data was then exported to Microsoft Excel.
[0142] SBC calculation Since the test protein input to the reaction is known, the static binding capacity (SBC) can be calculated indirectly by the following formula:
[0143]
number
[0144] Figure 3 shows the static binding capacity of the N-intein ligands of the present invention. Amino acid analysis (AAA) performed by conventional methods. The A48 prototype was coupled to porous agarose particles by epoxy chemistry.
[0145] (Experiment 5: Purification of elongation factor G without and with the Zn protocol) Elongation factor G (Ef-G) from Thermoanaerobacter tengcongensis was purified in this example using a resin prototype with immobilized ligand A48. C-intein (SEQ ID NO: 3)-tagged Ef-G was expressed intracellularly in Escherichia coli (E. coli) strain BL21(DE3).
[0146] Frozen cell pellets harvested after fermentation were thawed and resuspended in extraction buffer (20 mM Tris-HCl, pH 8.0) by magnetic stirring. DNase I (bovine pancreas) and 1 mM MgSO4 were added, followed by lysozyme (hen's egg). After stirring for 30 minutes at room temperature, the resuspended, lysozyme-treated cell suspension was warmed to 70-75°C in a water bath and maintained at this temperature for 5 minutes. After briefly cooling the extract on ice, the extract was clarified by centrifugation.
[0147] Purification using the Zn-free protocol was performed on an AKTA™ Avant system at 2 ml / min during sample loading and washes, and 1 ml / min thereafter. A 1 ml HiTrap™ column containing immobilized A48 ligand was used. Equilibration and binding of the C-intein-labeled target protein were performed in 20 mM MES buffer supplemented with 100 mM NaCl at pH 6.3, and the sample was adjusted to pH 6.3 using 2 M acetic acid. Column washes after sample application and subsequent elution were performed with 20 mM Tris-HCl buffer supplemented with 400 mM NaCl at pH 8.0. After column wash, the flow was stopped for a 4-hour incubation at room temperature, after which the cleaved EfG was eluted. A second flow stop was added to allow for a second elution, which was performed after an additional 16 hours of incubation.
[0148] 17.8 mg of pure, untagged EfG was eluted from a HiTrap™ column after 4 hours of incubation. The mass difference between the eluted protein and the CIPed protein was equal to the mass of the C-intein tag by mass spectrometry. Purity by SDS-PAGE was similarly high as in SEC analysis on a Superdex™ 200 Increase column. Total protein amount was calculated from the theoretical ultraviolet absorption coefficient and UV signal at 280 nm in the diluted eluted and CIPed fractions.
[0149] The purification was repeated using the protocol with Zn ions in the equilibration buffer and the clarified sample. The final Zn concentration was 1.6 mM. The flow rate was reduced to 0.5 ml / min during sample application and then increased to 1 ml / min during wash and elution. Wash and elution were performed with 50 mM Tris-HCl, 20 mM imidazole buffer, pH 7.5. Only one elution peak was collected in this purification, which was after a 4-hour incubation period following column wash.
[0150] 16.6 mg of pure, untagged EfG was eluted from a HiTrap™ column after 4 hours of incubation. Purity was 92% by SEC analysis on Superdex™ 200 Increase. Total protein amount was calculated from the theoretical ultraviolet absorption coefficient and UV signal at 280 nm in the diluted elution fractions.
[0151] (Experiment 6: Purification of IL-1β) A 1 ml HiTrap™ column containing immobilized A48 ligand was used to purify the C-intein-tagged target protein IL-1β (SEQ ID NO: 5) expressed intracellularly in E. coli BL21(DE3) and solubilized by sonication. The soluble protein was collected by centrifugation and loaded onto a 1 ml HiTrap™ column with immobilized A48 ligand. During sample loading and washes, the Zn-free protocol (as in Experiment 4) was used on the AKTA™ Avant system at 4 ml / min (linear flow rate of 600 cm / hr). Flow was then stopped for 4 hours, after which flow was restarted at 1 ml / min to elute the cleaved protein (4-hour cleavage fraction). Flow was then stopped again for another 12 hours, after which flow was restarted at 1 ml / min to elute the uncleaved protein after 4 hours. Equilibration and binding for washes and elution were performed with a single buffer. The chromatogram from the purification is shown in Figure 4A. The starting material, flow-through, wash fractions, 4-hour and 16-hour elution fractions were subjected to SDS-PAGE and Coomassie staining and subsequent analysis using IQTL software (Figure 4B).
[0152] 9.4 mg of cleaved IL-1β was eluted from the HiTrap™ column after 4 hours of incubation, followed by an additional 1.1 mg eluted 16 hours later. Purity was 99.5% (4 hours) and 99.8% (16 hours) by SDS-PAGE analysis. Total protein content was calculated from the theoretical ultraviolet absorption coefficient of the cleaved protein at 280 nm.
[0153] (Experiment 7: Purification of the SARS-COV-2 receptor-binding domain) The SARS-COV-2 NCBI receptor binding domain (RBD) tagged with C-intein was expressed in ExpiHEK cells and secreted into the cell culture medium. Approximately 210 mL of cell culture supernatant was loaded onto a 1 mL HiTrap column with immobilized A48 ligand using an AKTA™ Avant FPLC system, without adding any salts or other additives. Sample application and washing were performed at 4 mL / min (load time approximately 52.5 min (linear flow rate of 600 cm / hr)), followed by a 6 column volume wash and a 4-hour rest / hold step. The elution phase was performed at 1 mL / min. The column was left for an additional 68 hours, after which a second elution followed. A single 40 mM phosphate buffer, pH 7.4, supplemented with 300 mM NaCl, was used for all chromatographic steps.
[0154] To determine protein concentration and yield, a 0.1% theoretical absorbance factor was used in Unicorn™ software (Cytiva, Sweden). Purity was determined by densitometric SDS-PAGE analysis. For this experiment, a total of 14.1 mg of cleaved protein was obtained with a purity of >96%. The theoretical molecular weight was approximately 25 kDa, but experimental SDS-PAGE analysis indicated a molecular weight of 33 kDa, which was accounted for by two glycosylations, also determined by mass spectrometry.
[0155] The CCT-RBD protein has the following sequence:
[0156] [ka]
[0157] Signal sequence - bold underlined. CCT tag - dotted underline. The RBD domain is double underlined. His tag - dashed underline
[0158] The purity results from the cleaved proteins are found in Table 3.
[0159] [Table 3]
[0160] (Experiment 8: Tandem labeling and affinity purification on two columns) E. coli BL21(DE3) was transformed with the A43 expression plasmid TwinStrep™ and C-intein (SEQ ID NO: 3)-tagged IL-1b and plated on agar plates containing 50 μg / ml kanamycin. The next day, single colonies were picked and grown in 5 ml of Luria-Bertani (LB) broth to an OD of 0.6. The culture was transferred to 200 ml of LB broth containing the same antibiotic and grown at 37°C to an OD of 0.6. Protein expression was induced by the addition of isopropyl bD-1-thiogalactopyranoside (IPTG, 0.5 mM) at 22°C for 16 hours. After expression, cells were harvested by centrifugation at 4,000 × g for 15 minutes and stored at -80°C until use.
[0161] For purification, cell pellets were resuspended in Buffer A1 (100 mM Tris-HCl, 150 mM NaCl, 1 mM EDTA, pH 8.0) at 10 ml per gram of wet mass and disrupted by sonication (Sonics Vibracell, fine tip, 30% amplitude, 2 seconds on, 4 seconds off, total 3 minutes).
[0162] After centrifugation at 40,000 × g for 20 min at 4 °C, the supernatant containing the soluble fraction was collected and passed through a 5 ml HiTrap™ column, Streptactin™ XT (GE Healthcare, Sweden). The column was washed with the same buffer A1 until the UV absorbance at 280 nm was less than 20 mAU. The bound C-intein-tagged IL-1b was eluted with buffer B1 (100 mM Tris-HCl, 150 mM NaCl, 1 mM EDTA, 50 mM biotin, pH 8.0) and collected.
[0163] The purified protein was immediately applied to a 1 ml HiTrap™ column packed with resin containing immobilized N-intein ligand A48 without the addition of the inhibitor ZnCl. Cleaved, tag-free IL-1b was collected in the flow-through.
[0164] [ka]
[0165] TwinStrep-dotted underline CCT - Bold Underline IL1b (test protein) - underlined
[0166] The patent and scientific literature referred to herein establishes knowledge that is available to those skilled in the art. All U.S. patents and published or unpublished U.S. patent applications cited herein are incorporated by reference. All published foreign patents and patent applications cited herein are hereby incorporated by reference. All other published references, documents, manuscripts, and scientific literature cited herein are hereby incorporated by reference.
[0167] While the present invention has been shown and described in detail with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail can be made therein without departing from the scope of the invention as encompassed by the appended claims. It should also be understood that the embodiments described herein are not exclusive, and that features from various embodiments can be combined in whole or in part in accordance with the present invention.
Claims
1. The following array: ALSYDTEILTVEYGFLPIGXIVEEXIEXTVYSVDXXGFVYTQPIAQWHNRGEQEVFEYXLEDGSIIRATXDHXFMTTDGXMLPIDEIFEXGLDLXQV (SEQ ID NO: 2) 1. An N-intein protein mutant comprising an amino acid sequence comprising: During the ceremony, X at positions 20, 35, 70, 73, and 95 are each independently selected from K, R, or A; X at position 28 is C or A; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q, alkaline stability is increased compared to the wild-type N-intein domain of Nostoc punctiforme (Npu) consisting of the amino acid sequence of SEQ ID NO: 1; N-intein protein mutants.
2. X at positions 20, 35, 70, 73 and 95 is R; or X at positions 20, 35, 70, 73 and 95 is A; The X in the 28th position is an A; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q. The N-intein protein mutant of claim 1.
3. X at positions 20, 35, 70, 73 and 95 is K; X in 28th position is C; X in 36th position is H; The 25th place X is N; X in 59th place is D; The X in the 80th place is E; The 90th X is Q. The N-intein protein mutant of claim 1.
4. 4. The N-intein protein mutant of claim 1, which is coupled to a solid phase.
5. The N-intein protein mutant of claim 4, wherein the solid phase is a chromatography resin.
6. 5. The N-intein protein variant of claim 4, wherein the solid phase is a chromatography resin of natural or synthetic origin and the solid phase is provided with embedded magnetic particles.
7. 5. The N-intein protein variant of claim 4, wherein the solid phase is a non-diffusion-limiting resin / fibrous material.
8. 5. The N-intein protein mutant of claim 4, wherein the N-intein is coupled to a solid phase at its C-terminus via a Lys tail or Cys tail containing one or more Lys.
9. 4. A split intein system for affinity purification of a protein of interest (POI), comprising the N-intein protein mutant of any one of claims 1 to 3.
10. 10. The split intein system of claim 9, comprising a C intein mutant comprising the amino acid sequence of SEQ ID NO:
3.
11. 10. Use of the split intein system of claim 9 for affinity purification of a protein of interest (POI), wherein the split intein system comprises a C-intein, the C-intein is co-expressed with the POI, the N-intein protein mutant is immobilized on a solid phase, and the solid phase is regenerated under alkaline conditions after cleavage of the POI from the solid phase.
12. The use described in claim 11, wherein the split intein system comprises an additional tag, and the C-intein and the additional tag are co-expressed with the POI.
13. 10. A chromatography column comprising a chromatography resin comprising one or more N-intein protein variant ligands, wherein the N-intein protein variants are as defined in one or more of claims 1 to 3.
14. 10. A method for purifying a protein of interest (POI) labeled with a C-intein using the split intein system of claim 9, wherein the N-intein protein mutant is immobilized on a solid phase, the method comprising the steps of contacting the C-intein and the N-intein protein mutant at neutral pH and in the presence of divalent cations; washing the solid phase in the presence of divalent cations; adding a chelator that enables spontaneous cleavage between the C-intein and the POI; collecting the untagged POI; and regenerating the solid phase under alkaline conditions.
Citation Information
Patent Citations
Split Intein and its Use
JP2014528720A
Split-inteins, complexes, and their uses
JP2015522020A
Intein-mediated protein purification
JP2016504417A
chromatography carrier
JP2016534373A
Soluble intein fusion protein and method for purifying biomolecules
JP2017533701A