TEV proteases with amphiphilic and tagging
By fusing CBD and His tags in TEV proteases, the problem of difficulty in purifying target proteins with high histidine content in the prior art is solved, and a wider purification application and more efficient purification effect are achieved.
Patent Information
- Application Number
- CN202510099701.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-22
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the CBD tag has not been combined with the His tag for purification of TEV protease, making it difficult to efficiently purify target proteins with high histidine content, limiting the scope of application of TEV protease.
A TEV protease was designed to fuse chitin-binding domain (CBD) tags and histidine tags to form parent and tags, extending the application areas of TEV proteases to enable purification by chitin- and nickel binding.
The extensive purification of target proteins with high histidine content has been achieved, the application range of TEV proteases has been expanded, and the purification efficiency and flexibility have been improved.
Smart Images

Figure CN120400286A_ABST
Abstract
Description
Background Art
[0001] Protein proteases are a group of enzymes that cleave proteins and peptides. Proteases are widely used in industries and biotechnology including the production of Klenow fragments, peptide synthesis, digestion of unwanted proteins during nucleic acid purification, cell culture, and tissue dissociation.
[0002] The protease trypsin cleaves peptides at specific sites. Proteinase K cleaves in a non-specific manner. TEV protease is a widely used cysteine protease derived from tobacco etch virus that recognizes and efficiently cleaves at the cleavage site between Q and S in the protein sequence E-N-L-Y-F-Q / S (SEQ ID NO:7). TEV protease can also recognize and cleave between Q and X in the sequence E-N-L-Y-F-Q / X (SEQ ID NO:8), where X can be any one of the amino acids G, A, M, C, or H.
[0003] The chitin-binding domain (CBD) is a polypeptide that specifically binds N-acetylglucosamine. CBD is derived from a small domain of the chitinase A1 gene of Bacillus circulans. The CBD binding affinity is very high, so much so that it is almost irreversible, making it useful for tagging proteins for purification purposes. CBD binds tightly to chitin, which is poly-N-acetylglucosamine and is one of the most common polymers in nature. Chitin is present in the shells of all crustaceans, the exoskeletons of insects, and many fungi, algae, and yeasts. Chitin is commercially available as resin, as chitin beads, or as chitin magnetic beads (New England Biolabs, Ipswich, MA). Proteins expressed with a CBD tag or fusion proteins in which CBD is fused to the target protein can be purified using chitin beads or other immobilized chitin. Proteins immobilized to chitin resin or chitin beads via the CBD tag can be separated from crude cell lysates by an affinity purification protocol. The purified protein is eluted from the CBD by chemically induced internal cleavage between the tag and the fusion protein.
[0004] Another commonly used tag in affinity chromatography is the polyhistidine tag (His tag), which involves adding an uninterrupted string of about four to ten histidine residues to the N- or C-terminus of the target protein. Proteins with a His tag can bind to metal ions immobilized on a carrier such as resin or magnetic beads. After purification, increasing concentrations of imidazole can be used to elute the bound protein, as imidazole competes with the target protein for binding to the metal ions, thus displacing the purified protein into solution for recovery.
[0005] Methionine is the typical starting amino acid for each whole protein. Proteins can be engineered to present an N-terminal amino acid tag such that the starting methionine appears immediately after the TEV cleavage site, e.g., E-N-L-Y-F-Q / M (SEQ ID NO:9). TEV protease can cleanly cleave off the amino acid tag, leaving the intact protein without any additional unwanted N-terminal amino acid residues. However, after digestion, TEV is usually removed because the polyhistidine tag fused to TEV is part of the TEV fusion protein. The TEV fusion protein can be removed by binding to nickel beads or nickel magnetic beads via the polyhistidine tag. However, if the target protein also has a His tag, or if the target also binds to nickel beads due to a high intrinsic histidine content, it is desirable to have another method to remove it after the TEV protease has carried out the required cleavage reaction.
[0006] To date, the combination of a CBD tag and a His tag has not been used in a purification scheme for purifying TEV protease that retains both the His tag and the CBD tag after purification and uses the bound nickel and / or chitin for purification by binding to the polyhistidine tag or the CBD tag, respectively. Engineering the TEV protease to display both a 6-mer His tag and a CBD tag broadens the application field compared to TEV protease with only a His tag, thus enabling the purification of a wider range of targets (including targets with a high histidine content). Summary of the Invention
[0007] The present invention relates to a TEV protease (wild-type or mutant) that displays dual affinity tags, including a chitin binding domain (CBD) tag and a histidine tag, preferably where the CBD tag is before the N-terminus of the protease and the histidine tag is before the CBD tag, and the histidine tag is preferably a 6-mer histidine tag; and optionally, a first linker, preferably a Gly-Ser linker, and more preferably a 6-mer Gly-Ser linker, is also included between the tags. A second linker, also preferably a Gly-Ser linker, can also be after the CBD tag and before the protease portion. In certain embodiments, no such linker is present, while in other embodiments, only one such linker is present.
[0008] The present invention also includes the amino acid sequence of the fusion protein, wherein: the wild-type TEV protease displaying the dual affinity tags is arranged such that the chitin-binding domain (CBD) tag is before the N-terminus of the protease, and the His tag (preferably a 6-mer histidine tag) is before the CBD tag, and preferably a Gly-Ser linker (preferably a 6-mer Gly-Ser linker) is between the tags; and the amino acid sequence of such a fusion protein, wherein the TEV protease in the fusion protein is a mutant TEV protease whose amino acid sequence has only conservative substitutions such that it has at least 70%, 80%, 90%, 95%, 96%, 97%, 98% or 99% identity with the wild-type TEV protease mutant amino acid sequence (such mutant TEV protease sequences are hereinafter referred to as "variant sequences").
[0009] The present invention also includes RNA or DNA (collectively referred to as "degenerate nucleic acid sequences") encoding any fusion protein, the fusion protein including the fusion protein containing the variant sequence (collectively referred to as "fusion proteins" or individually as "fusion protein").
[0010] The present invention also includes a vector bound to any degenerate nucleic acid sequence; and a cell transformed with any such vector or degenerate nucleic acid sequence and capable of expressing any fusion protein (including the fusion protein containing the variant sequence).
[0011] The present invention also includes a composition or kit that encodes any fusion protein, or contains any degenerate nucleic acid sequence or a vector bound to such a degenerate nucleic acid sequence. The present invention also includes a method of using one or more fusion proteins in a reaction mixture designed to cleave a target protein to cleave the target protein.
[0012] The present invention also includes using the fusion protein to remove TEV protease from a reaction mixture, the reaction mixture including a reaction mixture in which the target or other product in the reaction mixture has a high histidine content.
[0013] Based on the following detailed description and the drawings, additional aspects and advantages of the present disclosure will become apparent to those skilled in the art. Only illustrative embodiments of the present disclosure are shown and described in the following detailed implementation. The present disclosure can have other different embodiments, and several details thereof can be modified in various obvious aspects, all of which do not depart from the present disclosure. Therefore, the descriptions and examples in this summary of the invention are only for illustrative purposes and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1: The construction scheme of His-CBD-TEV starts from the N-terminus as follows: a 6-mer His tag, followed by a CBD tag, and then TEV protease. The linker is not shown, but if present, it is preferably between the His tag and the CBD tag and / or after the CBD tag and before the protease.
[0015] Figure 2 : A Bis-tris protein gel showing 1 μg of eGFP (tagged MBP-TEV site-eGFP) digested by His-CBD-TEV in 2-fold serial dilutions in 1X TEV protease buffer (New England Biolabs, Inc) with imidazole added at 30 °C for 30 minutes. Lanes 1 to 11 show the digestion of MBP-TEV site-eGFP by His-CBD-TEV protease, with the protease serially diluted 2-fold from 0.32 μg in lane 1 to decreasing concentrations until lane 11. For each digestion lane, four main bands are visible: the top band is full-length eCFP with MBP attached via the TEV site; the next band below is the MBP tag after cleavage from the eGFP protein; the second lowest band is the His-CBD-TEV protease, which weakens from left to right as the protease concentration decreases; and the lowest band represents eGFP after cleavage of the MBP tag at the TEV site. Lane 12 is ColorMixed Protein Marker 180 (10 - 180 kDa) (ABclonal, Woburn MA).
[0016] Figure 3 : A Bis-tris protein gel of TEV protease with a CBD / 6-mer His tag, demonstrating the binding ability to nickel magnetic beads or chitin magnetic beads. Lane 1 is ColorMixed Protein Marker 180 (10 - 180 kDa) protein marker (ABclonal, Woburn MA). Lane 2 is 1.23 μg of purified His-CBD-TEV protease. Lane 3 shows the sample containing His-CBD-TEV protease after incubation with 10 μl of nickel magnetic beads. Lane 4 is the same sample after incubation with 10 μl of chitin magnetic beads. Detailed Description
[0017] The non-limiting embodiments shown in the reference drawings and detailed in the following description more fully explain the embodiments herein and their various features and advantageous details. Without departing from the present invention, various variations, changes, and substitutions can be envisioned by those skilled in the art. It should be understood that various alternatives can be employed for the embodiments of the present disclosure.
[0018] Initially, for ease of reference, certain terms used in this application and their meanings as used in context are set forth. If a term used herein is not defined below, the broadest definition given to the term by persons in the relevant art shall be given, as reflected in at least one printed publication or issued patent. Additionally, the technology is not limited by the use of the terms shown below, as all equivalents, synonyms, newly developed, and terms or technologies used for the same or similar purposes are considered to be within the scope of the claims of the present invention.
[0019] As used herein, the articles "a" and "an" when applied to any feature in the embodiments of the invention described in the specification and claims mean one or more. The use of "a" and "an" does not limit their meaning to a single feature unless such a limitation is expressly stated. The article "the" before a singular or plural noun or noun phrase denotes one or more specifically identified features and may have a singular or plural meaning depending on the context in which it is used. The adjective "any" means one, some, or all, regardless of quantity.
[0020] The term "bioactive fragment" refers to any fragment, derivative, homolog, or analogue of a TEV protease mutant that has in vivo or in vitro activity characteristic of a biomolecule; including, for example, protease activity. In some embodiments, the bioactive fragment, derivative, homolog, or analogue of the mutant TEV protease has any degree of the bioactivity of the mutant TEV protease in any in vivo or in vitro assay of interest.
[0021] In some embodiments, the bioactive fragment can optionally include any number of contiguous amino acid residues of the mutant TEV protease. The present invention also includes polynucleotides encoding any such bioactive fragment.
[0022] The bioactive fragment can result from post-transcriptional processing or translation of alternatively spliced RNA, or alternatively can be produced by engineering, bulk synthesis, or other suitable operations. Bioactive fragments include fragments expressed in native or endogenous cells, as well as fragments produced in expression systems such as bacteria, yeast, plants, insects, or mammalian cells.
[0023] As used herein, the phrase "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid by another amino acid having common properties. A functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz (1979) Principles of Protein Structure, Springer-Verlag). Based on such analysis, groups of amino acids can be defined such that amino acids within a group preferentially exchange with each other and are thus most similar to each other in terms of their effect on the overall protein structure (Schulz (1979) supra). Examples of amino acid groups defined in this way can include: "charged / polar group", including Glu, Asp, Asn, Gln, Lys, Arg, and His; "aromatic or cyclic group", including Pro, Phe, Tyr, and Trp; and "aliphatic group", including Gly, Ala, Val, Leu, Ile, Met, Ser, Thr, and Cys. Within each group, subgroups can also be identified. For example, the group of charged / polar amino acids can be subdivided into multiple subgroups, including: "positively charged subgroup", including Lys, Arg, and His; "negatively charged subgroup", including Glu and Asp; and "polar subgroup", including Asn and Gln. In another example, the aromatic or cyclic group can be subdivided into multiple subgroups, including: "nitrogen ring subgroup", including Pro, His, and Trp; "phenyl subgroup", including Phe and Tyr. In yet another further example, the aliphatic group can be subdivided into multiple subgroups, including: "large aliphatic non-polar subgroup", including Val, Leu, and Ile; "aliphatic weakly polar subgroup", including Met, Ser, Thr, and Cys; and "small residue subgroup", including Gly and Ala. Examples of conservative mutations include amino acid substitutions within the above-described subgroups, such as but not limited to: Lys substituting for Arg or vice versa such that a positive charge can be maintained; Glu substituting for Asp or vice versa such that a negative charge can be maintained; Ser substituting for Thr or vice versa such that a free -OH can be maintained; and Gln substituting for Asn or vice versa such that a free -NH2 can be maintained. A "conservative variant" is a polypeptide that includes one or more amino acids that have been substituted to replace one or more amino acids of a reference polypeptide (e.g., a polypeptide whose sequence is disclosed in a publication or sequence database or whose sequence has been determined by nucleic acid sequencing) with amino acids having common properties (e.g., belonging to the same amino acid group or subgroup as described above).
[0024] When referring to a gene, "mutant / mutant form" means that the gene has at least one base (nucleotide) change, deletion, or insertion relative to the native or wild-type gene. The mutation (change, deletion, and / or insertion of one or more nucleotides) can be in the coding region of the gene, or can be in an intron, 3'UTR, 5'UTR, or promoter region. As a non-limiting example, a mutant gene can be a gene having an insertion within the promoter region that can increase or decrease the expression of the gene; can be a gene having a deletion that results in the production of a non-functional protein, a truncated protein, a dominant-negative protein, or no protein; or, can be a gene having one or more point mutations that result in an amino acid change in the encoded protein or in aberrant splicing of the gene transcript.
[0025] When used in the Detailed Description section, depending on the context, the terms "mutant TEV protease of the present invention" and "mutant TEV protease" refer jointly or separately to mutants that exhibit protease activity and / or mutants having variant sequences and / or degenerate nucleic acid sequences, as defined in the Summary of the Invention section. The term "fusion protein" in this Detailed Description section refers to the use of this term in the Summary of the Invention section.
[0026] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence that exists in an organism and has not been deliberately modified by human manipulation.
[0027] The terms "percent identity" or "homology" with respect to a nucleic acid or polypeptide sequence are defined as the percentage of nucleotide or amino acid residues in a candidate sequence that are identical to a known polypeptide after aligning the sequences to obtain the maximum percent identity and introducing gaps (if necessary) to achieve the maximum percent homology. N-terminal or C-terminal insertions or deletions should not be construed as affecting homology. Homology or identity at the nucleotide or amino acid sequence level can be determined by BLAST (Basic Local Alignment Search Tool) analysis, which uses the algorithms employed by the programs blastp, blastn, blastx, tblastn, and tblastx (Altschul (1997), Nucleic Acids Res. 25, 3389-3402 and Karlin (1990), Proc. Natl. Acad. Sci. USA 87, 2264-2268), which are customized for sequence similarity searches. The method used by the BLAST programs is to first consider similar segments (with or without gaps) between the query sequence and the database sequence, then to evaluate the statistical significance of all the identified matches, and finally to summarize only those matches that meet a preselected significance threshold. For a discussion of the basic issues in sequence database similarity searches, see Altschul (1994), Nature Genetics 6, 119-129. The search parameters for histograms, descriptions, alignments, expectations (i.e., the statistical significance threshold used to report matches to database sequences), cutoffs, matrices, and filters (low complexity) can be the default settings. The default scoring matrix used by blastp, blastx, tblastn, and tblastx is the BLOSUM62 matrix (Henikoff (1992), Proc. Natl. Acad. Sci. USA 89, 10915-10919), which is recommended for query sequences longer than 85 units (nucleotide bases or amino acids).
[0028] Preparation of a fusion protein
[0029] The fusion proteins of the present invention can be expressed in any suitable host system, including bacterial, yeast, fungal, baculovirus, plant or mammalian host cells. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the following: the Escherichia coli (E. coli) lactose operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis α-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens α-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and the prokaryotic β-lactamase gene (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA 75:3727-3731), as well as the tac promoter (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA 80:21-25).
[0030] For filamentous fungal host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the following genes: Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α-amylase, Aspergillus niger acid-stable α-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the Aspergillus niger neutral α-amylase and Aspergillus oryzae triose phosphate isomerase genes), and its mutant, truncated, and hybrid promoters.
[0031] In yeast hosts, useful promoters can be from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters of yeast host cells are described by Romanos et al., 1992, Yeast 8:423-488.
[0032] For baculovirus expression, insect cell lines from Lepidoptera (moths and butterflies), such as Spodoptera frugiperda, are used as hosts. Gene expression is controlled by strong promoters (e.g., pPolh).
[0033] Plant expression vectors are based on the Ti plasmid of Agrobacterium tumefaciens, or on tobacco mosaic virus (TMV), potato virus X, or cowpea mosaic virus. A commonly used constitutive promoter in plant expression vectors is the cauliflower mosaic virus (CaMV) 35S promoter.
[0034] For mammalian expression, cultured mammalian cell lines such as Chinese hamster ovary (CHO), COS, including human cell lines such as HEK and HeLa, can be used to produce fusion proteins. Examples of mammalian expression vectors include adenovirus vectors, pSV and pCMV series plasmid vectors, vaccinia virus, and retroviral vectors, as well as baculovirus. Promoters of cytomegalovirus (CMV) and SV40 are commonly used in mammalian expression vectors to drive gene expression. Non-viral promoters, such as elongation factor (EF)-1 promoter, are also known.
[0035] The control sequence for expression can also be a suitable transcription terminator sequence, i.e., a sequence recognized by the host cell to terminate transcription. The terminator sequence is operably linked to the 3' end of the nucleic acid sequence encoding the polypeptide. Any terminator functional in the selected host cell can be used.
[0036] For example, exemplary transcription terminators for filamentous fungal host cells can be obtained from the following genes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger α-glucosidase, and Fusarium oxysporum trypsin-like protease.
[0037] Exemplary terminators for yeast host cells can be obtained from the following genes: Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Terminators for insect, plant, and mammalian host cells are also well known.
[0038] The control sequence may also be a suitable leader sequence, i.e., an untranslated region of the mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that is functional in the selected host cell can be used. Exemplary leader sequences for filamentous fungal host cells are obtained from the genes for Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leader sequences for yeast host cells are obtained from the genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0039] The control sequence may also be a polyadenylation sequence, which is a sequence that is operably linked to the 3'-end of the nucleic acid sequence and is recognized by the host cell as a signal to add polyadenylate residues to the transcribed mRNA when transcribed. Any polyadenylation sequence that is functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells may be derived from the genes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger α-glucosidase.
[0040] The control sequence may also be a signal peptide coding region that encodes an amino acid sequence linked to the amino terminus of the polypeptide and directs the encoded polypeptide into the secretory pathway of the cell. The 5'-end of the coding sequence of the nucleic acid sequence may inherently contain a signal peptide coding region that is naturally linked in the translation reading frame to a segment of the coding region encoding the secreted polypeptide. Alternatively, the 5'-end of the coding sequence may contain a signal peptide coding region that is foreign to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding region, a foreign signal peptide coding region may be required.
[0041] Alternatively, the foreign signal peptide coding region may simply replace the native signal peptide coding region to enhance polypeptide secretion. However, any signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the selected host cell can be used.
[0042] Effective signal peptide coding regions for bacterial host cells are signal peptide coding regions obtained from the genes: Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus α-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are described by Simonen and Palva, 1993, Microbiol Rev 57:109-137.
[0043] An effective signal peptide coding region for a filamentous fungal host cell can be a signal peptide coding region obtained from the genes of: Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.
[0044] Useful signal peptides for yeast host cells can be from the genes of Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase. Signal peptides for other host cell systems are also well-known.
[0045] The control sequence can also be a propeptide coding region encoding an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a proenzyme or a pro-polypeptide (or in some cases a zymogen). The pro-polypeptide is usually inactive and can be converted into a mature active polypeptide by catalytic cleavage or autocatalytic cleavage of the propeptide from the pro-polypeptide. The propeptide coding region can be obtained from the genes of: Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae α-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (WO 95 / 33836).
[0046] In the case where both a signal peptide region and a propeptide region are present at the amino terminus of the polypeptide, the propeptide region is located adjacent to the amino terminus of the polypeptide, and the signal peptide region is located adjacent to the amino terminus of the propeptide region.
[0047] It is also desirable to add regulatory sequences that allow the expression of the fusion protein to be regulated relative to the growth of the host cell. Examples of regulatory systems are those that cause gene expression to be turned on or off in response to chemical or physical stimuli (including the presence of regulatory compounds). In prokaryotic host cells, suitable regulatory sequences include the lac, tac, and trp operon systems. In yeast host cells, suitable regulatory systems include, for example, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include the TAKA α-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter. Regulatory systems for other host cells are also well-known.
[0048] Other examples of regulatory sequences are those that allow gene amplification. In eukaryotic systems, these include the dihydrofolate reductase gene amplified in the presence of methotrexate and the metallothionein gene amplified with heavy metals. In these cases, the nucleic acid sequence encoding the polypeptide of the present invention will be operably linked to the regulatory sequence.
[0049] Another embodiment includes a recombinant expression vector that contains a polynucleotide encoding an engineered mutant fusion protein, and one or more expression regulatory regions such as a promoter and a terminator, as well as an origin of replication, depending on the type of host into which they are to be introduced. The various nucleic acids and control sequences described above can be ligated together to generate a recombinant expression vector that can include one or more convenient restriction sites to allow the insertion or substitution of the nucleic acid sequence encoding the fusion protein at such sites. Alternatively, the nucleic acid sequence of the fusion protein can be expressed by inserting the nucleic acid sequence or a nucleic acid construct containing the sequence into an appropriate vector for expression. When generating the expression vector, the coding sequence is positioned in the vector such that the coding sequence is operably linked to an appropriate control sequence for expression.
[0050] The recombinant expression vector can be any vector (e.g., a plasmid or a virus) that can be conveniently subjected to recombinant DNA procedures and can cause the expression of the polynucleotide sequence of the fusion protein. The choice of the vector will generally depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be a linear plasmid or a closed circular plasmid.
[0051] The expression vector can be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity and whose replication is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector can contain any means for ensuring self-replication. Alternatively, the vector can be a vector that, when introduced into a host cell, is integrated into the genome and replicated with one or more chromosomes into which it has been integrated. In addition, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell can be used, or a transposon can be used.
[0052] The expression vector of the present invention preferably contains one or more selectable markers that allow for easy selection of transformed cells. A selectable marker is a gene whose product provides biocide resistance or virus resistance, resistance to heavy metals, prototrophy for auxotrophs, etc. Examples of bacterial selectable markers are the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers that confer antibiotic resistance such as ampicillin, kanamycin, chloramphenicol (Example 1), or tetracycline resistance. Suitable markers for yeast host cells are ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for filamentous fungal host cells include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and their equivalents. Embodiments for Aspergillus cells include the amdS and pyrG genes of Aspergillus nidulans or Aspergillus oryzae, and the bar gene of Streptomyces hygroscopicus. Selectable markers for insect, plant, and mammalian cells are also well known.
[0053] The expression vector of the present invention preferably contains one or more elements that allow the vector to integrate into the genome of the host cell or to replicate autonomously in the cell independently of the genome. For integration into the genome of the host cell, the vector can rely on the nucleic acid sequence encoding the polypeptide or any other element of the vector used for integrating the vector into the genome by homologous or non-homologous recombination.
[0054] Alternatively, the expression vector can contain additional nucleic acid sequences for guiding integration into the genome of the host cell by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into one or more precise positions of one or more chromosomes in the genome of the host cell. The integration element can be any sequence homologous to a target sequence within the genome of the host cell. In addition, the integration element can be a non-coding or coding nucleic acid sequence. On the other hand, the vector can integrate into the genome of the host cell by non-homologous recombination.
[0055] For autonomous replication, the vector may also contain an origin of replication that enables the vector to replicate autonomously in the host cell under discussion. Examples of bacterial origins of replication are the P15A ori, or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which plasmids have the P15A ori) or pACYC184 that permit replication in Escherichia coli, and the origins of replication of plasmids pUB110, pE194, pTA1060 or pAM31 that permit replication in Bacillus. Examples of origins of replication for use in yeast host cells are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication may be an origin of replication having a mutation that renders its function in the host cell temperature-sensitive (see, for example, Ehrlich, 1978, Proc Natl Acad Sci. USA 75:1433).
[0056] More than one copy of the nucleic acid sequence of the fusion protein can be inserted into the host cell to increase the production of the gene product. An increased copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene having the nucleic acid sequence, where cells containing the amplified copy of the selectable marker gene and thus the additional copy of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selective reagent.
[0057] Expression vectors for fusion protein polynucleotides are commercially available. Suitable commercial expression vectors include the p3xFLAGTM expression vector from Sigma-Aldrich Chemicals, St. Louis, Mo., which includes a CMV promoter and an hGH polyadenylation site for expression in mammalian host cells, and a pBR322 origin of replication and an ampicillin resistance marker for amplification in Escherichia coli. Other suitable expression vectors are pBluescriptII SK(-) and pBK-CMV commercially available from Stratagene, La Jolla, Calif., and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (Lathe et al., 1987, Gene 57:193-201).
[0058] Suitable host cells for expressing polynucleotides encoding fusion proteins are well known in the art and include, but are not limited to: bacterial cells such as Escherichia coli, Lactobacillus kefir, Lactobacillus brevis, Lactobacillus minutus, Streptomyces, and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)); insect cells such as Drosophila S2 and Spodoptera frugiperda Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Suitable media and growth conditions for the above host cells are well known in the art.
[0059] The polynucleotides for expressing fusion proteins can be introduced into cells by various methods known in the art. These techniques include electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion, etc. The various methods for introducing polynucleotides into cells are known to those skilled in the art.
[0060] The polynucleotides encoding fusion proteins can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then ligated (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired continuous sequence. For example, polynucleotides can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described by Beaucage et al., 1981, Tet Lett 22:1859-69, or the method described by Matthes et al., 1984, EMBO J. 3:801-05 (e.g., as it is commonly applied in automated synthesis methods). According to the phosphoramidite method, oligonucleotides are synthesized, for example, in an automated DNA synthesizer, purified, annealed, ligated, and cloned into a suitable vector. Additionally, substantially any nucleic acid can be obtained from a variety of commercial sources such as Midland Certified Reagent Company in Midland, Texas, Great American Gene Company in Ramona, California, ExpressGen in Chicago, Illinois, and Operon Technologies in Alameda, California.
[0061] The engineered fusion protein expressed in a host cell can be recovered from the cells and / or culture medium using any one or more well-known protein purification techniques, including lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography, etc. A suitable solution for lysing bacteria (such as Escherichia coli) and efficiently extracting proteins can be commercially obtained from Sigma-Aldrich in St. Louis under the trade name CelLytic B.TM.
[0062] Chromatographic techniques for separating the fusion protein include reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, gel electrophoresis, and affinity chromatography, etc. The purification conditions will depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, and molecular shape, and will be obvious to those skilled in the art.
[0063] In some embodiments, affinity techniques can be used to separate the fusion protein. For affinity chromatography purification, any antibody that specifically binds to the fusion protein can be used. To generate antibodies, various host animals (including but not limited to rabbits, mice, rats, etc.) can be immunized by injecting a compound. The compound can be attached to a suitable carrier, such as BSA, through a side-chain functional group or a linker attached to the side-chain functional group. Various adjuvants can be used to enhance the immune response, depending on the host species, including but not limited to Freund's (complete and incomplete), mineral gels such as aluminum hydroxide, surface-active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (Bacillus Calmette-Guérin) and Corynebacterium parvum.
[0064] The TEV in the fusion protein may be wild-type or one of any several functional mutants. The TEV mutant used in the experiments described herein is the TEV mutant shown in patent application number CN201010204707A (incorporated herein by reference), which carries the following mutations: T17S / L56V / N68D / I77V / S135G. The DNA sequence in SEQ ID NO:1 encodes Figure 1 the protein in, where after the 6-mer His tag is the CBD tag, followed by the N-terminus of the TEV protease. In some embodiments, a linker containing Gly-Ser residues, including GSGSSG (SEQ ID NO:7), can be used after either or both tags.
[0065] SEQ ID NO:1
[0066]
[0067]
[0068] The single-underlined portion of SEQ ID NO:1 encodes a 6-mer His tag, followed by a linker. The double-underlined portion in the sequence encodes CBD, followed by another linker.
[0069] The DNA in SEQ ID NO:1 encodes a protein without the MKI leader sequence, which is the first three amino acids in SEQ ID NO:2:
[0070]
[0071] In SEQ ID NO:2, after the N-terminal MKI is a 6-mer His tag, followed by a 6-mer Gly-Ser linker (underlined in SEQ ID NO:2), followed by a CBD tag, followed by a 6-mer Gly-Ser linker (double-underlined in SEQ ID NO:2), followed by TEV protease.
[0072] Thus, the fusion protein is His-CBD-TEV, which has a 3-mer leader sequence as shown in SEQ ID NO:2 and schematically shown in Figure 1 and optionally has a linker. After expression, the 3-mer N-terminal MKI tag in front of His-CBD-TEV is cleaved off. As shown, His-CBD-TEV is expressed in C2566 (New England Biolabs, MA) E. coli competent cells, where the plasmid construct is used as a kanamycin-resistant pBAD vector. The enzyme is purified from the cell lysate using the 6-mer His tag.
[0073] Example I
[0074] The His-CBD-TEV activity was confirmed by digestion of eGFP (enhanced green fluorescent protein) using an MBP tag linked to the protein via an engineered TEV site (SEQ ID NO:4, Appendix I below). Figure 2 A protein gel is shown, which shows the results of this digestion in serial dilutions of His-CBD-TEV. In the presence of His-CBD-TEV, MBP-TEV site-eGFP was cleaved into an MBP tag (43 kDa) and eGFP (33 kDa), showing that His-CBD-TEV has specific protease activity.
[0075] The successful binding of His-CBD-TEV to nickel magnetic beads (Beaver, China) or chitin magnetic beads (New England Biolabs, Ipswich) is shown in Figure 3 . This protein gel of the His-CBD-TEV sample after magnetic bead incubation shows that His-CBD-TEV can be successfully removed using either tag added for cleavage in the fusion protein.
[0076] Appendix I
[0077] The fusion protein substrate of TEV in the Experiment of Example I has the following corresponding DNA (SEQ ID NO:3) and protein (SEQ ID NO:4) sequences: His-CBD-TEV protease.
[0078] SEQ ID NO:3
[0079] ATGAAAATCCATCACCATCATCATCATGGGTCAGGGAGTTCCGGTTTGACAACGAACCCGGGGGTGAGTGCTTGGCAGGTGAACACAGCGTACACCGCTGGTCAGCTGGTGACTTACAATGGTAAAACCTACAAATGCCTGCAACCACACACTTCACTGGCTGGATGGGAGCCAAGCAATGTCCCGGCACTTTGGCAGTTACAGGGGTCTGGTAGCTCGGGAGAATCCCTGTTTAAAGGGCCACGTGATTATAACCCGATCTCCTCTAGTATATGTCACCTCACAAACGAATCCGATGGACATACCACCAGTCTTTATGGGATTGGCTTTGGTCCGTTTATAATCACCAACAAACATCTGTTCCGCCGCAACAATGGGACCCTGGTTGTACAATCACTGCACGGTGTTTTCAAAGTAAAAGATACGACCACACTGCAGCAGCATTTAGTTGATGGGCGGGACATGATAATCATTCGCATGCCAAAGGATTTTCCACCATTTCCACAGAAACTGAAGTTCCGCGAACCACAACGTGAGGAACGCATTTGTCTTGTGACGACTAATTTCCAGACCAAATCAATGAGTTCTATGGTTAGTGATACCTCGTGTACCTTCCCGAGCGGTGATGGGATTTTCTGGAAACATTGGATTCAGACAAAAGATGGACAGTGCGGTTCCCCGTTGGTATCCACAAGAGACGGTTTTATAGTTGGTATTCATTCTGCATCCAATTTTACCAATACAAACAACTACTTCACATCAGTTCCGAAGAACTTTATGGAACTGTTAACTAATCAGGAGGCCCAGCAGTGGGTATCAGGGTGGCGATTGAACGCCGACAGTGTTCTGTGGGGCGGGCATAAAGTGTTTATGTCTAAACCGGAAGAACCTTTTCAGCCGGTTAAAGAAGCCACTCAGTTAATGAACTGA。
[0080] The TEV fusion protein with the recognition site and cleavage site underlined (SEQ ID NO:4 below) is:
[0081] MKIHHHHHHEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GENLYFQS MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGMDELYK。
[0082] After TEV digestion, it is separated into two polypeptides. The first one is SEQ ID NO:5:
[0083] MKIHHHHHHEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNL GENLYFQ
[0084] And the second one is SEQ ID NO:6:
[0085] SMVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGMDELYK。
[0086] The specific processes, methods, and compositions described herein are representative of preferred embodiments and are exemplary and not intended to limit the scope of the invention. Given this specification, those skilled in the art will envision other objectives, aspects, and embodiments, and these are included within the spirit of the invention as defined by the scope of the claims. It will be apparent to those skilled in the art that various substitutions and modifications can be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention illustratively described herein can be practiced appropriately without the presence of any one or more elements, or any one or more limitations, not expressly disclosed herein as necessary. Thus, for example, in each instance in the embodiments or examples of the invention herein, any one of the terms “comprising,” “including,” “containing,” etc. should be read broadly and without limitation. The methods and processes illustratively described herein can be appropriately implemented in a different order of steps, and they need not be limited to the order of steps indicated herein or in the claims. It should also be noted that, unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents, and the plural includes the singular form. In no event should this patent application be construed as limited to the specific examples or embodiments or methods specifically disclosed herein. In no event should this patent application be construed as being limited by any statement made by any examiner or any other official or employee of the Patent and Trademark Office, unless such statement is expressly and unconditionally or specifically adopted by the applicant in responsive written material. The invention has been described herein broadly and generically. Each of the narrower species and subgeneric classifications falling within the scope of this general disclosure also forms part of the invention.
[0087] The terms and expressions that have been employed are used in a descriptive sense and not in a limiting sense, and are not intended to exclude any equivalents or portions of the features shown and described, but it will be recognized that various modifications within the scope of the claimed invention are possible. Accordingly, it is to be understood that, although the invention has been specifically disclosed by way of preferred embodiments and optional features, those skilled in the art can adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
Claims
1. A method for cleaving a target protein using a fusion protein having a sequence of a wild-type or mutant TEV protease with two affinity tags and then removing the fusion protein after cleavage by binding to either or both of the affinity tags, the method comprising: generating the fusion protein, wherein a chitin-binding domain is before the residue representing the N-terminus of the protease, and a polyhistidine tag is before the chitin-binding domain; combining the fusion protein with a substrate for cleavage by the fusion protein in a reaction solution under reaction conditions; and binding the fusion protein to a solid support having chitin on the support surface thereof or wherein the solid support carries a metal.
2. The method according to claim 1, wherein the solid support is a resin or magnetic beads.
3. The method according to claim 2, wherein the magnetic beads contain nickel.
4. The method according to claim 2, wherein the solid support is magnetic beads, and the method further comprises removing the magnetic beads carrying the bound fusion protein from the reaction solution using magnetic attraction.
5. The method according to claim 1, wherein the fusion protein has a first linker between the polyhistidine tag and the chitin-binding domain, the first linker being a series of glycine and serine residues.
6. The method according to claim 5, wherein the fusion protein further comprises a second linker between the chitin-binding domain and the residue at the N-terminus of the protease, the second linker being a series of glycine and serine residues.
7. The method according to claim 6, wherein both the first linker and the second linker are 6-mer.
8. The method according to claim 6, wherein both the first linker and the second linker are: GSGSSG (SEQ ID NO:7).
9. The method according to claim 1, wherein the fusion protein contains an MKI leader sequence.
10. The method according to claim 1, wherein the fusion protein has the sequence of SEQ ID NO:2.
Citation Information
Patent Citations
TEV protease mutant and coding gene and application thereof
CN101864407A
Phosphonyldipeptides useful in the treatment of cardiovascular diseases
WO1995033836A1
Non-toxic, non-toxigenic, non-pathogenic fusarium expression system and promoters and terminators for use therein
WO1996000787A1