Engineered protease for the removal of n-terminal purification tags

Novel proteases with high solubility and minimal C-terminus residue restriction address the limitations of TEV protease by enabling efficient, residue-preserving tag removal, ensuring the structural and biological integrity of target proteins.

WO2026027896A1PCT designated stage Publication Date: 2026-02-05THE UNIV COURT OF THE UNIV OF EDINBURGH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/051714
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-07-31
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing proteases like TEV protease exhibit low solubility and add non-native amino acids to the N-terminus of proteins during purification tag removal, which can affect the structural and biological characteristics of biologically active therapeutics.

Method used

Development of novel proteases with high solubility and minimal restriction on the identity of the C-terminus residue at the site of cleavage, allowing 'scarless' cleavage to maintain the natural N-terminus of target proteins.

Benefits of technology

The novel proteases effectively remove purification tags without leaving non-native amino acids, preserving the natural N-terminus and maintaining the structural and biological integrity of the target proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051714_05022026_PF_FP_ABST
    Figure GB2025051714_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure provides novel proteases having desirable cleavage specificity and high solubility. The disclosure also provides uses of these proteases for catalysing the proteolysis or cleavage of a substrate, methods of catalysing the proteolysis or cleavage of a substrate, said methods comprising contacting a substrate with any of the disclosed proteases and kits comprising the proteases.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ENZYMES FIELD The present disclosure provides novel enzymes that exhibit advantageous characteristics, such as improved solubility and minimal restriction on the identity of the C-terminus residue at the site of cleavage, and uses thereof. Also provided are methods of using the disclosed enzymes and kits comprising the enzymes. BACKGROUNDThe use of protein purification tags – such as polyhistidine (nHis), glutathione S transferase(GST) or maltose binding protein (MBP) has greatly simplified the purification of recombinantlyexpressed proteins [1-3]. The fusion of a purification tag to a protein of interest (POI) meansthat a standard purification method can be used for a variety of different POI [4]. The fusion tag is often placed at the N-terminus of the POI so it can be enzymatically removed afterpurification [5, 6]. Several enzymes have been developed for such purification tag removal,one of the most widely used is Tobacco Etch Virus nuclear-inclusion-A endopeptidase (TEV protease), which is favoured because of its high specificity [7]. TEV protease has a recognition sequence ENLYFQ-G / S, where (-) represents the site of cleavage [8]. The naming of the residues follows the convention P6-P5-P4-P3-P2-P1-P1’. For TEV, the preferred residue at the P1’ position is Gly or Ser, which means that after cleavage a Gly or Ser residue is added to the N-terminus of the POI [9]. For many laboratory applications, the addition of an extra residue to the N-terminus of the POI is not problematic, however, if a recombinant POI is to be a biologically active therapeutic, the addition of extra amino acids to the N-terminus is not desirable. Although TEV protease is widely used, it is well known to exhibit low solubility thatdiminishes the purification yield, high concentration storage and use [10, 11].With these attributes as motivation, it is amongst one of the objectives of the present disclosure to develop novel proteases with high substrate specificity, with little restriction on the identity of the P1’ residue, and with good solubility. SUMMARY The present disclosure is based in part on the development of novel proteases havingdesirable cleavage specificity and high solubility. An initial starting point of the present studywas based on an alignment of candidate viral proteases; these are known in the art to exhibitstringent cleavage site sequence specificity. In contrast, the proteases detailed hereinunexpectedly exhibit minimal restriction on the identity of the C-terminus residue at the site ofcleavage. Moreover, the various design elements that contribute to the structural andsequence elements of the disclosed proteases are not merely additive, but rather aresynergistic with regard to the tolerability of the identity of the P1’ residue. As used herein, the terms “comprising”, “comprise” and / or “comprises” are used to denote aspects and embodiments of this disclosure that “comprise” a particular feature or features. Itshould be understood that these terms may also encompass aspects and / or embodimentswhich “consists of” or “consists essentially of” the relevant feature(s). Moreover, in thefollowing text any reference to an aspect of the present disclosure may refer to anyembodiment or teaching of that aspect. Furthermore, and for the avoidance of doubt, allembodiments, disclosures and / or teachings associated with the first aspect of this disclosure,apply mutatis mutandis to the second aspect (and its embodiments, disclosures and / orteachings) and any further relevant aspects (embodiments, disclosures and / or teachings) described herein. The term “protease” typically refers to enzymes that catalyse the cleavage / hydrolysis of one or more peptide bonds in a protein or polypeptide. In the context of the present disclosure, the terms “enzyme” and “protease” may be used interchangeably.A protease may cleave a protein at (or within) a specific recognition site. As such, the term“recognition site” as used herein typically refers to a sequence of amino acids recognised bya protease and typically includes the protease cleavage site. By way of example, a recognitionsite may comprise a plurality of specific amino acids and a protease may cleave at or after a specific amino acid within said recognition sequence. The naming of the residues of a protease recognition sequence follows the convention P6-P5-P4-P3-P2-P1-P1’, where the protease cleaves between residues P1 and P1’. Many proteases are highly specific in that they only recognise sequences in which certain positions within the recognition site are occupied by specific amino acids. By way of example, Tobacco Etch Virus nuclear-inclusion-A endopeptidase (TEV protease), TEV protease has a recognition sequence ENLYFQ-G / S (where (-) represents the site of cleavage). As such, the preferred residue at the P1’ position is Gly or Ser. One of skill will know that affinity tags are important tools for target protein expression and purification. However, because they can interfere with the structure and / or function of the target protein, it is usually necessary to remove them. Proteases may be used to remove affinity purification tags from target proteins. However, the specificity of a protease often results in target proteins being left with a non-native amino acid at the N-terminus. For example, where tagged target protein is to be de-tagged using a TEV protease, the targetprotein may be left with a non-native Ser or Gly residue on its N-terminus. This may affect thestructural and / or biological characteristics of the target protein.The present disclosure provides proteases which are not restricted by the identity of a residue(e.g. the P1’ residue) in a recognition sequence. Without wishing to be bound by theory, thisfeature is restricted to the P1’ position and permits "scarless" cleavage to maintain the natural N-terminal of the target protein. A protease with this property may find particular application in a method of removing a tag (e.g. an affinity tag) from a tagged-target protein. A particular advantage of using a protease of this disclosure to remove a tag from a tagged target protein is that the released target protein may not be left with a non-native amino acid at or on its N- terminus. A protease of this disclosure may recognise and cleave a protein comprising a recognition sitecomprising the following amino acid structure:P6–P5–P4–P3–P2–P1–P1’ wherein each P independently represents an amino acid residue and the protease cleaves between the residues P1 and P1’.In one teaching, with respect to the amino acid structure detailed above,P6 is typically glutamate (E) or a conservative amino acid substitution thereof; P5 is typically alanine (A) or a conservative amino acid substitution thereof; P4 is typically valine (V) or a conservative amino acid substitution thereof; P3 is typically tyrosine (Y) or a conservative amino acid substitution thereof; P2 is typically histidine (H) or a conservative amino acid substitution thereof; P1 is typically glutamine (Q) or a conservative amino acid substitution thereof; P1’ may be any amino acid (X). In one teaching, the P1’ residue of the recognition sequence may be alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine,leucine, lysine, methionine, phenylalanine, serine, threonine, tryptophan, tyrosine or valine.In one teaching, the P1’ residue of the recognition sequence of the protease disclosed herein may be alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, threonine, tryptophan, tyrosine or valine. In one teaching, the P1’ residue of the recognition sequence may be alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, serine, threonine, tryptophan, tyrosine or valine. In one teaching, the P1’ residue of the recognition sequence may be alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, threonine, tryptophan, tyrosine or valine. In one teaching, the P1’ residue of the recognition sequence may be alanine, phenylalanine, isoleucine, methionine, tryptophan, glutamine, serine, threonine, arginine, aspartic acid, glutamic acid or glycine.In one teaching, the P1’ residue of the recognition sequence may not be any one of thefollowing: proline, asparagine, cysteine, histidine, leucine, lysine, tyrosine or valine.In one teaching, the P1’ residue of the recognition sequence is not proline. In one teaching, the P1’ residue of the recognition sequence is not glycine and / or serine. In one teaching, the P1’ residue of the recognition sequence is not proline, glycine and / or serine. As stated, a protease of this disclosure may target a substrate having a specific recognition sequence characterised by a P1’ of that sequence being ‘any amino acid’. In this case, the term ‘substrate’ may embrace any protein, peptide, glycoprotein or molecule comprising the amino acid structure: P6–P5–P4–P3–P2–P1–P1’. The term “substrate” of the proteases disclosed herein may encompass any peptide, protein,glycoprotein or molecule comprising a recognition sequence detailed herein. By way ofexample, the substrate may be a recombinant protein or a tagged protein. Where the substrate is a tagged (for example affinity-tagged) protein, the protease of this disclosure may be used to remove the tag from the protein.. The term “conservative amino acid substitution” in this context may refer to substitution of with an amino acid with similar physico-chemical and / or kinetic properties. Within the context ofthis disclosure the conservative substitution of any one or more of the amino acid residue(s)of the recognition site will not significantly affect or compromise the ability of the protease ofthe present disclosure to cleave its substrate at the recognition site. In other words, despite one or more conservative substitutions, the sequence will still serve as a recognition site for a protease of this disclosure. In one teaching, the recognition site of a protease of the present disclosure comprises the amino acid sequence EAVYHQX, wherein X is any amino acid. The proteases of this disclosure may exhibit optimum enzymatic (protease) activity at or withina specific temperature and / or pH. For example, at a specific temperature and / or pH or withina specific temperature and / or pH range, a protease of this disclosure may exhibit a maximal, near maximum or at least 50%, 60%, 70%, 80% or 90% of the maximal rate of reaction. The reaction in this context refers to the hydrolysis of a peptide bond between the P1 and P1’ amino acid residues of the recognition sequence.Any of the proteases of this disclosure may exhibit optimum enzymatic (protease) activity at atemperature between 0 and 70°C. In some examples, the optimum temperature may be selected from between 0 and 60°C, 0 and 50°C or 0 and 40°C. In some embodiments, the optimum temperature may be selected from 4, 5, 6, 7, 8, 9, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 37, 40, 45, 50, 55, 60, 65 or 70°C. In some embodiments, the optimum temperature range is between 20 and 40°C. In one embodiment, a preferred optimum temperature is 24°C.Additionally or alternatively, any of the proteases of this disclosure may exhibit optimalenzymatic (protease) activity at a pH between 4 and 10. In some embodiments, the optimum pH may be selected between pH 4 and 10, pH 5 and 10 or pH 6 and 10. In some embodiments, the optimum pH may be selected from pH 4, 5, 6, 7, 8, 9 or 10. In some embodiments, the optimum pH is between pH 6 and 9. In one embodiment, the preferred optimum pH is between 7 and 8. In one embodiment, the preferred optimum pH is 7.7.The proteases of this disclosure exhibit high solubility and low protein aggregation. One of skillwill appreciate various methods known in the art to determine the solubility of a protein, such as by using protein solubility prediction methods or software (e.g. CamSol method), for example. Protein aggregation may be determined by using methods detailed herein or othermethods known in the art, such as size exclusion chromatography, intrinsic tryptophanfluorescence detection, extrinsic dye-binding fluorescence assays, determining the aggregation index ratio, for example. In one teaching a protease of this disclosure may comprise, consist essentially of or consistof the following sequence (SEQ ID NO: 1):SX SLFRGLRD YNPIX X X ICH LTNX SDGX SX SLYGX GFGPLIX TNX HLFX R NNGELX X X SR HGEFVVKNTT QLX LLPX X GRDX X X IRLPKD FPPFPQKLX F RQPX KX ERIC X VGSNFQTKS YX Q] wherein: X1is K or N; X26is T or S; X51is Q or I; X2 is S or A; X27 is T, M or V; X52 is Q, H or D; X3 is S or N; X28 is T or V; X53 is E or S; X4is N or A; X29is T or I; X54is P, T or A; X5 is E or V; X30 is M or T; X55 is A or E; X6is H or A; X31is E or D; X56is I or V; X7is N or E; X32is Q or F; X57is S or G; X8 is I or V; X33 is H or Q; X58 is I or L; X9is I or L; X34is S or L; X59is E or A; X10is Q or R; X35is L or M; X60is L, T, Q or K; X11 is R or E; X36 is K or R; X61 is E or R; X12is T or R; X37is K or F; X62is E or S; X13 is I or V; X38 is L or V; X63 is P or L; X14 is Q or K; X39 is I or L; X64 is I or V; X15is K, H or Q; X40is L or H; X65is V, I or S; X16 is I or C; X41 is S, T or A; X66 is V or I; X17 is D, E or P; X42 is F or S; X67 is T or S; X18is I or L; X43is T or S; X68is F or D; X19 is L or I; X44 is N, Q or I; X69 is S or G; X20 is I, V or L; X45 is E or D; X70 is D, T or E; X21is K or G; X46is D or N; X71is A or L; X22 is E, I or T; X47 is E, H, A or M; X72 is A or F. X23 is G or E; X48 is T or K; X24is L or M; X49is H or A; X25 is I or V; X50 is T, D or N; It should be noted that the amino acid residues of SEQ ID NO: 1 that are enclosed withinsquare brackets (namely residues 222-243) may each (independently) be present or absentin a protease sequence of this disclosure or absent.As detailed herein, a protease comprising, consisting essentially of or consisting of the aminoacid sequence of SEQ ID NO: 1 typically recognises a substrate comprising the sequence EAVYHQX (recognition site) and cleaves between amino acid residues Q and X, wherein X may be any amino acid.A protease of this disclosure may comprise, consist essentially of or consist of SEQ ID NO: 1or 2, or a variant thereof, wherein said protease recognises a substrate comprising the sequence EAVYHQX and cleaves between amino acid residues Q and X, wherein X may be any amino acid. SEQ ID NO: 2SKSLFRGLRD YNPIASNICH LTNESDGHSN SLYGIGFGPL IITNQHLFRRNNGELTIQSR HGEFVVKNTT QLKLLPIDGR DILIIRLPKD FPPFPQKLKFRQPEKGERIC LVGSNFQTKS ITSTVSETST TMPVENSQFW KHWISTKDGHCGLPLVSTKD GKILGIHSLA NFTNTINYFA AFPEDFEETY LHTQEAQEWVKHWKYNPDAI SWGSLNLQES QXX X X XX X X X XX X X X X XX X X X X X Wherein X is P or absentX is E, R or absentX is E, S or absentX is P, L or absentX is F or absent X is K or absentX is I, V or absentX is V, I, S or absentX is K or absentX is L or absentabsentor absentabsent absent absentor absent E or absentX is A, L or absentX is V or absent X is Y or absentX is A, F or absentX is Q or absent A protease of this disclosure may comprise, consist essentially of or consist of SEQ ID NO: 3 or a variant thereof, wherein said protease recognises a substrate comprising the sequence EAVYHQX and cleaves between amino acid residues Q and X, wherein X may be any amino acid. SEQ ID NO: 3SKSLFRGLRD YNPIASNICH LTNESDGHSN SLYGIGFGPL IITNQHLFRRNNGELTIQSR HGEFVVKNTT QLKLLPIDGR DILIIRLPKD FPPFPQKLKFRQPEKGERIC LVGSNFQTKS ITSTVSETST TMPVENSQFW KHWISTKDGHCGLPLVSTKD GKILGIHSLA NFTNTINYFA AFPEDFEETY LHTQEAQEWVKHWKYNPDAI SWGSLNLQES QPEEPFKA protease of this disclosure may comprise, consist essentially of or consist of SEQ ID NO: 4 or a variant thereof, wherein said protease recognises a substrate comprising the sequence EAVYHQX and cleaves between amino acid residues Q and X, wherein X may be any amino acid. SEQ ID NO: 4SKSLFRGLRD YNPIASNICH LTNESDGHSN SLYGIGFGPL IITNQHLFRRNNGELTIQSR HGEFVVKNTT QLKLLPIDGR DILIIRLPKD FPPFPQKLKFRQPEKGERIC LVGSNFQTKS ITSTVSETST TMPVENSQFW KHWISTKDGHCGLPLVSTKD GKILGIHSLA NFTNTINYFA AFPEDFEETY LHTQEAQEWVKHWKYNPDAI SWGSLNLQES QA protease of this disclosure may comprise, consist essentially of or consist of SEQ ID NO: 5 or a variant thereof, wherein said protease recognises a substrate comprising the sequence EAVYHQX and cleaves between amino acid residues Q and X, wherein X may be any amino acid. SEQ ID NO: 5SKSLFRGLRD YNPIASNICH LTNESDGHSN SLYGIGFGPL IITNQHLFRRNNGELTIQSR HGEFVVKNTT QLKLLPIDGR DILIIRLPKD FPPFPQKLKFRQPEKGERIC LVGSNFQTKS ITSTVSETST TMPVENSQFW KHWISTKDGHCGLPLVSTKD GKILGIHSLA NFTNTINYFA AFPEDFEETY LHTQEAQEWVKHWKYNPDAI SWGSLNLQES QPEEPFKIVK LVTDLFSDAV YAQIn one example, where a protease of the present disclosure comprises the amino acid sequence of SEQ ID NO: 5, one or more of the following residues may be modified: Residue 2; Residue 91; Residue 197; Residue 4; Residue 99; Residue 207; Residue 16; Residue 111; Residue 208; Residue 24; Residue 124; Residue 217; Residue 28; Residue 131; Residue 219; Residue 35; Residue 150; Residue 223; Residue 42; Residue 155; Residue 224; Residue 45; Residue 166; Residue 225; Residue 49; Residue 173; Residue 228; Residue 56; Residue 175; Residue 229; Residue 58; Residue 184; Residue 232; Residue 66; Residue 187; Residue 233; Residue 73; Residue 189; Residue 236; Residue 78; Residue 194; Residue 238. Residue 79; Residue 82; Residue 84; Any of the abovementioned residues may be modified by: (i) an amino acid substitution (where a wild type or native amino acid is swappedor changed for another (different) amino acid – the term “substitutions” wouldinclude conservative amino acid substitutions); and / or (ii) an amino acid deletion (where a wild type or native amino acid residue isremoved); and / or (iii) amino acid addition(s) / insertion(s) (where additional amino acid residue(s) areadded to the wild type (or reference) primary sequence (e.g. any of SEQ IDNOS: 1-5); and / or (iv) amino acid / sequence inversions (usually where two or more consecutive aminoacids in a primary sequence are reversed; and / or (v) amino acid / sequence duplications (where an amino acid or a part of the primaryamino acid sequence (for example a stretch of 5-10 amino acids) is repeated). To expand on the term ‘conservative substitution’, the reader should understand that thisembraces modifications in which any amino acid residue is replaced with a different aminoacid having similar biochemical properties (e.g. charge, hydrophobicity and size). By way ofexample, any of glycine, alanine, valine, leucine and / or isoleucine may be used in place of one another. Also, any one of serine, cysteine, threonine or methionine may be used in place of one another. Also, phenylalanine, tyrosine and tryptophan may be substituted with one another. Histidine, arginine and lysine may also be substituted with one another as could aspartate, glutamate, asparagine and glutamate residues. One of skill will appreciate that theprecise choice of amino acid residue may depend on the exact properties that are to beconserved in the final protease. In view of the above, a protease of the present disclosure comprises, consists essentially ofor consists of the amino acid sequence of any of SEQ ID NOS: 2-5, or a variant thereof,wherein said protease recognises a recognition sequence EAVYHQX and cleaves between amino acid residues Q and X, wherein X may be any amino acid. The term ‘variant’ (e.g.variant of any of the sequences disclosed herein) may comprise, consist essentially of orconsist of any one or more of the abovementioned modifications relative to a referencesequence. Within the context of this disclosure, the term ‘reference sequence’ may embraceany of the disclosed protease sequences – for example the protease sequence of SEQ IDNO: 1-5. A variant protease sequence of the present disclosure may comprise, consist essentially of orconsist of one or more additional amino acid modifications (as noted above) but may stillfunction as a protease. A variant protease sequence may be derived from a reference sequence of this disclosure by additionally or alternatively substituting one or more residues with an un-natural amino acid, including for example non-proteinogenic amino acids. For example, the un-natural amino acid ornithine may be used in place of a lysine residue. One of skill will be aware of the various un- natural amino acids and how they may be substituted for at least some of the (natural) amino acids residues present in the (reference) protease sequences described herein. A disclosure of available un-natural amino acids for use as substitutions can be found in Narancic et al (World Journal of microbiology and biotechnology, 35, Article number: 67 (2019)).In one teaching, a protease of the present disclosure may comprise, consist essentially of orconsist of L- and / or D- amino acids. Additionally, or alternatively, a protease may comprise β-amino acids. In some embodiments, a protease of the disclosure may comprise N-methylatedamino acid or N-acetyl amino acids. Further additionally or alternatively, a protease of thisdisclosure may be lipidated and / or PEG-ylated. A variant protease sequence of this disclosure may comprise, consist essentially of or consistof a functional fragment or portion of any of the protease sequences described herein. By wayof example, the disclosure provides functional fragments of a protease comprising the aminoacid sequence of and of SEQ ID NOS: 1-5. A functional fragment of any of SEQ ID NOS: 1-5may comprise, consist essentially of or consist of anywhere between 10 and n-1 of any of theprotease sequences described herein, wherein ‘n’=the total number of amino acids in the fullprotease sequence. By way of example, a functional fragment of SEQ ID NO: 5 may comprise,consist essentially of or consist of 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85,90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180,185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240 or 241 of the amino acids of SEQID NO: 5. The term ‘functional’ (as used herein (e.g. ‘functional fragment’)) means that thefragment or portion retains protease activity – that is, it should catalyse proteolysis of a targetprotein comprising a recognition sequence of said protease. In one example, the amino acid sequence of the protease of the present disclosure maycomprise, consist essentially of or consist of SEQ ID NO: 6:SEQ ID NO: 6 (Con1)SKSLFRGLRD YNPIASNICH LTNESDGHSN SLYGIGFGPL IITNQHLFRRNNGELTIQSR HGEFVVKNTT QLKLLPIDGR DILIIRLPKD FPPFPQKLKFRQPEKGERIC LVGSNFQTKS ITSTVSETST TMPVENSQFW KHWISTKDGHCGLPLVSTKD GKILGIHSLA NFTNTINYFA AFPEDFEETY LHTQEAQEWVKHWKYNPDAI SWGSLNLQES QPEEPFKIVK LVTDIn one embodiment, a variant protease sequence, may comprise, consist essentially of orconsist of a sequence which is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%,91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical or homologous to all or part of any (reference) sequence of this disclosure. By way of example, a variant protease of thisdisclosure may comprise, consist essentially of or consist of an amino acid sequence which isat least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%or 99% identical to any one of SEQ ID NOS: 1-6 or a (functional) fragment thereof.As it will be appreciated by the skilled person in the art, the amount of sequence identitybetween any two sequences (for example two protease sequences) may be calculated byaligning the sequences and determining the number of identical residues as a proportion or percentage of the total number of residues. One of skill will appreciate that it is straightforward to test any variant protease or protease fragment (test protease) derived from any of the protease sequences disclosed herein for proteolytic activity by contacting the test protease with a substrate comprising the recognition / cleavage sequence, and observing or determining whether the test protease cleaves the substrate. In another aspect the present disclosure, provides the use of any of the proteases disclosed herein for catalysing proteolysis or cleavage of a substrate. In a further aspect, the disclosure provides a method of catalysing proteolysis or cleavage of a substrate, said method comprising contacting the substrate with the protease under conditions which permit proteolysis or cleavage of the substrate.The substrate may comprise, consist essentially of or consist of a protein or a peptide.The protease may catalyse the protein or peptide component of the substrate.In one teaching, the substrate may comprise a tagged protein comprising, for example, apeptide tag fused to a target peptide or protein. The protease may be used to remove the tagfrom the protein. The tagged protein may be an affinity tagged protein – in which case theprotease is used to remove the affinity tag part of the protein. The tag may be present at theN-terminus of the substrate; for example, at the N-terminus of the recognition sequenceEAVYHQX. It should be noted that any affinity tags known in the art to be suitable for protein purification may be removed using a protease of this disclosure; such tags include (but arenot limited to) His-, FLAG-, HA-, V5-, Myc-, Step-, GST-, MBP-, SUMO-, CBP- or FATT- tag.The “conditions which permit proteolysis or cleavage” may encompasses one or more factors,such as temperature, pH, time and / or molar ratio of protease:substrate. In one teaching, the ‘conditions’ may comprise contacting the protease and its substrate a temperature between 0 and 70°C. In some examples, the temperature may be selected from between 0 and 60°C, 0 and 50°C or 0 and 40°C. In some examples, the temperature may be selected from 4, 5, 6, 7, 8, 9, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 37, 40, 45, 50, 55, 60, 65 or 70°C. In some examples, the temperature range is between 20 and 40°C. In one example, the temperature is 24°C.Additionally, or alternatively, the “conditions which permit proteolysis or cleavage” maycomprise contacting the protease and its substrate at a pH of between 4 and 10. In some examples, the pH may be selected between pH 4 and 10, pH 5and 10 or pH 6 and 10. In some examples, the pH may be selected from pH 4, 5, 6, 7, 8, 9 or 10. In some examples, the pH is between pH 6 and 9. In one example, the pH is between 7 and 8. In one example, the pH is 7.7. Further additionally or alternatively, the “conditions which permit proteolysis or cleavage” maycomprise contacting the protease and its substrate for a minimum reaction time of, forexample, 5 min, 10 min, 15 min, 20 min, 25 min, 30 min, 35 min, 40 min, 45 min, 50 min, 1h,2h, 3h, 4h, 5h, 6h, 7h, 8h, 9h, 10h, 11h or 12h. Further additionally or alternatively, the “conditions which permit proteolysis or cleavage” may comprise contacting the protease and its substrate at a molar ratio of 1:100 (protease: substrate). In some examples, said method may comprise contacting a protease and a targetprotein or protein of interest at a molar ratio of 1:10, 1:50, 1:100, 1:200, 1:300, 1:400, 1:500,1:600, 1:1000, 1:200, 1:400, 1:500, 1:1000, 1:2000 or 1:3000. In another example, said method may comprise contacting a protease and a target protein or protein of interest at a molar ratio of 1:10, 1:50, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600 or 1:1000. In one teaching, the “conditions which permit proteolysis or cleavage” may comprise or consist of: (i) temperature between 20 and 40°C;(ii) pH between 6 and 9; (iii) reaction time between 0.5 h and 17 h; and / or(iv) molar ratio of 1:10, 1:50, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600 or 1:1000(protease:substrate).In some examples, the substrate (for example a protein to be subject to proteolysis or cleavageusing a protease of this disclosure) may be immobilised (for example bound, adhered orconjugated) to a surface. In a specific example, the substrate may be immobilised to apurification column, such as an affinity column.In one aspect, the present disclosure provides a kit comprising, consisting essentially of orconsisting of a protease of the present disclosure.In one teaching, a kit of this disclosure may include a protease comprising, consistingessentially of or consisting of an amino acid sequence of any of SEQ ID NO: 1-6. In a preferredembodiment, the kit comprises a protease comprising, consisting essentially of or consistingof an amino acid sequence of SEQ ID NO: 3, 4, 5 or 6.A kit of this disclosure may further comprise a protease substrate (e.g. a protein or peptide tobe subject to proteolysis and / or cleavage using a protease of the kit) and / or one or more buffers for use with the relevant protease. By way of example, a buffer for inclusion in a kit of this disclosure may be aimed at providing any of the optimal pH and / or solubility of the protease detailed herein. By way of example only, a suitable buffer may comprise, consistessentially of or consist of PBS, Tris-HCl, EDTA, NaCl, glycerol, or a combination thereof; andoptionally includes DTT. In one example, the buffer may comprise, consist essentially of orconsist of Tris-HCl, NaCl and glycerol. In some examples, the buffer may further comprise aprotease inhibitor. In some examples, the buffer may further comprise DTT. In a specificexample, the suitable buffer comprises or consists of PBS (pH 7.0), and optionally, DTT provided as a separate solution.The protease substrate may comprise, consist essentially of, or consist of a protein comprisingthe sequence EAVYHQX, wherein X may be any amino acid. Typically, the substrate will comprise the sequence EAVYHQX followed by an amino acid sequence of a protein or peptide of interest. In a preferred embodiment, the substrate comprises the sequence EAVYHQX at the N-terminal of the protein followed by an amino acid sequence of a protein or peptide to be purified. By way of example, the protein to be purified may be a recombinant protein, a therapeutic agent or an antibody (or a fragment thereof) of clinical relevance. A kit of the present disclosure may further comprise one or more probes for the detection of proteolysis and / or the detection of product(s) resulting therefrom. By way of example only, afluorescent probe may be conjugated or attached to the N- and / or C-terminal of the substratesuch that a detectable signal is only generated once the protease cleaves the substrate of interest. In a further aspect, the disclosure a nucleic acid sequence encoding any of the proteases described herein. In one teaching, the disclosure provides a nucleic acid sequence encoding proteases comprising, consisting essentially of or consisting of SEQ ID NOS:1-6. The disclosure also provides an expression construct comprising a nucleotide sequence of this disclosure, for example a nucleic acid sequence encoding any of the disclosed proteases. By way of example, an expression construct may comprise, consist essentially of or consist of a nucleotide sequence which encodes a protease with an amino acid sequence of SEQ IDNOS: 1-6. In one example, said expression construct may comprise, consist essentially of orconsist of an expression vector. In some examples, the expression vector is a pMALexpression vector harbouring a Tac promoter, wherein the pMAL expression vector preferablymay not have a MBP tag.DETAILED DESCRIPTION The present disclosure will now be further described by way of example and with reference to the Figures, which show:Figure 1. Design of Consensus protease. (A) Alignment of the full-length amino acidsequences of the indicated proteases, using Clustal Omega, with a 75% identity threshold coloured by the default Clustal scheme (light blue for hydrophobic, red for +ve charge, magenta for –ve charge, green for polar, pink for Cysteine, orange for Glycine, yellow for Proline, cyan for aromatic and colorless for non-conserved residues). The consensussequence was automatically defined by the JalView viewer. The asterisk annotations indicatethe TEV residues which interact with the substrate residues (P1, P3, P6and P1’). The hashtag annotations indicate the catalytic triad (H46, D81 and C151). (B) Intrinsic solubility prediction of the full-length (1-243) and C-terminal truncated (1-234; Δ235-243) versions of the designedlibrary along with TEV and TuMV_J proteases as references. The bluer, the more soluble. (C)Ribbon representation of TEV protease structure (pink) complexed with uncleaved substrate(orange; PDB ID: 1LVB). (D) AlphaFold2 model of the full-length Con1 (SEQ ID No: 5)protease predicted by ColabFold. (E) Close-up view of the active site of a superposed Con1 (SEQ ID No: 5) (light blue) onto TEV protease (pink) showing the longer C-terminal of Con1(SEQ ID No: 5) around the substrate (orange).Figure 2. Substrate specificity of Con1 protease. (A and B) Substrate sequencepreferences for cleavage by TEV protease; n=19 (A) or TuMV_J protease; n=7 (B). The amino acid preferences at each position (P6-P5-P4-P3-P2-P1-P1’) are indicated as logo plots (WebLogo online application). (C) Schematic illustration describing the FRET-based assay. The substrates contain cyan fluorescent protein (CFP; donor) and yellow fluorescent protein (YFP; acceptor) separated by the potential cleavage site. Before cleavage, by exciting substrate at 430 nm, the emission of YFP is high at 520 nm while CFP emission is 480 nm is low. If the substrate is cleaved by the protease, the YFP emission decreases and CFP emissionincreases. (D and E) Screening Con1 (SEQ ID No: 6), TuMV_J or TEV proteases againstFRET substrates with different cleavage sites, ENLYFQ-G (D) or EAVYHQ-S (E). All the results are the mean of 2 technical repeats ± standard deviation. Ordinary one-way ANOVAwas used: ****P < 0.0001; ***P < 0.001; **P < 0.01; *P < 0.05; ns, P > 0.05 with Tukey’s testcorrection.Figure 3. Kinetic studies of Con1 and TuMV_J. (A and B) Representative plots of theproduct formation versus time for the indicated concentrations of EAVYHQ-S FRET substrateincubated with 0.4 μM Con1 (SEQ ID No: 6) (A) or TuMV_J (B). The results show the meanof 2 technical replicates (solid lines) ± standard deviation (dotted lines). The curves were fitted by the exponential plateau model on GraphPad Prism (black lines) to determine the reactionrates. See materials and methods for data analysis and equations. (C and D) Rate versussubstrate concentration plots for Con1 (SEQ ID No: 6) (C) and TuMV_J (D) against thedifferent FRET substrates (P1’=Ser, Ala, Asp or Met). The results are the mean of 2 technical repeats ± standard deviation. The curves are fitted by the Michaelis-Menten non-linearregression using GraphPad Prism software. The curves' colors are indicated in (D).Figure 4. Characterization of the pH and temperature dependency of Con1 activity. (A)Con1 (SEQ ID No: 6) and FRET substrate (EAVYHQ-S) were mixed to initiate the reactionafter 10 min equilibration at the desired pH. The amount of FRET substrate cleaved after 2 hours of incubation at room temperature at a given pH was determined. Results are expressed as a percentage of the cleavage that occurred at the optimal pH (pH 7.7). (B) Con1 (SEQ ID No: 6) and substrate were equilibrated in buffer A at the specified temperature for 10 min and the reaction was initiated by mixing them. The amount of FRET substrate cleaved after 2 hours of incubation was determined. Results are expressed as a percentage of the cleavage that occurred at the optimal temperature (24°C). All the results are the mean of 2 technical repeats ± standard deviation.Figure 5. (A) Schematic illustrations of the different substrates. His tag is separated from thePOI by the P6-P1(EAVYHQ) residues of the Con1 (SEQ ID No: 6) cleavage site. Each POI (DARPin, uricase, GFP) has a different N-terminal residue (as indicated in red) which forms the P1’ site. (B) SDS PAGE analysis of the proteins before (–) and after (+) incubation withCon1 (SEQ ID No: 6). Complete removal of the purification tag is evident by the decrease inMW of the protein. The reaction mixture contained 0.5 µM Con1 (SEQ ID No: 6) and 5 µM substrate was incubated overnight at room temperature.Figure 6. The conserved sequence of the cleavage site of TuMV_Q or PPV proteases; n=5(A), TVMV protease; n=9 (B) or PVA protease; n=5 (C). The cleavage sites were identified from the MEROPS database, and the logo plots were generated by the WebLogo online application.Figure 7. (A) Surface representation of inactive TEV protease (pink; PDB ID: 1LVB)complexed with uncleaved substrate (TENLYFQSGT; orange). (B) The binding residues ofP6, P3, P1 and P1’ were identified from the inactive TEV protease (pink; PDB ID: 1LVB) complexed with un-cleaved substrate (TENLYFQSGT; orange). The substrate-enzyme interaction was defined by ICM-BrowserPro software.Figure 8. SDS-PAGEs showing the purification steps of N-terminal His tagged proteases. (A- C) Gel photos showing the purification of Con1 (SEQ ID No: 6) (A), TuMV_J (B) or TEV (C)that have been expressed by pMAL vector and Tac promoter. The TEV plasmid (pRK793; Addgene #8827) contains maltose binding protein (MBP) upstream to the TEV gene. The purification was done with 5-ml Ni-Sepharose resin in home-packed columns. The molecular weights are around 27 kDa for Con1 (SEQ ID No: 6) and TuMV_J, and 28.6 for TEV. FT: flow- through; W: wash.Figure 9. (A) TEV or Con1 (SEQ ID No: 6) proteases concentration over a 3K MWCO Amiconconcentrator. Aliquots were collected every 30 min and the concentrations were measured by Nano-drop spectrophotometer in duplicate readings using the extension coefficient of 31970and 34950 M-1cm-1 for TEV (28560.45 Da) and Con1 (SEQ ID No: 6) (27660.22 Da),respectively. (B and C) The optical density tracking of 15 mg / ml proteases at 280 nm (B) and400 nm (C) was measured by Nano-drop in duplicate readings. (D) Aggregation index of TEV and Con1 (SEQ ID No: 6) proteases by dividing OD values at 280 nm by 400 nm.Figure 10. Purification of FRET(X) substrates (CFP-cleavage site-YFP) with different P1’ (X).(A) SDS-PAGE photo showing the purification of FRET-ENLYFQG substrate using 5-ml Ni- Sepharose resin in a home-packed column. (B) Affinity chromatography (AC) using HisTrap excel Ni column showing the eluted fractions on the SDS-PAGE. (C) Size exclusion chromatography (SEC) over Superdex 75 column showing the excluded peaks on SDS- PAGE. The purification steps were done by AKTA Go instrument. X = alanine, serine, aspartic, methionine or proline. The molecular weights are around 60 kDa as theoretically calculated. FT: flow-through; W: wash.Figure 11. Time-course screening of Con1 (SEQ ID No: 6) protease specificity againstdifferent FRET substrates. (A and B) Screening Con1 (SEQ ID No: 6), TuMV_J or TEVspecificity against the different cleavage sites, EAVYHQS (A) or ENLYFQG (B), introduced to FRET substrates. The reaction was composed of 1 µM enzyme and 2 µM substrate final concentrations and the fluorescence was measured over 2 hours period at room temperatureby the plate reader. Cleavage is a percentage of the maximum. The FRET ratio was calculatedby dividing the fluorescence intensity at 480 nm (donor) by that at 520 nm (acceptor), blank subtraction and percentage calculation. All the results are the mean of 2 technical repeats ± standard deviation.Figure 12. Time-course screening specificity of Con1 (SEQ ID No: 6) (A) and TuMV_J (B)against EAVYHQ-P. The reaction mixture contains 0.4 µM protease with the indicatedsubstrate concentration and the fluorescence was gained over 17 hours at room temperatureby the plate reader. The results are the mean of 2 technical repeats (solid lines) ± standard deviation (dotted lines). The curves were fitted by the exponential plateau equation on GraphPad Prism (black lines) and the reaction rates were calculated.Figure 13. Kinetic studies of Con1 (SEQ ID No: 6) protease activity against different P1’ FRETsubstrates where P1’ = alanine (A), aspartic (B), or methionine (C). The reaction mixture contains 0.4 µM Con1 (SEQ ID No: 6) with the indicated substrate concentration and the fluorescence was gained over 17 hours at room temperature by the plate reader. The results are the mean of 2 technical repeats (solid lines) ± standard deviation (dotted lines). The curves were fitted by the exponential plateau equation on GraphPad Prism (black lines) and the reaction rates were calculated.Figure 14. Kinetic studies of TuMV_J protease activity against different P1’ FRET substrateswhere P1’ = alanine (A), aspartic (B), or methionine (C). The reaction mixture contains 0.4 µM TuMV_J with the indicated substrate concentration and the fluorescence was gained over 17 hours at room temperature by the plate reader. The results are the mean of 2 technical repeats (solid lines) ± standard deviation (dotted lines). The curves were fitted by the exponential plateau equation on GraphPad Prism (black lines) and the reaction rates were calculated.Figure 15. Purification of His-tagged proteins of interest (POI) with Con1 (SEQ ID No: 6)cleavage site substrates (His-cleavage site-POI). The purification was done by one-step affinity chromatography (AC) using HisTrap excel Ni column on an AKTA Go instrument and the eluted fractions are shown on the SDS-PAGEs. The first residue of each POI is highlighted in red, and the molecular weights are indicated. Figure 16. (A) Con1 (SEQ ID No: 6) cleaved 12 different amino acids at the P1’ position representing different groups, including hydrophobic amino acids (A, F, I, M and W), polar uncharged amino acids (Q, S and T), positively charged amino acids (R), negatively charged amino acids (D and E), and G. It cleaved 100% of the His-tag in all cases after overnightincubation at a 1:100 (protease: substrate) molar ratio. (B) The C-terminal truncated versionof Con1 (Δ222-234Con1: SEQ ID NO: 4) has P1’ tolerance as it cleaved 12 different residues atthe P1’ position representing different groups, including hydrophobic amino acids (A, F, I, Mand W), polar uncharged amino acids (Q, S and T), positively charged amino acids (R), negatively charged amino acids (D and E), and G. It cleaved 100% of the His-tag in all cases after overnight incubation at a 1:100 (protease: substrate) molar ratio. Figure 17. Graph showing the effect of the deletion of seven residues (228-234) and thirteenresidues (222-234) from the C-terminal of Con1 (SEQ ID No: 6) on the activity.RESULTS Consensus protease design A consensus sequence was designed based on TEV protease and TEV-like proteases (Fig. 1A). The entire consensus sequence is 243 amino acids. In the sequence, 207 residues withmore than 40% conservation amongst the eight aligned sequences. For the remaining 36positions, the residues were chosen based on the degree of conservation and hydrophilicity.We created an exemplary protease sequence Con1 (SEQ ID No: 5). We further evaluated theeffect of C-terminal (235-243) deletion on solubility which was found to improve the solubilityof all the studied sequences (Fig.1B). This is consistent with that most of the TEV protease studies have been performed on versions in which the C-terminal residues are deleted [8, 11, 19]. The structural analysis of the TEV protease (PDB: 1LVB; Fig.1 C; Fig.7) and the predicted Alphafold2 model of the full-length Con1 (SEQ ID No: 5) (Fig.1 D), revealed the high degree of homology between Con1 (SEQ ID No: 5) and TEV protease with an extended C-terminal ofCon1 (SEQ ID No: 5) (Fig.1 E) [19, 20]. We designed genes encoding Con1 (SEQ ID No: 6) or TuMV_J with C-terminal (235-243) deleted, and codons optimized for expression in E. coli. For TEV protease, the expressionplasmid pRK793 (Addgene #8827

[0011] ) was used. All three proteases were expressed insoluble form and purified (Fig.8). Con1 solubility A protease with high solubility is a key goal of this study. Using Con1 (SEQ ID No: 6) as an exemplary protease, protein solubility was experimentally assessed. We tracked the maximum concentration (mg / ml) of either Con1 (SEQ ID No: 6) or TEV proteases that can be achieved by concentrating until the solubility threshold beyond which precipitation occurred

[0010] . Afterregular measurements, the plateau indicates that the Con1 (SEQ ID No: 6) solubility threshold(~30 mg / ml) is 2x higher than that of TEV (~15 mg / ml; Fig.9A) which is consistent with the predicted data. Furthermore, we observed faster aggregation of TEV protease than Con1 (SEQ ID No: 6) indicating the high solubility and stability of Con1 (SEQ ID No: 6) at highconcentrations (Fig. 9B-D).Con1 (SEQ ID No: 6) cleaves EAVYHQ-S but not ENLYFQ-G The consensus sites for substrate recognition and cleavage are ENLYFQ-G for TEV (Fig.2A) and EAVYHQ-S for TuMV_J (Fig. 2B) [9, 19, 21, 22]. Using Con1 (SEQ ID No: 6) as an exemplary protease, we assessed the specificity of TEV, TuMV_J and Con1 (SEQ ID No: 6) using a FRET-based assay in which different potential substrate cleavage sequences were incorporated between cyan fluorescent protein (CFP, donor) and yellow fluorescent protein (YFP, acceptor; Fig.2C; Fig.10) [23, 24]. FRET substrate with EAVYHQ-S site (1593 bp): (6xHis-linker-CFP-linker-cleavage site-linker-YFP) atgcggggttctcatcatcatcatcatcatGGTATGGCTAGCATGACTGGTGGACAGCAAATGGGT CGGGATCTGTACGACGATGACGATAAGGATCCGGGCCGCATGGTGAGCAAGGGCGAGGAGCTGTTC ACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCC GGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAG CTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTGGGGCGTGCAGTGCTTCAGCCGCTAC CCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGC ACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACC CTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAG CTGGAGTACAACTACATCAGCCACAACGTCTATATCACCGCCGACAAGCAGAAGAACGGCATCAAG GCCAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAG AACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCC CTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGG ATCGGTACCGTAGGATTTCTAACAGCGACCGAAGCAGTGTATCATCAATCCCTGGATTCCATCACC GTTAACGGTACCGGTGGAATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTG GTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCC ACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACC CTCGTGACCACCTTCGGCTACGGCCTGCAGTGCTTCGCCCGCTACCCCGACCACATGAAGCAGCAC GACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGAC GGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTG AAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGC CACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCAC AACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGC CCCGTGCTGCTGCCCGACAACCACTACCTGAGCTACCAGTCCGCCCTGAGCAAAGACCCCAACGAG AAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAG CTGTACAAG ENLYFQ-G site: GAAAATCTTTATTTTCAAGGT See materials and methods for more details. Con1 (SEQ ID No: 6) and TuMV_J cleave EAVYHQ-S but not ENLYFQ-G (Fig.2D and E; Fig.11) which is consistent with the reported TuMV_J specificity [21, 25, 26]. Con1 and TuMV_J cleavage of EAVYHQ-X substrates We prepared FRET substrates with different potential cleavage sites, EAVYHQ-X (X= alanine, methionine, aspartic or proline; Fig. 10). We used the FRET-based assay to determine the kinetic parameters of Con1 (SEQ ID No: 6) and TuMV_J against substrates with different residues at the P1’ position. As expected, neither enzyme cleaves a substrate with proline at the P1’ position (Fig.12).In the present example, Con1 (SEQ ID No: 6) and TuMV_J cleaved substrates alanine, methionine and aspartic acid, with Con1 (SEQ ID No: 6) exhibiting a higher catalytic rate than TuMV_J against them (Fig.3; Fig.13 and 14). The kinetic parameters of Con1 are indicated in Table 2, but the low catalytic activity of TuMV_J precluded accurate calculations of kinetic parameters. Substrates with different residues at the P1’ position arecleaved at similar rates by Con1 (SEQ ID No: 6), differences are primarily in Km.Table 2. Summary of the kinetic parameters of Con1(SEQ ID No: 6). Substrate P1’ Km (µM) Kcat (s-1) Kcat / Km (s-1 M-1)Ala 12.94 ± 3.0 0.012 ± 0.002 924.16 ± 45.16Asp 23.40 ± 3.7 0.011 ± 0.004 454.42 ± 70.00Met 23.13 ± 6.5 0.011 ± 0.003 476.80 ± 03.03Ser 10.60 ± 2.4 0.013 ± 0.002 1247.65 ± 66.33The results are mean ± standard deviation of 2 independent experiments with 2 technical replicates each. Dependency of Con1 (SEQ ID No: 6) activity on pH and temperature To ascertain the conditions in which Con1 (SEQ ID No: 6) could be used in practice, we characterized Con1 (SEQ ID No: 6) activity at different pHs and different temperatures. For the pH studies, Con1 (SEQ ID No: 6) protease and substrate were pre-incubated in buffers of the desired pHs including Sodium Citrate (pH 4-6), Tris (pH 7-9), or Sodium Borate (pH 10), then the reaction was initiated by mixing the substrate with enzyme. This optimal pH for Con1 (SEQ ID No: 6) protease is pH 7.7 (Fig.4A). To investigate the effect of temperature on the proteolysis, Con1 (SEQ ID No: 6) and substrate were pre-incubated for 10 minutes at various temperatures (4, 15, 25, 37, 50 or 70°C) in buffer A. Then, the proteolytic reaction was initiated by mixing the substrate with enzyme. The results indicate that Con1 (SEQ ID No: 6) has optimum activity at a temperature of 24°C (Fig.4B). Con1 (SEQ ID No: 6) protease cleaves His purification tag of POI independent of P1’ A key goal of this study was to develop a protease that could be used to remove N-terminal affinity purification tags from a protein of interest (POI) after the affinity purification. Therefore, we designed different His tagged proteins of interest (POI) with P6-P1 of the Con1 (SEQ ID No: 6) cleavage site (EAVYHQ) directly upstream of the POI. The first amino acid of the POI contributes the P1’ residue. These POI are a designed ankyrin repeat protein (DARPin)

[0027] ,urate oxidase enzyme (Uricase)

[0028] and green fluorescent protein (GFP)

[0029] which start withglycine, serine or methionine, respectively (Fig.5A; Fig.15). We additionally prepared 12 different His tagged DARPins by mutating the N-terminal to broad range of amino acids covering different amino acids grouping including hydrophobic amino acids (A, F, I, M and W), polar uncharged amino acids (Q, S and T), positively charged aminoacids (R), negatively charged amino acids (D and E), and G. This enabled us to assess theP1’ tolerance of Con1 (SEQ ID No: 6) by cleaving at all the 12 residues indicating the broadP1’ specificity of Con1 (SEQ ID No: 6) (Fig.16A).DARPin (528 bp): (7xHis-linker-cleavage site-DARPin) atgcatcatcatcatcatcatcacGGCGGTAGCGAAGCAGTGTATCATCAAGGTTCGGACTTAGGCCG CAAGTTACTGGAAGCTGCACGCGCGGGTCAAGACGACGAGGTACGCATCCTTATGGCCAATGGTGCGG ACGTCAACGCCGCCGATAATACCGGAACAACTCCTTTGCACCTCGCAGCATATTCCGGCCATTTAGAG ATTGTAGAGGTACTGCTTAAGCACGGTGCCGATGTAGATGCGTCGGACGTATTCGGGTACACGCCTCT CCACCTCGCCGCTTATTGGGGCCACTTAGAAATTGTAGAGGTGCTCTTGAAGAATGGTGCCGACGTGA ATGCTATGGATTCAGACGGCATGACACCGCTCCATCTGGCCGCGAAATGGGGTTATCTGGAGATCGTC GAGGTTCTCTTGAAACATGGCGCCGACGTCAATGCACAAGATAAGTTCGGCAAGACTGCCTTCGACAT CAGCATCGACAACGGTAATGAGGACCTGGCCGAGATCCTTCAAAAGTTGAAT Uricase (957 bp): (7xHis-linker-cleavage site-Uricase) atgcatcatcatcatcatcatcacGGCGGTAGCGAAGCAGTGTATCATCAAAGCACTACACTGTCTTC ATCGACGTACGGGAAGGATAACGTTAAGTTTCTCAAGGTTAAGAAGGACCCTCAGAATCCGAAGAAGC AGGAAGTTATGGAAGCCACGGTTACGTGCTTGTTAGAGGGTGGATTCGATACCAGTTACACTGAGGCG GACAACTCGTCGATTGTGCCCACAGATACCGTGAAGAACACTATCCTCGTTCTTGCGAAAACGACCGA GATCTGGCCTATCGAGCGCTTTGCCGCGAAGCTTGCCACTCATTTTGTAGAGAAATACTCCCATGTCA GCGGCGTTTCTGTGAAAATCGTTCAGGATCGTTGGGTCAAGTACGCCGTAGACGGCAAGCCCCATGAT CACAGCTTCATCCACGAGGGCGGAGAAAAGCGTATTACTGACCTTTACTATAAGCGCAGCGGCGATTA TAAGTTGTCCAGTGCGATCAAAGATCTGACCGTGCTGAAGTCAACCGGCAGTATGTTCTATGGGTATA ATAAATGTGATTTTACGACGTTACAGCCAACCACTGATCGTATCCTTAGTACGGACGTCGACGCAACA TGGGTGTGGGACAACAAGAAAATCGGAAGCGTTTATGACATCGCTAAAGCCGCGGACAAGGGCATCTT TGATAATGTCTACAATCAGGCGCGCGAGATTACACTGACAACGTTCGCTCTGGAAAACTCCCCGAGCG TCCAAGCTACAATGTTTAATATGGCGACGCAAATCCTGGAGAAAGCATGTTCGGTTTACTCGGTTTCG TACGCGTTACCGAATAAGCATTATTTCCTGATTGACTTAAAATGGAAGGGCCTCGAAAACGATAATGA ACTGTTTTACCCGTCTCCGCATCCGAATGGCCTGATCAAGTGTACCGTGGTACGTAAGGAAAAGACGA AGTTA GFP (1182 bp): (6xHis-linker-cleavage site-GFP-TRAP) atgcatcaccatcaccatcacGATTACGATATCCCAACGACCGAAGCAGTGTATCATCAAATGCGTAA AGGCGAAGAACTGTTCACGGGCGTAGTTCCGATTCTGGTCGAGCTGGACGGCGATGTGAACGGTCATA AGTTTAGCGTTCGCGGTGAAGGTGAGGGCGACGCGACCAACGGCAAACTGACCCTGAAGTTCATCTGC ACCACCGGTAAACTGCCGGTGCCTTGGCCGACCTTGGTGACGACGTTGACGTATGGCGTGCAGTGTTT TGCGCGTTATCCGGACCACATGAAACAACACGATTTCTTCAAATCTGCGATGCCGGAGGGTTACGTCC AGGAGCGTACCATTTCCTTCAAGGATGATGGCTACTACAAAACTCGCGCAGAGGTTAAGTTTGAAGGT GACACGCTGGTCAATCGTATCGAATTGAAGGGTATCGACTTTAAAGAGGATGGTAACATTCTGGGCCA TAAACTGGAGTATAACTTCAACAGCCATAATGTTTACATTACGGCAGACAAGCAAAAGAACGGCATCA AGGCCAATTTCAAGATTCGCCACAATGTTGAGGACGGTAGCGTCCAACTGGCCGACCATTACCAGCAG AACACCCCAATTGGTGACGGTCCGGTTTTGCTGCCGGATAATCACTATCTGAGCACCCAAAGCGTGCT GAGCAAAGATCCGAACGAAAAACGTGATCACATGGTCCTGCTGGAATTTGTGACCGCTGCGGGCATCA CCCACGGTATGGACGAGCTGTATAAAGCCGGTGGTGGTTCTGGTGGTGGTTCGAAGCAGGCACTGAAA GAAAAAGAGCTGGGGAACGATGCCTACAAGAAGAAAGACTTTGACACAGCCTTGAAGCATTACGACAA AGCCAAGGAGCTGGACCCCACTAACATGACTTACATTATCAATCAAGCAGCGGTATACTTTGAAAAGG GCGACTACAATAAGTGCCGGGAGCTTTGTGAGAAGGCCATTGAAGTGGGGAGAGAAAACCGAGAGGAC TATCGATGGATTGCCATTGCATATGCTCGAATTGGCAACTCCTACTTCAAAGAAGAAAAGTACAAGGA TGCCATCCATTTCTATAACAAGTCTCTGGCAGAGCACCGAACCCCAAAGGTGCTAAAAAAGTGCCAAC AGGCGGAGAAAATCCTGAAGGAGCAA After affinity purification using Ni-NTA, the His tagged POI was incubated overnight with Con1 (SEQ ID No: 6). We assessed removal of the N-terminal purification tag by separating the proteins using SDS PAGE. As seen in Figure 5B, the tags are all completely removed, and not undesired additional cleavage is observed. Con1 is C-terminal sensitive We evaluated the effect of the C-terminal on Con1 (SEQ ID No: 6) activity by truncating 7 or13 residues, respectively. The truncated versions are Δ228-234Con1 (by deleting residues 228- 234: SEQ ID NO: 3) andΔ222-234Con1 (by deleting residues 222-234: SEQ ID NO: 4). Deleting7 residues (228-234) showed to slow down the catalysis process compared to Con1 (SEQ IDNo: 6), on the other hand, deleting 13 residues (222-234) enhanced the catalytic activity (Fig. 17). Since the C-terminal is close to the active site, this explains how much sensitive it is onthe substrate binding / releasing. However, this deletion didn’t affect the P1’ tolerance of theΔ222-234Con1 (SEQ ID NO: 4) as shown in Figure 16B by cleaving all the studied 12 His tagged POI. Discussion Affinity purification of recombinant proteins is powerful tool especially for functional products [1-4]. However, the affinity tags are favourable to be removed for maximising the functionality of the target POI by incorporating a protease cleavage site between the affinity tag and the POI [4]. TEV protease is one of the most highly specific proteases that is favoured for this purpose, however, it exhibits poor solubility that reduce the yield, optimal storage, and usage [10, 11]. In addition, its restriction to Gly or Ser at the P1’ position of the cleavage site, impairs its applicability for tag removal from functional POIs as an extra Gly or Ser to be added to the POI that may reduce their functionality [9]. Therefore, we have been motivated to design,develop, and characterize a new protease, Con1 (SEQ ID No: 6), with high solubility, highsubstrate specificity and P1’ residue tolerance. Con1 behaved higher solubility than TEV protease which is consistent with the prediction approach. To date, great effort has been done to improve TEV protease solubility, however, the solubility still limited [10, 30]. Con1 (SEQ ID No: 6) was designed to comprise the substrate binding site of TuMV_J protease which specifically recognize EAVYHQ but not ENLYFQ cleavage site. The specificity arises from the P1 (Q), P2 (H) and P4 (V) residues of the cleavage site as previously reported for TuMV proteases [21, 22]. We demonstrate that, in addition to the high substrate specificity, Con1 (SEQ ID No: 6) cleaved substrates with different residues at the P1’ position. Such P1’ cleavage independency would hire Con1 (SEQ ID No: 6) for various applications that requires free POI without additional residues. This is consistent with the efficiency of His tag removal by Con1 (SEQ ID No: 6) from different POIs that comprise different N-terminus residues without incorporating a P1’ residue at the cleavage site. In the kinetic terms, Con1 (SEQ ID No: 6) show a faster catalytic turnover than TuMV_J protease against the studied substrates. The optimal conditions for Con1 (SEQ ID No: 6) activity were demonstrated at pH 7.7 and temperature of 24°C that shows the efficient cleavage of substrates. The protease we report represents a powerful tool for the downstream applications comprising the release of functional proteins. This versatile protease is expressive, soluble, stable, active, substrate-specific, and P1’ independent. Materials and Methods FASTA search for TEV-like proteases The full-length sequence of the TEV protease (1-242), also known as Nuclear Inclusion proteina (NIa), was taken from the UniProt database (UniProt code of the poly-protein origin: P04157;Genome position in TEV virus: 2038-2279)

[0013] . We performed a sequence similarity search of this sequence against the UniProt Knowledgebase and UniProtKB / Swiss-Prot isoformsdatabases using the FASTA suite of programmes

[0012] . We shortlisted the results by excludingvariants with ≤50% (which is the desired threshold) and ≥90% (which are mutant TEV proteaseversions) identity to the TEV protease sequence. Intrinsic solubility prediction We used the CamSol, web server to predict protein solubility based on the amino acid sequence [15-17]. Scores >0 indicate soluble proteins, whereas scores of <0 indicate poorly soluble proteins. The higher the score, the more soluble a protein is predicted to be. TEV structure analysis We used the co-crystal structure of an inactive mutant of TEV protease (C15A) in complex with a peptide substrate (TENLYFQSGT) for our analyses (PDB ID: 1LVB)

[0031] . We used ICM- BrowserPro software to identify the substrate-enzyme interactions based on distances between atoms in the receptor and atoms in the substrate

[0032] . Substrate-enzyme interactions were identified by encountering on the van der Waals interaction map of the binding site and the surfaces

[0033] . Using this method, we identified the residues on the protease that are involved in binding the substrate P1’ (S), P1 (Q), P3 (Y), and P6 (E) positions which are responsible for the specificity were identified [9]. These are: T30, L32, H46, D81, T146, D148, G149, C151, H167, S168, S170, N171, N174, N176, Y178, H214, S219, and K220. Alignment and consensus We aligned the full-length sequence of the selected proteases (TEV, TVMV, TuMV_Q, TuMV_J, PPV, OMV, LMVE and LMV0) using the Clustal Omega tool using the default settings

[0012] . The alignment was then exported and analyzed by the Jalview 2.11.2.6 software

[0018] . The alignment was coloured by the Clustal scheme with >75% identity threshold. The software automatically generates the consensus sequence as a percentage (>40% conservation) of the modal residue per column. A clear consensus backbone that comprised 207 conserved residues was identified. At this point, the above-mentioned substrate binding positions were retained in the consensus designs to the equivalent TuMV_J residues. Theseare: N30, A170, I176, S214, E219, and S220. The remaining non-conserved 36 positions werefilled by choosing residues of the original sequences based on their hydrophilicity (Cys residues were excluded). Finally, C-terminal (235-243) was subjected to deletion to improve solubility since most of TEV protease studies were performed using a C-terminal truncated versions [8, 11, 19]. Plasmid construction: ProteasesWe obtained synthetic genes encoding 7xHis tagged TuMV_J and Con1 (SEQ ID No: 6) (IDT,Belgium). Con1 (726 bp; IDT, Belgium): (7xHis-Con1) atgcatcatcatcatcatcatcacAGCAAATCTTTATTCCGTGGATTACGCGATTACAACCCGATC GCCAGCAATATTTGTCATTTGACTAATGAGAGCGACGGTCATAGTAACTCGTTGTATGGCATCGGG TTTGGGCCGCTTATTATTACAAACCAACATTTGTTTCGTCGCAACAACGGGGAACTTACGATTCAA TCTCGCCATGGCGAGTTCGTAGTTAAGAACACCACGCAACTTAAGCTTCTGCCAATCGATGGGCGT GATATCCTTATCATTCGTTTGCCCAAAGATTTCCCACCCTTTCCACAGAAACTTAAATTCCGCCAA CCTGAAAAAGGCGAACGCATTTGCTTGGTTGGCTCGAATTTTCAGACCAAGTCCATCACCAGTACA GTCAGCGAAACATCAACAACCATGCCGGTGGAGAATAGTCAATTTTGGAAGCACTGGATTTCAACC AAAGACGGTCATTGTGGCCTGCCTCTTGTTTCCACGAAGGATGGCAAGATCCTGGGCATCCATTCT TTGGCGAATTTTACCAATACGATCAACTACTTCGCCGCATTCCCAGAAGATTTTGAGGAGACGTAT CTTCATACCCAAGAAGCTCAAGAATGGGTAAAGCACTGGAAATATAACCCGGATGCGATCTCCTGG GGGAGCCTTAACCTTCAAGAAAGCCAACCAGAAGAGCCTTTCAAAATCGTAAAATTAGTAACGGAC TuMV_J (732 bp; IDT, Belgium): (7xHis-TuMV_J) atgcatcatcatcatcatcatcacTCGAATTCTATGTTCCGTGGGTTACGCGATTACAACCCGATT GCCAATAATATTTGCCATTTAACGAACGTGTCCGATGGCGCATCGAACTCTCTGTATGGAGTCGGA TTTGGGCCGTTAATCTTAACCAATCGTCATCTGTTCGAACGCAATAATGGGGAATTGGTAATTAAG TCTCGTCACGGTGAATTTGTCATCAAGAATACTACCCAGTTGCACCTTTTACCGATTCCCGATCGT GACTTGCTGCTGATTCGCCTGCCCAAGGACATCCCCCCGTTCCCACAAAAACTGGGGTTCCGCCAA CCCGAAAAGGGGGAGCGCATTTGTATGGTTGGCTCCAATTTTCAGACAAAGAGTATCACAAGTGTA GTCTCGGAAACGAGCACAATCATGCCCGTCGAGAACAGTCAGTTTTGGAAGCACTGGATTAGCACG AAGGATGGGCAATGCGGTCTGCCTATGGTTAGCACCAAGGACGGTAAAATTCTGGGGTTACATAGT TTAGCCAATTTTCAAAACTCCATCAATTATTTCGCCGCCTTTCCAGACGACTTTGCTGAAAAATAT CTGCATACGATCGAGGCACACGAGTGGGTCAAGCATTGGAAATACAACACATCTGCAATTTCATGG GGATCGCTGAACATTCAAGCATCACAACCGGGTGGCTTATTTAAGGTCAGTAAACTTATTTCCGAC TTAGAT These were cloned by PCR to a pMAL expression vectors using AQUA cloning

[0034] .Following transformation into the E. coli strain Top 10, the desired plasmids were identifiedby DNA sequencing. See supplementary materials for genes and primers sequences. Plasmid construction: FRET substrates The synthetic genes encoding CFP-linker-cleavage-site-linker-YFP (IDT, Belgium) were cloned by digestion / ligation into the backbone of a pET28 expression vector. FRET substrate with EAVYHQ-S site (1593 bp): (6xHis-linker-CFP-linker-cleavage site-linker-YFP) atgcggggttctcatcatcatcatcatcatGGTATGGCTAGCATGACTGGTGGACAGCAAATGGGT CGGGATCTGTACGACGATGACGATAAGGATCCGGGCCGCATGGTGAGCAAGGGCGAGGAGCTGTTC ACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCC GGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAG CTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTGGGGCGTGCAGTGCTTCAGCCGCTAC CCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGC ACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACC CTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAG CTGGAGTACAACTACATCAGCCACAACGTCTATATCACCGCCGACAAGCAGAAGAACGGCATCAAG GCCAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAG AACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCC CTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGG ATCGGTACCGTAGGATTTCTAACAGCGACCGAAGCAGTGTATCATCAATCCCTGGATTCCATCACC GTTAACGGTACCGGTGGAATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTG GTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCC ACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACC CTCGTGACCACCTTCGGCTACGGCCTGCAGTGCTTCGCCCGCTACCCCGACCACATGAAGCAGCAC GACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGAC GGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTG AAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGC CACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCAC AACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGC CCCGTGCTGCTGCCCGACAACCACTACCTGAGCTACCAGTCCGCCCTGAGCAAAGACCCCAACGAG AAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAG CTGTACAAG The inserts and vector were digested by BamI / BsrGI enzymes and ligated by T4 ligase (NEB). Correct inserts were verified by DNA sequencing. Plasmid construction: His-tagged POI We designed constructs encoding N-terminal His tag separated from the POI by the P6-P1 residues of the Con1 (SEQ ID No: 6) cleavage site. The N-terminal residue of the POI contributes the P1’ position of the cleavage site. The correct clones were identified by DNA sequencing. Protein expressionProteins were expressed using standard protocols of BL21(DE3) E. coli strain. Briefly, colonieswere inoculated into 5-ml of LB plus antibiotic and grown overnight at 37°C. The next day, the overnight cultures were added to LB plus antibiotic media and grown with shaking at 37°C until the OD600was 0.6. IPTG was then added to a final concentration of 1mM, the temperature was dropped to 20°C and the incubation continued overnight. The following day cells were harvested by centrifugation and stored at -80°C. Protein purification: proteases Cell paste (from 100 ml culture) was re-suspended in 10 mL buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5% (v / v) glycerol) supplemented with protease inhibitor cocktail tablets (cOmpleteTM, Roche). The suspensions were kept on ice, and cells disrupted by sonication. After sonication, the solution was centrifuged to remove insoluble material, and the pellet was discarded. Supernatants were applied to a 5ml Ni-NTA agarose column pre-equilibrated in buffer A. After sample loading, the resin was washed with buffer A plus 20 mM imidazole. Finally, proteins were eluted by buffer A with an imidazole concentration from 75 to 500mM. The fractions were analysed by SDS–PAGE, pooled, buffer exchanged into buffer A and concentrated using 3 kDa MW-CO Amicon (Merck-Millipore). Protein purification: FRET substrates & His-tagged POI Cell lysates were prepared as described above. The supernatant was applied to a HisTrap excel column (Cytiva) equilibrated in buffer A and eluted by a gradient of 0–1M Imidazole, in buffer A. The peak fractions were analysed by SDS–PAGE, pooled, buffer exchanged and concentrated using a 3 kDa MW-CO Amicon, then stored stored at −80°C. FRET substrates were additionally purified by gel filtration using a Superdex-75 column, equilibrated, and run in buffer A. After checking the purity on SDS–PAGE, appropriate fractions were pooled, concentrated using a 3 kDa MW-CO Amicon, and then stored at −80°C. Protease solubility The solubility of TEV and Con1 (SEQ ID No: 6) proteases were studied by concentrating proteins and collecting aliquots over the time as previously described

[0010] . The 5 ml aliquots of 1.5 mg / ml proteins were spun down at high speed to pellet any aggregates and the protein concentrations were measured by the Nano-drop spectrophotometer using the extinction coefficient of 31970 and 34950 M-1cm-1for TEV (MW= 28560.45 Da) and Con1 (SEQ ID No:6) (MW=27660.22 Da), respectively. The proteins were loaded onto a 5-ml 3 kDa MW-COAmicon and spun at speed of 4000 rpm at 4°C until the aggregation observed. A 100µl of 15 mg / ml of each protease were incubated overnight at room temperature, aliquots were collected over the time, and optical density (OD) at 280 nm and 400 nm were measured. The aggregation index was calculated by dividing OD at 280 nm by OD at 400 nm. Proteolytic activity Substrate specificityWe assessed the specificity of Con1 (SEQ ID No: 6), TEV and TuMV_J for the cleavage sitesEAVYHQ-S and ENLYFQ-G. Protease (1 µM) and substrate (2 µM) were incubated in Buffer A at room temperature. The mixtures were excited at 430 nm, and fluorescence emission intensity at 480 nm and 520 nm were measured every minute over 2 hours. The amount of cleavage was determined from the change in FRET ratio, calculated by intensity at 480 nm (donor) divided by the intensity at 520 nm (acceptor), after blank subtraction. Cleavage is a percentage of end point after saturation. All the experiments were performed in duplicate.To screen the P1’ specificity of Con1 (SEQ ID No: 6), FRET substrates with different cleavagesequences between the donor and acceptor fluorescent proteins were prepared: EAVYHQ-X (X= S, D, M or P). A 0.5 µM Con1 (SEQ ID No: 6) was mixed with 5 µM substrate in Buffer A and the fluorescence was measured over 5 hours at room temperature. Protease kinetics Kinetic parameters for the proteases Con1 (SEQ ID No: 6) and TuMV_J were determined by incubating 0.4 µM enzyme with the substrates at range of (1-24 µM) in Buffer A. Different FRET substrates EAVYHQ-X (X = Ala, Ser, Met or Asp) between the FRET donor (CFP) and acceptor (YFP) proteins were used to assess the influence of the identity of the P1’ position. The data was collected and analysed as previously described

[0023] . The full-scale range (FSR) was calculated for each wavelength by subtracting minimum emission (Emmin; which is at time zero intensity at 480 nm and the end-point intensity at 520 nm) from the maximum emission (Emmax; which is at end-point intensity at 480 nm and the time zero intensity at 520 nm by the following equations: FSR480 nm = Em480 nm max – Em480 nm minFSR520 nm = Em520 nm max – Em520 nm minThen, the emission values were normalised by subtracting the emission at time zero by the following equations: ΔEm480 nm = Em480 nm – Em480 nm t=0ΔEm520 nm = -1(Em520 nm – Em520 nm t=0)Then, the cleavage percentage (C%) were calculated by dividing the normalised values by the FSR values by the following equations: C480 nm%= 100 (ΔEm480 nm / FSR480 nm) C520 nm%= 100 (ΔEm520 nm / FSR520 nm) The cleaved product concentration [P] was calculated as a percent of the total substrate concentration at time zero by the following equation: [P] = (C%. [S]t=0) / 100 After that, the product concentrations (on the Y axis) were plotted vs. time (on the X axis) and the curves were fitted by the exponential plateau equation on GraphPad Prism software that properly fits frequency changes versus time: Where, Pmax is the maximum product concentration; Pmin is the minimum productconcentration; k is the rate constant; and t is the time. From the calculated rate constants, thereaction rates (V) were calculated by the following equation: V = K. PmaxFinally, the reaction rates (Y axis) were plotted vs. substrate concentration (X axis) and the curves were fitted by the Michaelis-Menten equation on GraphPad Prism software: V = (Vmax . [S]) / (Km + [S]) Where, V is the reaction rate; Vmax is the maximum velocity; [S] is the substrate concentration;and Km is the Michaelis constant. The Vmax and Km were obtained from the non-linearregression and the Kcat was calculated by the following equation: Kcat = Vmax / [E] Where, Kcat is the catalytic turnover number and [E] is the total enzyme concentration. pH dependency of Con1 (SEQ ID No: 6) activity To evaluate the pH dependence of the Con1 (SEQ ID No: 6) activity, we performed the FRET cleavage assay in buffers of different pH: 0.1 M Sodium citrate (pH 4, 5 or 6), 0.1 M Tris-HCl (pH 7, 8 or 9) and 0.1 M sodium borate (pH 10)

[0035] . The Con1 (SEQ ID No: 6) protease (0.5 µM) and the FRET substrate (2.5 µM) were pre-incubated in the different pH buffers at room temperature for 10 min. The reaction was initiated by mixing protease with substrate, and incubation continued for a further 2 hours at room temperature. The end-point fluorescence was measured at 480 and 520 nm and the amount of cleavage calculated from the change in FRET. Temperature dependency of Con1 (SEQ ID No: 6) activity We assessed the effect of different temperature on Con1 (SEQ ID No: 6) activity. The Con1 (SEQ ID No: 6) protease and the FRET substrate were pre-incubated in buffer A at different temperatures (4, 15, 25, 37, 50 and 70°C) for 10 min. The reaction was initiated by mixing Con1 (SEQ ID No: 6) with substrate to final concentrations of 0.5 µM and 2.5 µM, respectively and incubation continued for a further 2 hours. The end-point fluorescence was measured at 480 and 520 nm and the amount of cleavage was calculated. His tag removal from different substrates We used Con1 (SEQ ID No: 6) to cleave an N-terminal His purification tag from different POI.We mixed Con1 (SEQ ID No: 6) (0.5 µM) with the His tagged POI (5 µM) in Buffer A. Themixture was incubated at room temperature, overnight with gentle shaking. The cleavage was assessed by separating proteins on SDS-PAGE. C-terminal truncation and kineticsWe assessed the effect of the C-terminal truncation on the catalytic turnover of Con1 (SEQ IDNo: 6). We generated two versions:Δ228-234Con1 (by deleting residues 228-234: SEQ ID NO: 3) andΔ222-234Con1 (by deleting residues 222-234: SEQ ID NO: 4). Conventional PCR was used to generate the truncated versions and the protein expression and purification was performed as mentioned above. The FRET method was used to evaluate the activity of the truncated versions. P1’ tolerance screening To assess the P1’ tolerance of the proteases, we prepared 12 different His tagged POI by mutating the N-terminal of the DARPins to broad range of amino acids covering different amino acids grouping including hydrophobic amino acids (A, F, I, M and W), polar uncharged amino acids (Q, S and T), positively charged amino acids (R), negatively charged amino acids (Dand E), and G. The P1’ tolerance of Con1 (SEQ ID No: 6) and the truncated version, Δ222-234Con1 (SEQ ID NO: 4), were evaluated by the SDS-PAGE experiments where the His taggedPOIs were incubated with or without protease at 1:100 enzyme: substrate molar ratio for overnight at ambient temperature and then resolved on the gel. REFERENCES1. Waugh DS. Making the most of affinity tags. Trends Biotechnol. 2005;23(6):316-20. doi:10.1016 / j.tibtech.2005.03.012. PubMed PMID: 15922084.Raran-Kurussi S, Waugh DS. The ability to enhance the solubility of its fusion partners isan intrinsic property of maltose-binding protein but their folding is either spontaneous or chaperone-mediated. PLoS One. 2012;7(11):e49589. Epub 20121116. doi: 10.1371 / journal.pone.0049589. PubMed PMID: 23166722; PubMed Central PMCID: PMCPMC3500312.Schafer F, Seip N, Maertens B, Block H, Kubicek J. Purification of GST-Tagged Proteins.Methods Enzymol.2015;559:127-39. Epub 20150516. doi: 10.1016 / bs.mie.2014.11.005. PubMed PMID: 26096507.Costa S, Almeida A, Castro A, Domingues L. Fusion tags for protein solubility, purificationand immunogenicity in Escherichia coli: the novel Fh8 system. Front Microbiol.2014;5:63. Epub 20140219. doi: 10.3389 / fmicb.2014.00063. PubMed PMID: 24600443; PubMed Central PMCID: PMCPMC3928792.Le NTP, Phan TTP, Phan HTT, Truong TTT, Schumann W, Nguyen HD. Influence of N-terminal His-tags on the production of recombinant proteins in the cytoplasm of Bacillus subtilis. Biotechnol Rep (Amst). 2022;35:e00754. Epub 20220719. doi: 10.1016 / j.btre.2022.e00754. PubMed PMID: 35911505; PubMed Central PMCID: PMCPMC9326129.Block H, Maertens B, Spriestersbach A, Kubicek J, Schafer F. Proteolytic Affinity TagCleavage. Methods Enzymol. 2015;559:71-97. Epub 20150516. doi: 10.1016 / bs.mie.2014.11.009. PubMed PMID: 26096504.Waugh DS. An overview of enzymatic reagents for the removal of affinity tags. ProteinExpr Purif.2011;80(2):283-93. Epub 20110819. doi: 10.1016 / j.pep.2011.08.005. PubMed PMID: 21871965; PubMed Central PMCID: PMCPMC3195948.Raran-Kurussi S, Cherry S, Zhang D, Waugh DS. Removal of Affinity Tags with TEVProtease. Methods Mol Biol. 2017;1586:221-30. doi: 10.1007 / 978-1-4939-6887-9_14. PubMed PMID: 28470608; PubMed Central PMCID: PMCPMC7974378.Kapust RB, Tozser J, Copeland TD, Waugh DS. The P1' specificity of tobacco etch virusprotease. Biochem Biophys Res Commun. 2002;294(5):949-55. doi: 10.1016 / S0006- 291X(02)00574-0. PubMed PMID: 12074568.Cabrita LD, Gilis D, Robertson AL, Dehouck Y, Rooman M, Bottomley SP. Enhancing thestability and solubility of TEV protease using in silico design. Protein Sci. 2007;16(11):2360-7. Epub 20070928. doi: 10.1110 / ps.072822507. PubMed PMID: 17905838; PubMed Central PMCID: PMCPMC2211701.Kapust RB, Tozser J, Fox JD, Anderson DE, Cherry S, Copeland TD, et al. Tobacco etchvirus protease: mechanism of autolysis and rational design of stable mutants with wild- type catalytic proficiency. Protein Eng. 2001;14(12):993-1000. doi: 10.1093 / protein / 14.12.993. PubMed PMID: 11809930.Madeira F, Pearce M, Tivey ARN, Basutkar P, Lee J, Edbali O, et al. Search andsequence analysis tools services from EMBL-EBI in 2022. Nucleic Acids Res. 2022;50(W1):W276-9. Epub 20220412. doi: 10.1093 / nar / gkac240. PubMed PMID:35412617; PubMed Central PMCID: PMCPMC9252731.UniProt C. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res.2023;51(D1):D523-D31. doi: 10.1093 / nar / gkac1052. PubMed PMID: 36408920; PubMed Central PMCID: PMCPMC9825514.Rawlings ND, Barrett AJ, Thomas PD, Huang X, Bateman A, Finn RD. The MEROPSdatabase of proteolytic enzymes, their substrates and inhibitors in 2017 and a comparison with peptidases in the PANTHER database. Nucleic Acids Res.2018;46(D1):D624-D32.doi: 10.1093 / nar / gkx1134. PubMed PMID: 29145643; PubMed Central PMCID:PMCPMC5753285.Sormanni P, Aprile FA, Vendruscolo M. The CamSol method of rational design of proteinmutants with enhanced solubility. J Mol Biol.2015;427(2):478-90. Epub 20141014. doi: 10.1016 / j.jmb.2014.09.026. PubMed PMID: 25451785.Sormanni P, Amery L, Ekizoglou S, Vendruscolo M, Popovic B. Rapid and accurate insilico solubility screening of a monoclonal antibody library. Sci Rep.2017;7(1):8200. Epub 20170815. doi: 10.1038 / s41598-017-07800-w. PubMed PMID: 28811609; PubMed Central PMCID: PMCPMC5558012.Sormanni P, Vendruscolo M. Protein Solubility Predictions Using the CamSol Method inthe Study of Protein Homeostasis. Cold Spring Harb Perspect Biol. 2019;11(12). Epub 20191202. doi: 10.1101 / cshperspect.a033845. PubMed PMID: 30833455; PubMed Central PMCID: PMCPMC6886446.Waterhouse AM, Procter JB, Martin DM, Clamp M, Barton GJ. Jalview Version 2--amultiple sequence alignment editor and analysis workbench. Bioinformatics. 2009;25(9):1189-91. Epub 20090116. doi: 10.1093 / bioinformatics / btp033. PubMed PMID: 19151095; PubMed Central PMCID: PMCPMC2672624.Phan J, Zdanov A, Evdokimov AG, Tropea JE, Peters HK, 3rd, Kapust RB, et al. Structuralbasis for the substrate specificity of tobacco etch virus protease. J Biol Chem. 2002;277(52):50564-72. Epub 20021010. doi: 10.1074 / jbc.M207224200. PubMed PMID: 12377789.Mirdita M, Schutze K, Moriwaki Y, Heo L, Ovchinnikov S, Steinegger M. ColabFold:making protein folding accessible to all. Nat Methods. 2022;19(6):679-82. Epub 20220530. doi: 10.1038 / s41592-022-01488-1. PubMed PMID: 35637307; PubMed Central PMCID: PMCPMC9184281.Kang H, Lee YJ, Goo JH, Park WJ. Determination of the substrate specificity of turnipmosaic virus NIa protease using a genetic method. J Gen Virol.2001;82(Pt 12):3115-7. doi: 10.1099 / 0022-1317-82-12-3115. PubMed PMID: 11714990.Ji Seon Han D-HK, Kwan Yong Choi. Potyvirus NIa Protease. In: Alan Barrett NR, J.Woessner, editor. Handbook of Proteolytic Enzymes.3rd Edition: Elsevier; 2012. p.2427- 32.Hay MHTRT. FRET-Based In Vitro Assays for the Analysis of SUMO Protease Activities.In: Ulrich HD, editor. METHODS IN MOLECULAR BIOLOGY: SUMO Protocols. 497. Totowa, NJ.: Humana Press; 2009. p.253–68.Gu H, Lalonde S, Okumoto S, Looger LL, Scharff-Poulsen AM, Grossman AR, et al. Anovel analytical method for in vivo phosphate tracking. FEBS Lett. 2006;580(25):5885- 93. Epub 20061002. doi: 10.1016 / j.febslet.2006.09.048. PubMed PMID: 17034793; PubMed Central PMCID: PMCPMC2748124.Kim DH, Park YS, Kim SS, Lew J, Nam HG, Choi KY. Expression, purification, andidentification of a novel self-cleavage site of the Nla C-terminal 27-kDa protease of turnip mosaic potyvirus C5. Virology. 1995;213(2):517-25. doi: 10.1006 / viro.1995.0024. PubMed PMID: 7491776.Han HE, Sellamuthu S, Shin BH, Lee YJ, Song S, Seo JS, et al. The nuclear inclusion a(NIa) protease of turnip mosaic virus (TuMV) cleaves amyloid-beta. PLoS One. 2010;5(12):e15645. Epub 20101220. doi: 10.1371 / journal.pone.0015645. PubMed PMID: 21187975; PubMed Central PMCID: PMCPMC3004936.Stumpp MT, Binz HK, Amstutz P. DARPins: a new generation of protein therapeutics.Drug Discov Today. 2008;13(15-16):695-701. Epub 20080711. doi: 10.1016 / j.drudis.2008.04.013. PubMed PMID: 18621567.Pierzynowska K, Deshpande A, Mosiichuk N, Terkeltaub R, Szczurek P, Salido E, et al.Oral Treatment With an Engineered Uricase, ALLN-346, Reduces Hyperuricemia, and Uricosuria in Urate Oxidase-Deficient Mice. Front Med (Lausanne).2020;7:569215. Epub 20201124. doi: 10.3389 / fmed.2020.569215. PubMed PMID: 33330529; PubMed Central PMCID: PMCPMC7732547.Remington SJ. Green fluorescent protein: a perspective. Protein Sci. 2011;20(9):1509-19. Epub 20110719. doi: 10.1002 / pro.684. PubMed PMID: 21714025; PubMed Central PMCID: PMCPMC3190146.Tropea JE, Cherry S, Waugh DS. Expression and purification of soluble His(6)-taggedTEV protease. Methods Mol Biol. 2009;498:297-307. doi: 10.1007 / 978-1-59745-196- 3_19. PubMed PMID: 18988033.Sun P, Austin BP, Tozser J, Waugh DS. Structural determinants of tobacco vein mottlingvirus protease substrate specificity. Protein Sci. 2010;19(11):2240-51. doi: 10.1002 / pro.506. PubMed PMID: 20862670; PubMed Central PMCID: PMCPMC3005794.Neves MA, Totrov M, Abagyan R. Docking and scoring with ICM: the benchmarkingresults and strategies for improvement. J Comput Aided Mol Des. 2012;26(6):675-86. Epub 20120509. doi: 10.1007 / s10822-012-9547-0. PubMed PMID: 22569591; PubMedCentral PMCID: PMCPMC3398187.An J, Totrov M, Abagyan R. Pocketome via comprehensive identification andclassification of ligand binding envelopes. Mol Cell Proteomics.2005;4(6):752-61. Epub 20050309. doi: 10.1074 / mcp.M400159-MCP200. PubMed PMID: 15757999.Beyer HM, Gonschorek P, Samodelov SL, Meier M, Weber W, Zurbriggen MD. AQUACloning: A Versatile and Simple Enzyme-Free Cloning Approach. PLoS One. 2015;10(9):e0137652. Epub 20150911. doi: 10.1371 / journal.pone.0137652. PubMed PMID: 26360249; PubMed Central PMCID: PMCPMC4567319.Kim DH, Hwang DC, Kang BH, Lew J, Choi KY. Characterization of Nla protease fromturnip mosaic potyvirus exhibiting a low-temperature optimum catalytic activity. Virology. 1996;221(1):245-9. doi: 10.1006 / viro.1996.0372. PubMed PMID: 8661434.

Claims

CLAIMS:

1. A protease comprising an amino acid sequence of any one of SEQ ID NOS: 1-6 or avariant or fragment thereof.

2. The protease according to any preceding claim, wherein the amino acid sequence ofsaid protease comprises or consists of the sequence SEQ ID NOS: 3, 4, 5 or 6.

3. The protease according to any preceding claim, wherein said protease cleaves thepeptide bond between Q and X of the recognition sequence EAVYHQX.

4. Use of the protease of any preceding claims for catalysing the proteolysis or cleavageof a substrate.

5. A method of catalysing the proteolysis or cleavage of a substrate, said methodcomprising contacting the substrate with a protease according to any preceding claimunder conditions which permit catalysis of the proteolysis or cleavage of the substrate.

6. The use or method of any one of claims 4 or 5, wherein the substrate comprise aprotein or a peptide.

7. The use of method of claim 6, wherein the protein or peptide is a tagged or affinity-tagged protein or peptide.

8. Use of the protease of any preceding claims for catalysing the removal of peptide tagfrom a peptide tagged protein.

9. A method of for catalysing the removal of a peptide tag from a peptide tagged protein,said method comprising contacting the substrate with a protease according to any preceding claim under conditions which permit catalysis of the removal of the peptide tag from the peptide tagged protein.

10. A kit comprising a protease according to any of one of claims 1-3.

Citation Information

Patent Citations

  • Phosphate sensing microbial gene switch

    WO2025166128A1