Polypeptide having collagenase activity and the use thereof

A polypeptide with collagenase activity enhances collagen tripeptide yield to ≥ 3.5 wt%, addressing the inefficiencies of current methods and providing improved health benefits.

WO2026108830A1PCT designated stage Publication Date: 2026-05-28NOVOZYMES AS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NOVOZYMES AS
Filing Date
2025-11-19
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current enzymatic methods yield less than 3 wt% of Glycine-Proline-Hydroxyproline (GPH) collagen tripeptides, failing to meet the increasing demand for a GPH-rich product that provides enhanced health benefits.

Method used

A polypeptide with collagenase activity, having specific sequence identity or structural similarity to SEQ ID NO: 2 or SEQ ID NO: 1, is used to enhance the hydrolysis of collagen sources, producing GPH-rich collagen tripeptides.

Benefits of technology

The polypeptide significantly increases the yield of GPH-rich collagen tripeptides to ≥ 3.5 wt%, offering improved health benefits through enhanced absorption in collagen-related human organs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025135889-FTAPPB-I100001
    Figure PCTCN2025135889-FTAPPB-I100001
  • Figure PCTCN2025135889-FTAPPB-I100002
    Figure PCTCN2025135889-FTAPPB-I100002
  • Figure PCTCN2025135889-FTAPPB-I100003
    Figure PCTCN2025135889-FTAPPB-I100003
Patent Text Reader

Abstract

A composition comprising a polypeptide having collagenase activity, a polypeptide having collagenase activity and method of using said polypeptide for producing GPH-rich collagen tripeptides are provided.
Need to check novelty before this filing date? Find Prior Art

Description

POLYPEPTIDE HAVING COLLAGENASE ACTIVITY AND THE USE THEREOF

[0001] REFERENCE TO SEQUENCE LISTING

[0002] This application contains a Sequence Listing in computer readable form. The computer readable form is incorporated herein by reference.FIELD OF THE INVENTION

[0003] The present invention relates to compositions comprising a polypeptide having collagenase activity, polypeptides having collagenase activity, polynucleotides encoding said polypeptides, nucleic acid constructs and expression vectors comprising said polynucleotides, recombinant host cells comprising said nucleic acid constructs or expressions vectors as well as methods for producing said polypeptides. The present invention further relates to method of using said polypeptide for producing collagen tripeptides as well as compositions comprising said collagen tripeptides.BACKGROUND OF THE INVENTION

[0004] Collagen is composed of (glycine-XY) n. When n=1, (glycine-XY) is called collagen tripeptide (CTP) and is the smallest structural unit of collagen. In this structure, Glycine (G) is an important amino acid that composes the peptide chain of collagen. X in this structure is usually proline or hydroxyproline, and Y is other kinds of amino acids (such as serine, glutamic acid, etc. ) . So when collagen is degraded, different kinds of tripeptides are generated.

[0005] When ingested, collagen tripeptide is rapidly absorbed from digestive tract directly or after enzymatic decomposition. In particular, compared to other CTP, GPH (Glycine-Proline-Hydroxyproline) is taken more into collagen-related human organs, such as skin, bones, cartilage and tendons, and thereby provides more benefits to human health (e.g., better skin health and bone / joint health) than other types of CTP.

[0006] Collagen tripeptide can be produced by enzymatic hydrolysis of collagen source such as skin or cartilage, fish and squid. Enzymes suitable for this process are typically proteases or collagenases. However, the yield of GPH by currently availably enzymes is often less than 3 wt%, failing to meet the increasing needs for an GPH-rich (≥ 3.5 wt%) collagen tripeptide product.

[0007] Therefore, one object of the present invention is to address the above need by providing an enzymatic solution that can improve the yield of GPH by enzymatic hydrolysis of collagen source.SUMMARY OF THE INVENTION

[0008] In one aspect, the present invention relates to a composition comprising a polypeptide having collagenase activity, selected from the group consisting of:

[0009] (i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;

[0010] (ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and

[0011] (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0012] In another aspect, the present invention relates to an isolated polypeptide having collagenase activity, selected from the group consisting of:

[0013] (i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;

[0014] (ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and

[0015] (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0016] In a furthermore aspect, the present invention relates to an isolated polynucleotide encoding the polypeptide of the present invention.

[0017] In a furthermore aspect, the present invention relates to a nucleic acid construct or an expression vector comprising of the present polynucleotide.

[0018] In a furthermore aspect, the present invention relates to a recombinant host cell comprising the nucleic acid construct or expression vector of the present invention.

[0019] In a furthermore aspect, the present invention relates to a method for producing a polypeptide of the present invention, comprising (a) cultivating a recombinant host cell of the present invention under conditions conducive for expression of the polypeptide; and (b) optionally recovering the polypeptide.

[0020] In a furthermore aspect, the present invention relates to methods of producing a polypeptide having at least 80%sequence identity to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1 and having collagenase activity, comprising (a) cultivating a recombinant host cell under conditions conducive for expression of the polypeptide; and optionally, (b) recovering the polypeptide.

[0021] In a fourth aspect, the present invention relates to methods for producing GPH-containing tripeptides, comprising a hydrolysis step by contacting a collagen source with the polypeptide or the composition of the present invention; and optionally recovering the tripeptides.

[0022] In a fifth aspect, the present invention relates to collagen tripeptide compositions which are produced according to the method described in the fourth aspect.

[0023] In a sixth aspect, the present invention relates to use of said collagen tripeptide compositions for preparing healthcare products.BRIEF DESCRIPTION OF DRAWINGS

[0024] Fig. 1 shows sequence alignment across three collagenases: SEQ ID NO: 2, SEQ ID NO: 5 and SEQ ID NO: 7.

[0025] SEQUENCE OVERVIEW

[0026] SEQ ID NO: 1 is full length protein sequence of the polypeptide having collagenase activity obtained from Paenibacillus azoreducens.

[0027] SEQ ID NO: 2 is a mature polypeptide of the polypeptide having collagenase activity shown in SEQ ID NO: 1.

[0028] SEQ ID NO: 3 is a fusion polypeptide in which the signal peptide MQVKSIVNLLLACSLAVA (SEQ ID NO: 4) is operably linked to the polypeptide having collagenase activity.

[0029] SEQ ID NO: 4 is the secretion signal peptide MQVKSIVNLLLACSLAVA.

[0030] SEQ ID NO: 5 is the amino acid sequence from BFH62464.1 collagenase ColA.

[0031] SEQ ID NO: 6 is the amino acid sequence from WP_212977976.1 collagenase.

[0032] SEQ ID NO: 7 is a purified WP_212977976.1 collagenase.

[0033] DEFINITIONS

[0034] As used herein, the singular forms "a" , "an" , and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0035] If not indicated otherwise, all references to percentages in relation to the disclosed compositions relate to “%” or “wt%” is relative to the total weight of the respective composition.

[0036] The dimensions and values disclosed herein are not to be understood as being strictly limited to the exact numerical values recited. Instead, unless otherwise specified, each such dimension is intended to mean both the recited value and a functionally equivalent range surrounding that value. For example, a dimension disclosed as "5 wt%" is intended to mean "about 5 wt%" .

[0037] GPH-rich collagen tripeptides: The collagen tripeptides are normally obtained by hydrolysing a collagen source (e.g., fish skin raw material or fish skin gelatin) and are a mixture of different types of collagen tripeptides which include a particular tripeptide, i.e., GPH (Glycine-Proline-Hydroxyproline) . The content of GPH in the dry matter (often further dehydrated) obtained through such hydrolysis process is often below 3 wt%. When the content of GPH is ≥ 3.5wt%(preferably at least 4wt%, at least 4.5wt%or at least 5wt%) in the dry matter, the resulting collagen tripeptides are called “GPH-rich collagen tripeptides” . GPH-rich collagen tripeptides are more bioactive compared to other types of collagen tripeptides and are therefore believed to provide more benefits to human health.

[0038] cDNA: The term "cDNA" means a DNA molecule that can be prepared by reverse transcription from a mature, spliced, mRNA molecule obtained from a eukaryotic or prokaryotic cell. cDNA lacks intron sequences that may be present in the corresponding genomic DNA. The initial, primary RNA transcript is a precursor to mRNA that is processed through a series of steps, including splicing, before appearing as mature spliced mRNA.

[0039] Coding sequence: The term “coding sequence” means a polynucleotide, which directly specifies the amino acid sequence of a polypeptide. The boundaries of the coding sequence are generally determined by an open reading frame, which begins with a start codon, such as ATG, GTG, or TTG, and ends with a stop codon, such as TAA, TAG, or TGA. The coding sequence may be a genomic DNA, cDNA, synthetic DNA, or a combination thereof.

[0040] Control sequences: The term “control sequences” means nucleic acid sequences involved in regulation of expression of a polynucleotide in a specific organism or in vitro. Each control sequence may be native (i.e., from the same gene) or heterologous (i.e., from a different gene) to the polynucleotide encoding the polypeptide, and native or heterologous to each other. Such control sequences include, but are not limited to leader, polyadenylation, prepropeptide, propeptide, signal peptide, promoter, terminator, enhancer, and transcription or translation initiator and terminator sequences. At a minimum, the control sequences include a promoter, and transcriptional and translational stop signals. The control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the polynucleotide encoding a polypeptide.

[0041] Expression: The term “expression” means any step involved in the production of a polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0042] Expression vector: An "expression vector" refers to a linear or circular DNA construct comprising a DNA sequence encoding a polypeptide, which coding sequence is operably linked to a suitable control sequence capable of effecting expression of the DNA in a suitable host. Such control sequences may include a promoter to effect transcription, an optional operator sequence to control transcription, a sequence encoding suitable ribosome binding sites on the mRNA, enhancers and sequences which control termination of transcription and translation.

[0043] Heterologous: The term "heterologous" means, with respect to a host cell, that a polypeptide or nucleic acid does not naturally occur in the host cell. The term "heterologous" means, with respect to a polypeptide or nucleic acid, that a control sequence, e.g., promoter, of a polypeptide or nucleic acid is not naturally associated with the polypeptide or nucleic acid, i.e., the control sequence is from a gene other than the gene encoding the mature polypeptide.

[0044] Mature polypeptide: The term “mature polypeptide” or “mature polypeptide of interest” means a polypeptide in its mature form following N-terminal and / or C-terminal processing, such as the removal of the signal peptide from the SP-linker by a signal peptidase, resulting in a matured polypeptide of interest comprising the SP-linker at the N-terminal end of the amino acid sequence. Additionally or alternatively, the N-and / or C-terminal processing may include cleavage by a further peptidase when the polypeptide of interest comprises an additional cleavage site, e.g. a cleavage site located between the SP-linker and the amino acid sequence of the mature polypeptide of interest, resulting in a matured polypeptide of interest not comprising the SP-linker at the N-terminal end of the amino acid sequence. Additionally or alternatively, when the polypeptide of interest comprises a pro-peptide, the N-and / or C-terminal processing may include cleavage by a further peptidase when the polypeptide of interest comprises an additional cleavage site, e.g. a cleavage site located between the pro-peptide and the amino acid sequence of the mature polypeptide of interest, resulting in a matured polypeptide of interest lacking the SP-linker and the pro-peptide at the N-terminal end of the amino acid sequence.

[0045] Host Strain or Host Cell: A "host strain" or "host cell" is an organism into which an expression vector, phage, virus, or other DNA construct, including a polynucleotide encoding a polypeptide of the present invention has been introduced. Exemplary host strains are microorganism cells (e.g., bacteria, filamentous fungi, and yeast) capable of expressing the polypeptide of interest and / or fermenting saccharides. The term "host cell" includes protoplasts created from cells.

[0046] Isolated: The term “isolated” means a polypeptide, nucleic acid, cell, or other specified material or component that has been separated from at least one other material or component, including but not limited to, other proteins, nucleic acids, cells, etc. An isolated polypeptide, nucleic acid, cell or other material is thus in a form that does not occur in nature. An isolated polypeptide includes, but is not limited to, a culture broth containing the secreted polypeptide expressed in a host cell.

[0047] Native: The term "native" means a nucleic acid or polypeptide naturally occurring in a host cell.

[0048] Nucleic acid: The term "nucleic acid" encompasses DNA, RNA, heteroduplexes, and synthetic molecules capable of encoding a polypeptide. Nucleic acids may be single stranded or double stranded and may be chemically modified. The terms "nucleic acid" and "polynucleotide" are used interchangeably. Because the genetic code is degenerate, more than one codon may be used to encode a particular amino acid, and the present invention encompass nucleotide sequences that encode a particular amino acid sequence. Unless otherwise indicated, nucleic acid sequences are presented in 5'-to-3' orientation.

[0049] Nucleic acid construct: The term "nucleic acid construct" means a nucleic acid molecule, either single-or double-stranded, which is isolated from a naturally occurring gene or is modified to contain segments of nucleic acids in a manner that would not otherwise exist in nature, or which is synthetic, and which comprises one or more control sequences operably linked to the nucleic acid sequence.

[0050] Operably linked: The term "operably linked" means that specified components are in a relationship (including but not limited to juxtaposition) permitting them to function in an intended manner. For example, a regulatory sequence is operably linked to a coding sequence such that expression of the coding sequence is under control of the regulatory sequence.

[0051] Collagenase: The term “collagenase” or “polypeptide having collagenase activity” can be used interchangeably in the present invention. In one embodiment, the collagenase of the present invention has enzymatic activity on gelatin as shown in Example 6. Gelatin is a collection of peptides and proteins produced by partial hydrolysis of collagen extracted from e.g., the skin, bones, and connective tissues of animals such as domesticated cattle, chicken, pigs, and fish.

[0052] In another embodiment, the collagenase of the present invention cleaves the substrate Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) at a position between L and G, i.e., the collagenase of the present invention has enzymatic activity on FALGPA, which may be determined according to Assay I described in the Examples herein.

[0053] Purified: The term “purified” means a nucleic acid, polypeptide (e.g., protease) or cell that is substantially free from other components as determined by analytical techniques well known in the art (e.g., a purified polypeptide or nucleic acid may form a discrete band in an electrophoretic gel, chromatographic eluate, and / or a media subjected to density gradient centrifugation) . A purified nucleic acid or polypeptide is at least about 50%pure, usually at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8%, about 99.9%, or more, pure (e.g., percent by weight or on a molar basis) . In a related sense, a composition is enriched for a molecule when there is a substantial increase in the concentration of the molecule after application of a purification or enrichment technique. The term "enriched" refers to a compound, polypeptide, cell, nucleic acid, amino acid, or other specified material or component that is present in a composition at a relative or absolute concentration that is higher than a starting composition.

[0054] ln one aspect, the term "purified" as used herein refers to the polypeptide (e.g., collagenase) or cell being essentially free from components (especially insoluble components) from the production organism. In other aspects. the term "purified" refers to the polypeptide being essentially free of insoluble components (especially insoluble components) from the native organism from which it is obtained. In one aspect, the polypeptide is separated from some of the soluble components of the organism and culture medium from which it is recovered. The polypeptide may be purified (i.e., separated) by one or more of the unit operations filtration, precipitation, or chromatography.

[0055] Accordingly, the polypeptide (e.g., collagenase) may be purified such that only minor amounts of other proteins, in particular, other polypeptides, are present. The term "purified" as used herein may refer to removal of other components, particularly other proteins and most particularly other enzymes present in the cell of origin of the polypeptide. The polypeptide may be "substantially pure" , i.e., free from other components from the organism in which it is produced, e.g., a host organism for recombinantly produced polypeptide. In one aspect, the polypeptide is at least 40%pure by weight of the total polypeptide material present in the preparation. In one aspect, the polypeptide is at least 50%, 60%, 70%, 80%or 90%pure by weight of the total polypeptide material present in the preparation. As used herein, a "substantially pure polypeptide" may denote a polypeptide preparation that contains at most 10%, preferably at most 9%, preferably at most 8%, preferably at most 7%, more preferably at most 6%, more preferably at most 5%, more preferably at most 4%, more preferably at most 3%, more preferably at most 2%, more preferably at most 1%, more preferably at most 0.5, more preferably at most 0.1%, more preferably at most 0.05%, more preferably at most 0.01%, even more preferably at most 0.005%, and most preferably at most 0.001%by weight of other polypeptide material with which the polypeptide is natively or recombinantly associated.

[0056] It is, therefore, preferred that the substantially pure polypeptide (e.g., collagenase) is at least 90%pure, preferably at least 91%, more preferably at least 92%pure, more preferably at least 93%pure, more preferably at least 94%pure, more preferably at least 95%pure, more preferably at least 96%pure, more preferably at least 97%pure, more preferably at least 98%pure, more preferably at least 99%pure, more preferably at least 99.5%pure, more preferably at least 99.9%pure, more preferably at least 99.95%, more preferably at least 99.99%pure, even more preferably at least 99.995%pure, and most preferably at least 99.999%pure by weight of the total polypeptide material present in the preparation. The polypeptide of the present invention is preferably in a substantially pure form (i.e., the preparation is essentially free of other polypeptide material with which it is natively or recombinantly associated) . This can be accomplished, for example by preparing the polypeptide by well-known recombinant methods or by classical purification methods.

[0057] Recombinant: The term "recombinant" is used in its conventional meaning to refer to the manipulation, e.g., cutting and rejoining, of nucleic acid sequences to form constellations different from those found in nature. The term recombinant refers to a cell, nucleic acid, polypeptide or vector that has been modified from its native state. Thus, for example, recombinant cells or hosts express genes that are not found within the native (non-recombinant) form of the cell, or express native genes at different levels or under different conditions than found in nature. The term “recombinant” is synonymous with “genetically modified” and “transgenic” .

[0058] Recover: The terms "recover" or “recovery” means the removal of a polypeptide from at least one fermentation broth component selected from the list of a cell, a nucleic acid, or other specified material, e.g., recovery of the polypeptide from the whole fermentation broth, or from the cell-free fermentation broth, by polypeptide crystal harvest, by chromatography, by filtration, e.g., depth filtration (by use of filter aids or packed filter medias, cloth filtration in chamber filters, rotary-drum filtration, drum filtration, rotary vacuum-drum filters, candle filters, horizontal leaf filters or similar, using sheet or pad filtration in framed or modular setups) or membrane filtration (using sheet filtration, module filtration, candle filtration, microfiltration, ultrafiltration in either cross flow, dynamic cross flow or dead end operation) , or by centrifugation (using decanter centrifuges, disc stack centrifuges, hydro cyclones or similar) , or by precipitating the polypeptide and using relevant solid-liquid separation methods to harvest the polypeptide from the broth media by use of classification separation by particle sizes. Recovery encompasses isolation and / or purification of the polypeptide.

[0059] Sequence identity: The relatedness between two amino acid sequences or between two nucleotide sequences is described by the parameter “sequence identity” .

[0060] For purposes of the present invention, the sequence identity between two amino acid sequences is determined as the output of “longest identity” using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) , preferably version 6.6.0 or later. The parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. In order for the Needle program to report the longest identity, the -nobrief option must be specified in the command line. The output of Needle labeled “longest identity” is calculated as follows:

[0061] (Identical Residues x 100)  /  (Length of Alignment –Total Number of Gaps in Alignment) .

[0062] For purposes of the present invention, the sequence identity between two polynucleotide sequences is determined as the output of “longest identity” using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, supra) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, supra) , preferably version 6.6.0 or later. The parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. In order for the Needle program to report the longest identity, the nobrief option must be specified in the command line. The output of Needle labeled “longest identity” is calculated as follows:

[0063] (Identical Deoxyribonucleotides x 100)  /  (Length of Alignment –Total Number of Gaps in Alignment) .

[0064] AlphaFold structure prediction: AlphaFold 2 is a computational method for calculating the three-dimensional structure of a polypeptide from its amino acid sequence (Jumper et al., 2021, Nature 596: 583-589) . Predicted structures for millions of polypeptides deposited in the UniProt database have been deposited in the AlphaFold Protein Structure Database, using the AlphaFold Monomer v2.0 algorithm (Varadi et al., 2021, Nucleic Acids Res. 50 (D1) : D439-D444) . In the AlphaFold Protein Structure Database, the three-dimensional structure of a polypeptide can be obtained by searching for the UniProt accession number of the polypeptide.

[0065] In addition to the many three-dimensional structures that are already publicly available, code is available for reproducing and predicting structures of new polypeptides at source code repositories such as Github. com under deepmind / alphafold / , using notebooks / AlphaFold. ipynb, which uses AlphaFold v2.3.1 or newer. Additionally, it can be found in Github. com under sokrypton / ColabFold using v1.5.2 or newer, using AlphaFold2. ipynb. For technical details, please see Jumper et al. (vide supra) .

[0066] AlphaFold 2 produces a per-residue estimate of its confidence on a scale from 0 to 100. This confidence measure is called pLDDT and corresponds to the model’s predicted score on the lDDT-Cα metric. It is stored in the B-factor fields of the mmCIF and PDB files available for download (although unlike a B-factor, higher pLDDT is better) . Regions with pLDDT score of more than 90 are expected to be modelled to high accuracy. These should be suitable for any application that benefits from high accuracy (e.g., characterization of binding sites) . Regions with a pLDDT score between 70 and 90 are expected to be modelled well, corresponding to a generally good backbone prediction.

[0067] Structural Similarity: For purposes of the present invention, the relatedness between the three-dimensional structure of two polypeptides is described by the parameter “structural similarity” .

[0068] A three-dimensional structure of any polypeptide may be obtained experimentally via, e.g., X-ray crystallography or using in silico methods such as AlphaFold 2 (vide supra) . The structural similarity between three-dimensional structures may then be determined by the TM-score, which is calculated using the following general formula (Zhang &Skolnick, 2004, Proteins 57: 702–710) :

[0069] where LN is the length of the native structure, LT is the length of the aligned residues to the template structure, di is the distance between pair i of aligned residues and d0 is a scale to normalize the match difference. ‘Max’ denotes the maximum value after optimal spatial superposition.

[0070] For the purposes of the present invention, LN is the length of the reference polypeptide.

[0071] A structural alignment of the three-dimensional structures of two polypeptides is necessary before the TM-score can be calculated. This is achieved via algorithms that optimize the structural overlap, and several methods are available, such as CEalign (Shindyalov and Bourne, 1998, Protein Eng., 11: 739-747) , DALI (Holm and Sander, 1995, Trends Biochem. Sci., 20: 478-480) , or TM-align (Zhang and Skolnick, 2005, Nucleic Acids Res. 33 (7) : 2302-2309) .

[0072] For the purposes of the present invention, TM-align is applied. For convenience, TM-score is integrated in the TM-align software, which is available from the author’s website (zhanggroup. org / TM-score / ) . The version of TM-align is preferably updated 2019-08-22 or later, and the TM-score between a reference and a query protein is determined by running this command:

[0073] TMalign <query. pdb> <reference. pdb> -L <length of reference>where <query. pdb> is the name of the PDB file containing coordinates of the query polypeptide, <reference. pdb> is the name of the PDB file containing coordinates of the reference polypeptide. The TM-score is calculated and reported in the output, along with several other parameters from the alignment.

[0074] The maximal TM-score is 1, e.g., 1.0, corresponding to identical three-dimensional structures.DETAILED DESCRIPTION OF THE INVENTION

[0075] The present invention provides a composition comprising a polypeptide having collagenase activity, selected from the group consisting of:

[0076] (i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;

[0077] (ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and

[0078] (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0079] Polypeptides

[0080] In one aspect, the present invention relates to a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1. In furthermore aspect, the invention relates to a polypeptide having collagenase activity and a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by Alphafold. In a furthermore aspect, the present invention relates to a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0081] In one embodiment, the polypeptide comprises of consists of SEQ ID NO: 1, a mature polypeptide of SEQ ID NO: 1. In one embodiment, the polypeptide comprises or consists of SEQ ID NO: 2.

[0082] The mature polypeptide of SEQ ID NO: 1 may correspond to amino acids of 22-1039 of SEQ ID NO: 1, or amino acids of 23-1039 of SEQ ID NO: 1, or amino acids of 141-1039 of SEQ ID NO: 1, or amino acids of 142-1039 of SEQ ID NO: 1, or amino acids of 143-1039 of SEQ ID NO: 1, or amino acids of 144-1039 of SEQ ID NO: 1, or amino acids of 145-1039 of SEQ ID NO: 1, or amino acids of 146-1039 of SEQ ID NO: 1, or amino acids of 147-1039 of SEQ ID NO: 1, or amino acids of 148-1039 of SEQ ID NO: 1, or amino acids of 149-1039 of SEQ ID NO: 1, or amino acids of 150-1037 of SEQ ID NO: 1, or amino acids of 151-828 of SEQ ID NO: 1, or 152-1039 of SEQ ID NO: 1, or amino acids of 153-1039 of SEQ ID NO: 1, or amino acids of 154-1039 of SEQ ID NO: 1, or amino acids of 155-1039 of SEQ ID NO: 1, or amino acids of 156-1039 of SEQ ID NO: 1.

[0083] In one embodiment, the mature polypeptide of SEQ ID NO: 1 may correspond to amino acids 1-888 of SEQ ID NO: 2, or amino acids of 1-887 of SEQ ID NO: 2, or amino acids of 1-886 of SEQ ID NO: 2, or amino acids of 1-885 of SEQ ID NO: 2, or amino acids of 1-884 of SEQ ID NO: 2, or amino acids of 1-883 of SEQ ID NO: 2.

[0084] In one embodiment, the polypeptide has enzymatic activity on gelatin (e.g., fish gelatin or pig gelatin) , which may be determined according to the method shown in Example 6.

[0085] In another embodiment, the polypeptide has enzymatic activity of cleaving the amino acid sequence Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) at a position between L and G. In a specific embodiment, collagenase activity may be determined by Assay I described in Example 5.

[0086] In one embodiment, the polypeptide having collagenase activity is obtained from Paenibacillus sp. In a specific embodiment, the polypeptide is obtained from Paenibacillus azoreducens.

[0087] In another aspect, the polypeptide is derived from SEQ ID NO: 2 or SEQ ID NO: 1 by substitution, deletion or addition of one or several amino acids. In some embodiments, the polypeptide is a variant of SEQ ID NO: 2 or SEQ ID NO: 1 comprising a substitution, deletion, and / or insertion at one or more positions. In a furthermore embodiment, the polypeptide is derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions. In a furthermore embodiment, the number of amino acid substitutions, deletions and / or insertions introduced into the polypeptide of SEQ ID NO: 2 or SEQ ID NO: 1 is up to 20, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20. The amino acid changes may be of a minor nature, that is conservative amino acid substitutions or insertions that do not significantly affect the folding and / or activity of the protein; small deletions, typically of 1-30 amino acids; small amino-or carboxyl-terminal extensions, such as an amino-terminal methionine residue; a small linker peptide of up to 20-25 residues; or a small extension that facilitates purification by changing net charge or another function, such as a poly-histidine tract, an antigenic epitope or a binding module.

[0088] Essential amino acids in a polypeptide can be identified according to procedures known in the art, such as site-directed mutagenesis or alanine-scanning mutagenesis (Cunningham and Wells, 1989, Science 244: 1081-1085) . In the latter technique, single alanine mutations are introduced at every residue in the molecule, and the resultant molecules are tested for protease activity and / or P1 preference to identify amino acid residues that are critical to the activity and / or the specificity of the molecule (see also Hilton et al., 1996, J. Biol. Chem. 271: 4699-4708) . The active site of a polypeptide can also be determined by physical analysis of structure, as determined by such techniques as nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, in conjunction with mutation of putative contact site amino acids. See, for example, de Vos et al., 1992, Science 255: 306-312; Smith et al., 1992, J. Mol. Biol. 224: 899-904; Wlodaver et al., 1992, FEBS Lett. 309: 59-64. The identity of essential amino acids can also be inferred from an alignment with a related polypeptide, and / or be inferred from sequence homology and conserved catalytic machinery with a related polypeptide or within a polypeptide or protein family with polypeptides / proteins descending from a common ancestor, typically having similar three-dimensional structures, functions, and significant sequence similarity. Additionally, or alternatively, protein structure prediction tools can be used for protein structure modelling to identify essential amino acids and / or active sites of polypeptides. See, for example, Jumper et al., 2021, “Highly accurate protein structure prediction with AlphaFold” , Nature 596: 583-589.

[0089] Single or multiple amino acid substitutions, deletions, and / or insertions can be made and tested using known methods of mutagenesis, recombination, and / or shuffling, followed by a relevant screening procedure, such as those disclosed by Reidhaar-Olson and Sauer, 1988, Science 241: 53-57; Bowie and Sauer, 1989, Proc. Natl. Acad. Sci. USA 86: 2152-2156; WO 95 / 17413; or WO 95 / 22625. Other methods that can be used include error-prone PCR, phage display (e.g., Lowman et al., 1991, Biochemistry 30: 10832-10837; US 5, 223, 409; WO 92 / 06204) , and region-directed mutagenesis (Derbyshire et al., 1986, Gene 46: 145; Ner et al., 1988, DNA 7: 127) .

[0090] Mutagenesis / shuffling methods can be combined with high-throughput, automated screening methods to detect activity of cloned, mutagenized polypeptides expressed by host cells (Ness et al., 1999, Nature Biotechnology 17: 893-896) . Mutagenized DNA molecules that encode active polypeptides can be recovered from the host cells and rapidly sequenced using standard methods in the art. These methods allow the rapid determination of the importance of individual amino acid residues in a polypeptide.

[0091] In one aspect, the polypeptide is isolated.

[0092] In another aspect, the polypeptide is purified.

[0093] Compositions comprising said polypeptide

[0094] The present invention further relates to a composition comprising the polypeptide of the present invention as described in the above section.

[0095] In one embodiment, the polypeptide is present in an amount of from about 0.1 mg / g to about 200 mg / g enzyme protein; preferably in an amount of from about 1 mg / g to about 100 mg / g, from about 2 mg / g to about 50 mg / g, from about 3 mg / g to about 15 mg / g, from about 5 mg / g to about 18 mg / g, or from about 5 mg / g to about 10 mg / g or more, based on the total amount of the composition.

[0096] In one embodiment, the composition is a granulate composition.

[0097] In one embodiment, the composition is a powder composition.

[0098] In one embodiment, the polypeptide is present in an amount of from about 0.01 to about 99 wt%; preferably in an amount of from about 1 to about 90 wt%, or from about 2 to about 85 wt%, or from about 3 to about 80 wt%, or from about 4 to about 70 wt%, or from about 5 to about 60 wt%, based on the total amount of the composition.

[0099] In one embodiment, the composition is a liquid composition, e.g., an aqueous composition.

[0100] In one embodiment, the composition is a liquid composition comprising an aqueous buffer; preferably wherein the aqueous buffer comprises 4- (2-hydroxyethyl) -1-piperazineethanesulfonic acid (HEPES) , tris (hydroxymethyl) aminomethane (TRIS) , enzyme stabilizers, phosphate, or bicarbonate; most preferably wherein the aqueous buffer comprises monopropylene glycol, phosphate or bicarbonate.

[0101] In one embodiment, the composition is a liquid composition having a pH value of about 5 to about 9; e.g., of about 6.5 to about 8.5, of about 7 to about 8, or of about 7 to about 7.5.

[0102] The composition may further comprise one or more enzyme stabilizers. Examples of enzyme stabilizer may be selected from propylene glycol or glycerol, sugar or sugar alcohol, lactic acid, reversible protease inhibitor, boric acid, or a boric acid derivative, e.g., an aromatic borate ester, or a phenyl boronic acid derivative such as 4-formylphenyl boronic acid.

[0103] In some embodiments, filler (s) or carrier material (s) are included to increase the volume of the liquid composition. Suitable filler or carrier materials include, but are not limited to, various salts of sulfate, carbonate and silicate as well as talc, clay and the like. Suitable filler or carrier materials for liquid compositions include, but are not limited to, water or low molecular weight primary and secondary alcohols including polyols and diols. Examples of such alcohols include, but are not limited to, methanol, ethanol, propanol and isopropanol. In some embodiments, the compositions contain from about 5%to about 90%of such materials.

[0104] In one embodiment, the liquid composition comprises 20-80%w / w of polyol. In one embodiment, the liquid composition comprises 0.001-2%w / w preservative.

[0105] In another embodiment, the invention relates to liquid compositions comprising:

[0106] (a) 0.001-25%w / w of a polypeptide of the present invention (e.g., SEQ ID NO: 2) ;

[0107] (b) 20-80%w / w of polyol (e.g., monopropylene glycol) ;

[0108] (c) optionally 0.001-2%w / w preservative; and

[0109] (d) water.

[0110] In another embodiment, the invention relates to liquid compositions comprising:

[0111] (a) 0.001-25%w / w of a polypeptide of the present invention (e.g., SEQ ID NO: 2) ;

[0112] (b) 0.001-2%w / w preservative;

[0113] (c) optionally 20-80%w / w of polyol (e.g., monopropylene glycol) ; and

[0114] (d) water.

[0115] In another embodiment, the liquid composition comprises one or more formulating agents, such as a formulating agent selected from the group consisting of polyol, sodium chloride, sodium benzoate, potassium sorbate, sodium sulfate, potassium sulfate, magnesium sulfate, sodium thiosulfate, calcium carbonate, sodium citrate, dextrin, glucose, sucrose, sorbitol, lactose, starch, PVA, acetate and phosphate, preferably selected from the group consisting of sodium sulfate, dextrin, cellulose, sodium thiosulfate, kaolin and calcium carbonate. In one embodiment, the polyols is selected from the group consisting of glycerol, sorbitol, propylene glycol (MPG) , ethylene glycol, diethylene glycol, triethylene glycol, 1, 2-propylene glycol or 1, 3-propylene glycol, dipropylene glycol, polyethylene glycol (PEG) having an average molecular weight below about 600 and polypropylene glycol (PPG) having an average molecular weight below about 600, more preferably selected from the group consisting of glycerol, sorbitol and propylene glycol (MPG) or any combination thereof.

[0116] In one embodiment, the liquid composition comprises glucose in an amount of from about 0.1 g / L to about 10 g / L, e.g., about 0.1 g / L, about 0.2 g / L, about 0.3 g / L, about 0.4 g / L, about 0.5 g / L, about 0.6 g / L, about 0.7 g / L, about 0.8 g / L, about 0.9 g / L, about 1 g / L, about 2 g / L, about 3 g / L, about 4 g / L, about 5 g / L, about 6 g / L, about 7 g / L, about 8 g / L, about 9 g / L, or about 10 g / L. In a preferred embodiment, the liquid composition comprises glucose in an amount of from about 0.5 g / L to about 5 g / L, most preferably in an amount of about 1 g / L.

[0117] In another embodiment, the liquid composition comprises 20-80%polyol (i.e., total amount of polyol) , e.g., 25-75%polyol, 30-70%polyol, 35-65%polyol, or 40-60%polyol. In one embodiment, the liquid formulation comprises 20-80%polyol, e.g., 25-75%polyol, 30-70%polyol, 35-65%polyol, or 40-60%polyol, wherein the polyol is selected from the group consisting of glycerol, sorbitol, propylene glycol (MPG) , ethylene glycol, diethylene glycol, triethylene glycol, 1, 2-propylene glycol or 1, 3-propylene glycol, dipropylene glycol, polyethylene glycol (PEG) having an average molecular weight below about 600 and polypropylene glycol (PPG) having an average molecular weight below about 600. In one embodiment, the liquid formulation comprises 20-80%polyol (i.e., total amount of polyol) , e.g., 25-75%polyol, 30-70%polyol, 35-65%polyol, or 40-60%polyol, wherein the polyol is selected from the group consisting of glycerol, sorbitol and propylene glycol (MPG) .

[0118] In another embodiment, the preservative may be included in the composition. Examples of such preservative may be selected from the group consisting of sodium sorbate, potassium sorbate, sodium benzoate and potassium benzoate or any combination thereof. In one embodiment, the liquid compositions comprise 0.02-1.5%w / w preservative, e.g., 0.05-1%w / w preservative or 0.1-0.5%w / w preservative. In one embodiment, the liquid formulation composition 0.001-2%w / w preservative (i.e., total amount of preservative) , e.g., 0.02-1.5%w / w preservative, 0.05-1%w / w preservative, or 0.1-0.5%w / w preservative, wherein the preservative is selected from the group consisting of sodium sorbate, potassium sorbate, sodium benzoate and potassium benzoate or any combination thereof.

[0119] In one aspect, the composition further comprises one or more additional enzymes, e.g., hydrolase, isomerase, ligase, lyase, oxidoreductase, and transferase. The one or more additional enzymes are preferably selected from the group consisting of acetylxylan esterase, acylglycerol lipase, amylase, alpha-amylase, beta-amylase, arabinofuranosidase, cellobiohydrolases, cellulase, DNase, feruloyl esterase, galactanase, alpha-galactosidase, beta-galactosidase, beta-glucanase, beta-glucosidase, lysophospholipase, lysozyme, alpha-mannosidase, beta-mannosidase (mannanase) , phytase, phospholipase A1, phospholipase A2, phospholipase D, pullulanase, pectin esterase, triacylglycerol lipase, xylanase, beta-xylosidase or any combination thereof. In one embodiment, the composition further comprises a DNase and / or protease. In a preferred embodiment, the composition comprises at least two different collagenases.

[0120] Polynucleotides

[0121] The present invention further relates to polynucleotides encoding the polypeptide of the present invention. The polynucleotide may be a genomic DNA, a cDNA, a synthetic DNA, a synthetic RNA, an mRNA, or a combination thereof. In one embodiment, the polynucleotide encodes the polypeptide of SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 2.

[0122] The polynucleotide may also be mutated by introduction of nucleotide substitutions that do not result in a change in the amino acid sequence of the polypeptide, but which correspond to the codon usage of the host organism intended for production of the enzyme, or by introduction of nucleotide substitutions that may give rise to a different amino acid sequence. For a general description of nucleotide substitution, see, e.g., Ford et al., 1991, Protein Expression and Purification 2: 95-107.

[0123] In an aspect, the polynucleotide is isolated.

[0124] In another aspect, the polynucleotide is purified.

[0125] Nucleic Acid Constructs

[0126] The present invention further relates to nucleic acid constructs comprising a polynucleotide of the present invention. The polynucleotide may be operably linked to one or more control sequences that direct the expression of the coding sequence in a suitable host cell under conditions compatible with the control sequences.

[0127] The polynucleotide may be manipulated in a variety of ways to provide for expression of the polypeptide. Manipulation of the polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides utilizing recombinant DNA methods are well known in the art.

[0128] In one embodiment, the present invention relates to a nucleic acid construct comprising or consisting of:

[0129] a) optionally a first polynucleotide encoding a signal peptide (SP) ,

[0130] b) at least a second polynucleotide located downstream of the first polynucleotide and encoding a polypeptide of interest or of the present invention;

[0131] wherein the first polynucleotide and the second polynucleotide are operably linked in translational fusion.

[0132] Promoters

[0133] The control sequence may be a promoter, a polynucleotide that is recognized by a host cell for expression of a polynucleotide encoding a polypeptide of the present invention. The promoter contains transcriptional control sequences that mediate the expression of the polypeptide. The promoter may be any polynucleotide that shows transcriptional activity in the host cell including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.

[0134] Examples of suitable promoters for directing transcription of the polynucleotide of the present invention in a bacterial host cell are described in Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Lab., NY, Davis et al., 2012, supra, and Song et al., 2016, PLOS One 11 (7) : e0158447.

[0135] Examples of suitable promoters for directing transcription of the polynucleotide of the present invention in a filamentous fungal host cell are promoters obtained from Aspergillus, Fusarium, Rhizomucor and Trichoderma cells, such as the promoters described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology.

[0136] For expression in a yeast host, examples of useful promoters are described by Smolke et al., 2018, “Synthetic Biology: Parts, Devices and Applications” (Chapter 6: Constitutive and Regulated Promoters in Yeast: How to Design and Make Use of Promoters in S. cerevisiae) , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology.

[0137] Terminators

[0138] The control sequence may also be a transcription terminator, which is recognized by a host cell to terminate transcription. The terminator is operably linked to the 3’ -terminus of the polynucleotide encoding the polypeptide. Any terminator that is functional in the host cell may be used in the present invention.

[0139] Preferred terminators for bacterial host cells may be obtained from the genes for Bacillus clausii alkaline protease (aprH) , Bacillus licheniformis alpha-amylase (amyL) , and Escherichia coli ribosomal RNA (rrnB) .

[0140] Preferred terminators for filamentous fungal host cells may be obtained from Aspergillus or Trichoderma species, such as obtained from the genes for Aspergillus niger glucoamylase, Trichoderma reesei beta-glucosidase, Trichoderma reesei cellobiohydrolase I, and Trichoderma reesei endoglucanase I, such as the terminators described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology.

[0141] Preferred terminators for yeast host cells may be obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1) , and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described by Romanos et al., 1992, Yeast 8: 423-488.

[0142] mRNA Stabilizers

[0143] The control sequence may also be an mRNA stabilizer region downstream of a promoter and upstream of the coding sequence of a gene which increases expression of the gene.

[0144] Examples of suitable mRNA stabilizer regions are obtained from a Bacillus thuringiensis cryIIIA gene (WO 94 / 25612) and a Bacillus subtilis SP82 gene (Hue et al., 1995, J. Bacteriol. 177: 3465-3471) .

[0145] Examples of mRNA stabilizer regions for fungal cells are described in Geisberg et al., 2014, Cell 156 (4) : 812-824, and in Morozov et al., 2006, Eukaryotic Cell 5 (11) : 1838-1846.

[0146] Leader Sequences

[0147] The control sequence may also be a leader, a non-translated region of an mRNA that is important for translation by the host cell. The leader is operably linked to the 5’ -terminus of the polynucleotide encoding the polypeptide. Any leader that is functional in the host cell may be used.

[0148] Suitable leaders for bacterial host cells are described by Hambraeus et al., 2000, Microbiology 146 (12) : 3051-3059, and by Kaberdin and  2006, FEMS Microbiol. Rev. 30 (6) : 967-979.

[0149] Preferred leaders for filamentous fungal host cells may be obtained from the genes for Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase.

[0150] Suitable leaders for yeast host cells may be obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1) , Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP) .

[0151] Polyadenylation Sequences

[0152] The control sequence may also be a polyadenylation sequence, a sequence operably linked to the 3’ -terminus of the polynucleotide which, when transcribed, is recognized by the host cell as a signal to add polyadenosine residues to transcribed mRNA. Any polyadenylation sequence that is functional in the host cell may be used.

[0153] Preferred polyadenylation sequences for filamentous fungal host cells are obtained from the genes for Aspergillus nidulans anthranilate synthase, Aspergillus niger glucoamylase, Aspergillus niger alpha-glucosidase, Aspergillus oryzae TAKA amylase, and Fusarium oxysporum trypsin-like protease.

[0154] Useful polyadenylation sequences for yeast host cells are described by Guo and Sherman, 1995, Mol. Cellular Biol. 15: 5983-5990.

[0155] Signal Peptides

[0156] The control sequence may also be a signal peptide coding region that encodes a signal peptide linked to the N-terminus of a polypeptide and directs the polypeptide into the cell’s secretory pathway. The 5’ -end of the coding sequence of the polynucleotide may inherently contain a signal peptide coding sequence naturally linked in translation reading frame with the segment of the coding sequence that encodes the polypeptide. Alternatively, the 5’ -end of the coding sequence may contain a signal peptide coding sequence that is heterologous to the coding sequence. A heterologous signal peptide coding sequence may be required where the coding sequence does not naturally contain a signal peptide coding sequence. Alternatively, a heterologous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance secretion of the polypeptide. Any signal peptide coding sequence that directs the expressed polypeptide into the secretory pathway of a host cell may be used.

[0157] Effective signal peptide coding sequences for bacterial host cells are the signal peptide coding sequences obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus alpha-amylase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM) , and Bacillus subtilis prsA. Further signal peptides are described by Freudl, 2018, Microbial Cell Factories 17: 52.

[0158] Effective signal peptide coding sequences for filamentous fungal host cells are the signal peptide coding sequences obtained from the genes for Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Aspergillus oryzae TAKA amylase, Humicola insolens cellulase, Humicola insolens endoglucanase V, Humicola lanuginosa lipase, and Rhizomucor miehei aspartic proteinase, such as the signal peptide described by Xu et al., 2018, Biotechnology Letters 40: 949-955.

[0159] Useful signal peptides for yeast host cells are obtained from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase. Other useful signal peptide coding sequences are described by Romanos et al., 1992, supra.

[0160] In one embodiment, the signal peptide corresponds to the amino acids 1-21 of SEQ ID NO: 1.

[0161] A native signal peptide may be replaced with a different signal peptide to facilitate e.g., protein expression in different host cells.

[0162] Propeptides

[0163] The control sequence may also be a propeptide coding sequence that encodes a propeptide positioned at the N-terminus of a polypeptide. The resultant polypeptide is known as a proenzyme or propolypeptide (or a zymogen in some cases) . A propolypeptide is generally inactive and can be converted to an active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. The propeptide coding sequence may be obtained from the genes for Bacillus subtilis alkaline protease (aprE) , Bacillus subtilis neutral protease (nprT) , Myceliophthora thermophila laccase (WO 95 / 33836) , Rhizomucor miehei aspartic proteinase, and Saccharomyces cerevisiae alpha-factor.

[0164] Where both signal peptide and propeptide sequences are present, the propeptide sequence is positioned next to the N-terminus of a polypeptide and the signal peptide sequence is positioned next to the N-terminus of the propeptide sequence. Additionally, or alternatively, when both signal peptide and propeptide sequences are present, the polypeptide may comprise only a part of the signal peptide sequence and / or only a part of the propeptide sequence. Alternatively, the final or isolated polypeptide may comprise a mixture of mature polypeptides and polypeptides which comprise, either partly or in full length, a propeptide sequence and / or a signal peptide sequence.

[0165] Regulatory Sequences

[0166] It may also be desirable to add regulatory sequences that regulate expression of the polypeptide relative to the growth of the host cell. Examples of regulatory sequences are those that cause expression of the gene to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. Regulatory sequences in prokaryotic systems include the lac, tac, and trp operator systems. In yeast, the ADH2 system or GAL1 system may be used. In filamentous fungi, the Aspergillus niger glucoamylase promoter, Aspergillus oryzae TAKA alpha-amylase promoter, and Aspergillus oryzae glucoamylase promoter, Trichoderma reesei cellobiohydrolase I promoter, and Trichoderma reesei cellobiohydrolase II promoter may be used. Other examples of regulatory sequences are those that allow for gene amplification. In fungal systems, these regulatory sequences include the dihydrofolate reductase gene that is amplified in the presence of methotrexate, and the metallothionein genes that are amplified with heavy metals.

[0167] Transcription Factors

[0168] The control sequence may also be a transcription factor, a polynucleotide encoding a polynucleotide-specific DNA-binding polypeptide that controls the rate of the transcription of genetic information from DNA to mRNA by binding to a specific polynucleotide sequence. The transcription factor may function alone and / or together with one or more other polypeptides or transcription factors in a complex by promoting or blocking the recruitment of RNA polymerase. Transcription factors are characterized by comprising at least one DNA-binding domain which often attaches to a specific DNA sequence adjacent to the genetic elements which are regulated by the transcription factor. The transcription factor may regulate the expression of a protein of interest either directly, i.e., by activating the transcription of the gene encoding the protein of interest by binding to its promoter, or indirectly, i.e., by activating the transcription of a further transcription factor which regulates the transcription of the gene encoding the protein of interest, such as by binding to the promoter of the further transcription factor. Suitable transcription factors for fungal host cells are described in WO 2017 / 144177. Suitable transcription factors for prokaryotic host cells are described in Seshasayee et al., 2011, Subcellular Biochemistry 52: 7-23, as well in Balleza et al., 2009, FEMS Microbiol. Rev. 33 (1) : 133-151.

[0169] Expression Vectors

[0170] The present invention further relates to recombinant expression vectors comprising a polynucleotide encoding the polypeptide of the present invention (e.g., SEQ ID NO: 3) of the present invention, a promoter, and transcriptional and translational stop signals. The various nucleotide and control sequences may be joined together to produce a recombinant expression vector that may include one or more convenient restriction sites to allow for insertion or substitution of the polynucleotide encoding the polypeptide at such sites. Alternatively, the polynucleotide may be expressed by inserting the polynucleotide or a nucleic acid construct comprising the polynucleotide into an appropriate vector for expression. In creating the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.

[0171] The recombinant expression vector may be any vector (e.g., a plasmid or virus) that can be conveniently subjected to recombinant DNA procedures and can bring about expression of the polynucleotide. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may be a linear or closed circular plasmid.

[0172] The vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector may contain any means for assuring self-replication. Alternatively, the vector may be one that, when introduced into the host cell, is integrated into the genome and replicated together with the chromosome (s) into which it has been integrated. Furthermore, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell, or a transposon, may be used.

[0173] The vector preferably contains one or more selectable markers that permit easy selection of transformed, transfected, transduced, or the like cells. A selectable marker is a gene the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like.

[0174] The vector preferably contains at least one element that permits integration of the vector into the host cell's genome or autonomous replication of the vector in the cell independent of the genome.

[0175] For integration into the host cell genome, the vector may rely on the polynucleotide’s sequence encoding the polypeptide or any other element of the vector for integration into the genome by homologous recombination, such as homology-directed repair (HDR) , or non-homologous recombination, such as non-homologous end-joining (NHEJ) .

[0176] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. The origin of replication may be any plasmid replicator mediating autonomous replication that functions in a cell. The term “origin of replication” or “plasmid replicator” means a polynucleotide that enables a plasmid or vector to replicate in vivo.

[0177] More than one copy of a polynucleotide of the present invention may be inserted into a host cell to increase production of a polypeptide. For example, 2 or 3 or 4 or 5 or more copies are inserted into a host cell. An increase in the copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the polynucleotide where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the polynucleotide, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0178] Recombinant Host Cells

[0179] The present invention further relates to recombinant host cells comprising a polynucleotide of the present invention. In one embodiment, the polynucleotide of the present invention is operably linked to one or more control sequences that direct the production of a polypeptide of the present invention.

[0180] Thus, in one embodiment, the recombinant host cell comprises a polynucleotide encoding a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, or 100%, to SEQ ID NO: 2 or SEQ ID NO: 1, or the mature polypeptide of SEQ ID NO: 1. In a furthermore embodiment, the recombinant host cell comprises a polynucleotide encoding a polypeptide having collagenase activity and a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by Alphafold. In a furthermore embodiment, the recombinant host cell comprises a polynucleotide encoding a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions. In a preferred embodiment, the recombinant host cell comprises a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 2. In another embodiment, the recombinant host cell comprises a polynucleotide encoding a polypeptide comprising or consisting of SEQ ID NO: 3.

[0181] A construct or vector comprising a polynucleotide is introduced into a host cell so that the construct or vector is maintained as a chromosomal integrant or as a self-replicating extra-chromosomal vector as described earlier. The choice of a host cell will to a large extent depend upon the gene encoding the polypeptide and its source. The polypeptide can be native or heterologous to the recombinant host cell. Also, at least one of the one or more control sequences can be heterologous to the polynucleotide encoding the polypeptide. The recombinant host cell may comprise a single copy, or at least two copies, e.g., three, four, five, or more copies of the polynucleotide of the present invention.

[0182] The host cell may be any microbial cell useful in the recombinant production of a polypeptide of the present invention, e.g., a prokaryotic cell or a fungal cell.

[0183] The prokaryotic host cell may be any Gram-positive or Gram-negative bacterium. Gram-positive bacteria include, but are not limited to, Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, and Streptomyces. Gram-negative bacteria include, but are not limited to, Campylobacter, E. coli, Flavobacterium, Fusobacterium, Helicobacter, Ilyobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma.

[0184] The prokaryotic host cell may be any Bacillus cell including, but not limited to, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, and Bacillus thuringiensis cells. In an embodiment, the Bacillus cell is a Bacillus amyloliquefaciens, Bacillus licheniformis, or Bacillus subtilis cell.

[0185] In a preferred embodiment, the recombinant host cell is a Bacillus licheniformis cell.

[0186] In a preferred embodiment, the recombinant host cell is a Bacillus subtilis cell.

[0187] For purposes of this invention, Bacillus classes / genera / species shall be defined as described in Patel and Gupta, 2020, Int. J. Syst. Evol. Microbiol. 70: 406-438.

[0188] The bacterial host cell may also be any Streptococcus cell including, but not limited to, Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. zooepidemicus cells.

[0189] The bacterial host cell may also be any Streptomyces cell including, but not limited to, Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans cells.

[0190] Methods for introducing DNA into prokaryotic host cells are well-known in the art, and any suitable method can be used including but not limited to protoplast transformation, competent cell transformation, electroporation, conjugation, transduction, with DNA introduced as linearized or as circular polynucleotide. Persons skilled in the art will be readily able to identify a suitable method for introducing DNA into a given prokaryotic cell depending, e.g., on the genus. Methods for introducing DNA into prokaryotic host cells are for example described in Heinze et al., 2018, BMC Microbiology 18: 56, Burke et al., 2001, Proc. Natl. Acad. Sci. USA 98: 6289-6294, Choi et al., 2006, J. Microbiol. Methods 64: 391-397, and Donald et al., 2013, J. Bacteriol. 195 (11) : 2612-2620.

[0191] The host cell may be a fungal cell. “Fungi” as used herein includes the phyla Ascomycota, Basidiomycota, Chytridiomycota, and Zygomycota as well as the Oomycota and all mitosporic fungi (as defined by Hawksworth et al., In, Ainsworth and Bisby’s Dictionary of The Fungi, 8th edition, 1995, CAB International, University Press, Cambridge, UK) .

[0192] Fungal cells may be transformed by a process involving protoplast-mediated transformation, Agrobacterium-mediated transformation, electroporation, biolistic method and shock-wave-mediated transformation as reviewed by Li et al., 2017, Microbial Cell Factories 16: 168 and procedures described in EP 238023, Yelton et al., 1984, Proc. Natl. Acad. Sci. USA 81: 1470-1474, Christensen et al., 1988, Bio / Technology 6: 1419-1422, and Lubertozzi and Keasling, 2009, Biotechn. Advances 27: 53-75. However, any method known in the art for introducing DNA into a fungal host cell can be used, and the DNA can be introduced as linearized or as circular polynucleotide.

[0193] The fungal host cell may be a yeast cell. “Yeast” as used herein includes ascosporogenous yeast (Endomycetales) , basidiosporogenous yeast, and yeast belonging to the Fungi Imperfecti (Blastomycetes) . For purposes of this invention, yeast shall be defined as described in Biology and Activities of Yeast (Skinner, Passmore, and Davenport, editors, Soc. App. Bacteriol. Symposium Series No. 9, 1980) .

[0194] The yeast host cell may be a Candida, Hansenula, Kluyveromyces, Pichia, Saccharomyces, Schizosaccharomyces, or Yarrowia cell, such as a Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, or Yarrowia lipolytica cell. In a preferred embodiment, the yeast host cell is a Pichia or Komagataella cell, e.g., a Pichia pastoris cell (Komagataella phaffii) .

[0195] In a preferred embodiment, the recombinant host cell is a Pichia pastoris (Komagataella phaffii) cell.

[0196] The fungal host cell may be a filamentous fungal cell. “Filamentous fungi” include all filamentous forms of the subdivision Eumycota and Oomycota (as defined by Hawksworth et al., 1995, supra) . The filamentous fungi are generally characterized by a mycelial wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth is by hyphal elongation and carbon catabolism is obligately aerobic. In contrast, vegetative growth by yeasts such as Saccharomyces cerevisiae is by budding of a unicellular thallus and carbon catabolism may be fermentative.

[0197] The filamentous fungal host cell may be an Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Chrysosporium, Coprinus, Coriolus, Cryptococcus, Filibasidium, Fusarium, Humicola, Magnaporthe, Mucor, Myceliophthora, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Phanerochaete, Phlebia, Piromyces, Pleurotus, Schizophyllum, Talaromyces, Thermoascus, Thielavia, Tolypocladium, Trametes, or Trichoderma cell. In a preferred embodiment, the filamentous fungal host cell is an Aspergillus, Trichoderma or Fusarium cell. In a further preferred embodiment, the filamentous fungal host cell is an Aspergillus niger, Aspergillus oryzae, Trichoderma reesei, or Fusarium venenatum cell.

[0198] For example, the filamentous fungal host cell may be an Aspergillus awamori, Aspergillus foetidus, Aspergillus fumigatus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, Aspergillus oryzae, Bjerkandera adusta, Ceriporiopsis aneirina, Ceriporiopsis caregiea, Ceriporiopsis gilvescens, Ceriporiopsis pannocinta, Ceriporiopsis rivulosa, Ceriporiopsis subrufa, Ceriporiopsis subvermispora, Chrysosporium inops, Chrysosporium keratinophilum, Chrysosporium lucknowense, Chrysosporium merdarium, Chrysosporium pannicola, Chrysosporium queenslandicum, Chrysosporium tropicum, Chrysosporium zonatum, Coprinus cinereus, Coriolus hirsutus, Fusarium bactridioides, Fusarium cerealis, Fusarium crookwellense, Fusarium culmorum, Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, Fusarium venenatum, Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora thermophila, Neurospora crassa, Penicillium purpurogenum, Phanerochaete chrysosporium, Phlebia radiata, Pleurotus eryngii, Talaromyces emersonii, Thielavia terrestris, Trametes villosa, Trametes versicolor, Trichoderma harzianum, Trichoderma koningii, Trichoderma longibrachiatum, Trichoderma reesei, or Trichoderma viride cell.

[0199] In a preferred embodiment, the recombinant host cell is an Aspergillus niger cell.

[0200] In a preferred embodiment, the recombinant host cell is an Aspergillus oryzae cell.

[0201] In a preferred embodiment, the recombinant host cell is a Pichia pastoris cell.

[0202] In a preferred embodiment, the recombinant host cell is a Trichoderma reesei cell.

[0203] In an aspect, the recombinant host cell is isolated.

[0204] In another aspect, the recombinant host cell is purified.

[0205] Methods of Production

[0206] The present invention further relates to a method for producing a polypeptide of the present invention, comprising (a) cultivating a recombinant host cell of the present invention under conditions conducive for expression of the polypeptide; and (b) optionally recovering the polypeptide.

[0207] In one embodiment, the present invention relates to a method of producing a polypeptide of the present invention, comprising (a) cultivating a host cell, which in its wild-type form produces a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, or 100%, to SEQ ID NO: 2 or SEQ ID NO: 1, or the mature polypeptide of SEQ ID NO: 1 under conditions conducive for production of the polypeptide; and optionally (b) recovering the polypeptide. In a preferred embodiment, the polypeptide comprises or consists of SEQ ID NO: 2.

[0208] In a furthermore embodiment, the present invention relates to a method of producing a polypeptide of the present invention, comprising (a) cultivating a recombinant host cell of the present invention under conditions conducive for production of a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1; and optionally, (b) recovering the polypeptide. In a preferred embodiment, the polypeptide comprises or consists of SEQ ID NO: 2.

[0209] In one aspect, the recombinant host cell is a prokaryotic host cell, preferably a bacterial cell. In one embodiment, the recombinant host cell is a Bacillus cell, preferably a Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, or Bacillus thuringiensis cells. In a preferred embodiment, the recombinant host cell is a Bacillus amyloliquefaciens, Bacillus licheniformis, or Bacillus subtilis cell.

[0210] In a preferred embodiment, the recombinant host cell is a Bacillus licheniformis cell.

[0211] In a preferred embodiment, the recombinant host cell is a Bacillus subtilis cell.

[0212] In one aspect, the recombinant host cell is a filamentous fungal host cell. In one embodiment, the recombinant host cell is selected from the group consisting of Aspergillus awamori, Aspergillus foetidus, Aspergillus fumigatus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, Aspergillus oryzae, Bjerkandera adusta, Ceriporiopsis aneirina, Ceriporiopsis caregiea, Ceriporiopsis gilvescens, Ceriporiopsis pannocinta, Ceriporiopsis rivulosa, Ceriporiopsis subrufa, Ceriporiopsis subvermispora, Chrysosporium inops, Chrysosporium keratinophilum, Chrysosporium lucknowense, Chrysosporium merdarium, Chrysosporium pannicola, Chrysosporium queenslandicum, Chrysosporium tropicum, Chrysosporium zonatum, Coprinus cinereus, Coriolus hirsutus, Fusarium bactridioides, Fusarium cerealis, Fusarium crookwellense, Fusarium culmorum, Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, Fusarium venenatum, Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora thermophila, Neurospora crassa, Penicillium purpurogenum, Phanerochaete chrysosporium, Phlebia radiata, Pleurotus eryngii, Talaromyces emersonii, Thielavia terrestris, Trametes villosa, Trametes versicolor, Trichoderma harzianum, Trichoderma koningii, Trichoderma longibrachiatum, Trichoderma reesei, and Trichoderma viride cell. In a preferred embodiment, the recombinant host cell is an Aspergillus niger, Aspergillus oryzae, or Trichoderma reesei cell.

[0213] In a preferred embodiment, the recombinant host cell is an Aspergillus niger cell.

[0214] In a preferred embodiment, the recombinant host cell is an Aspergillus oryzae cell.

[0215] In a preferred embodiment, the recombinant host cell is a Trichoderma reesei cell.

[0216] In one aspect, the recombinant host cell is a yeast host cell. In one embodiment, the recombinant host cell is selected from the group consisting of Candida, Hansenula, Kluyveromyces, Pichia, Saccharomyces, Schizosaccharomyces, or Yarrowia cell, such as a Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, or Yarrowia lipolytica cell. In a preferred embodiment, the yeast host cell is a Pichia or Komagataella cell, e.g., a Pichia pastoris cell (Komagataella phaffii) .

[0217] In a preferred embodiment, the recombinant host cell is a Pichia pastoris (Komagataella phaffii) cell.

[0218] The host cell or the recombinant host cell is cultivated in a nutrient medium suitable for production of the polypeptide using methods known in the art. For example, the cell may be cultivated by shake flask cultivation, or small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state, and / or microcarrier-based fermentations) in laboratory or industrial fermentors in a suitable medium and under conditions allowing the polypeptide to be expressed and / or isolated. Suitable media are available from commercial suppliers or may be prepared according to published compositions (e.g., in catalogues of the American Type Culture Collection) . If the polypeptide is secreted into the nutrient medium, the polypeptide can be recovered directly from the medium. If the polypeptide is not secreted, it can be recovered from cell lysates.

[0219] The polypeptide may be detected using methods known in the art that are specific for the polypeptide, including, but not limited to, the use of specific antibodies, formation of an enzyme product, disappearance of an enzyme substrate, or an assay determining the relative or specific activity of the polypeptide.

[0220] The polypeptide may be recovered from the medium using methods known in the art, including, but not limited to, collection, centrifugation, filtration, extraction, spray-drying, freeze-drying, evaporation, or precipitation. In one aspect, a whole fermentation broth comprising the polypeptide is recovered. In another aspect, a cell-free fermentation broth comprising the polypeptide is recovered.

[0221] The polypeptide may be purified by a variety of procedures known in the art to obtain substantially pure polypeptides and / or polypeptide fragments (see, e.g., Wingfield, 2015, Current Protocols in Protein Science; 80 (1) : 6.1.1-6.1.35; Labrou, 2014, Protein Downstream Processing, 1129: 3-10) .

[0222] In an alternative aspect, the polypeptide is not recovered.

[0223] Collagenase Granules

[0224] The present invention also relates to enzyme granules / particles comprising a polypeptide of the present invention, selected from the group consisting of: (i) a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1; (ii) a polypeptide having collagenase activity and a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by Alphafold; and (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0225] In an embodiment, the granule comprises a core, and optionally one or more coatings (outer layers) surrounding the core.

[0226] The core may have a diameter, measured as equivalent spherical diameter (volume based average particle size) , of 20-2000 μm, particularly 50-1500 μm, 100-1500 μm or 250-1200 μm. The core diameter, measured as equivalent spherical diameter, can be determined using laser diffraction, such as using a Malvern Mastersizer and / or the method described under ISO13320 (2020) .

[0227] In an embodiment, the core comprises a polypeptide of the present invention.

[0228] The core may include additional materials such as fillers, fiber materials (cellulose or synthetic fibers) , stabilizing agents, solubilizing agents, suspension agents, viscosity regulating agents, light spheres, plasticizers, salts, lubricants and fragrances.

[0229] The core may include a binder, such as synthetic polymer, wax, fat, or carbohydrate.

[0230] The core may include a salt of a multivalent cation, a reducing agent, an antioxidant, a peroxide decomposing catalyst and / or an acidic buffer component, typically as a homogenous blend.

[0231] The core may include an inert particle with the polypeptide absorbed into it, or applied onto the surface, e.g., by fluid bed coating.

[0232] The core may have a diameter of 20-2000 μm, particularly 50-1500 μm, 100-1500 μm or 250-1200 μm.

[0233] The core may be surrounded by at least one coating, e.g., to improve the storage stability, to reduce dust formation during handling, or for coloring the granule. The optional coating (s) may include a salt coating, or other suitable coating materials, such as polyethylene glycol (PEG) , methyl hydroxy-propyl cellulose (MHPC) and polyvinyl alcohol (PVA) .

[0234] The coating may be applied in an amount of at least 0.1%by weight of the core, e.g., at least 0.5%, at least 1%, at least 5%, at least 10%, or at least 15%. The amount may be at most 100%, 70%, 50%, 40%or 30%.

[0235] The coating is preferably at least 0.1 μm thick, particularly at least 0.5 μm, at least 1 μm or at least 5 μm. In some embodiments, the thickness of the coating is below 100 μm, such as below 60 μm, or below 40 μm.

[0236] The coating should encapsulate the core unit by forming a substantially continuous layer. A substantially continuous layer is to be understood as a coating having few or no holes, so that the core unit has few or no uncoated areas. The layer or coating should, in particular, be homogeneous in thickness.

[0237] The coating can further contain other materials as known in the art, e.g., fillers, antisticking agents, pigments, dyes, plasticizers and / or binders, such as titanium dioxide, kaolin, calcium carbonate or talc.

[0238] A salt coating may comprise at least 60%by weight of a salt, e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%or at least 99%by weight.

[0239] To provide acceptable protection, the salt coating is preferably at least 0.1 μm thick, e.g., at least 0.5 μm, at least 1 μm, at least 2 μm, at least 4 μm, at least 5 μm, or at least 8 μm. In a particular embodiment, the thickness of the salt coating is below 100 μm, such as below 60 μm, or below 40 μm.

[0240] The salt may be added from a salt solution where the salt is completely dissolved or from a salt suspension wherein the fine particles are less than 50 μm, such as less than 10 μm or less than 5 μm.

[0241] The salt coating may comprise a single salt or a mixture of two or more salts. The salt may be water soluble, in particular, having a solubility at least 0.1 g in 100 g of water at 20℃, preferably at least 0.5 g per 100 g water, e.g., at least 1 g per 100 g water, e.g., at least 5 g per 100 g water.

[0242] The salt may be an inorganic salt, e.g., salts of sulfate, sulfite, phosphate, phosphonate, nitrate, chloride or carbonate or salts of simple organic acids (less than 10 carbon atoms, e.g., 6 or less carbon atoms) such as citrate, malonate or acetate. Examples of cations in these salts are alkali or earth alkali metal ions, the ammonium ion or metal ions of the first transition series, such as sodium, potassium, magnesium, calcium, zinc or aluminum. Examples of anions include chloride, bromide, iodide, sulfate, sulfite, bisulfite, thiosulfate, phosphate, monobasic phosphate, dibasic phosphate, hypophosphite, dihydrogen pyrophosphate, tetraborate, borate, carbonate, bicarbonate, metasilicate, citrate, malate, maleate, malonate, succinate, lactate, formate, acetate, butyrate, propionate, benzoate, tartrate, ascorbate or gluconate. In particular, alkali-or earth alkali metal salts of sulfate, sulfite, phosphate, phosphonate, nitrate, chloride or carbonate or salts of simple organic acids such as citrate, malonate or acetate may be used.

[0243] The salt in the coating may have a constant humidity at 20℃ above 60%, particularly above 70%, above 80%or above 85%, or it may be another hydrate form of such a salt (e.g., anhydrate) . The salt coating may be as described in WO 00 / 01793 or WO 2006 / 034710.

[0244] Specific examples of suitable salts are NaCl (CH20℃=76%) , Na2CO3 (CH20℃=92%) , NaNO3 (CH20℃=73%) , Na2HPO4 (CH20℃=95%) , Na3PO4 (CH25℃=92%) , NH4Cl (CH20℃ =79.5%) , (NH4) 2HPO4 (CH20℃ = 93, 0%) , NH4H2PO4 (CH20℃ = 93.1%) , (NH4) 2SO4 (CH20℃=81.1%) , KCl (CH20℃=85%) , K2HPO4 (CH20℃=92%) , KH2PO4 (CH20℃=96.5%) , KNO3 (CH20℃=93.5%) , Na2SO4 (CH20℃=93%) , K2SO4 (CH20℃=98%) , KHSO4 (CH20℃=86%) , MgSO4 (CH20℃=90%) , ZnSO4 (CH20℃=90%) and sodium citrate (CH25℃=86%) . Other examples include NaH2PO4, (NH4) H2PO4, CuSO4, Mg (NO3) 2 and magnesium acetate.

[0245] The salt may be in anhydrous form, or it may be a hydrated salt, i.e., a crystalline salt hydrate with bound water (s) of crystallization, such as described in WO 99 / 32595. Specific examples include anhydrous sodium sulfate (Na2SO4) , anhydrous magnesium sulfate (MgSO4) , magnesium sulfate heptahydrate (MgSO4·7H2O) , zinc sulfate heptahydrate (ZnSO4·7H2O) , sodium phosphate dibasic heptahydrate (Na2HPO4·7H2O) , magnesium nitrate hexahydrate (Mg(NO3) 2 (6H2O) ) , sodium citrate dihydrate and magnesium acetate tetrahydrate.

[0246] Preferably the salt is applied as a solution of the salt, e.g., using a fluid bed.

[0247] The coating materials can be waxy coating materials and film-forming coating materials. Examples of waxy coating materials are poly (ethylene oxide) products (polyethyleneglycol, PEG) with mean molar weights of 1000 to 20000; ethoxylated nonylphenols having from 16 to 50 ethylene oxide units; ethoxylated fatty alcohols in which the alcohol contains from 12 to 20 carbon atoms and in which there are 15 to 80 ethylene oxide units; fatty alcohols; fatty acids; and mono-and di-and triglycerides of fatty acids. Examples of film-forming coating materials suitable for application by fluid bed techniques are given in GB 1483591.

[0248] The granule may optionally have one or more additional coatings. Examples of suitable coating materials are polyethylene glycol (PEG) , methyl hydroxy-propyl cellulose (MHPC) and polyvinyl alcohol (PVA) . Examples of enzyme granules with multiple coatings are described in WO 93 / 07263 and WO 97 / 23606.

[0249] The core can be prepared by granulating a blend of the ingredients, e.g., by a method comprising granulation techniques such as crystallization, precipitation, pan-coating, fluid bed coating, fluid bed agglomeration, rotary atomization, extrusion, prilling, spheronization, size reduction methods, drum granulation, and / or high shear granulation.

[0250] Methods for preparing the core can be found in the Handbook of Powder Technology; Particle size enlargement by C. E. Capes; Vol. 1; 1980; Elsevier. Preparation methods include known feed and granule formulation technologies, e.g.

[0251] (a) Spray dried products, wherein a liquid polypeptide-containing solution is atomized in a spray drying tower to form small droplets which during their way down the drying tower dry to form a polypeptide-containing particulate material. Very small particles can be produced this way (Michael S. Showell (editor) ; Powdered detergents; Surfactant Science Series; 1998; Vol. 71; pages 140-142; Marcel Dekker) .

[0252] (b) Layered products, wherein the polypeptide is coated as a layer around a pre-formed inert core particle, wherein a polypeptide-containing solution is atomized, typically in a fluid bed apparatus wherein the pre-formed core particles are fluidized, and the polypeptide-containing solution adheres to the core particles and dries up to leave a layer of dry polypeptide on the surface of the core particle. Particles of a desired size can be obtained this way if a useful core particle of the desired size can be found. This type of product is described in, e.g., WO 97 / 23606.

[0253] (c) Absorbed core particles, wherein rather than coating the polypeptide as a layer around the core, the polypeptide is absorbed onto and / or into the surface of the core. Such a process is described in WO 97 / 39116.

[0254] (d) Extrusion or pelletized products, wherein a polypeptide-containing paste is pressed to pellets or under pressure is extruded through a small opening and cut into particles which are subsequently dried. Such particles usually have a considerable size because of the material in which the extrusion opening is made (usually a plate with bore holes) sets a limit on the allowable pressure drop over the extrusion opening. Also, very high extrusion pressures when using a small opening increase heat generation in the polypeptide paste, which is harmful to the polypeptide (Michael S. Showell (editor) ; Powdered detergents; Surfactant Science Series; 1998; Vol. 71; pages 140-142; Marcel Dekker) .

[0255] (e) Prilled products, wherein a polypeptide-containing powder is suspended in molten wax and the suspension is sprayed, e.g., through a rotating disk atomizer, into a cooling chamber where the droplets quickly solidify (Michael S. Showell (editor) ; Powdered detergents; Surfactant Science Series; 1998; Vol. 71; pages 140-142; Marcel Dekker) . The product obtained is one wherein the polypeptide is uniformly distributed throughout an inert material instead of being concentrated on its surface. US 4, 016, 040 and US 4, 713, 245 describe this technique.

[0256] (f) Mixer granulation products, wherein a polypeptide-containing liquid is added to a dry powder composition of conventional granulating components. The liquid and the powder in a suitable proportion are mixed and as the moisture of the liquid is absorbed in the dry powder, the components of the dry powder will start to adhere and agglomerate and particles will build up, forming granulates comprising the polypeptide. Such a process is described in US 4, 106, 991, EP 170360, EP 304332, EP 304331, WO 90 / 09440 and WO 90 / 09428. In a particular aspect of this process, various high-shear mixers can be used as granulators. Granulates consisting of polypeptide, fillers and binders etc. are mixed with cellulose fibers to reinforce the particles to produce a so-called T-granulate. Reinforced particles are more robust and release less enzymatic dust.

[0257] (g) Size reduction, wherein the cores are produced by milling or crushing of larger particles, pellets, tablets, briquettes etc. containing the polypeptide. The wanted core particle fraction is obtained by sieving the milled or crushed product. Over and undersized particles can be recycled. Size reduction is described in Martin Rhodes (editor) ; Principles of Powder Technology; 1990; Chapter 10; John Wiley &Sons.

[0258] (h) Fluid bed granulation. Fluid bed granulation involves suspending particulates in an air stream and spraying a liquid onto the fluidized particles via nozzles. Particles hit by spray droplets get wetted and become tacky. The tacky particles collide with other particles and adhere to them to form a granule.

[0259] (i) The cores may be subjected to drying, such as in a fluid bed drier. Other known methods for drying granules in the feed or enzyme industry can be used by the skilled person. The drying preferably takes place at a product temperature of from 25 to 90℃. For some polypeptides, it is important the cores comprising the polypeptide contain a low amount of water before coating with the salt. If water sensitive polypeptides are coated with a salt before excessive water is removed, the excessive water will be trapped within the core and may affect the activity of the polypeptide negatively. After drying, the cores preferably contain 0.1-10%w / w water.

[0260] Non-dusting granulates may be produced, e.g., as disclosed in US 4,106,991 and US 4,661,452 and may optionally be coated by methods known in the art.

[0261] The granulate may further comprise one or more additional enzymes, e.g., hydrolase, isomerase, ligase, lyase, oxidoreductase, and transferase. The one or more additional enzymes are preferably selected from the group consisting of acetylxylan esterase, acylglycerol lipase, amylase, alpha-amylase, beta-amylase, arabinofuranosidase, cellobiohydrolases, cellulase, feruloyl esterase, galactanase, alpha-galactosidase, beta-galactosidase, beta-glucanase, beta-glucosidase, lysophospholipase, lysozyme, alpha-mannosidase, beta-mannosidase (mannanase) , phytase, phospholipase A1, phospholipase A2, phospholipase D, protease, pullulanase, pectin esterase, triacylglycerol lipase, xylanase, beta-xylosidase or any combination thereof. Each enzyme will then be present in more granules securing a more uniform distribution of the enzymes, and also reduces the physical segregation of different enzymes due to different particle sizes. Methods for producing multi-enzyme co-granulates is disclosed in the ip. com disclosure IPCOM000200739D.

[0262] Another example of formulation of polypeptides by the use of co-granulates is disclosed in WO 2013 / 188331.

[0263] The present invention also relates to protected polypeptides prepared according to the method disclosed in EP 238216.

[0264] Fermentation Broth Formulations

[0265] The present invention further relates to a fermentation broth formulation comprising a polypeptide of the present invention, selected from the group consisting of: (i) a polypeptide having collagenase activity and a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1; (ii) a polypeptide having collagenase activity and a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by Alphafold; and (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.

[0266] The fermentation broth formulation further comprises additional ingredients used in the fermentation process, such as, for example, cells (including, the host cells containing the gene encoding the polypeptide of the present invention which are used to produce the polypeptide of interest) , cell debris, biomass, fermentation media and / or fermentation products. In some embodiments, the composition is a cell-killed whole broth containing organic acid (s) , killed cells and / or cell debris, and culture medium.

[0267] The term "fermentation broth" as used herein refers to a preparation produced by cellular fermentation that undergoes no or minimal recovery and / or purification. For example, fermentation broths are produced when microbial cultures are grown to saturation, incubated under carbon-limiting conditions to allow protein synthesis (e.g., expression of enzymes by host cells) and secretion into cell culture medium. The fermentation broth can contain unfractionated or fractionated contents of the fermentation materials derived at the end of the fermentation. Typically, the fermentation broth is unfractionated and comprises the spent culture medium and cell debris present after the microbial cells (e.g., filamentous fungal cells) are removed, e.g., by centrifugation. In some embodiments, the fermentation broth contains spent cell culture medium, extracellular enzymes, and viable and / or nonviable microbial cells.

[0268] In one aspect, the composition contains an organic acid (s) , and optionally further contains killed cells and / or cell debris. In some embodiments, the killed cells and / or cell debris are removed from a cell-killed whole broth to provide a composition that is free of these components.

[0269] The fermentation broth formulation may further comprise a preservative and / or anti-microbial (e.g., bacteriostatic) agent, including, but not limited to, sorbitol, sodium chloride, potassium sorbate, and others known in the art.

[0270] A whole as described herein is typically a liquid, but may contain insoluble components, such as killed cells, cell debris, culture media components, and / or insoluble enzyme (s) . In some embodiments, insoluble components may be removed to provide a clarified liquid composition.

[0271] The whole broth formulations of the present invention may be produced by a method described in WO 90 / 15861 or WO 2010 / 096673.

[0272] Methods and Uses

[0273] The present invention also relates to methods for producing GPH-containing tripeptides, comprising a hydrolysis step by contacting a collagen source with the polypeptide of the invention, or with the above composition comprising said polypeptide of the invention; and optionally recovering the tripeptides. The tripeptides produced by the method is a mixture of different collagen tripeptides. To achieve a better bioactive effect (such as skincare, bone / joint care effect) , the molecular weight of the GPH-containing tripeptide mixture is preferably in the range of 250-800 Da, more preferably in the range of 280-600, e.g., 300-500 Da, or 300-450 Da.

[0274] In one embodiment, the contacting of a collagen source with the polypeptide is carried out under a pH of from about 5 to about 9. For example, the contacting is carried out under a pH of about 5.5, or about 6, or about 6.5, or about 7, or about 7.5, or about 8, or about 8.5.

[0275] In one embodiment, the contacting of a collagen source with the polypeptide is carried out at a temperature of about 40℃-65℃, preferably at about 45℃-55℃, about 48℃-60℃, about 50℃-58℃, about 50℃-55℃, or about 52℃-56℃.

[0276] In one embodiment, the contacting of a collagen source with the polypeptide may be carried out for at least 1 hours, e.g., at least 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5 or at least 8 hours. Preferably, the contacting may be carried out for 3-5 hours at about 45℃-55℃ under pH 6.5-7.5.

[0277] In general, the temperature, pH and the duration of contacting the collagen source (or raw collagen materials) with the collagenase of the invention can be determined by a skilled person in the art based on the present disclosure.

[0278] In one embodiment, the method of the invention further comprising a step of purification of the hydrolysed solution. The means of purification may be a known method in the art, such as centrifugation of the hydrolysed solution to removal insoluble impurities and followed by ion column purification of the centrifuged solution. Further filtration may be carried out if needed.

[0279] In another embodiment, the method of the invention further comprises a step of spraying dry to obtain a dry matter comprising the resulted tripetides. The condition of spraying dry may be carried out by known methods in the art. The dry matter may be further dehydrated to obtain a dehydrate dry matter by e.g., overnight incubation of the dry matter in an oven under 100-105℃.

[0280] In one embodiment, the content of GPH (glycine-proline-hydroxyproline) tripeptide relative to the total weight of the dehydrate dry matter is at least 3.5 wt%, at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, wherein the dehydrate dry matter is obtained by overnight incubation of the dry matter in a 105℃ oven. Dehydration of the dry matter may be carried out by other method in the art.

[0281] In one embodiment, the yield of GPH (glycine-proline-hydroxyproline) tripeptide of the present method is at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, wherein the yield of GPH is calculated based on the method described in Example 6.

[0282] In one embodiment, the collagen source or raw material is selected from skin and / or cartilage, fish and / or squid; and preferably selected from cattle skin and / or cartilage, pig skin and / or cartilage, fish skin, fish scale and / or squid.

[0283] The present invention further relates to a use of the polypeptide of the invention or the composition comprising said polypeptide in a process of producing GPH-containing collagen tripeptides (e.g., GPH-rich collagen tripeptides) .

[0284] Tripeptide Compositions and Uses thereof

[0285] The present invention further relates to collagen tripeptide compositions (e.g., food composition) , comprising the tripeptides produced by the method of the present invention.

[0286] In one embodiment, relative to the total weight of tripeptides comprised in the composition, the collagen tripeptide composition of the present invention comprises at least 2wt %of GPH tripeptides, preferably at least 2 wt%, at least 3.5 wt%, at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%.

[0287] In one embodiment, the collagen tripeptide composition of the invention is a liquid composition.

[0288] In one embodiment, the collagen tripeptide composition of the invention is a powder composition.

[0289] In one embodiment, the collagen tripeptide composition of the present invention is a healthcare product e.g., a skincare product, a dietary supplement or a functional food / drink.

[0290] Examples of the skincare product may include but not limited to, a facial or body cream, a facial or body lotion, a facial or body soap, and a shampoo or hair conditioner.

[0291] Examples of the dietary supplement may include but not limited to, bone and joint care supplements, muscle-gain supplements, anti-wrinkle supplements, skin-moisture supplements, and fascial repair supplements.

[0292] The present invention also relates to a use of the collagen tripeptide composition of present invention for preparing a healthcare product, such as skincare product (e.g., skincare lotion / cream) , or a dietary supplement (e.g., bone and joint care supplements) and / or functional food or functional drink.

[0293] The present invention further relates to a healthcare product comprising the collagen tripeptide composition. Preferred embodiments of the healthcare product include but not limit to skincare products, dietary supplements, and / or functional food / drink.

[0294] Preferred embodiments

[0295] 1) A composition comprising a polypeptide having collagenase activity, selected from the group consisting of:

[0296] (i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;

[0297] (ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and

[0298] (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitution.

[0299] 2) The composition of embodiment 1, wherein the polypeptide has enzymatic activity on gelatin.

[0300] 3) The composition of embodiment 1 or 2, wherein the polypeptide has enzymatic activity on Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) .

[0301] 4) The composition of any of the preceding embodiments, wherein the polypeptide comprises or consists of SEQ ID NO: 2 or SEQ ID NO: 1, or the mature polypeptide of SEQ ID NO: 1.

[0302] 5) The composition of any of the preceding embodiments, wherein the polypeptide having collagenase activity is obtained from Paenibacillus, e.g., Paenibacillus azoreducens.

[0303] 6) The composition of any of the preceding embodiments, wherein the polypeptide is a variant of SEQ ID NO: 2 or SEQ ID NO: 1.

[0304] 7) The composition of the preceding embodiment, wherein the variant comprises a substitution, deletion, and / or insertion at one or more (e.g., up to 20, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) positions of SEQ ID NO: 2 or SEQ ID NO: 1, or a mature polypeptide of SEQ ID NO: 1.

[0305] 8) The composition of any of the preceding embodiments, wherein the polypeptide is present in an amount of from about 0.1 mg / g to about 200 mg / g enzyme protein; preferably in an amount of from about 1 mg / g to about 100 mg / g, from about 2 mg / g to about 50 mg / g, from about 3 mg / g to about 15 mg / g, from about 5 mg / g to about 18 mg / g, or from about 5 mg / g to about 10 mg / g or more, based on the total amount of the composition.

[0306] 9) The composition of any of the preceding embodiments, which is a granulate composition.

[0307] 10) The composition of any of the preceding embodiments, which is a powder composition.

[0308] 11) The composition of any of embodiments 1-10, wherein the polypeptide is present in an amount of about 0.01 to about 99 wt%; preferably in an amount of from about 1 to about 90 wt%, or from about 2 to about 85 wt%, or from about 3 to about 80 wt%, or from about 4 to about 70 wt%, or from about 5 to about 60 wt%, based on the total amount of the composition.

[0309] 12) The composition of any of embodiments 1-8, which is a liquid composition; preferably an aqueous composition.

[0310] 13) The composition of embodiment 12, wherein the liquid composition comprises an aqueous buffer; preferably wherein the aqueous buffer comprises 4- (2-hydroxyethyl) -1-piperazineethanesulfonic acid (HEPES) , tris (hydroxymethyl) aminomethane (TRIS) , enzyme stabilizers, phosphate, or bicarbonate; most preferably wherein the aqueous buffer comprises enzyme stabilizers, phosphate, or bicarbonate.

[0311] 14) The composition of any of embodiments 12-13, which has a pH value of about 5 to about 9; more preferably of about 6.5 to about 8.5, of about 7 to about 8 or of about 7 to about 7.5.

[0312] 15) An isolated polypeptide having collagenase activity, selected from the group consisting of:

[0313] (i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;

[0314] (ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold;

[0315] (iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitution.

[0316] 16) The polypeptide of embodiment 15, wherein the polypeptide has enzymatic activity on gelatin.

[0317] 17) The polypeptide of embodiment 15 or 16, wherein the polypeptide has enzymatic activity on Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) .

[0318] 18) The polypeptide of any of embodiments 15-17, which comprises or consists of SEQ ID NO: 2 or the mature polypeptide of SEQ ID NO: 1.

[0319] 19) The polypeptide of any embodiments 15-18, which is a variant of SEQ ID NO: 2 or of the mature polypeptide of SEQ ID NO: 1, preferably the variant comprises a substitution, deletion, and / or insertion at one or more (e.g., up to 20, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20) positions of SEQ ID NO: 2 or SEQ ID NO: 1, or a mature polypeptide of SEQ ID NO: 1.

[0320] 20) An isolated polynucleotide encoding the polypeptide of any of embodiments15-19.

[0321] 21) A nucleic acid construct or an expression vector comprising a polynucleotide of any preceding embodiments 20.

[0322] 22) A recombinant host cell comprising in its genome the nucleic acid construct or expression vector of embodiment 21.

[0323] 23) The recombinant host cell of embodiment 22, which is a B. subtilis cell, B. licheniformis cell, A. niger cell, A. oryzae cell, T. reesei cell, or P. pastoris (K. phaffii) cell.

[0324] 24) A method for producing a polypeptide of any of embodiments 15-19, comprising (a) cultivating a recombinant host cell of embodiment 23 under conditions conducive for expression of the polypeptide; and (b) optionally recovering the polypeptide.

[0325] 25) A method for producing GPH-containing tripeptides, comprising a hydrolysis step by contacting a collagen source with the polypeptide of any of embodiments 15-19, or with the composition of any of embodiments 1-14, preferably, said tripeptides is GPH-rich tripeptides.

[0326] 26) The method of embodiment 25, wherein the contacting is carried out under a pH of about 5-9 (e.g., at a pH of about 5.5, about 6, about 6.5, about 7, about 7.5, about 8, or about 8.5) and preferably at a temperature of about 40℃-65℃, more preferably at a temperature of about 45℃-55℃, about 48℃-60℃, about 50℃-58℃, about 50℃-55℃, or about 52℃-56℃, for at least 1 hours (e.g., at least 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5 or 8 hours) .

[0327] 27) The method of embodiment 25 or 26, further comprising a step of purification (e.g., by centrifugation and / or filtration) of the hydrolysed solution; and a step of spraying dry of the purified solution to obtain a dry matter.

[0328] 28) The method of embodiment 27, wherein the content of GPH (glycine-proline-hydroxyproline) tripeptide relative to the total weight of a dehydrate dry matter is at least 3.5 wt%, at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, wherein the dehydrate dry matter is obtained by overnight incubation of the dry matter obtained in embodiment 28 in an 105℃ oven.

[0329] 29) The method of any one of embodiments 25-28, wherein the yield of GPH (glycine-proline-hydroxyproline) tripeptide is at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, wherein the yield of GPH is calculated based on the method described in Example 6.

[0330] 30) The method of any one of embodiments 25-29, wherein the molecular weight of the GPH-containing tripeptides is in the range of 250-800 Da, preferably in the range of 280-600, e.g., 300-500 Da, or 300-450 Da.

[0331] 31) The method of any of embodiments 25-30, wherein the collagen source is selected from cattle skin and / or cartilage, pig skin and / or cartilage, fish and / or squid; and preferably selected from cow skin and / or cartilage, pig skin and / or cartilage, fish skin, fish scale and / or squid.

[0332] 32) A collagen tripeptide composition, comprising the tripeptides produced according to any of embodiments 25-31.

[0333] 33) A collagen tripeptide composition, wherein the GPH is present at a level of at least 2 wt%, e.g., at least 3 wt%, at least 3.5 wt%, at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, relative to the total weight of the collagen tripeptides (CTP) comprised in the composition.

[0334] 34) The composition of any of embodiments 32-33, which is a liquid or a powder composition.

[0335] 35) A healthcare product comprising the collagen tripeptide composition of any of embodiments 32-34.

[0336] 36) The healthcare product of embodiment 35, which is a skincare product, a dietary supplement or a functional food / drink.

[0337] 37) Use of the collagen tripeptide composition of any of embodiments 32-34 for preparing a healthcare product, such as skincare products, dietary supplements and / or a functional food / drink.

[0338] 38) Use of the composition any of embodiments 1-14 or the polypeptide of any of embodiments 15-19 in a process of producing collagen tripeptides.

[0339] EXAMPLES

[0340] Materials and methods

[0341] Assay I: Collagenase Activity Assay

[0342] Collagenase hydrolyzes the substrate N- (3- [2-Furyl] acryloyl) -Leu-Gly-Pro-Ala (FALGPA) . This reaction produces an absorption decrease at 340 nm, which is proportional to the enzyme activity. The reactions were performed at pH 7.5 in 50mM tricine buffer containing 10mM CaCl2 and 400mM NaCl, 1.076mM substrate FALGPA at 37℃. Upon mixture of the substrate with the enzyme, the decrease in absorbance at 340 nm was monitored every 30 seconds for 8 minutes by a microtiter plate reader (EnSight system from PerkinElmer Inc., Waltham, MA, USA) ) and obtain linear rate (ΔA340 / minute) for all the tested samples and the blank samples. The sample should be diluted to a level where the slope is linear.

[0343] Activity calculation: 1 U equals to 1μmol FALGPA degradation in 1 min

[0344] U / mL= (A340test / min-A340blank / min)  / 0.53 / L* (Vr / Ve) *DF

[0345] L -optical length

[0346] 0.53 -millimolar extinction coefficient of FALGPA

[0347] Vr –Reaction volumn

[0348] Ve –Enzyme volumn

[0349] DF -Dilution factor

[0350] Assay II: Determination the amount or yield of GPH by HPLC

[0351] Collagen tripeptide e.g., GPH can be measured by any method known in the art, such as High-performance Liquid Chromatography (HPLC) . In the present invention, the yield or the amount of GPH is determined by HPLC (Agilgent Technologies, Santa Clara, California, USA) . The amount of GPH is calculated by comparing the peak area of GPH produced by the method of the present invention with the peak area of a Gly-Pro-Hyp peptide standard solution (PEPSTD-12007, commercially available from XIYU (Shanghai) Technology Co., Ltd, Shanghai, China) .

[0352] Detailed HPLC conditions are shown as below:

[0353] ○ Column ZORBAX SB-Aq, 4.6*250mm, 5μm

[0354] ○ Flow rate: 1 ml / min

[0355] ○ Mobile Phase: 0.1%trifluoroacetic acid (TFA)

[0356] ○ Diluted sample with 0.1%TFA

[0357] ○ Detection wavelength: 220nm

[0358] ○ Column temperature: 50℃

[0359] ○ 10μl inject volume.

[0360] Example 1: Cloning and expression of targeted collagenase having SEQ ID NO: 1

[0361] The target gene (SEQ ID NO: 1) , originally from Paenibacillus azoreducens, was cloned and transformed into Komagataella phaffii (formerly named Pichia pastoris) GLM chassis strain. The construction of the chassis strain is described in Chinese patent grant publication No. CN108949869B. Briefly, restriction enzyme BlnI was used to linearized plasmid pGGLacIMit1AD, followed by transformation into the Komagataella phaffii GS115 strain (InvitrogenTM, ThermoFisher Scientific Inc., Waltham, MA, USA) . The resulting chassis strain was referred to here as Komagataella. phaffii GLM strain.

[0362] An expression construct of a fusion protein (SEQ ID NO: 3) was prepared in which the signal peptide MQVKSIVNLLLACSLAVA (SEQ ID NO: 4) was operably linked to the polypeptide having collagenase activity. The expression construct was transformed into the competent cells of Komagataella phaffii GLM strain and integrated at the HIS location of the genome. Transformed cultures were coated onto YND (Yeast Nitrogen Dextrose) agar plates (0.67%Yeast Nitrogen Base, 1%glucose, 2%agar) and incubated at 30℃ for 48-72 h. The Komagataella phaffii clone harboring the correct expression construct was confirmed by colony PCR and DNA sequencing.

[0363] Example 2: Collagenase strain fermentation

[0364] Fed-batch fermentation media: YNB (Yeast Nitrogen Base) 13.4 g / L, peptone 20.0 g / L, yeast extract: 10.0g / L, K2HPO4 3.915 g / L, KH2PO4 10.54g / L, biotin 0.0004 g / L, glucose 40 g / L.

[0365] PTM1 stock: CuSO4·5H2O 6.0g / L, NaI 0.08 g / L, MnSO4·H2O 3.0g / L, Na2MoO4·2H2O 0.2 g / L, H3BO3 0.02 g / L, CoCl2·6H2O 0.914g / L, ZnCl2 20.0 g / L, FeSO4·7H2O 65.0 g / L, H2SO4 5 mL / L, biotin 0.2 g / L.

[0366] Fermentation process control: under ambient state, stirring paddle speed at 200 rpm, ventilation 0.5 L /  (L·min) , temperature 30℃, calibrated dissolved oxygen electrode 100%.

[0367] When collagenase expressed Komagataella phaffii strain reached OD600 ≈1.0, it was inoculated into the fermentation media. Then 4.5ml / L PTM1 stock was added into fermentation media and strain mixture, and automatic control of pH and dissolved oxygen was set. The pH was controlled within the range between 5.2 to 6.5 using ammonia solution. Dissolved oxygen, ventilation and stirring control were maintained at 30%. After the carbon source was completely depleted in the batch stage, the glucose feeding medium was added into fermentation process, with feeding rate was controlled at 2~6 g / L / h.

[0368] Example 3: Purification of collagenase from 3 liter fermentation broth

[0369] A total of 200 mL of Komagataella phaffii culture broth was harvested, and the conductivity was adjusted to approximately 140 mS / cm by the addition of ammonium sulfate (AMS) . The supernatant, which was filtered against 0.2 μm membrane, underwent initial purification via hydrophobic interaction chromatography (HIC) using Phenyl Sepharose High Performance column (17-1082-03, GE, Boston, Massachusetts, USA) . The column was equilibrated with Buffer A (20 mM Tris-HCl, pH 7.0, containing 1.2 M AMS) , and elution was performed using a linear gradient of Buffer B (20 mM Tris-HCl, pH 7.0) . Fractions containing the target protein were pooled and dialyzed to reduce the salt concentration. The dialyzed sample was further purified by ion exchange chromatography (IEC) on MonoQ column (GE, 17-0506-01) . Buffer C (20 mM Tris-HCl, pH 7.5) was used for column equilibration, and Buffer D (20 mM Tris-HCl, pH 7.5, containing 1 M NaCl) was applied for gradient elution. Protein-containing fractions from the IEC step were subjected to a second round of HIC purification under the same conditions as described above. Following three successive purification steps, the purity of the target protein exceeded 95%, as confirmed by SDS-PAGE analysis. The purified protein fractions were then concentrated and exchanged with 20 mM Tris-HCl buffer (pH 7.0) by ultrafiltration. Protein concentration was determined using the protein assay kit (QubitTM, Thermo Fisher Scientific, Waltham, MA, USA) .

[0370] The purified protein was digested into peptides using trypsin. The resulting peptides were analyzed by liquid chromatography-mass spectrometry (nanoLC-MS / MS) to determine the sequence. The mature polypeptide was shown in SEQ ID NO: 2.

[0371] Example 4: Cloning expression and purification of sequences from GenBank: BFH62464.1, NCBI reference sequence: WP_212977976.1

[0372] Protein sequences from BFH62464.1 (SEQ ID NO: 5) and WP_212977976.1 (SEQ ID NO: 7) with C terminal 6x His tag were codon optimized and cloned into E. coli BL21 (DE3) cells (Zoonbio Biotechnology, Nanjing, China) . Plasmid Pet30a (Zoonbio Biotechnology) and restriction sites NdeI and XhoI were used for recombinant expression. Transformed cultures were coated onto LB agar plates (5g / L yeast extract, 10g / L tryptone, 5g / L NaCl and 12g / L agar) overnight and Kanamycin resistance was adopted to select positive transformations. E. coli BL21 (DE3) cells harboring correct expression construct was confirmed by colony PCR and DNA sequencing.

[0373] Isolated BL21 (DE3) colonies with target enzymes were inoculated and cultivated in TB media (tryptone 12g / L, yeast extract 24g / L, glycerol 0.4%v / v, potassium dihydrogen phosphate 2.3g / L and potassium phosphate dibasic 12.54g / L) for 4 hours at 37℃ until OD600 reached 0.6-0.8. Then 0.2mM Isopropyl β-D-1-thiogalactopyranoside (IPTG) was added into culture to induce recombinant enzyme expression at 15℃. After overnight induction, cell cultures were centrifuged and harvested for cell lysis. At last, the supernatant from cell lysate were collected and transferred to sterile container for purification by Ni-NTA Resin.

[0374] Low-pressure chromatography was used. The supernatant was loaded at a flow rate of 0.5 mL / min to a Ni-IDA-Sepharose Cl-6B affinity column from Novagen, Madison, WI, USA pre-balanced by Ni-IDA Binding-Buffer (20 mM Tris-HCl, 0.15 M NaCl, pH 8.0) . Then affinity column was first rinsed with Ni-IDA Binding-Buffer at a flow rate of 0.5 mL / min until the effluent OD280 value reached baseline, followed by rinse with Ni-IDA Washing-Buffer (20 mM Tris-HCl, 30 mM imidazole, 0.15 M NaCl, pH 8.0) at a flow rate of 1 mL / min until effluent OD280 value reached baseline. The effluent was then collected by eluting the protein of interest at a flow rate of 1 mL / min with a Ni-IDA Elution-Buffer (20 mM Tris-HCl, 250 mM imidazole, 0.15 M NaCl, pH 8.0) . The protein solution collected above was added to the dialysis bag and dialysis was performed overnight using phosphate-buffered saline (PBS) . Finally a 12%SDS-PAGE analysis was performed to confirm the purity of target protein.

[0375] The purified proteins were digested into peptides using trypsin. The resulting peptides were analyzed by liquid chromatography-mass spectrometry (nanoLC-MS / MS) to determine the sequence. The purified BFH62464.1 collagenase ColA was shown in SEQ ID NO: 5 and the purified WP_212977976.1 collagenase was shown in SEQ ID NO: 7.

[0376] Example 5: Specific activities of collagenase on substrate N- (3- [2-Furyl] acryloyl) -Leu-Gly-Pro-Ala (FALGPA)

[0377] The specific activities of purified SEQ ID NO: 2, SEQ ID NO: 5 and SEQ ID NO: 7 were tested according to collagenase activity assay in Assay I. The enzyme samples were diluted to a level where the decrease in absorbance at 340 nm was linear for 8 minutes. The results were shown in the table below.

[0378] Table 1 Specific activity of purified SEQ ID NO: 2, SEQ ID NO: 5 and SEQ ID NO: 7

[0379] Compared to SEQ ID NO: 5 and SEQ ID NO: 7, SEQ ID NO: 2 exhibited significantly higher specific activity.

[0380] Example 6: Producing GPH-containing tripeptides by degrading pig skin gelatin with collagenases

[0381] Collagenase was mixed with 10%pig skin gelatin in 100 mM Tris-HCI buffer having a pH of 7, with a dosage of 0.06 mg enzyme protein  / g substrate. The reaction was carried out at 50℃ with 700rpm shaking for 5 hours. The reaction was terminated by heating the mixture for 15min at 85 ℃. GPH amount was measured by HPLC according to Assay II. The yield of GPH was calculated as: measured amount (g) of GPH  / amount (g) of gelatin substrate (i.e., pig skin gelatin) x 100%. Results were shown in Table 2.

[0382] Table 2. GPH yield from pig skin gelatin by collagenases

[0383] Compared to SEQ ID NO: 5 and SEQ ID NO: 7, SEQ ID NO: 2 produced significantly higher GPH yield.

Claims

1.A composition comprising a polypeptide having collagenase activity, wherein the polypeptide is selected from the group consisting of:(i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;(ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and(iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.2.The composition of claim 1, wherein the polypeptide has enzymatic activity on gelatin or the substrate Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) .3.The composition of claim 1 or 2, wherein the polypeptide comprises or consists of SEQ ID NO: 2, SEQ ID NO: 1, or the mature polypeptide of SEQ ID NO: 1.4.The composition of any of the preceding claims, wherein the polypeptide having collagenase activity is obtained from Paenibacillus, e.g., Paenibacillus azoreducens.5.The composition of any of the preceding claims, wherein the polypeptide having collagenase activity is present in an amount of from about 0.1 mg / g to about 200 mg / g enzyme protein; preferably in an amount of from about 1 mg / g to about 100 mg / g, from about 2 mg / g to about 50 mg / g, from about 3 mg / g to about 15 mg / g, from about 5 mg / g to about 18 mg / g, or from about 5 mg / g to about 10 mg / g or more, based on the total amount of the composition.6.An isolated polypeptide having collagenase activity, selected from the group consisting of:(i) a polypeptide having a sequence identity of at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%, to SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1;(ii) a polypeptide having a TM-score of at least 0.80, at least 0.85, at least 0.90, at least 0.905, at least 0.910, at least 0.915, at least 0.920, at least 0.925, at least 0.930, at least 0.935, at least 0.940, at least 0.945, at least 0.950, at least 0.955, at least 0.960, at least 0.965, at least 0.970, at least 0.975, at least 0.980, at least 0.985, at least 0.990, at least 0.995, or even 1.0, to the three-dimensional structure of the polypeptide of SEQ ID NO: 2, wherein the three-dimensional structure is calculated by AlphaFold; and(iii) a polypeptide derived from SEQ ID NO: 2, SEQ ID NO: 1 or a mature polypeptide of SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions.7.The polypeptide of claim 6, wherein the polypeptide has enzymatic activity on gelatin or the substrate Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) .8.The polypeptide of claim 6 or 7, which comprises or consists of SEQ ID NO: 2, SEQ ID NO: 1 or the mature polypeptide of SEQ ID NO: 1.9.An isolated polynucleotide encoding the polypeptide of any of claims 6-8.10.A nucleic acid construct or an expression vector comprising a polynucleotide of claim 9.11.A recombinant host cell comprising the nucleic acid construct or expression vector of claim 10.12.The recombinant host cell of claim 11, which is a B. subtilis cell, B. licheniformis cell, A. niger cell, A. oryzae cell, T. reesei cell, or P. pastoris (K. phaffii) cell.13.A method for producing a polypeptide of any of claims 6-8, comprising (a) cultivating a recombinant host cell of claim 11 or 12 under conditions conducive for expression of the polypeptide; and (b) optionally recovering the polypeptide.14.A method for producing GPH-containing tripeptides, comprising a hydrolysis step by contacting a collagen source with the polypeptide of any of claims 6-8, or with the composition of any of claims 1-5.15.The method of claim 14, wherein the yield of GPH (glycine-proline-hydroxyproline) tripeptide is at least 4 wt%, e.g., at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%.16.A collagen tripeptide composition, comprising the tripeptides produced according to claim 14 or 15.17.A collagen tripeptide composition, wherein the GPH is present at a level of at least 2 wt%, e.g., at least 3 wt%, at least 3.5 wt%, at least 4 wt%, at least 4.5 wt%, at least 5 wt%, at least 5.5 wt%, at least 6 wt%, at least 6.5 wt%, at least 7 wt%, at least 8 wt%, at least 9 wt%, at least 10 wt%, at least 11 wt%, at least 12 wt%, at least 13 wt%, at least 14 wt%, at least 15 wt%, at least 16 wt%, at least 17 wt%, at least 18 wt%, at least 19 wt%, or at least 20 wt%, relative to the total weight of collagen tripeptides (CTP) comprised in the composition.18.Use of the collagen tripeptide composition of claim 16 or 17 for preparing healthcare products, such as skincare products, dietary supplements and / or a functional food / drink.