Codon optimized collagenase expression

WO2026179860A1PCT designated stage Publication Date: 2026-09-03NOVOZYMES AS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/079783
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-25
Filing Date
2026-02-24
Publication Date
2026-09-03

Smart Images

  • Figure PCTCN2026079783-FTAPPB-I100001
    Figure PCTCN2026079783-FTAPPB-I100001
  • Figure PCTCN2026079783-FTAPPB-I100002
    Figure PCTCN2026079783-FTAPPB-I100002
  • Figure PCTCN2026079783-FTAPPB-I100003
    Figure PCTCN2026079783-FTAPPB-I100003
Patent Text Reader

Abstract

Provided are synthetic polynucleotides encoding a collagenase, nucleic acid constructs, vectors and host cells comprising the synthetic polynucleotides. Also provided are methods of producing the collagenase.
Need to check novelty before this filing date? Find Prior Art

Description

CODON OPTIMIZED COLLAGENASE EXPRESSION

[0001] Reference to a Sequence Listing

[0002] This application contains a Sequence Listing in computer readable form, which is incorporated herein by reference.Background of the InventionField of the Invention

[0003] The present invention relates to synthetic polynucleotides encoding a collagenase, and to nucleic acid constructs, vectors, and host cells comprising the synthetic polynucleotides as well as methods of producing the collagenase.

[0004] Description of the Related Art

[0005] In the highly competitive industrial manufacture of enzymes it is of vital importance to constantly improve yield or productivity. Genetic manipulation or engineering has been put to use for this purpose for many years, where genes encoding polypeptides of interest have been placed under the transcriptional control of heterologous or synthetic promoters, expressed with heterologous signal peptides in various host cells and integrated in the host cell genomes in multiple copies in order to achieve so-called mRNA-saturation.

[0006] Another well-known technique to increase enzyme productivity has been to optimize the codon-usage in enzyme-encoding DNA sequence based on that of the host cell intended for its expression and based on various theoretical mRNA melting point or tertiary structure calculations.

[0007] Even so, it remains of significant interest to identify new ways to improve the expression of an enzyme of interest. Due to the highly competitive environment in the enzyme manufacturing industry, even minor improvements are desirable.Summary of the Invention

[0008] The present invention provides means and methods to increase recombinant collagenase production. Testing two codon-optimized collagenase-coding sequences the present inventors identified a synthetic DNA sequence which significantly increased collagenase expression compared to the other optimized synthetic sequence. For the synthetic optimized sequence of design 1 the collagenase yield was significantly increased compared to the synthetic sequence of design 2, which result was totally unexpected.

[0009] Accordingly, in a 1st aspect the present invention relates to synthetic polynucleotides encoding a collagenase, selected from the group consisting of:

[0010] (a) a polynucleotide having at least 80%sequence identity to SEQ ID NO: 1,

[0011] (b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0012] (c) a polynucleotide derived from the polynucleotide of (a) , or (b) , wherein the 3’ -and / or 5’ -end has been extended by addition of one or more nucleotides; and

[0013] (d) a fragment of the polynucleotide of (a) , (b) , or (c) .

[0014] In a 2nd aspect the invention relates to nucleic acid constructs of expression vectors comprising the polynucleotide of the 1st aspect.

[0015] In a 3rd aspect the invention relates to a recombinant host cell comprising the nucleic acid construct or expression vector of the 2nd aspect.

[0016] In a 4th aspect the invention relates to a composition, cell composition or fermentation broth comprising the polynucleotide of the 1st aspect and / or the cell of the 3rd aspect.

[0017] In a 5th aspect the invention relates to methods of producing a collagenase comprising cultivating the host cell of the 3rd aspect under conditions conducive for production of the collagenase.Brief Description of the Drawings

[0018] Figure 1 shows a SDS PAGE for the collagenase yield of synthetic DNA design 1 and design 2.

[0019] Definitions

[0020] In accordance with this detailed description, the following definitions apply. Note that the singular forms "a, " "an, " and "the" include plural references unless the context clearly dictates otherwise.

[0021] Unless defined otherwise or clearly indicated by context, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0022] Collagenase: The term “collagenase” or “polypeptide having collagenase activity” can be used interchangeably in the present invention. In one embodiment, the collagenase of the present invention has enzymatic activity on gelatin. Gelatin is a collection of peptides and proteins produced by partial hydrolysis of collagen extracted from e.g., the skin, bones, and connective tissues of animals such as domesticated cattle, chicken, pigs, and fish. The collagenase of the present invention cleaves the substrate Phe-Ala-Leu-Gly-Pro-Ala (FALGPA) at a position between L and G, i.e., the collagenase of the present invention has enzymatic activity on FALGPA. Non-limiting examples of collagenase polypeptides comprise the polypeptide of SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0023] Collagenases are catalytic proteins (enzymes) , and the term “active (collagenase) enzyme protein” is defined herein as the amount of catalytic protein (s) , which exhibits collagenase activity. This can be determined using an activity based analytical enzyme assay. This technique is well-known in the art.

[0024] Collagenase activity: For the purpose of the present invention, collagenase activity is determined by the FALGPA assay (Van Wart and Steinbrink, 1981, 113 (2) : 356-365) . Briefly, collagenase is capable of hydrolyzing the substrate N- (3- [2-Furyl] acryloyl) -Leu-Gly-Pro-Ala (FALGPA; CAS. 78832-65-2,  USA) . This reaction produces an absorption decrease at 340-345 nm which is proportional to the enzyme activity, where one unit of collagenase hydrolyzes 1.0 umole of FALGPA per minute at 25 ℃ at pH 7.5 in the presence of calcium ions. The assay is also used to generate a linear slope so that the collagenase concentration in a given sample may also be determined.

[0025] cDNA: The term "cDNA" means a DNA molecule that can be prepared by reverse transcription from a mature, spliced, mRNA molecule obtained from a eukaryotic or prokaryotic cell. cDNA lacks intron sequences that may be present in the corresponding genomic DNA. The initial, primary RNA transcript is a precursor to mRNA that is processed through a series of steps, including splicing, before appearing as mature spliced mRNA.

[0026] Coding sequence: The term “coding sequence” means a polynucleotide, which directly specifies the amino acid sequence of a polypeptide. The boundaries of the coding sequence are generally determined by an open reading frame, which begins with a start codon, such as ATG, GTG, or TTG, and ends with a stop codon, such as TAA, TAG, or TGA. The coding sequence may be a genomic DNA, cDNA, synthetic DNA, or a combination thereof.

[0027] Control sequences: The term “control sequences” means nucleic acid sequences involved in regulation of expression of a polynucleotide in a specific organism or in vitro. Each control sequence may be native (i.e., from the same gene) or heterologous (i.e., from a different gene) to the polynucleotide encoding the polypeptide, and native or heterologous to each other. Such control sequences include, but are not limited to leader, polyadenylation, prepropeptide, propeptide, signal peptide, promoter, terminator, enhancer, and transcription or translation initiator and terminator sequences. At a minimum, the control sequences include a promoter, and transcriptional and translational stop signals. The control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the polynucleotide encoding a polypeptide.

[0028] Expression: The term “expression” means any step involved in the production of a polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0029] Expression vector: An "expression vector" refers to a linear or circular DNA construct comprising a DNA sequence encoding a polypeptide, which coding sequence is operably linked to a suitable control sequence capable of effecting expression of the DNA in a suitable host. Such control sequences may include a promoter to effect transcription, an optional operator sequence to control transcription, a sequence encoding suitable ribosome binding sites on the mRNA, enhancers and sequences which control termination of transcription and translation.

[0030] Extension: The term “extension” means an addition of one or more amino acids to the amino and / or carboxyl terminus of a polypeptide, wherein the “extended” polypeptide has collagenase activity.

[0031] Fragment: The term “fragment” means a polypeptide having one or more amino acids absent from the amino and / or carboxyl terminus of the mature polypeptide wherein the fragment has collagenase activity.

[0032] Fusion polypeptide: The term “fusion polypeptide” is a polypeptide in which one polypeptide is fused at the N-terminus and / or the C-terminus of a polypeptide of the present invention. A fusion polypeptide is produced by fusing a polynucleotide encoding another polypeptide to a polynucleotide of the present invention, or by fusing two or more polynucleotides of the present invention together. Techniques for producing fusion polypeptides are known in the art, and include ligating the coding sequences encoding the polypeptides so that they are in frame and that expression of the fusion polypeptide is under control of the same promoter (s) and terminator. Fusion polypeptides may also be constructed using intein technology in which fusion polypeptides are created post-translationally (Cooper et al., 1993, EMBO J. 12: 2575-2583; Dawson et al., 1994, Science 266: 776-779) . A fusion polypeptide can further comprise a cleavage site between the two polypeptides. Upon secretion of the fusion protein, the site is cleaved releasing the two polypeptides. Examples of cleavage sites include, but are not limited to, the sites disclosed in Martin et al., 2003, J. Ind. Microbiol. Biotechnol. 3: 568-576; Svetina et al., 2000, J. Biotechnol. 76: 245-251; Rasmussen-Wilson et al., 1997, Appl. Environ. Microbiol. 63: 3488-3493; Ward et al., 1995, Biotechnology 13: 498-503; and Contreras et al., 1991, Biotechnology 9: 378-381; Eaton et al., 1986, Biochemistry 25: 505-512; Collins-Racie et al., 1995, Biotechnology 13: 982-987; Carter et al., 1989, Proteins: Structure, Function, and Genetics 6: 240-248; and Stevens, 2003, Drug Discovery World 4: 35-48.

[0033] Heterologous: The term "heterologous" means, with respect to a host cell, that a polypeptide or nucleic acid does not naturally occur in the host cell. The term "heterologous" means, with respect to a polypeptide or nucleic acid, that a control sequence, e.g., promoter, of a polypeptide or nucleic acid is not naturally associated with the polypeptide or nucleic acid, i.e., the control sequence is from a gene other than the gene encoding the mature polypeptide.

[0034] Host Strain or Host Cell: A "host strain" or "host cell" is an organism into which an expression vector, phage, virus, or other DNA construct, including a polynucleotide of the present invention has been introduced. Exemplary host strains are microorganism cells (e.g., bacteria, filamentous fungi, and yeast, for example, Pichia) capable of expressing the polypeptide of interest and / or fermenting saccharides. The term "host cell" includes protoplasts created from cells.

[0035] Introduced: The term "introduced" in the context of inserting a nucleic acid sequence into a cell, means "transfection" , "transformation" or "transduction, " as known in the art.

[0036] Isolated: The term “isolated” means a polypeptide, nucleic acid, cell, or other specified material or component that has been separated from at least one other material or component, including but not limited to, other proteins, nucleic acids, cells, etc. An isolated polypeptide, nucleic acid, cell or other material is thus in a form that does not occur in nature. An isolated polypeptide includes, but is not limited to, a culture broth containing the secreted polypeptide expressed in a host cell.

[0037] Mature polypeptide: The term “mature polypeptide” means a polypeptide in its mature form following N-terminal and / or C-terminal processing (e.g., removal of signal peptide) . In one aspect, the mature polypeptide is SEQ ID NO: 4 (after removal of the signal peptide) . In one aspect, the mature polypeptide is SEQ ID NO: 5 or SEQ ID NO: 6 (after removal of the signal peptide and different pro-peptide region) .

[0038] Mature polypeptide coding sequence: The term “mature polypeptide coding sequence” means a polynucleotide that encodes a mature polypeptide having collagenase activity. In one aspect, the mature polypeptide coding sequence is nucleotides 421 to 3120 of any one of SEQ ID NO: 1, 2, or 3. In another aspect, the mature polypeptide coding sequence is nucleotides 454 to 3120 of any one of SEQ ID NO: 1, 2, or 3.

[0039] Native: The term "native" means a nucleic acid or polypeptide naturally occurring in a host cell.

[0040] Nucleic acid: The term "nucleic acid" encompasses DNA, RNA, heteroduplexes, and synthetic molecules capable of encoding a polypeptide. Nucleic acids may be single stranded or double stranded, and may be chemical modifications. The terms "nucleic acid" and "polynucleotide" are used interchangeably. Because the genetic code is degenerate, more than one codon may be used to encode a particular amino acid, and the present compositions and methods encompass nucleotide sequences that encode a particular amino acid sequence. Unless otherwise indicated, nucleic acid sequences are presented in 5'-to-3'orientation.

[0041] Nucleic acid construct: The term "nucleic acid construct" means a nucleic acid molecule, either single-or double-stranded, which is isolated from a naturally occurring gene or is modified to contain segments of nucleic acids in a manner that would not otherwise exist in nature or which is synthetic, and which comprises one or more control sequences operably linked to the nucleic acid sequence.

[0042] Operably linked: The term "operably linked" means that specified components are in a relationship (including but not limited to juxtaposition) permitting them to function in an intended manner. For example, a regulatory sequence is operably linked to a coding sequence such that expression of the coding sequence is under control of the regulatory sequence.

[0043] Purified: The term “purified” means a nucleic acid, polypeptide or cell that is substantially free from other components as determined by analytical techniques well known in the art (e.g., a purified polypeptide or nucleic acid may form a discrete band in an electrophoretic gel, chromatographic eluate, and / or a media subjected to density gradient centrifugation) . A purified nucleic acid or polypeptide is at least about 50%pure, usually at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8%or more pure (e.g., percent by weight or on a molar basis) . In a related sense, a composition is enriched for a molecule when there is a substantial increase in the concentration of the molecule after application of a purification or enrichment technique. The term "enriched" refers to a compound, polypeptide, cell, nucleic acid, amino acid, or other specified material or component that is present in a composition at a relative or absolute concentration that is higher than a starting composition.

[0044] ln one aspect, the term "purified" as used herein refers to the polypeptide or cell being essentially free from components (especially insoluble components) from the production organism. In other aspects. the term "purified" refers to the polypeptide being essentially free of insoluble components (especially insoluble components) from the native organism from which it is obtained. In one aspect, the polypeptide is separated from some of the soluble components of the organism and culture medium from which it is recovered. The polypeptide may be purified (i.e., separated) by one or more of the unit operations filtration, precipitation, or chromatography.

[0045] Accordingly, the polypeptide may be purified such that only minor amounts of other proteins, in particular, other polypeptides, are present. The term "purified" as used herein may refer to removal of other components, particularly other proteins and most particularly other enzymes present in the cell of origin of the polypeptide. The polypeptide may be "substantially pure" , i.e., free from other components from the organism in which it is produced, e.g., a host organism for recombinantly produced polypeptide. In one aspect, the polypeptide is at least 40%pure by weight of the total polypeptide material present in the preparation. In one aspect, the polypeptide is at least 50%, at least 60%, at least 70%, at least 80%or at least 90%pure by weight of the total polypeptide material present in the preparation. As used herein, a "substantially pure polypeptide" may denote a polypeptide preparation that contains at most 10%, preferably at most 8%, more preferably at most 6%, more preferably at most 5%, more preferably at most 4%, more preferably at most 3%, even more preferably at most 2%, most preferably at most 1%, and even most preferably at most 0.5%by weight of other polypeptide material with which the polypeptide is natively or recombinantly associated.

[0046] It is, therefore, preferred that the substantially pure polypeptide is at least 92%pure, preferably at least 94%pure, more preferably at least 95%pure, more preferably at least 96%pure, more preferably at least 97%pure, more preferably at least 98%pure, even more preferably at least 99%pure, most preferably at least 99.5%pure by weight of the total polypeptide material present in the preparation. The polypeptide of the present invention is preferably in a substantially pure form (i.e., the preparation is essentially free of other polypeptide material with which it is natively or recombinantly associated) . This can be accomplished, for example by preparing the polypeptide by well-known recombinant methods or by classical purification methods.

[0047] Recombinant: The term "recombinant" is used in its conventional meaning to refer to the manipulation, e.g., cutting and rejoining, of nucleic acid sequences to form constellations different from those found in nature. The term recombinant refers to a cell, nucleic acid, polypeptide or vector that has been modified from its native state. Thus, for example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell, or express native genes at different levels or under different conditions than found in nature. The term “recombinant” is synonymous with “genetically modified” and “transgenic” .

[0048] Recover: The terms "recover" or “recovery” means the removal of a polypeptide from at least one fermentation broth component selected from the list of a cell, a nucleic acid, or other specified material, e.g., recovery of the polypeptide from the whole fermentation broth, or from the cell-free fermentation broth, by polypeptide crystal harvest, by filtration, e.g., depth filtration (by use of filter aids or packed filter medias, cloth filtration in chamber filters, rotary-drum filtration, drum filtration, rotary vacuum-drum filters, candle filters, horizontal leaf filters or similar, using sheed or pad filtration in framed or modular setups) or membrane filtration (using sheet filtration, module filtration, candle filtration, microfiltration, ultrafiltration in either cross flow, dynamic cross flow or dead end operation) , or by centrifugation (using decanter centrifuges, disc stack centrifuges, hyrdo cyclones or similar) , or by precipitating the polypeptide and using relevant solid-liquid separation methods to harvest the polypeptide from the broth media by use of classification separation by particle sizes. Recovery encompasses isolation and / or purification of the polypeptide.

[0049] Sequence difference: The term "sequence difference" means the percent of amino acid differences between a polypeptide and the polypeptide of a given SEQ ID NO: , and is calculated as follows (Method 1a) :

[0050] (Different Residues x 100)  /  (Length of given SEQ ID NO: )

[0051] wherein the different residues comprise any substitution, deletion, or insertion (e.g., an extension at the N-terminus and / or C-terminus) in the sequence.

[0052] Method 1b: For purposes of the present invention, the sequence identity between two polynucleotide sequences is determined as the output of “longest identity” using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, supra) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, supra) , preferably version 6.6.0 or later. The parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. In order for the Needle program to report the longest identity, the nobrief option must be specified in the command line. The output of Needle labeled “longest identity” is calculated as follows:

[0053] (Identical Deoxyribonucleotides x 100)  /  (Length of Alignment - Total Number of Gaps in Alignment)

[0054] Method 1c: For purposes of the present invention, the sequence identity between two amino acid sequences is determined as the output of “longest identity” using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) , preferably version 6.6.0 or later. The parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. In order for the Needle program to report the longest identity, the nobrief option must be specified in the command line. The output of Needle labeled “longest identity” is calculated as follows:

[0055] (Identical Residues x 100)  /  (Length of Alignment - Total Number of Gaps in Alignment)

[0056] Signal Peptide: A "signal peptide" is a sequence of amino acids attached to the N-terminal portion of a protein, which facilitates the secretion of the protein outside the cell. The mature form of an extracellular protein lacks the signal peptide, which is cleaved off during the secretion process.

[0057] Subsequence: The term “subsequence” means a polynucleotide having one or more nucleotides absent from the 5'a nd / or 3'end of a mature polypeptide coding sequence, wherein the subsequence encodes a fragment having collagenase activity.

[0058] Variant: The term “variant” means a synthetic polynucleotide encoding a polypeptide having collagenase activity, the polynucleotide comprising a man-made mutation, i.e., a nucleotide or codon substitution, insertion (including extension) , and / or deletion, at one or more positions and / or codons. A substitution means replacement of the nucleotide occupying a position with a different nucleotide; a deletion means removal of the nucleotide occupying a position; and an insertion means adding 1-5 nucleotides (e.g., 1-3 nucleotides) adjacent to and immediately following the nucleotide occupying a position.

[0059] Wild-type: The term "wild-type" in reference to an amino acid sequence or nucleic acid sequence means that the amino acid sequence or nucleic acid sequence is a native or naturally-occurring sequence. As used herein, the term "naturally-occurring" refers to anything (e.g., proteins, amino acids, or nucleic acid sequences) that is found in nature. Conversely, the term "non-naturally occurring" refers to anything that is not found in nature (e.g., recombinant nucleic acids and protein sequences produced in the laboratory or modification of the wild-type sequence) .Detailed Description of the Invention

[0060] The present invention is based on the surprising and inventive finding that expression of difficult-to-express proteins, namely bacterial collagenases, with codon optimized sequences increased yield when expressed in fungal host cells. The present invention finds that the codon optimized sequence of SEQ ID NO: 1 unexpectedly has excellent expression in a heterologous Pichia expression system.

[0061] Using the codon optimized DNA sequence of the invention, an improved yield of a bacterial collagenase is achieved.

[0062] Synthetic Polynucleotides

[0063] The present invention relates to synthetic polynucleotides encoding a collagenase as described herein.

[0064] Thus, in a 1st aspect the invention relates a synthetic polynucleotide encoding a collagenase, selected from the group consisting of:

[0065] (a) a polynucleotide having at least 80%sequence identity to SEQ ID NO: 1;

[0066] (b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0067] (c) a polynucleotide derived from the polynucleotide of (a) , or (b) , wherein the 3’ -and / or 5’ -end has been extended by addition of one or more nucleotides; and

[0068] (d) a fragment of the polynucleotide of (a) , (b) , or (c) .

[0069] In a preferred embodiment the collagenase has collagenase activity.

[0070] In one embodiment the polynucleotide has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 1.

[0071] In one embodiment, the polynucleotide has at least 85%, e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to the sequence of nucleotides 421 - 3120 of SEQ ID NO: 1 or nucleotides 454 - 3120 of SEQ ID NO: 1.

[0072] In one embodiment the polynucleotide is comprising, consisting essentially of, or consisting of SEQ ID NO: 1.

[0073] In one embodiment the polynucleotide is comprising, consisting essentially of, or consisting of the sequence of nucleotides 421-3120 of SEQ ID NO: 1.

[0074] In one embodiment the polynucleotide is comprising, consisting essentially of, or consisting of the sequence of nucleotides 454-3120 of SEQ ID NO: 1.

[0075] In one embodiment the synthetic polynucleotide is encoding a mature collagenase.

[0076] In one embodiment the polynucleotide is a fragment of SEQ ID NO: 1, wherein the fragment preferably contains at least 2700 nucleotides (e.g., nucleotides 421 to 3120 of SEQ ID NO: 1) , at least 2600 nucleotides (e.g., nucleotides 454 to 3120 of SEQ ID NO: 1) , at least 2500 nucleotides (e.g., nucleotides 521 to 3020 of SEQ ID NO: 1) , or at least 2400 nucleotides (e.g., nucleotides 571 to 2970 of SEQ ID NO: 1) , preferably wherein the fragment encodes a collagenase having collagenase activity.

[0077] In one embodiment the collagenase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 4.

[0078] In one embodiment the mature collagenase comprises or consists of an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 5.

[0079] In one embodiment the mature collagenase comprises or consists of an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 6.

[0080] In one embodiment the collagenase is comprising, consisting essentially of, or consisting of SEQ ID NO: 4.

[0081] In one embodiment the collagenase is comprising, consisting essentially of, or consisting of SEQ ID NO: 5.

[0082] In one embodiment the collagenase is comprising, consisting essentially of, or consisting of SEQ ID NO: 6.

[0083] In one embodiment sequence identity between two polypeptides is determined by Sequence Identity Determination Method 1a.

[0084] In one embodiment sequence identity between two polypeptides is determined by Sequence Identity Determination Method 1c.

[0085] In one embodiment sequence identity between two polynucleotides is determined by Sequence Identity Determination Method 1 b.

[0086] In one embodiment the polynucleotide is isolated.

[0087] In one embodiment the polynucleotide is purified.

[0088] The synthetic polynucleotide may be a synthetic cDNA, a synthetic DNA, a synthetic RNA, a synthetic mRNA, or a combination thereof. The synthetic polynucleotide may be derived from a native collagenase coding sequence from a strain of Paenibacillus azoreducens or a related organism and thus, for example, may be a variant of the synthetic polynucleotide sequence of the invention.

[0089] In an embodiment, the synthetic polynucleotide is a subsequence of the synthetic polynucleotide of the present invention encoding a fragment having collagenase activity. In an aspect, the subsequence contains at least 2700 nucleotides (e.g., nucleotides 421 to 3120 of SEQ ID NO: 1) , at least 2600 nucleotides (e.g., nucleotides 454 to 3120 of SEQ ID NO: 1) , at least 2500 nucleotides (e.g., nucleotides 521 to 3020 of SEQ ID NO: 1) , or at least 2400 nucleotides (e.g., nucleotides 571 to 2970 of SEQ ID NO: 1) , preferably wherein the fragment encodes a collagenase having collagenase activity.

[0090] In one embodiment the synthetic polynucleotide of the invention encoding the collagenase is derived from a bacterial cell, e.g., from Paenibacillus azoreducens.

[0091] The synthetic polynucleotide of the invention may also be mutated by introduction of nucleotide substitutions that do not result in a change in the amino acid sequence of the collagenase, but which further improve the yield of the collagenase in the host organism intended for production of the collagenase. For a general description of nucleotide substitution, see, e.g., Ford et al., 1991, Protein Expression and Purification 2: 95-107.

[0092] Nucleic Acid Constructs

[0093] In a 2nd aspect the present invention also relates to nucleic acid constructs comprising a synthetic polynucleotide of the 1st aspect, wherein the synthetic polynucleotide is operably linked to one or more control sequences that direct the expression of the coding sequence in a suitable host cell under conditions compatible with the control sequences.

[0094] In one embodiment the one or more control sequences comprises an inducible PICL1 promoter or a PICL1-based promoter.

[0095] In one embodiment the one or more control sequences comprises or consists of a PICL1 promoter with a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 7.

[0096] In one embodiment the control sequences comprise a polynucleotide region that encodes a signal peptide fused to the N-terminus of the collagenase which directs the collagenase into the secretory pathway of the host cell.

[0097] In one embodiment the signal peptide comprises or consists of a polypeptide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to any one of SEQ ID NOs: 9-12.

[0098] In one embodiment the signal peptide is encoded by a polynucleotide comprising or consisting of a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 8.

[0099] The synthetic polynucleotide may be manipulated in a variety of ways to provide for expression of the collagenase. Manipulation of the synthetic polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides utilizing recombinant DNA methods are well known in the art.

[0100] Promoters

[0101] The control sequence may be a promoter, a polynucleotide that is recognized by a host cell for expression of a synthetic polynucleotide of the present invention. The promoter contains transcriptional control sequences that mediate the expression of the collagenase. The promoter may be any polynucleotide that shows transcriptional activity in the host cell including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.

[0102] In an embodiment, the promoter is a heterologous promoter. In some embodiments, the promoter is a tandem promoter. In some embodiments, the promoter is an inducible promoter. An inducible promoter may be activated by the presence or absence of its inducer, or by the concentration levels of its inducer. In some embodiments, the promoter is a synthetic inducible promoter. A synthetic inducible promoter may comprise heterologous elements, such as operator or promoter elements that are derived from bacterial sources, such as for example the lac operon of Escherichia coli.

[0103] Suitable promoters for directing transcription of a polynucleotide of the present invention in a filamentous fungal host cell are promoters obtained from Aspergillus, Fusarium, Rhizomucor and Trichoderma cells, such as the promoters described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology.

[0104] For expression in a yeast host, examples of useful promoters are described by Smolke et al., 2018, “Synthetic Biology: Parts, Devices and Applications” (Chapter 6: Constitutive and Regulated Promoters in Yeast: How to Design and Make Use of Promoters in S. cerevisiae) , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology. Examples of promoters suitable for directing transcription of the polynucleotide of the present invention in a Komagataella phaffi yeast host cell include glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter, for constitutive expression, or the alcohol oxidase 1 (AOX1) promoter, for inducible expression by methanol.

[0105] Recently, expression systems of the fungal yeast K. phaffi have been developed to allow for flexible inducible expression of proteins of interest (see Zhu et al., Nucleic Acids Res. 2022, 50 (17) : 10187-10199; Wu et al, Engineering Microbiology, 2023, 3 (2023) : 100094, each incorporated by reference in its entirety herein) . In some embodiments, expression of a polynucleotide of the invention may be driven by an expression system, which may be a two-part system involving a sensor protein and a transcriptional activator (also referred to as a transcription factor) . In some embodiments, the expression system may be methanol-inducible, ethanol-inducible, or glucose-inducible. In some embodiments, the promoter is constitutive.

[0106] In some embodiments, the expression system utilizes a chimeric transcriptional activator LacI-Mit1AD (Liu et al., 2019, Metab Eng, 54 (2019) : 275-284, incorporated by reference in its entirety herein) driven by either the constitutive promoter PGAP in plasmid pGGLacI-Mit1AD or the glucose concentration responsive promoter PGAL in plasmid pGALLacI-Mit1AD (construction of each plasmid is described in Chinese patent publication No. CN108949869B (incorporated by reference in its entirety herein) ) .

[0107] In one embodiment the promoter is an inducible PICL1 promoter or a PICL1-based promoter. For example, the PICL1 promoter comprises or consists of SEQ ID NO: 7.

[0108] Terminators

[0109] The control sequence may also be a transcription terminator, which is recognized by a host cell to terminate transcription. The terminator is operably linked to the 3’ -terminus of the synthetic polynucleotide encoding the collagenase. Any terminator that is functional in the host cell may be used in the present invention.

[0110] Preferred terminators for filamentous fungal host cells may be obtained from Aspergillus or Trichoderma species, such as obtained from the genes for Aspergillus niger glucoamylase, Trichoderma reesei beta-glucosidase, Trichoderma reesei cellobiohydrolase I, and Trichoderma reesei endoglucanase I, such as the terminators described in Mukherjee et al., 2013, “Trichoderma: Biology and Applications” , and by Schmoll and  2016, “Gene Expression Systems in Fungi: Advancements and Applications” , Fungal Biology.

[0111] Preferred terminators for yeast host cells may be obtained from the genes for Saccharomyces cerevisiae enolase, S. cerevisiae cytochrome C (CYC1) , S. cerevisiae glyceraldehyde-3-phosphate dehydrogenase, and Komagataella phaffialcohol oxidase 1 (AOX1) . Other useful terminators for yeast host cells are described by Romanos et al., 1992, Yeast 8: 423-488 and by Kajiwara et al., 2018, Biotechnology Journal 13 (8) : 1700409.

[0112] mRNA Stabilizers

[0113] The control sequence may also be an mRNA stabilizer region downstream of a promoter and upstream of the coding sequence of a gene which increases expression of the gene.

[0114] Examples of suitable mRNA stabilizer regions are obtained from a Bacillus thuringiensis cryIIIA gene (WO 94 / 25612) and a Bacillus subtilis SP82 gene (Hue et al., 1995, J. Bacteriol. 177: 3465-3471) .

[0115] Examples of mRNA stabilizer regions for fungal cells are described in Geisberg et al., 2014, Cell 156 (4) : 812-824, and in Morozov et al., 2006, Eukaryotic Cell 5 (11) : 1838-1846.

[0116] Leader Sequences

[0117] The control sequence may also be a leader, a non-translated region of an mRNA that is important for translation by the host cell. The leader is operably linked to the 5’ -terminus of the synthetic polynucleotide encoding the collagenase. Any leader that is functional in the host cell may be used.

[0118] Suitable leaders for filamentous fungal host cells may be obtained from the genes for Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase.

[0119] Suitable leaders for yeast host cells may be obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1) , Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae alpha-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP) .

[0120] Polyadenylation Sequences

[0121] The control sequence may also be a polyadenylation sequence, a sequence operably linked to the 3’ -terminus of the synthetic polynucleotide which, when transcribed, is recognized by the host cell as a signal to add polyadenosine residues to transcribed mRNA. Any polyadenylation sequence that is functional in the host cell may be used.

[0122] Suitable polyadenylation sequences for filamentous fungal host cells are obtained from the genes for Aspergillus nidulans anthranilate synthase, Aspergillus niger glucoamylase, Aspergillus niger alpha-glucosidase, Aspergillus oryzae TAKA amylase, and Fusarium oxysporum trypsin-like protease.

[0123] Useful polyadenylation sequences for yeast host cells are described by Guo and Sherman, 1995, Mol. Cellular Biol. 15: 5983-5990.

[0124] Signal Peptides

[0125] The control sequence may also be a signal peptide coding region that encodes a signal peptide linked to the N-terminus of a collagenase and directs the collagenase into the cell’s secretory pathway. The 5’ -end of the coding sequence of the synthetic polynucleotide may inherently contain a signal peptide coding sequence naturally linked in translation reading frame with the segment of the coding sequence that encodes the polypeptide. Alternatively, the 5’ -end of the coding sequence may contain a signal peptide coding sequence that is heterologous to the coding sequence. A heterologous signal peptide coding sequence may be required where the coding sequence does not naturally contain a signal peptide coding sequence. Alternatively, a heterologous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance secretion of the collagenase. Any signal peptide coding sequence that directs the expressed collagenase into the secretory pathway of a host cell may be used.

[0126] Effective signal peptide coding sequences for filamentous fungal host cells are the signal peptide coding sequences obtained from the genes for Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Aspergillus oryzae TAKA amylase, Humicola insolens cellulase, Humicola insolens endoglucanase V, Humicola lanuginosa lipase, and Rhizomucor miehei aspartic proteinase, such as the signal peptide described by Xu et al., 2018, Biotechnology Letters 40: 949-955.

[0127] Useful signal peptides for yeast host cells are obtained from the genes for Saccharomyces cerevisiae alpha-factor and Saccharomyces cerevisiae invertase. Other useful signal peptide coding sequences are described by Romanos et al., 1992, supra.

[0128] Propeptides

[0129] The control sequence may also be a propeptide coding sequence that encodes a propeptide positioned at the N-terminus of a collagenase. The resultant polypeptide is known as a proenzyme or propolypeptide (or a zymogen in some cases) . A propolypeptide is generally inactive and can be converted to an active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. As a non-limiting example the propeptide is encoded by a polynucleotide comprising nucleotides 1-420 of any one of SEQ ID NO: 1, 2, or 3. As a non-limiting example the propeptide is encoded by a polynucleotide comprising nucleotides 1-453 of any one of SEQ ID NO: 1, 2, or 3.

[0130] Where both signal peptide and propeptide sequences are present, the propeptide sequence is positioned next to the N-terminus of a polypeptide and the signal peptide sequence is positioned next to the N-terminus of the propeptide sequence. Additionally, or alternatively, when both signal peptide and propeptide sequences are present, the polypeptide may comprise only a part of the signal peptide sequence and / or only a part of the propeptide sequence. Alternatively, the final or isolated polypeptide may comprise a mixture of mature polypeptides and polypeptides which comprise, either partly or in full length, a propeptide sequence and / or a signal peptide sequence.

[0131] In one example, the propeptide consists of amino acids corresponding to amino acids at position 1-140 of SEQ ID NO: 4. In another example, the propeptide consists of amino acids corresponding to amino acids at position 1-151 of SEQ ID NO: 4.

[0132] Expression Vectors

[0133] The present invention also relates to recombinant expression vectors comprising a synthetic polynucleotide of the present invention, a promoter, and transcriptional and translational stop signals. The various nucleotide and control sequences may be joined together to produce a recombinant expression vector that may include one or more convenient restriction sites to allow for insertion or substitution of the synthetic polynucleotide encoding the collagenase at such sites. Alternatively, the polynucleotide may be expressed by inserting the polynucleotide or a nucleic acid construct comprising the polynucleotide into an appropriate vector for expression. In creating the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.

[0134] The recombinant expression vector may be any vector (e.g., a plasmid or virus) that can be conveniently subjected to recombinant DNA procedures and can bring about expression of the polynucleotide. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may be a linear or closed circular plasmid.

[0135] The vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector may contain any means for assuring self-replication. Alternatively, the vector may be one that, when introduced into the host cell, is integrated into the genome and replicated together with the chromosome (s) into which it has been integrated. Furthermore, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell, or a transposon, may be used.

[0136] The vector preferably contains one or more selectable markers that permit easy selection of transformed, transfected, transduced, or the like cells. A selectable marker is a gene the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like.

[0137] The vector preferably contains at least one element that permits integration of the vector into the host cell's genome or autonomous replication of the vector in the cell independent of the genome.

[0138] For integration into the host cell genome, the vector may rely on the polynucleotide’s sequence encoding the polypeptide or any other element of the vector for integration into the genome by homologous recombination, such as homology-directed repair (HDR) , or non-homologous recombination, such as non-homologous end-joining (NHEJ) .

[0139] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. The origin of replication may be any plasmid replicator mediating autonomous replication that functions in a cell. The term “origin of replication” or “plasmid replicator” means a polynucleotide that enables a plasmid or vector to replicate in vivo.

[0140] More than one copy of a synthetic polynucleotide of the present invention may be inserted into a host cell to increase production of a collagenase. For example, 2 or 3 or 4 or 5 or more copies are inserted into a host cell. An increase in the copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the polynucleotide where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the polynucleotide, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0141] Host Cells

[0142] In a 3rd aspect, the present invention also relates to recombinant host cells, comprising a synthetic polynucleotide of the present invention operably linked to one or more control sequences that direct the production of a collagenase, e.g. a polynucleotide according to the 1st aspect of the invention. Alternatively the host cell is comprising a nucleic acid construct or expression vector according to the 2nd aspect of the invention.

[0143] A construct or vector comprising a polynucleotide is introduced into a host cell so that the construct or vector is maintained as a chromosomal integrant or as a self-replicating extra-chromosomal vector as described earlier. The choice of a host cell will to a large extent depend upon the gene encoding the polypeptide and its source. The collagenase can be native or heterologous to the recombinant host cell. Also, at least one of the one or more control sequences can be heterologous to the synthetic polynucleotide encoding the collagenase. The recombinant host cell may comprise a single copy, or at least two copies, e.g., three, four, five, or more copies of the synthetic polynucleotide of the present invention.

[0144] The host cell may be any microbial cell useful in the recombinant expression of a polynucleotide of the present invention, e.g., a fungal cell.

[0145] In one embodiment the collagenase is heterologous to the host cell.

[0146] In one embodiment at least one of the one or more control sequences is heterologous to the synthetic polynucleotide.

[0147] In one embodiment the host cell comprises at least two copies, e.g., three, four, or five, or more copies of the synthetic polynucleotide, nucleic acid construct or expression vector.

[0148] The host cell may be a fungal cell. “Fungi” as used herein includes the phyla Ascomycota, Basidiomycota, Chytridiomycota, and Zygomycota as well as the Oomycota and all mitosporic fungi (as defined by Hawksworth et al., 1995, In, Ainsworth and Bisby’s Dictionary of The Fungi, 8th edition, CAB International, University Press, Cambridge, UK) .

[0149] Fungal cells may be transformed by a process involving protoplast-mediated transformation, Agrobacterium-mediated transformation, electroporation, biolistic method and shock-wave-mediated transformation as reviewed by Li et al., 2017, Microbial Cell Factories 16: 168 and procedures described in EP 238023, Yelton et al., 1984, Proc. Natl. Acad. Sci. USA 81: 1470-1474, Christensen et al., 1988, Bio / Technology6: 1419-1422, and Lubertozzi and Keasling, 2009, Biotechn. Advances 27: 53-75. However, any method known in the art for introducing DNA into a fungal host cell can be used, and the DNA can be introduced as linearized or as circular polynucleotide.

[0150] The fungal host cell may be a filamentous fungal cell. “Filamentous fungi” include all filamentous forms of the subdivision Eumycota and Oomycota (as defined by Hawksworth et al., 1995, supra) . The filamentous fungi are generally characterized by a mycelial wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth is by hyphal elongation and carbon catabolism is obligately aerobic. In contrast, vegetative growth by yeasts such as Saccharomyces cerevisiae is by budding of a unicellular thallus and carbon catabolism may be fermentative.

[0151] The filamentous fungal host cell may be an Acremonium, Aspergillus, Aureobasidium, Bjerkandera, Ceriporiopsis, Chrysosporium, Coprinus, Coriolus, Cryptococcus, Filibasidium, Fusarium, Humicola, Magnaporthe, Mucor, Myceliophthora, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Phanerochaete, Phlebia, Piromyces, Pleurotus, Schizophyllum, Talaromyces, Thermoascus, Thielavia, Tolypocladium, Trametes, or Trichoderma cell. In a preferred embodiment, the filamentous fungal host cell is an Aspergillus, Trichoderma or Fusarium cell. In a further preferred embodiment, the filamentous fungal host cell is an Aspergillus niger, Aspergillus oryzae, Trichoderma reesei, or Fusarium venenatum cell.

[0152] For example, the filamentous fungal host cell may be an Aspergillus awamori, Aspergillus foetidus, Aspergillus fumigatus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, Aspergillus oryzae, Bjerkandera adusta, Ceriporiopsis aneirina, Ceriporiopsis caregiea, Ceriporiopsis gilvescens, Ceriporiopsis pannocinta, Ceriporiopsis rivulosa, Ceriporiopsis subrufa, Ceriporiopsis subvermispora, Chrysosporium inops, Chrysosporium keratinophilum, Chrysosporium lucknowense, Chrysosporium merdarium, Chrysosporium pannicola, Chrysosporium queenslandicum, Chrysosporium tropicum, Chrysosporium zonatum, Coprinus cinereus, Coriolus hirsutus, Fusarium bactridioides, Fusarium cerealis, Fusarium crookwellense, Fusarium culmorum, Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, Fusarium venenatum, Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora thermophila, Neurospora crassa, Penicillium purpurogenum, Phanerochaete chrysosporium, Phlebia radiata, Pleurotus eryngii, Talaromyces emersonii, Thielavia terrestris, Trametes villosa, Trametes versicolor, Trichoderma harzianum, Trichoderma koningii, Trichoderma longibrachiatum, Trichoderma reesei, or Trichoderma viride cell.

[0153] The fungal host cell may be a yeast cell. “Yeast” as used herein includes ascosporogenous yeast (Endomycetales) , basidiosporogenous yeast, and yeast belonging to the Fungi Imperfecti (Blastomycetes) . For purposes of this invention, yeast shall be defined as described in Biology and Activities of Yeast (Skinner, Passmore, and Davenport, editors, Soc. App. Bacteriol. Symposium Series No. 9, 1980) .

[0154] The yeast host cell may be a Candida, Hansenula, Kluyveromyces, Komagataella (formerly Pichia) , Ogataea, Blastobotrys, Saccharomyces, Schizosaccharomyces, or Yarrowia cell, such as a Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, Komagataella phaffii (formerly Pichia pastoris) , Ogataea polymorpha, Blastobotrys adeninivorans (formerly Arxula adeninivorans) , or Yarrowia lipolytica cell. In some embodiments, the yeast host cell is Komagataella kurtzmanii, K. mondaviorum, K. pastoris, K. phaffi, K. populi, K. pseudopastoris, or K. ulmi. In some embodiments, the yeast host cell is a Komagataella (Pichia) cell, e.g., a Komagataella phaffii (Pichia pastoris) cell.

[0155] Komagataella is a methylotrophic yeast and is used for production of recombinant proteins and enzymes. In some embodiments, a fusion polypeptide of the invention is expressed in a Komagataella cell using a constitutive expression system. Recently, expression systems of K. phaffi have been developed to allow for flexible inducible expression of proteins of interest (see Zhu et al., Nucleic Acids Res. 2022, 50 (17) : 10187-10199, incorporated by reference in its entirety herein) . In some embodiments, a fusion polypeptide of the invention is expressed in a Komagataella cell using a methanol-inducible expression system. In other embodiments, a fusion polypeptide of the invention is expressed in a Komagataella cell using an ethanol-inducible expression system. In some embodiments, a fusion polypeptide of the invention is expressed in a Komagataella cell using a glucose-inducible expression system.

[0156] In some embodiments, the expression system utilizes a chimeric transcriptional activator LacI-Mit1AD (Liu et al., 2019, Metab Eng, 54 (2019) : 275-284) driven by either the constitutive promoter PGAP in plasmid pGGLacI-Mit1AD or the glucose concentration responsive promoter PGAL in plasmid pGALLacI-Mit1AD (construction of both plasmids are described in Chinese patent publication No. CN108949869B (incorporated by reference in its entirety herein) ) . The LacI-Mit1AD transcriptional activator activates a downstream promoter for downstream target gene expression.

[0157] In one embodiment the host cell is isolated.

[0158] In one embodiment the host cell is purified.

[0159] Methods of Production

[0160] The present invention also relates to methods of producing a collagenase, comprising (a) cultivating a cell according to the 3rd aspect under conditions conducive for production of the collagenase; and optionally, (b) recovering the collagenase.

[0161] In one aspect, the cell is a Pichia cell. In another aspect, the cell is a Pichia pastoris cell.

[0162] The host cell is cultivated in a nutrient medium suitable for production of the collagenase using methods known in the art. For example, the cell may be cultivated by shake flask cultivation, or small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state, and / or microcarrier-based fermentations) in laboratory or industrial fermentors in a suitable medium and under conditions allowing the polypeptide to be expressed and / or isolated. Suitable media are available from commercial suppliers or may be prepared according to published compositions (e.g., in catalogues of the American Type Culture Collection) . If the polypeptide (collagenase) is secreted into the nutrient medium, the polypeptide can be recovered directly from the medium. If the polypeptide is not secreted, it can be recovered from cell lysates.

[0163] The polypeptide may be detected using methods known in the art that are specific for the polypeptide, including, but not limited to, the use of specific antibodies, formation of an enzyme product, disappearance of an enzyme substrate, or an assay determining the relative or specific activity of the polypeptide.

[0164] The polypeptide may be recovered from the medium using methods known in the art, including, but not limited to, collection, centrifugation, filtration, extraction, spray-drying, evaporation, or precipitation. In one aspect, a whole fermentation broth comprising the polypeptide is recovered. In another aspect, a cell-free fermentation broth comprising the polypeptide is recovered.

[0165] The polypeptide may be purified by a variety of procedures known in the art to obtain substantially pure polypeptides and / or polypeptide fragments (see, e.g., Wingfield, 2015, Current Protocols in Protein Science; 80 (1) : 6.1.1-6.1.35; Labrou, 2014, Protein Downstream Processing, 1129: 3-10) .

[0166] In an alternative aspect, the polypeptide is not recovered.

[0167] Generation of synthetic polynucleotides

[0168] A synthetic polynucleotide of the present invention may be a sequence obtained from microorganisms of any genus, optimized for expression in a host cell of choice. In other words, microorganisms of any genus may be a source of a native polynucleotide encoding a collagenase, which native polynucleotide is then altered and optimized towards expression in a recombinant host. For purposes of the present invention, the term “obtained from” as used herein in connection with a given source shall mean that the synthetic polynucleotide is a variant of the native sequence obtained from that microorganism.

[0169] In one aspect, the polynucleotide is an optimized sequence of a collagenase coding sequence obtained from a bacterial species.

[0170] In one aspect, the polynucleotide is an optimized sequence of a collagenase coding sequence obtained from a Paenibacillus spp., e.g., Paenibacillus azoreducens.

[0171] It will be understood that for the aforementioned species, the invention encompasses both the perfect and imperfect states, and other taxonomic equivalents, e.g., anamorphs, regardless of the species name by which they are known. Those skilled in the art will readily recognize the identity of appropriate equivalents.

[0172] The native polynucleotides to be optimized may be identified and obtained from other sources including microorganisms isolated from nature (e.g., soil, composts, water, etc. ) or DNA samples obtained directly from natural materials (e.g., soil, composts, water, etc. ) using the above-mentioned probes. Techniques for isolating microorganisms and DNA directly from natural habitats are well known in the art. A polynucleotide encoding the collagenase may then be obtained by similarly screening a genomic DNA or cDNA library of another microorganism or mixed DNA sample. Once a polynucleotide encoding a collagenase has been detected with the probe (s) , the polynucleotide can be isolated or cloned by utilizing techniques that are known to those of ordinary skill in the art (see, e.g., Davis et al., 2012, Basic Methods in Molecular Biology, Elsevier) .

[0173] Fermentation Broth Formulations or Cell Compositions

[0174] The present invention also relates to a fermentation broth formulation, a fermentation broth, a composition or a cell composition comprising a synthetic polynucleotide of the present invention. The fermentation broth formulation or the cell composition further comprises additional ingredients used in the fermentation process, such as, for example, cells (including, the host cells containing the gene encoding a collagenase which are used to produce the collagenase) , cell debris, biomass, fermentation media and / or fermentation products. In some embodiments, the composition is a cell-killed whole broth containing organic acid (s) , killed cells and / or cell debris, and culture medium.

[0175] The term "fermentation broth" as used herein refers to a preparation produced by cellular fermentation that undergoes no or minimal recovery and / or purification. For example, fermentation broths are produced when microbial cultures are grown to saturation, incubated under carbon-limiting conditions to allow protein synthesis (e.g., expression of enzymes by host cells) and secretion into cell culture medium. The fermentation broth can contain unfractionated or fractionated contents of the fermentation materials derived at the end of the fermentation. Typically, the fermentation broth is unfractionated and comprises the spent culture medium and cell debris present after the microbial cells (e.g., fungal cells) are removed, e.g., by centrifugation. In some embodiments, the fermentation broth contains spent cell culture medium, extracellular enzymes, and viable and / or nonviable microbial cells.

[0176] In some embodiments, the fermentation broth formulation or the cell composition comprises a first organic acid component comprising at least one 1-5 carbon organic acid and / or a salt thereof and a second organic acid component comprising at least one 6 or more carbon organic acid and / or a salt thereof. In some embodiments, the first organic acid component is acetic acid, formic acid, propionic acid, a salt thereof, or a mixture of two or more of the foregoing and the second organic acid component is benzoic acid, cyclohexanecarboxylic acid, 4-methylvaleric acid, phenylacetic acid, a salt thereof, or a mixture of two or more of the foregoing.

[0177] In one aspect, the composition contains an organic acid (s) , and optionally further contains killed cells and / or cell debris. In some embodiments, the killed cells and / or cell debris are removed from a cell-killed whole broth to provide a composition that is free of these components.

[0178] The fermentation broth formulation or cell composition may further comprise a preservative and / or anti-microbial (e.g., bacteriostatic) agent, including, but not limited to, sorbitol, sodium chloride, potassium sorbate, and others known in the art.

[0179] The cell-killed whole broth or cell composition may contain the unfractionated contents of the fermentation materials derived at the end of the fermentation. Typically, the cell-killed whole broth or cell composition contains the spent culture medium and cell debris present after the microbial cells (e.g., Pichia cells) are grown to saturation, incubated under carbon-limiting conditions to allow protein synthesis. In some embodiments, the cell-killed whole broth or cell composition contains the spent cell culture medium, extracellular enzymes, and killed filamentous fungal cells. In some embodiments, the microbial cells present in the cell-killed whole broth or cell composition can be permeabilized and / or lysed using methods known in the art.

[0180] A whole broth or cell composition as described herein is typically a liquid, but may contain insoluble components, such as killed cells, cell debris, culture media components, and / or insoluble enzyme (s) . In some embodiments, insoluble components may be removed to provide a clarified liquid composition.

[0181] The whole broth formulations and cell compositions of the present invention may be produced by a method described in WO 90 / 15861 or WO 2010 / 096673.

[0182] The present invention is further described by the following examples that should not be construed as limiting the scope of the invention.

[0183] Examples

[0184] Overview of Sequences SEQ ID NO: 1               synthetic DNA sequence of design 1 SEQ ID NO: 2               synthetic DNA sequence of design 2 SEQ ID NO: 3               wt collagenase DNA sequence of Paenibacillus azoreducens SEQ ID NO: 4               collagenase polypeptide (after removal of the signal peptide) SEQ ID NOs: 5 and 6        collagenase polypeptide (after removal of the signal peptide and different pro-peptide region) SEQ ID NO: 7               PICL1 promoter sequence SEQ ID NO: 8               α-mating factor signal peptide DNA sequence SEQ ID NO: 9               α-mating factor signal peptide amino acid sequence SEQ ID NOs: 10-12          alternative signal peptide amino acid sequences for collagenase expression

[0185] Example 1: Generation of several synthetic designs for collagenase expression

[0186] Several synthetic sequences were generated aiming to increase collagenase yield during expression in Komagataella phaffii, without changing the amino acid sequence of the collagenase polypeptide (SEQ ID NO: 4) and its matured form (SEQ ID NO: 6) . Collagenase DNA design 1 (SEQ ID NO: 1) and design 2 (SEQ ID NO: 2) encode the same bacterial collagenase polypeptide and both synthetic designs were selected for evaluation.

[0187] The starting sequence was the native collagenase coding sequence from Paenibacillus azoreducens (SEQ ID NO: 3) based on which synthetic designs 1-2 were generated for optimized expression in Komagataella phaffii.

[0188] Table 1. Percent-identity matrix (PIM) with sequence homology in %.

[0189] Table 1 shows a percent identity matrix (PIM) comparing the sequence homologies of the wildtype (wt) and synthetic sequences to another. As shown in Table 1, design 1 and design 2 have a low sequence homology to another with only with 79.46%. Both synthetic designs 1 and 2 show only around 75-77%sequence homology to the wildtype coding sequence, indicating that significant codon changes are present in the synthetic designs.

[0190] In other words, the percentage of recoded nucleotides relative to the wildtype sequence is 25.16%for design 1, and 22.80%for design 2.

[0191] Example 2: Design 1 showed significantly increased collagenase expression

[0192] To test the collagenase expression of the 2 synthetic designs in Komagataella phaffii cells (formerly named Pichia pastoris) , expression cassettes were prepared. Parent strain was K. phaffii GS115 (InvitrogenTM, ThermoFisher Scientific Inc. ) .

[0193] Each codon optimized sequence (design 1 = SEQ ID NO: 1; and design 2 = SEQ ID NO: 2) was fused to α-mating factor signal peptide coding sequence (SEQ ID NO: 8; encoding the signal peptide of SEQ ID NO: 9) located upstream of the collagenase coding sequence and operably linked thereto. Each cassette comprised the PICL1 promoter coding sequence with the sequence shown in SEQ ID NO: 7 operably linked to the signal peptide coding sequence. Synthetic gene cloning was made by seamless assembly.

[0194] K. phaffii cell transformation was made according to the procedure described in CN108949869B. For each design, three strains were generated with stable integration of the expression cassette in the genome. Following fermentation in 24-well plate and separation of the cells from surrounding media, the supernatant was collected and analyzed via SDS-PAGE (Fig. 1) for collagenase expression analysis. Procedure for 24-well plate cultivation is described below.

[0195] As shown in Fig. 1, the three strains with design 1 (SEQ ID NO: 1) showed significantly improved collagenase expression (rectangle enclosed) when compared to the three strains using design 2 (SEQ ID NO: 2) . The unmatured collagenase is present at ~112 kDa (marker shown to the very left of Fig. 1) , whereas the matured collagenase is present at ~95 kDa. Surprisingly, out of the two synthetic and optimized DNA sequences only design 1 could increase collagenase expression significantly whereas design 2 gave little to no expression.

[0196] 24-well plate assays for collagenase expression

[0197] Seed medium: YPD medium: 2%peptone, 1%yeast powder, add 4 ml of 50%glucose solution per 100 ml YPD medium.

[0198] Fermentation medium: BMDY media: 2%peptone、1%yeast extract、1.34%YNB、100 mmol / L pH6.0 Potassium phosphate buffer (2.3 mg / ml K2HPO4, 11.8 mg / ml KH2PO4) , 1%glucose.

[0199] Seed culture: 20 μL glycerol strain stock was inoculated to 5 ml YPD medium, 30℃, 200 rpm for 2 days. Then 50 μL strain broth from first inoculation was transferred into 10 mL of YPD medium for cultivation at 30℃ and 200 rpm for 16~20 h. It was then inoculated into 24-well plate with 3 mL of BMDY medium, and cultivated at 30℃, 200 rpm for 96h. Initial OD600 of fermentation broth from each well after inoculation was 1.60 μL of 50%glucose stock was added to each well every 24 h, until 96 h for sampling and analysis.

[0200] The invention described and claimed herein is not to be limited in scope by the specific aspects herein disclosed, since these aspects are intended as illustrations of several aspects of the invention. Any equivalent aspects are intended to be within the scope of this invention. Indeed, various modifications of the invention in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. In the case of conflict, the present disclosure including definitions will control.

[0201] The invention is further defined by the following numbered paragraphs:

[0202] 1. A synthetic polynucleotide encoding a collagenase, selected from the group consisting of:

[0203] (a) a polynucleotide having at least 80%sequence identity to SEQ ID NO: 1;

[0204] (b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0205] (c) a polynucleotide derived from the polynucleotide of (a) , or (b) , wherein the 3’ -and / or 5’ -end has been extended by addition of one or more nucleotides; and

[0206] (d) a fragment of the polynucleotide of (a) , (b) , or (c) .

[0207] 2. The polynucleotide according to paragraph 1, wherein the collagenase has collagenase activity.

[0208] 3. The polynucleotide of any one of paragraphs 1-2, wherein the polynucleotide has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 1.

[0209] 4. The polynucleotide of any one of paragraphs 1-2, wherein the polynucleotide has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to the sequence of nucleotides 421 -3120 of SEQ ID NO: 1.

[0210] 5. The polynucleotide of any one of paragraphs 1-2, wherein the polynucleotide has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to the sequence of nucleotides 454 -3120 of SEQ ID NO: 1.

[0211] 6. The polynucleotide of any one of the preceding paragraphs, the polynucleotide comprising, consisting essentially of, or consisting of SEQ ID NO: 1.

[0212] 7. The polynucleotide of any one of the preceding paragraphs, wherein the polynucleotide is comprising, consisting essentially of, or consisting of the sequence of nucleotides 421 -3120 of SEQ ID NO: 1.

[0213] 8. The polynucleotide of any one of the preceding paragraphs, wherein the polynucleotide is comprising, consisting essentially of, or consisting of the sequence of nucleotides 454 -3120 of SEQ ID NO: 1.

[0214] 8. The polynucleotide of any one of the preceding paragraphs, wherein the synthetic polynucleotide is encoding a mature collagenase.

[0215] 9. The polynucleotide of any one of the preceding paragraphs, wherein the mature collagenase comprises or consists of a polypeptide having at least 80%, e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 5 or SEQ ID NO: 6.

[0216] 10. The polynucleotide of any one of the preceding paragraphs, which is a fragment of SEQ ID NO: 1, wherein the fragment contains at least 2700 nucleotides (e.g., nucleotides 421 to 3120 of SEQ ID NO: 1) , at least 2600 nucleotides (e.g., nucleotides 454 to 3120 of SEQ ID NO: 1) , at least 2500 nucleotides (e.g., nucleotides 521 to 3020 of SEQ ID NO: 1) , or at least 2400 nucleotides (e.g., nucleotides 571 to 2970 of SEQ ID NO: 1) , preferably wherein the fragment encodes a collagenase having collagenase activity.

[0217] 11. The polynucleotide of any one of the preceding paragraphs, wherein the collagenase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 4.

[0218] 12. The polynucleotide of any one of the preceding paragraphs, wherein the collagenase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 5.

[0219] 13. The polynucleotide of any one of the preceding paragraphs, wherein the collagenase comprises, consists essentially of, or consists of the amino acid sequence of SEQ ID NO: 5.

[0220] 14. The polynucleotide of any one of the preceding paragraphs, wherein the collagenase comprises or consists of a polypeptide having at least 80%, e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 6.

[0221] 15. The polynucleotide of any one of the preceding paragraphs, wherein the collagenase comprises, consists essentially of, or consists of the amino acid sequence of SEQ ID NO: 6.

[0222] 16. The polynucleotide of any one of the preceding paragraphs, wherein polynucleotide sequence identity is determined by Sequence Identity Determination Method 1 b.

[0223] 17. The polynucleotide of any one of the preceding paragraphs, which is isolated.

[0224] 18. The polynucleotide of any one of the preceding paragraphs, which is purified.

[0225] 19. A nucleic acid construct or expression vector comprising the polynucleotide of any one of the preceding paragraphs operably linked to one or more control sequences that direct the production of the collagenase in an expression host.

[0226] 20. The nucleic acid construct or expression vector according to paragraph 19, wherein the one or more control sequence comprises an inducible PICL1 promoter or a PICL1-based promoter.

[0227] 21. The nucleic acid construct or expression vector according to any one of the preceding paragraphs, wherein the one or more control sequences comprises or consists of a promoter with a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 7.

[0228] 22. The nucleic acid construct or expression vector according to any one the preceding paragraphs, wherein the control sequences comprise a polynucleotide region that encodes a signal peptide fused to the N-terminus of the collagenase which directs the collagenase into the secretory pathway of the host cell.

[0229] 23. The nucleic acid construct or expression vector according to any one of the preceding paragraphs, wherein the signal peptide comprises or consists of a polypeptide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to any one of the the sequences of SEQ ID NO: 9 -12.

[0230] 24. A recombinant host cell comprising the nucleic acid construct or expression vector of any one of the preceding paragraphs.

[0231] 25. The host cell according to paragraph 24, wherein the collagenase is heterologous to the host cell.

[0232] 26. The host cell of any one of paragraph 24-25, wherein at least one of the one or more control sequences is heterologous to the synthetic polynucleotide.

[0233] 27. The host cell of any one of paragraphs 24-26, which comprises at least two copies, e.g., three, four, or five, or more copies of the polynucleotide, nucleic acid construct or expression vector according to any one of paragraphs 1-23.

[0234] 28. The host cell of any one of paragraphs 24-27, which is a fungal recombinant host cell.

[0235] 29. The host cell of any one of paragraphs 24-28, which is a Candida, Hansenula, Kluyveromyces, Komagataella (formally Pichia) , Ogataea, Blastobotrys, Saccharomyces, Schizosaccharomyces, or Yarrowia cell, such as a Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, Komagataella phaffii (formerly Pichia pastoris) , Ogataea polymorpha, Blastobotrys adeninivorans (formerly Arxula adeninivorans) , or Yarrowia lipolytica cell.

[0236] 30. The host cell of any one of paragraphs 24-29, which is a yeast cell, such as Komagataella kurtzmanii, K. mondaviorum, Komagataella pastoris, Komagataella phaffi, Komagataella populi, Komagataella pseudopastoris, or Komagataella ulmi cell.

[0237] 31. The host cell of any one of paragraphs 24-30, which is a Komagataella (Pichia) cell, e.g., a Komagataella phaffii (Pichia pastoris) cell.

[0238] 32. The host cell of any one of paragraphs 24-31, which is isolated.

[0239] 33. The host cell of any one of paragraphs 24-32, which is purified.

[0240] 34. A method of producing a collagenase, comprising cultivating the recombinant host cell of any one of paragraphs 24-32 under conditions conducive for production of the collagenase; and optionally recovering the collagenase.

[0241] 35. The method of paragraph 34, comprising contacting the collagenase with a protease to provide a matured collagenase.

[0242] 36. A composition, cell composition or fermentation broth comprising the polynucleotide of any one of paragraphs 1-18, a nucleic acid construct or expression vector of any one of paragraphs 19-23, and / or the cell of any one of paragraphs 24-35.

Claims

A synthetic polynucleotide encoding a collagenase, selected from the group consisting of:(a) a polynucleotide having at least 80%sequence identity to SEQ ID NO: 1,(b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;(c) a polynucleotide derived from the polynucleotide of (a) , or (b) , wherein the 3’ -and / or 5’ -end has been extended by addition of one or more nucleotides; and(d) a fragment of the polynucleotide of (a) , (b) , or (c) .The polynucleotide of claim 1, wherein the polynucleotide has at least 85%, e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 1.The polynucleotide of claim 1 or 2, wherein the polynucleotide has at least 85%, e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to the sequence of nucleotides 421 -3120 of SEQ ID NO: 1.The polynucleotide of any one of claims 1 to 3, wherein the polynucleotide has at least 85%, e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to the sequence of nucleotides 454 -3120 of SEQ ID NO: 1.The polynucleotide of any one of claims 1 to 4, the polynucleotide comprising, consisting essentially of, or consisting of SEQ ID NO: 1, or the sequence of nucleotides 421-3120 of SEQ ID NO: 1, or the sequence of nucleotides 454 -3120 of SEQ ID NO: 1.The polynucleotide of any one of claims 1 to 5, wherein the collagenase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to any one of SEQ ID NO: 4 or SEQ ID NO: 5 or SEQ ID NO: 6.A nucleic acid construct or expression vector comprising the polynucleotide of any one of the claims 1 to 6 operably linked to one or more control sequences that direct the production of the collagenase in an expression host.A recombinant host cell comprising the nucleic acid construct or expression vector of claim 7.The host cell of claim 8, which is a Candida, Hansenula, Kluyveromyces, Komagataella (formally Pichia) , Ogataea, Blastobotrys, Saccharomyces, Schizosaccharomyces, or Yarrowia cell, such as a Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, Komagataella phaffii (formerly Pichia pastoris) , Ogataea polymorpha, Blastobotrys adeninivorans (formerly Arxula adeninivorans) , or Yarrowia lipolytica cell.The host cell of claim 8 or 9, which is a yeast cell, such as Komagataella kurtzmanii, Komagataella mondaviorum, Komagataella pastoris, Komagataella phaffi, Komagataella populi, Komagataella pseudopastoris, or Komagataella ulmi cell.The host cell of claim 10, which is a Komagataella phaffii (Pichia pastoris) cell.The host cell according to any one of claims 8 to 11, wherein the collagenase is heterologous to the host cell.A method of producing a collagenase, comprising cultivating the recombinant host cell of any one of claims 8 to 12 under conditions conducive for production of the collagenase; and optionally recovering the collagenase.A composition, cell composition or fermentation broth comprising the polynucleotide of any one of claims 1 to 6, a nucleic acid construct or expression vector of claim 7, and / or the host cell of any one of claims 8 to 12.