Codon optimized deamidase expression

Codon-optimized deamidase-coding sequences address the limitations of existing enzyme productivity methods by achieving a 69% increase in deamidase yield, enhancing enzyme expression through targeted synthetic DNA modifications.

WO2026073574A1PCT designated stage Publication Date: 2026-04-09NOVOZYMES AS
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-02
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing methods for improving enzyme productivity, such as genetic manipulation and codon-usage optimization, have limitations, and there is a need for further enhancements in enzyme yield and expression.

Method used

The development of codon-optimized deamidase-coding sequences, particularly design 6, which significantly increases deamidase expression by up to 69% compared to native sequences, through synthetic DNA sequences with specific alterations and extensions.

Benefits of technology

The codon-optimized deamidase-coding sequences enhance deamidase yield by 69%, providing a substantial improvement in enzyme production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000023_0001
    Figure IMGF000023_0001
  • Figure IMGF000023_0002
    Figure IMGF000023_0002
  • Figure IMGF000024_0001
    Figure IMGF000024_0001
Patent Text Reader

Abstract

The present invention relates to synthetic polynucleotides encoding a deamidase, and to nucleic acid constructs, vectors, and host cells comprising the synthetic polynucleotides as well as methods of producing the deamidase.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CODON OPTIMIZED DEAMIDASE EXPRESSION

[0002] Reference to a Sequence Listing

[0003] This application contains a Sequence Listing in computer readable form, which is incorporated herein by reference.

[0004] Background of the Invention

[0005] Field of the Invention

[0006] The present invention relates to synthetic polynucleotides encoding a deamidase, and to nucleic acid constructs, vectors, and host cells comprising the synthetic polynucleotides as well as methods of producing the deamidase.

[0007] Description of the Related Art

[0008] In the highly competitive industrial manufacture of enzymes it is of vital importance to constantly improve yield or productivity. Genetic manipulation or engineering has been put to use for this purpose for many years, where genes encoding polypeptides of interest have been placed under the transcriptional control of heterologous or synthetic promoters, they have been expressed with heterologous signal peptides in various host cells and they have been integrated in the host cell genomes in multiple copies in order to achieve so-called mRNA-saturation.

[0009] Another well-known technique to increase enzyme productivity has been to optimize the codon-usage in enzyme-encoding DNA sequence based on that of the host cell intended for its expression and based on various theoretical mRNA melting point or tertiary structure calculations.

[0010] Even so, it remains of significant interest to identify new ways to improve the expression of an enzyme of interest. Due to the highly competitive environment in the enzyme manufacture industry, even minor improvements are desirable.

[0011] Summary of the Invention

[0012] The present invention provides means and methods to increase recombinant deamidase production. Testing multiple codon-optimized deamidase-coding sequences the present inventors identified a synthetic DNA sequence which significantly increased deamidase expression compared to other synthetic sequences and compared to the wildtype deamidase coding sequence. Surprisingly, only one of the generated DNA sequence designs (design 6) achieved a significant yield increase. For design 6 the deamidase yield was increased by circa 69% compared to the native sequence, which result was totally unexpected.

[0013] Accordingly, in a 1staspect the present invention relates to synthetic polynucleotides encoding a deamidase, selected from the group consisting of:

[0014] (a) a polynucleotide having at least 80% sequence identity to SEQ ID NO:1 , (b) a polynucleotide derived from SEQ ID NO:1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0015] (c) a polynucleotide derived from the polynucleotide of (a), or (b), wherein the 3’- and / or 5’- end has been extended by addition of one or more nucleotides; and

[0016] (d) a fragment of the polynucleotide of (a), (b), or (c).

[0017] In a 2ndaspect the invention relates to nucleic acid constructs of expression vectors comprising the polynucleotide of the 1staspect.

[0018] In a 3rdaspect the invention relates to a recombinant host cell comprising the nucleic acid construct or expression vector of the 2ndaspect.

[0019] In a 4thaspect the invention relates to a composition, cell composition or fermentation broth comprising the polynucleotide of the 1staspect and / or the cell of the 3rdaspect.

[0020] In a 5thaspect the invention relates to methods of producing a deamidase comprising cultivating the host cell of the 3rdaspect under conditions conducive for production of the deamidase.

[0021] Brief Description of the Drawings

[0022] Figure 1 shows a phylogenetic tree for the 10 sequence designs.

[0023] Definitions

[0024] In accordance with this detailed description, the following definitions apply. Note that the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.

[0025] Unless defined otherwise or clearly indicated by context, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0026] Deamidase: The term “deamidase” means a protein-glutamine glutaminase (also known as glutaminylpeptide glutaminase, or protein deamidase), as described in EC 3.5.1.44, which catalyzes the hydrolysis of the gamma-amide of glutamine substituted at the carboxyl position or both the alpha-amino and carboxyl positions, e.g., L-glutaminylglycine and L-phenylalanyl-L- glutaminylglycine. Thus, deamidases can deamidate glutamine residues in proteins to glutamate residues and are also referred to as protein glutamine deamidase. Deamidases comprise a Cys- His-Asp catalytic triad (e.g., Cys-156, His-197, and Asp-217, as shown in Hashizume et al. “Crystal structures of protein glutaminase and its pro forms converted into enzyme-substrate complex”, Journal of Biological Chemistry, vol. 286, no. 44, pp. 38691-38702) and belong to the InterPro entry IPR041325. In a preferred embodiment, the deamidases of the present invention belong to PFAM domain PF18626.

[0027] Deamidases are catalytic proteins (enzymes), and the term “active (deamidase) enzyme protein” is defined herein as the amount of catalytic protein(s), which exhibits deamidase activity. This can be determined using an activity based analytical enzyme assay. This technique is well-known in the art.

[0028] Deamidase activity can be determined using the assay described in Example 1 of WO23170177 (Novozymes A / S), referred to as “Method 1”. The activity assay of Method 1 consists of two separate de-coupled parts: (1) an enzymatic step wherein ammonia is formed by the catalytic action of the protein deamidase; and (2) a non-enzymatic detection step, wherein the ammonia formed in step (1) is derivatized to a blue indophenol compound with an absorption maximum at 630 nm. The amount of enzyme producing 1 pmol ammonia per minute at 37°C is defined as 1 unit (given in Indophenol Assay Unit: IPA(U)). The activity may be determined relative to a standard of declared strength.

[0029] Deamidase activity may also be measured by deamidating a glutamine substrate (for example Cbz-GIn-Gly) and generate ammonia in the process, herewith referred to as “Method 2”. The ammonia is used as substrate for a glutamate dehydrogenase in combination with a- ketoglutarate to produce glutamate. This latter enzymatic reaction requires NADH as a coenzyme. The depletion of NADH can be followed by kinetic absorbance measurement at 340 nm and is directly proportional to the deamidase activity. The reaction is carried out at pH 7 and 37 degrees centigrade. cDNA: The term "cDNA" means a DNA molecule that can be prepared by reverse transcription from a mature, spliced, mRNA molecule obtained from a eukaryotic or prokaryotic cell. cDNA lacks intron sequences that may be present in the corresponding genomic DNA. The initial, primary RNA transcript is a precursor to mRNA that is processed through a series of steps, including splicing, before appearing as mature spliced mRNA.

[0030] Coding sequence: The term “coding sequence” means a polynucleotide, which directly specifies the amino acid sequence of a polypeptide. The boundaries of the coding sequence are generally determined by an open reading frame, which begins with a start codon, such as ATG, GTG, or TTG, and ends with a stop codon, such as TAA, TAG, or TGA. The coding sequence may be a genomic DNA, cDNA, synthetic DNA, or a combination thereof.

[0031] Control sequences: The term “control sequences” means nucleic acid sequences involved in regulation of expression of a polynucleotide in a specific organism or in vitro. Each control sequence may be native ( / .e., from the same gene) or heterologous ( / .e., from a different gene) to the polynucleotide encoding the polypeptide, and native or heterologous to each other. Such control sequences include, but are not limited to leader, polyadenylation, prepropeptide, propeptide, signal peptide, promoter, terminator, enhancer, and transcription or translation initiator and terminator sequences. At a minimum, the control sequences include a promoter, and transcriptional and translational stop signals. The control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the polynucleotide encoding a polypeptide.

[0032] Expression: The term “expression” means any step involved in the production of a polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0033] Expression vector: An "expression vector" refers to a linear or circular DNA construct comprising a DNA sequence encoding a polypeptide, which coding sequence is operably linked to a suitable control sequence capable of effecting expression of the DNA in a suitable host. Such control sequences may include a promoter to effect transcription, an optional operator sequence to control transcription, a sequence encoding suitable ribosome binding sites on the mRNA, enhancers and sequences which control termination of transcription and translation.

[0034] Extension: The term “extension” means an addition of one or more amino acids to the amino and / or carboxyl terminus of a polypeptide, wherein the “extended” polypeptide has deamidase activity.

[0035] Fragment: The term “fragment” means a polypeptide having one or more amino acids absent from the amino and / or carboxyl terminus of the mature polypeptide wherein the fragment has deamidase activity.

[0036] Fusion polypeptide: The term “fusion polypeptide” is a polypeptide in which one polypeptide is fused at the N-terminus and / or the C-terminus of a polypeptide of the present invention. A fusion polypeptide is produced by fusing a polynucleotide encoding another polypeptide to a polynucleotide of the present invention, or by fusing two or more polynucleotides of the present invention together. Techniques for producing fusion polypeptides are known in the art, and include ligating the coding sequences encoding the polypeptides so that they are in frame and that expression of the fusion polypeptide is under control of the same promoter(s) and terminator. Fusion polypeptides may also be constructed using intein technology in which fusion polypeptides are created post-translationally (Cooper et al., 1993, EMBO J. 12: 2575-2583; Dawson et al., 1994, Science 266: 776-779). A fusion polypeptide can further comprise a cleavage site between the two polypeptides. Upon secretion of the fusion protein, the site is cleaved releasing the two polypeptides. Examples of cleavage sites include, but are not limited to, the sites disclosed in Martin et al., 2003, J. Ind. Microbiol. Biotechnol. 3: 568-576; Svetina et al., 2000, J. Biotechnol. 7Q: 245-251 ; Rasmussen-Wilson et al., 1997, Appl. Environ. Microbiol. 63: 3488-3493; Ward et al., 1995, Biotechnology 13: 498-503; and Contreras et al., 1991 , Biotechnology 9: 378-381 ; Eaton etal., 1986, Biochemistry 25: 505-512; Collins-Racie etal., 1995, Biotechnology 13: 982-987; Carter et al., 1989, Proteins: Structure, Function, and Genetics 6: 240-248; and Stevens, 2003, Drug Discovery World 4: 35-48. Heterologous: The term heterologous means, with respect to a host cell, that a polypeptide or nucleic acid does not naturally occur in the host cell. The term "heterologous" means, with respect to a polypeptide or nucleic acid, that a control sequence, e.g., promoter, of a polypeptide or nucleic acid is not naturally associated with the polypeptide or nucleic acid, i.e., the control sequence is from a gene other than the gene encoding the mature polypeptide.

[0037] Host Strain or Host Cell: A "host strain" or "host cell" is an organism into which an expression vector, phage, virus, or other DNA construct, including a polynucleotide encoding a polypeptide of the present invention has been introduced. Exemplary host strains are microorganism cells (e.g., bacteria, filamentous fungi, and yeast) capable of expressing the polypeptide of interest and / or fermenting saccharides. The term "host cell" includes protoplasts created from cells.

[0038] Introduced: The term "introduced" in the context of inserting a nucleic acid sequence into a cell, means "transfection", "transformation" or "transduction," as known in the art.

[0039] Isolated: The term “isolated” means a polypeptide, nucleic acid, cell, or other specified material or component that has been separated from at least one other material or component, including but not limited to, other proteins, nucleic acids, cells, etc. An isolated polypeptide, nucleic acid, cell or other material is thus in a form that does not occur in nature. An isolated polypeptide includes, but is not limited to, a culture broth containing the secreted polypeptide expressed in a host cell.

[0040] Mature polypeptide: The term “mature polypeptide” means a polypeptide in its mature form following N-terminal and / or C-terminal processing (e.g., removal of signal peptide). In one aspect, the mature polypeptide is SEQ ID NO: 5 (after removal of the signal peptide). In one aspect, the mature polypeptide is SEQ ID NO: 6 (after removal of the signal peptide and the propeptide).

[0041] Mature polypeptide coding sequence: The term “mature polypeptide coding sequence” means a polynucleotide that encodes a mature polypeptide having deamidase activity. In one aspect, the mature polypeptide coding sequence is nucleotides 1 to 882 of SEQ ID NO: 1.

[0042] Native: The term "native" means a nucleic acid or polypeptide naturally occurring in a host cell.

[0043] Nucleic acid: The term "nucleic acid" encompasses DNA, RNA, heteroduplexes, and synthetic molecules capable of encoding a polypeptide. Nucleic acids may be single stranded or double stranded, and may be chemical modifications. The terms "nucleic acid" and "polynucleotide" are used interchangeably. Because the genetic code is degenerate, more than one codon may be used to encode a particular amino acid, and the present compositions and methods encompass nucleotide sequences that encode a particular amino acid sequence. Unless otherwise indicated, nucleic acid sequences are presented in 5'-to-3' orientation.

[0044] Nucleic acid construct: The term "nucleic acid construct" means a nucleic acid molecule, either single- or double-stranded, which is isolated from a naturally occurring gene or is modified to contain segments of nucleic acids in a manner that would not otherwise exist in nature or which is synthetic, and which comprises one or more control sequences operably linked to the nucleic acid sequence.

[0045] Operably linked: The term "operably linked" means that specified components are in a relationship (including but not limited to juxtaposition) permitting them to function in an intended manner. For example, a regulatory sequence is operably linked to a coding sequence such that expression of the coding sequence is under control of the regulatory sequence.

[0046] Purified: The term “purified” means a nucleic acid, polypeptide or cell that is substantially free from other components as determined by analytical techniques well known in the art (e.g., a purified polypeptide or nucleic acid may form a discrete band in an electrophoretic gel, chromatographic eluate, and / or a media subjected to density gradient centrifugation). A purified nucleic acid or polypeptide is at least about 50% pure, usually at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91 %, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8% or more pure (e.g., percent by weight or on a molar basis). In a related sense, a composition is enriched for a molecule when there is a substantial increase in the concentration of the molecule after application of a purification or enrichment technique. The term "enriched" refers to a compound, polypeptide, cell, nucleic acid, amino acid, or other specified material or component that is present in a composition at a relative or absolute concentration that is higher than a starting composition.

[0047] In one aspect, the term "purified" as used herein refers to the polypeptide or cell being essentially free from components (especially insoluble components) from the production organism. In other aspects, the term "purified" refers to the polypeptide being essentially free of insoluble components (especially insoluble components) from the native organism from which it is obtained. In one aspect, the polypeptide is separated from some of the soluble components of the organism and culture medium from which it is recovered. The polypeptide may be purified ( / .e., separated) by one or more of the unit operations filtration, precipitation, or chromatography.

[0048] Accordingly, the polypeptide may be purified such that only minor amounts of other proteins, in particular, other polypeptides, are present. The term "purified" as used herein may refer to removal of other components, particularly other proteins and most particularly other enzymes present in the cell of origin of the polypeptide. The polypeptide may be "substantially pure", i.e., free from other components from the organism in which it is produced, e.g., a host organism for recombinantly produced polypeptide. In one aspect, the polypeptide is at least 40% pure by weight of the total polypeptide material present in the preparation. In one aspect, the polypeptide is at least 50%, 60%, 70%, 80% or 90% pure by weight of the total polypeptide material present in the preparation. As used herein, a "substantially pure polypeptide" may denote a polypeptide preparation that contains at most 10%, preferably at most 8%, more preferably at most 6%, more preferably at most 5%, more preferably at most 4%, more preferably at most 3%, even more preferably at most 2%, most preferably at most 1%, and even most preferably at most 0.5% by weight of other polypeptide material with which the polypeptide is natively or recombinantly associated.

[0049] It is, therefore, preferred that the substantially pure polypeptide is at least 92% pure, preferably at least 94% pure, more preferably at least 95% pure, more preferably at least 96% pure, more preferably at least 97% pure, more preferably at least 98% pure, even more preferably at least 99% pure, most preferably at least 99.5% pure by weight of the total polypeptide material present in the preparation. The polypeptide of the present invention is preferably in a substantially pure form ( / .e., the preparation is essentially free of other polypeptide material with which it is natively or recombinantly associated). This can be accomplished, for example by preparing the polypeptide by well-known recombinant methods or by classical purification methods.

[0050] Recombinant: The term "recombinant" is used in its conventional meaning to refer to the manipulation, e.g., cutting and rejoining, of nucleic acid sequences to form constellations different from those found in nature. The term recombinant refers to a cell, nucleic acid, polypeptide or vector that has been modified from its native state. Thus, for example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell, or express native genes at different levels or under different conditions than found in nature. The term “recombinant” is synonymous with “genetically modified” and “transgenic”.

[0051] Recover: The terms "recover" or “recovery” means the removal of a polypeptide from at least one fermentation broth component selected from the list of a cell, a nucleic acid, or other specified material, e.g., recovery of the polypeptide from the whole fermentation broth, or from the cell-free fermentation broth, by polypeptide crystal harvest, by filtration, e.g., depth filtration (by use of filter aids or packed filter medias, cloth filtration in chamber filters, rotary-drum filtration, drum filtration, rotary vacuum-drum filters, candle filters, horizontal leaf filters or similar, using sheed or pad filtration in framed or modular setups) or membrane filtration (using sheet filtration, module filtration, candle filtration, microfiltration, ultrafiltration in either cross flow, dynamic cross flow or dead end operation), or by centrifugation (using decanter centrifuges, disc stack centrifuges, hyrdo cyclones or similar), or by precipitating the polypeptide and using relevant solidliquid separation methods to harvest the polypeptide from the broth media by use of classification separation by particle sizes. Recovery encompasses isolation and / or purification of the polypeptide.

[0052] Sequence difference: The term "sequence difference" means the percent of amino acid differences between a polypeptide and the polypeptide of SEQ ID NO: 5, and is calculated as follows (Method 1a):

[0053] (Different Residues x 100) / (Length of SEQ ID NO: 5) wherein the different residues comprise any substitution, deletion, or insertion (e.g., an extension at the N-terminus and / or C-terminus) in the sequence. Method 1b: For purposes of the present invention, the sequence identity between two polynucleotide sequences is determined as the output of “longest identity” using the Needleman- Wunsch algorithm (Needleman and Wunsch, 1970, supra) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, supra), preferably version 6.6.0 or later. The parameters used are a gap open penalty of 10, a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version of NCBI NLIC4.4) substitution matrix. In order for the Needle program to report the longest identity, the nobrief option must be specified in the command line. The output of Needle labeled “longest identity” is calculated as follows:

[0054] (Identical Deoxyribonucleotides x 100) / (Length of Alignment- Total Number of Gaps in Alignment) Signal Peptide: A "signal peptide" is a sequence of amino acids attached to the N- terminal portion of a protein, which facilitates the secretion of the protein outside the cell. The mature form of an extracellular protein lacks the signal peptide, which is cleaved off during the secretion process.

[0055] Subsequence: The term “subsequence” means a polynucleotide having one or more nucleotides absent from the 5' and / or 3' end of a mature polypeptide coding sequence, wherein the subsequence encodes a fragment having deamidase activity.

[0056] Variant: The term “variant” means a synthetic polynucleotide encoding a polypeptide having deamidase activity, the polynucleotide comprising a man-made mutation, i.e., a nucleotide or codon substitution, insertion (including extension), and / or deletion, at one or more positions and / or codons. A substitution means replacement of the nucleotide occupying a position with a different nucleotide; a deletion means removal of the nucleotide occupying a position; and an insertion means adding 1-5 nucleotides (e.g., 1-3 nucleotides) adjacent to and immediately following the nucleotide occupying a position.

[0057] Wild-type: The term "wild-type" in reference to an amino acid sequence or nucleic acid sequence means that the amino acid sequence or nucleic acid sequence is a native or naturally- occurring sequence. As used herein, the term "naturally-occurring" refers to anything (e.g., proteins, amino acids, or nucleic acid sequences) that is found in nature. Conversely, the term "non-naturally occurring" refers to anything that is not found in nature (e.g., recombinant nucleic acids and protein sequences produced in the laboratory or modification of the wild-type sequence).

[0058] Detailed Description of the Invention

[0059] Synthetic Polynucleotides

[0060] The present invention relates to synthetic polynucleotides encoding a deamidase as described herein.

[0061] Thus, in a 1staspect the invention relates a synthetic polynucleotide encoding a deamidase, selected from the group consisting of:

[0062] (a) a polynucleotide having at least 80% sequence identity to SEQ ID NO: 1 ; (b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0063] (c) a polynucleotide derived from the polynucleotide of (a), or (b), wherein the 3’- and / or 5’- end has been extended by addition of one or more nucleotides; and

[0064] (d) a fragment of the polynucleotide of (a), (b), or (c).

[0065] In a preferred embodiment the deamidase has deamidase activity.

[0066] In one embodiment the polynucleotide has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:1.

[0067] In one embodiment the polynucleotide is comprising, consisting essentially of, or consisting of SEQ ID NO: 1.

[0068] In one embodiment the polynucleotide is comprising a GC content below 52%, e.g., below 51%, below 50%, below 49%, below 48%, of below 47%.

[0069] In one embodiment the polynucleotide comprises a GC content in the range of 40-52 %, such as 40-51 %, 40-50%, 41-49%, 40-49%, 42-48%, 42-47%, 43-48%, 43-47%, 44-47%, 44-46%, or 46- 47%.

[0070] In one embodiment the synthetic polynucleotide is encoding a mature deamidase.

[0071] In one embodiment the mature deamidase comprises or consists of a polypeptide having at least 80%, e.g., at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:6.

[0072] In one embodiment the polynucleotide is a fragment of SEQ ID NO: 1 , wherein the fragment preferably contains at least 555 nucleotides (e.g., nucleotides 328 to 882 of SEQ ID NO: 1), at least 500 nucleotides (e.g., nucleotides 300 to 800 of SEQ ID NO: 1), or at least 400 nucleotides (e.g., nucleotides 250 to 650 of SEQ ID NO: 1), preferably wherein the fragment encodes a deamidase having deamidase activity. In one embodiment the deamidase comprises or consists of an ammo acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 5.

[0073] In one embodiment the mature deamidase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 6.

[0074] In one embodiment sequence identity is determined by Sequence Identity Determination Method la.

[0075] In one embodiment sequence identity is determined by Sequence Identity Determination Method l b.

[0076] In one embodiment the polynucleotide is isolated.

[0077] In one embodiment the polynucleotide is purified.

[0078] The synthetic polynucleotide may be a synthetic cDNA, a synthetic DNA, a synthetic RNA, a synthetic mRNA, or a combination thereof. The synthetic polynucleotide may be derived from a native deamidase coding sequence from a strain of Chryseobacterium viscerum or a related organism and thus, for example, may be a variant of the synthetic polynucleotide sequence of the invention.

[0079] In an embodiment, the synthetic polynucleotide is a subsequence of the synthetic polynucleotide of the present invention encoding a fragment having deamidase activity. In an aspect, the subsequence contains at least 555 nucleotides (e.g., nucleotides 328 to 882 of SEQ ID NO: 1), at least 500 nucleotides (e.g., nucleotides 300 to 800 of SEQ ID NO: 1), or at least 400 nucleotides (e.g., nucleotides 250 to 650 of SEQ ID NO: 1).

[0080] In one embodiment the synthetic polynucleotide of the invention encoding the deamidase is derived from a Chryseobacterium cell.

[0081] The synthetic polynucleotide of the invention may also be mutated by introduction of nucleotide substitutions that do not result in a change in the amino acid sequence of the deamidase, but which further improve the yield of the deamindase in the host organism intended for production of the deamidase. For a general description of nucleotide substitution, see, e.g., Ford et al., 1991 , Protein Expression and Purification 2: 95-107. Nucleic Acid Constructs

[0082] In a 2ndaspect the present invention also relates to nucleic acid constructs comprising a synthetic polynucleotide of the 1staspect, wherein the synthetic polynucleotide is operably linked to one or more control sequences that direct the expression of the coding sequence in a suitable host cell under conditions compatible with the control sequences.

[0083] In one embodiment the one or more control sequences comprise a P3 promoter or a P3-based promoter, preferably the P3 promoter is a tandem promoter comprising the P3 promoter or is a tandem promoter derived from the P3 promoter.

[0084] In one embodiment the one or more control sequences comprises or consists of a promoter with a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:7.

[0085] In one embodiment the control sequences comprise a polynucleotide region that encodes a signal peptide fused to the N-terminus of the deamidase which directs the deamidase into the secretory pathway of the host cell; preferably the signal peptide is the amyL signal peptide.

[0086] In one embodiment the signal peptide comprises or consists of a polypeptide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:9.

[0087] In one embodiment the signal peptide is encoded by a polynucleotide comprising or consisting of a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:8.

[0088] The synthetic polynucleotide may be manipulated in a variety of ways to provide for expression of the deamidase. Manipulation of the synthetic polynucleotide prior to its insertion into a vector may be desirable or necessary depending on the expression vector. Techniques for modifying polynucleotides utilizing recombinant DNA methods are well known in the art.

[0089] Promoters

[0090] The control sequence may be a promoter, a polynucleotide that is recognized by a host cell for expression of a synthetic polynucleotide of the present invention. The promoter contains transcriptional control sequences that mediate the expression of the deamidase. The promoter may be any polynucleotide that shows transcriptional activity in the host cell including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.

[0091] Examples of suitable promoters for directing transcription of the polynucleotide of the present invention in a bacterial host cell are described in Sambrook et al. , 1989, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Lab., NY, Davis et al., 2012, supra, and Song et al., 2016, PLOS One 11(7): e0158447.

[0092] In one embodiment the promoter is a P3 promoter or a P3-based promoter, preferably the heterologous promoter is a tandem promoter comprising the P3 promoter or is a tandem promoter derived from the P3 promoter. For example, the P3 promoter comprises or consists of SEQ ID NO: 7.

[0093] Terminators

[0094] The control sequence may also be a transcription terminator, which is recognized by a host cell to terminate transcription. The terminator is operably linked to the 3’-terminus of the synthetic polynucleotide encoding the deamidase. Any terminator that is functional in the host cell may be used in the present invention.

[0095] Preferred terminators for bacterial host cells may be obtained from the genes for Bacillus clausii alkaline protease (aprH), Bacillus licheniformis alpha-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB). mRNA Stabilizers

[0096] The control sequence may also be an mRNA stabilizer region downstream of a promoter and upstream of the coding sequence of a gene which increases expression of the gene.

[0097] Examples of suitable mRNA stabilizer regions are obtained from a Bacillus thuringiensis crylllA gene (WO 94 / 25612) and a Bacillus subtilis SP82 gene (Hue etal., 1995, J. Bacterid. 177: 3465-3471).

[0098] Examples of mRNA stabilizer regions for fungal cells are described in Geisberg et al., 2014, Cell 156(4): 812-824, and in Morozov et al., 2006, Eukaryotic Ce / / 5(11): 1838-1846.

[0099] Leader Sequences

[0100] The control sequence may also be a leader, a non-translated region of an mRNA that is important for translation by the host cell. The leader is operably linked to the 5’-terminus of the synthetic polynucleotide encoding the deamidase. Any leader that is functional in the host cell may be used.

[0101] Suitable leaders for bacterial host cells are described by Hambraeus et al., 2000, Microbiology 146(12): 3051-3059, and by Kaberdin and Blasi, 2006, FEMS Microbiol. Rev. 30(6): 967-979. Polyadenylation Sequences

[0102] The control sequence may also be a polyadenylation sequence, a sequence operably linked to the 3’-terminus of the synthetic polynucleotide which, when transcribed, is recognized by the host cell as a signal to add polyadenosine residues to transcribed mRNA. Any polyadenylation sequence that is functional in the host cell may be used.

[0103] Signal Peptides

[0104] The control sequence may also be a signal peptide coding region that encodes a signal peptide linked to the N-terminus of a deamidase and directs the deamidase into the cell’s secretory pathway. The 5’-end of the coding sequence of the synthetic polynucleotide may inherently contain a signal peptide coding sequence naturally linked in translation reading frame with the segment of the coding sequence that encodes the polypeptide. Alternatively, the 5’-end of the coding sequence may contain a signal peptide coding sequence that is heterologous to the coding sequence. A heterologous signal peptide coding sequence may be required where the coding sequence does not naturally contain a signal peptide coding sequence. Alternatively, a heterologous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance secretion of the deamidase. Any signal peptide coding sequence that directs the expressed deamidase into the secretory pathway of a host cell may be used.

[0105] Effective signal peptide coding sequences for bacterial host cells are the signal peptide coding sequences obtained from the genes for Bacillus NCIB 11837 maltogenic amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis beta-lactamase, Bacillus stearothermophilus alphaamylase, Bacillus stearothermophilus neutral proteases (nprT, nprS, npr / VT), and Bacillus subtilis prsA. Further signal peptides are described by Freudl, 2018, Microbial Cell Factories 17: 52.

[0106] Propeptides

[0107] The control sequence may also be a propeptide coding sequence that encodes a propeptide positioned at the N-terminus of a deamidase. The resultant polypeptide is known as a proenzyme or propolypeptide (or a zymogen in some cases). A propolypeptide is generally inactive and can be converted to an active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the propolypeptide. As a non-limiting example the propeptide is encoded by a polynucleotide comprising nucleotides 1-327 of SEQ ID NO:1.

[0108] Where both signal peptide and propeptide sequences are present, the propeptide sequence is positioned next to the N-terminus of a polypeptide and the signal peptide sequence is positioned next to the N-terminus of the propeptide sequence. Additionally, or alternatively, when both signal peptide and propeptide sequences are present, the polypeptide may comprise only a part of the signal peptide sequence and / or only a part of the propeptide sequence. Alternatively, the final or isolated polypeptide may comprise a mixture of mature polypeptides and polypeptides which comprise, either partly or in full length, a propeptide sequence and / or a signal peptide sequence.

[0109] In one example, the propeptide consists of amino acids corresponding to amino acids at position 1-109 of SEQ ID NO: 5.

[0110] Expression Vectors

[0111] The present invention also relates to recombinant expression vectors comprising a synthetic polynucleotide of the present invention, a promoter, and transcriptional and translational stop signals. The various nucleotide and control sequences may be joined together to produce a recombinant expression vector that may include one or more convenient restriction sites to allow for insertion or substitution of the synthetic polynucleotide encoding the deamidase at such sites. Alternatively, the polynucleotide may be expressed by inserting the polynucleotide or a nucleic acid construct comprising the polynucleotide into an appropriate vector for expression. In creating the expression vector, the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.

[0112] The recombinant expression vector may be any vector (e.g., a plasmid or virus) that can be conveniently subjected to recombinant DNA procedures and can bring about expression of the polynucleotide. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector may be a linear or closed circular plasmid.

[0113] The vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector may contain any means for assuring self-replication. Alternatively, the vector may be one that, when introduced into the host cell, is integrated into the genome and replicated together with the chromosome(s) into which it has been integrated. Furthermore, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell, or a transposon, may be used.

[0114] The vector preferably contains one or more selectable markers that permit easy selection of transformed, transfected, transduced, or the like cells. A selectable marker is a gene the product of which provides for biocide or viral resistance, resistance to heavy metals, prototrophy to auxotrophs, and the like.

[0115] The vector preferably contains at least one element that permits integration of the vector into the host cell's genome or autonomous replication of the vector in the cell independent of the genome.

[0116] For integration into the host cell genome, the vector may rely on the polynucleotide’s sequence encoding the polypeptide or any other element of the vector for integration into the genome by homologous recombination, such as homology-directed repair (HDR), or non- homologous recombination, such as non-homologous end-joining (NHEJ).

[0117] For autonomous replication, the vector may further comprise an origin of replication enabling the vector to replicate autonomously in the host cell in question. The origin of replication may be any plasmid replicator mediating autonomous replication that functions in a cell. The term “origin of replication” or “plasmid replicator” means a polynucleotide that enables a plasmid or vector to replicate in vivo.

[0118] More than one copy of a synthetic polynucleotide of the present invention may be inserted into a host cell to increase production of a deamidase. For example, 2 or 3 or 4 or 5 or more copies are inserted into a host cell. An increase in the copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene with the polynucleotide where cells containing amplified copies of the selectable marker gene, and thereby additional copies of the polynucleotide, can be selected for by cultivating the cells in the presence of the appropriate selectable agent.

[0119] Host Cells

[0120] In a 3rdaspect, the present invention also relates to recombinant host cells, comprising a synthetic polynucleotide of the present invention operably linked to one or more control sequences that direct the production of a deamidase, e.g. a polynucleotide according to the 1staspect of the invention. Alternatively the host cell is comprising a nucleic acid construct or expression vector according to the 2ndaspect of the invention.

[0121] In one embodiment the deamidase is heterologous to the host cell.

[0122] In one embodiment at least one of the one or more control sequences is heterologous to the synthetic polynucleotide.

[0123] In one embodiment the host cell comprises at least two copies, e.g., three, four, or five, or more copies of the synthetic polynucleotide, nucleic acid construct or expression vector.

[0124] In one embodiment the host cell is a prokaryotic recombinant host cell, e.g., a Gram-positive cell selected from the group consisting of Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, or Streptomyces cells, or a Gram-negative bacteria selected from the group consisting of Campylobacter, E. coli, Flavobacterium, Fusobacterium, Helicobacter, llyobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma cells, such as Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, Bacillus thuringiensis, Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. Zooepidemicus, Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans cells.

[0125] In one embodiment the host cell is comprising in its genome one or more protease genes selected from the group consisting of: alkaline protease (aprL), glu-specific protease (mprL), bacillopeptidase F (bprAB), a first minor extracellular serine protease (epr), second minor extracellular serine protease (vpr), cell-wall associated protease (wprA), and intracellular serine protease (ispA), wherein at least six of the one or more protease genes are modified rendering at least six proteases truncated, partly or fully inactivated, present at reduced level or eliminated compared to the parent Bacillus strain when cultivated under identical conditions.

[0126] In one embodiment the host cell is comprising in its genome one or more protease genes selected from the group consisting of: alkaline protease (aprL), glu-specific protease (mprL), bacillopeptidase F (bprAB), a first minor extracellular serine protease (epr), second minor extracellular serine protease (vpr), cell-wall associated protease (wprA), and intracellular serine protease (ispA), wherein all seven of the one or more protease genes are modified rendering the seven proteases truncated, partly or fully inactivated, present at reduced level or eliminated compared to the parent Bacillus strain when cultivated under identical conditions.

[0127] In one embodiment the host cell is isolated.

[0128] In one embodiment the host cell is purified.

[0129] In one embodiment the host cell is a Bacillus cell; preferably a Bacillus subtilis cell, a Bacillus licheniformis cell, or a Bacillus amyloliquefaciens cell.

[0130] In one embodiment the host cell is a Bacillus subtilis cell.

[0131] In one embodiment the host cell is a Bacillus licheniformis cell. A construct or vector comprising a polynucleotide is introduced into a host cell so that the construct or vector is maintained as a chromosomal integrant or as a self-replicating extra- chromosomal vector as described earlier. The choice of a host cell will to a large extent depend upon the gene encoding the polypeptide and its source. The deamidase can be native or heterologous to the recombinant host cell. Also, at least one of the one or more control sequences can be heterologous to the synthetic polynucleotide encoding the deamidase. The recombinant host cell may comprise a single copy, or at least two copies, e.g., three, four, five, or more copies of the synthetic polynucleotide of the present invention.

[0132] The host cell may be any microbial cell useful in the recombinant expression of a polynucleotide of the present invention, e.g., a prokaryotic cell.

[0133] The prokaryotic host cell may be any Gram-positive or Gram-negative bacterium. Grampositive bacteria include, but are not limited to, Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, and Streptomyces. Gram-negative bacteria include, but are not limited to, Campylobacter, E. coli, Flavobacterium, Fusobacterium, Helicobacter, llyobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma.

[0134] The bacterial host cell may be any Bacillus cell including, but not limited to, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, and Bacillus thuringiensis cells. In an embodiment, the Bacillus cell is a Bacillus amyloliquefaciens, Bacillus licheniformis and Bacillus subtilis cell.

[0135] For purposes of this invention, Bacillus classes / genera / species shall be defined as described in Patel and Gupta, 2020, Int. J. Syst. Evol. Microbiol. 70: 406-438.

[0136] The bacterial host cell may also be any Streptococcus cell including, but not limited to, Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. Zooepidemicus cells.

[0137] The bacterial host cell may also be any Streptomyces cell including, but not limited to, Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans cells.

[0138] Methods for introducing DNA into prokaryotic host cells are well-known in the art, and any suitable method can be used including but not limited to protoplast transformation, competent cell transformation, electroporation, conjugation, transduction, with DNA introduced as linearized or as circular polynucleotide. Persons skilled in the art will be readily able to identify a suitable method for introducing DNA into a given prokaryotic cell depending, e.g., on the genus. Methods for introducing DNA into prokaryotic host cells are for example described in Heinze et al., 2018, BMC Microbiology 18:56, Burke et al., 2001 , Proc. Natl. Acad. Sci. USA 98: 6289-6294, Choi et al., 2006, J. Microbiol. Methods 64: 391-397, and Donald et al., 2013, J. Bacteriol. 195(11): 2612- 2620. Methods of Production

[0139] The present invention also relates to methods of producing a deamidase, comprising (a) cultivating a cell according to the 3rdaspect under conditions conducive for production of the deamidase; and optionally, (b) recovering the deamidase.

[0140] In one aspect, the cell is a Bacillus cell. In another aspect, the cell is a Bacillus subtilis cell. In another aspect, the cell is a Bacillus licheniformis cell. In another aspect, the cell is a Bacillus amyloliquefaciens cell.

[0141] In one embodiment, the method is comprising contacting the deamidase with a protease to provide a matured deamidease.

[0142] The host cell is cultivated in a nutrient medium suitable for production of the deamidase using methods known in the art. For example, the cell may be cultivated by shake flask cultivation, or small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state, and / or microcarrier-based fermentations) in laboratory or industrial fermentors in a suitable medium and under conditions allowing the polypeptide to be expressed and / or isolated. Suitable media are available from commercial suppliers or may be prepared according to published compositions (e.g., in catalogues of the American Type Culture Collection). If the polypeptide (deamidase) is secreted into the nutrient medium, the polypeptide can be recovered directly from the medium. If the polypeptide is not secreted, it can be recovered from cell lysates.

[0143] The polypeptide may be detected using methods known in the art that are specific for the polypeptide, including, but not limited to, the use of specific antibodies, formation of an enzyme product, disappearance of an enzyme substrate, or an assay determining the relative or specific activity of the polypeptide.

[0144] The polypeptide may be recovered from the medium using methods known in the art, including, but not limited to, collection, centrifugation, filtration, extraction, spray-drying, evaporation, or precipitation. In one aspect, a whole fermentation broth comprising the polypeptide is recovered. In another aspect, a cell-free fermentation broth comprising the polypeptide is recovered.

[0145] The polypeptide may be purified by a variety of procedures known in the art to obtain substantially pure polypeptides and / or polypeptide fragments (see, e.g., Wingfield, 2015, Current Protocols in Protein Science’, 80(1): 6.1.1-6.1.35; Labrou, 2014, Protein Downstream Processing, 1129: 3-10).

[0146] In an alternative aspect, the polypeptide is not recovered.

[0147] Generation of synthetic polynucleotides

[0148] A synthetic polynucleotide of the present invention may be a sequence obtained from microorganisms of any genus, optimized for expression in a host cell of choice. In other words, microorganisms of any genus may be a source of a native polynucleotide encoding a deamidase, which native polynucleotide is then altered and optimized towards expression in a recombinant host. For purposes of the present invention, the term obtained from as used herein in connection with a given source shall mean that the synthetic polynucleotide is a variant of the native sequence obtained from that microorganism.

[0149] In one aspect, the polynucleotide is an optimized sequence of a deamidase coding sequence obtained from Chryseobacterium, e.g., from Chryseobacterium viscerum, e.g., Chryseobacterium sp-62563, or Chryseobacterium proteloyticum.

[0150] It will be understood that for the aforementioned species, the invention encompasses both the perfect and imperfect states, and other taxonomic equivalents, e.g., anamorphs, regardless of the species name by which they are known. Those skilled in the art will readily recognize the identity of appropriate equivalents.

[0151] The native polynucleotides to be optimized may be identified and obtained from other sources including microorganisms isolated from nature (e.g., soil, composts, water, etc.) or DNA samples obtained directly from natural materials (e.g., soil, composts, water, etc.) using the above-mentioned probes. Techniques for isolating microorganisms and DNA directly from natural habitats are well known in the art. A polynucleotide encoding the deamidase may then be obtained by similarly screening a genomic DNA or cDNA library of another microorganism or mixed DNA sample. Once a polynucleotide encoding a deamidase has been detected with the probe(s), the polynucleotide can be isolated or cloned by utilizing techniques that are known to those of ordinary skill in the art (see, e.g., Davis et al., 2012, Basic Methods in Molecular Biology, Elsevier).

[0152] Fermentation Broth Formulations or Cell Compositions

[0153] The present invention also relates to a fermentation broth formulation, a fermentation broth, a composition or a cell composition comprising a synthetic polynucleotide of the present invention. The fermentation broth formulation or the cell composition further comprises additional ingredients used in the fermentation process, such as, for example, cells (including, the host cells containing the gene encoding a deamidase which are used to produce the deamidase), cell debris, biomass, fermentation media and / or fermentation products. In some embodiments, the composition is a cell-killed whole broth containing organic acid(s), killed cells and / or cell debris, and culture medium.

[0154] The term "fermentation broth" as used herein refers to a preparation produced by cellular fermentation that undergoes no or minimal recovery and / or purification. For example, fermentation broths are produced when microbial cultures are grown to saturation, incubated under carbon-limiting conditions to allow protein synthesis (e.g., expression of enzymes by host cells) and secretion into cell culture medium. The fermentation broth can contain unfractionated or fractionated contents of the fermentation materials derived at the end of the fermentation. Typically, the fermentation broth is unfractionated and comprises the spent culture medium and cell debris present after the microbial cells (e.g., bacterial cells) are removed, e.g., by centrifugation. In some embodiments, the fermentation broth contains spent cell culture medium, extracellular enzymes, and viable and / or nonviable microbial cells.

[0155] In some embodiments, the fermentation broth formulation or the cell composition comprises a first organic acid component comprising at least one 1-5 carbon organic acid and / or a salt thereof and a second organic acid component comprising at least one 6 or more carbon organic acid and / or a salt thereof. In some embodiments, the first organic acid component is acetic acid, formic acid, propionic acid, a salt thereof, or a mixture of two or more of the foregoing and the second organic acid component is benzoic acid, cyclohexanecarboxylic acid, 4-methylvaleric acid, phenylacetic acid, a salt thereof, or a mixture of two or more of the foregoing.

[0156] In one aspect, the composition contains an organic acid(s), and optionally further contains killed cells and / or cell debris. In some embodiments, the killed cells and / or cell debris are removed from a cell-killed whole broth to provide a composition that is free of these components.

[0157] The fermentation broth formulation or cell composition may further comprise a preservative and / or anti-microbial (e.g., bacteriostatic) agent, including, but not limited to, sorbitol, sodium chloride, potassium sorbate, and others known in the art.

[0158] The cell-killed whole broth or cell composition may contain the unfractionated contents of the fermentation materials derived at the end of the fermentation. Typically, the cell-killed whole broth or cell composition contains the spent culture medium and cell debris present after the microbial cells (e.g., Bacillus cells) are grown to saturation, incubated under carbon-limiting conditions to allow protein synthesis. In some embodiments, the cell-killed whole broth or cell composition contains the spent cell culture medium, extracellular enzymes, and killed filamentous fungal cells. In some embodiments, the microbial cells present in the cell-killed whole broth or cell composition can be permeabilized and / or lysed using methods known in the art.

[0159] A whole broth or cell composition as described herein is typically a liquid, but may contain insoluble components, such as killed cells, cell debris, culture media components, and / or insoluble enzyme(s). In some embodiments, insoluble components may be removed to provide a clarified liquid composition.

[0160] The whole broth formulations and cell compositions of the present invention may be produced by a method described in WO 90 / 15861 or WO 2010 / 096673.

[0161] The present invention is further described by the following examples that should not be construed as limiting the scope of the invention.

[0162] Examples

[0163] Strains

[0164] The parental B. licheniformis strain SJ 13633 used in example 3 contains protease deletions in the genes encoding alkaline protease (encoded by aprL gene), Glu-specific protease (encoded by mprL gene), bacillopeptidase F (encoded by bprAB gene), two minor extracellular serine proteases (encoded by epr gene, and by vpr gene), cell-wall associated protease (encoded by wprA gene), and intracellular serine protease (encoded by the ispA gene).

[0165] Overview of

[0166] SEQ ID NO: 1 synthetic DNA sequence of design 6

[0167] SEQ ID NO: 2 synthetic DNA sequence of design 2

[0168] SEQ ID NO: 3 synthetic DNA sequence of design 8

[0169] SEQ ID NO: 4 wt deamidase DNA sequence of Chryseobacterium viscerum (design 1)

[0170] SEQ ID NO: 5 deamidase AA

[0171] SEQ ID NO: 6 deamidase AA sequence (matured form)

[0172] SEQ ID NO: 7 promoter sequence P3

[0173] SEQ ID NO: 8 amyL SP (3A) coding sequence

[0174] SEQ ID NO: 9 amyL SP (3A)

[0175] SEQ ID NO: 10 6x His tag coding sequence

[0176] Example 1 : Generation of several synthetic designs for deamidase expression

[0177] Several synthetic sequences were generated aiming to increase deamidase expression in Bacillus hosts, without changing the AA sequence of the deamidase polypeptide (SEQ ID NO: 5) and its matured form (SEQ ID NO: 6). Deamidase DNA designs 2-10 encode the same deamidase as DNA design 1 , except that designs 2-10 encode a deamidase comprising a F99G substitution (described in WO23170177).

[0178] The starting sequence was the native deamidase coding sequence from Chryseobacterium viscerum (SEQ ID NO: 4, i.e. design 1) based on which synthetic designs 2- 10 were generated for optimized expression in Bacillus cells.

[0179] Figure 1 shows the phylogenetic relationship between the 10 designs and the diversity of the generated sequences using the Neighbor-Joining method (Saitou N. and Nei M. (1987), Molecular Biology and Evolution 4:406-425).

[0180] The percentage of replicate trees in which the associated designs clustered together in the bootstrap test (1000 replicates) are shown next to the branches (Felsenstein J. 1985, Evolution 39:783-791).

[0181] The tree is drawn to scale, with branch lengths in the same units as those of the evolutionary distances used to infer the phylogenetic tree. The evolutionary distances were computed using the Maximum Composite Likelihood method (Tamura K. et al., 2004, Proceedings of the National Academy of Sciences (SA) 101 :11030-11035) and are in the units of the number of base substitutions per site. The analysis involved 11 nucleotide sequences. Codon positions included were 1st+2nd+3rd+Noncoding. All positions containing gaps and missing data were eliminated. There was a total of 987 positions in the final dataset. Evolutionary analyses were conducted in MEGA7 (Kumar S. et al., 2016, Molecular Biology and Evolution 33:1870- 1874).

[0182] As shown in Fig. 1 , design 6 is closest related to design 7 forming a cluster with 80%, whereas the remaining designs form separate clusters which are more distinct from design 6. Also shown in Fig. 1 is a cluster formed by design 2 and design 8.

[0183] Table 1 shows a comparison of the GC content between synthetic designs 2, 6, and 8 relative to the GC content of the native reference sequence (design 1). As shown in Table 1 , the native sequence has a GC content of around 40.22%, whereas the tested designs 2, 6 and 8 all have a higher GC content. Designs 2 and 8 both have GC contents around 52%. Interestingly, design 6 has a GC content of 46.5%, meaning a lower GC content than designs 2 and 8, but a higher GC content compared to the native sequence (design 1). Thus, the GC content of design 6 is distinct from the other 3 tested designs.

[0184] Table 1 . GC-content of designs and wildtype.

[0185] Table 2. Percent-identity matrix (PIM) with sequence homology in %.

[0186] Table 2 shows a percent identity matrix (PIM) comparing the sequence homologies of the 4 tested sequences to another. As shown in Table 2, design 2 and design 8 form a cluster with 95.92% sequence homology, which cluster is also indicated in Figure 1. All 3 synthetic designs 2, 6 and 8 show only around 75% sequence homology to the native sequence of design 1 , indicating that significant codon changes are present in the synthetic designs. In other words, the percentage of recoded nucleotides relative to design 1 is 24.08% for design 2, 23.72% for design 6, and 25.03% for design 8. Also, design 6 has only 78.35% and 78.91% sequence homology to design 2 and design 8, respectively, further confirming that the sequence of design 6 is very distinct from the sequences of design 2 and design 8. Example 2: Design 6 showed significantly increased deamidase expression in B. subtilis

[0187] To test the deamidase expression of the 4 designs of Tables 1-2 in Bacillus subtilis cells, expression cassettes were prepared. The cassettes comprise the different designs fused to the amyL signal peptide coding sequence (SEQ ID NO: 8; encoding the signal peptide of SEQ ID NO:9) located upstream of the deamidase coding sequence and operably linked thereto. Downstream of the deamidase coding sequence we cloned a 6x His-tag operably linked to the different designs. The 6x His-tag coding sequence is shown in SEQ ID NO: 10. Each cassette comprised the P3 triple promoter sequence of SEQ ID NO: 7 operably linked to the signal peptide coding sequence.

[0188] Synthetic genes cloning was made by overlap SOE PCR as described previously in Example 5 of WQ23170177 (Novozymes A / S). Bacillus subtilis cells transformation, and fermentation was made according to procedure described in Example 5 of WQ23170177 (Novozymes A / S). The tested strains each comprised a single copy of the expression cassette.

[0189] Protein deamidase was purified via HisTag purification and the deamidase was analysed by SDS PAGE as described in Example 6 and Example 7 of WQ23170177 (Novozymes A / S), respectively. Deamidase quantification by densitometry was carried out as described in Example 7 of WQ23170177 (Novozymes A / S).

[0190] Deamidase quantification for the 4 designs is shown in Table 3. Relative to the yield of design 1 , designs 2 and 6 both improved deamidase yield with 115% and 172%, respectively. Design 8 showed poor expression with a yield of only 87% relative to the wildtype sequence.

[0191] Surprisingly, out of the 3 synthetic sequences only design 6 could increase deamidase expression significantly.

[0192] Table 3. Deamidase yield after expressing the 4 different designs.

[0193] Example 3: Design 6 shows increased deamidase expression in B. licheniformis

[0194] To test the deamidase expression of designs 1 , 6, and 8 of Tables 1-2 in Bacillus licheniformis cells, synthetic DNA was ordered with the three gene designs combined with the same amyL signal peptide as when examining expression in Bacillus subtilis (SEQ ID NO: 8; encoding the signal peptide of SEQ ID NO:9), but a stop codon instead of a fusion to a downstream His-tag. The different synthetic DNAs were joined by POE PCR with plasmid elements needed to generate flp-FRT donor plasmids (WO 2018 / 077796 A1). The resulting plasmids were transformed into a Bacillus subtilis donor and conjugated into the Bacillus licheniformis strain SJ 13633. Using the procedure described in WO 2018 / 077796 A1 , this will enable insertion of the deamidase genes, linked to the signal peptide coding sequence, in the host chromosome under control of the triple promoter P3 (SEQ ID NO: 7, and as described in WO 99 / 43835) in one or more integration loci.

[0195] The following strains were obtained:

[0196] BT13231 : design 1 (2 gene copies inserted) BT13273: design 6 (1 gene copy inserted) BT13275: design 6 (2 gene copies inserted) BT13277: design 8 (1 gene copy inserted)

[0197] The strains were fermented and deamidase expression was calculated. A suitable fermentation protocol is the lab-scale fermentation method disclosed in WO 2023 / 152220. The deamidase expression was calculated by activity analysis measurement according to “Method 2” described in the section “Definition” of the present application. The expression levels for the different designs are presented in Table 4.

[0198] Table 4. Expression levels for different sequence designs.

[0199] As shown in Table 4, strain BT13273 comprising one copy of design 6 showed a 9% increased yield compared to strain BT13231 comprising two copies of the wild type gene. Design 6 can thus reduce the need for additional copy numbers in the genome and consequently shorten strain construction times and efforts without compromising deamidase yield, in fact while increasing deamidase yield at the same time.

[0200] As also shown in T able 4, strain BT 13275 comprising two gene copies of design 6 showed ca. a 42% increase of yield relative to BT13231 comprising two copies of the wild type gene, confirming that design 6 is leading to improved expression. Using design 6 for expression of deamidase, yield can be increased without increasing gene copy number, allowing efficient strain construction.

[0201] Table 4 also shows that the expression level with one copy of design 8 (BT13277) is significantly lower than with one copy of design 6, again demonstrating that the increase in yield obtained with design 6 is surprising and indicating that generating and identifying optimized DNA sequences is a challenging task.

[0202] The invention described and claimed herein is not to be limited in scope by the specific aspects herein disclosed, since these aspects are intended as illustrations of several aspects of the invention. Any equivalent aspects are intended to be within the scope of this invention. Indeed, various modifications of the invention in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. In the case of conflict, the present disclosure including definitions will control.

[0203] The invention is further defined by the following numbered paragraphs:

[0204] 1 . A synthetic polynucleotide encoding a deamidase, selected from the group consisting of:

[0205] (a) a polynucleotide having at least 80% sequence identity to SEQ ID NO: 1 ;

[0206] (b) a polynucleotide derived from SEQ ID NO: 1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;

[0207] (c) a polynucleotide derived from the polynucleotide of (a), or (b), wherein the 3’- and / or 5’- end has been extended by addition of one or more nucleotides; and

[0208] (d) a fragment of the polynucleotide of (a), (b), or (c).

[0209] 2. The polynucleotide according to paragraph 1 , wherein the deamidase has deamidase activity.

[0210] 3. The polynucleotide of any one of paragraphs 1-2, wherein the polynucleotide has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:1. 4. The polynucleotide of any one of the preceding paragraphs, the polynucleotide comprising, consisting essentially of, or consisting of SEQ ID NO: 1.

[0211] 5. The polynucleotide of any one of the preceding paragraphs, the polynucleotide comprising a GC content below 52%, e.g., below 51%, below 50%, below 49%, below 48%, of below 47%.

[0212] 6. The polynucleotide of any one of the preceding paragraphs, the polynucleotide comprising a GC content in the range of 40-52 %, such as 40-51%, 40-50%, 41-49%, 40-49%, 42-48%, 42- 47%, 43-48%, 43-47%, 44-47%, 44-46%, or 46-47%.

[0213] 7. The polynucleotide of any one of the preceding paragraphs, wherein the synthetic polynucleotide is encoding a mature deamidase.

[0214] 8. The polynucleotide of paragraph 7, wherein the mature deamidase comprises or consists of a polypeptide having at least 80%, e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:6.

[0215] 9. The polynucleotide of any one of the preceding paragraphs, which is a fragment of SEQ ID NO: 1 , wherein the fragment preferably contains at least 555 nucleotides (e.g., nucleotides 328 to 882 of SEQ ID NO: 1), at least 500 nucleotides (e.g., nucleotides 300 to 800 of SEQ ID NO: 1), or at least 400 nucleotides (e.g., nucleotides 250 to 650 of SEQ ID NO: 1), preferably wherein the fragment encodes a deamidase having deamidase activity.

[0216] 10. The polynucleotide of any one of the preceding paragraphs, wherein the deamidase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 5.

[0217] 11. The polynucleotide of any one of the preceding paragraphs, wherein the comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 6. 12. The polynucleotide of any one of the preceding paragraphs, wherein sequence identity is determined by Sequence Identity Determination Method 1.

[0218] 13. The polynucleotide of any one of the preceding paragraphs, which is isolated.

[0219] 14. The polynucleotide of any one of the preceding paragraphs, which is purified.

[0220] 15. A nucleic acid construct or expression vector comprising the polynucleotide of any one of the preceding paragraphs operably linked to one or more control sequences that direct the production of the deamidase in an expression host.

[0221] 16. The nucleic acid construct or expression vector according to paragraph 15, wherein the one or more control sequences comprise a P3 promoter or a P3-based promoter, preferably the P3 promoter is a tandem promoter comprising the P3 promoter or is a tandem promoter derived from the P3 promoter.

[0222] 17. The nucleic acid construct or expression vector according to any one of the preceding paragraphs, wherein the one or more control sequences comprises or consists of a promoter with a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:7.

[0223] 18. The nucleic acid construct or expression vector according to any one the preceding paragraphs, wherein the control sequences comprise a polynucleotide region that encodes a signal peptide fused to the N-terminus of the deamidase which directs the deamidase into the secretory pathway of the host cell; preferably the signal peptide is the amyL signal peptide.

[0224] 19. The nucleic acid construct or expression vector according to any one of the preceding paragraphs, wherein the signal peptide comprises or consists of a polypeptide sequence having at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:9.

[0225] 20. The nucleic acid construct or expression vector according to any one of the preceding paragraphs, wherein the signal peptide is encoded by a polynucleotide comprising or consisting of a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:8. 21. A recombinant host cell comprising the nucleic acid construct or expression vector of any one of the preceding paragraphs.

[0226] 22. The host cell according to paragraph 21 , wherein the deamidase is heterologous to the host cell.

[0227] 23. The host cell of any one of paragraph 21 or 22, wherein at least one of the one or more control sequences is heterologous to the synthetic polynucleotide.

[0228] 24. The host cell of any one of paragraphs 22-23, which comprises at least two copies, e.g., three, four, or five, or more copies of the polynucleotide, nucleic acid construct or expression vector according to any one of paragraphs 1-20.

[0229] 25. The host cell of any one of paragraphs 21-24, which is a prokaryotic recombinant host cell, e.g., a Gram-positive cell selected from the group consisting of Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, or Streptomyces cells, or a Gram-negative bacteria selected from the group consisting of Campylobacter, E. coli, Flavobacterium, Fusobacterium, Helicobacter, llyobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma cells, such as Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, Bacillus thuringiensis, Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. Zooepidemicus, Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans cells.

[0230] 26. The host cell of any one of paragraphs 21-25, comprising in its genome one or more protease genes selected from the group consisting of: alkaline protease (aprL), glu-specific protease (mprL), bacillopeptidase F (bprAB), a first minor extracellular serine protease (epr), second minor extracellular serine protease (vpr), cell-wall associated protease (wprA), and intracellular serine protease (ispA), wherein at least six of the one or more protease genes are modified rendering at least six proteases truncated, partly or fully inactivated, present at reduced level or eliminated compared to the parent Bacillus strain when cultivated under identical conditions.

[0231] 27. The host cell of any one of paragraphs 21-26, comprising in its genome one or more protease genes selected from the group consisting of: alkaline protease (aprL), glu-specific protease (mprL), bacillopeptidase F (bprAB), a first minor extracellular serine protease (epr), second minor extracellular serine protease (vpr), cell-wall associated protease (wprA), and intracellular serine protease (ispA), wherein all seven of the one or more protease genes are modified rendering the seven proteases truncated, partly or fully inactivated, present at reduced level or eliminated compared to the parent Bacillus strain when cultivated under identical conditions.

[0232] 28. The host cell of any one of paragraphs 21-27, which is isolated.

[0233] 29. The host cell of any one of paragraphs 21-28, which is purified.

[0234] 30. The host cell according to any one of paragraphs 21-29, wherein the host cell is a Bacillus cell; preferably a Bacillus subtilis cell, a Bacillus licheniformis cell, or a Bacillus amyloliquefaciens cell.

[0235] 31 . The host cell according to any one of paragraphs 21-30, wherein the host cell is a Bacillus subtilis cell.

[0236] 32. The host cell according to any one of paragraphs 21-30, wherein the host cell is a Bacillus licheniformis cell.

[0237] 33. A method of producing a deamidase, comprising cultivating the recombinant host cell of any one of paragraphs 21-32 under conditions conducive for production of the deamidase; and optionally recovering the deamidase.

[0238] 34. The method of paragraph 33, comprising contacting the deamidase with a protease to provide a matured deamidease.

[0239] 35. A composition, cell composition or fermentation broth comprising the polynucleotide of any one of paragraphs 1-20, and / or the cell of any one of paragraphs 21-32.

Claims

Claims1 . A synthetic polynucleotide encoding a deamidase, selected from the group consisting of:(a) a polynucleotide having at least 80% sequence identity to SEQ ID NO:1 ,(b) a polynucleotide derived from SEQ ID NO:1 by having 1-30 alterations (e.g., substitutions, deletions and / or insertions at one or more positions, e.g., 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15 or 16 or 17 or 18 or 19 or 20 or 21 or 22 or 23 or 24 or 25 or 26 or 27 or 28 or 29 or 30 alterations, in particular substitutions;(c) a polynucleotide derived from the polynucleotide of (a), or (b), wherein the 3’- and / or 5’- end has been extended by addition of one or more nucleotides; and(d) a fragment of the polynucleotide of (a), (b), or (c).

2. The polynucleotide of claim 1 , wherein the polynucleotide has at least 85%, e.g., at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO:1.

3. The polynucleotide of any one of claim 1 to 2, the polynucleotide comprising, consisting essentially of, or consisting of SEQ ID NO: 1.

4. The polynucleotide of any one of claims 1 to 3, wherein the polynucleotide comprises a GC content in the range of 40-52 %, such as 40-51%, 40-50%, 41-49%, 40-49%, 42-48%, 42- 47%, 43-48%, 43-47%, 44-47%, 44-46%, or 46-47%.

5. The polynucleotide of any one of claims 1 to 4, wherein the deamidase comprises or consists of an amino acid sequence having at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NO: 5 or 6.

6. A nucleic acid construct or expression vector comprising the polynucleotide of any one of the claims 1 to 5 operably linked to one or more control sequences that direct the production of the deamidase in an expression host.

7. A recombinant host cell comprising the nucleic acid construct or expression vector of claim 6.

8. The host cell according to claim 7, wherein the host cell is a Bacillus cell.

9. The host cell according to any one of claims 7 to 8, wherein the deamidase is heterologous to the host cell.

10. A method of producing a deamidase, comprising cultivating the recombinant host cell of any one of claims 7 to 9 under conditions conducive for production of the deamidase; and optionally recovering the deamidase.

Citation Information

Patent Citations

  • A method for killing cells without cell lysis

    WO1990015861A1

  • Nucleotide sequences for the control of the expression of DNA sequences in a cellular host

    WO1994025612A2

  • Methods for producing a polypeptide in a bacillus cell

    WO1999043835A2

  • Fermentation broth formulations

    WO2010096673A1

  • FLP-mediated genomic integration in bacillus licheniformis

    WO2018077796A1