Novel thermostable enzymes

A Thermosyntropha lipolytica Tl_Est47 variant with improved thermostability, achieved through specific amino acid modifications, addresses the need for high-temperature enzymes, maintaining activity up to 99°C and improving industrial applicability.

JP2025535521APending Publication Date: 2025-10-24ENZIDE TECHNOLOGIES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525025
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-28
Filing Date
2023-10-25
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

There is an increasing need for enzymes with high temperature thermotolerance properties for use in industrial processes and products that involve high temperatures.

Method used

A novel polypeptide variant of Thermosyntropha lipolytica Tl_Est47 with improved thermostability, characterized by specific amino acid substitutions at positions 221 and 377, is developed, along with associated nucleotide sequences and expression systems for production in host cells.

Benefits of technology

The variant enzyme maintains enzymatic activity at higher temperatures, exceeding 90°C and retaining activity even after heating at 99°C, enhancing its suitability for industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535521000003
    Figure 2025535521000003
  • Figure 2025535521000004
    Figure 2025535521000004
  • Figure 2025535521000005
    Figure 2025535521000005
Patent Text Reader

Abstract

Novel enzyme variants of naturally occurring wild-type Thermosyntropha lipolytica Tl_Est47 are provided, which have improved thermotolerance properties that do not adversely affect the stability or activity of the enzyme variants over a range of temperatures or pH values.
Need to check novelty before this filing date? Find Prior Art

Description

Detailed Description of the Invention

[0001] FIELD OF THE INVENTION The present invention relates to novel variants of the Thermosyntropha lipolytica Tl_Est47 enzyme that have improved thermotolerance properties compared to the naturally occurring wild-type Thermosyntropha lipolytica Tl_Est47.

[0002] [Background technology] The following background art discussion is intended solely to facilitate an understanding of the present invention. It should be understood that this discussion is not an admission or admission that any of the material referred to was part of the common general knowledge at the priority date of this application.

[0003] Enzymes have been used for many years in products in the detergent, textile, and starch industries, and in food manufacturing, including cheese, bread, beer, wine, leather, and linen. These industrial processes utilize enzymes produced and isolated by specific microorganisms, or enzymes present in natural sources (e.g., papaya fruit).

[0004] Recent advances in protein engineering have led to the development of mutant enzymes for established applications and new, custom-made enzymes for application areas where enzymes have not previously been used. More than half of the enzymes used in industrial processes are derived from fungi, more than one-third from bacteria, and the remainder from animals and plants. Recombinant DNA technology has made it possible to isolate and clone enzyme-encoding genes from all possible sources and heterologously express them at high yields. This has resulted in improved production levels and the availability of enzymes from bacterial strains not normally suited for industrial use, including Aspergillus, Saccharomyces, and Bacillus, among others.

[0005] The global market for industrial enzymes is currently estimated at approximately US$10 billion, and commonly used enzymes include lipase, polyphenol oxidase, lignin peroxidase, horseradish peroxidase, amylase, nitrite reductase, and urease.

[0006] Enzymes are selected for use in these industrial processes because of the reactions they initiate and catalyze, as well as various properties they possess when operating within the environment of these processes, such as the ability of the enzyme to maintain enzymatic activity in the various temperature and pH environments required.

[0007] For use in many industrial processes that involve high temperatures at certain stages, or in products intended for use under conditions involving high temperatures, enzymes with a high level of thermostability are of substantial benefit.

[0008] Thus, there is an increasing need for enzymes with high temperature thermotolerant properties for use in products and in certain industrially useful applications and processes.

[0009] Summary of the Invention The present inventors have designed a novel polypeptide that is a variant of Thermosyntropha lipolytica Tl_Est47 that has improved thermostability in terms of maintaining enzymatic activity at higher temperatures when compared to naturally occurring (wild-type) Thermosyntropha lipolytica Tl_Est47.

[0010] In a first aspect, the present invention provides a novel polypeptide comprising: i) the amino acid sequence set forth in SEQ ID NO: 4; ii) an amino acid sequence that is at least 90% identical to i); or iii) A biologically active fragment of ii).

[0011] In one embodiment, the novel polypeptide of the present invention comprises: i) a glutamine (Q) at a position corresponding to amino acid number 221 of SEQ ID NO:4; and ii) an aspartic acid (D) at a position corresponding to amino acid number 377 of SEQ ID NO:4.

[0012] In one embodiment, the novel polypeptides of the present invention comprise fusion proteins that include at least one other polypeptide sequence.

[0013] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the present invention comprising a sequence selected from the following: a. the sequence of nucleotides set forth in SEQ ID NO:2; b. a nucleotide sequence encoding a novel polypeptide of the invention; or c. A sequence of nucleotides complementary to either i) or ii).

[0014] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the present invention as described herein, wherein said polynucleotide is operably linked to a promoter capable of directing expression of a novel polypeptide of the present invention in a cell.

[0015] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the invention described herein, wherein said polynucleotide is operably linked to a promoter capable of directing expression of a novel polypeptide of the invention in an expression host cell.

[0016] In one embodiment, the present invention provides a vector comprising a polynucleotide of the invention described herein.

[0017] In one embodiment, the present invention provides a nucleic acid construct or expression vector comprising a polynucleotide of the invention described herein.

[0018] In one embodiment, the present invention provides a nucleic acid construct or expression vector comprising a polynucleotide of the invention described herein, wherein said polynucleotide is operably linked to one or more control sequences that direct the production of a novel polypeptide of the invention in an expression host cell.

[0019] In one embodiment, the present invention provides a recombinant expression host cell comprising a polynucleotide encoding a novel polypeptide of the present invention, wherein said polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide.

[0020] In one embodiment, the present invention provides a host cell comprising a polynucleotide of the invention described herein.

[0021] In one embodiment, the host cell preferably comprises a bacterial cell, a fungal cell, or a plant cell.

[0022] In one embodiment, the invention provides a transgenic non-human organism comprising at least one host cell described herein.

[0023] In one embodiment, the present invention provides an extract of a host cell described herein, wherein the extract comprises: a. the amino acid sequence set forth in SEQ ID NO:4; b. an amino acid sequence that is at least 90% identical to i), or c. a biologically active fragment of ii); The polypeptide comprises:

[0024] In one embodiment, the invention provides an extract of a host cell described herein, wherein the polypeptide is: a. a glutamine (Q) at a position corresponding to amino acid 221 of SEQ ID NO:4; and b. one or both of the aspartic acids (D) at the position corresponding to amino acid number 377 of SEQ ID NO:4.

[0025] In one embodiment, the present invention provides compositions comprising a novel polypeptide of the present invention and one or more acceptable carriers.

[0026] In one embodiment, the present invention provides a composition comprising an extract as described herein and one or more acceptable carriers.

[0027] In one embodiment, the present invention provides: a. the amino acid sequence set forth in SEQ ID NO:4; b. an amino acid sequence that is at least 90% identical to i), or c. a biologically active fragment of ii); The present invention provides a method for producing a novel polypeptide comprising:

[0028] In one embodiment, the novel polypeptide preferably comprises: a. a glutamine (Q) at a position corresponding to amino acid 221 of SEQ ID NO:4; and b. one or both of the aspartic acids (D) at the position corresponding to amino acid number 377 of SEQ ID NO:4.

[0029] In one embodiment, the present invention provides a method for producing a novel polypeptide of the present invention, comprising culturing a recombinant expression host cell comprising a novel polypeptide of the present invention, wherein said polynucleotide is operably linked to one or more control sequences that direct the production of the novel polypeptide of the present invention under conditions conducive to the production of said polypeptide.

[0030] In one embodiment, the method preferably comprises recovering the novel polypeptide of the invention.

[0031] In one embodiment, the present invention provides an enzyme comprising the novel polypeptide of the present invention.

[0032] In one embodiment, the present invention provides a thermostable enzyme comprising the novel polypeptide of the present invention.

[0033] In one embodiment, the present invention provides a hyperthermophilic enzyme comprising a novel polypeptide of the present invention.

[0034] In one embodiment, the invention provides an enzyme described herein, wherein the enzyme maintains enzymatic activity at a temperature higher than the temperature at which an enzyme comprising a wild-type Thermosyntropha lipolytica Tl_Est47 polypeptide comprising the amino acid sequence set forth in SEQ ID NO:3 loses substantially all enzymatic activity.

[0035] In one embodiment, the enzymes described herein maintain enzymatic activity at temperatures above 90°C.

[0036] In one embodiment, the enzymes described herein retain some enzymatic activity at temperatures above 95°C.

[0037] In one embodiment, the enzymes described herein comprise esterase activity.

[0038] In one embodiment, the invention provides an enzyme as described herein that comprises lipase activity.

[0039] In one embodiment, the present invention provides a novel polypeptide of the present invention comprising a polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47, wherein said polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47 comprises the amino acid sequence set forth in SEQ ID NO:3.

[0040] In one embodiment, the invention provides a polynucleotide of the invention described herein capable of expressing a polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47, wherein said wild-type Thermosyntropha lipolytica Tl_Est47 comprises an amino acid sequence expressed by an isolated and / or exogenous polynucleotide comprising a sequence selected from the sequence of nucleotides set forth in SEQ ID NO:1.

[0041] [Brief explanation of the figure] The invention will now be described, by way of example only, with reference to the accompanying drawings in which: Figure 1 shows an unrooted maximum likelihood phylogenetic tree of bacterial lipase family 1.5 constructed using IQ-tree and visualized with iTOL v5. UF-Boot values ​​are shown for nodes with SH-aLRT support values ​​≥ 80%. Chemical analysis records of macroalgal hydrolysates are shown. Figure 2 shows the esterase activity of lipase homologs using pNP substrate. (A) A bar graph shows an activity test using 5 μl of cell culture with 0.75 mM pNP acetate to confirm protein expression. Error bars represent SEM from three replicate experiments. (B) A photograph of an SDS-PAGE gel analyzing the purity of the protein obtained from small-scale nickel affinity purification. The expected band sizes are indicated at the bottom of the gel. (C) A bar graph shows the comparison of the activity of 10 nM purified protein using pNP substrates of different lengths. Error bars represent SEM from two replicate experiments. Figure 3 shows a line graph depicting the residual activity of heat-stable proteins for pNP acetate and pNP propionate. Samples were heated at the indicated temperatures for 10 minutes, cooled to 4°C, and then the residual activity was measured. The residual activity is provided as a percentage of the activity of the unheated sample. The six enzymes shown were identified from the initial assay using all 10 enzymes presented in this study. Error bars represent the SEM from three replicate experiments. Figure 4 is a graph showing EqAD activity in 185 μl of supernatant from a reaction containing 100 nM Cl_EstA and 5 mg / mL PBAT and incubated for 48 hours at 40°C. The EqAD reaction contained 250 μM NAD in a final reaction volume of 200 μL. + and 0.1 U / mL EqAD, and the reaction progress was monitored by NADH + +H + Measurement was carried out by tracking the change in absorbance at 340 nm during production. Figure 5 shows the thermotolerance of the P4G12 Tl_Est47 variant measured using protein-expressing cell cultures. (A) Represents an individual graph of a plate screened for activity by heating cultures at 89°C for 10 minutes, heating overnight, then cooling, and then testing for activity against paranitrophenyl butyrate. (B) Represents a line graph showing the residual activity of the improved variant TL_Est47 compared to the wild type, heated from 75°C to 99°C for 10 minutes, cooled, and then tested for activity against paranitrophenyl butyrate. Figure 6 shows the properties of WT Tl_Est47 and its L221Q and G377D mutants, which have improved thermostability. (A) A line graph showing the esterase activity of the WT and mutant proteins measured using pNP-butyrate as a substrate. (B) A table showing the kinetic parameters of the WT and mutant proteins for the curves shown in (A). Figure 7 shows a line graph depicting the thermostability of purified WT and mutant Tl_Est47 in LB medium. Reactions to measure residual activity contained 200 nM enzyme and 300 μM pNP-butyrate, and the amount of liberated pNP was measured by absorbance at 405 nm. FIG. 8 shows a table showing the experimental log of conditions in the 2 liter fermentor. FIG. 9 shows a table showing the experimental log of conditions in a 500 mL flask. FIG. 10 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with wild-type Tl_Est47 at pH 7 to pH 9. FIG. 11 depicts a graph showing the alcohol dehydrogenase activity of PBSA incubated with Tl_Est47 variants from pH 7 to pH 9 (Tris buffer). FIG. 12 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with wild-type Tl_Est47 from 4° C. to 37° C. FIG. 13 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with Tl_Est47 variants from 4° C. to 37° C. FIG. 14 shows SDS PAGE gels of whole cell and soluble (blue arrow indicates Tl_Est47) after 2, 4, 5.5, and 22 hours of incubation time after induction in either flasks or fermentors. Figure 15 shows FPLC traces and SDS-PAGE gels of (A) pellets from 500 ml flask cultures and (B) 2 L fermentations. The absorbance at 280 nM is shown as a blue line. SDS / PAGE shows molecular size markers (M, sizes are indicated to the left of the gel), whole cell (W), and soluble (S) fractions. Fractions 1-9 correspond to fractions from the FPLC traces analyzed by SDS / PAGE.

[0042] DESCRIPTION OF THE PREFERRED EMBODIMENTS To provide a more accurate understanding of the gist of the present invention, features of the present invention will be described with reference to the following preferred embodiments.

[0043] Sequence search and creation of sequence similarity networks (SSNs)

[0044] In search of potential enzymes with improved thermostability, the sequences of Cl_EstA (accession number: WP_011948553.1), Cl_EstB (accession number: WP_011986581.1), and PfL1 (accession number: EIW29778.1) were used to perform pBLAST searches against the NCBI Refseq_Select_proteins database using default parameters. All sequences identified with an E-value greater than 0.005 were searched and submitted to the Enzyme Function Initiative Enzyme Similarity Tool (EFI-EST) to generate a sequence similarity network (SSN) including only sequences between 200 and 1000 amino acids in length, with an initial alignment score cutoff of 7. In SSN, nodes represent individual proteins, and edges represent alignment scores, which are calculated by the EFI-EST algorithm using bit scores from an all-vs-all BLAST run of the provided protein sequences. The alignment score is approximately the negative logarithm of the BLAST E-value, which approximates protein similarity.

[0045] SSNs were visualized using Cytoscape 3.9.0 by applying the yfiles Organic Layout. Clustering was analyzed by gradually increasing the alignment score cutoff value and then reapplying the layout. No significant changes were observed until a large jump in the cutoff value was observed, from alignment scores of 20 to 40 for this SSN. This visual clustering method allows for the identification of potentially homofunctional groups from other related proteins. Only sequences belonging to large clusters containing the query protein were then used for further SSN and phylogenetic analysis. Further adjustments to the alignment score cutoff value and subsequent reapplication of the yfiles Organic Layout were performed to visualize clades within larger protein families.

[0046] Phylogenetic analysis identifying sequences from thermophiles and extremophiles

[0047] The refined sequence set from SSN was then aligned using the structure-based protein sequence alignment algorithm PROMALS3D, using the structures of Cl_EstA (PDB ID: 5AH1) and PfL1 (PDB ID: 5AH0) as templates. Regions with poor alignment at the N- and C-termini, as well as long insertions in otherwise well-aligned sequences, were removed, and sequences with poor overall alignment or large deletions were completely deleted. This curated sequence set was realigned using MUSCLE with default settings on the EMBL-EBI web server (https: / / www.ebi.ac.uk / Tools / msa / muscle / ). The resulting alignment was further trimmed to remove poorly aligned regions at the N-terminus and large gaps or insertions, and then used to generate an initial approximate maximum likelihood (ML) tree using FastTree 2.1 with the WAG+CAT evolutionary model. The multiple sequence alignment (MSA) was further curated to remove any long branches along with highly similar sequences with very short branch lengths. This process was repeated until the SH-like local support value exceeded 80% at most nodes. The resulting alignment was used to estimate a maximum likelihood tree using the IQ-TREE web service (https: / / www.hiv.lanl.gov / content / sequence / IQTREE / iqtree.html). The best evolutionary model calculated was the following: LG model from the data + FreeRate heterogeneity (#rate categories=8) + ML-optimized AA frequency. Ultrafast bootstrap (UF-Boot) values ​​and SH-aLRT branch test values ​​were calculated for 1,000 replicates. The tree was visualized and annotated using the Interactive Tree of Life (iTOL) web tool (https: / / itol.embl.de / ).

[0048] A maximum likelihood phylogenetic tree was constructed using sequences from bacterial lipase family 1.5 curated using the SSN (Figure 1). This analysis supported the SSN findings, suggesting that Cl_EstA, Cl_EstB, and PfL1 belong to a larger clade containing proteins from bacteria in the Clostridiaceae family. Within this clade, two subclades, including Cl_EstA and Cl_EstB, respectively, suggest that they arose from a gene duplication event in the last common ancestor of the Clostridiaceae family. As with the SSN, the tree supports horizontal gene transfer from the Cl_EstA clade to the last common ancestor of Pelosinus species, resulting in PfL1. As observed in the SSN, sequences from Clostridiaceae are most closely related to sequences from bacteria in the Thermoactinomycetaceae, Paenibacillaceae, and Alicyclobacillaceae families. Among these, Clostridia and Bacilli are also sequences from thermophilic and acidophilic Clostridia in the families Peptococcaceae and Syntrophomonadaceae. For Bacillus and Geobacillus species, a third large clade was observed, including Bacilli in the order Bacillales (e.g., Bacillus and Caldibacillus species) and Lactobacilli in the order Lactobacillales. Sequences from all other phyla, except for Firmicutes, form distinct clades, likely acquired via horizontal gene transfer, with the exception of three sequences from Betaproteobacteria and Bacteroidetes.

[0049] Based on the phylogenetic tree and SSN, functional analysis was limited to sequences from Clostridia and Bacilli of the order Bacillales, which showed closest homology to Cl_EstA, Cl_EstB, and PfL1. Sequences were selected from six thermophilic and one acidophilic species: Ct_Est from Caldibacillus thermoamylovorans (accession: WP_152032401.1), Gk_Est from Geobacillus kaustophilus (accession: WP_044733155.1), Gs_Est from Geobacillus stearothermophilus (accession: WP_095860225.1), Tl_Est47 from Thermosyntropha lipolytica (accession: WP_073088947.1), Tl_Est64 from Thermosyntropha lipolytica (accession: WP_014826614.1), and Desulfurispora Dt_Est from Desulfosporosinus thermophila (accession: WP_018085325.1), and Da_Est from Desulfosporosinus acidophilus (accession: WP_014826614.1).

[0050] The polyesterase homologue was expressed in E. coli and showed activity with the pNP substrate.

[0051] Protein expression

[0052] The seven identified protein sequences, along with the sequences of Cl_EstA, Cl_EstB, and PfL1, were entered into the SignalP-5.0 web server. After trimming the predicted signal sequences, expression vectors were designed and ordered from GenScript (Singapore). Each truncated sequence was placed between the Nde1 and Xho1 restriction sites of pET-29b(+), ensuring that the expressed protein contained an N-terminal His tag. Plasmids were transformed into NEB T7 expression cells (New England Biolabs) using the manufacturer's recommended protocol and plated on Luria Broth (LB) agar plates containing 50 μg / mL kanamycin. Plates were incubated overnight at 37°C and then stored at 4°C for up to two weeks. A negative control plasmid (gfasPurple-S125R-F162R-V44A-L123T_pETcc2) was also transformed and plated on LB agar plates containing 100 μg / mL ampicillin.

[0053] For protein expression, a single colony from each strain was inoculated into 10 ml of autoinduction medium (5 g yeast extract, 20 g tryptone, 85.5 mM NaCl, 22 mM KH2PO4, 42 mM Na2HPO4, 0.6% glycerol, 0.05% glucose, and 0.2% lactose) containing 50 μg / mL kanamycin (100 μg / mL ampicillin for the control strain) in a 50 ml tube. The cultures were grown at 37°C for 3-6 hours with shaking at 200 rpm, then incubated overnight at 30°C. The resulting cultures were either stored at 4°C for up to 2 weeks for whole-cell assays or spun down at 4000 x g for 10 minutes at 4°C, discarding the supernatant and storing the resulting pellet at -20°C until protein purification.

[0054] Selected proteins were expressed with an N-terminal His tag in the E. coli strain NEB T7 Express. The protein profiles of the total cell fractions (soluble and insoluble proteins) and the soluble fraction isolated from the cell culture were evaluated on SDS-PAGE gels.

[0055] Testing protein expression by SDS-PAGE

[0056] 500 μL from each culture was spun down in a 1.5 mL microfuge tube at 4000 × g for 10 minutes at 4°C, and the supernatant was discarded. The resulting cell pellet was suspended in 100 μL of lysis solution (50 mM Tris pH 8, 1× BugBuster Protein Extraction Reagent (Millipore), and approximately 33 nL of DNAse I) and left on ice for approximately 10 minutes. After lysis, 5 μL from each sample was mixed with 10 μL of 50 mM Tris H8 and 5 μL of 4× NuPAGE™ LDS Sample Buffer (Invitrogen). The remaining sample was spun down at 20,000 × g for 10 minutes at 4°C, and 15 μL of the supernatant was mixed with 5 μL of 4× NuPAGE™ LDS Sample Buffer (Invitrogen). Samples were heated to 90°C for 3 minutes, loaded onto precast NuPAGE™ 4–12% Bis-Tris gels (Invitrogen), and run in MES SDS running buffer (Invitrogen) for 30–40 minutes at 150 V. Gels were stained with AcquaStain Protein Gel Stain (Bulldog) for 30 minutes and destained in water.

[0057] Clear bands were only observed for Cl_EstA, Cl_EstB, Dt_Est, and Tl_Est64, due to the strong background bands of the expected size also observed in the negative control. Therefore, to confirm protein expression, an initial esterase activity assay was performed using p-nitrophenyl (pNP) acetate (Figure 2A). Activity was detected for all 10 proteins, confirming successful expression in E. coli.

[0058] To compare the relative activities of the proteins, the proteins were partially purified by small-scale nickel affinity chromatography (Fig. 2B).

[0059] Protein purification

[0060] For small-scale purification, pellets from 10 mL cultures were resuspended in 1 mL of lysis buffer containing 50 mM Tris, 300 mM NaCl, pH 8, and transferred to 2 mL microfuge tubes. Cells were lysed by sonication (Fisher Scientific, 5-second pulse, 1-second break, 30-second intervals, 3 times). The lysate was then spun at 20,000 × g for 20 minutes at 4°C, and the supernatant was loaded onto NEBExpress® Ni Spin Columns (New England Biolabs) prewashed with 250 μL of the same lysis buffer. The columns were washed with a total of 750 μL of wash buffer (50 mM Tris, 300 mM NaCl, 5 mM imidazole, pH 8), and the sample was eluted with 2 × 200 μL of elution buffer (50 mM Tris, 300 mM NaCl, 500 mM imidazole, pH 8). From each eluate, 15 μL was mixed with 5 μL of 4× NuPAGE™ LDS Sample Buffer (Invitrogen) and subjected to SDS-PAGE analysis as described above. Purified proteins were stored at 4°C for up to 3 weeks.

[0061] Bands with approximately 90% or greater purity were observed for Cl_EstA, PfL1, Da_Est, Dt_Est, Tl_Est47, and Tl_Est64. For Cl_EstB and Ct_Est, only approximately 40-50% purity was obtained, but clear bands of the expected size were still observed. In contrast, for Gs_Est and Gk_Est, although pNP acetate activity was observed in whole-cell samples, the protein purified with this protocol was very low, almost undetectable.

[0062] Esterase activity assay using pNP-substrate.

[0063] The purified proteins were then used to characterize the substrate preferences of these proteins and compare their activity with pNP acetate, pNP propionate, pNP butyrate, pNP verarate, and pNP octonoate (Figure 2C).

[0064] To confirm protein activity using pNP substrate, 90 μL of 50 mM Tris (pH 8.0) was first mixed with 5 μL of cell culture for whole-cell assays or 5 μL of 200 nM purified protein. To initiate the reaction, 5 μL of 15 mM pNP-acetate, pNP-propionate, pNP-butyrate, pNP-velerate, or pNP-octonoate in 100% methanol was added, resulting in a final reaction solution containing 5% methanol. Absorbance at 405 nm was measured at 3-4 second intervals over 10 minutes, using a rate within the linear portion of the curve (typically within the first 30-100 seconds) to obtain a 16,853 nm absorbance over a 0.25 cm path length. -1 cm -1 The amount of pNP produced was calculated using the extinction coefficient.

[0065] A general trend of increasing activity with increasing substrate length was observed for all proteins except Cl_EstA, which showed a contrasting trend with highest activity with pNP acetate. Dt_Est showed the highest activity with the longest substrate tested. Tl_Est64 showed the highest activity with pNP butyrate and pNP valerate among all proteins. As expected, Cl_EstB, Ct_Est, Gk_Est, and Gk_Est, which were poorly purified, showed very low activity with all substrates. Despite successful purification, Da_Est from an acidophilic organism showed only slight activity with pNP octonoate and no activity with other substrates (Figure 2B). Furthermore, when cell cultures were tested, it also showed only slight activity with pNP acetate.

[0066] Cl_EstA, PfL1, Dt_Est, Tl_Est47 and Tl_Est64 have high heat resistance.

[0067] The thermostability of these enzymes was then tested by examining the residual activity of cell cultures heated at various temperatures for 10 min.

[0068] For each enzyme (and negative control culture), 8 x 15 μL of cell culture was added to 0.2 ml tubes in strips and heated to various temperatures for 10 minutes using the temperature gradient function of a thermocycler. The tubes were then cooled on ice for an additional 10 minutes. Residual activity was measured using the pNP-substrate assay described above using either pNP-acetate, pNP-propionate, or pNP-butyrate. Five μL of heated and cooled cell culture samples were used; a positive control containing 5 μL of unheated cell culture and a negative control containing 5 μL of buffer were also set up for each protein. Reaction rates were determined in mOD / min and used to calculate the percentage of remaining activity compared to the unheated positive control sample.

[0069] Initial studies showed clear residual activity with pNP acetate for Cl_EstA, PfL1, Dt_Est, Gk_Est, Tl_Est47, and Tl_Est64 after heating above 60°C. Cells expressing these proteins were further evaluated over a range of temperatures to determine their melting temperatures (T m ) was estimated (Fig. 3). The highest temperatures at which at least 10% residual activity was detected with pNP acetate were as follows: >10% for Dt_Est, Pfl1, and Tl_Est64 after heating at 78°C, >50% for Cl_EstA after heating at 70°C, and >20% for Gk_Est after heating at 65.75°C.

[0070] Interestingly, Tl_Est47 showed >10% residual activity with pNP-acetate even after heating at 85°C, which was further evaluated by repeating the assay for this enzyme using pNP-propionate. Although Tl_Est47 had almost no detectable activity with pNP-acetate under these reaction conditions, the reaction rate increased approximately fourfold with pNP-propionate as a substrate (Figure 3), increasing the sensitivity of the assay. For Tl_Est47, >10% residual activity with pNP-propionate was observed after heating at 90°C, making it the most thermostable enzyme evaluated in these experiments.

[0071] Evolution of Tl_Est47B for improved heat resistance.

[0072] We next investigated the possibility of further improving the thermostability of Tl_Est47, which retains 10% of its enzymatic activity after heating at 90°C for 10 min. To do this, we used error-prone PCR to generate a random mutation library, resulting in an estimated library size of 13,000 variants with an average of 2–12 mutations per variant. The mutants were grown overnight in LB medium in a 96-well growth block and heated at 89°C, 90°C, and 91°C for 10 min. Residual activity was measured using pNP-butyrate (Figure 4). The 10 most active mutants were then further screened, and mutant P4G12 was found to retain residual activity even after heating at 99°C for 10 min. This represents an improvement of approximately 5–10°C compared to the WT Tl_Est47 protein (Figure 5).

[0073] Large-scale purification of Tl_Est47 (WT) and the P4G12 mutant was performed to compare the two proteins. First, their activity with pNP-butyrate (Figures 6A and 6B) was compared, revealing a significantly higher turnover rate (k cat ) was reduced threefold. M The two-fold decrease in k suggests an improved substrate affinity by the mutant and therefore an overall catalytic efficiency (k ) of the WT. cat / KM ) is only 1.5 times higher than the mutant.

[0074] Generally, higher thermostability is achieved by mutations that stabilize proteins, which can reduce their overall movement in solution and decrease their activity by affecting the rates of substrate binding and diffusion as well as the overall catalytic mechanism.

[0075] We also confirmed the thermostability of the purified WT and mutant proteins (Figure 7). Similar to the results reported for the whole-cell assay, we observed that the mutants exhibited improved thermostability of approximately 5–7°C when comparing residual activity after heating in LB medium. The overall thermostability of the WT protein was improved by approximately 5°C in LB medium compared to buffer.

[0076] Performance of Tl_Est47 mutants at different pH and temperatures

[0077] The performance of enzymes was investigated over a range of pH and temperature. Some enzymes, when transferred to industrial-scale production systems (i.e., fermenters), can result in a significant decrease in yield and an increase in the cost of goods (CoG) of the enzyme-derived product. Therefore, the production of enzymes by fermentation was investigated and compared with shake-flask production to determine the enzyme yield under these conditions.

[0078] method

[0079] Protein expression

[0080] For protein expression, a single colony of each strain was inoculated into 10 ml of autoinduction medium (5 g yeast extract, 20 g tryptone, 85.5 mM NaCl, 22 mM KH2PO4, 42 mM Na2HPO4, 0.6% glycerol, 0.05% glucose, and 0.2% lactose) containing 50 μg / mL kanamycin (100 μg / mL ampicillin for the control strain) in a 50 ml tube. The cultures were grown at 37°C for 3-6 hours with shaking at 200 rpm, then incubated overnight at 30°C. The resulting cultures were either stored at 4°C for up to 2 weeks for whole-cell assays or spun down at 4000 x g for 10 minutes at 4°C, the supernatant discarded, and the resulting pellet stored at -20°C until protein purification.

[0081] Testing protein expression by SDS-PAGE

[0082] 500 μL from each culture was spun down in a 1.5 mL microfuge tube at 4000 × g for 10 minutes at 4 °C, and the supernatant was discarded. The resulting cell pellet was resuspended in 100 μL of lysis solution (50 mM Tris pH 8, 1 × BugBuster Protein Extraction Reagent (Millipore), and approximately 33 nL DNAse I) and left on ice for approximately 10 minutes. After lysis, 5 μL from each sample was mixed with 10 μL of 50 mM Tris H8 and 5 μL of 4 × NuPAGE™ LDS Sample Buffer (Invitrogen). The remaining sample was spun down at 20,000 × g for 10 minutes at 4 °C, and 15 μL of the supernatant was mixed with 5 μL of 4 × NuPAGE™ LDS Sample Buffer (Invitrogen). Samples were heated to 90°C for 3 minutes, loaded onto precast NuPAGE™ 4–12% Bis-Tris gels (Invitrogen), and run in MES SDS running buffer (Invitrogen) for 30–40 minutes at 150 V. Gels were stained with AcquaStain Protein Gel Stain (Bulldog) for 30 minutes and destained in water.

[0083] Protein purification

[0084] For small-scale purification, pellets from 10 mL cultures were resuspended in 1 mL of lysis buffer containing 50 mM Tris, 300 mM NaCl, pH 8, and transferred to 2 mL microfuge tubes. Cells were lysed by sonication (Fisher Scientific, 5-second pulse, 1-second break, 30-second intervals, 3 times). The lysate was then spun at 20,000 × g for 20 minutes at 4°C, and the supernatant was loaded onto NEBExpress® Ni Spin Columns (New England Biolabs) prewashed with 250 μL of the same lysis buffer. The columns were washed with a total of 750 μL of wash buffer (50 mM Tris, 300 mM NaCl, 5 mM imidazole, pH 8), and the samples were eluted with 2 × 200 μL of elution buffer (50 mM Tris, 300 mM NaCl, 500 mM imidazole, pH 8). From each eluate, 15 μL was mixed with 5 μL of 4× NuPAGE™ LDS Sample Buffer (Invitrogen) for SDS-PAGE analysis as described above. Purified proteins were stored at 4°C for up to 3 weeks.

[0085] PBSA degradation assay

[0086] First, a 10 mg / mL suspension of PBSA was prepared in buffer. 250 μL of the suspension was then dispensed into 1.5 mL or 2 mL tubes and assayed for each enzyme. For whole cell assays, an additional 250 μL of buffer was added to dilute the PBSA to 5 mg / mL, and 5 μL or 10 μL of cell suspension was added to initiate the reaction. For assays using purified enzymes, 250 μL of a 200 nM solution of each enzyme in the same buffer was added to the suspension, resulting in a final PBSA concentration of 5 mg / mL and a final enzyme concentration of 100 nM. A negative control using only buffer was also included.

[0087] Reactions were incubated with shaking at either room temperature or 37°C for 7 days. For the temperature assay, solutions were made using a buffer containing 50 mM Tris pH 8.0. Reactions were incubated at 4, 15, 25, and 37°C. The pH of each stock was adjusted to the desired pH using NaOH, as appropriate. Buffer stocks were diluted (1 in 10) with distilled water for the reactions.

[0088] After the incubation period, samples were spun down at 16,000 x g for 1 minute at room temperature to pellet any remaining plastic. Two 200 μL aliquots of supernatant from each tube were transferred to a 96-well UV plate (Grenier). To each aliquot, 10 μl of 50 mM NAD+ and 5 μl of 4 U / ml equine alcohol dehydrogenase were added, and the change in absorbance at 340 nm was measured at 30-second intervals for 30 minutes. Technical replicate values ​​were averaged, and experiments were repeated two to three times to calculate the standard error of the mean (SEM).

[0089] Comparison of fermentor and flask growth

[0090] To compare growth in batch and fermentor cultures, cultures were grown in TB medium or 2YT medium, respectively. 2YT medium (1 L): 5 g of yeast extract, 16 g of tryptone, and 5 g of NaCl were dissolved in 600 mL of distilled water, then made up to 1 L and sterilized at 121°C for 20 min. TB medium (2.5 L): Dissolve 12.5 g yeast extract (5 g / L), 50 g tryptone (20 g / L), 12.5 g NaCl (5 g / L), 7.5 g KH2PO4 (3 g / L), 14.9 g Na2HPO4 (5.96 g / L), and 0.6% glycerol (12.5 mL) in 2.5 L of distilled water. Then, distribute 500 mL into a sterile, baffled 2 L Erlenmeyer flask (with vent cap) or 2 L into a Sartorius Biostat B fermenter and sterilize at 121 °C for 60 minutes. After sterilization, add kanamycin to a final concentration of 50 mg / mL.

[0091] 500 mL flasks and 2 L fermentors were inoculated with a 10 mL overnight culture of the same medium inoculated with a single colony from an agar plate streaked with E. coli BL21 DE3 transformed with the appropriate expression plasmid.

[0092] Fermenter control parameters: pO2 setpoint = 30%, cascade: stirrer / airflow / O2 enrichment, initial settings: stirrer = 500 rpm, airflow = 0.3 L / min, O2 = 0.07 L / min, temperature setpoint = 37°C, pH setpoint = 7.0, acid / base setpoint: 10% H3PO4 / 10% NH3.

[0093] Samples for OD and 1 mL samples for enzyme analysis were taken throughout the process. At harvest, 20 mL cell pellets and 36 g total harvested pellets were stored frozen at -80°C along with 1 mL samples. Experimental records of the flask and fermenter conditions are shown in Figures 8 and 9.

[0094] Performance of variant Tl_Est47 at different pH and temperatures

[0095] Tl_Est47 and its variants were purified from E. coli BL21 DE3 cultures and used in this study. The purified proteins were incubated with PBSA for 7 days, and estimation of alcohol dehydrogenase activity was performed daily.

[0096] Both the wild-type and variants responded to pH as expected (Figures 10 and 11). Serine hydrolases tend to have pH optima above pH 8 due to the pKa of the serine nucleophile at the active site. Activity decreases at lower pH values.

[0097] The variant and wild-type enzymes exhibited similar activity at temperatures between 4°C and 37°C, indicating that the amino acid substitutions in the variant enzymes do not affect enzyme activity at low temperatures, and the enzymes exhibit activity at temperatures likely to reflect those encountered in their applications (Figures 12 and 13). As expected, the activity of both enzymes increased with increasing temperature, with activity approximately doubling when incubated at 37°C compared to 4°C.

[0098] Production of Tl_Est47 by fermentation

[0099] E. coli BL21 DE3 expressing the Tl_Est47 variant was cultured in a 500 mL flask and a 2 L fermenter, and the amount of Tl_Est47 produced under both conditions was estimated. The enzyme production capacity in the fermenter (2 L scale) was compared. The two culture systems were compared by sampling 1 mL of culture at 2, 4, 5.5, and 22 hours of expression and flash-frozen to -80°C in liquid nitrogen. The total protein content in each 1 mL sample was estimated (Table 1). After 5.5 hours, the total protein reached a plateau. After 22 hours of expression, 4.9 g of protein was harvested from the flask incubation and 22 g from the fermenter. [Table 1]

[0100] Table 1. Measurement of total protein produced during incubation in 500 mL flasks and 2 L fermentors.

[0101] After thawing on ice, samples were sonicated for 30 seconds and total cellular and soluble proteins were analyzed by SDS / PAGE (Figure 14).

[0102] After 22 hours of flask culture or fermentation, 6.5 g of pellet was obtained in the 500 mL flask and 36 g of pellet was obtained in the 2 L fermenter. Protein was purified by FPLC using 6.5 g of pellet from the flask and 6 g of pellet from the fermenter, using a His-trap column on an AKTA pure FPLC. Analysis of fractions by SDS-PAGE and traces are shown in Figure 15A / B.

[0103] Fractions 3-9 from each system were pooled and concentrated on an Amicon column (10 kDa cutoff), followed by buffer exchange to remove residual imidazole. Protein concentrations were estimated using Nanodrop (extinction coefficient at 280 nM: 124,470). Enzyme purified from flasks yielded 14 mL of purified enzyme at 0.9 mg / mL from 6.5 g of treated pellets. Enzyme produced in fermentors yielded 16 mL of purified enzyme at 2.6 mg / mL from 6 g of treated pellets. A summary of the enzyme yields from the two systems can be seen in Table 2. [Table 2]

[0104] Table 2. Purification overview

[0105] Protein yields were good, and production of the Tl_Est47 variant in fermentors improved both the percentage and quantity compared to the same protein produced in shake flasks. However, in both cases, the percentage of Tl_Est47 variant produced was lower than that observed for some heterologously expressed proteins with high production yields, suggesting that there is room for improvement in the expression level and production yield of the Tl_Est47 variant, if desired. This could ultimately reduce the cost of the commodity and final product for enzyme production. An additional 10 L fermentation was performed (estimated yield of Tl_Est47 variant: approximately 1.2 g), and the cell pellet was stored at -70°C.

[0106] The engineering techniques used to create the thermostable Tl_Est47 variant did not adversely affect the protein's stability or activity over a range of temperatures or pH values. Tl_Est47 maintained significant activity at low temperatures, suggesting that it will retain activity under conditions likely to be encountered during use. Heterologous production of the Tl_Est47 variant was significantly improved in fermentors compared to shake flasks. While expression levels were good, protein yields could potentially be further improved to reduce production costs.

[0107] Unless otherwise defined, all technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art (e.g., cell culture, molecular genetics, immunology, immunohistochemistry, protein chemistry, and biochemistry).

[0108] Unless otherwise indicated, nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation.

[0109] Unless otherwise indicated, the recombinant protein, cell culture, and immunological techniques utilized in the present invention are standard procedures, well known to those skilled in the art. Such techniques are detailed and explained throughout the literature, for example in the following sources: J. Perbal, A Practical Guide to Molecular Cloning, John Wiley and Sons (1984); J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press (1989); T. A. Brown (editor), Essential Molecular Biology: A Practical Approach, Volumes 1 and 2, IRL Press (1991); D. M. Glover and B. D. Hames (editors), DNA Cloning: A Practical Approach, Volumes 1-4, IRL Press (1995 and 1996); and F. M. Ausubel et al. (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all updates to date); Ed Harlow and David Lane (editors), Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, (1988); and J. E. Coligan et al. (editors) Current Protocols in Immunology, John Wiley & Sons (including all updates to date).

[0110] The term "and / or," e.g., "X and / or Y," shall be understood to mean either "X and Y" or "X or Y," and shall explicitly endorse both or either meanings. The definitions provided herein shall be interpreted in the context of the specification as a whole. As used herein, the singular forms "a," "an," "said," and "the" include the plural unless the context clearly indicates otherwise. For example, reference to "a protein" includes a plurality of proteins unless the context clearly indicates otherwise. As used herein, the term "protein" includes proteins, polypeptides, and peptides. In some embodiments, the terms "protein," "polypeptide," and "peptide" may be used interchangeably.

[0111] As used herein, the term about, unless stated to the contrary, refers to ±20%, more preferably ±10%, and even more preferably ±5% of the specified value. As used herein, each numerical range includes all smaller numerical ranges that fall within such broader numerical range, as if all such narrower numerical ranges were expressly written herein.

[0112] In this specification and the appended claims, when used in connection with one or more numerical values ​​or numerical ranges, the terms "about" or "substantially" should be understood to refer to all such numerical values, including every numerical value within the range, and to modify that range by extending the boundaries above and below the recited numerical values. The recitation of numerical ranges by endpoints includes all numbers subsumed within that range, e.g., integers including fractions thereof (e.g., recitation of 1 to 5 includes 1, 2, 3, 4, 5, and fractions thereof, e.g., 1.5, 2.25, 3.75, 4.1, etc.), and any range within that range.

[0113] The terms "polypeptide" and "protein" are generally used interchangeably and refer to a single polypeptide chain (a polymeric sequence of amino acid residues) that may or may not be modified by the addition of non-amino acid groups. It is understood that such a polypeptide chain may be linked to other molecules, such as other polypeptides or proteins or co-factors. As used herein, the terms "protein" and "polypeptide" also include variants, mutants, biologically active fragments, and / or modifications of the polypeptides described herein. One-letter and three-letter codes for amino acids, as defined in accordance with the IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN), are used throughout this disclosure. The single letter X refers to any of the 20 amino acids. It is also understood that a polypeptide may be encoded by more than one nucleotide sequence due to the degeneracy of the genetic code.

[0114] The percent identity of polypeptides may be determined by GAP (Needleman and Wunsch, 1970) analysis (GCG program) with a gap creation penalty of 5 and a gap extension penalty of 0.3. The query sequence is at least 250 amino acids in length, and the GAP analysis aligns the two sequences over a region of at least 250 amino acids. More preferably, the query sequence is at least 300 amino acids in length, and the GAP analysis aligns the two sequences over a region of at least 300 amino acids. Even more preferably, the query sequence is at least 350 amino acids in length, and the GAP analysis aligns the two sequences over a region of at least 350 amino acids. Even more preferably, the GAP analysis aligns the two sequences over their entire lengths.

[0115] As used herein, the phrase "at a position corresponding to an amino acid number" refers to the relative position of an amino acid compared to surrounding amino acids with reference to a defined amino acid sequence. For example, in some embodiments, polypeptides of the invention may have additional N-terminal amino acids to aid in intracellular localization or extracellular secretion, which alters the relative position of the amino acid when aligned to, for example, SEQ ID NO:3 or SEQ ID NO:4.

[0116] The term "mature" form of a protein, polypeptide, or peptide refers to the functional form of the protein, polypeptide, or peptide that does not include signal peptide and propeptide sequences.

[0117] As used herein with respect to an amino acid residue position, "corresponding to" or "corresponds to" or "corresponds" refers to the amino acid residue at the recited position in a protein or peptide, or an amino acid residue that is similar, homologous, or equivalent to the recited residue in a protein or peptide. As used herein, a "corresponding region" generally refers to an analogous position in a related or reference protein.

[0118] The term "wild-type" with respect to an amino acid sequence or a nucleic acid sequence indicates that the amino acid sequence or nucleic acid sequence is a native or naturally-occurring sequence. As used herein, the term "naturally occurring" refers to something that exists in nature (e.g., a protein or polynucleotide sequence). Conversely, the term "non-naturally occurring" refers to something that is not found in nature (e.g., recombinant polynucleotide and protein sequences produced in the laboratory, or modifications of a wild-type sequence).

[0119] As used herein, the term "improved thermostability" or "enhanced thermotolerance" or "increased thermotolerance" refers to novel polypeptides that exhibit increased retention of enzymatic activity after a period of incubation at elevated temperatures, particularly relative to the wild-type form of a similar enzyme or polypeptide. Furthermore, the terms "improved thermotolerance properties" and "thermostability" are used interchangeably in the specification and claims.

[0120] The term "variant" with respect to a polypeptide amino acid sequence refers to an amino acid sequence of the polypeptide that differs from the amino acid sequence of a specified wild-type, parental, or reference polypeptide by containing one or more artificial substitutions, insertions, or deletions of amino acids. Similarly, the term "variant" with respect to a nucleic acid sequence of a polynucleotide refers to the nucleic acid sequence of the polynucleotide that differs from the nucleic acid sequence of a specified wild-type, parental, or reference polynucleotide by containing one or more artificial substitutions, insertions, or deletions of nucleic acids. The identity of the amino acid sequence of the wild-type, parental, or reference polypeptide or the nucleic acid sequence of the polynucleotide will be clear from the context.

[0121] As used herein, the term "mutation" or "engineering" refers to an artificial alteration to a reference amino acid or nucleic acid sequence. The term is intended to encompass artificial substitutions, insertions, and deletions.

[0122] As used herein, the term "vector" refers to a nucleic acid construct used to introduce or transfer nucleic acids into target cells or tissues. Vectors are typically used to introduce foreign DNA into cells or tissues. Vectors include plasmids, cloning vectors, bacteriophages, viruses (e.g., viral vectors), cosmids, expression vectors, shuttle vectors, and the like. Vectors typically contain an origin of replication, a multicloning site, and a selectable marker. The process of inserting a vector into a target cell is typically called transformation.

[0123] As used herein, in the context of introducing a nucleic acid sequence into a cell, the term "introduced" refers to any suitable method for incorporating a nucleic acid sequence into a cell. Such introduction methods include, but are not limited to, protoplast fusion, transfection, transformation, electroporation, conjugation, and transduction. Transformation refers to the genetic alteration of a cell resulting from the uptake, optional genomic integration, and expression of genetic material (e.g., DNA).

[0124] An "expression cassette" or "expression vector" refers to a nucleic acid construct or vector produced recombinantly or synthetically to express a nucleic acid of interest (e.g., a foreign nucleic acid or a transgene) in a target cell. The nucleic acid of interest typically expresses a protein of interest. An expression vector or expression cassette typically contains a promoter nucleotide sequence that drives or promotes the expression of the foreign nucleic acid. An expression vector or expression cassette also typically contains other specific nucleic acid elements that enable transcription of the specific nucleic acid in the target cell. Recombinant expression cassettes can be incorporated into plasmids, chromosomes, mitochondrial DNA, plastid DNA, viruses, or nucleic acid fragments. Some expression vectors are capable of integrating and expressing heterologous DNA fragments in a host cell or the genome of a host cell. Many prokaryotic and eukaryotic expression vectors are commercially available. Selection of an appropriate expression vector for protein expression from a nucleic acid sequence incorporated into the expression vector is within the knowledge of one of ordinary skill in the art.

[0125] As used herein, a nucleic acid is "operably linked" to another nucleic acid sequence when it is placed into a functional relationship with the other nucleic acid sequence. For example, a promoter or enhancer is operably linked to a nucleotide coding sequence if the promoter affects the transcription of the coding sequence. A ribosome binding site may be operably linked to a coding sequence if it is positioned so as to promote translation of the coding sequence. Typically, "operably linked" DNA sequences are contiguous; however, enhancers need not be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adapters or linkers can be used in accordance with conventional methods.

[0126] As used herein, the term "gene" refers to a polynucleotide (e.g., a DNA segment) that encodes a polypeptide and includes regions before and after the coding region. In some cases, a gene contains intervening sequences (introns) between individual coding segments (exons).

[0127] As used herein, "recombinant" when used with respect to cells typically indicates that the cell has been modified by the introduction of an exogenous nucleic acid sequence or that the cell is derived from a cell so modified. For example, a recombinant cell may contain a gene not found in the same form within the native (non-recombinant) form of the cell, or a recombinant cell may contain a native gene (found in the native form of the cell) that has been altered and reintroduced into the cell. A recombinant cell may contain nucleic acid endogenous to the cell that has been modified without removing the nucleic acid from the cell; such modifications include those obtained by gene replacement, site-specific mutagenesis, and related techniques known to those of skill in the art. Recombinant DNA technology includes techniques for producing recombinant DNA in vitro and introducing recombinant DNA into cells where it can be expressed or propagated, thereby producing a recombinant polypeptide. "Recombination" and "recombining" of polynucleotides or nucleic acids generally refer to the assembly or joining of two or more nucleic acids or polynucleotide strands or fragments to produce a new polynucleotide or nucleic acid.

[0128] A nucleic acid or polynucleotide is said to "encode" a polypeptide if, in its natural state, or when manipulated by methods known to those of skill in the art, it can be transcribed and / or translated to produce a polypeptide or fragment thereof. The antisense strand of such a nucleic acid is also said to be encoding the sequence.

[0129] The terms "host strain" and "host cell" refer to a suitable host for an expression vector containing a DNA sequence of interest.

[0130] The term "precursor" form of a protein or peptide refers to the mature form of the protein having a prosequence operably linked to the amino- or carboxyl-terminus of the protein. A precursor may also have a "signal" sequence operably linked to the amino-terminus of the prosequence. A precursor may also have additional polypeptides involved in post-translational activity (e.g., polypeptides that are cleaved from them to leave the mature form of the protein or peptide).

[0131] The terms "derived" and "obtained from" refer not only to proteins produced or producible by the strain of organism in question, but also to proteins encoded by DNA sequences isolated from such strains and produced in host organisms containing such DNA sequences. Additionally, the terms refer to proteins encoded by DNA sequences of synthetic and / or cDNA origin and having the identifying characteristics of the protein in question.

[0132] The term "identical" in the context of two polynucleotide or polypeptide sequences refers to nucleic acids or amino acids in the two sequences that are identical when aligned for maximum correspondence, as measured using sequence comparison or analysis algorithms known in the art, as described below.

[0133] The term "% identity" or "percent identity" or "PID" refers to protein sequence identity. Percent identity can be determined using standard techniques known in the art. The percent identity of amino acids shared by sequences of interest can be determined by aligning sequences to directly compare sequence information, for example, by using programs such as BLAST, MUSCLE, or CLUSTAL. The BLAST algorithm is described, for example, in Altschul et al., J Mol Biol, 215:403-410 (1990) and Karlin et al., Proc Natl Acad Sci USA, 90:5873-5787 (1993). The percent (%) amino acid sequence identity value is determined by dividing the number of matching identical residues by the total number of residues in the "reference" sequence, including gaps created by the optimal / maximal alignment program. The BLAST algorithm refers to the "reference" sequence as the "query" sequence.

[0134] The CLUSTAL W algorithm is another example of a sequence alignment algorithm (see Thompson et al., Nucleic Acids Res, 22:4673-4680, 1994). The default parameters for the CLUSTAL W algorithm are as follows: Gap opening penalty = 10.0; Gap extension penalty = 0.05; Protein weight matrix = BLOSUM series; DNA weight matrix = IUB; Delay divergent sequences % = 40; Gap separation distance = 8; DNA transitions weight = 0.50; List hydrophilic residues = GPSNDQEKR; Use negative matrix = OFF; Toggle Residue specific penalties = ON; Toggle hydrophilic penalties = ON; and Toggle end gap separation penalty = OFF. The CLUSTAL algorithm also includes deletions that occur at either end. For example, a variant with 5 amino acids deleted at either end of a 500 amino acid polypeptide (or within the polypeptide) would have a percent sequence identity of 99% (495 / 500 identical residues x 100) to the "reference" polypeptide. Such a variant is encompassed by mutants having "at least 99% sequence identity" to the polypeptide.

[0135] Understanding homology between molecules can reveal information about a molecule's evolutionary history and function. If a newly sequenced protein is homologous to a previously characterized protein, it strongly suggests the biochemical function of the new protein. The most fundamental relationship between two entities is homology; two molecules are said to be homologous if they are derived from a common ancestor. Homologous molecules, or homologs, are divided into two classes: paralogs and orthologs. Paralogs are homologs that exist within a single species. Paralogs often differ in their detailed biochemical function. Orthologs are homologs that exist in different species and have very similar or identical functions. Protein superfamilies are the largest groups (clades) of proteins for which a common ancestor can be inferred. This common ancestor is usually based on sequence alignment and mechanistic similarity. Superfamilies typically contain several protein families that show sequence similarity within the family. The term "protein clan" is commonly used to refer to protease superfamilies based on the MEROPS protease classification system.

[0136] A nucleic acid or polynucleotide is "isolated" if it is at least partially or completely separated from other components, including, but not limited to, other proteins, nucleic acids, cells, etc. Similarly, a polypeptide, protein, or peptide is "isolated" if it is at least partially or completely separated from other components, including, but not limited to, other proteins, nucleic acids, cells, etc. On a molar basis, an isolated species is more abundant than other species in a composition. For example, an isolated species may comprise at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% (on a molar basis) of all macromolecular species present. Preferably, the species of interest is purified to essential homogeneity (i.e., contaminating species cannot be detected in the composition by conventional detection methods). Purity and homogeneity can be determined using a number of techniques well known in the art, such as subjecting nucleic acid or protein samples to electrophoresis on agarose or polyacrylamide gels, respectively, followed by visualization by staining. If desired, high-resolution techniques, such as high-performance liquid chromatography (HPLC) or similar means, can be used to purify materials.

[0137] The term "purified," as applied to a nucleic acid or polypeptide, generally refers to a nucleic acid or polypeptide that is essentially free from other components, as determined by analytical techniques well known in the art (e.g., a purified polypeptide or polynucleotide forms a distinct band in an electrophoretic gel, a chromatographic eluate, and / or a medium subjected to density gradient centrifugation). For example, a nucleic acid or polypeptide that gives rise to essentially one band in an electrophoretic gel is "purified." A purified nucleic acid or polypeptide is at least about 50% pure, usually at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8% or more pure (e.g., percent by weight on a molar basis). In a related sense, a composition is enriched for a molecule if the concentration of the molecule is substantially increased after application of a purification or concentration technique. The term "enriched" refers to a compound, polypeptide, cell, nucleic acid, amino acid, or other particular substance or component that is present in a composition at a higher relative or absolute concentration than in the starting composition.

[0138] One or more novel polypeptide variants described herein may be modified in various ways, such as conservative or non-conservative insertions, deletions, and / or substitutions of one or more amino acids, including cases where such modifications do not substantially alter the enzymatic activity of the variant. Similarly, the nucleic acids of the present invention may also be modified in various ways, such as: one or more substitutions of one or more nucleotides in one or more codons, such that a particular codon encodes the same or a different amino acid, resulting in either a silent mutation (e.g., when the encoded amino acid is unchanged by the nucleotide mutation) or a non-silent mutation; one or more deletions of one or more nucleic acids (or codons) in the sequence; one or more additions or insertions of one or more nucleic acids (or codons) in the sequence; and / or one or more truncations of one or more nucleic acids (or codons) in the sequence. Many such modifications in nucleic acid sequences may not substantially alter the enzymatic activity of the resulting encoded polypeptide enzyme compared to the polypeptide enzyme encoded by the original nucleic acid sequence. The nucleic acid sequences described herein can also be modified to include one or more codons that provide for optimal expression in an expression system (e.g., a bacterial expression system), while, if desired, one or more of the codons described above still encode the same amino acid.

[0139] One or more nucleic acid sequences described herein can be produced by using any suitable synthesis, manipulation, and / or isolation technique, or a combination thereof. For example, one or more polynucleotides described herein can be produced using standard nucleic acid synthesis techniques, such as solid-phase synthesis techniques well known to those of skill in the art. In such techniques, fragments of typically up to 50 or more nucleotide bases are synthesized and then joined (e.g., by enzymatic or chemical ligation) to form essentially any desired contiguous nucleic acid sequence. Synthesis of one or more polynucleotides described herein can also be facilitated by any suitable method known in the art, including, but not limited to, chemical synthesis using the classical phosphoramidite method (see, e.g., Beaucage et al. Tetrahedron Letters 22:1859-69 (1981)) or the method described in Matthes et al., EMBO J.3:801-805 (1984), as typically performed in automated synthesis. One or more polynucleotides described herein can also be produced using an automated DNA synthesizer. Customized nucleic acids can be ordered from various commercial sources (e.g., Midland Certified Reagent Company, Great American Gene Company, Operon Technologies Inc., and DNA 2.0). Other techniques and related principles for synthesizing nucleic acids are described, for example, in Itakura et al., Ann. Rev. Biochem. 53:323 (1984) and Itakura et al., Science 198:1056 (1984).

[0140] Further embodiments are directed to one or more vectors comprising one or more novel polypeptide variants described herein (e.g., polynucleotides encoding one or more novel polypeptide variants described herein); expression vectors or expression cassettes comprising one or more nucleic acid or polynucleotide sequences described herein; isolated, substantially pure, or recombinant DNA constructs comprising one or more nucleic acid or polynucleotide sequences described herein; isolated or recombinant cells comprising one or more polynucleotide sequences described herein; and compositions comprising one or more of the above vectors, nucleic acids, expression vectors, expression cassettes, DNA constructs, cells, cell cultures, or any combination or mixture thereof.

[0141] Some embodiments are directed to one or more recombinant cells comprising one or more vectors (e.g., expression vectors or DNA constructs) described herein, which contain one or more nucleic acid or polynucleotide sequences described herein. Some such recombinant cells are transformed or transfected with at least one such vector, although other methods are available and known in the art. Such cells are typically referred to as host cells. Some such cells include bacterial cells. Other embodiments are directed to recombinant cells (e.g., recombinant host cells) comprising one or more novel polypeptides described herein.

[0142] In some embodiments, one or more vectors described herein are expression vectors or expression cassettes comprising one or more polynucleotide sequences described herein operably linked to one or more additional nucleic acid segments required for efficient gene expression (e.g., a promoter operably linked to one or more polynucleotide sequences described herein). The vector may contain a transcription terminator and / or a selection gene (e.g., an antibiotic resistance gene) that allows for continued culture maintenance of plasmid-infected host cells by growth in antimicrobial-containing medium. Expression vectors may be derived from plasmid or viral DNA, or in alternative embodiments, contain elements of both.

[0143] For expression and production of a protein of interest (e.g., one or more novel polypeptides described herein) in a cell, one or more expression vectors containing one or more copies, and in some cases multiple copies, of a polynucleotide encoding one or more novel polypeptides described herein are transformed into the cell under conditions suitable for expression of the novel polypeptides. In some embodiments, the polynucleotide sequence encoding one or more novel polypeptides described herein (as well as other sequences contained in the vector) integrates into the genome of the host cell, while in other embodiments, the plasmid vector containing the polynucleotide sequence encoding one or more novel polypeptides described herein remains within the cell as an autonomous extrachromosomal element. Some embodiments provide both extrachromosomal nucleic acid elements as well as incoming nucleotide sequences that integrate into the host cell genome. The vectors described herein are useful for producing the novel polypeptides described herein. In some embodiments, the polynucleotide construct encoding one or more novel polypeptides described herein is present on an integrating vector, which allows for integration and, optionally, amplification of the polynucleotide encoding the novel polypeptide into the host chromosome. Exemplary sites for integration are well known to those of skill in the art. In some embodiments, transcription of the polynucleotide encoding the novel polypeptide described herein is driven by a promoter that is the wild-type promoter for the wild-type polypeptide. In some other embodiments, the promoter is heterologous to one or more of the novel polypeptides described herein but is functional in the host cell.

[0144] In addition to commonly used methods, in some embodiments, host cells are directly transformed with a DNA construct or vector containing a nucleic acid encoding one or more novel polypeptides described herein (i.e., no intermediate cells are used to amplify or otherwise manipulate the DNA construct or vector before introduction into the host cell). Introduction of the DNA construct or vector described herein into a host cell includes physical and chemical methods known in the art for introducing a nucleic acid sequence (e.g., a DNA sequence) into a host cell without inserting it into the host genome. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, and liposomes. In a further embodiment, the DNA construct or vector is co-transformed with a plasmid without being inserted into the plasmid. In a further embodiment, the selectable marker is removed from the modified bacterial strain by methods known in the art (see Stahl et al., J. Bacteriol. 158:411-418 (1984); and Palmeros et al., Gene 247:255-264 (2000)). In some embodiments, the transformed cells are cultured in conventional nutrient media. Suitable specific culture conditions, such as temperature, pH, etc., are known to those of skill in the art and are well described in the scientific literature.

[0145] As used herein, the term "extract" refers to any part of a host cell or non-human transgenic organism of the present invention that contains a polypeptide of the present invention, and preferably also contains a polynucleotide or vector of the present invention. This term includes parts secreted from host cells, and thus encompasses culture supernatants. Preferably, the extract is a relatively crude extract that has not undergone a purification process to purify the polypeptide of the present invention while avoiding other polypeptides produced together with the polypeptide of the present invention. The extract may also be a composition containing the polypeptide of the present invention.

[0146] As used herein, a "biologically active fragment" refers to a portion of a polypeptide described herein that maintains a defined activity of the full-length polypeptide. Biologically active fragments can be of any size, so long as they maintain the defined activity.

[0147] Thus, where applicable, in light of minimum % identity figures, it is preferred that the polypeptide comprises an amino acid sequence that is at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, and even more preferably at least 99.9% identical to the corresponding designated SEQ ID NO. It will be understood that, with respect to defined polypeptides, % identity figures higher than those provided herein encompass preferred embodiments.

[0148] "Substantially purified" or "purified" refers to a polypeptide that has been separated from one or more lipids, nucleic acids, other polypeptides, or other contaminating molecules with which it is naturally associated. Preferably, a substantially purified polypeptide is at least 60% free, more preferably at least 75%, and even more preferably at least 90% free from other components with which it is naturally associated. Although there is currently no evidence that the polypeptides of the invention exist in nature, the terms native state and naturally associated also encompass polypeptides produced in the host cells of the invention.

[0149] The term "recombinant" in the context of a polypeptide refers to a polypeptide when produced by a cell or in a cell-free expression system in an altered amount or at an altered rate compared to its native state. In one embodiment, the cell is a cell that does not naturally produce the polypeptide. Recombinant polypeptides of the invention include polypeptides that have not been separated from other components of the transgenic (recombinant) cells or cell-free expression system in which they are produced, as well as polypeptides produced in such cells or cell-free systems and subsequently separated and purified from at least some other components.

[0150] Mutants or variants of the amino acid sequences of the polypeptides described herein can be prepared by introducing appropriate nucleotide changes into the nucleic acids defined herein or by in vitro synthesis of the desired polypeptide. Such variants include, for example, deletions, insertions, or substitutions of residues within the amino acid sequence. A combination of deletions, insertions, and substitutions can be used to arrive at a final construct, provided that the final polypeptide product possesses the desired properties.

[0151] Mutant or variant polypeptides can be prepared using any technique known in the art, for example, using directed evolution or rational design strategies (see below). Products derived from mutated / altered DNA can be readily screened to determine whether they have enzymatic activity using the techniques described herein.

[0152] When designing variants of an amino acid sequence, the location of the mutation site and the nature of the mutation can vary depending on the property to be modified. The mutation sites can be modified individually or sequentially, for example, by (1) initially selecting conservative amino acids, followed by substitution with more radical choices depending on the results obtained, (2) deleting the target residue, or (3) inserting other residues adjacent to the localized site.

[0153] Amino acid sequence deletions generally range from about 1 to 15 contiguous residues, more preferably about 1 to 10 residues, and typically about 1 to 5 contiguous residues.

[0154] Substitutional variants have at least one amino acid residue in the polypeptide molecule removed and replaced with another residue. Interesting sites are those where specific residues from various strains or species are identical. These sites may be important for biological activity. These sites, particularly those within a sequence of at least three other identically conserved sites, are preferably substituted in a relatively conservative manner.

[0155] In preferred embodiments, mutant / variant polypeptides have only conservative substitutions compared to the novel polypeptides specifically defined herein, hi preferred embodiments, mutant / variant polypeptides have one or two or three or four conservative amino acid changes when compared to the novel polypeptides specifically defined herein.

[0156] Preferably, unless otherwise specified, at a given amino acid position, the novel polypeptide contains the amino acid found at the corresponding position in the polypeptide provided as SEQ ID NO:4.

[0157] Also included within the scope of the present invention are novel polypeptides of the present invention that have been differentially modified during or after synthesis, for example, by biotinylation, benzylation, glycosylation, acetylation, phosphorylation, amidation, derivatization with known protecting / blocking groups, proteolytic cleavage, linkage to antibody molecules or other cellular ligands, etc. These modifications may serve to increase the stability and / or biological activity of the polypeptide.

[0158] The novel polypeptides described herein can be produced in a variety of ways, including recombinant polypeptide production and recovery and chemical polypeptide synthesis. In one embodiment, an isolated polypeptide of the present invention is produced by culturing cells capable of expressing the polypeptide under conditions effective to produce the polypeptide and recovering the polypeptide. Preferred cells for culturing are recombinant cells of the present invention. Effective culture conditions include, but are not limited to, effective media, bioreactors, temperature, pH, and oxygen conditions that permit polypeptide production. An effective medium refers to any medium in which cells can be cultured to produce a polypeptide of the present invention. Such media typically comprise an aqueous medium having assimilable carbon, nitrogen, and phosphate sources, as well as appropriate salts, minerals, metals, and other nutrients, such as vitamins. Cells of the present invention can be cultured in conventional fermentation bioreactors, shake flasks, test tubes, microtiter dishes, or Petri plates. Culturing can be carried out at temperatures, pH, and oxygen content appropriate for the recombinant cells. Such culture conditions are within the expertise of one of ordinary skill in the art.

[0159] In one embodiment, the novel polypeptides of the present invention contain a signal sequence capable of directing secretion of the polypeptide from a cell. As one skilled in the art will appreciate, the signal sequence may or may not be cleaved, or may be partially cleaved, while still being partially exported from the cell nucleus. However, when the signal sequence is removed, the cell may produce a heterogeneous population of polypeptides with slightly different, e.g., N-terminal sequences. Thus, the term "consisting of" encompasses such variants generated by signal sequence removal. Many such signal sequences have been isolated, including N- and C-terminal signal sequences. N-terminal signal sequences in prokaryotes and eukaryotes are similar, and it has been shown that eukaryotic N-terminal signal sequences can function as secretory sequences in bacteria. One example of such an N-terminal signal sequence is the bacterial β-lactamase signal sequence, which is a well-studied sequence and is widely used to promote secretion of polypeptides into the external environment. One example of a C-terminal signal sequence is the hemolysin A (hlyA) signal sequence from Escherichia coli. Further examples of signal sequences include, but are not limited to, aerolysin, alkaline phosphatase gene (phoA), chitinase, endochitinase, α-hemolysin, MIpB, pullulanase, Yops, and TAT signal peptides.

[0160] As used herein, "isolated polynucleotide" refers to a polynucleotide that is at least partially separated from polynucleotide sequences with which it is naturally associated or linked. Isolated polynucleotides include DNA and RNA molecules, as well as molecules that are combinations of DNA and RNA. They may be single-stranded, double-stranded, or partially double-stranded, and may be in a sense or antisense orientation relative to the promoter. Preferably, an isolated polynucleotide is at least 60%, preferably at least 75%, and most preferably at least 90% free from other components with which it is naturally associated. Furthermore, the term "polynucleotide" is used interchangeably with the term "nucleic acid" herein.

[0161] The term "exogenous" in the context of a polynucleotide refers to a polynucleotide when present in a cell or in a cell-free expression system in an altered amount compared to its native state. In one embodiment, the cell is a cell that does not naturally contain the polynucleotide. However, the cell may contain a non-endogenous polynucleotide that results in an altered, preferably increased, production of the encoded polypeptide. Exogenous polynucleotides of the invention include polynucleotides that have not been separated from other components of the transgenic (recombinant) cell or cell-free expression system in which they are present, as well as polynucleotides produced in such cells or cell-free systems that have subsequently been purified from at least some of the other components.

[0162] The percent identity of polynucleotides is determined by GAP (Needleman and Wunsch, 1970) analysis (GCG program) with a gap creation penalty of 5 and a gap extension penalty of 0.3. Unless otherwise specified, the query sequence is at least 45 nucleotides in length, and the GAP analysis aligns the two sequences over a region of at least 45 nucleotides. Preferably, the query sequence is at least 150 nucleotides in length, and the GAP analysis aligns the two sequences over a region of at least 150 nucleotides. More preferably, the query sequence is at least 300 nucleotides in length, and the GAP analysis aligns the two sequences over a region of at least 300 nucleotides. Even more preferably, the GAP analysis aligns the two sequences over their entire length.

[0163] The polynucleotides of the present invention may have one or more mutations that are deletions, insertions, or substitutions of nucleotide residues when compared to the molecules provided herein. Variants can be either naturally occurring (i.e., isolated from a natural source) or synthetic (e.g., by performing site-directed mutagenesis on the nucleic acid).

[0164] Typically, the monomers of a polynucleotide are linked by phosphodiester bonds or their analogs, including phosphorothioates, phosphorodithioates, phosphoroselenoates, phosphorodiselenoates, phosphoroanilothioates, phosphoroanilidates, and phosphoramidates.

[0165] One embodiment of the present invention includes a recombinant vector comprising at least one isolated / exogenous polynucleotide of the present invention inserted into any vector capable of delivering a polynucleotide molecule to a host cell. Such vectors contain heterologous polynucleotide sequences, i.e., polynucleotide sequences not naturally found adjacent to the polynucleotide molecule of the present invention, preferably polynucleotide sequences derived from a species other than that from which the polynucleotide molecule is derived. Vectors can be either RNA or DNA, prokaryotic or eukaryotic, and typically are transposons, viruses, or plasmids.

[0166] One type of recombinant vector comprises a polynucleotide operably linked to an expression vector. The term operably linked refers to inserting a polynucleotide molecule into an expression vector in such a manner that the molecule can be expressed when transformed into a host cell. As used herein, an expression vector is a DNA or RNA vector capable of transforming a host cell and causing expression of a particular polynucleotide molecule. Preferably, the expression vector is also capable of replicating within the host cell. Expression vectors can be either prokaryotic or eukaryotic and are typically viruses or plasmids. Expression vectors include any vector that functions (i.e., directs gene expression) in recombinant cells, including bacterial, fungal, endoparasitic, arthropod, animal, and plant cells. The vectors of the present invention can also be used to produce polypeptides in cell-free expression systems, and such systems are well known in the art.

[0167] As used herein, "operably linked" refers to a functional relationship between two or more nucleic acid (e.g., DNA) segments. Typically, it refers to the functional relationship between a transcriptional control element and a transcribed sequence. For example, a promoter is operably linked to a coding sequence, such as a polynucleotide defined herein, if it stimulates or regulates the transcription of the coding sequence in an appropriate host cell and / or cell-free expression system. Generally, a promoter transcriptional control element operably linked to a transcribed sequence is physically contiguous to the transcribed sequence, i.e., cis-acting. However, some transcriptional control elements, such as enhancers, need not be physically contiguous or located in close proximity to the coding sequence whose transcription they enhance.

[0168] In particular, expression vectors according to the present invention contain control sequences, such as transcriptional control sequences, translational control sequences, origins of replication, and other control sequences compatible with recombinant cells that control expression of polynucleotide molecules of the present invention. In particular, recombinant molecules of the present invention contain transcriptional control sequences. Transcriptional control sequences are sequences that control the initiation, elongation, and termination of transcription. Particularly important transcriptional control sequences are sequences that control transcription initiation, such as promoter, enhancer, operator, and repressor sequences. Suitable transcriptional control sequences include any transcriptional control sequence that can function in at least one of the recombinant cells of the present invention. A variety of such transcriptional control sequences are known to those skilled in the art. Preferred transcription control sequences include those that function in bacteria, yeast, arthropods, nematodes, plant or animal cells, such as, but not limited to, tac, lac, tip, trc, oxy-pro, omp / lpp, rmB, bacteriophage λ, bacteriophage T7, T71ac, bacteriophage T3, bacteriophage SP6, bacteriophage SP01, metallothionein, alpha mating factor, Pichia alcohol oxidase, alphavirus subgenomic promoter (Sindbis virus subgenomic promoter), and the like. motors, etc.), antibiotic resistance genes, baculovirus, Heliothis tea insect virus, vaccinia virus, herpesvirus, raccoon poxvirus, other poxviruses, adenovirus, cytomegalovirus (e.g., intermediate early promoters), simian virus 40, retrovirus, actin, retroviral long terminal repeats, Rous sarcoma virus, heat shock, phosphate and nitrate transcriptional control sequences, and other sequences capable of controlling gene expression in prokaryotic or eukaryotic cells.

[0169] Another embodiment of the present invention includes host cells transformed with one or more recombinant molecules described herein, or their progeny. Transformation of a polynucleotide molecule into a cell can be accomplished by any method by which a polynucleotide molecule can be inserted into a cell. Transformation techniques include, but are not limited to, transfection, electroporation, microinjection, lipofection, adsorption, and protoplast fusion. Recombinant cells may remain unicellular or may grow into tissues, organs, or multicellular organisms. The transformed polynucleotide molecules of the present invention may be maintained extrachromosomally or may be integrated into one or more sites within the chromosome of the transformed (i.e., recombinant) cell such that expression competence is maintained.

[0170] Suitable host cells for transformation include any cell that can be transformed with a polynucleotide of the present invention. Host cells of the present invention can either be endogenous (i.e., naturally occurring) capable of producing a polypeptide described herein or can produce such a polypeptide after being transformed with at least one polynucleotide molecule described herein. Host cells of the present invention can be any cell capable of producing at least one protein defined herein, including bacterial, fungal (including yeast), parasitic, nematode, arthropod, animal, and plant cells. Exemplary host cells include Salmonella, Escherichia, Bacillus, Listeria, Saccharomyces, Spodoptera, Mycobacteria, Trichoplusia, BHK (baby hamster kidney) cells, MDCK cells, CRFK cells, CV-1 cells, COS (e.g., COS-7) cells, and Vero cells. Further examples of host cells include Escherichia coli, including K-12 E. coli derivatives; Salmonella typhi; Salmonella typhimurium, including attenuated strains; Spodoptera frugiperda; Trichoplusia ni; and non-tumorigenic mouse myoblast G8 cells (e.g., ATCC CRL 1246). Useful yeast cells include Pichia sp., Aspergillus sp., and Saccharomyces sp. Particularly preferred host cells are bacterial, fungal, or plant cells.

[0171] In one embodiment, the cells are suitable for fermentation. Examples of bacterial cells useful for fermentation include, but are not limited to, Escherichia sp. (such as Escherichia coli), Bacillus sp. (such as Bacillus subtilis and Bacillus licheniformis), Lactobacillus sp. (such as Lactobacillus brevis), Pseudomonas sp. (such as Pseudomonas aeruginosa), and Streptomyces sp. (such as Streptomyces lividans). Examples of fungal cells useful for fermentation include, but are not limited to, Candida sp. (such as Candida albicans), Hansenula sp. (such as Hansenula polymorpha), Pichia sp. (such as Pichia pastoris), Kluyveromyces sp. (such as Kluyveromyces marxianus), and Saccharomyces sp. (such as Saccharomyces cerevisiae).

[0172] Recombinant DNA technology can be used to improve the expression of transformed polynucleotide molecules, for example, by manipulating the copy number of the polynucleotide molecules in the host cell, the efficiency with which those polynucleotide molecules are transcribed, the efficiency with which the resulting transcripts are translated, and the efficiency of post-translational modifications. Recombinant techniques useful for increasing the expression of the polynucleotide molecules of the invention include, but are not limited to, operatively linking the polynucleotide molecules to high-copy-number plasmids, integrating the polynucleotide molecules into one or more host cell chromosomes, adding vector stability sequences to the plasmid, replacing or modifying transcriptional control signals (e.g., promoters, operators, enhancers), replacing or modifying translational control signals (e.g., ribosome binding sites, Shine-Dalgarno sequences), modifying the polynucleotide molecules of the invention to correspond to the codon usage of the host cell, and deleting sequences that destabilize the transcript.

[0173] The term "plant" as used herein as a noun refers to a whole plant, such as, for example, a commercial plant or a plant grown in a field for crop production. "Plant parts" refers to vegetative structures (e.g., leaves, stems), roots, floral organs / structures, seeds (including embryos, endosperms, and seed coats), plant tissues (e.g., vascular tissue, aboveground tissues, etc.), cells, and progeny thereof.

[0174] A "transgenic plant" refers to a plant that contains a genetic construct (a "transgene") that is not found in a wild-type plant of the same species, variety, or cultivar. As used herein, "transgene" has its ordinary meaning in the art of biotechnology and includes a genetic sequence that has been generated or modified by recombinant DNA or RNA techniques and introduced into a plant cell. A transgene may also include a genetic sequence derived from a plant cell. Typically, a transgene has been introduced into a plant by artificial manipulation, such as by transformation, although one of skill in the art will recognize that any method may be used.

[0175] The polynucleotides of the present invention may be constitutively expressed in transgenic plants during all developmental stages. Depending on the intended use of the plant or plant organ, the polypeptides may be expressed in a stage-specific manner. Furthermore, the polynucleotides may be expressed in a tissue-specific manner.

[0176] The compositions of the present invention contain an excipient, also referred to herein as an "acceptable carrier." The excipient can be any material acceptable to the animal, plant, animal or plant material, or environment (including soil and water samples) being treated. Examples of such excipients include water, saline, Ringer's solution, dextrose solution, Hank's solution, and other physiologically balanced salt solutions. Non-aqueous vehicles such as fixed oils, sesame oil, ethyl oleate, or triglycerides can also be used. Other useful formulations include suspensions containing viscosity-enhancing agents such as sodium carboxymethylcellulose, sorbitol, or dextran. The excipient may also contain minor amounts of additives, such as substances that enhance isotonicity or chemical stability. Examples of buffers include phosphate buffer, bicarbonate buffer, and Tris buffer, and examples of preservatives include thimerosal or o-cresol, formalin, and benzyl alcohol. Excipients may also be used to increase the half-life of the composition, including, but not limited to, polymeric controlled release vehicles, biodegradable implants, liposomes, bacteria, viruses, other cells, oils, esters, and glycols.

[0177] Each document, reference, patent application, or patent cited in this text is expressly incorporated herein by reference in its entirety, meaning that it should be read and considered by the reader as part of this text. Documents, references, patent applications, or patents cited in this text are not repeated in this text solely for the sake of brevity. This statement is not an admission that any of the cited documents is prior art or part of the common general knowledge of those working in the art to which this invention pertains.

[0178] Any embodiment of the present invention may also be broadly said to consist of any or all combinations of two or more of the parts, elements and features, individually or collectively, referred to or shown in this specification, and where specific integers that have known equivalents in the art to which the present invention pertains are referred to herein, such known equivalents are deemed to be incorporated herein as if individually defined.

[0179] It should be understood that references to "one example" or "an example" of the invention are not made in an exclusive sense. Thus, one example illustrates a particular embodiment of the invention, while other embodiments may be illustrated in separate examples. These examples are intended to aid one of ordinary skill in the art in practicing the invention, and are not intended to limit the overall scope of the invention in any way, unless the context clearly indicates otherwise.

[0180] It should be understood that the terminology used above is for the purpose of description and should not be regarded as limiting. The described embodiments are intended to illustrate the present invention without limiting its scope. The present invention can be embodied with various modifications and additions that will be readily apparent to those skilled in the art.

[0181] Other definitions of selected terms used herein can be found in the detailed description of the invention and apply throughout. Unless otherwise defined, all other scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0182] Various substantial and particularly practical and useful exemplary embodiments of the claimed subject matter are described in the text and / or drawings herein, including the best mode, if any, known to the inventors for carrying out the claimed subject matter.

[0183] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications. The invention also includes all steps, features, compositions, and compounds referred to or shown in this specification, individually or collectively, and any and all combinations of those steps or features, or any two or more thereof.

[0184] The inventors anticipate that those skilled in the art will employ such variations as appropriate, and the inventors intend for the claimed subject matter to be practiced otherwise than as specifically described herein. Accordingly, the claimed subject matter includes and encompasses all equivalents of and all modifications to the claimed subject matter, as permitted by law. Moreover, all combinations of the above-described elements, activities, and all possible variations thereof are encompassed by the claimed subject matter unless expressly indicated herein, clearly and specifically disclaimed, or clearly contradicted by context.

[0185] The present invention is not to be limited in scope by the specific embodiments described herein, which are for the purpose of illustration only, and functionally equivalent products, compositions and methods are clearly within the scope of the invention described herein.

[0186] Any examples provided herein, or the use of exemplary language (e.g., "such as" or "for example"), are intended merely to better illuminate one or more embodiments and do not limit the scope of the claimed subject matter, unless otherwise specified. No language in this specification should be construed as indicating any claimed subject matter as essential to the practice of the claimed subject matter.

[0187] Throughout this specification and claims, unless the context requires otherwise, the word "comprise" or variations such as "comprises" or "comprising" are understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers.

[0188] Throughout this specification, unless the context clearly indicates otherwise, the word "include" or variations such as "includes" or "including" will be understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers.

[0189] Furthermore, when a numerical value or range is described herein, unless otherwise specified, the numerical value or range is approximate. The recitation of a range of values ​​herein is intended to serve merely as a shorthand method of individually referring to each distinct value falling within the range, unless otherwise specified herein, and each distinct value and each distinct portion of the range defined by such each distinct value is incorporated herein as if it were individually recitation herein. For example, when a range of 1 to 10 is described, the range includes all values ​​therebetween, such as 1.1, 2.5, 3.335, 5, 6.179, 8.9999, etc., and all subranges therebetween, such as 1 to 3.65, 2.8 to 8.14, 1.93 to 9, etc.

[0190] Accordingly, all portions of this application other than the claims (e.g., title, field, background, summary, description, abstract, drawings, etc.) shall be deemed exemplary in nature and not limiting, and the scope of the subject matter protected by any patent issuing based on this application shall be defined solely by the claims of such patent.

[0191] While additional embodiments of the present application are shown and described herein, it is to be expressly understood that the present application is not limited thereto and is capable of various other embodiments and implementations within the scope of the following claims.

[0192] Any feature of an embodiment of an aspect is applicable to all other aspects and embodiments identified herein. Any feature of an embodiment can be independently, partially, or in whole combined in any way with other embodiments described herein, e.g., one, two, or more embodiments can be combined in whole or in part. [Brief explanation of the drawings]

[0193] [Figure 1] Figure 1 shows an unrooted maximum likelihood phylogenetic tree of bacterial lipase family 1.5 constructed using IQ-tree and visualized in iTOL v5. UF-Boot values ​​are shown for nodes with SH-aLRT support values ​​≥ 80%. Chemical analysis records of macroalgal hydrolysates are shown. [Figure 2] Esterase activity of lipase homologs using pNP substrate. (A) A bar graph of activity assay using 5 μl of cell culture with 0.75 mM pNP acetate to confirm protein expression. Error bars represent SEM from three replicate experiments. (B) A photograph of an SDS-PAGE gel analyzing the purity of the protein obtained from small-scale nickel affinity purification. The expected band sizes are indicated at the bottom of the gel. (C) A bar graph comparing the activity of 10 nM purified protein using pNP substrates of different lengths. Error bars represent SEM from two replicate experiments. [Figure 3]A line graph showing the residual activity of heat-stable proteins for pNP acetate and pNP propionate is shown. Samples were heated at the indicated temperature for 10 minutes, cooled to 4°C, and then the residual activity was measured. Residual activity is provided as a percentage of the activity of the unheated sample. The six enzymes shown were identified from the initial assay using all 10 enzymes presented in this study. Error bars represent the SEM from three replicate experiments. [Figure 4] This graph shows the EqAD activity of 185 μl of supernatant from a reaction containing 100 nM Cl_EstA and 5 mg / mL PBAT, incubated for 48 hours at 40° C. The EqAD reaction contained 250 μM NAD+ and 0.1 U / mL EqAD in a final reaction volume of 200 μL, and the progress of the reaction was measured by tracking the change in absorbance at 340 nm as NADH++H+ was produced. [Figure 5A] Figure 1 shows the thermotolerance of the P4G12 Tl_Est47 mutant measured using protein-expressing cell cultures. (A) Individual graphs are shown for plates screened for activity by heating cultures at 89°C for 10 min, heating overnight, then cooling, and then testing for activity against paranitrophenyl butyrate. [Figure 5B] (B) Thermotolerance of the P4G12 Tl_Est47 variant measured using protein-expressing cell cultures. (C) Line graph showing the residual activity of the improved variant TL_Est47 compared to the wild type, heated from 75°C to 99°C for 10 minutes, cooled, and then tested for activity against paranitrophenyl butyrate. [Figure 6A] Characterization of WT Tl_Est47 and its L221Q and G377D mutants with improved thermostability. (A) Line graph showing the esterase activity of the WT and mutant proteins measured using pNP-butyrate as a substrate. [Figure 6B](B) Characterization of the thermostable WT Tl_Est47 and its L221Q and G377D mutants. (B) Table showing the kinetic parameters of the WT and mutant proteins for the curves shown in (A). [Figure 7] A line graph showing the thermostability of purified WT and mutant Tl_Est47 in LB medium. Reactions to measure residual activity contained 200 nM enzyme and 300 μM pNP-butyrate, and the amount of liberated pNP was measured by absorbance at 405 nm. [Figure 8] A table showing the experimental record of conditions in a 2 liter fermenter is shown. [Figure 9] A table showing the experimental record of conditions in a 500 mL flask is shown. [Figure 10] 1 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with wild-type Tl_Est47 at pH 7 to pH 9. [Figure 11] 1 depicts a graph showing the alcohol dehydrogenase activity of PBSA incubated with Tl_Est47 variants at pH 7 to pH 9 (Tris buffer). [Figure 12] 10 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with wild-type Tl_Est47 from 4° C. to 37° C. [Figure 13] 10 depicts a graph showing alcohol dehydrogenase activity from PBSA incubated with Tl_Est47 variants from 4° C. to 37° C. [Figure 14] Whole cell and soluble SDS PAGE gels are shown after 2 h, 4 h, 5.5 h, and 22 h incubation times after induction in either flasks or fermentors (blue arrow indicates Tl_Est47). [Figure 15A]FPLC traces and SDS-PAGE gels are shown for (A) pellets from 500 ml flask cultures and (B) 2 L fermentations. The absorbance at 280 nM is indicated by the blue line. SDS / PAGE shows molecular size markers (M, sizes indicated to the left of the gel), whole cell (W), and soluble (S) fractions. Fractions 1–9 correspond to fractions from the FPLC traces analyzed by SDS / PAGE. [Figure 15B] FPLC traces and SDS-PAGE gels are shown for (A) pellets from 500 ml flask cultures and (B) 2 L fermentations. The absorbance at 280 nM is indicated by the blue line. SDS / PAGE shows molecular size markers (M, sizes indicated to the left of the gel), whole cell (W), and soluble (S) fractions. Fractions 1–9 correspond to fractions from the FPLC traces analyzed by SDS / PAGE.

Claims

1. i) the amino acid sequence set forth in SEQ ID NO: 4; ii) an amino acid sequence that is at least 90% identical to i); or iii) a biologically active fragment of ii); A polypeptide comprising:

2. i) a glutamine (Q) at a position corresponding to amino acid number 221 of SEQ ID NO:4; and ii) an aspartic acid (D) at a position corresponding to amino acid number 377 of SEQ ID NO:4; The polypeptide of claim 1, comprising one or both of:

3. 3. The polypeptide of claim 1 or claim 2, comprising a fusion protein comprising at least one other polypeptide sequence.

4. i) the nucleotide sequence set forth in SEQ ID NO:2; ii) a nucleotide sequence encoding a polypeptide according to claim 1 or claim 2; or iii) a sequence of nucleotides complementary to either i) or ii); An isolated and / or exogenous polynucleotide comprising a sequence selected from:

5. 5. The isolated and / or exogenous polynucleotide of claim 4, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the polypeptide of claim 1 or claim 2 in a cell.

6. 5. The isolated and / or exogenous polynucleotide of claim 4, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the polypeptide of claim 1 or claim 2 in an expression host cell.

7. A vector comprising the polynucleotide of claim 4.

8. A nucleic acid construct or expression vector comprising the polynucleotide of claim 4.

9. A nucleic acid construct or expression vector comprising the polynucleotide of claim 4, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide of claim 1 or claim 2 in an expression host cell.

10. 3. A recombinant expression host cell comprising a polynucleotide encoding a polypeptide according to claim 1 or claim 2, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide.

11. A host cell comprising the polynucleotide of claim 4.

12. The host cell of claim 11 , comprising a bacterial cell, a fungal cell, or a plant cell.

13. A transgenic non-human organism comprising at least one cell according to claim 11 or claim 12.

14. 13. An extract of a host cell according to claim 11 or claim 12, wherein the extract comprises: i) the amino acid sequence set forth in SEQ ID NO:4; ii) an amino acid sequence that is at least 90% identical to i); or iii) a biologically active fragment of ii); An extract comprising a polypeptide comprising:

15. the polypeptide comprising: i) a glutamine (Q) at a position corresponding to amino acid number 221 of SEQ ID NO:4; and ii) an aspartic acid (D) at a position corresponding to amino acid number 377 of SEQ ID NO:

4.

16. A composition comprising a polypeptide according to claim 1 or claim 2 and one or more acceptable carriers.

17. 16. A composition comprising the extract of claim 14 or claim 15 and one or more acceptable carriers.

18. i) the amino acid sequence set forth in SEQ ID NO:4; ii) an amino acid sequence that is at least 90% identical to i); or iii) a biologically active fragment of ii); A method for producing a polypeptide comprising:

19. the polypeptide comprising: i) a glutamine (Q) at a position corresponding to amino acid number 221 of SEQ ID NO:4; and ii) an aspartic acid (D) at a position corresponding to amino acid number 377 of SEQ ID NO:

4.

20. A method for producing a polypeptide described in claim 1 or claim 2, comprising culturing a recombinant expression host cell containing the polynucleotide described in claim 4, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide described in claim 1 or claim 2 under conditions that promote the production of the polypeptide.

21. 22. The method of claim 21, further comprising recovering the polypeptide of claim 1 or claim 2.

22. An enzyme comprising the polypeptide of claim 1 or claim 2.

23. A thermostable enzyme comprising the polypeptide of claim 1 or 2.

24. A hyperthermophilic enzyme comprising the polypeptide of claim 1 or claim 2.

25. 25. The enzyme of any one of claims 22 to 24, wherein the enzyme maintains enzymatic activity at a temperature higher than the temperature at which an enzyme comprising a wild-type Thermosyntropha lipolytica Tl_Est47 polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 3 loses substantially all enzymatic activity.

26. 25. The enzyme of any one of claims 22 to 24, wherein the enzyme maintains its enzymatic activity at temperatures above 90°C.

27. 25. The enzyme of any one of claims 22 to 24, wherein the enzyme maintains some enzymatic activity at temperatures above 95°C.

28. 28. An enzyme according to any one of claims 22 to 27, comprising esterase activity.

29. 28. An enzyme according to any one of claims 22 to 27, comprising lipase activity.

30. 3. The polypeptide of claim 1 or claim 2, comprising a polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47, wherein the polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47 comprises the amino acid sequence set forth in SEQ ID NO:

3.

31. 5. The polynucleotide of claim 4, capable of expressing a polypeptide variant of wild-type Thermosyntropha lipolytica Tl_Est47, wherein the wild-type Thermosyntropha lipolytica Tl_Est47 comprises an amino acid sequence expressed by an isolated and / or exogenous polynucleotide comprising a sequence selected from the sequence of nucleotides set forth in SEQ ID NO:1.