Novel heat-resistant enzyme

By modifying the amino acid sequence of Thermotrophic Bacillus lipolytica Tl_Est47, a polypeptide variant Tl_Est47 with higher thermal stability at high temperatures was developed, which solved the problem of unstable enzyme activity, achieved the maintenance of enzyme activity at high temperatures, and expanded its application range in industrial processes.

CN120603941APending Publication Date: 2025-09-05ENZIDE TECHNOLOGIES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088346.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-28
Filing Date
2023-10-25
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The enzyme activity of the existing Thermotrophic Bacillus lipolytica Tl_Est47 is not stable enough at high temperatures, which limits its application in high-temperature industrial processes.

Method used

By modifying the amino acid sequence of Thermotrophic Bacillus lipolytica Tl_Est47, especially introducing glutamine (Q) at amino acid position 221 and aspartic acid (D) at amino acid position 377, a polypeptide variant Tl_Est47 with higher thermal stability was developed.

Benefits of technology

The improved polypeptide variant Tl_Est47 maintains enzyme activity at high temperatures and can maintain enzyme activity at 90°C or higher, significantly increasing its application potential in high-temperature industrial processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005461196480000141
    Figure BDA0005461196480000141
  • Figure BDA0005461196480000151
    Figure BDA0005461196480000151
  • Figure HDA0005461196490000011
    Figure HDA0005461196490000011
Patent Text Reader

Abstract

The provided novel enzyme variant of the naturally existing wild type thermophilic syntrophy lipolytica TlEst47 has improved thermal stability, and has no negative effect on the stability or activity of the enzyme variant at a certain temperature or pH value range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to novel variants of Thermotrophicus lipolytica Tl_Est47 having improved thermostability compared to the naturally occurring wild-type Thermotrophicus lipolytica Tl_Est47. Background Art

[0002] The following discussion of the background art is intended only to facilitate an understanding of the present invention and it should be understood that the discussion does not constitute an acknowledgement or admission that any of the material referred to was part of the common general knowledge as at the priority date of the application.

[0003] Enzymes have been used for many years in the production of detergents, products for the textile and starch industries, and foods including cheese, bread, beer, wine, leather and linen. These industrial processes utilize enzymes produced and isolated by specific microorganisms or present in natural products such as papaya fruit.

[0004] Recent advances in protein engineering have enabled the development of mutant enzymes for established applications or new, customized enzymes for applications where enzymes have not previously been used. Over half of the enzymes used in industrial processes are derived from fungi, over a third from bacteria, and the remainder from animal and plant sources. Recombinant DNA technology enables the isolation and cloning of enzyme-encoding genes from all possible sources and enables high-yield heterologous expression. As a result, bacterial strains initially unsuitable for industrial applications, including Aspergillus, Saccharomyces, and Bacillus, have shown improved production levels and enzymes.

[0005] Currently, the global industrial enzyme market is estimated to be around $10 billion, with some of the more commonly used enzymes including lipase, polyphenol oxidase, lignin peroxidase, horseradish peroxidase, amylase, nitrite reductase, and urease.

[0006] Enzymes for use in these industrial processes are selected based on the reactions they initiate and catalyze, as well as various properties that enable them to operate in the industrial process environment, for example, properties that enable the enzyme to maintain enzymatic activity over the various required temperature and pH environments.

[0007] Enzymes are highly thermostable, which is a great benefit for their use in many industrial processes that involve high temperatures at certain stages, or for products that are used under high temperature conditions.

[0008] Therefore, there is an increasing demand for enzymes with thermostable properties at high temperatures for use in products and specific industrial applications and processes. Summary of the Invention

[0009] The inventors have engineered novel polypeptides that are variants of T. lipolytica Tl_Est47 that have improved thermostability in maintaining enzyme activity at high temperatures compared to naturally occurring (wild-type) T. lipolytica Tl_Est47.

[0010] In a first aspect, the present invention provides a novel polypeptide comprising:

[0011] i) the amino acid sequence provided in SEQ ID NO: 4,

[0012] ii) an amino acid sequence that is at least 90% identical to i), or

[0013] iii) biologically active fragments of ii).

[0014] In one embodiment, the novel polypeptide of the present invention comprises one or both of the following:

[0015] i) glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and

[0016] ii) Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO: 4.

[0017] In one embodiment, the novel polypeptides of the present invention comprise a fusion protein comprising at least one other polypeptide sequence.

[0018] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the present invention comprising a sequence selected from the group consisting of:

[0019] a. the nucleotide sequence provided in SEQ ID NO: 2;

[0020] b. a nucleotide sequence encoding a novel polypeptide of the present invention; or

[0021] c. A nucleotide sequence complementary to i) or ii).

[0022] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the present invention as described herein, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the novel polypeptide of the present invention in a cell.

[0023] In one embodiment, the present invention provides an isolated and / or exogenous polynucleotide of the present invention as described herein, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the novel polypeptide of the present invention in an expression host cell.

[0024] In one embodiment, the present invention provides a vector comprising a polynucleotide of the present invention as described herein.

[0025] In one embodiment, the present invention provides a nucleic acid construct or expression vector comprising a polynucleotide of the present invention as described herein.

[0026] In one embodiment, the present invention provides a nucleic acid construct or expression vector comprising a polynucleotide of the present invention as described herein, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the novel polypeptide of the present invention in an expression host cell.

[0027] In one embodiment, the present invention provides a recombinant expression host cell comprising a polynucleotide encoding a novel polypeptide of the present invention, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide.

[0028] In one embodiment, the invention provides a host cell comprising a polynucleotide of the invention as described herein.

[0029] In one embodiment, the host cell preferably includes a bacterial cell, a fungal cell or a plant cell.

[0030] In one embodiment, the present invention provides a transgenic non-human organism comprising at least one host cell as described herein.

[0031] In one embodiment, the present invention provides an extract of a host cell as described herein, wherein the extract comprises a polypeptide comprising:

[0032] a. the amino acid sequence provided in SEQ ID NO: 4;

[0033] b) an amino acid sequence at least 90% identical to i), or

[0034] c. biologically active fragments of ii).

[0035] In one embodiment, the present invention provides an extract of the host cell described herein, wherein the polypeptide comprises one or both of the following:

[0036] a. glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and

[0037] b. Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO: 4.

[0038] In one embodiment, the present invention provides a composition comprising the novel polypeptide of the present invention and one or more acceptable carriers.

[0039] In one embodiment, the present invention provides a composition comprising an extract as described herein and one or more acceptable carriers.

[0040] In one embodiment, the present invention provides a method for producing a novel polypeptide, comprising:

[0041] a. the amino acid sequence provided in SEQ ID NO: 4;

[0042] b) an amino acid sequence at least 90% identical to i), or

[0043] c. biologically active fragments of ii).

[0044] In one embodiment, the novel polypeptide preferably comprises one or both of the following:

[0045] a. glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and

[0046] b. Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO: 4.

[0047] In one embodiment, the present invention provides a method for producing a novel polypeptide of the present invention, comprising culturing a recombinant expression host cell comprising the novel polypeptide of the present invention, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the novel polypeptide of the present invention under conditions conducive to the production of the polypeptide.

[0048] In one embodiment, the method preferably comprises recovering the novel polypeptide of the present invention.

[0049] In one embodiment, the present invention provides an enzyme comprising the novel polypeptide of the present invention.

[0050] In one embodiment, the present invention provides a thermostable enzyme comprising the novel polypeptide of the present invention.

[0051] In one embodiment, the present invention provides a hyperthermophilic enzyme comprising the novel polypeptide of the present invention.

[0052] In one embodiment, the invention provides an enzyme as described herein, wherein the enzyme retains enzymatic activity at a temperature above the temperature at which the enzyme comprising the wild-type Thermotrophic bacterium lipolytica Tl_Est47 polypeptide loses substantially all enzymatic activity, the enzyme comprising the amino acid sequence provided in SEQ ID NO:3.

[0053] In one embodiment, the enzymes described herein retain enzymatic activity at 90°C or higher.

[0054] In one embodiment, the enzymes described herein retain some enzymatic activity at 95°C or higher.

[0055] In one embodiment, the enzymes described herein have esterase activity.

[0056] In one embodiment, the present invention provides an enzyme as described herein having lipase activity.

[0057] In one embodiment, the present invention provides a novel polypeptide of the present invention comprising a polypeptide variant of wild-type Thermotrophic Bacillus lipolytica Tl_Est47, wherein the polypeptide variant of wild-type Thermotrophic Bacillus lipolytica Tl-Est47 comprises the amino acid sequence provided in SEQ ID NO:3.

[0058] In one embodiment, the present invention provides a polynucleotide of the present invention as described herein, which is capable of expressing a polypeptide variant of wild-type Thermotrophic bacterium lipolytica Tl_Est47, wherein the wild-type Thermotrophic bacterium lipolytica Tl-Est47 comprises an amino acid sequence expressed by an isolated and / or exogenous polynucleotide, the polynucleotide comprising a sequence selected from the nucleotide sequence provided in SEQ ID NO: 1. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The present invention will now be described by way of example with reference to the accompanying drawings, in which:

[0060] Figure 1 A schematic diagram showing an unrooted maximum likelihood phylogenetic tree of the bacterial lipase family 1.5, constructed using IQ-tree and visualized using iTOL v5. UF-Boot values ​​for nodes with SH-aLRT support ≥80% are shown. A chemical analysis report of the macroalgae hydrolysate is also included.

[0061] Figure 2 Esterase activity of lipase homologs against pNP substrates. (A) Bar graph comparing activity assays using 5 μl of cell culture medium with 0.75 mM pNP acetate to verify protein expression. Error bars represent the standard error mean of three replicates. (B) SDS-PAGE gel electrophoresis analysis of protein purity obtained by small-scale nickel affinity purification. Expected band sizes are indicated at the bottom of the gel. (C) Bar graph comparing the activity of 10 nM purified protein against pNP substrates of varying lengths. Error bars represent the standard error mean of two replicates.

[0062] Figure 3Line graph showing the residual activity of thermostable proteins towards pNP acetate and pNP propionate. Samples were heated at the indicated temperatures for 10 minutes and then cooled at 4°C. Residual activity is expressed as a percentage of the activity of the unheated sample. The six enzymes shown were identified from preliminary assays of all ten enzymes included in this study. Error bars represent the standard error of the mean from three replicates.

[0063] Figure 4 . Shown is the EqAD activity in 185 μl of supernatant after a 48-hour incubation at 40°C containing 100 nM Cl_EstA and 5 mg / mL PBAT. The EqAD reaction contained 250 μM NAD in a final reaction volume of 200 μl. + and 0.1U / mL EqAD, by NADH at 340nm + +H + The resulting absorbance change measures the reaction progress.

[0064] Figure 5. Thermostability of the P4G12 Tl_Est47 mutant measured using protein-expressing cell cultures. (A) Single image of a plate screened for activity by heating overnight cultures at 89°C for 10 minutes, cooling, and then testing for activity against p-nitrophenol butyrate. (B) Line graph showing the residual activity of the improved variant TL_Est47 compared to the wild-type after heating at temperatures between 75°C and 99°C for 10 minutes, allowing to cool, and then testing for activity against p-nitrophenol butyrate.

[0065] Figure 6. Characterization of WT Tl_Est47 and its L221Q and G377D mutants, demonstrating increased thermostability. (A) Line graph showing esterase activity of WT and mutant proteins measured using pNP butyrate as a substrate. (B) Table showing kinetic parameters for WT and mutant proteins for the curves shown in (A).

[0066] Figure 7 Line graph showing the thermal stability of purified wild-type and mutant Tl_Est47 in LB medium. The reaction for measuring residual activity contained 200 nM enzyme and 300 μM pNP butyrate, and the amount of released pNP was measured at an absorbance of 405 nM.

[0067] Figure 8 .A table showing the conditions recorded in a 2-liter fermenter.

[0068] Figure 9 .Show the condition record table in the 500mL flask.

[0069] Figure 10 . Graph showing the alcohol dehydrogenase activity of PBSA incubated with wild-type Tl_Est47 between pH 7 and pH 9 (Tris buffer).

[0070] Figure 11 . Graph showing the alcohol dehydrogenase activity of PBSA incubated with Tl_Est47 variants between pH 7 and pH 9 (Tris buffer).

[0071] Figure 12 Graph showing alcohol dehydrogenase activity of PBSA incubated with wild-type Tl_Est47 at 4 to 37°C.

[0072] Figure 13 Graph showing alcohol dehydrogenase activity of PBSA incubated with Tl_Est47 variants at 4 to 37°C.

[0073] Figure 14 SDS-PAGE gel electrophoresis of whole-cell and soluble fractions after 2 h, 4 h, 5.5 h, and 22 h incubation times after induction in flasks or fermentors (blue arrows indicate Tl_Est47).

[0074] Figure 15. FPLC chromatograms (FPLC traces) and SDS-PAGE gel electrophoresis of pellets from (A) 500 mL flask and (B) 2 L fermentation cultures. The blue line shows absorbance at 280 nM. The SDS-PAGE shows molecular size markers (M, sizes given on the left side of the gel), whole cell (W), and soluble (S) fractions. 1-9 correspond to the fractions of the FPLC chromatograms analyzed by SDS-PAGE. DETAILED DESCRIPTION

[0075] DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0076] In order to more accurately understand the subject matter of the present invention, the features of the invention will now be discussed with reference to the following preferred embodiments.

[0077] Sequence retrieval and sequence similarity network (SSN) generation.

[0078] In an initial search for potential enzymes with improved thermostability, sequences of Cl_EstA (accession number: WP_011948553.1), Cl_EstB (accession number: WP_011986581.1), and PfL1 (accession number: EIW29778.1) were used to perform a pBLAST search against the NCBI Refseq_Select_proteins database using default parameters. All sequences with an e-value greater than 0.005 were retrieved and submitted to the Enzyme Function Initiative Enzyme Similarity Tool (EFI-EST) to generate a sequence similarity network (SSN) containing only sequences between 200 and 1000 amino acids in length, with an initial alignment score threshold of 7. In the SSN, nodes represent individual proteins, and edges represent alignment scores, calculated by the EFI-EST algorithm using the bit scores obtained from an all-vs-all BLAST of the provided protein sequences. Alignment scores are close in magnitude to the negative logarithm of the BLAST E-value, which approximates protein similarity.

[0079] The SSNs were visualized using Cytoscape 3.9.0 by applying the yfile Organic layout. Clusters were analyzed by gradually increasing the alignment score threshold and then reapplying the layout until no significant changes were observed when the threshold jumped significantly (from 20 to 40 alignment scores in this SSN). This visual cluster analysis method can identify potential functional groups from other related proteins. Only sequences belonging to large clusters containing the query protein were then used for further SSN and phylogenetic analysis. Further adjustments to the alignment score threshold and subsequent reapplication of the yfile Organic layout were performed to visualize branches within larger protein families.

[0080] Phylogenetic analysis to identify sequences from thermophiles and extremophiles

[0081] The refined sequence set in SSN was then aligned using the structure-based protein sequence alignment algorithm PROMALS3D, using the structures of Cl_EstA (PDB ID: 5AH1) and PfL1 (PDB ID: 5AH0) as templates. Regions with poor alignment at the N-terminus and C-terminus, as well as long insertions in otherwise well-aligned sequences, were deleted, and sequences with poor overall alignment or large deletions were completely removed. Subsequently, the trimmed sequence set was realigned using MUSCLE on the EMBL-EBI web server (https: / / www.ebi.ac.uk / Tools / msa / muscle / ) with default settings. The alignment results were further fine-tuned to remove regions with poor N-terminal alignment and any large gaps or insertions, and then an initial approximate maximum likelihood (ML) tree was generated using FastTree2.1 and the WAG+CAT evolutionary model. The multiple sequence alignment (MSA) was further trimmed to remove any long branches and highly similar sequences with very short branch lengths, and the process was repeated until the SH-like local support values ​​for most nodes were greater than 80%. The resulting alignment was used to infer a maximum likelihood tree using the IQ-TREE web service (https: / / www.hiv.lanl.gov / content / sequence / IQTREE / iqtree.html). The best evolutionary model calculated was as follows: LG model + FreeRate heterogeneity (# rate classes = 8) + amino acid frequencies optimized from the data using maximum likelihood. Ultra-rapid bootstrap (UF-Boot) values ​​and SH-aLRT branch test values ​​were calculated for 1000 replicates. The tree was visualized and annotated using the interactive tree of life (iTOL) web tool (https: / / itol.embl.de / ).

[0082] Sequences of bacterial lipase family 1.5 screened using SSN were used to generate a maximum likelihood phylogenetic tree ( Figure 1This analysis supports the observations from SSN, where Cl_EstA, Cl_EstB, and PfL1 belong to a large clade containing proteins from bacteria belonging to the Clostridiaceae family. Within this clade, two subclades contain Cl_EstA and Cl_EstB, respectively, suggesting that they arose from a gene duplication event in the last common ancestor of the Clostridiaceae family. As with SSN, the phylogenetic tree supports horizontal gene transfer from the Cl_EstA clade to the last common ancestor of the genus Pelosin, which gave rise to PfL1. Sequences from the Clostridiaceae family observed on SSN were most closely related to Bacilli from the Thermoactinomycetes, Paenibacillusaceae, and Cycloacidiaceae families. Among these Clostridia and Bacilli, sequences from thermophilic and acidophilic Clostridia from the Peptococcusaceae and Syntrophomonasaceae families were also present. Distinct clades were observed for the genera Bacillus and Geobacillus, as well as a third large clade containing Bacilli from the orders Bacilales (e.g., Bacillus sp. and Caldibacillus sp.) and Lactobacilliales. With the exception of Firmicutes, sequences from all other phyla formed a clearly separate branch, with the exception of three sequences from Betaproteobacteria and Bacteroidetes, which were likely acquired by horizontal gene transfer.

[0083] Based on the phylogenetic tree and SSN, further functional analysis was limited to sequences from the genera Clostridium and Bacillus of the order Bacillales, as they had the closest homology to Cl_EstA, Cl_EstB, and PfL1. Sequences were selected from six thermophiles and one acidophile, namely: Ct_Est from Caldibacillus thermoaylovorans (accession number: WP_152032401.1), Gk_Est from Geobacillus kaustophilus (accession number: WP_044733155.1), Gs_Est from Geobacillus stearothermophilus (accession number: WP_095860225.1), and Thermosyntropha lipolytica. lipolytica (accession number: WP_073088947.1), Tl_Est64 from Thermosyntrophalipolytica (accession number: WP_014826614.1), Dt_Est from Desulfurispora thermophila (accession number: WP_018085325.1), and Da_Est from Desulfosporosinus acidophilus (accession number: WP_014826614.1).

[0084] The polyesterase homolog was expressed in Escherichia coli and showed activity towards the pNP substrate.

[0085] Protein expression

[0086] The sequences of the seven identified target proteins, as well as those of Cl_EstA, Cl_EstB, and PfL1, were entered into the SignalP-5.0 web server, and the predicted signal sequences were trimmed before designing expression vectors ordered from GenScript (Singapore). Each truncated sequence was placed between the Nde1 and Xho1 restriction sites in pET-29b(+) so that the expressed protein contained an N-terminal His tag. The plasmids were transformed into NEB T7 expression cells (New England Biolabs) according to the manufacturer's recommended protocol and plated on Luria Broth (LB) agar plates containing 50 μg / mL kanamycin. The plates were incubated at 37°C overnight and stored at 4°C for up to 2 weeks. A negative control plasmid (gfasPurple-S125R-F162R-V44A-L123T_pETcc2) was also transformed and plated on LB agar plates containing 100 μg / mL ampicillin.

[0087] For protein expression, a single colony of each strain was inoculated into a 50 ml tube containing 10 ml of autoinduction medium (5 g yeast extract, 20 g tryptone, 85.5 mM NaCl, 22 mM KH2PO4, 42 mM Na2HPO4, 0.6% glycerol, 0.05% glucose and 0.2% lactose) containing 50 μg / ml kanamycin (100 μg / ml ampicillin for the control strain). The culture was grown at 37°C for 3-6 hours while shaking at 200 rpm and then incubated at 30°C overnight. The resulting culture was stored at 4°C for up to 2 weeks for whole cell analysis or spun at 4°C for 10 minutes at 4000 x g, the supernatant discarded, and the resulting pellet stored at -20°C until protein purification.

[0088] Selected proteins were expressed in E. coli NEB T7 expression strain using an N-terminal His tag. Protein profiles of whole cell fractions (soluble and insoluble proteins) and soluble fractions isolated from cell cultures were evaluated by SDS-PAGE gel electrophoresis.

[0089] SDS-PAGE detection of protein expression

[0090] 500 μL of each culture was centrifuged at 4000 x g for 10 minutes at 4°C in a 1.5 mL microcentrifuge tube, and the supernatant was discarded. The resulting cell pellet was resuspended in 100 μL of lysis solution (50 mM Tris pH 8, 1x BugBuster protein extraction reagent (Millipore) and approximately 33 nL DNase I) and placed on ice for approximately 10 minutes. After lysis, 5 μL of each sample was mixed with 10 μL of 50 mM Tris H 8 and 5 μL of 4x NuPAGE. TM The remaining samples were centrifuged at 20,000 x g for 10 min at 4°C, and 15 μL of the supernatant was mixed with 5 μL of 4x NuPAGE TM Mix with LDS sample buffer (Invitrogen). The samples were heated at 90°C for 3 minutes and then loaded onto a pre-loaded NuPAGE TM The gels were loaded onto 4-12% Bis-Tris gels (Invitrogen) and run in MES SDS running buffer (Invitrogen) for 30-40 minutes at 150 V. The gels were stained with AcquaStain protein gel stain (Bulldog) for 30 minutes and then destained in water.

[0091] Since a strong background band of the expected size was also observed in the negative control, clear bands were only observed in Cl_EstA, Cl_EstB, Dt_Est, and Tl_Est64. Therefore, to confirm protein expression, an initial esterase activity assay was performed using p-nitrophenyl (pNP) acetate ( Figure 2 A) The activities of all 10 proteins were detected, confirming their successful expression in E. coli.

[0092] To compare the relative activities of the proteins, the proteins were partially purified by small-scale nickel affinity chromatography ( Figure 2 B).

[0093] Protein purification

[0094] For small-scale purification, pellets of 10 mL of culture were resuspended in 1 ml of lysis buffer containing 50 mM Tris, 300 mM NaCl, pH 8, and transferred to 2 ml microcentrifuge tubes. Cells were lysed by sonication (Fisher Scientific, 5 s pulse, 1 s interval, 30 s duration, repeated 3 times). The lysate was then spun at 20,000 x g and 4°C for 20 min, and the supernatant was loaded onto a microcentrifuge tube. Ni spin column (New England biolabs) was pre-washed with 250 μL of the same lysis buffer. The column was washed with a total of 750 μL of wash buffer (50 mM Tris, 300 mM NaCl, 5 mM imidazole, pH 8) and the sample was eluted with 2 x 200 μL of elution buffer (50 mM Tris, 300 mM NaCl, 500 mM imidazole, pH 8). For each eluate, 15 μL of the eluate was mixed with 5 μL of 4x NuPAGE. TM The purified protein was mixed with LDS sample buffer (Invitrogen) and used for SDS-PAGE analysis as described above. The purified protein was stored at 4°C for up to 3 weeks.

[0095] The purity of the bands for Cl_EstA, PfL1, Da_Est, Dt_Est, Tl_Est47, and Tl_Est64 was approximately 90% or higher. Although the purity of Cl_EstB and Ct_Est was only approximately 40-50%, clear bands of the expected size were observed. In contrast, the protein levels of Gs_Est and Gk_Est purified using this protocol were very low and barely detectable, although pNP acetate activity was observed in whole cell samples.

[0096] Esterase activity assay was performed using pNP substrate.

[0097] Purified proteins were then used to characterize the substrate preferences of these proteins and compared with the activities of pNP acetate, pNP propionate, pNP butyrate, pNP valerate (pNP-velarate), and pNP octanoate ( Figure 2 C).

[0098] To check the activity of the protein with the pNP substrate, first mix 90 μL of 50mM Tris (pH 8.0) with 5 μL of cell culture for whole-cell assays, or mix with 5 μL of 200nM purified protein. To start the reaction, add 5 μL of a 100% methanol solution containing 15mM pNP acetate, pNP propionate, pNP butyrate, pNP valerate or pNP-octanoate so that the final reaction contains 5% methanol. The absorbance at 405nm is measured at intervals of 3-4s over 10 minutes, and the amount of pNP produced is calculated using the rate of the linear portion of the curve (usually within the first 30-100s), with an extinction coefficient of 16853M. -1 cm -1 , the path length is 0.25cm.

[0099] The general trend was that the activity of all proteins increased with increasing substrate length, with the exception of Cl_EstA, which showed the opposite trend with the best activity towards pNP acetate. Dt_Est showed the highest activity with the longest substrate tested. Tl_Est64 showed the highest activity of all proteins with pNP butyrate and pNP valerate. As expected, the poorly purified Cl_EstB, Ct_Est, Gk_Est, and Gk_Est had low activity towards all substrates. Despite successful purification, Da_Est from the acidophilic organism showed only slight activity towards pNP octanoate and no activity towards any other substrates ( Figure 2 B) It also showed only slight activity against pNP acetate when tested in cell culture.

[0100] Cl_EstA, PfL1, Dt_Est, Tl_Est47 and Tl_Est64 have high thermal stability.

[0101] Next, the thermostability of these enzymes was tested by measuring the residual activity of cell cultures heated at different temperatures for 10 minutes.

[0102] For each enzyme (and negative control culture), 8 x 15 μL of cell culture was added to a 0.2 ml test tube and heated to different temperatures for 10 minutes using the temperature gradient function on the thermal cycler (instrument). The test tube was then cooled on ice for another 10 minutes. Residual activity was measured using the above-mentioned pNP substrate assay method using pNP acetate, pNP propionate, or pNP butyrate. 5 μL of heated and cooled cell culture samples were used, and a positive control containing 5 μL of unheated cell culture and a negative control containing 5 μL of buffer were set up for each protein. The reaction rate in mOD / min was obtained and used to calculate the percentage of residual activity compared to the unheated positive control sample.

[0103] Preliminary testing showed that Cl_EstA, PfL1, Dt_Est, Gk_Est, Tl_Est47, and Tl_Est64 had significant residual activity after heating above 60°C using pNP acetate. Cells expressing these proteins were further evaluated over a range of temperatures to estimate their melting temperatures (T m )( Figure 3 The highest temperatures at which at least 10% residual activity was detected with pNP acetate were as follows: residual activity of Dt_Est, Pfl1, and Tl_Est64 >10% after heating at 78°C, residual activity of Cl_EstA >50% after heating at 70°C, and residual activity of Gk_Est >20% when heated at 65.75°C.

[0104] Interestingly, with pNP acetate, Tl_Est47 retained >10% residual activity even after heating at 85°C, a finding that was further assessed by repeating the assay with pNP propionate. While Tl_Est47 showed little detectable activity with pNP acetate under these reaction conditions, the reaction rate increased approximately 4-fold with pNP propionate as a substrate ( Figure 3 For Tl_Est47, we observed >10% residual activity after heating at 90°C with pNP propionate, making it the most thermostable enzyme evaluated in these experiments.

[0105] Improvements to Tl_Est47B to increase thermal stability.

[0106] An investigation was then conducted to determine whether the thermostability of Tl_Est47 could be further improved. Tl_Est47, which retains 10% of its enzyme activity after heating at 90°C for 10 minutes, was found to retain 10% of its enzyme activity. To this end, a random mutation library was generated using error-prone PCR with an estimated library size of 13,000 variants, each containing an average of 2-12 mutations. The mutants were cultured overnight in LB medium in 96-well growth plates and heated at 89, 90, and 91°C for 10 minutes. Residual activity was measured using pNP butyrate. Figure 4 ). The 10 most active mutants were then further screened, and it was observed that mutant P4G12 still had residual activity even after heating at 99°C for 10 minutes, which was an improvement of about 5-10°C compared to the WT Tl_Est47 protein ( Figure 5 ).

[0107] Tl_Est47 (WT) and the P4G12 mutant were purified on a large scale to compare the two proteins. Their activities with pNP butyrate were first compared ( Figure 6A and 6B ), the results showed that the loss rate of the mutant (k cat ) was reduced by 3-fold. However, the K M The 2-fold decrease indicates that the substrate affinity of the mutant is improved, so the overall catalytic efficiency (k cat / K M ) was only 1.5 times higher than that of the mutant.

[0108] Often, greater thermostability is achieved through protein-stabilizing mutations, which can reduce the overall motion of the protein in solution, decreasing protein activity by affecting the rates of substrate binding and diffusion and the rate of the overall catalytic mechanism.

[0109] The thermal stability of purified WT and mutant proteins was also examined ( Figure 7Consistent with the observations in the whole-cell assay, the thermostability of the mutants was found to be approximately 5-7°C higher when comparing the residual activity after heating in LB medium. The overall thermostability of the WT protein in LB medium was increased by approximately 5°C compared to buffer.

[0110] Performance of variant Tl_Est47 at different pH values ​​and temperatures

[0111] The performance of enzymes was investigated over a range of pH values ​​and temperatures. When transferred to industrial-scale production systems (i.e., fermenters), the yield of some enzymes can also be significantly reduced, thereby increasing the cost of goods (CoG) of enzyme-based products. Therefore, enzyme production via fermentation was also investigated and compared to production in shake flasks to determine enzyme yields under these conditions.

[0112] method

[0113] Protein expression

[0114] For protein expression, a single colony of each strain was inoculated into a 50 ml tube containing 10 ml of autoinduction medium (5 g yeast extract, 20 g trypsin, 85.5 mM NaCl, 22 mM KH2PO4, 42 mM Na2HPO4, 0.6% glycerol, 0.05% glucose and 0.2% lactose) containing 50 μg / ml kanamycin (100 μg / ml ampicillin for the control strain). The culture was grown at 37°C for 3-6 hours while shaking at 200 rpm and then incubated at 30°C overnight. The resulting culture was stored at 4°C for up to 2 weeks for whole cell analysis or centrifuged at 4000 x g for 10 minutes at 4°C, the supernatant discarded, and the resulting pellet stored at -20°C until protein purification.

[0115] SDS-PAGE detection of protein expression

[0116] 500 μL of each culture was centrifuged at 4000 x g for 10 minutes at 4°C in a 1.5 mL microcentrifuge tube, and the supernatant was discarded. The resulting cell pellet was resuspended in 100 μL of lysis solution (50 mM Tris pH 8, 1x BugBuster protein extraction reagent (Millipore) and approximately 33 nL DNase I) and placed on ice for approximately 10 minutes. After lysis, 5 μL of each sample was mixed with 10 μL of 50 mM Tris H 8 and 5 μL of 4x NuPAGE. TM The remaining samples were centrifuged at 20,000 x g for 10 min at 4°C, and 15 μL of the supernatant was mixed with 5 μL of 4xNuPAGE TMMix with LDS sample buffer (Invitrogen). The samples were heated at 90°C for 3 minutes and then loaded onto a pre-loaded NuPAGE TM The gels were loaded onto 4-12% Bis-Tris gels (Invitrogen) and run in MESSDS running buffer (Invitrogen) for 30-40 minutes at 150 V. The gels were stained with AcquaStain protein gel stain (Bulldog) for 30 minutes and then destained in water.

[0117] Protein purification

[0118] For small-scale purification, pellets of 10 mL of culture were resuspended in 1 mL of lysis buffer containing 50 mM Tris, 300 mM NaCl, pH 8, and transferred to 2 mL microcentrifuge tubes. Cells were lysed by sonication (Fisher Scientific, 5-second pulses, 1-second intervals, 30-second durations, repeated 3 times). The lysate was then spun at 20,000 x g and 4°C for 20 minutes, and the supernatant was loaded onto a microcentrifuge tube. Ni spin column (New England biolabs) was pre-washed with 250 μL of the same lysis buffer. The column was washed with a total of 750 μL of wash buffer (50 mM Tris, 300 mM NaCl, 5 mM imidazole, pH 8) and the sample was eluted with 2 x 200 μL of elution buffer (50 mM Tris, 300 mM NaCl, 500 mM imidazole, pH 8). For each eluate, 15 μL of the eluate was mixed with 5 μL of 4x NuPAGE. TM The purified protein was mixed with LDS sample buffer (Invitrogen) and used for SDS-PAGE analysis as described above. The purified protein was stored at 4°C for up to 3 weeks.

[0119] PBSA degradation test

[0120] First, prepare a 10 mg / mL suspension of PBSA in buffer. Then, aliquot 250 μL of this suspension into 1.5 mL or 2 mL tubes, one for each enzyme to be tested. For whole-cell assays, add an additional 250 μL of buffer to dilute PBSA to 5 mg / mL and add 5 or 10 μL of cell suspension to start the reaction. For assays of purified enzymes, add 250 μL of the same buffer containing 200 nM of each enzyme solution to the suspension, resulting in a final concentration of 5 mg / mL PBSA and 100 nM enzyme. Also include a negative control containing buffer alone.

[0121] The reactions were shaken at room temperature at 37°C for 7 days. The temperature determination solution was prepared using a buffer containing 50 mM Tris pH 8.0. Reactions were incubated at 4, 15, 25, and 37°C. The pH of each starting material was adjusted to the desired pH using NaOH as needed. The buffer was diluted (1 / 10) in distilled water for use in the reaction.

[0122] At the end of the incubation period, the samples were centrifuged at 16,000 x g for 1 minute at room temperature to pellet the remaining plastic, and 2 x 200 μL aliquots of the supernatant were transferred from each tube to a 96-well UV plate (Grenier). 10 μL of 50 mM NAD+ and 5 μL of 4 U / mL equine alcohol dehydrogenase were added to each aliquot, and the absorbance change at 340 nm was read at 30-second intervals for 30 minutes. The values ​​of the technical replicates were averaged, and the standard error of the mean (SEM) was calculated for 2 or 3 experimental replicates.

[0123] Comparison of fermenter and flask growth

[0124] To compare growth in batch and fermentor cultures, cultures were grown in either TB or 2YT medium. 2YT medium (1 L) consisted of 5 g yeast extract, 16 g tryptone, and 5 g sodium chloride dissolved in 600 mL of distilled water, then made up to 1 L and sterilized at 121°C for 20 min. TB medium (2.5 L) consisted of 12.5 g yeast extract (5 g / L), 50 g tryptone (20 g / L), 12.5 g NaCl (5 g / L), 7.5 g KH2PO4 (3 g / L), 14.9 g Na2HPO4 (5.96 g / L), and 0.6% glycerol (12.5 mL) dissolved in 2.5 L of distilled water. 500 mL was dispensed into sterile, baffled 2-L Erlenmeyer flasks with vented caps, or 2 L was dispensed into a pre-assembled Sartorius Bioreactor B fermenter and sterilized at 121°C for 60 min. After sterilization, kanamycin was added at a final concentration of 50 mg / mL.

[0125] The 500 mL flasks and 2 L fermentors were inoculated with 10 mL of overnight culture into the same medium inoculated with a single colony from an agar plate streaked with E. coli BL21DE3 transformed with the appropriate expression plasmid.

[0126] Fermenter control parameters were: pO2 setpoint = 30%, cascade: agitator / airflow / O2 enrichment, initial setpoints: agitator = 500 rpm, airflow = 0.3 L / min, O2 = 0.07 L / min, temperature setpoint = 37°C, pH setpoint = 7.0, acid / base configuration, 10% H3PO4 / 10% NH3.

[0127] Samples for OD measurement and 1 mL samples for enzyme analysis were collected throughout the process. At harvest, 20 mL of cell pellet and total harvest pellet (weighing 36 g) were retained and frozen at -80°C along with the 1 mL samples. Conditions in the flasks and fermenters were recorded as follows: Figure 8 and Figure 9 shown.

[0128] Performance of variant Tl_Est47 at different pH values ​​and temperatures

[0129] Tl_Est47 and its variants were purified from E. coli BL21 DE3 cultures and used in this study. The purified proteins were incubated with PBSA for seven days, and alcohol dehydrogenase activity was assessed daily.

[0130] The wild type and variants responded to pH as expected ( Figure 10 and 11 Due to the pKa of the active site serine nucleophile, the pH optimum for serine hydrolases is often above pH 8. At lower pH values, activity decreases.

[0131] The variants and wild-type enzymes had similar activities at temperatures between 4 and 37°C, indicating that the amino acid substitutions in the variants did not affect the enzyme's activity at low temperatures and that the enzymes remained active at temperatures that likely reflect those encountered in their actual applications ( Figure 12 and 13 As expected, the activity of both enzymes increased with increasing temperature, approximately doubling when incubated at 37°C compared to 4°C.

[0132] Fermentation production of Tl_Est47

[0133] The E. coli BL21 DE3 expressing the Tl_Est47 variant was grown in 500 mL flask cultures and 2 L fermentors to assess the yield of Tl_Est47 under these two conditions. Enzyme productivity in the fermentor (2 L scale). During the expression process of 2, 4, 5.5 and 22 hours, the two culture systems were compared with 1 mL of culture sampled and quickly frozen in liquid nitrogen at -80 ° C. The total protein content in each 1 mL sample was assessed (Table 1). After 5.5 hours, the total protein content was stable. After 22 hours of expression, the culture was harvested, and 4.9 g of protein was harvested from the flask incubation and 22 g of protein was harvested from the fermentor.

[0134]

[0135] Table 1. Estimation of the total amount of protein produced during cultivation in 500 mL flasks and 2 L fermentors.

[0136] After thawing on ice, samples were sonicated for 30 seconds and analyzed by SDS-PAGE for total cell and soluble proteins ( Figure 14 ).

[0137] After 22 hours of flask culture or fermentation, the precipitate was harvested. 6.5 g of precipitate was obtained from a 500 mL flask and 36 g of precipitate was obtained from a 2 L fermentor. We used 6.5 g and 6 g of flask and fermentor precipitates to purify the protein by FPLC using a His-trap column on an AKTApure FPLC. The fractions were analyzed by SDS-PAGE. The results and related figures are shown in Figure 15A / B.

[0138] Fractions 3-9 from each system were combined and concentrated on an Amicon column with a 10 kDa molecular weight cutoff, followed by buffer exchange to remove residual imidazole. Protein concentration was estimated by Nanodrop (extinction coefficient at 280 nM: 124470). Enzyme purified from the flask yielded 14 mL of 0.9 mg / mL purified enzyme from 6.5 g of processed pellet. Enzyme produced in the fermentor yielded 16 mL of 2.6 mg / mL purified enzyme from 6 g of processed pellet. Table 2 summarizes the enzyme amounts produced by our two systems.

[0139]

[0140] Table 2. Purification summary.

[0141] Protein yields were good, with increased production and quantity of the Tl_Est47 variant in the fermentor compared to the same protein produced in shake flasks. However, in both cases, the percentage of Tl_Est47 variant produced was lower than that observed for some high-yield heterologously expressed proteins, suggesting that the expression level and yield of the Tl_Est47 variant could be improved if desired. This would ultimately reduce the cost of goods for enzyme production and the final product. An additional 10 L fermentation was performed (estimated yield of approximately 1.2 g of the Tl_Est47 variant), and the cell pellet was stored at -70°C.

[0142] The engineering efforts to generate Tl_Est47 variants with increased thermostability did not negatively impact protein stability or activity within a range of temperatures or pH values. Tl_Est47 retained considerable activity at low temperatures, suggesting that it will remain active under conditions likely to be encountered in use. Heterologous production of Tl_Est47 variants in fermenters was significantly improved compared to shake flasks. While expression levels were good, protein yields could be further increased to reduce production costs.

[0143] Unless explicitly defined otherwise, all technical and scientific terms used herein should be interpreted as having the meaning commonly understood by one of ordinary skill in the art (eg, in cell culture, molecular genetics, immunology, immunohistochemistry, protein chemistry, and biochemistry).

[0144] Unless otherwise indicated, nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxyl orientation.

[0145] Unless otherwise indicated, the recombinant protein, cell culture, and immunological techniques utilized in the present invention are standard procedures well known to those skilled in the art. These techniques are described and explained throughout the literature in the following sources: J. Perbal, A Practical Guide to Molecular Cloning, John Wiley and Sons (1984), J. Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbour Laboratory Press (1989), T. A. Brown (editor), Essential Molecular Biology: A Practical Approach, Volumes 1 and 2, IRL Press (1991), D. M. Glover and B. D. Humes (editors), DNA Cloning: A Practical Approach, Volumes 1-4, IRL Press (1995 and 1996), and F. M. Ausubel et al. (editors), Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience (1988, including all updates to date), Ed Harlow and David Lane (editors), Antibodies: A Laboratory Manual, Cold Spring Harbour Laboratory Press (1989). Harbour Laboratory, (1988), and JE Coligan et al. (editors) Current Protocols in Immunology, John Wiley & Sons (including all updates to date).

[0146] The term "and / or", e.g., "X and / or Y" should be understood as "X and Y" or "X or Y", and should be considered as explicit support for either or both meanings. Any definitions provided herein should be interpreted in the context of the entire specification. As used herein, unless the context clearly indicates otherwise, the singular "a", "an", "said", and "the" include the plural. For example, unless the context clearly indicates otherwise, reference to a "protein" includes a plurality of proteins. As used herein, the term "protein" includes proteins, polypeptides, and peptides. In some embodiments, the terms "protein", "polypeptide", and "peptide" can be used interchangeably.

[0147] As used herein, unless otherwise indicated to the contrary, the term "about" refers to + / - 20%, more preferably + / - 10%, even more preferably + / - 5% of the specified value. Every numerical range used herein includes every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were expressly recited herein.

[0148] For purposes of this specification and the appended claims, when used in conjunction with one or more numbers or numerical ranges, the term "about" or "substantially" should be understood to refer to all such numbers, including all numbers within a range, and to modify the range by extending the boundaries above and below the numerical values. The recitation of numerical ranges by endpoints includes all numbers contained in that range, such as integers, including fractions thereof (e.g., the recitation of 1 to 5 includes 1, 2, 3, 4, and 5, and fractions thereof, such as 1.5, 2.25, 3.75, 4.1, etc.), as well as any range within that range.

[0149] The terms "polypeptide" and "protein" are generally used interchangeably to refer to a single polypeptide chain (a polymeric sequence of amino acid residues) that may or may not be modified by the addition of non-amino acid groups. It should be understood that such polypeptide chains may be bound to other polypeptides or proteins or other molecules such as cofactors. The terms "protein" and "polypeptide" as used herein also include variants, mutants, biologically active fragments and / or modifications of the polypeptides described herein. Single-letter and three-letter codes for amino acids as defined by the IUPAC-IUB Joint Committee on Biochemical Nomenclature (JCBN) are used in this disclosure. The single letter X represents any one of the twenty amino acids. It will also be understood that due to the degeneracy of the genetic code, a polypeptide may be encoded by a plurality of nucleotide sequences.

[0150] The percent identity of a polypeptide can be determined by GAP (Needleman and Wunsch, 1970) analysis (GCG program) with a gap creation penalty of 5 and a gap extension penalty of 0.3. The query sequence is at least 250 amino acids in length, and the GAP analysis compares the two sequences over a region of at least 250 amino acids. More preferably, the query sequence is at least 300 amino acids in length, and the GAP analysis compares the two sequences over a region of at least 300 amino acids. Even more preferably, the query sequence is at least 350 amino acids in length, and the GAP analysis compares the two sequences over a region of at least 350 amino acids. Even more preferably, the GAP analysis compares the two sequences over their entire length.

[0151] As used herein, the phrase "position corresponding to an amino acid number" refers to the relative position of the amino acid to the surrounding amino acids with reference to a particular amino acid sequence. For example, in some embodiments, a polypeptide of the invention may have additional N-terminal amino acids to aid in intracellular localization or extracellular secretion, which would alter the relative positioning of the amino acids when aligned with, for example, SEQ ID NO: 3 or SEQ ID NO: 4.

[0152] The term "mature" form of a protein, polypeptide or peptide refers to the functional form of the protein, polypeptide or peptide without the signal peptide sequence and the propeptide sequence.

[0153] As used herein with respect to amino acid residue positions, "corresponding to" or "corresponds" or "corresponds" refers to the amino acid residue at the recited position in a protein or peptide, or an amino acid residue that is similar, homologous, or identical to the recited residue in a protein or peptide. As used herein, a "corresponding region" generally refers to an analogous position in a related protein or reference protein.

[0154] The term "wild-type" with respect to an amino acid sequence or a nucleic acid sequence indicates that the sequence is a native or naturally occurring sequence. As used herein, the term "naturally occurring" refers to anything found in nature (e.g., a protein or polynucleotide sequence). In contrast, the term "non-naturally occurring" refers to anything not found in nature (e.g., recombinant polynucleotide and protein sequences produced in a laboratory or modifications of a wild-type sequence).

[0155] As used herein, the term "improved thermostability" or "enhanced thermostability" or "increased thermostability" refers to a novel polypeptide that exhibits increased retention of enzymatic activity after incubation at elevated temperatures for a period of time, particularly relative to a wild-type analogous enzyme or polypeptide. Furthermore, for the purposes of the specification and claims, the terms "improved thermostability" and "thermostability" are used interchangeably herein.

[0156] With respect to polypeptide amino acid sequences, the term "variant" refers to a polypeptide amino acid sequence that differs from a specific wild-type, parent, or reference polypeptide amino acid sequence, including one or more artificially substituted, inserted, or deleted amino acids. Similarly, with respect to polynucleotide nucleic acid sequences, the term "variant" refers to a polynucleotide nucleic acid sequence that differs from a specific wild-type, parent, or reference polynucleotide, including one or more artificially substituted, inserted, or deleted nucleic acids. The identity of the wild-type, parent, or reference polypeptide amino acid sequence or polynucleotide nucleic acid sequence will be apparent from the context.

[0157] As used herein, the term "mutation" or "engineering" refers to an artificial change to a reference amino acid or nucleic acid sequence. The term is intended to encompass artificial substitutions, insertions, and deletions.

[0158] As used herein, the term "vector" refers to a nucleic acid construct for introducing or transferring a nucleic acid into a target cell or tissue. Vectors are typically used to introduce exogenous DNA into cells or tissues. Vectors include plasmids, cloning vectors, phages, viruses (e.g., viral vectors), cosmids, expression vectors, shuttle vectors, etc. Vectors typically include an origin of replication, a multiple cloning site, and a selective marker. The process of inserting a vector into a target cell is commonly referred to as transformation.

[0159] As used herein, the term "introduction" in the context of introducing a nucleic acid sequence into a cell refers to any method suitable for transferring a nucleic acid sequence into a cell. Such methods of introduction include, but are not limited to, protoplast fusion, transfection, transformation, electroporation, conjugation, and transduction. Transformation refers to the genetic alteration of a cell resulting from the uptake, optional genomic incorporation, and expression of genetic material (e.g., DNA).

[0160] "Expression cassette" or "expression vector" refers to a nucleic acid construct or vector produced recombinantly or synthetically for expressing a target nucleic acid (e.g., an exogenous nucleic acid or a transgene) in a target cell. The target nucleic acid typically expresses a target protein. The expression vector or expression cassette typically comprises a promoter nucleotide sequence that drives or promotes the expression of the exogenous nucleic acid. The expression vector or expression cassette typically also includes other specific nucleic acid elements that allow transcription of the specific nucleic acid in the target cell. The recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plasmid DNA, virus, or nucleic acid fragment. Some expression vectors have the ability to incorporate and express heterologous DNA fragments in a host cell or host cell genome. Many prokaryotic and eukaryotic expression vectors are commercially available. It is within the knowledge of those skilled in the art to select a suitable expression vector to express a protein from the nucleic acid sequence incorporated into the expression vector.

[0161] As used herein, when a nucleic acid is in a functional relationship with another nucleic acid sequence, it is "operably connected" to another nucleic acid sequence. For example, if a promoter affects the transcription of a coding sequence, a promoter or enhancer is operably connected to a nucleotide coding sequence. If a ribosome binding site is positioned to promote the translation of a coding sequence, it can be operably connected to a coding sequence. Typically, a DNA sequence that is "operably connected" is continuous. However, an enhancer does not have to be continuous. Engagement is achieved by connecting at a convenient restriction site. If there is no such site, a synthetic oligonucleotide adapter or linker can be used according to conventional practice.

[0162] As used herein, the term "gene" refers to a polynucleotide (e.g., DNA fragment) encoding a polypeptide, including regions before and after the coding region. In some cases, a gene includes intervening sequences (introns) between individual coding segments (exons).

[0163] As used herein, when applied to cells, "recombinant" generally means that the cell has been modified by the introduction of an exogenous nucleic acid sequence, or that the cell is derived from a cell so modified. For example, a recombinant cell may contain genes that do not exist in the same form in the natural (non-recombinant) form of the cell, or a recombinant cell may contain native genes that have been modified and reintroduced into the cell (existing in the natural form of the cell). A recombinant cell may contain nucleic acids endogenous to the cell that have been modified without removing the nucleic acid from the cell; such modifications include modifications obtained by gene replacement, site-specific mutations, and related techniques known to those of ordinary skill in the art. Recombinant DNA technology includes techniques for producing recombinant DNA in vitro and transferring the recombinant DNA into cells for expression or propagation, thereby producing recombinant polypeptides. "Recombination" and "recombining" of polynucleotides or nucleic acids generally refer to the assembly or combination of two or more nucleic acids or polynucleotide chains or fragments to produce new polynucleotides or nucleic acids.

[0164] If a nucleic acid or polynucleotide can be transcribed and / or translated to produce a polypeptide or fragment thereof in its native state or when manipulated by methods known to those skilled in the art, it is said to "encode" a polypeptide. The antisense strand of such a nucleic acid is also said to encode the sequence.

[0165] The terms "host strain" and "host cell" refer to a suitable host for an expression vector containing a DNA sequence of interest.

[0166] A "precursor" form of a protein or peptide refers to the mature form of the protein, the sequence of which is operably linked to the amino or carbonyl terminus of the protein. A precursor may also have a "signal" sequence operably linked to the amino terminus of the precursor sequence. A precursor may also have additional polypeptides that participate in post-translational activity (e.g., polypeptides that are cleaved from it, leaving the mature form of the protein or peptide).

[0167] The terms "derived from" and "obtained from" refer not only to proteins produced or producible by the strain of the organism in question, but also to proteins encoded by a DNA sequence isolated from such a strain and produced in a host organism containing such DNA sequence. In addition, the terms refer to proteins encoded by DNA sequences of synthetic and / or cDNA origin and having the identifying characteristics of the protein in question.

[0168] The term "identical" in the context of two polynucleotide or polypeptide sequences means that the nucleic acids or amino acids in the two sequences are the same when aligned for maximum correspondence, as measured using a sequence comparison or analysis algorithm described below and known in the art.

[0169] "% identity" or "percent identity" or "PID" refers to protein sequence identity. Percent identity can be determined using standard techniques known in the art. By aligning sequences to directly compare sequence information, for example, using programs such as BLAST, MUSCLE, or CLUSTAL, the percentage of amino acid identity shared by the target sequences can be determined. For example, the BLAST algorithm is described in Altschul et al., J Mol Biol, 215:403-410 (1990) and Karlin et al., Proc Natl Acad Sci USA, 90:5873-5787 (1993). The percentage (%) amino acid sequence identity value is determined by the number of identical residues that match divided by the total number of residues in the "reference" sequence, including any gaps created by the program for optimal / maximal alignment. The BLAST algorithm refers to the "reference" sequence as the "query" sequence.

[0170] The CLUSTALW algorithm is another example of a sequence alignment algorithm (see Thompson et al., Nucleic Acids Res, 22:4673-4680, 1994). The default parameters of the CLUSTALW algorithm include: gap open penalty = 10.0; gap extension penalty = 0.05; protein weight matrix = BLOSUM series; DNA weight matrix = IUB; delayed divergent sequence % = 40; gap separation distance = 8; DNA conversion weight = 0.50; list hydrophilic residues = GPSNDQEKR; use negative matrix = off; enable residue specific penalty = on; enable hydrophilicity penalty = on; terminal gap separation penalty = off. In the CLUSTAL algorithm, deletions occurring at both ends are included in the calculation. For example, a variant with 5 amino acids deleted at either end of a 500 amino acid polypeptide (or within the polypeptide) has a sequence identity percentage relative to the "reference" polypeptide of 99% (495 / 500 identical residues x 100). Such variants are to be included within the scope of variants having "at least 99% sequence identity" to the polypeptide.

[0171] Understanding the homology between molecules can reveal information about their evolutionary history and their functions; if a newly sequenced protein is homologous to an already characterized protein, this can strongly suggest the biochemical function of the new protein. The most fundamental relationship between two entities is homology; two molecules are said to be homologous if they originate from a common ancestor. Homologous molecules, or homologs, can be divided into two categories: paralogs and orthologs. Paralogs are homologs that exist within a species. The specific biochemical functions of paralogs are typically different. Orthologs are homologs that exist in different species and have very similar or identical functions. A protein superfamily is the largest group (branch) of proteins for which a common ancestor can be inferred. This common ancestor is typically based on sequence alignments and mechanistic similarities. A superfamily typically contains several protein families that display sequence similarity within the family. The term "protein family" is often applied to protease superfamilies based on the MEROPS protease classification system.

[0172] When a nucleic acid or polynucleotide is at least partially or completely separated from other components (including but not limited to other proteins, nucleic acids, cells, etc.), it is "isolated". Similarly, when a polypeptide, protein or peptide is at least partially or completely separated from other components (including but not limited to other proteins, nucleic acids, cells, etc.), it is "isolated". On a molar basis, the isolated species is more abundant than the species in the other compositions. For example, the isolated species can contain at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or about 100% (on a molar basis) of all macromolecular species present. Preferably, the target species is purified to substantially uniformity (i.e., contaminant species cannot be detected in the composition by conventional detection methods). Purity and homogeneity can be determined using a variety of techniques well known in the art, such as agarose or polyacrylamide gel electrophoresis of nucleic acid or protein samples, respectively, followed by visualization after staining. If desired, high resolution techniques such as high performance liquid chromatography (HPLC) or similar methods can be used to purify the material.

[0173] The term "purified" as applied to nucleic acids or polypeptides generally refers to a nucleic acid or polypeptide that is substantially free of other components as determined by analytical techniques well known in the art (e.g., a purified polypeptide or polynucleotide forms a discrete band in an electrophoretic gel, a chromatographic eluate, and / or a medium subjected to density gradient centrifugation). For example, a nucleic acid or polypeptide that produces essentially a single band in an electrophoretic gel is "purified." A purified nucleic acid or polypeptide is at least about 50% pure, typically at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more (e.g., by weight on a molar basis). In a related sense, a molecule in a composition is enriched when its concentration is significantly increased following application of a purification or enrichment technique. The term "enriched" refers to the presence of a compound, polypeptide, cell, nucleic acid, amino acid or other specific substance or component in a composition at a relative or absolute concentration higher than that of the starting composition.

[0174] One or more novel polypeptide variants described herein may undergo various changes, such as one or more amino acid insertions, deletions, and / or substitutions, whether conservative or non-conservative, including changes that do not substantially change the enzymatic activity of the variant. Similarly, the nucleic acids of the present invention may also undergo various changes, such as replacing one or more nucleotides in one or more codons so that specific codons encode the same or different amino acids, resulting in silent variations (e.g., when the encoded amino acid is not changed by a nucleotide mutation) or non-silent variations; one or more deletions of one or more nucleic acids (or codons) in the sequence; one or more nucleic acids (or codons) added or inserted into the sequence; and / or cutting of one or more nucleic acids (or codons) in the sequence, or one or more truncations. Compared to the polypeptide enzyme encoded by the original nucleic acid sequence, many such changes in the nucleic acid sequence may not substantially change the enzymatic activity of the resulting encoded polypeptide enzyme. The nucleic acid sequences described herein may also be modified to include one or more codons that provide optimal expression in an expression system (e.g., a bacterial expression system), while if desired, the one or more codons still encode the same amino acid.

[0175] One or more nucleotide sequences as herein described can be produced by using any suitable synthesis, operation and / or separation technology or its combination.For example, one or more polynucleotides as herein described can be produced using standard nucleic acid synthesis technology, such as solid phase synthesis technology well known to those skilled in the art.In these technologies, usually synthesize up to 50 or more nucleotide bases of fragments, then connect (for example, by enzyme or chemical connection method) to form basically any required continuous nucleic acid sequence.The synthesis of one or more polynucleotides as herein described can also be promoted by any suitable method known in the art, including but not limited to chemical synthesis (see, for example, Beaucage et al. Tetrahedron Letters 22:1859-69 (1981)) using the classical phosphoramidite method, or the method described in Matthes et al. EMBO J.3:801-805 (1984), such as the method commonly used in the automatic synthesis method.One or more polynucleotides as herein described can also be produced by using an automatic DNA synthesizer. Custom nucleic acids can be ordered from various commercial sources (for example, Midland Certified Reagent Company, Great American Gene Company, Operon Technologies Co., Ltd. and DNA 2.0). Other techniques and related principles for synthesizing nucleic acids are described, for example, in Itakura et al., Ann. Rev. Biochem. 53:323 (1984) and Itakura et al., Science 198:1056 (1984).

[0176] Another embodiment relates to one or more vectors comprising one or more novel polypeptide variants described herein (e.g., polynucleotides encoding one or more novel polypeptide variants described herein); expression vectors or expression cassettes comprising one or more nucleic acid or polynucleotide sequences described herein; isolated, substantially pure or recombinant DNA constructs comprising one or more nucleic acid or polynucleotide sequences described herein; isolated or recombinant cells comprising one or more polynucleotide sequences described herein; and compositions comprising one or more such vectors, nucleic acids, expression vectors, expression cassettes, DNA constructs, cells, cell cultures, or any combination or mixture thereof.

[0177] Some embodiments relate to one or more recombinant cells comprising one or more vectors (e.g., expression vectors or DNA constructs) described herein comprising one or more nucleic acid or polynucleotide sequences as described herein. Some such recombinant cells are transformed or transfected with at least one such vector, although other methods known in the art may also be used. These cells are generally referred to as host cells. Some such cells include bacterial cells. Other embodiments relate to recombinant cells (e.g., recombinant host cells) comprising one or more novel polypeptides described herein.

[0178] In some embodiments, one or more vectors described herein are expression vectors or expression cassettes comprising one or more polynucleotide sequences described herein operably linked to one or more additional nucleic acid fragments required for efficient gene expression (e.g., a promoter operably linked to one or more polynucleotide sequences described herein). The vector may include a transcription terminator and / or a selection gene (e.g., an antibiotic resistance gene) that enables sustained culture of plasmid-infected host cells by growth in a culture medium containing an antimicrobial agent. The expression vector may be derived from plasmid or viral DNA, or, in alternative embodiments, comprise elements of both.

[0179] In order to express and produce a target protein (e.g., one or more novel polypeptides described herein) in a cell, one or more expression vectors are transformed into the cell under conditions suitable for expressing the novel polypeptide, the expression vector comprising one or more copies, and in some cases multiple copies, of a polynucleotide encoding one or more novel polypeptides described herein. In some embodiments, the polynucleotide sequence encoding one or more novel polypeptides described herein (as well as other sequences contained in the vector) is integrated into the genome of the host cell, while in other embodiments, a plasmid vector comprising a polynucleotide sequence encoding one or more novel polypeptides described herein is retained in the cell as an autonomous extrachromosomal element. Some embodiments provide extrachromosomal nucleic acid elements and input nucleotide sequences integrated into the host cell genome. The vectors described herein can be used to produce the novel polypeptides described herein. In some embodiments, the polynucleotide construct encoding one or more novel polypeptides described herein is present on an integration vector that is capable of integrating and optionally amplifying the nucleotide sequence encoding the novel polypeptide into the host chromosome. Examples of integration sites are well known to those skilled in the art. In some embodiments, transcription of the polynucleotide encoding the novel polypeptides described herein is achieved by a promoter that is a wild-type promoter of a wild-type polypeptide. In some other embodiments, the promoter is heterologous to the one or more novel polypeptides described herein, but is functional in the host cell.

[0180] In addition to commonly used methods, in some embodiments, host cells are directly transformed with a DNA construct or vector comprising a nucleic acid encoding one or more novel polypeptides described herein (i.e., before introduction into the host cell, the intermediate cell is not used to amplify or otherwise process the DNA construct or vector). Introducing the DNA construct or vector described herein into the host cell includes physical and chemical methods known in the art for introducing a nucleic acid sequence (e.g., a DNA sequence) into the host cell without inserting it into the host genome. These methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, and liposomes. In another embodiment, the DNA construct or vector is co-transformed with a plasmid without inserting it into the plasmid. In a further embodiment, the selective marker is deleted from the altered bacterial strain by methods known in the art (see Stahl et al., J. Bacteriol. 158: 411-418 (1984); Palmeros et al., Gene 247: 255-264 (2000)). In some embodiments, the transformed cells are cultured in a conventional nutrient medium. Suitable specific culture conditions, such as temperature, pH, etc., are known to those skilled in the art and are well described in the scientific literature.

[0181] As used herein, the term "extract" refers to any portion of a host cell or non-human transgenic organism of the present invention that contains a polypeptide of the present invention, and preferably also contains a polynucleotide or vector of the present invention. The term includes portions secreted from the host cell, and thus includes culture supernatants. Preferably, the extract is a relatively crude extract that has not undergone purification steps to purify the polypeptide of the present invention from other polypeptides co-produced with the polypeptide of the present invention. The extract may also be a composition comprising the polypeptide of the present invention.

[0182] As used herein, a "biologically active fragment" is a portion of a polypeptide described herein that retains a defined activity of the full-length polypeptide. Biologically active fragments can be of any size, as long as they retain a defined activity.

[0183] Thus, where applicable, based on the minimum % identity figures, preferred polypeptides comprise an amino acid sequence that is at least 90%, more preferably at least 91%, more preferably at least 92%, more preferably at least 93%, more preferably at least 94%, more preferably at least 95%, more preferably at least 96%, more preferably at least 97%, more preferably at least 98%, more preferably at least 99%, more preferably at least 99.1%, more preferably at least 99.2%, more preferably at least 99.3%, more preferably at least 99.4%, more preferably at least 99.5%, more preferably at least 99.6%, more preferably at least 99.7%, more preferably at least 99.8%, even more preferably at least 99.9% identical to the relevant designated SEQ ID NO. With respect to defined polypeptides, it will be understood that percentage identity figures higher than those provided herein will encompass preferred embodiments.

[0184] "Substantially purified" or "purified" refers to a polypeptide that is separated from one or more lipids, nucleic acids, other polypeptides, or other contaminating molecules with which it is naturally associated. Preferably, a substantially purified polypeptide is at least 60%, more preferably at least 75%, and even more preferably at least 90% free from other components with which it is naturally associated. Although there is no evidence that the polypeptides of the present invention exist in nature, the terms "native state" and "naturally associated" also include polypeptides produced in the host cells of the present invention.

[0185] In the context of polypeptides, the term "recombinant" refers to a polypeptide that is produced by a cell or cell-free expression system in an altered amount or at an altered rate compared to the native state. In one embodiment, the cell is one that does not naturally produce the polypeptide. Recombinant polypeptides of the invention include polypeptides that have not been separated from other components of the transgenic (recombinant) cells or cell-free expression systems in which they were produced, as well as polypeptides produced in these cells or cell-free systems that have subsequently been purified to remove at least some other components.

[0186] Amino acid sequence mutants or variants of the polypeptides described herein can be prepared by introducing appropriate nucleotide changes into the nucleic acids defined herein or by in vitro synthesis of the desired polypeptide. For example, such mutants include deletions, insertions, or substitutions of residues in the amino acid sequence. A combination of deletions, insertions, and substitutions can be made to obtain the final construct, provided that the final polypeptide product possesses the desired characteristics.

[0187] Mutant or variant polypeptides can be prepared using any technique known in the art, for example using directed evolution or rational design strategies (see below). Products derived from mutated / altered DNA can be readily screened using the techniques described herein to determine whether they have enzymatic activity.

[0188] In designing amino acid sequence variants, the location of the mutation site and the nature of the mutation will depend on the feature to be modified. Mutation sites can be modified individually or in tandem, for example, by (1) first replacing with a conservative amino acid choice and then, depending on the results obtained, replacing with a more radical choice, (2) deleting the target residue, or (3) inserting additional residues adjacent to the targeted site.

[0189] Amino acid sequence deletions typically range from about 1 to 15 residues, more preferably about 1 to 10 residues, and typically about 1 to 5 contiguous residues.

[0190] Substitution mutants have at least one amino acid residue removed from a polypeptide molecule and a different residue inserted in its place. Target sites are sites where a specific residue is identical across strains or species. These sites may be important for biological activity. These sites, particularly those located within a sequence of at least three other identical conserved sites, are preferably substituted in a relatively conservative manner.

[0191] In a preferred embodiment, the mutant / variant polypeptide has only conservative substitutions compared to the novel polypeptides specifically defined herein. In a preferred embodiment, the mutant / variant polypeptide has one or two or three or four conservative amino acid changes compared to the novel polypeptides specifically defined herein.

[0192] If not otherwise specified, preferably, at a given amino acid position, the novel polypeptide comprises the amino acid found at the corresponding position in the polypeptide set forth in SEQ ID NO:4.

[0193] The scope of the present invention also includes novel polypeptides of the present invention that are differentially modified during or after synthesis, for example, by biotinylation, benzylation, glycosylation, acetylation, phosphorylation, amidation, derivatization with known protecting / blocking groups, proteolytic cleavage, linkage to antibody molecules or other cellular ligands, etc. These modifications can be used to increase the stability and / or biological activity of the polypeptides.

[0194] Novel polypeptides as described herein can be produced in a variety of ways, including the production and recovery of recombinant polypeptides, and chemical synthesis of polypeptides. In one embodiment, isolated polypeptides of the present invention are produced by cultivating cells capable of expressing the polypeptide under conditions that effectively produce the polypeptide and recovering the polypeptide. Preferred cultured cells are recombinant cells of the present invention. Effective culture conditions include, but are not limited to, effective culture medium, bioreactor, temperature, pH, and oxygen conditions that allow polypeptide production. Effective culture medium refers to any culture medium in which cells are cultured to produce polypeptides of the present invention. This culture medium typically includes an aqueous culture medium with assimilable carbon, nitrogen, and phosphate sources and suitable salts, minerals, metals, and other nutrients (such as vitamins). Cells of the present invention can be cultured in conventional fermentation bioreactors, shake flasks, test tubes, microtiter dishes, and culture dishes. Cultivation can be carried out under temperature, pH, and oxygen content that are suitable for recombinant cells. This culture condition is within the professional knowledge of those of ordinary skill in the art.

[0195] In one embodiment, the novel polypeptides of the present invention comprise a signal sequence capable of directing secretion of the polypeptide from a cell. Those skilled in the art will appreciate that a signal sequence may or may not be cleaved, or may be partially cleaved while being partially exported from the cell. However, when a signal sequence is removed, the cell may produce a heterogeneous population of polypeptides with slightly different N-terminal sequences, for example. Therefore, the term "comprising" encompasses such variants produced by removing a signal sequence. Numerous such signal sequences have been isolated, including both N-terminal and C-terminal signal sequences. Prokaryotic and eukaryotic N-terminal signal sequences are similar, and eukaryotic N-terminal signal sequences have been shown to function as secretion sequences in bacteria. An example of such an N-terminal signal sequence is the bacterial β-lactamase signal sequence, a well-studied sequence that has been widely used to promote secretion of polypeptides into the external environment. An example of a C-terminal signal sequence is the Escherichia coli hemolysin A (hlyA) signal sequence. Other examples of signal sequences include, but are not limited to, aerolysin, alkaline phosphatase gene (phoA), chitinase, endochitinase, α-hemolysin, MIpB, pullulanase, Yops, and TAT signal peptides.

[0196] As used herein, an "isolated polynucleotide" refers to a polynucleotide that is at least partially separated from the polynucleotide sequence to which it is naturally associated or linked. Isolated polynucleotides include DNA and RNA molecules, as well as combinations of DNA and RNA molecules. They can be single-stranded, double-stranded, or partially double-stranded, and can be in the sense or antisense orientation relative to a promoter. Preferably, an isolated polynucleotide is at least 60%, preferably at least 75%, and most preferably at least 90% free from other components to which it is naturally associated. Furthermore, the term "polynucleotide" is used interchangeably herein with the term "nucleic acid."

[0197] In the context of polynucleotides, the term "exogenous" refers to a polynucleotide that is present in a cell or cell-free expression system in an altered amount compared to the native state. In one embodiment, the cell is a cell that does not naturally contain the polynucleotide. However, the cell may be a cell that contains a non-endogenous polynucleotide, resulting in an altered, preferably increased, amount of the encoded polypeptide produced. The exogenous polynucleotides of the present invention include polynucleotides that have not been separated from other components of the transgenic (recombinant) cell or the cell-free expression system in which they are present, and the polynucleotides produced in these cells or cell-free systems will subsequently be purified from at least some of the other components.

[0198] The percent identity of polynucleotides is determined by GAP (Needleman and Wunsch, 1970) analysis (GCG program), with a gap creation penalty=5 and a gap extension penalty=0.3. Unless otherwise indicated, the query sequence is at least 45 nucleotides in length, and the two sequences are compared in a region of at least 45 nucleotides by GAP analysis. Preferably, the query sequence is at least 150 nucleotides in length, and the two sequences are compared in a region of at least 150 nucleotides by GAP analysis. More preferably, the query sequence is at least 300 nucleotides in length, and the two sequences are compared in a region of at least 300 nucleotides by GAP analysis. Even more preferably, the two sequences are compared over their entire length by GAP analysis.

[0199] The polynucleotides of the present invention may have one or more mutations, i.e., deletions, insertions, or substitutions of nucleotide residues, compared to the molecules provided herein. Mutants may be naturally occurring (i.e., isolated from natural sources) or synthetic (e.g., by site-directed mutagenesis of the nucleic acid).

[0200] Typically, the monomers of a polynucleotide are linked by phosphodiester bonds or their analogs. Analogs of phosphodiester bonds include phosphorothioates, phosphorodithioates, phosphoroselenoates, phosphorodiselenoates, phosphoroanilothioates, phosphoranilidates, and phosphoramidates.

[0201] One embodiment of the present invention includes a recombinant vector comprising at least one isolated / exogenous polynucleotide of the present invention inserted into any vector capable of delivering the polynucleotide molecule into a host cell. Such vectors comprise heterologous polynucleotide sequences, i.e., naturally occurring polynucleotide sequences adjacent to the polynucleotide molecule of the present invention, and preferably derived from species other than the species from which the polynucleotide molecule was derived. The vector may be RNA or DNA, whether prokaryotic or eukaryotic, and is typically a transposon, virus, or plasmid.

[0202] A type of recombinant vector comprises a polynucleotide operably connected to an expression vector. The phrase operably connected refers to a certain way in which a polynucleotide molecule is inserted into an expression vector that can be expressed when transformed into a host cell. As used herein, an expression vector is a DNA or RNA vector that can transform a host cell and affect the expression of a specific polynucleotide molecule. Preferably, the expression vector can also be replicated in the host cell. The expression vector can be prokaryotic or eukaryotic, typically a virus or a plasmid. Expression vectors include any vector that works (i.e., direct gene expression) in recombinant cells, including bacteria, fungi, endoparasites, arthropods, animals, and plant cells. The vector of the present invention can also be used to produce polypeptides in a cell-free expression system, and this system is well known in the art.

[0203] As used herein, "operably linked" refers to the functional relationship between two or more nucleic acid (e.g., DNA) fragments. Typically, it refers to the functional relationship between a transcriptional regulatory element and a transcribed sequence. For example, if a promoter stimulates or regulates the transcription of a coding sequence in an appropriate host cell and / or cell-free expression system, the promoter is operably linked to the coding sequence (a polynucleotide as defined herein). Typically, promoter transcriptional regulatory elements operably linked to a transcribed sequence are physically adjacent to the transcribed sequence, i.e., they are cis-acting. However, some transcriptional regulatory elements, such as enhancers, do not need to be physically continuous or located near the coding sequence that they enhance transcription.

[0204] In particular, expression vectors according to the present invention comprise regulatory sequences that are compatible with recombinant cells and that control the expression of polynucleotide molecules of the present invention, such as transcriptional control sequences, translational control sequences, replication origins, and other regulatory sequences. In particular, recombinant molecules of the present invention include transcriptional control sequences. Transcriptional control sequences are sequences that control the initiation, elongation, and termination of transcription. Particularly important transcriptional control sequences are those that control the initiation of transcription, such as promoters, enhancers, operators, and repressor sequences. Suitable transcriptional control sequences include any transcriptional control sequence that can function in at least one of the recombinant cells of the present invention. A variety of such transcriptional control sequences are known to those skilled in the art. Preferred transcription control sequences include those that function in bacterial, yeast, arthropod, nematode, plant, or animal cells, such as, but not limited to, tac, lac, tip, trc, oxy-pro, omp / lpp, rmB, bacteriophage lambda, bacteriophage T7, T7lac, bacteriophage T3, bacteriophage SP6, bacteriophage SP01, metallothioneins, alpha mating factors, Pichia pastoris alcohol oxidase, alphavirus subgenomic promoters (e.g., Sindbis virus subgenomic promoters), antibiotic resistance genes, baculovirus, Heliothis tea insectvirus, vaccinia virus, herpes virus, raccoon pox virus, other poxviruses, adenovirus, cytomegalovirus (e.g., intermediate early promoter), simian virus 40, retrovirus, actin, retroviral long terminal repeats, Rous sarcoma virus, heat shock, phosphate and nitrate transcription control sequences, and other sequences capable of controlling gene expression in prokaryotic or eukaryotic cells.

[0205] Another embodiment of the present invention includes a host cell or its daughter cell transformed with one or more recombinant molecules as described herein. The conversion of polynucleotide molecules to cells can be achieved by any method by which the polynucleotide molecules are inserted into cells. Conversion techniques include, but are not limited to, transfection, electroporation, microinjection, lipofection, adsorption, and protoplast fusion. Recombinant cells can remain in a single cell state or can grow into tissues, organs, or multicellular organisms. The polynucleotide molecules of the conversion of the present invention can remain outside the chromosome, or can be integrated into one or more sites in the cell chromosome of the conversion (i.e., recombinant) in a manner that retains its expressivity.

[0206] Host cells suitable for transformation include any cells that can be transformed with the polynucleotides of the present invention. The host cells of the present invention can either be endogenous (i.e., natural) to produce polypeptides as described herein, or produce such polypeptides after transformation with at least one polynucleotide molecule as described herein. The host cells of the present invention can be any cells capable of producing at least one protein defined herein, including bacteria, fungi (including yeast), parasites, nematodes, arthropods, animals, and plant cells. The example of a host cell includes Salmonella, Escherichia, Bacillus, Listeria, yeast, Spodoptera, Mycobacterium, Trichoplusia, BHK (baby hamster kidney) cells, MDCK cells, CRFK cells, CV-1 cells, COS (e.g., COS-7) cells, and Vero cells. Other examples of host cells are Escherichia coli, including Escherichia coli K-12 derivatives; Salmonella typhi; Salmonella typhimurium, including attenuated strains; Spodoptera litura; Trichoplusia ni; and non-tumorigenic mouse myoblast G8 cells (e.g., ATCC CRL 1246). Useful yeast cells include Pichia, Aspergillus, and Saccharomyces. Particularly preferred host cells are bacterial cells, fungal cells, or plant cells.

[0207] In one embodiment, the cell is suitable for fermentation.The example of useful bacterial cells for fermentation includes but is not limited to Escherichia coli (such as Escherichia coli), Bacillus (such as Bacillus subtilis and Bacillus licheniformis), Lactobacillus (such as Lactobacillus brevis), Pseudomonas (such as Pseudomonas aeruginosa) and Streptomyces (Streptomyces lividans).The example of useful fungal cells for fermentation includes but is not limited to Candida (such as Candida albicans), Hansenula (such as Hansenula polymorpha), Pichia (Pichia pastoris), Kluyveromyces (such as Kluyveromyces maize) and Saccharomyces (Saccharomyces cerevisiae).

[0208] Recombinant DNA technology can be used to improve the expression of transformed polynucleotide molecules by manipulating the copy number of polynucleotide molecules in the host cell, the efficiency of transcription of these polynucleotide molecules, the efficiency of translation of the resulting transcripts, and the efficiency of post-translational modifications. Recombinant techniques that can be used to increase the expression of polynucleotide molecules of the present invention include, but are not limited to, operably linking polynucleotide molecules to high-copy number plasmids, integrating polynucleotide molecules into one or more host cell chromosomes, adding vector stabilizing sequences to plasmids, replacing or modifying transcriptional control signals (e.g., promoters, operators, enhancers), replacing or modifying translational control signals (e.g., ribosome binding sites, Shine-Dalgarno sequences), modifying the polynucleotide molecules of the present invention to correspond to the codon usage of the host cell, and deleting sequences that destabilize transcripts.

[0209] The term "plant" as used herein refers to a whole plant, e.g., a plant grown in a commercial plant or food production field. "Plant parts" refer to vegetative structures (e.g., leaves, stems), roots, floral organs / structures, seeds (including embryos, endosperms, and seed coats), plant tissues (e.g., vascular tissues, matrix tissues, etc.), cells, and their progeny.

[0210] "Transgenic plants" refers to plants that contain a genetic construct ("transgene") that is not found in wild-type plants of the same species, variety, or cultivar. "Transgenic," as used herein, has its usual meaning in the field of biotechnology and includes genetic sequences that have been created or altered and introduced into plant cells by recombinant DNA or RNA techniques. A transgene may include a genetic sequence derived from a plant cell. Typically, a transgene is introduced into a plant through human manipulation, such as transformation, but as will be appreciated by those skilled in the art, any method may be used.

[0211] The polynucleotides of the present invention can be constitutively expressed in all developmental stages of transgenic plants. Depending on the purpose of the plant or plant organ, the polypeptide can be expressed in a stage-specific manner. In addition, the polynucleotides can be expressed in a tissue-specific manner.

[0212] The compositions of the present invention include an excipient, also referred to herein as an "acceptable carrier." An excipient can be any material that can be tolerated by the animal, plant, plant or animal material, or environment (including soil and water samples) to be treated. Examples of such excipients include water, physiological saline, Ringer's solution, dextrose solution, Hank's solution, and other physiologically balanced saline solutions. Non-aqueous carriers such as fixed oils, sesame oil, ethyl oleate, or triglycerides can also be used. Other useful formulations include suspensions containing viscosity enhancers such as sodium carboxymethylcellulose, sorbitol, or dextran. Excipients may also contain small amounts of additives such as substances that enhance isotonicity and chemical stability. Examples of buffers include phosphate buffer, bicarbonate buffer, and Tris buffer, while examples of preservatives include thimerosal or o-cresol, formalin, and benzyl alcohol. Excipients may also be used to extend the half-life of the composition, such as, but not limited to, polymer controlled-release carriers, biodegradable implants, liposomes, bacteria, viruses, other cells, oils, esters, and glycols.

[0213] Each document, reference, patent application, or patent cited herein is expressly incorporated herein by reference, meaning that the reader should read and consider it as part of this disclosure. Documents, references, patent applications, or patents cited herein are not repeated herein solely for the sake of clarity. Inclusion does not constitute an admission that any reference constitutes prior art or is part of the common general knowledge of those working in the field relevant to the present invention.

[0214] Optional embodiments of the invention may also be considered to broadly include the parts, elements and features referred to or indicated herein, individually or collectively, and any or all combinations of two or more parts, elements or features, where specific integers referred to herein have known equivalents in the art to which the invention pertains, such known equivalents are deemed to be incorporated herein as if individually set forth.

[0215] It should be understood that references to "one example" or "an example" of the present invention are not exclusive. Thus, one example may illustrate certain aspects of the invention, while other aspects are illustrated in different examples. These examples are intended to aid those skilled in the art in practicing the invention and are not intended to limit the overall scope of the invention in any way, unless the context clearly indicates otherwise.

[0216] It should be understood that the above terms are used for descriptive purposes only and should not be considered as limiting. The described embodiments are intended to illustrate the present invention and not to limit its scope. The present invention can be implemented by various modifications and additions that are readily apparent to those skilled in the art.

[0217] Other definitions of selected terms used herein can be found in the detailed description of the invention and throughout. Unless otherwise defined, all other scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.

[0218] Various substantial and specific practical and useful exemplary embodiments of the claimed subject matter are described herein in text and / or graphical form, including the best mode, if any, known to the inventors for carrying out the claimed subject matter.

[0219] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications. The invention also includes all steps, features, compositions and compounds referred to or indicated in the specification, individually or collectively, as well as any step or feature and any combination of all steps or features or any two or more steps or features.

[0220] The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend that the claimed subject matter be practiced otherwise than as specifically described herein. Accordingly, the claimed subject matter includes and encompasses all equivalents of the claimed subject matter and all modifications thereof as permitted by law. Moreover, every combination of the above-described elements, activities, and all possible variations thereof are included within the scope of the claimed subject matter unless otherwise expressly indicated herein, expressly and specifically denied, or clearly contradicted by context.

[0221] The present invention is not to be limited in scope by the specific embodiments described herein, which are intended as illustrations only. Functionally equivalent products, compositions, and methods are clearly within the scope of the invention described herein.

[0222] Unless otherwise specified, the use of any and all examples or exemplary language (e.g., "such as" or "for example") provided herein is intended merely to better illuminate one or more embodiments and does not limit the scope of any claimed subject matter. No language in the specification should be construed as indicating any non-claimed subject matter is essential to the practice of the claimed subject matter.

[0223] Throughout the specification and claims, unless the context requires otherwise, the word "comprise" or variations such as "comprises" or "comprising", will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.

[0224] Unless the context requires otherwise, throughout this specification, the word "include" or variations such as "includes" or "including" will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers.

[0225] In addition, when any number or range is described herein, unless otherwise expressly stated, the number or range is approximate. Unless otherwise stated herein, reference to a range of values ​​herein is intended solely as a shorthand method of individually referring to each individual value falling within the range, and each individual value and each individual subrange defined by these individual values ​​is incorporated into the specification as if it were individually cited herein. For example, if a range of 1 to 10 is described, the range includes all values ​​therebetween, such as 1.1, 2.5, 3.335, 5, 6.179, 8.9999, etc., and includes all subranges therebetween, such as 1 to 3.65, 2.8 to 8.14, 1.93 to 9, etc.

[0226] Therefore, except the claims themselves, each part of this application (such as title, field, background, invention content, description, abstract, drawings, etc.) should be regarded as illustrative and not restrictive; and the scope of the subject matter protected by any patent issued based on this application shall be defined solely by the claims of that patent.

[0227] While other embodiments of the present application are shown and described, it is to be distinctly understood that the application is not limited thereto but may be otherwise embodied and practiced within the scope of the appended claims.

[0228] Any feature of the embodiment of one or more aspects is applicable to all other aspects and embodiments described herein.Any feature of an embodiment can be combined with other embodiments described herein independently, in part or in whole in any way, for example, one, two or three or more embodiments can be combined in whole or in part.

Claims

1. A polypeptide comprising: i) the amino acid sequence provided in SEQ ID NO: 4, ii) an amino acid sequence that is at least 90% identical to i), or iii) biologically active fragments of ii).

2. The polypeptide according to claim 1, wherein the polypeptide comprises one or both of the following: i) glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and ii) Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO:

4.

3. The polypeptide according to claim 1 or 2, comprising a fusion protein comprising at least one other polypeptide sequence.

4. An isolated and / or exogenous polynucleotide comprising a sequence selected from the group consisting of: i) the nucleotide sequence provided in SEQ ID NO: 2; ii) a nucleotide sequence encoding the polypeptide according to claim 1 or 2; or iii) a nucleotide sequence complementary to i) or ii).

5. The isolated and / or exogenous polynucleotide according to claim 4, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the polypeptide according to claim 1 or 2 in a cell.

6. The isolated and / or exogenous polynucleotide according to claim 4, wherein the polynucleotide is operably linked to a promoter capable of directing expression of the polypeptide according to claim 1 or 2 in an expression host cell. A vector comprising the polynucleotide according to claim 4 . A nucleic acid construct or expression vector comprising the polynucleotide according to claim 4 .

9. A nucleic acid construct or expression vector comprising the polynucleotide of claim 4, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide of claim 1 or 2 in an expression host cell.

10. A recombinant expression host cell comprising a polynucleotide encoding the polypeptide according to claim 1 or 2, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide. A host cell comprising the polynucleotide according to claim 4 .

12. The host cell according to claim 11, which comprises a bacterial cell, a fungal cell or a plant cell.

13. A transgenic non-human organism comprising at least one cell according to claim 11 or 12.

14. The host cell extract of claim 11 or 12, wherein the extract comprises a polypeptide comprising: i) the amino acid sequence provided in SEQ ID NO: 4; ii) an amino acid sequence that is at least 90% identical to i), or iii) biologically active fragments of ii).

15. The host cell extract of claim 14, wherein the polypeptide comprises one or both of the following: i) glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and ii) Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO:

4.

16. A composition comprising the polypeptide according to claim 1 or 2 and one or more acceptable carriers.

17. A composition comprising the extract according to claim 14 or 15 and one or more acceptable carriers.

18. A method for producing a polypeptide, comprising: i) the amino acid sequence provided in SEQ ID NO: 4; ii) an amino acid sequence that is at least 90% identical to i), or iii) biologically active fragments of ii).

19. The method of claim 18, wherein the polypeptide comprises one or both of the following: i) glutamine (Q) at amino acid position 221 corresponding to SEQ ID NO: 4; and ii) Aspartic acid (D) at the amino acid position corresponding to amino acid position 377 of SEQ ID NO:

4.

20. A method for producing a polypeptide according to claim 1 or 2, comprising culturing a recombinant expression host cell comprising a polynucleotide according to claim 4, wherein the polynucleotide is operably linked to one or more control sequences that direct the production of the polypeptide according to claim 1 or 2 under conditions that are conducive to the production of the polypeptide.

21. The method of claim 21, further comprising recovering the polypeptide of claim 1 or 2.

22. An enzyme comprising the polypeptide according to claim 1 or 2. A thermostable enzyme comprising the polypeptide according to claim 1 or 2.

24. A hyperthermophilic enzyme comprising the polypeptide according to claim 1 or 2.

25. The enzyme of any one of claims 22 to 24, wherein the enzyme retains enzymatic activity at a temperature above the temperature at which an enzyme comprising a wild-type Thermotrophic bacterium lipolytica T1_Est47 polypeptide loses substantially all enzymatic activity, the enzyme comprising the amino acid sequence provided in SEQ ID NO:

3.

26. The enzyme according to any one of claims 22 to 24, wherein the enzyme maintains enzymatic activity at 90°C or higher.

27. The enzyme according to any one of claims 22 to 24, wherein the enzyme maintains a certain enzyme activity at 95°C or higher.

28. The enzyme according to any one of claims 22 to 27, which has esterase activity.

29. The enzyme according to any one of claims 22 to 27, which has lipase activity.

30. The polypeptide of claim 1 or 2, comprising a polypeptide variant of wild-type Thermotrophic bacterium lipolytica T1-Est47, wherein the peptide variant of wild-type Thermotrophic bacterium lipolytica T1-Est47 comprises the amino acid sequence provided in SEQ ID NO:

3.

31. The polynucleotide of claim 4, capable of expressing a polypeptide variant of wild-type Thermotrophic bacterium lipolytica T1-Est47, wherein the wild-type Thermotrophic bacterium lipolytica T1-Est47 comprises an amino acid sequence expressed by an isolated and / or exogenous polynucleotide comprising a sequence selected from the nucleotide sequence provided in SEQ ID NO: 1.