Carbohydrate binding module variants and polynucleotides encoding the same

By substituting specific sites in the carbohydrate binding module and cellobiase, peptide variants with improved binding activity were developed, solving the problem of low cellulose conversion efficiency and achieving efficient conversion of cellulose materials and increased yield of fermentation products.

CN107002056BActive Publication Date: 2026-07-21NOVOZYMES AS
View PDF 150 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NOVOZYMES AS
Filing Date
2015-09-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing carbohydrate binding module variants have insufficient binding affinity during the conversion of cellulose to ethanol, resulting in low conversion efficiency of cellulose materials.

Method used

By substituting specific sites in the carbohydrate-binding module and cellobiase, peptide variants with improved binding activity were developed, and heterologous catalytic domains of cellulase were combined to form hybrid peptides for processing cellulose materials.

Benefits of technology

It improved the conversion efficiency of cellulose materials, enhanced the activity of cellulase, promoted the production of cellobiose and glucose, and increased the yield of fermentation products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0001238777810000011
    Figure HDA0001238777810000011
  • Figure HDA0001238777810000021
    Figure HDA0001238777810000021
  • Figure HDA0001238777810000031
    Figure HDA0001238777810000031
Patent Text Reader

Abstract

The present invention relates to cellobiohydrolase variants and carbohydrate binding module variants. The present invention also relates to polynucleotides encoding the variants; nucleic acid constructs, vectors, and host cells comprising the polynucleotides; and methods of using the variants.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 046,344, filed September 5, 2014. The entire contents of that application are incorporated herein by reference.

[0003] Reference sequence list

[0004] This application includes a sequence list in computer-readable form, which is incorporated herein by reference. Background of the Invention

[0006] Invention Field

[0007] The present invention relates to polypeptides including carbohydrate-binding module variants, polynucleotides encoding these variants, methods for generating these variants, and methods for using these variants.

[0008] Related technical descriptions

[0009] Cellulose is a polymer of glucose monosaccharides covalently linked by β-1,4-bonds. Many microorganisms produce enzymes that hydrolyze β-linked glucans. These enzymes include endoglucanases, cellobiases, and β-glucosidases. Endoglucanases digest the cellulose polymer at any point, opening it for attack by cellobiases. Cellobiases sequentially release cellobiose molecules from the ends of the cellulose polymer. Cellobiose is a water-soluble β-1,4-linked dimer of glucose. β-glucosidases hydrolyze cellobiose into glucose.

[0010] Converting lignocellulosic feedstocks into ethanol offers several advantages: readily available large quantities of raw materials, avoidance of the desire to burn or landfill materials, and the cleanliness of ethanol fuel. Wood, agricultural residues, herbaceous crops, and municipal solid waste are considered feedstocks for ethanol production. These materials are primarily composed of cellulose, hemicellulose, and lignin. Once lignocellulosic acid is converted into fermentable sugars (e.g., glucose), these sugars are readily fermented into ethanol using yeast.

[0011] Modified carbohydrate-binding modules with reduced lignin binding have been described (WO 2011 / 097713A1; Linder et al., 1995, Protein Science 4:1056-1064; and Linder et al., 1999, FEBS 447:13-16). Further variants of the carbohydrate-binding module have been described in WO2012 / 135719.

[0012] Hybrid polypeptides including cellobiase catalytic domains and carbohydrate-binding modules are described, for example, in WO 2010 / 060056, WO 2013 / 091577 and WO 2014 / 138672.

[0013] It will be advantageous in the art to provide peptides comprising carbohydrate-binding module variants (e.g., cellobiose hydrolase variants) with improved properties (e.g., increased binding affinity) for converting cellulosic materials into monosaccharides, disaccharides, and polysaccharides.

[0014] The present invention provides peptides comprising carbohydrate-binding module variants having improved properties compared to their parents. Invention Overview

[0016] This invention relates to carbohydrate-binding module variants comprising substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31, and 32 of the carbohydrate-binding module corresponding to SEQ ID NO:4, wherein these variants have carbohydrate-binding activity. In one aspect, cellulases include the carbohydrate-binding module variants of this invention. In some embodiments, these carbohydrate-binding module variants have improved binding activity.

[0017] The present invention also relates to isolated cellobiose hydrolase variants comprising substitutions at one or more (e.g., several) positions corresponding to positions 483, 491, 509 and 510 of SEQ ID NO:2, wherein these variants have cellobiose hydrolase activity.

[0018] The present invention also relates to isolated hybrid polypeptides comprising a heterologous catalytic domain of a carbohydrate-binding module variant described herein and a cellulase. In one aspect, this catalytic domain is a cellobiase catalytic domain.

[0019] The present invention also relates to hybrid polypeptides comprising heterologous catalytic domains of carbohydrate-binding module variants and cellulases described herein. In one aspect, this catalytic domain is a cellobiase catalytic domain.

[0020] The present invention also relates to isolated polynucleotides encoding these variants and hybrid polypeptides; nucleic acid constructs, vectors and host cells containing these polynucleotides; and methods for producing these variants and hybrid polypeptides.

[0021] The present invention also relates to a method for degrading or converting cellulosic materials, the method comprising treating the cellulosic material with an enzyme composition in the presence of a cellobiose hydrolase variant or hybrid polypeptide of the present invention. In one aspect, the method further comprises recovering the degraded or converted cellulosic material.

[0022] The present invention also relates to methods for producing fermentation products, the methods comprising: (a) saccharifying a cellulose material with an enzyme composition in the presence of a cellobiase variant or a hybrid polypeptide of the present invention; (b) fermenting the saccharified cellulose material with one or more (e.g., several) fermenting microorganisms to produce the fermentation product; and (c) recovering the fermentation product from the fermentation.

[0023] The present invention also relates to methods for fermenting cellulosic materials, the methods comprising: fermenting the cellulosic material with one or more (e.g., several) fermenting microorganisms, wherein the cellulosic material is saccharified with an enzyme composition in the presence of a cellobiose hydrolase variant or hybrid polypeptide of the present invention. In one aspect, the fermentation of the cellulosic material produces a fermentation product. In another aspect, the method further comprises recovering the fermentation product from the fermentation. Brief description of the attached diagram

[0025] Figure 1 The cDNA sequence (SEQ ID NO:31) and deduced amino acid sequence (SEQ ID NO:2) of the *Trichoderma reesei* cellobiose hydrolase I gene are shown. The signal peptide is shown in italics. The carbohydrate binding module is underlined.

[0026] Figure 2 The hydrolysis of microcrystalline cellulose by R. emersonii wild-type cellobiose hydrolase I and hybrid peptides PC1-147, PC1-499, and PC1-500 is shown. Values ​​in mM represent the cellobiose released after 24 hours at pH 5 and 50°C.

[0027] Figure 3 The hydrolysis of microcrystalline cellulose by R. emersonii wild-type cellobiose hydrolase I and hybrid peptides PC1-147, PC1-499, and PC1-500 is shown. Values ​​in mM represent the cellobiose released after 24 hours at pH 5 and 60°C.

[0028] Figure 4 This shows a comparison of the percentage of cellulose conversion of pretreated corn stalks at 35°C, 50°C, and 60°C using enzyme compositions comprising hybrid peptides PC1-147, PC1-499, or PC1-500.

[0029] Figure 5 This shows a comparison of the percentage of cellulose conversion of pretreated corn stalks at 35°C, 50°C, and 60°C using enzyme compositions comprising heterozygous peptides PC1-147, PC1-499, PC1-500, or PC1-668.

[0030] Figure 6This shows a comparison of the percentage of cellulose conversion of pretreated corn stalks at 35°C, 50°C, and 60°C using an enzyme composition comprising peptides AC1-596, AC1-660, or AC1-661.

[0031] Figure 7 This shows a comparison of the percentage of cellulose conversion of pretreated corn stalks at 35°C, 50°C, and 60°C using an enzyme composition including peptide PC1-147 or PC1-899.

[0032] definition

[0033] Acetylxylan esterase: The term "acetylxylan esterase" refers to a carboxylesterase (EC 3.1.1.72) that catalyzes the hydrolysis of acetyl groups from polyxylan, acetylated xylose, acetylated glucose, α-naphthyl acetate, and p-nitrophenyl acetate. It can be used in the presence of 0.01% TWEEN. TM Acetylxylan esterase activity was determined using 0.5 mM p-nitrophenyl acetate as a substrate in 50 mM sodium acetate (pH 5.0) of 20 (polyoxyethylene sorbitan monolaurate). One unit of acetylxylan esterase was defined as the amount of enzyme capable of releasing 1 μmol of p-nitrophenol anion per minute at pH 5 and 25°C.

[0034] Allelic variants: The term "allelic variant" refers to any of two or more alternative forms of a gene occupying the same chromosomal locus. Allelic variations arise naturally from mutations and can lead to polymorphism within a population. Gene mutations can be silent (without alteration in the encoded polypeptide) or can encode a polypeptide with a modified amino acid sequence. Allelic variants of a polypeptide are polypeptides encoded by allelic variants of a gene.

[0035] α-L-Arabofuranosaccharidase: The term "α-L-Arabofuranosaccharidase" refers to an α-L-arabinofuranoside arabinofuranoside hydrolase (EC 3.2.1.55) that catalyzes the hydrolysis of terminal non-reducing α-L-arabinofuranoside residues in α-L-arabinoside. This enzyme acts on α-L-arabinofuranosides, α-L-arabinanan containing (1,3)- and / or (1,5)- bonds, arabinosylxylan, and arabinogalactan. α-L-Arabofuranosaccharidase is also known as arabinosaccharidase, α-arabinosaccharidase, α-L-arabinosaccharidase, α-arabinofuranosaccharidase, polysaccharide α-L-arabinofuranosaccharidase, α-L-arabinofuranoside hydrolase, L-arabinosaccharidase, or α-L-arabinananase. Medium-viscosity wheat arabinosyl xylan (Megazyme International Ireland, Ltd., Bray, Co., Wicklow, Ireland) can be used at a total volume of 200 μl per ml of 100 mM sodium acetate (pH 5) for 30 minutes at 40°C, followed by... HPX-87H column chromatography (Bio-Rad Laboratories, Inc., Hercules, California, USA) was used for arabinose analysis to determine α-L-arabinofuranosidase activity.

[0036] α-Glucuronidase: The term "α-glucuronidase" refers to an α-D-glucuronide hydrolase (EC 3.2.1.139) that catalyzes the hydrolysis of α-D-glucuronide to D-glucuronide and alcohol. α-Glucuronidase activity can be determined according to de Vries, 1998, J. Bacteriol. 180:243-249. One unit of α-glucuronidase is equal to the amount of enzyme capable of releasing 1 micromole of glucuronic acid or 4-O-methylglucuronic acid per minute at pH 5 and 40°C.

[0037] Co-active 9 polypeptide: The term “co-active 9 polypeptide” or “AA9 polypeptide” refers to polypeptides classified as lytic polysaccharide monooxygenases (Quinlan et al., 2011, Proc. Natl. Acad. Sci. USA 208:15079-15084; Phillips et al., 2011, ACS Chem. Biol. 6:1399-1406; Lin et al., 2012, Structure 20:1051-1061). According to Henrissat, 1991, Biochem.J. 280:309-316; and Henrissat and Bairoch, 1996, Biochem.J. 316:695-696, the AA9 polypeptide was previously classified as glycoside hydrolase family 61 (GH61).

[0038] AA9 peptides enhance the hydrolysis of cellulose materials through enzymes with cellulose-degrading activity. Cellulose-enhancing activity can be determined by measuring the increase in reducing sugars or the total amount of cellobiose and glucose in the cellulose material hydrolyzed by cellulase, compared to a control hydrolysis with an equivalent total protein load (1-50 mg of cellulase protein / g of cellulose in PCS) without cellulose-enhancing activity, under the following conditions: 1-50 mg of total protein / g of cellulose in pretreated corn stalks (PCS), wherein the total protein consists of 50%-99.5% w / w cellulase protein and 0.5%-50% w / w AA9 polypeptide protein, for 1-7 days at suitable temperatures (e.g., 40°C-80°C, e.g., 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C) and suitable pH (e.g., 4-9, e.g., 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, or 9.0).

[0039] CELLUCLAST can be used TMAA9 polypeptide-enhancing activity was determined using a mixture of 1.5 L (Novozymes, Bagsvaerd, Denmark) and β-glucosidase as a source of cellulolytic activity, wherein the β-glucosidase was present at a weight of at least 2%-5% of the cellulase protein loaded with the protein. In one aspect, the β-glucosidase was Aspergillus oryzae β-glucosidase (e.g., recombinantly produced in Aspergillus oryzae according to WO 02 / 095014). In the other aspect, the β-glucosidase was Aspergillus fumigatus β-glucosidase (e.g., recombinantly produced in Aspergillus oryzae as described in WO 02 / 095014).

[0040] The enhanced activity of AA9 peptide can also be determined by reacting AA9 peptide with 0.5% phosphate-swellable cellulose (PASC), 100 mM sodium acetate (pH 5), 1 mM MnSO4, 0.1% gallic acid, 0.025 mg / ml Aspergillus fumigatus β-glucosidase, and 0.01% [unclear - possibly a specific ingredient or solution] at 40°C. Incubate with X-100 (4-(1,1,3,3-tetramethylbutyl)phenyl-polyethylene glycol) for 24-96 hours, then measure the glucose released from PASC.

[0041] The AA9 peptide enhancement activity of the high-temperature composition can also be determined according to WO 2013 / 028928.

[0042] The AA9 peptide enhances the hydrolysis of cellulose materials catalyzed by enzymes with cellulose-degrading activity by reducing the amount of cellulase required to achieve the same degree of hydrolysis by preferably at least 1.01 times, for example, at least 1.05 times, at least 1.10 times, at least 1.25 times, at least 1.5 times, at least 2 times, at least 3 times, at least 4 times, at least 5 times, at least 10 times, or at least 20 times.

[0043] According to WO 2008 / 151043 or WO 2012 / 122518, the AA9 peptide can also be used in the presence of a soluble activated divalent metal cation (e.g., manganese or copper).

[0044] This AA9 polypeptide can be used in the presence of dioxins, bicyclic compounds, heterocyclic compounds, nitrogen-containing compounds, quinone compounds, sulfur-containing compounds, or liquids obtained from pretreated cellulose or hemicellulose materials (such as pretreated corn stalks) (WO 2012 / 021394, WO 2012 / 021395, WO 2012 / 021396, WO 2012 / 021399, WO2012 / 021400, WO 2012 / 021401, WO 2012 / 021408, and WO 2012 / 021410).

[0045] β-Glucosidase: The term "β-glucosidase" refers to β-D-glucosidase (EC 3.2.1.21), which catalyzes the hydrolysis of terminal non-reducing β-D-glucose residues, releasing β-D-glucose. β-glucosidase activity can be determined using p-nitrophenyl-β-D-glucopyranoside as a substrate, according to the procedure described by Venturi et al., 2002, J. Basic Microbiol. 42:55-66. One unit of β-glucosidase is defined as the concentration of 0.01% glucosidase at 25°C and pH 4.8. 1.0 μmol of p-nitrophenol anion per minute is generated from 1 mM p-nitrophenyl-β-D-glucopyranoside as a substrate in 50 mM sodium citrate.

[0046] β-Xylosidase: The term "β-xylosidase" refers to β-D-xyloside xylose hydrolase (EC 3.2.1.37), which catalyzes the external hydrolysis of short β(1→4) xylooligosaccharides to remove consecutive D-xylose residues from the non-reducing end. It can be present in 0.01%... β-xylosidase activity was determined using 1 mM p-nitrophenyl-β-D-xyloside as a substrate in 20% sodium citrate at pH 5 and 40°C. One unit of β-xylosidase was defined as the activity of β-xylosidase at 40°C and pH 5 in a solution containing 0.01% sodium citrate. 1.0 μmol of p-nitrophenol anion is generated per minute from 1 mM p-nitrophenyl-β-D-xyloside in 20 oz.

[0047] Carbohydrate-binding modules: The term "carbohydrate-binding module" refers to the region within a carbohydrate-active enzyme that provides carbohydrate-binding activity (Boraston et al., 2004, Biochem.J. 383:769-781). Most known carbohydrate-binding modules (CBMs) are continuous amino acid sequences with discrete folds. Carbohydrate-binding modules (CBMs) are typically found at the N-terminus or C-terminus of an enzyme. Some CBMs are known to be specific for cellulose.

[0048] Catalase: The term "catalase" refers to hydrogen peroxide:hydrogen peroxide reductase (EC 1.11.1.6), which catalyzes the conversion of 2H₂O₂ to O₂ + 2H₂O. For the purposes of this invention, catalase activity was determined according to U.S. Patent No. 5,646,025. One unit of catalase activity is equal to the amount of enzyme that catalyzes 1 micromolar of hydrogen peroxide under the assay conditions.

[0049] Catalytic domain: The term "catalytic domain" refers to the region of an enzyme containing the catalytic machinery of that enzyme. In one aspect, the catalytic domain is amino acids 1 to 429 of SEQ ID NO:30. In another aspect, the catalytic domain is amino acids 1 to 437 of SEQ ID NO:36. In another aspect, the catalytic domain is amino acids 1 to 440 of SEQ ID NO:38. In another aspect, the catalytic domain is amino acids 1 to 437 of SEQ ID NO:40. In another aspect, the catalytic domain is amino acids 1 to 437 of SEQ ID NO:42. In another aspect, the catalytic domain is amino acids 1 to 438 of SEQ ID NO:44. In another aspect, the catalytic domain is amino acids 1 to 437 of SEQ ID NO:46. In another aspect, the catalytic domain is amino acids 1 to 430 of SEQ ID NO:48. In another aspect, the catalytic domain is amino acids 1 to 433 of SEQ ID NO:50.

[0050] Catalytic domain coding sequence: The term "catalytic domain coding sequence" refers to a polynucleotide encoding a catalytically active catalytic domain. In one aspect, the catalytic domain coding sequence is nucleotides 52 to 1469 of SEQ ID NO:29. In another aspect, the catalytic domain coding sequence is nucleotides 52 to 1389 of SEQ ID NO:31. In another aspect, the catalytic domain coding sequence is nucleotides 52 to 1389 of SEQ ID NO:32. In another aspect, the catalytic domain coding sequence is nucleotides 79 to 1389 of SEQ ID NO:35. In another aspect, the catalytic domain coding sequence is nucleotides 52 to 1371 of SEQ ID NO:37. In another aspect, the catalytic domain coding sequence is nucleotides 55 to 1482 of SEQ ID NO:39. In another aspect, the catalytic domain coding sequence is nucleotides 76 to 1386 of SEQ ID NO:41. In another aspect, the catalytic domain is nucleotides 76 to 1386 of SEQ ID NO:43. In another aspect, the catalytic domain coding sequence is nucleotides 55 to 1504 of SEQ ID NO:45. In another aspect, the catalytic domain coding sequence is nucleotides 61 to 1350 of SEQ ID NO:47. In another aspect, the catalytic domain coding sequence is nucleotides 55 to 1353 of SEQ ID NO:49.

[0051] cDNA: The term "cDNA" refers to a DNA molecule that can be prepared by reverse transcription from mature, spliced ​​mRNA molecules derived from eukaryotic or prokaryotic cells. cDNA lacks the intron sequences that can be present in the corresponding genomic DNA. Early initial RNA transcripts are precursors to mRNA, undergoing a series of processing steps, including splicing, before becoming mature, spliced ​​mRNA.

[0052] Cellobiose hydrolase: The term “cellobiose hydrolase” refers to a 1,4-β-D-glucan-cellobiose hydrolase (EC3.2.1.91 and EC3.2.1.176) that catalyzes the hydrolysis of 1,4-β-D-glycosidic bonds in cellulose, cellooligosaccharides, or any polymer containing β-1,4-linked glucose, thereby releasing cellobiose from either the reducing end (cellobiose hydrolase I) or the non-reducing end (cellobiose hydrolase II) of the chain (Teeri, 1997, Trends in Biotechnology 15:160-167; Teeri et al., 1998, Biochem.Soc.Trans. 26:173-178). Cellobiase activity can be determined according to the procedures described by Lever et al., 1972, Analytical Biochemistry, 47:273-279; van Tilbeurgh et al., 1982, FEBS Letters, 149:152-156; van Tilbeurgh and Claeyssens, 1985, FEBS Letters, 187:283-288; and Tomme et al., 1988, European Journal of Biochemistry, 170:575-581.

[0053] Cellulose-degrading enzymes or cellulases: The term “cellulose-degrading enzyme” or “cellulase” refers to one or more (e.g., several) enzymes that hydrolyze cellulose materials. Such enzymes include one or more endoglucanases, one or more cellobiases, one or more β-glucosidases, or combinations thereof. Two basic methods for measuring cellulose-degrading enzyme activity include (1) determining total cellulose-degrading enzyme activity and (2) determining the activity of individual cellulose-degrading enzymes (endoglucanases, cellobiases, and β-glucosidases), as described in Zhang et al., 2006, Biotechnology Advances 24:452-481. Total cellulose-degrading enzyme activity can be measured using insoluble substrates, including Whatman No. 1 filter paper, microcrystalline cellulose, bacterial cellulose, algal cellulose, cotton, pretreated lignocellulose, etc. The most common assay of total cellulose degrading activity is the filter paper assay using Whatman No. 1 filter paper as the substrate. This assay was established by the International Union of Pure and Applied Chemistry (IUPAC) (Ghose, 1987, Pure Appl. Chem. 59:257-68).

[0054] Cellulase activity can be determined by measuring the increase in sugars produced / released during the hydrolysis of cellulose material by one or more cellulases, compared to a control hydrolysis without added cellulase protein, under the following conditions: 1-50 mg of cellulase protein / g in pretreated corn stalks (PCS) containing cellulose (or other pretreated cellulose material), for 3-7 days at a suitable temperature (e.g., 40°C-80°C, such as 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C) and a suitable pH (e.g., 4-9, such as 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, or 9.0). Typical conditions are: 1 ml reaction, washed or unwashed PCS, 5% insoluble solids (dry weight), 50 mM sodium acetate (pH 5), 1 mM MnSO4, 50°C, 55°C, or 60°C, 72 hours, via Sugar analysis was performed using HPX-87H column chromatography (Bio-Rad Laboratories, Hercules, California, USA).

[0055] Cellulose Material: The term "cellulose material" refers to any material containing cellulose. The primary polysaccharide in the primary cell walls of biomass is cellulose, followed by hemicellulose, and then pectin. Secondary cell walls, formed after cell cessation of growth, also contain polysaccharides and are reinforced by polymeric lignin covalently cross-linked with hemicellulose. Cellulose is a homopolymer of dehydrated cellobiose and is therefore a linear β-(1-4)-D-glucan, while hemicellulose comprises a variety of compounds, such as xylan, xyloglucan, arabinoylxylan, and mannan, which exist with a series of substituents in complex branched structures. Although cellulose is generally polymorphic, it is found in plant tissues primarily as an insoluble crystalline matrix of parallel glucan chains. Hemicellulose is typically hydrogen-bonded to cellulose along with other hemicelluloses, which helps stabilize the cell wall matrix.

[0056] Cellulose is commonly found in, for example, the stems, leaves, shells, bark, and rachis of plants, or in the leaves, branches, and trees. Cellulose materials can be, but are not limited to: agricultural waste, herbaceous materials (including energy crops), municipal solid waste, pulp and paper mill waste, waste paper, and wood (including forestry waste) (see, for example, Wiselogel et al., 1995, Handbook on Bioethanol (Charles E. Wyman, ed.), pp. 105-118, Taylor & Francis, Washington, D.C.; Wyman, 1994, Bioresource Technology, 50:3-16; Lynd, 1990, Applied Biochemistry and Biotechnology, 24 / 25:695-719; Mosier et al., 1999, Recent Advances in the Bioconversion of Lignocellulose, Advances in Biochemical Engineering / Biotechnology). (Engineering / Biotechnology), T. Scheper, Editor, Vol. 65, pp. 23-40, Springer-Verlag, New York. It should be understood here that cellulose can be in the form of lignin cellulose, a plant cell wall material comprising lignin, cellulose, and hemicellulose in a mixed matrix. On one hand, this cellulose material is any biomass material. On the other hand, this cellulose material is lignocellulose, which comprises cellulose, hemicellulose, and lignin.

[0057] In one embodiment, the cellulosic material is agricultural waste, herbaceous material (including energy crops), municipal solid waste, pulp and paper mill waste, waste paper, or wood (including forestry waste).

[0058] In another embodiment, the cellulose material is reed, bagasse, bamboo, corn cob, corn fiber, corn stalk, awn, rice straw, sugarcane stalk, willow sorghum, or wheat straw.

[0059] In another embodiment, the cellulose material is aspen, eucalyptus, fir, pine, poplar, spruce, or willow.

[0060] In another embodiment, the cellulose material is seaweed cellulose, bacterial cellulose, cotton linters, filter paper, microcrystalline cellulose (e.g., ), or cellulose treated with phosphoric acid.

[0061] In another embodiment, the cellulose material is an aquatic biomass. As used herein, the term "aquatic biomass" means biomass produced in an aquatic environment through the process of photosynthesis. Aquatic biomass can be algae, emergent plants, floating-leaved plants, or submerged plants.

[0062] Cellulose materials can be used as is or pretreated using conventional methods known in the art, as described herein. In a preferred aspect, the cellulose material is pretreated.

[0063] Coding sequence: The term "coding sequence" refers to a polynucleotide that directly identifies the amino acid sequence of a variant. The boundaries of a coding sequence are generally determined by an open reading frame, which begins with a start codon (such as ATG, GTG, or TTG) and ends with a stop codon (such as TAA, TAG, or TGA). A coding sequence can be genomic DNA, cDNA, synthetic DNA, or a combination thereof.

[0064] Control sequences: The term "control sequence" refers to the nucleic acid sequence necessary for the expression of a polynucleotide encoding a variant of the present invention. Each control sequence may be native (i.e., from the same gene) or exogenous (i.e., from a different gene) for the polynucleotide encoding that variant, or native or exogenous relative to each other. These regulatory sequences include, but are not limited to, pro-leaders, polyadenylated sequences, propeptide sequences, promoters, signal peptide sequences, and transcription terminators. At a minimum, control sequences include promoters, as well as transcription and translation termination signals. These control sequences may be provided with multiple linkers for the purpose of introducing specific restriction enzyme sites that facilitate the linking of these control sequences to the coding regions of the polynucleotides encoding the variant.

[0065] Endoglucanase: The term "endoglucanase" refers to a 4-(1,3;1,4)-β-D-glucanase (EC 3.2.1.4) that catalyzes the endo-hydrolysis of β-1,4-β-D-glycosidic bonds in cellulose, cellulose derivatives (such as carboxymethyl cellulose and hydroxyethyl cellulose), lichen polysaccharides, and mixed β-1,3-1,4-glucans such as cereal β-D-glucan or xyloglucan, as well as other plant materials containing cellulose components. Endoglucanase activity can be determined by measuring a decrease in substrate viscosity or an increase in reducing ends as determined by reducing sugar assays (Zhang et al., 2006, Biotechnology Advances 24:452-481). Endoglucanase activity can also be determined using carboxymethyl cellulose (CMC) as a substrate at pH 5 and 40°C, according to the procedure described by Ghose, 1987, Pure and Appl. Chem. 59:257-268.

[0066] Expression: The term “expression” includes any step involved in variant generation, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0067] Expression vector: The term “expression vector” refers to a linear or circular DNA molecule that includes a polynucleotide encoding a variant and that the polynucleotide is operatively linked to a control sequence provided for its expression.

[0068] Ferulic acid esterase: The term "ferulic acid esterase" refers to 4-hydroxy-3-methoxycinnamoyl-sugar hydrolase (EC 3.1.1.73), which catalyzes the hydrolysis of the 4-hydroxy-3-methoxycinnamoyl (feruloyl) group from esterified sugars (which are typically arabinose in natural biomass substrates) to produce ferulic acid esterase (4-hydroxy-3-methoxycinnamoyl esterase). Ferulic acid esterase (FAE) is also known as ferulic acid esterase, hydroxycinnamoyl esterase, FAE-III, cinnamoyl esterase, FAEA, cinnAE, FAE-I, or FAE-II. Its activity can be determined in 50 mM sodium acetate (pH 5.0) using 0.5 mM p-nitrophenylferulate as a substrate. One unit of ferulic acid esterase is equal to the amount of enzyme capable of releasing 1 μmol of p-nitrophenol anion per minute at pH 5 and 25°C.

[0069] Fragment: The term "fragment" means a polypeptide that has one or more (e.g., several) amino acids deleted from the amino and / or carboxyl termini of a mature polypeptide; wherein the fragment has cellobiose hydrolase activity. In one aspect, a fragment comprises at least 425 amino acid residues, for example, at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:2. In another aspect, a fragment comprises at least 425 amino acid residues, for example, at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:6. In another aspect, a fragment comprises at least 425 amino acid residues, for example, at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:10. In another aspect, a fragment comprises at least 425 amino acid residues, for example, at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:14. In another aspect, a fragment comprises at least 425 amino acid residues, such as at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:18. In another aspect, a fragment comprises at least 425 amino acid residues, such as at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:22. In another aspect, a fragment comprises at least 425 amino acid residues, such as at least 450 amino acid residues or at least 475 amino acid residues, of the mature polypeptide of SEQ ID NO:26.

[0070] Hemicellulase or hemicellulase: The term “hemicellulase” or “hemicellulase” refers to one or more (e.g., several) enzymes that hydrolyze hemicellulose materials. See, for example, Sharlom and Shoham, 2003, Current Opinion in Microbiology 6(3):219-228). Hemicellulases are key components in the degradation of plant biomass. Examples of hemicellulases include, but are not limited to, acetylmannan esterase, acetylxylan esterase, arabinonanase, arabinofuranylase, coumarin esterase, ferulic esterase, galactosidase, glucuronidase, glucuronidase, mannanase, mannosidase, xylanase, and xylosidase. The substrate of these enzymes, hemicellulose, is a heterogeneous group of branched and linear polysaccharides that can be cross-linked into a robust network by hydrogen bonds to cellulose microfibers in the plant cell wall. Hemicellulose is also covalently attached to lignin, thus forming a highly complex structure together with cellulose. The variable structure and organization of hemicellulose require the synergistic action of many enzymes to achieve its complete degradation. The catalytic modules of hemicellulases are either glycosidases (GH) that hydrolyze glycosidic bonds, or carbohydrate esterases (CE) that hydrolyze ester bonds on the side groups of acetic acid or ferulic acid. These catalytic modules can be assigned to the GH and CE families based on their primary sequence homology. Some families with generally similar folds can be further grouped into alphabetically labeled clans (e.g., GH-A). The most comprehensive and up-to-date classification of these enzymes, as well as carbohydrate-active enzymes, is available in the Carbohydrate Active Enzymes (CAZY) database. Hemicellulase activity can be measured according to Ghose and Bisaria, 1987, Pure & Applied Chemistry 59:1739-1752, at suitable temperatures such as 40°C–80°C, for example 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C, and suitable pH such as 4–9, for example 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, or 9.0.

[0071] Hemicellulose Material: The term "hemicellulose material" refers to any material containing hemicellulose. Hemicellulose includes xylan, glucuronic acid xylan, arabinoyl xylan, glucomannan, and xyloglucan. These polysaccharides contain many different sugar monomers. Sugar monomers in hemicellulose can include xylose, mannose, galactose, rhamnose, and arabinose. Hemicellulose contains mostly D-pentose sugars. In most cases, xylose is the sugar monomer present in the largest quantity, although mannose can be the most abundant sugar in cork. Xylan contains a backbone of β-(1-4)-linked xylose residues. Terrestrial plant xylans are heteropolymers with a β-(1-4)-D-xylpyranose backbone branched by short carbohydrate chains. They include D-glucuronic acid or its 4-O-methyl ether, L-arabinose, and / or various oligosaccharides composed of D-xylose, L-arabinose, D- or L-galactose, and D-glucose. Xylan-type polysaccharides can be classified into homooxylans and heterooxylans, including glucuronide xylan, (arabinose)glucuronide xylan, (glucuronide)arabinosyl xylan, arabinosyl xylan, and complex heterooxylans. See, for example, Ebringerova et al., 2005, Adv. Polym. Sci. 186:1-67. Hemicellulose materials are also referred to herein as "xylan-containing materials".

[0072] The sources used for hemicellulose materials are essentially the same as those used for cellulose materials described herein.

[0073] In the method of the present invention, any material containing hemicellulose can be used. In a preferred aspect, the hemicellulose material is lignin cellulose.

[0074] Highly stringent conditions: The term "highly stringent conditions" refers to pre-hybridization and hybridization for probes of at least 100 nucleotides in length, following standard DNA blotting procedures at 42°C in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 50% formamide for 12 to 24 hours. Vector material is finally washed three times at 65°C for 15 minutes each time with 0.2X SSC and 0.2% SDS.

[0075] Host cell: The term "host cell" refers to any cell type that is readily transformed, transfected, transduced, etc., using a nucleic acid construct or expression vector containing the polynucleotides of the present invention. The term "host cell" also encompasses any offspring of the parent cell that is not identical to its parent cell due to mutations that occur during replication.

[0076] Heteropolymer polypeptide: The term "heteropolymer polypeptide" refers to a polypeptide in which a region of one polypeptide is fused to the N-terminus or C-terminus of a region of another (heterologous) polypeptide.

[0077] Increased specific performance: The term "increased specific performance" in the present invention means an improved conversion of cellulose material to product compared to the same degree of conversion performed on the parent. The increased specific performance per unit of protein (e.g., mg protein or micromolar protein) is determined. The increased specific performance of the variant relative to the parent can be evaluated, for example, under one or more (e.g., several) conditions of pH, temperature, and substrate concentration. In one aspect, the product is glucose. In another aspect, the product is cellobiose. In yet another aspect, the product is glucose + cellobiose.

[0078] One condition is pH. For example, the pH can be any pH in the range of 3 to 7, such as 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, or 7.0 (or between). Any suitable buffer solution can be used to achieve the desired pH.

[0079] On the other hand, the condition is temperature. For example, the temperature can be any temperature in the range of 25°C to 90°C, such as 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, or 90°C (or between).

[0080] On the other hand, the condition is the substrate concentration. Any cellulose material as defined herein can be used as the substrate. On one hand, the substrate concentration is measured as dry solids content. The dry solids content is preferably in the range of about 1 wt% to about 50 wt%, for example, about 5 wt% to about 45 wt%, about 10 wt% to about 40 wt%, or about 20 wt% to about 30 wt%. On the other hand, the substrate concentration is measured as insoluble dextran content. The insoluble dextran content is preferably in the range of about 2.5 wt% to about 25 wt%, for example, about 5 wt% to about 20 wt%, or about 10 wt% to about 15 wt%.

[0081] On the other hand, combinations of two or more (e.g., several) of the above conditions are used to determine the increased specific performance of the variant relative to the parent, such as at any temperature in the range of 25°C to 90°C, for example 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C or 90°C (or between thereafter), at pH in the range of 3 to 7, for example 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5 or 7.0 (or between thereafter).

[0082] The increased specific performance of the variant relative to the parent can be determined using any enzyme assay known in the art for cellobiases, as described herein. Alternatively, the increased specific performance of the variant relative to the parent can be determined using the assays described in Examples 9 and 12.

[0083] On the other hand, the specific performance of the variant is at least 1.01 times higher than that of the parent, for example, at least 1.02 times, at least 1.03 times, at least 1.04 times, at least 1.05 times, at least 1.06 times, at least 1.07 times, at least 1.08 times, at least 1.09 times, at least 1.1 times, at least 1.2 times, at least 1.3 times, at least 1.4 times, at least 1.5 times, at least 1.6 times, at least 1.7 times, at least 1.8 times, at least 1.9 times, at least 2 times, at least 2.1 times, at least 2.2 times, at least 2.3 times, at least 2.4 times, at least 2.5 times, at least 5 times, at least 10 times, at least 15 times, at least 20 times, at least 25 times, and at least 50 times.

[0084] Separate: The term “separate” means a substance in a form or environment not naturally occurring. Non-limiting examples of separated substances include (1) any substance not naturally occurring, (2) any substance including, but not limited to, any enzyme, variant, nucleic acid, protein, peptide, or cofactor, which is at least partially removed from one or more of the naturally occurring components associated with it; (3) any substance artificially modified relative to a naturally found substance; or (4) any substance modified by increasing the amount of the substance relative to other components naturally associated with it (e.g., recombinant production in a host cell; multiple copies of the gene encoding the substance; and the use of a promoter stronger than the promoter naturally associated with the gene encoding the substance). Any of the carbohydrate-binding module variants, cellobiase variants, or hybrid polypeptides described herein may be in a separated form.

[0085] Low stringency conditions: The term "low stringency conditions" refers to pre-hybridization and hybridization for probes of at least 100 nucleotides in length, following a standard DNA blotting procedure at 42°C in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 25% formamide for 12 to 24 hours. Vector material is finally washed three times at 50°C for 15 minutes each time with 0.2X SSC and 0.2% SDS.

[0086] Mature polypeptide: The term "mature polypeptide" refers to a polypeptide in its final form after translation and any post-translational modifications such as N-terminal processing, C-terminal truncation, glycosylation, phosphorylation, etc. On the one hand, based on the predicted amino acids 1 to 17 of SEQ ID NO:2, amino acids 1 to 18 of SEQ ID NO:6, amino acids 1 to 18 of SEQ ID NO:10, amino acids 1 to 25 of SEQ ID NO:14, amino acids 1 to 26 of SEQ ID NO:18, amino acids 1 to 17 of SEQ ID NO:22, amino acids 1 to 17 of SEQ ID NO:26, amino acids 1 to 18 of SEQ ID NO:61, amino acids 1 to 18 of SEQ ID NO:63, amino acids 1 to 18 of SEQ ID NO:73, amino acids 1 to 26 of SEQ ID NO:78, amino acids 1 to 26 of SEQ ID NO:90, amino acids 1 to 26 of SEQ ID NO:92, and amino acids 1 to 18 of SEQ ID NO:94, which are respectively the signal P (SignalP) program of the signal peptide (Nielsen et al., 1997, Protein Engineering 10:1-6), the mature polypeptide is SEQ ID NO:22. Amino acids 18 to 514 of SEQ ID NO:2, amino acids 19 to 525 of SEQ ID NO:6, amino acids 19 to 530 of SEQ ID NO:10, amino acids 26 to 537 of SEQ ID NO:14, amino acids 27 to 532 of SEQ ID NO:18, amino acids 18 to 526 of SEQ ID NO:22, amino acids 18 to 525 of SEQ ID NO:26, amino acids 19 to 519 of SEQ ID NO:61, amino acids 19 to 519 of SEQ ID NO:63, amino acids 19 to 519 of SEQ ID NO:73, amino acids 27 to 532 of SEQ ID NO:78, amino acids 27 to 532 of SEQ ID NO:90, amino acids 27 to 532 of SEQ ID NO:92, and amino acids 19 to 521 of SEQ ID NO:94. It is known in the art that host cells can produce mixtures of two or more different mature polypeptides (i.e., with different C-terminal and / or N-terminal amino acids) expressed from the same polynucleotide.

[0087] Mature polypeptide coding sequence: The term "mature polypeptide coding sequence" refers to a polynucleotide that encodes a mature polypeptide with cellobiose hydrolase activity. On one hand, based on the predicted signal P (SignalP) program (Nielsen et al., ibid.) encoding signal peptides by nucleotides 1 to 51 of SEQ ID NO:1, nucleotides 1 to 54 of SEQ ID NO:5, nucleotides 1 to 54 of SEQ ID NO:9, nucleotides 1 to 75 of SEQ ID NO:13, nucleotides 1 to 78 of SEQ ID NO:17, nucleotides 1 to 51 of SEQ ID NO:21 and nucleotides 1 to 51 of SEQ ID NO:25, the mature polypeptide encoding sequence is nucleotides 52 to 1542 of SEQ ID NO:1, nucleotides 55 to 1635 of SEQ ID NO:5, nucleotides 55 to 1590 of SEQ ID NO:9, nucleotides 76 to 1614 of SEQ ID NO:13, nucleotides 79 to 1596 of SEQ ID NO:17, nucleotides 52 to 1578 of SEQ ID NO:21 and nucleotides 52 to 1575 of SEQ ID NO:25, or their genomic DNA or cDNA sequences.

[0088] Medium-tough conditions: The term "medium-tough conditions" refers to pre-hybridization and hybridization at 42°C for 12 to 24 hours in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 35% formamide, following a standard DNA blotting procedure. Vector material is finally washed three times at 55°C for 15 minutes each time with 0.2X SSC and 0.2% SDS.

[0089] Medium-high stringent conditions: The term "medium-high stringent conditions" refers to pre-hybridization and hybridization for probes of at least 100 nucleotides in length, following standard DNA blotting procedures at 42°C in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 35% formamide for 12 to 24 hours. Vector material is finally washed three times at 60°C for 15 minutes each time with 0.2X SSC and 0.2% SDS.

[0090] Mutant: The term “mutant” refers to a polynucleotide that encodes a variant.

[0091] Nucleic acid constructs: The term “nucleic acid construct” refers to a single-stranded or double-stranded nucleic acid molecule that is isolated from a naturally occurring gene, or modified in a way that does not normally exist in nature to contain segments of nucleic acid, or is synthesized and includes one or more control sequences.

[0092] Operable ligation: The term “operable ligation” refers to a construction in which a control sequence is positioned relative to the coding sequence of a polynucleotide so that the control sequence directs the expression of the coding sequence.

[0093] Parental cellobiose hydrolase: The term "parental cellobiose hydrolase" refers to a cellobiose hydrolase that has been modified to produce the enzyme variant of the present invention. The parent can be a naturally occurring (wild-type) polypeptide or a variant or fragment thereof. The parental cellobiose hydrolase may include a carbohydrate-binding module.

[0094] Carbohydrate-binding module: The term "carbohydrate-binding module" refers to a carbohydrate-binding module that is modified to produce a variant of the carbohydrate-binding module of the present invention. The parent can be a naturally occurring (wild-type) polypeptide or a variant or fragment thereof.

[0095] Pretreated cellulose or hemicellulose material: The term “pretreated cellulose or hemicellulose material” means cellulose or hemicellulose material obtained from biomass by heat treatment and dilute sulfuric acid treatment, alkali pretreatment, neutral pretreatment, or any pretreatment known in the art.

[0096] Pretreated corn stalks: The term “pretreated corn stalks” or “PCS” means cellulose material obtained from corn stalks by heat and dilute sulfuric acid treatment, alkali pretreatment, neutral pretreatment, or any pretreatment known in the art.

[0097] Sequence consistency: The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter "sequence consistency".

[0098] For the purposes of this invention, the Niedleman-Wunsch algorithm (Needleman and Wunsch, 1970, J.Mol.Biol. 48:443-453) implemented in the Niedle program of the EMBOSS package (EMBOSS: European Open Software Suite for Molecular Biology, Rice et al., 2000, Trends Genet. 16:276-277) (preferably version 5.0.0 or later) is used to determine sequence consistency between two amino acid sequences. The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The Niedle output marked as “longest consistency” (obtained using the -nobrief option) is used as the percentage consistency and is calculated as follows:

[0099] (Consistent residues x 100) / (Alignment length - Total number of vacancies in the alignment)

[0100] For the purposes of this invention, the Niedle-Onsch algorithm (Niedleman and Onsch, 1970, ibid.) implemented in the Niedle program, such as in the EMBOSS package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, ibid.) (preferably version 5.0.0 or later), is used to determine sequence consistency between two deoxyribonucleotide sequences. The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. The Niedle output labeled "Longest Consistency" (obtained using the -nobrief option) is used as the percentage consistency and is calculated as follows:

[0101] (Consistent deoxyribonucleotides x 100) / (Alignment length - total number of vacancies in the alignment)

[0102] Subsequence: The term "subsequence" refers to a polynucleotide in which one or more (e.g., several) nucleotides are deleted from the 5' and / or 3' end of a mature polypeptide coding sequence; wherein the subsequence encodes a fragment having cellobiase activity. In one aspect, a subsequence comprises at least 1275 nucleotides, for example, at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:1. In another aspect, a subsequence comprises at least 1275 nucleotides, for example, at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:5. In another aspect, a subsequence comprises at least 1275 nucleotides, for example, at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:9. In another aspect, a subsequence comprises at least 1275 nucleotides, for example, at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:13. In another aspect, a subsequence comprises at least 1275 nucleotides, such as at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:17. In another aspect, a subsequence comprises at least 1275 nucleotides, such as at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:21. In another aspect, a subsequence comprises at least 1275 nucleotides, such as at least 1350 nucleotides or at least 1425 nucleotides, of the mature polypeptide coding sequence of SEQ ID NO:25.

[0103] Variants: The term "variant" refers to a polypeptide having altered (i.e., substituted, inserted, and / or deleted) cellobiase activity or carbohydrate-binding modules at one or more (e.g., several) positions. Substitution means that an amino acid occupying a position is replaced by a different amino acid; deletion means that an amino acid occupying a position is removed; and insertion means that an amino acid is added adjacent to or immediately next to the amino acid occupying a position.

[0104] The cellobiose hydrolase variants of the present invention have at least 20%, for example at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 100% cellobiose hydrolase activity of the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, or SEQ ID NO:26. The carbohydrate-binding module variants of the present invention have at least 20%, for example at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 100% carbohydrate-binding activity of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28.

[0105] Very High Tough Conditions: The term "very high tough conditions" refers to pre-hybridization and hybridization for probes of at least 100 nucleotides in length, following standard DNA blotting procedures at 42°C in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 50% formamide for 12 to 24 hours. Vector material is finally washed three times at 70°C for 15 minutes each time with 2X SSC and 0.2% SDS.

[0106] Very low stringency conditions: The term "very low stringency conditions" refers to pre-hybridization and hybridization for probes of at least 100 nucleotides in length, following standard DNA blotting procedures, at 42°C in 5X SSPE, 0.3% SDS, 200 μg / ml cleaved and denatured salmon sperm DNA, and 25% formamide for 12 to 24 hours. Vector material is finally washed three times at 45°C for 15 minutes each time with 0.2X SSC and 0.2% SDS.

[0107] Wild-type cellobiase: The term "wild-type" cellobiase refers to a cellobiase expressed by naturally occurring microorganisms (such as bacteria, yeast, or filamentous fungi found in nature).

[0108] Materials containing xylan: The term "materials containing xylan" refers to any material comprising plant cell wall polysaccharides containing a backbone of β-(1-4)-linked xylose residues. Terrestrial plant xylans are heteropolymers having a β-(1-4)-D-xylanose backbone branched by short carbohydrate chains. They include D-glucuronic acid or its 4-O-methyl ether, L-arabinose, and / or various oligosaccharides composed of D-xylose, L-arabinose, D- or L-galactose, and D-glucose. Xylan-type polysaccharides can be classified into homoxylans and heteroxylans, including glucuronide xylan, (arabinose)glucuronide xylan, (glucuronide)arabinosylxylan, arabinosylxylan, and complex heteroxylans. See, for example, Ebringerova et al., 2005, Adv. Polym. Sci. 186:1-67.

[0109] In the method of the present invention, any material containing xylan can be used. In a preferred aspect, the material containing xylan is lignocellulose.

[0110] Xylan degradation activity or xylan decomposition activity: The term “xylan degradation activity” or “xylan decomposition activity” refers to the biological activity of hydrolyzing materials containing xylan. Two basic methods for measuring xylan decomposition activity include: (1) measuring total xylan decomposition activity, and (2) measuring individual xylan decomposition activities (e.g., endoxylanase, β-xylosidase, arabinofuranylase, α-glucuronylase, acetylxylan esterase, ferulic acid esterase, and α-glucuronylase). Recent advances in the determination of xylanases have been summarized in several publications, including Biely and Puchard, 2006, Journal of the Science of Food and Agriculture 86(11):1636-1647; Spanikova and Biely, 2006, FEBS Letters 580(19):4597-4601; and Herrimann et al., 1997, Biochemical Journal 321:375-381.

[0111] Total xylan degradation activity can be measured by identifying reducing sugars formed from different types of xylans, including, for example, oat xylan, beech wood xylan, and larch wood xylan, or by spectrophotometric determination of stained xylan fragments released from different covalently stained xylans. A common assay for total xylan degradation activity is based on the production of reducing sugars from polymerized 4-O-methylglucuronic acid xylan, as described in Bailey et al., 1992, Interlaboratory testing of methods for assay of xylanase activity, Journal of Biotechnology 23(3):257-270. Xylanase activity can also be measured at 37°C at 0.01%. X-100 and 200 mM sodium phosphate (pH 6) were used with 0.2% AZCL-arabinosylxylan as a substrate. One unit of xylanase activity was defined as the production of 1.0 μmol of azurin per minute from 0.2% AZCL-arabinosylxylan as a substrate in 200 mM sodium phosphate (pH 6) at 37 °C and pH 6.

[0112] Xylan degradation activity can be determined by measuring the increase in hydrolysis of birch xylan (Sigma Chemical Co., Inc., St. Louis, MO, USA) caused by one or more xylan-degrading enzymes under the following typical conditions: 1 ml reaction, 5 mg / ml substrate (total solids), 5 mg xylan-degrading protein / g substrate, 50 mM sodium acetate (pH 5), 50 °C, 24 h, as described by Lever, 1972, Analytical Biochemistry 47:273-279, using p-hydroxybenzoic acid hydrazide (PHBAH) for sugar analysis.

[0113] Xylanase: The term "xylanase" refers to 1,4-β-D-xylan-xylohydrolase (EC 3.2.1.8), which catalyzes the internal hydrolysis of the 1,4-β-D-xylosidic bond in xylan. It can be synthesized at 37°C at a concentration of 0.01%. Xylanase activity was determined using 0.2% AZCL-arabinosylxylan as a substrate in X-100 and 200 mM sodium phosphate (pH 6). One unit of xylanase activity was defined as the production of 1.0 μmol of azurin per minute from 0.2% AZCL-arabinosylxylan as a substrate in 200 mM sodium phosphate (pH 6) at 37 °C and pH 6.

[0114] The reference to "about" a value or parameter here includes the aspect referring to that value or parameter itself. For example, a description of "about X" includes the aspect "X".

[0115] As used herein and in the appended claims, the singular forms “a”, “or”, and “the” include plural indicators unless the context clearly indicates otherwise. It should be understood that these aspects of the invention described herein include “consisting of aspects” and / or “substantially composed of aspects”.

[0116] Unless otherwise defined or clearly indicated by the context, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Invention Details

[0118] This invention relates to carbohydrate-binding module variants comprising substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31, and 32 of the carbohydrate-binding module corresponding to SEQ ID NO:4, wherein these variants have carbohydrate-binding activity. In one aspect, cellulases include carbohydrate-binding module variants (e.g., hybrid polypeptides) of the present invention.

[0119] The present invention also relates to isolated cellobiose hydrolase variants comprising substitutions at one or more (e.g., several) positions corresponding to positions 483, 491, 509 and 510 of SEQ ID NO:2, wherein these variants have cellobiose hydrolase activity.

[0120] Variant Naming Rules

[0121] For the purposes of this invention, the polypeptide sequence disclosed in SEQ ID NO:2 or the carbohydrate binding module (CBM) disclosed in SEQ ID NO:4 will be used to determine the corresponding amino acid residues in another cellobiase or CBM, respectively. The amino acid sequence of the other cellobiase or CBM will be compared with SEQ ID NO:2 or SEQ ID NO:4, respectively, and based on the comparison, the Niederman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48:443-453) implemented in the Nieder program of the EMBOSS package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16:276-277) (preferably version 3.0.0 or later) will be used to determine the amino acid position number corresponding to any amino acid residue in the polypeptide disclosed in SEQ ID NO:2 or the CBM disclosed in SEQ ID NO:4. The parameters used are an open space penalty of 10, an extended space penalty of 0.5, and an EBLOSUM62 (the EMBOSS version of BLOSUM62) replacement matrix.

[0122] The identification of the corresponding amino acid residues in another cellobiase or CBM can be determined by comparing multiple polypeptide sequences using several computer programs with their corresponding default parameters. These computer programs include, but are not limited to, MUSCLE (multiple sequence comparisons by logarithmic expectation; version 3.5 or later; Edgar, 2004, Nucleic Acids Research 32:1792-1797); MAFFT (version 6.857 or later; Katoh and Kuma, 2002, Nucleic Acids Research 30:3059-3066; Kato et al., 2005, Nucleic Acids Research 33:511-518; Kato and Toh, 2007, Bioinformatics 23:372-374; Kato et al., 2009, Methods in Molecular Biology). Biology 537:39-64; Kato and Asato, 2010, Bioinformatics 26:1899-1900; and EMBOSS EMMA using ClustalW (version 1.83 or later; Thompson et al., 1994, Nucleic Acid Research 22:4673-4680).

[0123] When other enzymes deviate from the polypeptide of SEQ ID NO:2 or other CBMs deviate from SEQ ID NO:4, making conventional sequence-based comparison methods unable to detect their relationship (Lindahl and Elofsson, 2000, Journal of Molecular Biology 295:613-615), other pairwise sequence comparison algorithms can be applied. Greater sensitivity in sequence-based searches can be achieved using search programs that utilize probabilistic representations (profiles) of polypeptide families to search a database. For example, the PSI-BLAST program generates multiple profiles through an iterative database search process and is capable of detecting distant homologs (Atschul et al., 1997, Nucleic Acids Res. 25:3389-3402). Even greater sensitivity can be achieved if the polypeptide family or superfamily has one or more representatives in a protein structure database. Procedures such as GenTHREADER (Jones, 1999, J.Mol.Biol. 287:797-815; McGuffin and Jones, 2003, Bioinformatics 19:874-881) utilize information from various sources (PSI-BLAST, secondary structure prediction, structural alignment spectra, and solvation potential) as input to neural networks that predict the structural folding of query sequences. Similarly, the method of Gough et al., 2000, J.Mol.Biol. 313:903-919 can be used to align sequences of unknown structures with superfamily models existing in the SCOP database. These alignments can then be used to generate homology models of peptides, and the accuracy of such models can be evaluated using various tools developed for this purpose.

[0124] For proteins with known structures, several tools and resources are available for retrieving and generating structure alignments. For example, the SCOP superfamily of proteins has already been structurally aligned, and those alignments are accessible and downloadable. Various algorithms, such as distance alignment matrices (Holm and Sander, 1998, Proteins 33:88-96) or combined extensions (Shindyalov and Bourne, 1998, Protein Engineering 11:739-747), can be used to align two or more protein structures, and implementations of these algorithms can also be used to query structure databases with structures of interest to discover possible structural homologs (e.g., Holm and Park, 2000, Bioinformatics 16:566-567).

[0125] In the description of variations of the invention, the following nomenclature is used for ease of reference. The accepted IUPAC single-letter and three-letter amino acid abbreviations are adopted.

[0126] replace For amino acid substitutions, the following nomenclature is used: initial amino acid, position, substituted amino acid. Therefore, the substitution of threonine at position 226 with alanine is represented as "Thr226Ala" or "T226A". Multiple mutations are separated by plus signs ("+"), for example, "Gly205Arg+Ser411Phe" or "G205R+S411F" represent the substitution of glycine (G) with arginine (R) at positions 205 and 411, respectively, and the substitution of serine (S) with phenylalanine (F).

[0127] Missing For amino acid deletions, the following nomenclature is used: initial amino acid, position, *. Therefore, a glycine deletion at position 195 is represented as "Gly195". * "or "G195 * Multiple missing characters are separated by plus signs ("+"), for example, "Gly195*+Ser411*" or "G195*+S411*".

[0128] insertFor amino acid insertions, the following nomenclature is used: initial amino acid, position, initial amino acid, inserted amino acid. Therefore, the insertion of lysine after glycine at position 195 is represented as "Gly195GlyLys" or "G195GK". Insertions of multiple amino acids are represented as [original amino acid, position, original amino acid, inserted amino acid #1, inserted amino acid #2, etc.]. For example, the insertion of lysine and alanine after glycine at position 195 is represented as "Gly195GlyLysAla" or "G195GKA".

[0129] In such cases, the inserted amino acid residues are numbered by adding lowercase letters to the position numbers of the amino acid residues preceding them. In the example above, the sequence would therefore be:

[0130] <![CDATA[ Parent: ]]> <![CDATA[ Variants: ]]> 195 195 195a 195b G GKA

[0131] Multiple changes Variants with multiple alterations are separated by a plus sign ("+"), such as "Arg170Tyr+Gly195Glu" or "R170Y+G195E", which represent that arginine and glycine at positions 170 and 195 are replaced by tyrosine and glutamic acid, respectively.

[0132] Different changes When different variations can be introduced at a single position, these variations are separated by a comma, for example, "Arg170Tyr,Glu" means that arginine at position 170 is replaced by either tyrosine or glutamic acid. Therefore, "Tyr167Gly,Ala+Arg170Gly,Ala" has the following variants: "Tyr167Gly+Arg170Gly", "Tyr167Gly+Arg170Ala", "Tyr167Ala+Arg170Gly", and "Tyr167Ala+Arg170Ala".

[0133] Carbohydrate-binding module variants

[0134] The present invention relates to a variant of the parental carbohydrate-binding module comprising substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31 and 32 of the carbohydrate-binding module of SEQ ID NO:4, wherein the variant has carbohydrate-binding activity.

[0135] On one hand, the carbohydrate-binding module variant has at least 60%, for example at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100%, sequence identity with the parental carbohydrate-binding module.

[0136] On the other hand, the carbohydrate-binding module variant has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, such as at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100% sequence identity with the polypeptide of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28.

[0137] In one aspect, the number of substitutions in the carbohydrate-binding module variant of the present invention is 1 to 4, such as 1, 2, 3, or 4 substitutions.

[0138] In one aspect, this carbohydrate-binding module variant includes or is composed of a substitution at position 5 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 5 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as being replaced by Tyr, Phe, or Trp. In another embodiment, the amino acid at position 5 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 5 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y5W of SEQ ID NO:4).

[0139] On the other hand, this carbohydrate-binding module variant includes or is composed of substitutions at position 13 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 13 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 13 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 13 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y13W of SEQ ID NO:4).

[0140] On the other hand, this carbohydrate-binding module variant includes or is composed of substitutions at position 31 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 31 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 31 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 31 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y31W of SEQ ID NO:4).

[0141] On the other hand, this carbohydrate-binding module variant includes or is composed of substitutions at position 32 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 32 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as being replaced by Tyr, Phe, or Trp. In another embodiment, the amino acid at position 32 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 32 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y32W of SEQ ID NO:4).

[0142] In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5, 13, 31, and 32 corresponding to SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 13 (e.g., substitution by Trp at positions 5 and 13, such as Y5W and / or Y13W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 31 (e.g., substitution by Trp at positions 5 and 31, such as Y5W and / or Y31W). In yet another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 32 (e.g., substitution by Trp at positions 5 and 32, such as Y5W and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13 and 31 (e.g., substitution by Trp at positions corresponding to positions 13 and 31, such as Y13W and / or Y31W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13 and 32 (e.g., substitution by Trp at positions corresponding to positions 13 and 32, such as Y13W and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 31 and 32 (e.g., substitution by Trp at positions corresponding to positions 31 and 32, such as Y31W and / or Y32W).

[0143] In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at three positions corresponding to positions 5, 13, 31, and 32 of SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 5, 13, and 31 (e.g., substitution by Trp at positions corresponding to positions 5, 13, and 31, such as Y5W, Y13W, and / or Y31W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 5, 13, and 32 (e.g., substitution by Trp at positions corresponding to positions 5, 13, and 32, such as Y5W, Y13W, and / or Y32W). In yet another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 5, 31, and 32 (e.g., substitution by Trp at positions corresponding to positions 5, 31, and 32, such as Y5W, Y31W, and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13, 31, and 32 (e.g., substitution by Trp at positions corresponding to positions 13, 31, and 32, such as Y13W, Y31W, and / or Y32W).

[0144] On the other hand, the carbohydrate-binding module variant includes substitutions or is composed of Trp at all four positions corresponding to positions 5, 13, 31, and 32 of SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes Trp substitutions or is composed of Trp at one or more positions corresponding to positions 5, 13, 31, and 32, such as Y5W, Y13W, Y31W, and / or Y32W.

[0145] This carbohydrate-binding module variant may further include substitutions, deletions, and / or insertions at one or more (e.g., several) other locations, such as one or more (e.g., several) substitutions at locations corresponding to those disclosed in WO 2012 / 135719, which is incorporated herein by reference. For example, in one aspect, this carbohydrate-binding module variant further includes substitutions at one or more (e.g., several) locations corresponding to positions 4, 6, and 29 of SEQ ID NO:4. In another aspect, this carbohydrate-binding module variant further includes substitutions at two locations corresponding to any one of positions 4, 6, and 29. In yet another aspect, this carbohydrate-binding module variant further includes substitutions at each location corresponding to positions 4, 6, and 29.

[0146] In another aspect, this carbohydrate-binding module variant includes or is composed of a substitution at the position corresponding to position 4. In another aspect, the amino acid at the position corresponding to position 4 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Glu, Leu, Lys, Phe, or Trp. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted H4L of SEQ ID NO:4. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted H4K of SEQ ID NO:4. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted H4E of SEQ ID NO:4. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted H4F of SEQ ID NO:4. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted H4W of SEQ ID NO:4.

[0147] In another aspect, this carbohydrate-binding module variant includes or is composed of a substitution at position 6. In another aspect, the amino acid at position 6 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted G6A of SEQ ID NO:4.

[0148] In another aspect, this carbohydrate-binding module variant includes or is composed of a substitution at the position corresponding to position 29. In another aspect, the amino acid at the position corresponding to position 29 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Asp. In another aspect, this carbohydrate-binding module variant includes or is composed of the substituted N29D of SEQ ID NO:4.

[0149] On the other hand, this carbohydrate-binding module variant further includes substitutions or components thereof at positions corresponding to positions 4 and 6, as described above.

[0150] On the other hand, this carbohydrate-binding module variant further includes substitutions or components thereof at positions corresponding to positions 4 and 29, as described above.

[0151] On the other hand, this carbohydrate-binding module variant further includes substitutions or components thereof at positions corresponding to positions 6 and 29, as described above.

[0152] On the other hand, the carbohydrate-binding module variant further includes substitutions or components thereof at positions corresponding to positions 4, 6, and 29, as described above.

[0153] On the other hand, the carbohydrate binding module variant further includes or consists of one or more (e.g., several) substitutions selected from the group consisting of H4L,K,E,F,W,G6A, and N29D; or one or more (e.g., several) substitutions selected from the group consisting of H4L,K,E,F,W,G6A, and N29D corresponding to SEQ ID NO:4 in other cellulose binding modules described herein.

[0154] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4L+G6A of SEQ ID NO:4.

[0155] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4K+G6A of SEQ ID NO:4.

[0156] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4E+G6A of SEQ ID NO:4.

[0157] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4F+G6A of SEQ ID NO:4.

[0158] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4W+G6A of SEQ ID NO:4.

[0159] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4L+N29D of SEQ ID NO:4.

[0160] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4K+N29D of SEQ ID NO:4.

[0161] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4E+N29D of SEQ ID NO:4.

[0162] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4F+N29D of SEQ ID NO:4.

[0163] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4W+N29D of SEQ ID NO:4.

[0164] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted G6A+N29D of SEQ ID NO:4.

[0165] On the other hand, the carbohydrate-binding module variant includes or is composed of SEQ ID NO:4 substituted H4L+G6A+N29D.

[0166] On the other hand, the carbohydrate-binding module variant includes or is composed of substituted H4K+G6A+N29D of SEQ ID NO:4.

[0167] On the other hand, the carbohydrate-binding module variant includes or is composed of substituted H4E+G6A+N29D of SEQ ID NO:4.

[0168] On the other hand, this carbohydrate-binding module variant includes or is composed of substituted H4F+G6A+N29D of SEQ ID NO:4.

[0169] On the other hand, the carbohydrate-binding module variant includes or is composed of SEQ ID NO:4 substituted H4W+G6A+N29D.

[0170] Amino acid changes can be of a minor nature, i.e., conserved amino acid substitutions or insertions that do not significantly affect protein folding and / or activity; small deletions typically of 1 to 30 amino acids; small amino-terminal or carboxyl-terminal elongations, such as amino-terminal methionine residues; small linker peptides of up to 20-25 residues; or small elongated portions that facilitate purification by altering net charge or another function, such as multihistidine bundles, antigenic epitopes, or binding modules.

[0171] Examples of conserved substitutions are found in the following group: basic amino acids (arginine, lysine, and histidine), acidic amino acids (glutamic acid and aspartic acid), polar amino acids (glutamine and asparagine), hydrophobic amino acids (leucine, isoleucine, and valine), aromatic amino acids (phenylalanine, tryptophan, and tyrosine), and small amino acids (glycine, alanine, serine, threonine, and methionine). Amino acid substitutions that do not typically alter specific activity are known in the art and are described, for example, by H. Neurath and RL Hill, 1979, in *The Proteins*, Academic Press, New York. Common substitutes are Ala / Ser, Val / Ile, Asp / Glu, Thr / Ser, Ala / Gly, Ala / Thr, Ser / Asn, Ala / Val, Ser / Gly, Tyr / Phe, Ala / Pro, Lys / Arg, Asp / Asn, Leu / Ile, Leu / Val, Ala / Glu, and Asp / Gly.

[0172] Alternatively, amino acid changes are a property that alters the physicochemical properties of a peptide. For example, amino acid changes can improve the peptide's thermal stability, change its substrate specificity, or alter its optimal pH.

[0173] Essential amino acids in peptides can be identified using methods known in the art, such as site-directed mutagenesis or alanine scanning mutagenesis (Cunningham and Wells, 1989, Science 244:1081-1085). In the latter technique, a single alanine mutation is introduced at each residue in the molecule, and the cellobiase activity of the resulting mutant molecule is tested to identify amino acid residues essential to the molecule's activity. See also Hilton et al., 1996, Journal of Biochemistry 271:4699-4708. The active site of an enzyme or other biological interaction can also be determined by physical analysis of the structure, such as by techniques including nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, along with mutation of the amino acid at the putative contract site. See, for example, de Vos et al., 1992, Science 255:306-312; Smith et al., 1992, Journal of Molecular Biology 224:899-904; Wlodaver et al., 1992, FEBS Lett. 309:59-64. The identity of essential amino acids can also be inferred from comparisons with related peptides.

[0174] In some respects, these carbohydrate-binding module variants can consist of 28 to 36 (inclusive) amino acids, such as 28, 29, 30, 31, 32, 33, 34, 35 or 36 amino acids.

[0175] As described in more detail below, the present invention also relates to polypeptides having cellulolytic activity, comprising the carbohydrate-binding module variants as described above. In one aspect, the polypeptide is derived from a “wild-type” cellulase (such as a “wild-type” cellobiase) having a carbohydrate-binding module, wherein the carbohydrate-binding module comprises substitutions at one or more (e.g., several) positions corresponding to SEQ ID NO:4, 5, 13, 31, and 32. In one aspect, the carbohydrate-binding module variants of the present invention can be fused with polypeptides lacking a carbohydrate-binding module. In another aspect, the carbohydrate-binding module contained in the polypeptide can be replaced by the carbohydrate-binding module variants of the present invention. In another aspect, the polypeptide is a cellulase selected from the group consisting of: endoglucanase, cellobiase, and GH61 polypeptide. In one embodiment, the cellulase is an endoglucanase. In another embodiment, the cellulase is a cellobiase. In yet another embodiment, the cellulase is GH61 polypeptide.

[0176] In some respects, this carbohydrate-binding module variant has improved binding activity. In some embodiments, the carbohydrate-binding module variant does not have reduced binding activity compared to the parent. In some embodiments, the carbohydrate-binding module variant has a binding activity at least 1.01 times higher than the parent, for example, at least 1.02 times, at least 1.03 times, at least 1.04 times, at least 1.05 times, at least 1.06 times, at least 1.07 times, at least 1.08 times, at least 1.09 times, at least 1.1 times, at least 1.2 times, at least 1.3 times, at least 1.4 times, at least 1.5 times, at least 1.6 times, at least 1.7 times, at least 1.8 times, at least 1.9 times, at least 2 times, at least 2.1 times, at least 2.2 times, at least 2.3 times, at least 2.4 times, at least 2.5 times, at least 5 times, at least 10 times, at least 15 times, at least 20 times, at least 25 times, and at least 50 times higher than the parent.

[0177] Cellobiose hydrolase variant

[0178] The present invention also relates to variants of parental cellobiose hydrolases comprising a carbohydrate-binding module, wherein the carbohydrate-binding module includes substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31, and 32 of SEQ ID NO:4. For example, in one aspect, there is a variant of a parental cellobiose hydrolase that includes substitutions at one or more (e.g., several) positions corresponding to positions 483, 491, 509, or 510 of SEQ ID NO:2, wherein the variant has cellobiose hydrolase activity.

[0179] In one embodiment, the variant has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100%, sequence identity with the parent cellobiase.

[0180] In another embodiment, the variant has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, such as at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100% sequence identity with the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78.

[0181] In one aspect, the number of substitutions in the variants of the present invention is 1 to 3, such as 1, 2, 3, or 3 substitutions.

[0182] On the other hand, this variant includes or consists of a substitution at position 483 corresponding to SEQ ID NO:2. In one embodiment, the amino acid at position 483 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 483 corresponding to SEQ ID NO:2 is replaced by Trp. In yet another embodiment, the amino acid at position 483 corresponding to SEQ ID NO:2 is Tyr replaced by Trp (e.g., Y483W of SEQ ID NO:4).

[0183] On the other hand, this variant includes or consists of a substitution at position 491 corresponding to SEQ ID NO:2. In one embodiment, the amino acid at position 491 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 491 corresponding to SEQ ID NO:2 is replaced by Trp. In yet another embodiment, the amino acid at position 491 corresponding to SEQ ID NO:2 is Tyr replaced by Trp (e.g., Y491W of SEQ ID NO:4).

[0184] On the other hand, this variant includes or consists of a substitution at position 509 corresponding to SEQ ID NO:2. In one embodiment, the amino acid at position 509 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 509 corresponding to SEQ ID NO:2 is replaced by Trp. In yet another embodiment, the amino acid at position 509 corresponding to SEQ ID NO:2 is Tyr replaced by Trp (e.g., Y509W of SEQ ID NO:4).

[0185] On the other hand, this variant includes or consists of a substitution at position 510 corresponding to SEQ ID NO:2. In one embodiment, the amino acid at position 510 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 510 corresponding to SEQ ID NO:2 is replaced by Trp. In yet another embodiment, the amino acid at position 510 corresponding to SEQ ID NO:2 is Tyr replaced by Trp (e.g., Y510W of SEQ ID NO:4).

[0186] On the other hand, the variant includes substitutions or is composed of substitutions at positions 483, 491, 509, and 510 corresponding to SEQ ID NO:2, as described above. In one embodiment, the variant includes substitutions or is composed of substitutions at positions 483 and 491 (e.g., substitution by Trp at positions 483 and 491, such as Y483W and / or Y491W). In another embodiment, the variant includes substitutions or is composed of substitutions at positions 483 and 509 (e.g., substitution by Trp at positions 483 and 509, such as Y483W and / or Y509W). In yet another embodiment, the variant includes substitutions or is composed of substitutions at positions 483 and 510 (e.g., substitution by Trp at positions 483 and 510, such as Y483W and / or Y510W). In another embodiment, the variant includes substitutions or components thereof at positions corresponding to positions 491 and 509 (e.g., substitution by Trp at positions corresponding to positions 491 and 509, such as Y491W and / or Y509W). In another embodiment, the variant includes substitutions or components thereof at positions corresponding to positions 491 and 510 (e.g., substitution by Trp at positions corresponding to positions 491 and 510, such as Y491W and / or Y32W). In another embodiment, the variant includes substitutions or components thereof at positions corresponding to positions 509 and 510 (e.g., substitution by Trp at positions corresponding to positions 509 and 510, such as Y509W and / or Y510W).

[0187] On the other hand, the variant includes substitutions or is composed of substitutions at positions 483, 491, 509, and 32 corresponding to SEQ ID NO:2, as described above. In one embodiment, the variant includes substitutions or is composed of substitutions at positions 483, 491, and 509 (e.g., substitution by Trp at positions 483, 491, and 509, such as Y483W, Y491W, and / or Y509W). In another embodiment, the variant includes substitutions or is composed of substitutions at positions 483, 491, and 510 (e.g., substitution by Trp at positions 483, 491, and 510, such as Y5W, Y491W, and / or Y510W). In another embodiment, the variant includes substitutions or components thereof at positions corresponding to positions 483, 509, and 510 (e.g., substitution by Trp at positions corresponding to positions 483, 509, and 510, such as Y483W, Y509W, and / or Y510W). In another embodiment, the variant includes substitutions or components thereof at positions corresponding to positions 491, 509, and 510 (e.g., substitution by Trp at positions corresponding to positions 491, 509, and 510, such as Y491W, Y509W, and / or Y510W).

[0188] On the other hand, the variant includes substitutions or components thereof at four positions corresponding to positions 483, 491, 509, and 510 of SEQ ID NO:2, as described above. In one embodiment, the variant includes substitutions or components thereof for Trp at one or more positions corresponding to positions 483, 491, 509, and 510, such as Y483W, Y491W, Y509W, and / or Y510W.

[0189] In one respect, the variant includes or consists of the following: SEQ ID NO:90 or SEQ ID NO:92, or a mature polypeptide sequence thereof.

[0190] These cellobiase variants may further include substitutions, deletions, and / or insertions at one or more (e.g., several) other locations, such as changes at one or more (e.g., several) locations corresponding to those disclosed in PCT / US 2014 / 022068, WO 2011 / 050037, WO 2005 / 028636, WO 2005 / 001065, WO 2004 / 016760, and U.S. Patent No. 7,375,197, the entire contents of which are incorporated herein by reference.

[0191] For example, in another aspect, a variant includes a change at one or more positions corresponding to positions 214, 215, 216, and 217 of SEQ ID NO:2, wherein the change at one or more positions corresponding to positions 214, 215, and 217 is a substitution and the change at position 216 is a deletion. In another aspect, a variant includes a change at two positions corresponding to any of positions 214, 215, 216, and 217 of SEQ ID NO:2, wherein the change at one or more positions corresponding to positions 214, 215, and 217 is a substitution and the change at position 216 is a deletion. In yet another aspect, a variant includes a change at three positions corresponding to any of positions 214, 215, 216, and 217 of SEQ ID NO:2, wherein the change at one or more positions corresponding to positions 214, 215, and 217 is a substitution and the change at position 216 is a deletion. On the other hand, one variant includes a substitution at each of the positions corresponding to positions 214, 215 and 217 and a deletion at the position corresponding to position 216.

[0192] In another aspect, this variant includes or consists of a substitution at the position corresponding to position 214. In another aspect, the amino acid at the position corresponding to position 214 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, this variant includes or consists of the substituted N214A of SEQ ID NO:2.

[0193] In another aspect, this variant includes or consists of a substitution at the position corresponding to position 215. In another aspect, the amino acid at the position corresponding to position 215 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, this variant includes or consists of substituted N215A of SEQ ID NO:2.

[0194] On the other hand, this variant includes a deletion or is composed of a term corresponding to position 216. On the other hand, the amino acid at position 216 is Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably Ala. On the other hand, this variant includes a deletion A216* of SEQ ID NO:2 or is composed of a term thereof.

[0195] In another aspect, this variant includes or consists of substitutions at the position corresponding to position 217. In another aspect, the amino acid at the position corresponding to position 217 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala, Gly, or Trp. In another aspect, this variant includes or consists of the substituted N217A, G, W of SEQ ID NO:2.

[0196] On the other hand, the variant includes a change or composition thereof at the positions corresponding to positions 214 and 215, as those described above.

[0197] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214 and 216, as described above.

[0198] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214 and 217, as described above.

[0199] On the other hand, the variant includes changes or components thereof at the positions corresponding to positions 215 and 216, as described above.

[0200] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 215 and 217, as described above.

[0201] On the other hand, the variant includes changes or components thereof at the positions corresponding to positions 216 and 217, as described above.

[0202] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214, 215 and 216, as described above.

[0203] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214, 215 and 217, as described above.

[0204] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214, 216 and 217, as described above.

[0205] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 215, 216 and 217, as described above.

[0206] On the other hand, the variant includes changes or components thereof at positions corresponding to positions 214, 215, 216 and 217, as described above.

[0207] On the other hand, the variant includes one or more changes selected from or composed of the group consisting of: N214A, N215A, A216*, and N217A,G,W.

[0208] On the other hand, the variant includes changes to SEQ ID NO:2 N214A+N215A or combinations thereof.

[0209] On the other hand, this variant includes changes to SEQ ID NO:2, N214A+A216*, or is composed of them.

[0210] On the other hand, the variant includes changes to SEQ ID NO:2 N214A+N217A,G,W or combinations thereof.

[0211] On the other hand, this variant includes changes to SEQ ID NO:2, N215A+A216*, or is composed of them.

[0212] On the other hand, the variant includes changes to SEQ ID NO:2 N215A+N217A,G,W or combinations thereof.

[0213] On the other hand, the variant includes changes to SEQ ID NO:2 A216*+N217A,G,W or combinations thereof.

[0214] On the other hand, the variant includes changes to SEQ ID NO:2 N214A+N215A+A216* or is composed of therewith.

[0215] On the other hand, the variant includes changes to SEQ ID NO:2 N214A+N215A+N217A,G,W or combinations thereof.

[0216] On the other hand, the variant includes changes to SEQ ID NO:2 N214A+A216*+N217A,G,W or combinations thereof.

[0217] On the other hand, this variant includes changes to the mature polypeptide of SEQ ID NO:2, such as N215A+A216*+N217A,G,W, or composition thereof.

[0218] On the other hand, the variant includes changes to SEQ ID NO:2 such as N214A+N215A+A216*+N217A,G,W or combinations thereof.

[0219] Essential amino acids in the parent can be identified according to procedures known in the art, as described herein.

[0220] In some respects, these cellobiase variants can consist of 310 to 537 amino acids (including the beginning and end), such as 320 to 330, 330 to 340, 340 to 350, 350 to 360, 360 to 370, 370 to 380, 380 to 390 to 390, 490 to 400, 400 to 415, 415 to 425, 425 to 435, 435 to 445, 445 to 455, 455 to 465, 465 to 475, 475 to 485, 485 to 495, 495 to 505, 505 to 515, 515 to 525, or 525 to 537 amino acids.

[0221] Parental cellobiose hydrolase and carbohydrate binding module

[0222] The parental carbohydrate binding module can be (a) a carbohydrate binding module having at least 60% sequence identity with the carbohydrate binding modules of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24 or SEQ ID NO:28; (b) a carbohydrate binding module encoded by a polynucleotide that hybridizes with the carbohydrate binding module coding sequence of SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23 or SEQ ID NO:27 or its full-length complement under at least low stringency conditions; or (c) a carbohydrate binding module encoded by a polynucleotide having at least 60% sequence identity with the carbohydrate binding module coding sequence of SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23 or SEQ ID NO:27.

[0223] The parental cellobiose hydrolase may be (a) a polypeptide having at least 60% sequence identity with the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44 or SEQ ID NO:78; (b) a polypeptide encoded by a polynucleotide hybridized under at least low stringency conditions to: (i) the coding sequence of the mature polypeptide of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43 or SEQ ID NO:77, (ii) its genomic DNA or cDNA sequence, or (iii) a full-length complement of (i) or (ii); (c) a polypeptide encoded by a polynucleotide hybridized to the mature polypeptide of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:77, (ii) its genomic DNA or cDNA sequence, or (iii) a full-length complement of (i) or (ii); A mature polypeptide encoding sequence of SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43 or SEQ ID NO:77 having at least 60% sequence identity with a polynucleotide-encoded polypeptide; or a fragment of (d)(a), (b) or (c) having cellobiose hydrolase activity.

[0224] In one aspect, the parental carbohydrate-binding module has at least 60%, for example at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical to the carbohydrate-binding module of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28, and has carbohydrate-binding activity. In one embodiment, the amino acid sequence of the parental carbohydrate binding module differs from that of the carbohydrate binding modules of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24 or SEQ ID NO:28 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.

[0225] In another embodiment, the parental carbohydrate binding module comprises or consists of the amino acid sequence of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24 or SEQ ID NO:28.

[0226] In another embodiment, the parental carbohydrate binding module is a fragment of a mature polypeptide of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28, the fragment containing at least 28 amino acid residues, such as at least 30, at least 32, or at least 34 amino acid residues.

[0227] In another embodiment, the parental carbohydrate-binding module is an allelic variant of the carbohydrate-binding module of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24 or SEQ ID NO:28.

[0228] In another first aspect, the parental cellobiose hydrolase has at least 60%, for example at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identical to the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78, and has cellobiose hydrolase activity. In one embodiment, the amino acid sequence of the parent cellobiose hydrolase differs from the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

[0229] In another embodiment, the parental cellobiose hydrolase comprises or is composed of the amino acid sequence of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78. In another embodiment, the parental cellobiose hydrolase comprises or is composed of the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78. In another embodiment, the parental cellobiose hydrolase comprises amino acids 18 to 514 of SEQ ID NO:2, amino acids 19 to 525 of SEQ ID NO:6, amino acids 19 to 530 of SEQ ID NO:10, amino acids 26 to 537 of SEQ ID NO:14, amino acids 27 to 532 of SEQ ID NO:18, amino acids 18 to 526 of SEQ ID NO:22, or amino acids 18 to 525 of SEQ ID NO:26, or is composed of the like.

[0230] In another embodiment, the parental cellobiose hydrolase is a fragment of at least 85% of the amino acid residues, such as at least 90% or at least 95% of the amino acid residues, of a mature polypeptide containing the parental cellobiose hydrolase.

[0231] In another embodiment, the parental cellobiose hydrolase is an allelic variant of the mature polypeptide of SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44 or SEQ ID NO:78.

[0232] In one second aspect, the parental carbohydrate-binding module is encoded by a polynucleotide that hybridizes with the carbohydrate-binding module encoding sequence of SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27, or their full-length complement, under very low stringent, low stringent, medium stringent, medium-high stringent, high stringent, or very high stringent conditions (Sambrook et al., 1989, Molecular Cloning, A Laboratory Manual, 2nd Edition, Cold Spring Harbor, New York).

[0233] In another second aspect, the parental cellobiose hydrolase is encoded by a polynucleotide that hybridizes under very low stringent, low stringent, moderate stringent, moderate-high stringent, high stringent, or very high stringent conditions with (i) the mature polypeptide coding sequence of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77, (ii) its genomic DNA or cDNA sequence, or (iii) the full-length complement of (i) or (ii) (Sambrook et al., 1989, ibid.).

[0234] SEQ ID NO: 1, SEQ ID NO: 5, SEQ ID NO: 9, SEQ ID NO: 13, SEQ ID NO: 17, SEQ ID NO: 21, SEQ ID NO: 25, SEQ ID NO: 41, SEQ ID NO: 43, SEQ ID NO: 77, SEQ ID NO: 3, SEQ ID NO: 7, SEQ ID NO: 11, SEQ ID NO: 15, SEQ ID The polynucleotide of NO:19, SEQ ID NO:23, or SEQ ID NO:27, or a subsequence thereof, together with SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78, SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID Nucleic acid probes are designed using the polypeptide or fragment thereof from NO:28 to identify and clone parental DNA encoding strains from different genera or species, according to methods well known in the art. Specifically, such probes can be hybridized with the genomic DNA or cDNA of cells of interest according to standard DNA blotting procedures to identify and isolate the corresponding genes therein. These probes can be significantly shorter than the complete sequence, but should be at least 15 nucleotides long, for example, at least 25, at least 35, or at least 70 nucleotides. Preferably, the nucleic acid probe is at least 100 nucleotides long, for example, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or at least 900 nucleotides. Both DNA and RNA probes can be used. Typically, the probes are labeled (e.g., with...). 32 P, 3 H, 35 This invention covers probes containing biotin (or avidin) to detect corresponding genes.

[0235] Genomic DNA or cDNA libraries prepared from other strains of this type can be screened for DNA that hybridizes with the probes described above and encodes the parental DNA. Genomic DNA or other DNA from these other strains can be separated by agarose or polyacrylamide gel electrophoresis, or other separation techniques. DNA from the library or separated DNA can be transferred and immobilized on nitrocellulose or other suitable vector materials. Vector materials are used in DNA blotting to identify clones or DNA that hybridize with SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, SEQ ID NO:77, SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27, or their subsequences.

[0236] For the purposes of this invention, the hybridization indicator allows the polynucleotide to hybridize under very low to very high stringency conditions with labeled nucleic acid probes corresponding to: (i) SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77; (ii) the mature polypeptide coding sequence of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77; (iii) its genomic DNA or cDNA sequence; (iv) its full-length complement; or (v) its daughter sequence; or (i) SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77; (iii) its genomic DNA or cDNA sequence; (iv) its full-length complement; or (v) its daughter sequence; or (i) SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77. NO:23 or SEQ ID NO:27; (ii) its full-length complement; or (iii) its subsequence. Molecules hybridized to nucleic acid probes under these conditions can be detected using, for example, X-ray film or any other detection method known in the art.

[0237] In one embodiment, the nucleic acid probe is a mature polypeptide coding sequence of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77. In another embodiment, the nucleic acid probe is nucleotides 52 to 1542 of SEQ ID NO:1, nucleotides 55 to 1635 of SEQ ID NO:5, nucleotides 55 to 1590 of SEQ ID NO:9, nucleotides 76 to 1614 of SEQ ID NO:13, nucleotides 79 to 1596 of SEQ ID NO:17, nucleotides 52 to 1578 of SEQ ID NO:21, or nucleotides 52 to 1575 of SEQ ID NO:25. In another embodiment, the nucleic acid probe is a polypeptide encoding SEQ ID NO:2, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:18, SEQ ID NO:22, SEQ ID NO:26, SEQ ID NO:42, SEQ ID NO:44, or SEQ ID NO:78; its mature polypeptide; or a fragment thereof, which is a polynucleotide. In another embodiment, the nucleic acid probe is a genomic DNA or cDNA sequence of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77, or the thereof.

[0238] In another embodiment, the nucleic acid probe is SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27. In another embodiment, the nucleic acid probe is a polynucleotide encoding a carbohydrate-binding module or fragment thereof of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28.

[0239] For short probes ranging from approximately 15 to approximately 70 nucleotides in length, stringent conditions are defined as optimally following standard DNA blotting procedures, compared to the T calculated using calculations based on those of Bolton and McCarthy (1962, Proceedings of the National Academy of Sciences of the United States of America (Proc. Natl. Acad. Sci. USA) 48:1390). m Pre-hybridization and hybridization were performed at approximately 5°C to approximately 10°C in solutions of 0.9 M NaCl, 0.09 M Tris-HCl (pH 7.6), 6 mM EDTA, 0.5% NP-40, 1X Denhardt's solution, 1 mM sodium pyrophosphate, 1 mM sodium dihydrogen phosphate, 0.1 mM ATP, and 0.2 mg of yeast RNA per ml for 12 to 24 hours. Finally, the vector material was subjected to pre-hybridization and hybridization at a temperature lower than the calculated T0. m Wash once (15 minutes) in 6X SCC with 0.1% SDS at a temperature 5°C to 10°C, and then wash twice (15 minutes each) in 6X SSC.

[0240] In a third aspect, the parental carbohydrate-binding module is encoded by a polynucleotide having at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the carbohydrate-binding module encoding sequence of SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27, and the polynucleotide encodes a polypeptide having carbohydrate-binding activity. In one embodiment, the carbohydrate-binding module encoding sequence is SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27. In another embodiment, the parent is encoded by a polynucleotide comprising or consisting of SEQ ID NO:3, SEQ ID NO:7, SEQ ID NO:11, SEQ ID NO:15, SEQ ID NO:19, SEQ ID NO:23, or SEQ ID NO:27.

[0241] In another third aspect, the parental cellobiase is encoded by a polynucleotide having at least 60%, for example at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the mature polypeptide encoding sequence of SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:25, SEQ ID NO:41, SEQ ID NO:43, or SEQ ID NO:77, which encodes a polypeptide having cellobiase activity. On one hand, the mature polypeptide encoding sequence is nucleotides 52 to 1542 of SEQ ID NO:1, nucleotides 55 to 1635 of SEQ ID NO:5, nucleotides 55 to 1590 of SEQ ID NO:9, nucleotides 76 to 1614 of SEQ ID NO:13, nucleotides 79 to 1596 of SEQ ID NO:17, nucleotides 52 to 1578 of SEQ ID NO:21, or nucleotides 52 to 1575 of SEQ ID NO:25, or their genomic DNA or cDNA sequence. On the other hand, the parental cellobiose hydrolase is encoded by polynucleotides including SEQ ID NO:1, SEQ ID NO:5, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:17, SEQ ID NO:21, or SEQ ID NO:25, or their genomic DNA or cDNA sequence, or composed of thereof.

[0242] The parent can be obtained from any genus of microorganisms. For the purposes of this invention, the term "obtained from" as used herein in conjunction with a given source should mean that the parent encoded by the polynucleotide is produced by that source or by a strain in which a polynucleotide from that source has been inserted. In one aspect, the parent is extracellularly secreted.

[0243] The parent can be a bacterial cellobiose hydrolase or a carbohydrate-binding module. For example, the parent can be a Gram-positive bacterial polypeptide, such as Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, or Streptomyces. Cellobiose hydrolase (Cycyl) or a Gram-negative bacterial polypeptide, such as Campylobacter, Escherichia coli, Flavobacterium, Fusobacterium, Helicobacter, Ilyobacter, Neisseria, Pseudomonas, Salmonella, or Ureaplasma polypeptide.

[0244] On one hand, the parent is a polypeptide of Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus scoagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, or Bacillus thuringiensis.

[0245] On the other hand, the parent is a polypeptide of Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, or Streptococcus equi subsp. Zooepidemicus.

[0246] On the other hand, the parent does not produce Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, or Streptomyces lividans polypeptides.

[0247] The parent can be a fungal cellobiose hydrolase or a carbohydrate-binding module. For example, the parent can be a yeast cellobiose hydrolase or a carbohydrate-binding module, such as a Candida, Kluyveromyces, Pichia, Saccharomyces, Schizosaccharomyces, or Yarrowia polypeptide.For example, the parent could be a cellobiose hydrolase or carbohydrate-binding module from filamentous fungi, such as *Acremonium*, *Agaricus*, *Alternaria*, *Aspergillus*, *Aureobasidium*, *Botryospaeria*, *Ceriporiopsis*, *Chaetomium*, *Chrysosporium*, *Claviceps*, *Cochliobolus*, and *Coprinops*. The genera *is*, *Coptotermes*, *Corynascus*, *Cryphonectria*, *Cryptococcus*, *Diplodia*, *Exidia*, *Fennellia*, *Filibasidium*, *Fusarium*, *Gibberella*, *Holomastigotoides*, *Humicola*, *Irpex*, *Lentinula*, and *Micrococcus* (…) Leptospaeria, Magnaporthe, Melanocarpus, Meripius, Mucor, Myceliophthora, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Phanerochaete, Piromyces, Poitrasia, Pseudoplectania Polypeptides from the genera *Pseudotrichonympha*, *Rhizomucor*, *Schizophyllum*, *Scytalidium*, *Talaromyces*, *Thermoascus*, *Thielavia*, *Tolypocladium*, *Trichoderma*, *Trichophaea*, *Verticillium*, *Volvariella*, or *Xylaria*.

[0248] On the other hand, the parent is a polypeptide from Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces kluyveri, Saccharomyces norbensis, or Saccharomyces oviformis.

[0249] On the other hand, the parent species are *Acremonium cellulolyticus*, *Aspergillus aculeatus*, *Aspergillus awamori*, *Aspergillus foetidus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidulans*, *Aspergillus niger*, *Aspergillus oryzae*, *Chrysosporium inops*, *Chrysosporium keratinophilum*, *Chrysosporium lucknowense*, *Chrysosporium merdarium*, *Chrysosporium pannicola*, and *Chrysosporium quercetinum*. Queenslandicum, Chrysosporium tropicum, Chrysosporium zonatum, Corynascus thermophilus, Fusarium bactridioides, Fusarium cerealis, Fusarium crookwellense, Fusarium culmorum, Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium arcochroum Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, FusariumTrichothecioides, Fusarium venenatum, Humicola grisea, Humicolainsolens, Humicola lanuginosa, Irpex lacteus, Mucor miehei, Myceliophthora thermophila, Neurosporacrassa, Penicillium emersonii, Penicillium funiculosum, Penicillium purpurogenum, Phanerochaete chrysosporium, Polyporus pinsitus, Thermoascus aurantiacus, Thermoascus scabra The following fungi are listed: crustaceus, Thievora achromatica, Thievora albomyces, Thievora albopilosa, Thievora australeinsis, Thievora fimeti, Thievora microspora, Thievora iaovispora, Thievora peruviana, Thievora spededonium, Thievora setosa, Thievora subthermophila, Thievora terrestris, Trichoderma harzianum, Trichoderma koningii, Trichoderma longibrachiatum, and Trichoderma reesei. (reesei), or Trichoderma viride polypeptide.

[0250] On the other hand, the parent is a cellobiose hydrolase of Trichoderma reesei, such as the cellobiose hydrolase of SEQ ID NO:2 or its mature polypeptide; or a carbohydrate-binding module of Trichoderma reesei, such as the carbohydrate-binding module of SEQ ID NO:4.

[0251] On the other hand, the parent is a cellobiase of *Humicola insolens*, such as the cellobiase of SEQ ID NO:6 or its mature polypeptide; or a carbohydrate-binding module of *Humicola insolens*, such as the carbohydrate-binding module of SEQ ID NO:8.

[0252] On the other hand, the parent is a cellobiase of *Chaetoceros thermophilus*, such as the cellobiase of SEQ ID NO:10 or its mature polypeptide; or a carbohydrate-binding module of *Chaetoceros thermophilus*, such as the carbohydrate-binding module of SEQ ID NO:12.

[0253] On the other hand, the parent is a cellobiase of Hymenopterus filamentosa, such as the cellobiase of SEQ ID NO:14 or its mature polypeptide; or a carbohydrate-binding module of Hymenopterus filamentosa, such as the carbohydrate-binding module of SEQ ID NO:16.

[0254] On the other hand, the parent is a cellobiase of Talamomyces leycettanus, such as a cellobiase of SEQ ID NO:42, SEQ ID NO:44 or its mature polypeptide; or a carbohydrate-binding module of Talamomyces leycettanus, such as amino acids 472 to 507 of SEQ ID NO:42 or amino acids 472 to 507 of SEQ ID NO:44.

[0255] On the other hand, the parent is a cellobiase of Aspergillus fumigatus, such as a cellobiase of SEQ ID NO:18, SEQ ID NO:78 or its mature polypeptide; or a carbohydrate-binding module of Aspergillus fumigatus, such as a carbohydrate-binding module of SEQ ID NO:20.

[0256] On the other hand, the parent is a cellobiose hydrolase of *Clostridium terrestris*, such as the cellobiose hydrolase of SEQ ID NO:22 or its mature polypeptide; or a carbohydrate-binding module of *Clostridium terrestris*, such as the carbohydrate-binding module of SEQ ID NO:24.

[0257] On the other hand, the parent is a cellobiose hydrolase of *Thermophilus cytotoxicus*, such as the cellobiose hydrolase of SEQ ID NO:26 or its mature polypeptide; or a carbohydrate-binding module of *Thermophilus cytotoxicus*, such as the carbohydrate-binding module of SEQ ID NO:28.

[0258] It will be understood that, for the species mentioned above, this invention covers both perfect and imperfect states, as well as other taxonomic equivalents, such as asexual forms, regardless of their known species names. Those skilled in the art will readily identify the appropriate equivalents.

[0259] Strains of these species are readily available to the public at many culture collections, such as the American Type Culture Collection (ATCC), the German Microbial Culture Collection (DSMZ), the Netherlands Culture Collection (CentraalbureauVoor Schimmelcultures, CBS), and the Northern Research Center (NRRL) of the Patent Culture Collection of the U.S. Agricultural Research Service.

[0260] The parent can be identified and obtained from other sources, including microorganisms isolated from nature (e.g., soil, compost, water, etc.) or DNA samples obtained directly from natural materials (e.g., soil, compost, water, etc.), using the probes mentioned above. Techniques for directly isolating microorganisms and DNA from their natural habitat are well known in the art. The polynucleotide encoding the parent can then be obtained by similarly screening a library of genomic DNA or cDNA from another microorganism or a mixed DNA sample. Once the polynucleotide encoding the parent has been detected with one or more probes, it can be isolated or cloned using techniques known to those skilled in the art (see, for example, Sambrook et al., 1989, above).

[0261] Preparation of variants

[0262] The present invention also relates to methods for obtaining variants having cellobiase activity, the methods comprising: (a) introducing a substitution into a parent cellobiase at one or more (e.g., several) positions corresponding to positions 483, 491, 509 and 510 of SEQ ID NO:2, wherein the variant has cellobiase activity; and (b) recovering the variant.

[0263] The present invention also relates to methods for obtaining carbohydrate-binding module variants, the methods comprising: (a) introducing a substitution into a parent carbohydrate-binding module at one or more (e.g., several) positions corresponding to positions 5, 13, 31 and 32 of the carbohydrate-binding module of SEQ ID NO:4, wherein the variant has carbohydrate-binding activity; and (b) recovering the variant.

[0264] These variants can be prepared using any mutagenesis procedure known in the art, such as site-directed mutagenesis, synthetic gene construction, semi-synthetic gene construction, random mutagenesis, shuffling, etc.

[0265] Site-directed mutagenesis is a technique that introduces one or more (e.g., several) mutations at one or more designated sites in a polynucleotide encoding the parent.

[0266] Site-directed mutagenesis can be achieved in vitro using PCR involving primers containing oligonucleotides with the desired mutation. Site-directed mutagenesis can also be performed in vitro via cassette mutagenesis, which involves cleavage by a restriction enzyme at a site in a plasmid containing a polynucleotide encoding the parent and subsequent ligation of the oligonucleotide containing the mutation into the polynucleotide. Typically, the restriction enzyme used to digest the plasmid is the same as that used to digest the oligonucleotide to allow the sticky ends of the plasmid and the insert to ligate to each other. See, for example, Scherer and Davis, 1979, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA) 76:4949-4955; and Barton et al., 1990, Nucleic Acids Res. 18:7349-4966.

[0267] Site-directed mutagenesis can also be achieved in vivo using methods known in the art. See, for example, U.S. Patent Application Publication No. 2004 / 0171154; Storici et al., 2001, Nature Biotechnol. 19:773-776; Kren et al., 1998, Nat. Med. 4:285-290; and Calissano and Macino, 1996, Fungal Genet. Newslett. 43:15-16.

[0268] Any site-directed mutagenesis procedure can be used in this invention. Many commercially available kits are available for preparing variants.

[0269] Synthetic gene construction requires the in vitro synthesis of designed polynucleotide molecules to encode polypeptides of interest. Gene synthesis can be performed using a variety of techniques, such as the multi-channel microchip-based technique described by Tian et al. (2004, Nature 432:1050-1054), and similar techniques involving the synthesis and assembly of oligonucleotides on optically programmable microfluidic chips.

[0270] Single or multiple amino acid substitutions, deletions, and / or insertions can be made and tested using known methods of mutagenesis, recombination, and / or truncation, followed by relevant screening procedures, such as those disclosed by Reidhaar-Olson and Sauer, 1988, Science 241:53-57; Bowie and Sauer, 1989, Proceedings of the National Academy of Sciences of the United States of America (Proc. Natl. Acad. Sci. USA) 86:2152-2156; WO 95 / 17413; or WO 95 / 22625. Other methods that can be used include error-prone PCR, phage display (e.g., Lowman et al., 1991, Biochemistry 30:10832-10837; US Patent No. 5,223,409; WO 92 / 06204), and region-directed mutagenesis (Derbyshire et al., 1986, Gene 46:145; Ner et al., 1988, DNA 7:127).

[0271] Mutagenesis / reorganization methods can be combined with high-throughput automated screening methods to detect the activity of cloned mutagenic peptides expressed by host cells (Ness et al., 1999, Nature Biotechnology 17:893-896). The mutagenic DNA molecules encoding the active peptides can be recovered from the host cells and rapidly sequenced using standard methods in the art. These methods allow for the rapid determination of the importance of individual amino acid residues within the peptide.

[0272] Semi-synthetic gene construction is achieved through a combination of various methods, including synthetic gene construction, and / or site-directed mutagenesis, and / or random mutagenesis, and / or shuffling. Semi-synthetic construction typically utilizes the process of synthesizing polynucleotide fragments in conjunction with PCR technology. Therefore, specific regions of the gene can be synthesized de novo, while other regions can be amplified using site-specific mutagenic primers, and still others can undergo error-prone or non-error-prone PCR amplification. The polynucleotide subsequence can then be shuffled.

[0273] hybrid polypeptides

[0274] The present invention also relates to hybrid polypeptides comprising the carbohydrate-binding module variants described herein and heterologous catalytic domains of cellulases. In some embodiments, the hybrid polypeptide has carbohydrate-binding activity. In some embodiments, the hybrid polypeptide has cellulolytic activity (e.g., cellobiose hydrolase activity). In some embodiments, the hybrid polypeptide has both carbohydrate-binding activity and cellulolytic activity (e.g., cellobiose hydrolase activity).

[0275] This hybrid polypeptide can be formed by fusing the catalytic domain of a cellulase lacking a carbohydrate-binding module with a variant of the carbohydrate-binding module described herein, or by replacing the existing catalytic domain of a cellulase including a variant of the carbohydrate-binding module (such as the cellobiose hydrolase variant described herein) with the catalytic domain of a different cellulase.

[0276] In one respect, the carbohydrate-binding module variant is fused to the N-terminus of the heterogeneous catalytic domain. In the other respect, the carbohydrate-binding module variant is fused to the C-terminus of the heterogeneous catalytic domain.

[0277] One aspect is a hybrid polypeptide with cellulose-degrading activity, the hybrid polypeptide comprising:

[0278] (a) A fragment at the N-terminus of a hybrid polypeptide including a heterologous catalytic domain of cellulase; and

[0279] (b) A fragment at the C-terminus of a first polypeptide fragment comprising a variant of a carbohydrate-binding module, wherein the variant comprises substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31 and 32 of the carbohydrate-binding module of SEQ ID NO:4.

[0280] The catalytic domain used in the hybrid polypeptide can be any suitable catalytic domain of any cellulase described herein (the catalytic domain of any cellulase described in the Enzyme Composition section below), and can be obtained from any genus of microorganisms as described above.

[0281] For example, the catalytic domain can be obtained, in particular, from endoglucanase, cellobiase, or the GH61 polypeptide. In one embodiment, the catalytic domain is derived from endoglucanase. In another embodiment, the catalytic domain is derived from cellobiase. In yet another embodiment, the catalytic domain is derived from the GH61 polypeptide.

[0282] On the one hand, the catalytic domain of the hybrid polypeptide is a cellobiose hydrolase catalytic domain, and the hybrid polypeptide has cellobiose hydrolase activity.

[0283] The catalytic domain of this hybrid polypeptide can be a cellobiase from filamentous fungi. For example, the parent can be a cellobiase from filamentous fungi such as Aspergillus, Chaetomium, Aureospora, Trichoderma, Penicillium, Arthrophyll, Thermophila, or Trichoderma.

[0284] On one hand, the catalytic domain of this hybrid polypeptide is *Aspergillus echinosporum*, *Aspergillus avocado*, *Aspergillus sulphureus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus oryzae*, *Chaetomium thermophilum*, *Chrysosporium inops*, *Chrysosporium keratinophilum*, *Chrysosporium lucknowense*, *Chrysosporium merdarium*, *Chrysosporium pannicola*, *Chrysosporium queenslandicum*, *Chrysosporium tropicum*, and *Chrysosporium zonalum*. zonatum), thermophilic pyridobacterium, Penicillium emersonii, Penicillium cordiformis, Penicillium purpureum, Talaromyces byssochlamydoides, Talaromyces leycettanus, Trichoderma harzianum, Trichoderma corningensis, Trichoderma longibranchii, Trichoderma reesei, or Trichoderma viride cellobiose hydrolase catalytic domain.

[0285] In one embodiment, the catalytic domain is a heterologous catalytic domain of Trichoderma reesei cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:30, such as amino acids 1 to 429 of SEQ ID NO:30.

[0286] In another embodiment, the catalytic domain is a heterologous catalytic domain of Aspergillus fumigatus cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:36, such as amino acids 1 to 437 of SEQ ID NO:36.

[0287] In another embodiment, the catalytic domain is a heterologous catalytic domain of Ascomycota aurea cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:38, such as amino acids 1 to 440 of SEQ ID NO:38.

[0288] In another embodiment, the catalytic domain is a heterologous catalytic domain of Penicillium emersonii cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:40, such as amino acids 1 to 437 of SEQ ID NO:40.

[0289] In another embodiment, the catalytic domain is a heterologous catalytic domain of cellobiose hydrolase from Talamoyces leycettanus, such as the catalytic domain of SEQ ID NO:42, or amino acids 1 to 438 of SEQ ID NO:44.

[0290] In another embodiment, the catalytic domain is a heterologous catalytic domain of cellobiose hydrolase from Hymenopterus hymenopterus, such as the catalytic domain of SEQ ID NO:46, such as amino acids 1 to 437 of SEQ ID NO:46.

[0291] In another embodiment, the catalytic domain is a heterologous catalytic domain of thermophilic cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:48, such as amino acids 1 to 430 of SEQ ID NO:48.

[0292] In another embodiment, the catalytic domain is a heterologous catalytic domain of thermophilic cellobiose hydrolase, such as the catalytic domain of SEQ ID NO:50, such as amino acids 1 to 433 of SEQ ID NO:50.

[0293] In another aspect, it is a hybrid polypeptide with cellulose-degrading activity, comprising:

[0294] (a) A fragment at the N-terminus of a hybrid polypeptide including a heterologous catalytic domain of cellulase; wherein the fragment

[0295] (i) having at least 60% similarity to amino acids 1 to 429 of SEQ ID NO:30, amino acids 1 to 437 of SEQ ID NO:36, amino acids 1 to 440 of SEQ ID NO:38, amino acids 1 to 437 of SEQ ID NO:40, amino acids 1 to 437 of SEQ ID NO:42, amino acids 1 to 438 of SEQ ID NO:44, amino acids 1 to 437 of SEQ ID NO:46, amino acids 1 to 430 of SEQ ID NO:48, or amino acids 1 to 433 of SEQ ID NO:50.

[0296] (ii) Encoded by the following catalytic domain coding sequence, which hybridizes under low stringency conditions with nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, nucleotides 52 to 1389 of SEQ ID NO:32, nucleotides 79 to 1389 of SEQ ID NO:35, nucleotides 52 to 1371 of SEQ ID NO:37, nucleotides 55 to 1482 of SEQ ID NO:39, nucleotides 76 to 1386 of SEQ ID NO:41, nucleotides 76 to 1386 of SEQ ID NO:43, nucleotides 55 to 1504 of SEQ ID NO:45, nucleotides 61 to 1350 of SEQ ID NO:47, or nucleotides 55 to 1353 of SEQ ID NO:49; their cDNA sequence; or the aforementioned full-length complement;

[0297] (iii) Encoded by a catalytic domain coding sequence having at least 60% identity with nucleotides 52 to 1469 of SEQ ID NO:29, 52 to 1389 of SEQ ID NO:31, 52 to 1389 of SEQ ID NO:32, 79 to 1389 of SEQ ID NO:35, 52 to 1371 of SEQ ID NO:37, 55 to 1482 of SEQ ID NO:39, 76 to 1386 of SEQ ID NO:41, 76 to 1386 of SEQ ID NO:43, 55 to 1504 of SEQ ID NO:45, 61 to 1350 of SEQ ID NO:47, or 55 to 1353 of SEQ ID NO:49; or the cDNA sequence thereof;

[0298] (iv) is a variant of amino acids 1 to 429 of SEQ ID NO:30, amino acids 1 to 437 of SEQ ID NO:36, amino acids 1 to 440 of SEQ ID NO:38, amino acids 1 to 437 of SEQ ID NO:40, amino acids 1 to 437 of SEQ ID NO:42, amino acids 1 to 438 of SEQ ID NO:44, amino acids 1 to 437 of SEQ ID NO:46, amino acids 1 to 433 of SEQ ID NO:48, or amino acids 1 to 433 of SEQ ID NO:50, which includes substitution, deletion, and / or insertion at one or more (e.g., several) positions; or

[0299] (v) comprising or consisting of amino acids 1 to 429 of SEQ ID NO:30, amino acids 1 to 437 of SEQ ID NO:36, amino acids 1 to 440 of SEQ ID NO:38, amino acids 1 to 437 of SEQ ID NO:40, amino acids 1 to 437 of SEQ ID NO:42, amino acids 1 to 438 of SEQ ID NO:44, amino acids 1 to 437 of SEQ ID NO:46, amino acids 1 to 430 of SEQ ID NO:48, or amino acids 1 to 433 of SEQ ID NO:50; and

[0300] (b) A fragment at the C-terminus of a first polypeptide fragment comprising a variant of a carbohydrate-binding module, wherein the variant comprises substitutions at one or more (e.g., several) positions corresponding to positions 5, 13, 31 and 32 of the carbohydrate-binding module of SEQ ID NO:4.

[0301] In one embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 429 of SEQ ID NO:30; (ii) is encoded by a catalytic domain coding sequence that, under low stringency conditions, is identical to nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, or nucleotides 52 to 1389 of SEQ ID NO:32; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, or nucleotides 52 to 1389 of SEQ ID NO:32.

[0302] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 437 of SEQ ID NO:36; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 79 to 1389 of SEQ ID NO:35 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 79 to 1389 of SEQ ID NO:35.

[0303] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 440 of SEQ ID NO:38; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 52 to 1371 of SEQ ID NO:37 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 52 to 1371 of SEQ ID NO:37.

[0304] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 437 of SEQ ID NO:40; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 55 to 1482 of SEQ ID NO:39 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 55 to 1482 of SEQ ID NO:39.

[0305] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 437 of SEQ ID NO:42; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 76 to 1386 of SEQ ID NO:41 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 76 to 1386 of SEQ ID NO:41.

[0306] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 438 of SEQ ID NO:44; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 76 to 1386 of SEQ ID NO:43 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 76 to 1386 of SEQ ID NO:43.

[0307] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 437 of SEQ ID NO:46; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 55 to 1504 of SEQ ID NO:45 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 55 to 1504 of SEQ ID NO:45.

[0308] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain includes (i) having at least 60% identity with amino acids 1 to 430 of SEQ ID NO:48; (ii) being encoded by a catalytic domain coding sequence that hybridizes with nucleotides 61 to 1350 of SEQ ID NO:47 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) being encoded by a catalytic domain coding sequence having at least 60% identity with nucleotides 61 to 1350 of SEQ ID NO:47.

[0309] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide (i) has at least 60% identity with amino acids 1 to 433 of SEQ ID NO: 50; (ii) is encoded by a catalytic domain coding sequence that hybridizes with nucleotides 55 to 1353 of SEQ ID NO: 49 under low stringency conditions; its cDNA sequence; or the aforementioned full-length complement; or (iii) is encoded by a catalytic domain coding sequence that has at least 60% identity with nucleotides 55 to 1353 of SEQ ID NO: 49.

[0310] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 429 of SEQ ID NO:30, and has cellobiase activity. On the other hand, the amino acid sequence of the parent differs from amino acids 1 to 429 of SEQ ID NO:30 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 429 of SEQ ID NO:30.

[0311] In another embodiment, a fragment at the N-terminus of a heterologous catalytic domain of a hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 437 of SEQ ID NO:36, and possesses cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 437 of SEQ ID NO:36 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 437 of SEQ ID NO:36.

[0312] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 440 of SEQ ID NO:38, and has cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 440 of SEQ ID NO:38 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 440 of SEQ ID NO:38.

[0313] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 437 of SEQ ID NO:40, and has cellobiase activity. On the other hand, the amino acid sequence of the parent differs from amino acids 1 to 437 of SEQ ID NO:40 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 437 of SEQ ID NO:40.

[0314] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 437 of SEQ ID NO:42, and has cellobiose hydrolase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 437 of SEQ ID NO:42 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 437 of SEQ ID NO:42.

[0315] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 438 of SEQ ID NO:44, and has cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 438 of SEQ ID NO:44 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 438 of SEQ ID NO:44.

[0316] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 437 of SEQ ID NO:46, and has cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 437 of SEQ ID NO:46 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 437 of SEQ ID NO:46.

[0317] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 430 of SEQ ID NO:48, and has cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 430 of SEQ ID NO:48 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the hybrid polypeptide including the heterocatalytic domain comprises or consists of amino acids 1 to 430 of SEQ ID NO:48. In some embodiments, the hybrid polypeptide comprises or consists of SEQ ID NO:61. In other embodiments, the hybrid polypeptide comprises or consists of SEQ ID NO:63.

[0318] In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with amino acids 1 to 433 of SEQ ID NO:50, and has cellobiase activity. On the other hand, the amino acid sequence of this parent differs from amino acids 1 to 433 of SEQ ID NO:50 by up to 10 amino acids, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises or consists of amino acids 1 to 433 of SEQ ID NO:50.

[0319] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, nucleotides 52 to 1389 of SEQ ID NO:32, their cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0320] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 79 to 1389 of SEQ ID NO:35, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0321] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 52 to 1371 of SEQ ID NO:37; its cDNA sequence; or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0322] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 55 to 1482 of SEQ ID NO:39, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0323] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 76 to 1386 of SEQ ID NO:41, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0324] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 76 to 1386 of SEQ ID NO:43, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0325] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 55 to 1504 of SEQ ID NO:45, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0326] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 61 to 1350 of SEQ ID NO:47, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0327] On the other hand, the N-terminal fragment of the heterologous catalytic domain of the hybrid polypeptide is hybridized with nucleotides 55 to 1353 of SEQ ID NO:49, its cDNA sequence, or the aforementioned full-length complement (Salahbrook et al., 1989, ibid.) under low-strict, medium-strict, medium-high-strict, high-strict, or very high-strict conditions.

[0328] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide is identical to nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, nucleotides 52 to 1389 of SEQ ID NO:32; or its cDNA sequence has at least 60%, for example at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain comprises nucleotides 52 to 1469 of SEQ ID NO:29, nucleotides 52 to 1389 of SEQ ID NO:31, and nucleotides 52 to 1389 of SEQ ID NO:32; or a cDNA sequence thereof or composed thereof.

[0329] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the nucleotides 79 to 1389 of SEQ ID NO:35; or the cDNA sequence of the nucleotides 79 to 1389; or constitutes thereof. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or is composed of the nucleotides 79 to 1389 of SEQ ID NO:35; or the cDNA sequence of the nucleotides 79 to 1389; or thereof.

[0330] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the nucleotides 52 to 1371 of SEQ ID NO:37; or the cDNA sequence of the nucleotides 52 to 1371 of the SEQ ID NO:37; or constitutes thereof.

[0331] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises nucleotides 55 to 1482 of SEQ ID NO:39; or its cDNA sequence, having at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or consists of nucleotides 55 to 1482 of SEQ ID NO:39; or its cDNA sequence.

[0332] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises nucleotides 76 to 1386 of SEQ ID NO:41; or its cDNA sequence, having at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or consists of nucleotides 76 to 1386 of SEQ ID NO:41; or its cDNA sequence.

[0333] On the other hand, the fragment at the N-terminus of the hybrid polypeptide including the heterologous catalytic domain has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the heterologous catalytic domain. In another embodiment, the fragment at the N-terminus of the hybrid polypeptide including the heterologous catalytic domain comprises or is composed of the nucleotides 76 to 1386 of SEQ ID NO:43; or its cDNA sequence.

[0334] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises nucleotides 55 to 1504 of SEQ ID NO:45; or its cDNA sequence, having at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or consists of nucleotides 55 to 1504 of SEQ ID NO:45; or its cDNA sequence.

[0335] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide has at least 60%, for example, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with respect to the heterologous catalytic domain of the hybrid polypeptide. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or is composed of the nucleotides 61 to 1350 of SEQ ID NO:47; or its cDNA sequence.

[0336] On the other hand, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises nucleotides 55 to 1353 of SEQ ID NO:49; or its cDNA sequence, having at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In another embodiment, the fragment at the N-terminus of the heterologous catalytic domain of the hybrid polypeptide comprises or consists of nucleotides 55 to 1353 of SEQ ID NO:49; or its cDNA sequence.

[0337] The N-terminal fragment of a heterozygous polypeptide containing a heterocatalytic domain may further include substitutions, deletions, and / or insertions at one or more (e.g., several) other positions, such as changes at one or more (e.g., several) positions corresponding to those disclosed in PCT / US2014 / 022068, WO 2011 / 050037, WO 2005 / 028636, WO 2005 / 001065, WO 2004 / 016760, and U.S. Patent No. 7,375,197, the entire contents of which are incorporated herein by reference.

[0338] For example, in one aspect, the fragment including a heterocatalytic domain at the N-terminus includes a modified cellobiase at one or more positions corresponding to positions 197, 198, 199, and 200 of SEQ ID NO:30, wherein the modification at one or more positions corresponding to positions 197, 198, and 200 is a substitution and the modification at position 199 is a deletion. In another aspect, the fragment includes a modification at two positions corresponding to any of positions 197, 198, 199, and 200 of SEQ ID NO:30, wherein the modification at one or more positions corresponding to positions 197, 198, and 200 is a substitution and the modification at position 199 is a deletion. In yet another aspect, the fragment includes a modification at three positions corresponding to any of positions 197, 198, 199, and 200 of SEQ ID NO:30, wherein the modification at one or more positions corresponding to positions 197, 198, and 200 is a substitution and the modification at position 199 is a deletion. On the other hand, the fragment includes a substitution at each of the positions corresponding to positions 197, 198 and 200 and a deletion at the position corresponding to position 199.

[0339] In another aspect, the fragment comprising a heterocatalytic domain at the N-terminus includes or is composed of a substitution at position 197 corresponding to SEQ ID NO:30. In another aspect, the amino acid at position 197 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, the fragment comprises or is composed of the substituted N197A of SEQ ID NO:30.

[0340] In another aspect, the fragment comprising a heterocatalytic domain at the N-terminus includes or is composed of a substitution at position 198 corresponding to SEQ ID NO:30. In another aspect, the amino acid at position 198 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, the fragment comprises or is composed of the substituted N198A of SEQ ID NO:30.

[0341] On the other hand, the fragment comprising a heterocatalytic domain at the N-terminus includes a deletion or is composed of such a domain at position 199 corresponding to SEQ ID NO:30. On the other hand, the amino acid at position 199 is Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably Ala. On the other hand, this variant includes a deletion A199* of SEQ ID NO:30 or is composed of such a domain.

[0342] In another aspect, the fragment comprising a heterocatalytic domain at the N-terminus includes or is composed of a substitution at position 200 corresponding to SEQ ID NO:30. In another aspect, the amino acid at position 200 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala, Gly, or Trp. In another aspect, the fragment comprises or is composed of the substituted N200A, G, W of the mature polypeptide of SEQ ID NO:30.

[0343] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include alterations or composition thereof at positions 197 and 198 corresponding to SEQ ID NO:30, as described above.

[0344] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include alterations or composition thereof at positions 197 and 199 corresponding to SEQ ID NO:30, as described above.

[0345] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of alterations at positions 197 and 200 corresponding to SEQ ID NO:30, as described above.

[0346] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include changes or composition thereof at positions 198 and 199 corresponding to SEQ ID NO:30, as described above.

[0347] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include changes or composition thereof at positions 198 and 200 corresponding to SEQ ID NO:30, as described above.

[0348] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of alterations at positions 199 and 200 corresponding to SEQ ID NO:30, as described above.

[0349] On the other hand, segments comprising heterogeneous catalytic domains at the N-terminus include alterations or composition thereof at positions 197, 198, and 199 corresponding to SEQ ID NO:30, as described above.

[0350] On the other hand, segments comprising heterogeneous catalytic domains at the N-terminus include alterations or composition thereof at positions 197, 198, and 200 corresponding to SEQ ID NO:30, as described above.

[0351] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include alterations or composition thereof at positions 197, 199, and 200 corresponding to SEQ ID NO:30, as described above.

[0352] On the other hand, the fragments comprising heterogeneous catalytic domains at the N-terminus include changes or composition thereof at positions 198, 199 and 200 corresponding to SEQ ID NO:30, as those described above.

[0353] On the other hand, the fragments comprising heterogeneous catalytic domains at the N-terminus include changes or composition thereof at positions 197, 198, 199 and 200 corresponding to SEQ ID NO:30, as those described above.

[0354] On the other hand, segments that include a heterogeneous catalytic domain at the N-terminus include one or more variations of or constitute the group consisting of N197A, N198A, A199*, and N200A,G,W.

[0355] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of changes to N197A+N198A of SEQ ID NO:30.

[0356] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of changes to N197A+A199* of SEQ ID NO:30.

[0357] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30 N197A+N200A,G,W.

[0358] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of changes to N198A+A199* of SEQ ID NO:30.

[0359] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30 N198A+N200A,G,W.

[0360] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30, such as A199*+N200A,G,W.

[0361] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of changes to SEQ ID NO:30 N197A+N198A+A199*.

[0362] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30, such as N197A+N198A+N200A, G, W, or the like.

[0363] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30, such as N197A+A199*+N200A,G,W.

[0364] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30, such as N198A+A199*+N200A,G,W.

[0365] On the other hand, fragments comprising heterogeneous catalytic domains at the N-terminus include or consist of variations of SEQ ID NO:30, such as N197A+N198A+A199*+N200A,G,W.

[0366] These hybrid polypeptide carbohydrate-binding module variants can be any suitable carbohydrate-binding module variants described above.

[0367] On one hand, the carbohydrate-binding module variant of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100%, sequence identity with the parental carbohydrate-binding module.

[0368] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide has at least 60%, such as at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, such as at least 96%, at least 97%, at least 98%, or at least 99%, but less than 100% sequence identity with the polypeptides of SEQ ID NO:4, SEQ ID NO:8, SEQ ID NO:12, SEQ ID NO:16, SEQ ID NO:20, SEQ ID NO:24, or SEQ ID NO:28.

[0369] On the one hand, the number of substitutions in the carbohydrate-binding module variants of this hybrid polypeptide is 1 to 4, such as 1, 2, 3, or 4 substitutions.

[0370] In one aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of a substitution at position 5 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 5 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 5 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 5 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y5W of SEQ ID NO:4).

[0371] On the other hand, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of substitutions at position 13 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 13 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 13 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 13 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y13W of SEQ ID NO:4).

[0372] On the other hand, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of substitutions at position 31 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 31 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 31 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 31 corresponding to SEQ ID NO:4 is Tyr replaced by Trp (e.g., Y31W of SEQ ID NO:4).

[0373] On the other hand, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of substitutions at position 32 corresponding to SEQ ID NO:4. In one embodiment, the amino acid at position 32 is replaced by Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, such as Tyr, Phe, or Trp. In another embodiment, the amino acid at position 32 corresponding to SEQ ID NO:4 is replaced by Trp. In yet another embodiment, the amino acid at position 32 corresponding to SEQ ID NO:4 is Tyr substituted with Trp (e.g., Y5W of SEQ ID NO:4).

[0374] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide includes substitutions or is composed of substitutions at positions 5, 13, 31, and 32 corresponding to SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 13 (e.g., substitution by Trp at positions 5 and 13, such as Y5W and / or Y13W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 31 (e.g., substitution by Trp at positions 5 and 31, such as Y5W and / or Y31W). In yet another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions 5 and 32 (e.g., substitution by Trp at positions 5 and 32, such as Y5W and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13 and 31 (e.g., substitution by Trp at positions corresponding to positions 13 and 31, such as Y13W and / or Y31W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13 and 32 (e.g., substitution by Trp at positions corresponding to positions 13 and 32, such as Y13W and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 31 and 32 (e.g., substitution by Trp at positions corresponding to positions 31 and 32, such as Y31W and / or Y32W).

[0375] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of substitutions at positions 5, 13, 31, and 32 corresponding to SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes or is composed of substitutions at positions 5, 13, and 31 (e.g., Trp substitution at positions 5, 13, and 31, such as Y5W, Y13W, and / or Y31W). In another embodiment, the carbohydrate-binding module variant includes or is composed of substitutions at positions 5, 13, and 32 (e.g., Trp substitution at positions 5, 13, and 32, such as Y5W, Y13W, and / or Y32W). In yet another embodiment, the carbohydrate-binding module variant includes or is composed of substitutions at positions 5, 31, and 32 (e.g., Trp substitution at positions 5, 31, and 32, such as Y5W, Y31W, and / or Y32W). In another embodiment, the carbohydrate-binding module variant includes substitutions or is composed of substitutions at positions corresponding to positions 13, 31, and 32 (e.g., substitution by Trp at positions corresponding to positions 13, 31, and 32, such as Y13W, Y31W, and / or Y32W).

[0376] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide includes substitutions or is composed of substitutions at all four positions corresponding to positions 5, 13, 31, and 32 of SEQ ID NO:4, as described above. In one embodiment, the carbohydrate-binding module variant includes Trp substitutions or is composed of Trp at one or more positions corresponding to positions 5, 13, 31, and 32, such as Y5W, Y13W, Y31W, and / or Y32W.

[0377] The carbohydrate-binding module variant of the hybrid polypeptide may further include substitutions, deletions, and / or insertions at one or more (e.g., several) other positions, such as one or more (e.g., several) substitutions at positions corresponding to those disclosed in WO 2012 / 135719, which is incorporated herein by reference. For example, in one aspect, the carbohydrate-binding module variant of the hybrid polypeptide further includes substitutions at one or more (e.g., several) positions corresponding to positions 4, 6, and 29 of SEQ ID NO:4. In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide further includes substitutions at two positions corresponding to any one of positions 4, 6, and 29. In yet another aspect, the carbohydrate-binding module variant of the hybrid polypeptide further includes substitutions at each position corresponding to positions 4, 6, and 29.

[0378] In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of a substitution at the position corresponding to position 4. In another aspect, the amino acid at the position corresponding to position 4 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Glu, Leu, Lys, Phe, or Trp. In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of the substituted H4L of SEQ ID NO:4. In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of the substituted H4K of SEQ ID NO:4. In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of the substituted H4E of SEQ ID NO:4. In another aspect, the carbohydrate-binding module variant of the hybrid polypeptide includes or is composed of the substituted H4F of SEQ ID NO:4. On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of the substituted H4W of SEQ ID NO:4.

[0379] In another aspect, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of a substitution at position 6. In another aspect, the amino acid at position 6 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Ala. In another aspect, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of the substituted G6A of SEQ ID NO:4.

[0380] In another aspect, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of substitutions at the position corresponding to position 29. In another aspect, the amino acid at the position corresponding to position 29 is substituted with Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val, preferably substituted with Asp. In another aspect, the carbohydrate-binding module variant of this hybrid polypeptide includes or is composed of the substituted N29D of SEQ ID NO:4.

[0381] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide further includes or consists of substitutions at positions corresponding to positions 4 and 6, as those described above.

[0382] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide further includes or consists of substitutions at positions corresponding to positions 4 and 29, as described above.

[0383] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide further includes or consists of substitutions at positions corresponding to positions 6 and 29, as described above.

[0384] On the other hand, the carbohydrate-binding module variants of the hybrid polypeptide further include or consist of substitutions at positions corresponding to positions 4, 6, and 29, as those described above.

[0385] On the other hand, the carbohydrate-binding module variant of the hybrid polypeptide further includes or consists of one or more (e.g., several) substitutions selected from the group consisting of H4L,K,E,F,W,G6A, and N29D; or one or more (e.g., several) substitutions selected from the group consisting of H4L,K,E,F,W,G6A, and N29D corresponding to SEQ ID NO:4 in the other cellulose-binding modules described herein.

[0386] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4L+G6A of SEQ ID NO:4.

[0387] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4K+G6A of SEQ ID NO:4.

[0388] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4E+G6A of SEQ ID NO:4.

[0389] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4F+G6A of SEQ ID NO:4.

[0390] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4W+G6A of SEQ ID NO:4.

[0391] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4L+N29D of SEQ ID NO:4.

[0392] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4K+N29D of SEQ ID NO:4.

[0393] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4E+N29D of SEQ ID NO:4.

[0394] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4F+N29D of SEQ ID NO:4.

[0395] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4W+N29D of SEQ ID NO:4.

[0396] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted G6A+N29D of SEQ ID NO:4.

[0397] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4L+G6A+N29D of SEQ ID NO:4.

[0398] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4K+G6A+N29D of SEQ ID NO:4.

[0399] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4E+G6A+N29D of SEQ ID NO:4.

[0400] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4F+G6A+N29D of SEQ ID NO:4.

[0401] On the other hand, carbohydrate-binding module variants of this hybrid polypeptide include or consist of substituted H4W+G6A+N29D of SEQ ID NO:4.

[0402] In some embodiments, the hybrid polypeptide comprises or consists of SEQ ID NO:61. In other embodiments, the hybrid polypeptide is encoded by the coding sequence of SEQ ID NO:60.

[0403] In other embodiments, the hybrid polypeptide comprises or is composed of SEQ ID NO:63. In other embodiments, the hybrid polypeptide is encoded by the coding sequence of SEQ ID NO:62.

[0404] In other embodiments, the hybrid polypeptide comprises or is composed of SEQ ID NO:73. In other embodiments, the hybrid polypeptide is encoded by the coding sequence of SEQ ID NO:72.

[0405] In other embodiments, the hybrid polypeptide comprises or is composed of SEQ ID NO:94. In other embodiments, the hybrid polypeptide is encoded by the coding sequence of SEQ ID NO:93.

[0406] Essential amino acids in the parent can be identified according to procedures known in the art, as described herein.

[0407] Techniques for generating fusion peptides are known in the art and involve linking coding sequences of the peptides such that they are within a frame and the expression of the fusion peptide is under the control of the same one or more promoters and terminators. Fusion peptides can also be constructed using integrin technology, wherein the fusion peptide is generated post-translationally (Cooper et al., 1993, EMBO J. 12:2575-2583; Dawson et al., 1994, Science 266:776-779).

[0408] Polynucleotides

[0409] The present invention also relates to the isolation of polynucleotides encoding the carbohydrate-binding module variants, cellobiase variants, and hybrid polypeptides of the present invention.

[0410] Nucleic acid constructs

[0411] The present invention also relates to a nucleic acid construct comprising a polynucleotide operatively linked to one or more control sequences encoding a carbohydrate-binding module variant, a cellobiase variant, and a hybrid polypeptide of the present invention, wherein the one or more control sequences guide the expression of the coding sequence in a suitable host cell under conditions compatible with the control sequences.

[0412] The polynucleotide can be manipulated in a variety of ways to provide expression of the variant. Depending on the expression vector, manipulation of the polynucleotide before its insertion into the vector may be desired or necessary. Techniques for modifying polynucleotides using recombinant DNA methods are well known in the art.

[0413] The control sequence can be a promoter, i.e., a polynucleotide recognized by the host cell to express a polynucleotide encoding the polypeptide of the present invention. The promoter contains a transcriptional control sequence that mediates polypeptide expression. The promoter can be any polynucleotide exhibiting transcriptional activity in the host cell, including mutant, truncated, and heterozygous promoters, and can be obtained from a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to that of the host cell.

[0414] Examples of suitable promoters for directing the transcription of the nucleic acid constructs of this invention in bacterial host cells are promoters obtained from the following genes: Bacillus amyloliquefaciens α-amylase gene (amyQ), Bacillus licheniformis α-amylase gene (amyL), Bacillus licheniformis penicillinase gene (penP), Bacillus thermophilus maltose amylase gene (amyM), Bacillus subtilis fructan sucrase gene (sacB), Bacillus subtilis xylA and xylB genes, Bacillus thuringiensis cryIIIA gene (Agaisse and Lereclus, 1994, Molecular Microbiology). Microbiology 13:97-107), Escherichia coli lac operon, Escherichia coli trc promoter (Egon et al., 1988, Gene 69:301-315), Streptomyces agar hydrolase gene (dagA), and prokaryotic β-lactamase gene (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA 75:3727-3731) and tac promoter (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA 80:21-25). Other promoters are described in Gilbert et al., 1980, Scientific American 242:74-94, “Useful proteins from recombinant bacteria”; and in Sambrook et al., 1989, ibid. Examples of tandem promoters are disclosed in WO 99 / 43835.

[0415] Examples of suitable promoters for guiding the transcription of the nucleic acid constructs of this invention in filamentous fungal host cells are promoters derived from the following genes: Aspergillus nidulans acetamase, Aspergillus niger neutral α-amylase, Aspergillus niger acid-stable α-amylase, Aspergillus niger or Aspergillus avocado glucoamylase (glaA), Aspergillus oryzae TAKA amylase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Fusarium oxysporum trypsin-like protease (WO 96 / 00787), Fusarium variegatum amyglucosidase (WO 00 / 56900), Fusarium variegatum Daria (WO 00 / 56900), Fusarium variegatum Quinn (WO 00 / 56900). 00 / 56900), *Rhizopus mirtii* lipase, *Rhizopus mirtii* aspartic protease, *Trichoderma reesei* β-glucosidase, *Trichoderma reesei* cellobiose hydrolase I, *Trichoderma reesei* cellobiose hydrolase II, *Trichoderma reesei* endoglucanase I, *Trichoderma reesei* endoglucanase II, *Trichoderma reesei* endoglucanase III, *Trichoderma reesei* endoglucanase V, *Trichoderma reesei* xylanase I, *Trichoderma reesei* xylanase II, *Trichoderma reesei* xylanase III, *Trichoderma reesei* β-xylosidase, and *Trichoderma reesei* translational elongation factor. The promoter, together with the NA2-tpi promoter (a modified promoter from an Aspergillus gene encoding neutral α-amylase, wherein the untranslated pre-progenitor has been replaced with an untranslated pre-progenitor from an Aspergillus gene encoding triose phosphate isomerase; non-limiting examples include a modified promoter from an Aspergillus niger gene encoding neutral α-amylase, wherein the untranslated pre-progenitor has been replaced with an untranslated pre-progenitor from an Aspergillus nidulans or Aspergillus oryzae gene encoding triose phosphate isomerase), and its mutant, truncated, and heterozygous promoters. Other promoters are described in U.S. Patent No. 6,011,147.

[0416] In yeast hosts, useful promoters are derived from genes targeting the following: *Saccharomyces cerevisiae* enolase (ENO-1), *Saccharomyces cerevisiae* galactokinase (GAL1), *Saccharomyces cerevisiae* alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH1, ADH2 / GAP), *Saccharomyces cerevisiae* triose phosphate isomerase (TPI), *Saccharomyces cerevisiae* metallothionein (CUP1), and *Saccharomyces cerevisiae* 3-phosphate glycerate kinase. Other useful promoters in yeast host cells are described in Romanos et al., 1992, Yeast 8:423-488.

[0417] The control sequence can also be a transcription terminator recognized by the host cell to terminate transcription. This terminator is operatively linked to the 3' end of the polynucleotide encoding the variant. Any terminator that functions within the host cell can be used in this invention.

[0418] Preferred terminators for bacterial host cells were obtained from the genes of Bacillus clausti alkaline protease (aprH), Bacillus licheniformis α-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB).

[0419] Preferred terminators for filamentous fungal host cells are derived from the genes of the following: Aspergillus nidulans acetamase, Aspergillus nidulans o-aminobenzoic acid synthase, Aspergillus niger glucosylamylase, Aspergillus niger α-glucosidase, Aspergillus oryzae TAKA amylase, Fusarium oxysporum trypsin-like protease, Trichoderma reesei β-glucosidase, Trichoderma reesei cellobiose hydrolase I, Trichoderma reesei cellobiose hydrolase II, Trichoderma reesei endoglucanase I, Trichoderma reesei endoglucanase II, Trichoderma reesei endoglucanase III, Trichoderma reesei endoglucanase V, Trichoderma reesei xylanase I, Trichoderma reesei xylanase II, Trichoderma reesei xylanase III, Trichoderma reesei β-xylosidase, and Trichoderma reesei translation elongation factor.

[0420] Preferred terminators for yeast host cells are derived from genes targeting: *Saccharomyces cerevisiae* enolase, *Saccharomyces cerevisiae* cytochrome C (CYC1), and *Saccharomyces cerevisiae* glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described above in Romanos et al., 1992.

[0421] Control sequences can also be mRNA stabilizing regions downstream of the promoter and upstream of the gene's coding sequence, which increase the expression of the gene.

[0422] Examples of suitable mRNA stable regions were obtained from the following: Bacillus thuringiensis cryIIIA gene (WO 94 / 25612) and Bacillus subtilis SP82 gene (Hue et al., 1995, Journal of Bacteriology 177:3465-3471).

[0423] The control sequence can also be a leader, a non-translated mRNA region that is important for translation in the host cell. This leader sequence is operatively linked to the 5' end of the polynucleotide encoding the variant. Any leader that is functional in the host cell can be used.

[0424] Preferred precursors for use in filamentous fungal host cells were obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase.

[0425] Precursors suitable for yeast host cells are obtained from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0426] The control sequence can also be a polyadenylation sequence, i.e., a sequence operatively linked to the 3' end of the variant encoding a polynucleotide and recognized by the host cell during transcription as a signal to add polyadenylation residues to the transcribed mRNA. Any polyadenylation sequence that is functional in the host cell can be used.

[0427] Preferred polyadenylation of filamentous fungal host cells is obtained from genes targeting the following: Aspergillus nidulans o-aminobenzoate synthase, Aspergillus niger glucosidase, Aspergillus niger glutamate-glucosidase, Aspergillus oryzae TAKA amylase, and Fusarium oxysporum trypsin-like protease.

[0428] A polyadenylation sequence for yeast host cells is described in Guo and Sherman, 1995, Molecular Cell Biology, 15:5983-5990.

[0429] The control sequence can also be a signal peptide coding region, encoding a signal peptide linked to the N-terminus of the variant and guiding the variant into the cell's secretory pathway. The 5' end of the polynucleotide coding sequence may inherently contain a signal peptide coding sequence naturally linked within the translation reading frame to a segment encoding the variant's coding sequence. Alternatively, the 5' end of the coding sequence may include a signal peptide coding sequence that is exogenous to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding sequence, an exogenous signal peptide coding sequence may be required. Alternatively, an exogenous signal peptide coding sequence may simply replace the native signal peptide coding sequence to enhance the variant's secretion. However, any signal peptide coding sequence that guides the expressed variant into the host cell's secretory pathway can be used.

[0430] Effective signal peptide coding sequences for bacterial host cells are obtained from the following genes: maltose amylase from Bacillus NCIB 11837, subtilisin from Bacillus licheniformis, β-lactamase from Bacillus licheniformis, α-amylase from Bacillus thermophilus, neutral proteases (nprT, nprS, nprM) from Bacillus thermophilus, and prsA from Bacillus subtilis. Additional signal peptides are described in Simonen and Palva, 1993, Microbiological Reviews 57:109-137.

[0431] The effective signal peptide coding sequences for filamentous fungal host cells are obtained from the following genes: Aspergillus niger neutral amylase, Aspergillus niger glucosylase, Aspergillus oryzae TAKA amylase, Aspergillus oryzae cellulase, Aspergillus oryzae endoglucanase V, Aspergillus pubescens lipase, and Aspergillus oryzae aspartic protease.

[0432] Useful signal peptides for yeast host cells were obtained from the genes of *Saccharomyces cerevisiae* saccharide and *Saccharomyces cerevisiae* invertase. Sequences encoding other useful signal peptides were described by Romanos et al. (1992, above).

[0433] The control sequence can also be a propeptide-coding sequence encoding a propeptide located at the N-terminus of the variant. The resulting polypeptide is called a proenzyme or propeptide progenitor (or, in some cases, a zymogen). The propeptide progenitor is usually inactive and can be converted into an active polypeptide by catalytic cleavage or autocatalytic cleavage of the propeptide progenitor. The propeptide-coding sequence can be obtained from the genes of Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Thermophilus laccase (WO 95 / 33836), Rhizopus oryzae aspartic protease, and Saccharomyces cerevisiae α-factor.

[0434] In the case where both the signal peptide and the propeptide sequence are present, the propeptide sequence is positioned immediately adjacent to the N-terminus of the variant, and the signal peptide sequence is positioned immediately adjacent to the N-terminus of the propeptide sequence.

[0435] It is also desirable to add regulatory sequences that modulate the expression of the variant relative to the growth of the host cell. Examples of regulatory sequences are those sequences that cause gene expression to be turned on or off in response to chemical or physical stimuli, including the presence of regulatory compounds. Regulatory sequences in prokaryotic systems include the lac, tac, and trp operon systems. In yeast, the ADH2 or GAL1 systems can be used. In filamentous fungi, the *Aspergillus niger* glucosylamylase promoter, the *Aspergillus oryzae* TAKA α-amylase promoter and *Aspergillus oryzae* glucosylamylase promoter, the *Trichoderma reesei* cellobiose hydrolase I promoter, and the *Trichoderma reesei* cellobiose hydrolase II promoter can be used. Other examples of regulatory sequences are those sequences that allow gene amplification. In eukaryotic systems, these regulatory sequences include dihydrofolate reductase genes amplified in the presence of methotrexate and metallothionein genes amplified with heavy metals. In these cases, the polynucleotide encoding the variant will be operatively linked to the regulatory sequence.

[0436] expression carrier

[0437] The present invention also relates to recombinant expression vectors comprising a polynucleotide encoding a carbohydrate-binding module variant, a cellobiase variant, or a hybrid polypeptide of the present invention, along with a promoter and transcription and translation termination signals. Various nucleotides and control sequences can be linked together to produce a recombinant expression vector, which may contain one or more suitable restriction sites to allow insertion or substitution of the polynucleotide encoding the variant sequence at such sites. Alternatively, the polynucleotide can be expressed by inserting the polynucleotide or a nucleic acid construct containing the polynucleotide into a suitable vector for expression. In producing the expression vector, the coding sequence is located within the vector, such that the coding sequence is operatively linked to the suitable control sequence for expression.

[0438] The recombinant expression vector can be any vector (e.g., plasmid or virus) that can readily undergo recombinant DNA procedures and induce polynucleotide expression. The choice of vector will typically depend on its compatibility with the host cell to which it will be introduced. The vector can be a linear or closed circular plasmid.

[0439] The vector can be a self-replicating vector, that is, a vector existing as an extrachromosomal entity whose replication is independent of chromosome replication, such as a plasmid, extrachromosomal element, microchromosome, or artificial chromosome. The vector can contain any elements necessary to ensure self-replication. Alternatively, the vector can be one that, when introduced into the host cell, is integrated into the genome and replicates along with one or more chromosomes in which it has been integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids (which together contain the total DNA of the genome to be introduced into the host cell), or transposons can be used.

[0440] The vector preferably contains one or more selective markers that allow for convenient selection of cells such as transformed cells, transfected cells, and transduced cells. A selective marker is a gene whose product provides resistance to biocides or viruses, heavy metal resistance, or auxotrophic prototrophs, etc.

[0441] Examples of bacterial selective markers include the dal gene in Bacillus licheniformis or Bacillus subtilis, or markers that confer antibiotic resistance (such as resistance to ampicillin, chloramphenicol, kanamycin, neomycin, spectinomycin, or tetracycline). Suitable markers for use in yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selective markers for use in filamentous fungal host cells include, but are not limited to, adeA (phosphogluconoylaminoimidazolium-succinate carboxylamine synthase), adeB (phosphogluconoylaminoimidazolium synthase), amdS (acetamipase), argB (ornithine carbamoyltransferase), bar (glufosinate acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotate nucleoside-5'-phosphate decarboxylase), sC (adenosyl sulfate transtransferase), and trpC (o-aminobenzoic acid synthase), along with their equivalents. The amdS and pyrG genes of *Aspergillus nidus* or *Aspergillus oryzae*, and the bar gene of *Streptomyces hygroscopicus* are preferably used in *Aspergillus* cells. In *Trichoderma* cells, the adeA, adeB, amdS, hph, and pyrG genes are preferred.

[0442] Selective labeling can be a biselective labeling system as described in WO 2010 / 039889. In one aspect, the biselective labeling is the hph-tk biselective labeling system.

[0443] The vector preferably contains one or more elements that allow the vector to integrate into the host cell’s genome or to replicate autonomously in the cell independently of the genome.

[0444] For integration into the host cell genome, the vector can rely on a polynucleotide sequence encoding the variant or any other element of the vector for integration into the genome via homologous or non-homologous recombination. Alternatively, the vector can contain additional polynucleotides to guide integration into one or more precise locations on one or more chromosomes within the host cell genome via homologous recombination. To increase the likelihood of integration at precise locations, these integrative elements should contain a sufficient number of nucleic acids, such as 100 to 10,000 base pairs, 400 to 10,000 base pairs, and 800 to 10,000 base pairs, that have high sequence identity with the corresponding target sequence to enhance the likelihood of homologous recombination. These integrative elements can be any sequence homologous to the target sequence within the host cell genome. Furthermore, these integrative elements can be non-coding or coding polynucleotides. On the other hand, the vector can integrate into the host cell genome via non-homologous recombination.

[0445] For autonomous replication, the vector may further include an origin of replication that enables the vector to replicate autonomously in the host cell in question. The origin of replication can be any plasmid replicon that mediates autonomous replication and functions within the cell. The terms "origin of replication" or "plasmid replicator" refer to a polynucleotide that enables a plasmid or vector to replicate in vivo.

[0446] Examples of bacterial origins of replication are the origins of replication of plasmids pBR322, pUC19, pACYC177, and pACYC184, which allow replication in Escherichia coli, and the origins of replication of plasmids pUB110, pE194, pTA1060, and pAMβ1, which allow replication in Bacillus.

[0447] Examples of replication origins used in yeast host cells include the 2-micron replication origin, ARS1, ARS4, a combination of ARS1 and CEN3, and a combination of ARS4 and CEN6.

[0448] Examples of useful origins of replication in filamentous fungal cells are AMA1 and ANS1 (Gems et al., 1991, Gene 98:61-67; Cullen et al., 1987, Nucleic Acids Res. 15:9163-9175; WO 00 / 24883). The isolation of the AMA1 gene and the construction of plasmids or vectors containing this gene can be performed according to the methods disclosed in WO 00 / 24883.

[0449] More than one copy of the polynucleotide of the present invention can be inserted into host cells to increase the generation of variants. An increased copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene along with the polynucleotide, wherein cells containing an amplified copy of the selectable marker gene, and thus additional copies of the polynucleotide, can be selected by culturing cells in the presence of a suitable selectivity reagent.

[0450] The procedures for connecting the above-described elements to construct the recombinant expression vector of the present invention are well known to those skilled in the art (see, for example, Sambrook et al., 1989, ibid.).

[0451] host cells

[0452] This invention also relates to recombinant host cells comprising a polynucleotide operably linked to one or more control sequences encoding a carbohydrate-binding module variant, a cellobiase variant, or a hybrid polypeptide of the invention, which directs the production of the desired variant polypeptide. A construct or vector comprising the polynucleotide is introduced into the host cell such that the construct or vector is maintained as a chromosomal integrase or as an autonomously replicating extrachromosomal vector, as previously described. The term "host cell" encompasses any progeny of a parent cell that differs from the parent cell due to mutations occurring during replication. The selection of the host cell will depend largely on the gene encoding the variant and its origin.

[0453] The host cell can be any cell that is useful in the recombinant-generated variants, such as prokaryotic or eukaryotic cells.

[0454] Prokaryotic host cells can be any Gram-positive or Gram-negative bacteria. Gram-positive bacteria include, but are not limited to: Bacillus, Clostridium, Enterococcus, Bacillus aeruginosa, Lactobacillus, Lactococcus, Marine Bacillus, Staphylococcus, Streptococcus, and Streptomyces. Gram-negative bacteria include, but are not limited to: Campylobacter, Escherichia coli, Flavobacterium, Clostridium, Helicobacter, Coliform, Neisseria, Pseudomonas, Salmonella, and Ureaplasma.

[0455] The bacterial host cell can be any Bacillus genus cell, including but not limited to: Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus croceus, Bacillus coagulans, Bacillus sclerosus, Bacillus splenium, Bacillus stenosis, Bacillus licheniformis, Bacillus megaterium, Bacillus brevis, Bacillus thermophilus, Bacillus subtilis, and Bacillus thuringiensis cells.

[0456] The bacterial host cell can also be any streptococcal cell, including but not limited to: Streptococcus equina, Streptococcus pyogenes, Streptococcus mammae, and Streptococcus equine subsp. veterinaryis.

[0457] The bacterial host cell can also be any Streptomyces cell, including but not limited to: non-chromogenic Streptomyces, insecticidal Streptomyces, sky blue Streptomyces, gray Streptomyces, and light blue Streptomyces cells.

[0458] DNA can be introduced into Bacillus cells via the following methods: protoplast transformation (see, for example, Chang and Cohen, 1979, Molecular Genetics and Genomics, 168:111-115), and competent cell transformation (see, for example, Young and Spizizen, 1961, Journal of Bacteriology, 81:823-829; or Dubnau and David Dubnau). Davidoff-Abelson, 1971, Journal of Molecular Biology 56:209-221, electroporation (see, e.g., Shigekawa and Dower, 1988, Biotechniques 6:742-751), or conjugation (see, e.g., Koehler and Thorne, 1987, Journal of Bacteriology 169:5271-5278). DNA can be introduced into *E. coli* cells via protoplast transformation (see, e.g., Hanahan, 1983, Journal of Molecular Biology 166:557-580) or electroporation (see, e.g., Dower et al., 1988, Nucleic Acids Res. 16:6127-6145). DNA can be introduced into Streptomyces cells via protoplast transformation, electroporation (see, for example, Gong et al., 2004, Folia Microbiol. (Praha) 49:399-405), conjugation (see, for example, Mazodier et al., 1989, Journal of Bacteriol. 171:3583-3585), or transduction (see, for example, Burke et al., 2001, Proceedings of the National Academy of Sciences of the United States of America (Proc. Natl. Acad. Sci. USA) 98:6289-6294). DNA can be introduced into Pseudomonas cells by electroporation (see, for example, Choi et al., 2006, Journal of Microbiological Methods, 64:391-397) or conjugation (see, for example, Pinedo and Smets, 2005, Appl. Environ. Microbiol., 71:51-57).DNA can be introduced into Streptococcus cells via the following methods: native competent cells (see, e.g., Perry and Kuramitsu, 1981, Infect. Immun. 32:1295-1297), protoplast transformation (see, e.g., Catt and Jollick, 1991, Microbios 68:189-207), electroporation (see, e.g., Buckley et al., 1999, Appl. Environ. Microbiol. 65:3800-3804), or conjugation (see, e.g., Clewell, 1981, Microbiol. Rev. 45:409-436). However, any method known in the art for introducing DNA into host cells can be used.

[0459] The host cell can also be a eukaryotic cell, such as a mammalian, insect, plant, or fungal cell.

[0460] The host cell can be a fungal cell. As used herein, “fungus” includes Ascomycota, Basidiomycota, Chytridiomycota, Zygomycota, Oomycota, and all mitotic fungi (as defined by Hawksworth et al., in: Dictionary of the Fungi, Ainsworth and Bisby, 8th ed., 1995, CAB International, University Press, Cambridge, UK).

[0461] The host cell for fungi can be a yeast cell. As used herein, "yeast" includes *Entomophycetes* (order Endosporales), *Basidiophycetes*, and yeasts belonging to the class Deuteromycetes (class Bacillus). Since the classification of yeast may change in the future, for the purposes of this invention, yeast should be defined as described in *Biology and Activities of Yeast* (Skinner, Passmore, and Davenport, eds., Soc. App. Bacteriol. Symposium Series No. 9, 1980).

[0462] Yeast host cells can be cells of the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia pastoris*, *Saccharomyces*, *Saccharomyces*, or *Yerovia*, such as *Kluyveromyces lactis*, *Saccharomyces cerevisiae*, *Saccharomyces cerevisiae*, *Saccharomyces sacchariformis*, *Saccharomyces douglasii*, *Kluyveromyces kluyveromyces*, *Nordiya*, *Ovo*, or *Yerovia lipolytica*.

[0463] The host cell of this fungus can be a filamentous fungal cell. "Filamentous fungi" includes all filamentous forms of the subphylum Eumycota and Oomycota (as defined above by Hawkesworth et al., 1995). Filamentous fungi are typically characterized by a hyphal wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth occurs through hyphal extension, and carbon metabolism is obligate aerobic. In contrast, yeast (such as Saccharomyces cerevisiae) grows vegetatively through budding of single-celled cells, and carbon metabolism can be fermentation.

[0464] The host cells of filamentous fungi can be cells from genera such as *Apertoire*, *Aspergillus*, *Bjerkandera*, *Pseudomonas*, *Aureospora*, *Coprinus*, *Coriolus*, *Cryptococcus*, *Filibasidium*, *Fusarium*, *Pyrophyllus*, *Magnaporthe*, *Mucor*, *Hydrophyllus*, *Neurophyllus*, *Penicillium*, *Penicillium*, *Phlebia*, *Ruminocyticum*, *Pleurotus*, *Schizophyllum*, *Basilella*, *Thermophila*, *Fusarium*, *Trametes*, or *Trichoderma*.

[0465] For example, the host cells of filamentous fungi can be *Aspergillus amblymorii*, *Aspergillus sulphureus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus oryzae*, *Bjerkandera adusta*, *Ceriporiopsis saneirina*, *Ceriporiopsis caregiea*, *Ceriporiopsis gilvescens*, *Ceriporiopsis pannocinta*, *Ceriporiopsis rivulosa*, *Ceriporiopsis subrufa*, *Ceriporiopsis subvermispora*, *Chrysosporium inops*, *Chrysosporium lucknowense*, and *Chrysosporium foetida*. merdarium), rent spores, Queensland golden spores (Chrysosporium queenslandicum), tropical golden spores, brown golden spores (Chrysosporium zonatum), gray-capped coprinus (Coprinus cinereus), hairy-skinned spores (Coriolushirsutus), rod-shaped spores (Fusarium), cereal spores (Fusarium), kuweiss spores (Fusarium), large-knife spores (Fusarium), grass spores (Fusarium), red spores (Fusarium), heterospores (Fusarium), albinofuss spores (Fusarium), acuminata spores (Fusarium), multibranched spores (Fusarium), pink spores (Fusarium), elderberry spores (Fusarium), skin-colored spores (Fusarium), pseudo-clastic spores (Fusarium), sulfur spores (Fusarium), round spores (Fusarium), pseudo-filamentous spores (Fusarium), patchy spores (Fusarium), specific humic molds (Fusarium), soft-haired humic molds (Fusarium), rice black molds (Mucor), thermophilic filamentous molds (Fusarium), rough spores (Nephrolepis), purpuric molds (Penicillium), Phanerochaete chrysosporium (Phanerochaete chrysosporium), radiata (Phlebia radiata), Pleurotus eryngii (Pleurotus) eryngii), terrestrial clostridium, Trametes villosa, Trametes versicolor, Trichoderma harzianum, Trichoderma corningensis, Trichoderma longibranchii, Trichoderma reesei, or green Trichoderma cells.

[0466] Fungal cells can be transformed in a manner known per se through methods involving protoplast formation, protoplast transformation, and cell wall reconstruction. Suitable procedures for transforming Aspergillus and Trichoderma host cells are described in EP 238023; Yelton et al., 1984, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA) 81:1470-1474; and Christensen et al., 1988, Bio / Technology 6:1419-1422. Suitable methods for transforming Fusarium species are described by Malardier et al., 1989, Gene 78:147-156, and WO 96 / 00787. Yeast can be transformed using procedures described in the following literature: Becker and Guarente, in Abelson, JN and Simon, MI (eds.), Guide to Yeast Genetics and Molecular Biology, Methods in Enzymology, Vol. 194, pp. 182-187, Academic Press, Inc., New York; Ito et al., 1983, J. Bacteriol. 153:163; and Hinnen et al., 1978, Proceedings of the National Academy of Sciences of the United States of America 75:1920.

[0467] Method of generation

[0468] The present invention also relates to methods for producing the carbohydrate-binding module variant, cellobiase variant, or hybrid polypeptide described herein, the methods comprising: (a) culturing the recombinant host cells of the present invention under conditions suitable for producing the carbohydrate-binding module variant, cellobiase variant, or hybrid polypeptide; and optionally (b) recovering the carbohydrate-binding module variant, cellobiase variant, or hybrid polypeptide.

[0469] These host cells are cultured in nutrient media suitable for generating variants using methods known in the art. For example, cells can be cultured by shake-flask culture in a suitable medium and under conditions that allow for the expression and / or isolation of the carbohydrate-binding module variants, cellobiase variants, or hybrid peptides described herein, or by small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state fermentation) in a laboratory or industrial fermenter. This culture occurs using procedures known in the art in a suitable nutrient medium comprising carbon and nitrogen sources and inorganic salts. Suitable media are available from a supplier or can be prepared according to disclosed compositions (e.g., the catalogue of the U.S. Center for Type Culture Collection). If the variant is secreted into the nutrient medium, it can be recovered directly from the medium. If the variant is not secreted, it can be recovered from cell lysates.

[0470] These variants can be detected using methods specific to them known in the art. These detection methods include, but are not limited to, the use of specific antibodies, the formation of enzyme products, or the disappearance of enzyme substrates. For example, enzyme assays can be used to determine the activity of the variant.

[0471] The variant can be recovered using methods known in the art. For example, the variant can be recovered from the nutrient medium through a variety of routine procedures, including but not limited to collection, centrifugation, filtration, extraction, spray drying, evaporation, or precipitation.

[0472] Variants can be purified to obtain substantially pure variants by a variety of procedures known in the art, including but not limited to: chromatography (e.g., ion exchange chromatography, affinity chromatography, hydrophobic interaction chromatography, chromatographic focusing and size exclusion chromatography), electrophoresis procedures (e.g., preparative isoelectric point focusing), differential solubility (e.g., ammonium sulfate precipitation), SDS-PAGE, or extraction (see, for example, Protein Purification, edited by Janson and Ryden, VCH Publishers, New York, 1989).

[0473] Alternatively, instead of recycling the variant, the host cell of the present invention expressing the variant is used as the source of the variant.

[0474] Fermentation broth preparations or cell compositions

[0475] The present invention also relates to fermentation broth formulations or cell compositions comprising variants of the carbohydrate-binding module, cellobiase variants, or hybrid polypeptides of the present invention. The fermentation broth product further includes additional components used in the fermentation process, such as cells (including host cells containing genes encoding variants of the present invention, which are used to generate variants of interest), cell debris, biomass, fermentation medium, and / or fermentation products. In some embodiments, the composition is a whole culture medium comprising one or more organic acids, killed cells and / or cell debris, and cell-killing culture medium.

[0476] As used herein, the term "fermentation broth" refers to a preparation produced by cell fermentation that undergoes little or no recovery and / or purification. For example, fermentation broth is produced when a microbial culture is incubated to saturation under carbon-limited conditions that allow protein synthesis (e.g., expression by enzymes of the host cell) and secretion of proteins into the cell culture medium. Fermentation broth may contain unfractionated or fractionated contents of the fermentation material obtained at the end of fermentation. Typically, fermentation broth is unfractionated and includes used culture medium and cell debris remaining after, for example, removal of microbial cells (e.g., filamentous fungal cells) by centrifugation. In some embodiments, fermentation broth contains used cell culture medium, extracellular enzymes, and viable and / or non-viable microbial cells.

[0477] In one embodiment, the fermentation broth formulation and cell composition comprise a first organic acid component (comprising at least one 1-5 carbon organic acid and / or its salt) and a second organic acid component (comprising at least one 6- or more carbon organic acid and / or its salt). In a specific embodiment, the first organic acid component is acetic acid, formic acid, propionic acid, its salt, or a mixture of two or more of the foregoing; and the second organic acid component is benzoic acid, cyclohexanecarboxylic acid, 4-methylvaleric acid, phenylacetic acid, its salt, or a mixture of two or more of the foregoing.

[0478] In one aspect, the composition comprises one or more organic acids and optionally further comprises killed cells and / or cell debris. In one embodiment, these killed cells and / or cell debris are removed from the cell-killing whole culture medium to provide a composition free of these components.

[0479] These fermentation broth formulations or cell compositions may further include preservatives and / or antimicrobial (e.g., bacteriostatic) agents, including but not limited to sorbitol, sodium chloride, potassium sorbate, and other agents known in the art.

[0480] The cell-killing whole culture or composition may contain the ungraded contents of the fermentation material obtained at the end of fermentation. Typically, the cell-killing whole culture or composition contains used culture medium and cell debris present after microbial cells (e.g., filamentous fungal cells) have been grown to saturation and incubated under carbon-limited conditions to allow protein synthesis. In some embodiments, the cell-killing whole culture or composition contains used cell culture medium, extracellular enzymes, and killed filamentous fungal cells. In some embodiments, methods known in the art can be used to permeate and / or lyse the microbial cells present in the cell-killing whole culture or composition.

[0481] The whole culture medium or cell composition described herein is typically a liquid, but may contain insoluble components, such as killed cells, cell debris, culture medium components, and / or one or more insoluble enzymes. In some embodiments, insoluble components may be removed to provide a clear liquid composition.

[0482] The whole culture medium formulations and cell compositions of the present invention can be produced by the methods described in WO 90 / 15861 or WO 2010 / 096673.

[0483] Enzyme composition

[0484] The present invention also relates to compositions comprising the carbohydrate-binding module variant, cellobiase variant, or hybrid polypeptide of the present invention. Preferably, these compositions are rich in such a polypeptide. The term "enrichment" indicates that the cellobiase activity or cellulase activity of the composition has increased, for example, an enrichment factor of at least 1.1.

[0485] The composition may include the carbohydrate-binding module variant, cellobiase variant, or hybrid polypeptide of the present invention as the main enzyme component, for example, a single-component composition. Alternatively, these compositions may include a variety of enzymatic activities, such as one or more (e.g., several) enzymes selected from the group consisting of: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as α-galactosidase, α-glucosidase, aminopeptidase, amylase, β-galactosidase, β-glucosidase, β-xylosidase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitosanase, keratinase, cyclodextrin glucosyltransferase, deoxyribonuclease, endoglucanase, esterase, patulin, glucosylamylase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, swelling agent, transglutaminase, or xylanase.

[0486] These compositions can be prepared according to methods known in the art and can be in the form of liquid or dry compositions. These compositions can be stabilized according to methods known in the art.

[0487] Examples of preferred uses of the compositions of the present invention are given below. The dosage of the composition and other conditions for using the composition can be determined based on methods known in the art.

[0488] use

[0489] The present invention also relates to the following methods for using cellobiase variants described herein or heterologous catalytic domains comprising carbohydrate-binding module variants and cellulases described herein, together with the following compositions.

[0490] This invention relates to methods for degrading or converting cellulosic materials, the methods comprising treating the cellulosic material with an enzyme composition in the presence of a cellobiose hydrolase variant comprising a carbohydrate-binding module variant of the present invention or a cellulase. In one aspect, these methods further comprise the recovery of the degraded or converted cellulosic material. The soluble products of the degradation or conversion of the cellulosic material can be separated from the insoluble cellulosic material using methods known in the art, such as centrifugation, filtration, or gravity sedimentation.

[0491] The present invention also relates to a method for producing a fermentation product, the method comprising: (a) saccharifying a cellulose material with an enzyme composition in the presence of a cellobiase variant comprising a carbohydrate-binding module variant of the present invention or a cellulase; (b) fermenting the saccharified cellulose material with one or more (e.g., several) fermenting microorganisms to produce the fermentation product; and (c) recovering the fermentation product from the fermentation.

[0492] The present invention also relates to methods for fermenting cellulosic materials, the methods comprising: fermenting the cellulosic material with one or more (e.g., several) fermenting microorganisms, wherein the cellulosic material is saccharified with an enzyme composition in the presence of a cellobiose hydrolase variant or cellulase comprising the carbohydrate-binding module of the present invention. In one aspect, the fermentation of the cellulosic material produces a fermentation product. In another aspect, the method further comprises recovering the fermentation product from the fermentation.

[0493] The method of the present invention can be used to saccharify cellulosic materials into fermentable sugars and convert the fermentable sugars into a variety of useful fermentation products, such as fuels, drinking ethanol, and / or platform compounds (e.g., acids, alcohols, ketones, gases, etc.). The production of desired fermentation products from the cellulosic material typically involves pretreatment, enzymatic hydrolysis (saccharification), and fermentation.

[0494] According to the present invention, the processing of cellulose materials can be accomplished using processes conventional in the art. Furthermore, the method of the present invention can be implemented using conventional biomass processing equipment configured to operate according to the present invention.

[0495] Separate or simultaneous hydrolysis (saccharification) and fermentation include, but are not limited to: separate hydrolysis and fermentation (SHF), simultaneous saccharification and fermentation (SSF), simultaneous saccharification and co-fermentation (SSCF), hybrid hydrolysis and fermentation (HHF), separate hydrolysis and co-fermentation (SHCF), hybrid hydrolysis and co-fermentation (HHCF), and direct microbial transformation (DMC), sometimes referred to as combined bioprocessing (CBP). SHF uses separate processing steps to first enzymatically hydrolyze cellulosic material into fermentable sugars (e.g., glucose, cellobiose, and pentose monomers), and then ferment the fermentable sugars into ethanol. In SSF, the enzymatic hydrolysis of cellulose material and the fermentation of sugars into ethanol are combined in one step (Philippidis, GP, 1996, Cellulose bioconversion technology, Handbook on Bioethanol: Production and Utilization, edited by Wyman, CE, Taylor & Francis, Washington, DC, 179-212). SSCF involves the co-fermentation of multiple sugars (Sheehan, J. and Himmel, M., 1999, Enzymes, energy and the environment: A strategic perspective on the USDepartment of Energy's research and development activities for bioethanol), Biotechnol. Prog. 15:817-827). HHF involves separate hydrolysis steps and additionally involves simultaneous saccharification and hydrolysis steps that can be carried out in the same reactor. The steps in the HHF process can be carried out at different temperatures, i.e., high-temperature enzymatic saccharification followed by SSF at a lower temperature that the fermentation strain can tolerate.DMC combines all three processes (enzyme production, hydrolysis, and fermentation) into one or more (e.g., several) steps, wherein the same organism is used to produce the enzymes for converting cellulosic material into fermentable sugars and to convert the fermentable sugars into the final product (Lynd, LR; Weimer, PJ; van Zyl, WH; and Pretorius, IS, 2002, Microbial cellulose utilization: Fundamentals and biotechnology, Microbiol.Mol.Biol.Reviews 66:506-577). It should be understood here that any method known in the art, including pretreatment, enzymatic hydrolysis (saccharification), fermentation, or combinations thereof, can be used to implement the methods of this invention.

[0496] Conventional equipment may include fed-batch stirred reactors, continuous flow stirred reactors with ultrafiltration, and / or continuous plug flow column reactors (Fernanda de Castilhos Corazza, Flavio Faria de Moraes, Gisella Maria Zanin, and Ivo Neitzel, 2003, Optimal control in fed-batch reactor for the cellobiose hydrolysis, Acta Scientiarum. Technology 25:33-38; Gusakov AV and Sinitsyn AP, 1985, Enzymatic hydrolysis kinetics of cellulose: 1. Mathematical model of the process in a batch reactor). Enzymatic hydrolysis of cellulose: 1. A mathematical model for a batch reactor process, Enzyme and Microbial Technology (Enz. Microb. Technol.) 7:346-352); Friction reactor (Ryu, SK and Lee, JM, 1983, Bioconversion of waste cellulose by using an attrition bioreactor, Biotechnology and Bioengineering (Biotechnol. Bioeng.) 25:53-65); or reactors with strong stirring caused by electromagnetic fields (Gusakov AV, Sinnetsi AP, Sergei IY (Davydkin, IY), Sergei VY, Protas OV (Protas, OV)(Applied Biochemistry and Biotechnology, 1996, "Enhancement of enzymatic cellulose hydrolysis using a novel type of bioreactor with intensive stirring induced by electromagnetic field," Applied Biochemistry and Biotechnology, 56:141-153). Other reactor types include fluidized beds, blanket reactors, immobilized reactors, and extruder-type reactors for hydrolysis and / or fermentation.

[0497] Preprocessing. In practice, any pretreatment method known in the art can be used to disrupt the cellulose material components of plant cell walls (Chandra et al., 2007, Substrate pretreatment: The key to effective enzymatic hydrolysis of lignocellulosics?, Adv. Biochem. Engin. / Biotechnol. 108:67-93; Galbe and Zacchi, 2007, Pretreatment of lignocellulosic materials for efficient bioethanol production, Adv. Biochem. Engin. / Biotechnol. 108:41-65; Hendriks and Zeeman, 2009, Pretreatments to enhance the digestibility of lignocellulosic biomass, Bioresource). Technol. 100:10-18; Mosier et al., 2005, Features of promising technologies for pretreatment of lignocellulosic biomass, Bioresource Technology 96:673-686; Taherzadeh and Karimi, 2008, Pretreatment of lignocellulosic wastes to improve ethanol and biogas production: A review, International Journal of Molecular Science.(9:1621-1651; Yang and Wyman, 2008, Pretreatment: the key to unlocking low-cost cellulosic ethanol, Biofuels, Bioproducts and Biorefining, Bioofpr. 2:26-40).

[0498] Cellulose materials can also be subjected to particle size reduction, sieving, pre-soaking, wetting, washing and / or conditioning using methods known in the art prior to pretreatment.

[0499] Conventional pretreatment methods include, but are not limited to: steam pretreatment (with or without explosion), dilute acid pretreatment, hot water pretreatment, alkali pretreatment, lime pretreatment, wet oxidation, wet explosion, ammonia fiber explosion, organic solvent pretreatment, and biological pretreatment. Other pretreatment methods include ammonia percolation, ultrasonication, electroporation, microwave treatment, supercritical CO2, supercritical H2O, ozone treatment, ionic liquid treatment, and gamma radiation pretreatment.

[0500] Cellulose materials can be pretreated before hydrolysis and / or fermentation. Pretreatment is preferred before hydrolysis. Alternatively, pretreatment can be carried out simultaneously with enzymatic hydrolysis to release fermentable sugars such as glucose, xylose, and / or cellobiose. In most cases, the pretreatment step itself results in the conversion of biomass into fermentable sugars (even in the absence of enzymes).

[0501] Steam pretreatment. In steam pretreatment, the cellulose material is heated to break down plant cell wall components, including lignin, hemicellulose, and cellulose, making cellulose and other fractions, such as hemicellulose, accessible to enzymes. The cellulose material is passed through or through a reaction vessel, into which steam is injected to increase the temperature to the desired temperature and pressure, and the steam is maintained therein for the desired reaction time. Steam pretreatment is preferably carried out at 140°C to 250°C, for example 160°C to 200°C, or 170°C to 190°C, with the optimal temperature range depending on the addition of a chemical catalyst. The residence time for steam pretreatment is preferably 1–60 minutes, for example 1–30 minutes, 1–20 minutes, 3–12 minutes, or 4–10 minutes, with the optimal residence time depending on the temperature range and the added chemical catalyst. Steam pretreatment allows for relatively high solid loadings, so that the cellulose material typically only becomes moist during pretreatment. Steam pretreatment is often combined with an explosive discharge of the pretreated material, known as a steam explosion, which involves rapid evaporation to atmospheric pressure and turbulence of the material to increase the accessible surface area through breakup (Duff and Murray, 1996, Bioresource Technology 855:1-33; Galbe and Zacchi, 2002, Applied Microbiology and Biotechnol. 59:618-628; US Patent Application No. 20020164730). During steam pretreatment, hemicellulose acetyl groups are cleaved, and the resulting acid autocatalytically hydrolyzes the hemicellulose into monosaccharides and oligosaccharides. Lignin is removed only to a limited extent.

[0502] Chemical pretreatment: The term "chemical treatment" refers to any chemical pretreatment that promotes the separation and / or release of cellulose, hemicellulose, and / or lignin. Such pretreatment can convert crystalline cellulose into amorphous cellulose. Examples of suitable chemical pretreatment processes include, for example, dilute acid pretreatment, lime pretreatment, wet oxidation, ammonia cellulose / freeze-explosion (AFEX), ammonia percolation (APR), ionic liquids, and organic solvent pretreatment.

[0503] A catalyst (e.g., H₂SO₄ or SO₂) (typically 0.3% to 5% w / w) is often added prior to steam pretreatment. This catalyst reduces time and temperature, increases recovery, and improves enzymatic hydrolysis (Ballesteros et al., 2006, Applied Biochemistry and Biotechnology 129-132:496-508; Varga et al., 2004, Applied Biochemistry and Biotechnology 113-116:509-523; Sassner et al., 2006, Enzyme Microb. Technol. 39:756-762). In dilute acid pretreatment, cellulose material is mixed with dilute acid (typically H₂SO₄) and water to form a slurry, heated by steam to the desired temperature, and flashed to atmospheric pressure after a residence time. Many reactor designs can be used for dilute acid pretreatment, such as plug flow reactors, countercurrent reactors, or continuous countercurrent shrink-bed reactors (Duff and Murray, 1996, see above; Schell et al., 2004, Bioresource Technology 91:179-188; Lee et al., 1999, Advances in Biochemical Engineering and Biotechnology 65:93-115).

[0504] Several pretreatment methods under alkaline conditions can also be used. These alkaline pretreatments include, but are not limited to, sodium hydroxide, lime, wet oxidation, ammonia percolation (APR), and ammonia fiber / freeze-explosion (AFEX).

[0505] Lime pretreatment using calcium oxide or calcium hydroxide at temperatures ranging from 85°C to 150°C, with residence times ranging from one hour to several days (Wyman et al., 2005, Bioresource Technol. 96:1959-1966; Mosier et al., 2005, Bioresource Technol. 96:673-686). Pretreatment methods using ammonia are disclosed in WO 2006 / 110891, WO 2006 / 110899, WO 2006 / 110900, and WO 2006 / 110901.

[0506] Wet oxidation is a thermal pretreatment typically carried out at 180°C to 200°C for 5–15 minutes with the addition of an oxidant (such as oxygen peroxide or superpressure oxygen) (Schmidt and Thomsen, 1998, Bioresource Technology 64:139-151; Palonen et al., 2004, Applied Biochemistry and Biotechnology 117:1-17; Varga et al., 2004, Biotechnology and Bioengineering (Biotechnol. Bioeng.) 88:567-574; Martin et al., 2006, Journal of Chemical Technology and Biotechnology (J. Chem. Technol. Biotechnol.) 81:1669-1677). Pretreatment is preferably carried out with 1%–40% dry matter, for example 2%–30% or 5%–20% dry matter, and the initial pH often increases due to the addition of a base such as sodium carbonate.

[0507] A modified version of the wet oxidation pretreatment method, known as wet explosion (a combination of wet oxidation and steam explosion), can handle up to 30% dry matter. In wet explosion, an oxidant is introduced during the pretreatment process after a certain residence time. The pretreatment is then terminated by rapid evaporation to atmospheric pressure (WO 2006 / 032282).

[0508] Ammonia Fiber Explosion (AFEX) involves treating cellulosic materials with liquid or gaseous ammonia for 5-10 minutes at moderate temperatures such as 90°C-150°C and high pressures such as 17-20 bar, where the dry matter content can be as high as 60% (Gollapalli et al., 2002, Applied Biochemistry and Biotechnology, 98:23-35; Chundawat et al., 2007, Biotechnology and Bioengineering, 96:219-231; Alizadeh et al., 2005, Applied Biochemistry and Biotechnology, 121:1133-1141; Teymouri et al., 2005, Bioresource Technol., 96:2014-2018). During AFEX pretreatment, cellulose and hemicellulose remain relatively intact. The lignin-carbohydrate complex is cleaved.

[0509] Organic solvent pretreatment removes lignin from cellulose materials by extraction with aqueous ethanol (40%-60% ethanol) at 160-200°C for 30-60 minutes (Pan et al., 2005, Biotechnology and Bioengineering, 90:473-481; Pan et al., 2006, Biotechnology and Bioengineering, 94:851-861; Kurabi et al., 2005, Applied Biochemistry and Biotechnology, 121:219-230). Sulfuric acid is typically added as a catalyst. During organic solvent pretreatment, most of the hemicellulose and lignin are removed.

[0510] Other examples of suitable pretreatment methods are described by Schell et al., 2003, Appl. Biochem. and Biotechnol., vols. 105-108, pp. 69-85; Mosier et al., 2005, Bioresource Technology, 96:673-686; and U.S. Publication 2002 / 0164730.

[0511] In one aspect, the chemical pretreatment is preferably carried out as a dilute acid treatment, and more preferably as a continuous dilute acid treatment. The acid is typically sulfuric acid, but other acids such as acetic acid, citric acid, nitric acid, phosphoric acid, tartaric acid, succinic acid, hydrogen chloride, or mixtures thereof may also be used. The weak acid treatment is preferably carried out in a pH range of 1 to 5, for example, 1 to 4 or 1 to 2.5. In another aspect, the acid concentration is preferably in the range of 0.01 wt% to 10 wt% acid, for example, 0.05 wt% to 5 wt% acid or 0.1 wt% to 2 wt% acid. The acid is brought into contact with the cellulose material and maintained at a temperature preferably in the range of 140°C to 200°C, for example, 165°C to 190°C, for a period ranging from 1 to 60 minutes.

[0512] In another aspect, pretreatment occurs in the aqueous slurry. In a preferred aspect, the cellulose material is present during pretreatment in an amount preferably between 10 wt% and 80 wt%, for example, 20 wt% to 70 wt% or 30 wt% to 60 wt%, such as about 40 wt%. The pretreated cellulose material may be left unwashed or washed using any method known in the art, for example, washing with water.

[0513] Mechanical or physical pretreatment: The terms “mechanical pretreatment” or “physical pretreatment” refer to any pretreatment that promotes a reduction in particle size. For example, such pretreatment can involve different types of grinding or milling (e.g., dry grinding, wet grinding, or vibratory ball milling).

[0514] Cellulose materials can be pretreated physically (mechanically) and chemically. Mechanical or physical pretreatment can be combined with steam / steam explosion, hydrothermolysis, dilute or weak acid treatment, high temperature, high pressure treatment, radiation (e.g., microwave radiation), or combinations thereof. On one hand, high pressure means a pressure preferably in the range of about 100 to about 400 psi, for example, about 150 to about 250 psi. On the other hand, high temperature means a temperature in the range of about 100°C to about 300°C, for example, about 140°C to about 200°C. In a preferred aspect, mechanical or physical pretreatment is carried out in batches using a steam gun hydrolyzer system, such as the Sunds Hydrolyzer available in Sweden from Sunds Defibrator AB, which uses the high pressure and high temperature as defined above. These physical and chemical pretreatments can be performed sequentially or simultaneously as needed.

[0515] Therefore, in a preferred aspect, the cellulose material is subjected to physical (mechanical) or chemical pretreatment, or any combination thereof, to promote the separation and / or release of cellulose, hemicellulose, and / or lignin.

[0516] Biological pretreatment: The term “biological pretreatment” refers to any biological pretreatment that promotes the separation and / or release of cellulose, hemicellulose, and / or lignin from cellulosic materials. Biological pretreatment techniques can involve the application of lignin-dissolving microorganisms and / or enzymes (see, for example, Hsu, T.-A., 1996, Pretreatment of biomass, Handbook on Bioethanol: Production and Utilization, edited by Wyman CE, Taylor-Francis Publishing Group, Washington, D.C., 179-212; Ghosh and Singh, 1993, Physicochemical and biological treatments for enzymatic / microbial conversion of cellulosic biomass, Adv. Appl. Microbiol. 39:295-333; McMillan, JD, 1994, Pretreating lignocellulosic biomass: a review). Review), Enzymatic Conversion of Biomass for Fuels Production, edited by Himmel ME, Baker JO, and Overend RP, ACS Symposium Series 566, American Chemical Society, Washington, D.C., Chapter 15; Gong CS, Cao NJ, Du J., and Tsao GT, 1999, Ethanol production from renewable resources, Advances in Biochemical Engineering / Biotechnology, Schepper T.Edited by Springer Publishers, Berlin Heidelberg, Germany, 65:207-241; Olsson and Hahn-Hagerdal, 1996, Fermentation of lignocellulosic hydrolysates for ethanol production, Enzyme and Microbial Technology, 18:312-331; and Vallander and Eriksson, 1990, Production of ethanol from lignocellulosic materials: State of the art, Advances in Biochemical Engineering / Biotechnology, 42:63-95.

[0517] Glycation. In the hydrolysis step (also known as saccharification), the cellulose material (e.g., pretreated) is hydrolyzed to break down cellulose and / or hemicellulose into fermentable sugars such as glucose, cellobiose, xylose, xylulose, arabinose, mannose, galactose, and / or soluble oligosaccharides. Hydrolysis is enzymatically propelled by an enzyme composition in the presence of a cellobiose hydrolase variant, including a carbohydrate-binding module variant of the present invention, or a cellulase. The enzymes in these compositions may be added simultaneously or sequentially.

[0518] Enzymatic hydrolysis is preferably performed in a suitable aqueous environment under conditions readily determined by those skilled in the art. In one aspect, hydrolysis is carried out under conditions suitable for the activity of one or more enzymes, i.e., optimal for those enzymes. Hydrolysis can be carried out as a batch or continuous process, wherein cellulose material is gradually fed into, for example, an enzyme-containing hydrolysis solution.

[0519] Saccharification is typically carried out in a stirred tank reactor or fermenter under controlled pH, temperature, and mixing conditions. Suitable processing times, temperatures, and pH conditions can be readily determined by those skilled in the art. For example, saccharification can last up to 200 hours, but is typically carried out preferably for about 12 to about 120 hours, for example, about 16 to about 72 hours or about 24 to about 48 hours. Temperatures are preferably in the range of about 25°C to about 70°C, for example, about 30°C to about 65°C, about 40°C to about 60°C, or about 50°C to 55°C. pH is preferably in the range of about 3 to about 8, for example, about 3.5 to about 7, about 4 to about 6, or about pH 5.0 to about pH 5.5. The dry solids content is preferably in the range of about 5 wt% to about 50 wt%, for example, about 10 wt% to about 40 wt% or about 20 wt% to about 30 wt%.

[0520] These enzyme compositions can include any protein used to degrade cellulose materials.

[0521] In one aspect, the enzyme composition comprises or further comprises one or more proteins selected from the group consisting of: cellulase, GH61 polypeptide with enhanced cellulose-degrading activity, hemicellulase, esterase, patulin, laccase, lignin-degrading enzyme, pectinase, peroxidase, protease, and swelling agent. In another aspect, the cellulase is preferably one or more enzymes selected from the group consisting of: endoglucanase, cellobiase, and β-glucosidase. In yet another aspect, the hemicellulase is preferably one or more enzymes selected from the group consisting of: acetylmannan esterase, acetylxylan esterase, arabinonanase, arabinofuranylase, coumarin esterase, ferulic acid esterase, galactosidase, glucuronidase, glucuronidase, mannanase, mannosidase, xylanase, and xylosidase.

[0522] In another aspect, the enzyme composition comprises one or more (e.g., several) cellulases. In another aspect, the enzyme composition comprises or further comprises one or more (e.g., several) hemicellulases. In another aspect, the enzyme composition comprises one or more (e.g., several) cellulases and one or more (e.g., several) hemicellulases. In another aspect, the enzyme composition comprises one or more (e.g., several) enzymes selected from the group consisting of cellulases and hemicellulases. In another aspect, the enzyme composition comprises an endoglucanase. In another aspect, the enzyme composition comprises a cellobiose hydrolase. In another aspect, the enzyme composition comprises a β-glucosidase. In another aspect, the enzyme composition comprises a polypeptide having enhanced cellulolytic activity. In another aspect, the enzyme composition comprises an endoglucanase and a polypeptide having enhanced cellulolytic activity. In another aspect, the enzyme composition comprises a cellobiose hydrolase and a polypeptide having enhanced cellulolytic activity. In another aspect, the enzyme composition comprises a β-glucosidase and a polypeptide having enhanced cellulolytic activity. In another aspect, the enzyme composition comprises an endoglucanase and a cellobiase. In another aspect, the enzyme composition comprises an endoglucanase and a β-glucosidase. In another aspect, the enzyme composition comprises a cellobiase and a β-glucosidase. In another aspect, the enzyme composition comprises an endoglucanase, a cellobiase, and a polypeptide with enhanced cellulose-degrading activity. In another aspect, the enzyme composition comprises an endoglucanase, a β-glucosidase, and a polypeptide with enhanced cellulose-degrading activity. In another aspect, the enzyme composition comprises a cellobiase, a β-glucosidase, and a polypeptide with enhanced cellulose-degrading activity. In another aspect, the enzyme composition comprises an endoglucanase, a cellobiase, and a β-glucosidase. In another aspect, the enzyme composition comprises an endoglucanase, a cellobiase, a β-glucosidase, and a polypeptide with enhanced cellulose-degrading activity.

[0523] In another aspect, the enzyme composition includes acetylmannan esterase. In another aspect, the enzyme composition includes acetylxylan esterase. In another aspect, the enzyme composition includes arabinogalactanase (e.g., α-L-arabinogalactanase). In another aspect, the enzyme composition includes arabinofuranase (e.g., α-L-arabinofuranase). In another aspect, the enzyme composition includes coumarate esterase. In another aspect, the enzyme composition includes ferulic acid esterase. In another aspect, the enzyme composition includes galactosidase (e.g., α-galactosidase and / or β-galactosidase). In another aspect, the enzyme composition includes glucuronidase (e.g., α-D-glucuronidase). In another aspect, the enzyme composition includes glucuronidase. In another aspect, the enzyme composition includes mannanase. In another aspect, the enzyme composition includes mannosidase (e.g., β-mannosidase). In another aspect, the enzyme composition includes xylanase. In a preferred aspect, the xylanase is a family 10 xylanase. On the other hand, the enzyme composition includes xylosidase (e.g., β-xylosidase).

[0524] In another aspect, the enzyme composition includes an esterase. In another aspect, the enzyme composition includes patulin. In another aspect, the enzyme composition includes a laccase. In another aspect, the enzyme composition contains a lignin-degrading enzyme. In a preferred aspect, the lignin-degrading enzyme is a manganese peroxidase. In another preferred aspect, the lignin-degrading enzyme is a lignin peroxidase. In another preferred aspect, the lignin-degrading enzyme is an H₂O₂-producing enzyme. In another aspect, the enzyme composition includes a pectinase. In another aspect, the enzyme composition includes a peroxidase. In another aspect, the enzyme composition includes a protease. In another aspect, the enzyme composition includes a swelling agent.

[0525] In the method of the present invention, one or more enzymes may be added before or during saccharification, saccharification and fermentation, or fermentation.

[0526] One or more (e.g., several) components of the enzyme composition may be wild-type proteins, recombinant proteins, or a combination of wild-type and recombinant proteins. For example, one or more (e.g., several) components may be a natural protein of a cell used as a host cell to recombinantly express one or more (e.g., several) other components of the enzyme composition. One or more (e.g., several) components of the enzyme composition may be produced as single components and then combined to form the enzyme composition. The enzyme composition may be a combination of multi-component and single-component protein formulations.

[0527] The enzymes used in the methods of this invention can be in any suitable form, such as fermentation broth formulations or cell compositions, cell lysates with or without cell debris, semi-purified or purified enzyme preparations, or host cells as the source of the enzyme. The enzyme composition can be a dry powder or granules, dust-free granules, liquid, stabilized liquid, or stabilized protected enzyme. Liquid enzyme preparations can be stabilized according to established methods, for example, by adding a stabilizer (such as a sugar, sugar alcohol, or other polyol), and / or lactic acid or another organic acid.

[0528] The optimal amount of enzymes, including carbohydrate-binding module variants and cellobiase variants or cellulases, depends on several factors, including but not limited to: the mixture of cellulase and / or hemicellulose-degrading enzyme components, the cellulose material, the concentration of the cellulose material, one or more pretreatments of the cellulose material, temperature, time, pH, and the fermentation organisms involved (e.g., yeast used for simultaneous saccharification and fermentation).

[0529] On the one hand, the effective amount of cellulase or hemicellulose-degrading enzyme for cellulosic materials is about 0.5 to about 50 mg, for example about 0.5 to about 40 mg, about 0.5 to about 25 mg, about 0.75 to about 20 mg, about 0.75 to about 15 mg, about 0.5 to about 10 mg, or about 2.5 to about 10 mg / g of cellulosic material.

[0530] In another preferred aspect, the effective amount of cellobiose hydrolase variants or cellulases, including carbohydrate-binding module variants, on cellulose material is about 0.01 to about 50.0 mg, preferably about 0.01 to about 40 mg, more preferably about 0.01 to about 30 mg, more preferably about 0.01 to about 20 mg, more preferably about 0.01 to about 10 mg, more preferably about 0.01 to about 5 mg, more preferably about 0.025 to about 1.5 mg, more preferably about 0.05 to about 1.25 mg, more preferably about 0.075 to about 1.25 mg, more preferably about 0.1 to about 1.25 mg, even more preferably about 0.15 to about 1.25 mg, and most preferably about 0.25 to about 1.0 mg / g of cellulose material.

[0531] In another preferred aspect, the effective amount of cellobiase variant or cellulase comprising the carbohydrate-binding module to cellulase or hemicellulase is about 0.005 to about 1.0 g, for example about 0.01 to about 1.0 g, about 0.15 to about 0.75 g, about 0.15 to about 0.5 g, about 0.1 to about 0.5 g, about 0.1 to about 0.25 g, or about 0.05 to about 0.2 g / g of cellulase or hemicellulase.

[0532] Polypeptides with cellulase or hemicellulase activity, as well as other proteins / peptides suitable for the degradation of cellulosic materials, such as the GH61 polypeptide with cellulase-enhancing activity (collectively referred to below as "enzyme-active polypeptides"), can be derived from or obtained from any suitable source, including bacterial, fungal, yeast, plant, or mammalian sources. The term "obtained" herein also means that the enzyme can be recombinantly produced in a host organism using the methods described herein, wherein the recombinantly produced enzyme is native or heterologous to the host organism, or has a modified amino acid sequence, for example, having one or more (e.g., several) deleted, inserted, and / or substituted amino acids; that is, the recombinantly produced enzyme is a mutant and / or fragment of the native amino acid sequence or an enzyme produced by amino acid rearrangement methods known in the art. The meaning of native enzyme includes natural variants, while the meaning of exogenous enzyme includes variants obtained through recombinant methods (such as by site-directed mutagenesis or tampering).

[0533] The enzymatic polypeptide can be a bacterial polypeptide. For example, the polypeptide can be a Gram-positive bacterial polypeptide with enzymatic activity, such as Bacillus, Streptococcus, Streptomyces, Staphylococcus, Enterococcus, Lactobacillus, Lactococcus, Clostridium, Bacillus aeruginosa, Pyrolytic Cellulobacillus, Thermobifidia, or Marine Bacillus polypeptide, or a Gram-negative bacterial polypeptide with enzymatic activity, such as Escherichia coli, Pseudomonas, Salmonella, Campylobacter, Helicobacter, Flavobacterium, Fusobacterium, Coleobacterium, Neisseria, or Ureaplasma polypeptide.

[0534] On one hand, the polypeptide is an enzymatically active Bacillus alkaliphilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus croceus, Bacillus coagulans, Bacillus sclerosus, Bacillus splenium, Bacillus stenosis, Bacillus licheniformis, Bacillus megaterium, Bacillus brevis, Bacillus thermophilus, Bacillus subtilis, or Bacillus thuringiensis polypeptide.

[0535] On the other hand, the polypeptide is an enzymatic polypeptide of Streptococcus equi, Streptococcus pyogenes, Streptococcus lactis, or Streptococcus equi subsp. veterinary.

[0536] On the other hand, the polypeptide is an enzymatic polypeptide from non-chromogenic Streptomyces, insecticidal Streptomyces, sky blue Streptomyces, gray Streptomyces, or light blue-purple Streptomyces.

[0537] The enzymatically active polypeptide can also be a fungal polypeptide, and more preferably a yeast polypeptide, such as an enzymatically active polypeptide from the genera Candida, Kluyveromyces, Pichia, Saccharomyces, Schizosaccharomyces, or Yarrowia.Or more preferably, it is a filamentous fungal polypeptide, such as those from the genera *Acremonium*, *Agaricus*, *Alternaria*, *Aspergillus*, *Aureobasidium*, *Botryospaeria*, *Ceriporiopsis*, *Chaetomidium*, *Chrysosporium*, *Claviceps*, *Cochliobolus*, and *Coprinopsis*, which possess enzymatic activity. * *Coptotermes*, *Corynascus*, *Cryphonectria*, *Cryptococcus*, *Diplodia*, *Exidia*, *Filibasidium*, *Fusarium*, *Gibberella*, *Holomastigotoides*, *Humicola*, *Irpex*, *Lentinula*, *Leptospaeria* The genera *Magnaporthe*, *Melanocarpus*, *Meripilus*, *Mucor*, *Myceliophthora*, *Neocallimastix*, *Neurospora*, *Paecilomyces*, *Penicillium*, *Phanerochaete*, *Piromyces*, *Poitrasia*, *Pseudoplectania*, and *P.* (seudotrichonympha), Rhizomucor, Schizophyllum, Scytalidium, Talaromyces, Thermoascus, Thievaria, Tolypocladium, Trichoderma, Trichophaea, Verticillium, Volvariella, or Xylaria polypeptides.

[0538] On the one hand, polypeptides are enzymatic polypeptides from Kelp yeast, Saccharomyces cerevisiae, Saccharomyces sacchariformis, Saccharomyces douglas, Saccharomyces crocinus, Nodizymes, or Oval yeast.

[0539] On the other hand, the polypeptide is an enzymatically active fungicide containing *Aquilaria sinensis*, *Aspergillus echinosporum*, *Aspergillus amblyceae*, *Aspergillus fumigatus*, *Aspergillus sulphureus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus oryzae*, *Aureobasidium perfringens*, *Aureobasidium laurylodes*, *Aureobasidium tropicalis*, *Aureobasidium fecalithii*, *Aureobasidium stenoptera*, *Aureobasidium stenoptera*, *Aureobasidium monnieri*, *Aureobasidium brevicornum*, *Fusarium moniliforme*, *Fusarium graminearum ... Fusarium sulfide, Fusarium rotundum, Fusarium pseudofilariae, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Fusarium variegatum, Trichoderma harzianum, Trichoderma koningii, Trichoderma longifolium, Trichoderma reesei, Trichoderma viride polypeptide, or Fusarium variegatum longifolium polypeptide.

[0540] Chemically modified or protein-engineered mutants of peptides with enzymatic activity can also be used.

[0541] One or more (e.g., several) components of the enzyme composition can be recombinant components, i.e., produced by cloning the DNA sequence encoding the single component and subsequently transforming cells with that DNA sequence and expressing it in a host (see, for example, WO 91 / 17243 and WO 91 / 17244). The preferred host is a heterologous host (the enzyme is heterologous to the host), but under certain conditions, the host can also be a homologous host (the enzyme is native to the host). A single-component cellulose-degrading protein can also be prepared by purifying such a protein from fermentation broth.

[0542] In one aspect, the one or more (e.g., several) cellulases include commercial cellulase formulations. Examples of commercial cellulase formulations suitable for use in this invention include: (e.g.) CTec (Novozymes A / S) CTec2 (Novozymes A / S) CTec3 (Novozymes A / S), CELLUCLAST TM (Novozymes A / S Company), NOVOZYM TM188 (Novozymes A / S), CELLUZYME TM (Novozymes A / S Company), CEREFLO TM (Novozymes A / S Company, and ULTRAFLO) TM (Novozymes A / S), ACCELERASE TM Genencor International, LAMINEX TM (Genetronics International), SPEZYME TM CP (Genetronics International), NL (DSM); S / L 100 (DSM), ROHAMENT TM 7069 W (ROHM Corporation) GmbH) LDI (Dyadic International, Inc.) LBR (Dyadic International Limited), or 150L (Dyadic International Ltd). Add cellulase in an effective amount from about 0.001 to about 5.0 wt% solids, for example from about 0.025 to about 4.0 wt% solids or from about 0.005 to about 2.0 wt% solids.

[0543] Examples of bacterial endoglucanases that can be used in the methods of the present invention include, but are not limited to: Acidothermus cellulolyticus endoglucanase (WO 91 / 05039; WO 93 / 15186; US Patent No. 5,275,944; WO 96 / 02551; US ​​Patent No. 5,536,655; WO 00 / 70031; WO 05 / 093050); Thermobifida fusca endoglucanase III (WO 05 / 093050); and Thermobifida fusca endoglucanase V (WO 05 / 093050).

[0544] Examples of fungal endoglucanases that can be used in this invention include, but are not limited to: *Trichoderma reesei* endoglucanase I (Penttila et al., 1986, Gene 45:253-263), *Trichoderma reesei* Cel7B endoglucanase I (GENBANK). TMAccession number M15665); Trichoderma reesei endoglucanase II (Saloheimo et al., 1988, gene 63:11-22), Trichoderma reesei Cel5A endoglucanase II (GENBANK) TM Accession number M19373); Trichoderma reesei endoglucanase III (Okada et al., 1988, Applied and Environmental Microbiology, 64:555-563, GENBANK) TM Accession number AB003694); Trichoderma reesei endoglucanase V (Salohemo et al., 1994, Molecular Microbiology 13:219-228, GENBANK) TM Accession number Z33381); Aspergillus echinosporum endoglucanase (Ooi et al., 1990, Nucleic Acids Research 18:5884); Aspergillus kawachii endoglucanase (Sakamoto et al., 1995, Current Genetics 27:435-439); Erwinia carotovara endoglucanase (Saarilahti et al., 1990, Gene 90:9-14); Fusarium oxysporum endoglucanase (GENBANK) TM Accession number L29381); Endoglucanase of *Gnaphalium affine* high-temperature variant (GENBANK) TM Accession number AB003107); endoglucanase from Melanocarpus albomyces (GENBANK) TM Accession number MAL515703); Streptococcus roughus endoglucanase (GENBANK) TMAccession number XM_324477); *Cladorrhinum foecundissimum* ATCC 62373 endoglucanase V; *Cladorrhinum foecundissimum* CBS117.65 endoglucanase; *Cladorrhinum foecundissimum* CBS 495.95 endoglucanase; *Cladorrhinum foecundissimum* CBS494.95 endoglucanase; *Cladorrhinum foecundissimum* NRRL 8126 CEL6B endoglucanase; *Cladorrhinum foecundissimum* NRRL 8126 CEL6C endoglucanase; *Cladorrhinum foecundissimum* NRRL 8126 CEL7C endoglucanase; *Cladorrhinum foecundissimum* ATCC 62373 CEL7A endoglucanase; and Trichoderma reesei strain VTT-D-80133 endoglucanase (GENBANK) TM Login number M15665).

[0545] Examples of cellobiose hydrolases that can be used in this invention include, but are not limited to: Aspergillus echinosporus cellobiose hydrolase II (WO 2011 / 059740), Chaetomium thermophilum cellobiose hydrolase I, Chaetomium thermophilum cellobiose hydrolase II, Specific humic mold cellobiose hydrolase I, Thermophilus thermophilum cellobiose hydrolase II (WO 2009 / 042871), Clostridium hyrcanie cellobiose hydrolase II (WO 2010 / 141325), Clostridium terrestris cellobiose hydrolase II (CEL6A, WO 2006 / 074435), Trichoderma reesei cellobiose hydrolase I, Trichoderma reesei cellobiose hydrolase II, and Pterygomycetes var. chrysospora cellobiose hydrolase II (WO 2010 / 057086).

[0546] Examples of β-glucosidases that can be used in this invention include, but are not limited to, β-glucosidases from the following: Aspergillus echinosporum (Kawaguchi et al., 1996, gene 173:287-288), Aspergillus fumigatus (WO 2005 / 047499), Aspergillus niger (Dan et al., 2000, J. Biol. Chem. 275:4973-4980), Aspergillus oryzae (WO 2002 / 095014), Penicillium brasiliensis IBT 20888 (WO 2007 / 019442 and WO 2010 / 088387), Clostridium perfringens (WO2011 / 035029), and Pterocarya spp. (WO 2007 / 019442).

[0547] The β-glucosidase may be a fusion protein. In one aspect, the β-glucosidase is either the Aspergillus oryzae β-glucosidase variant BG fusion protein (WO 2008 / 057637) or the Aspergillus oryzae β-glucosidase fusion protein (WO 2008 / 057637).

[0548] Other useful endoglucanases, cellobiases, and β-glucosidases are disclosed in many families of glycosyl hydrolases classified according to: Henrissat B., 1991, A classification of glycosyl hydrolases based on amino-acid sequence similarities, Biochem.J. 280:309-316; and Henrissat B. and Bairoch A., 1996, Updating the sequence-based classification of glycosyl hydrolases, Biochem.J. 316:695-696.

[0549] Other cellulases that can be used in this invention are described in WO 98 / 13465, WO 98 / 015619, WO98 / 015633, WO 99 / 06574, WO 99 / 10481, WO 99 / 025847, WO 99 / 031255, WO 2002 / 101078, WO2003 / 027306, WO 2003 / 052054, WO 2003 / 052055, WO 2003 / 052056, WO 2003 / 052057, WO2003 / 052118, WO 2004 / 016760, WO 2004 / 043980, WO 2004 / 048592, WO Among them are WO2005 / 001065, WO2005 / 028636, WO2005 / 093050, WO2005 / 093073, WO2006 / 074005, WO2006 / 117432, WO2007 / 071818, WO2007 / 071820, WO2008 / 008070, WO2008 / 008793, U.S. Patent No. 5,457,046, U.S. Patent No. 5,648,263, and U.S. Patent No. 5,686,593.

[0550] In the method of the present invention, any GH61 polypeptide with cellulolytic-enhancing activity can be used.

[0551] On one hand, the GH61 polypeptide with enhanced cellulolytic activity includes the following motif:

[0552] [ILMV]-PX(4,5)-GXY-[ILMV]-XRX-[EQ]-X(4)-[HNQ](SEQ ID NO:27 or SEQ ID NO:28) and [FW]-[TF]-K-[AIV],

[0553] Where X is any amino acid, X(4,5) is any amino acid in 4 or 5 consecutive positions, and X(4) is any amino acid in 4 consecutive positions.

[0554] Peptides that include the motifs mentioned above may further include:

[0555] HX(1,2)-GPX(3)-[YW]-[AILMV](SEQ ID NO:29 or SEQ ID NO:30),

[0556] [EQ]-XYX(2)-CX-[EHQN]-[FILV]-X-[ILV](SEQ ID NO:31), or

[0557] HX(1,2)-GPX(3)-[YW]-[AILMV](SEQ ID NO:32 or SEQ ID NO:33) and [EQ]-XYX(2)-CX-[EHQN]-[FILV]-X-[ILV](SEQ ID NO:34),

[0558] Where X represents any amino acid, X(1,2) represents any amino acid in one or two consecutive positions, X(3) represents any amino acid in three consecutive positions, and X(2) represents any amino acid in two consecutive positions. In the above motifs, acceptable IUPAC single-letter amino acid abbreviations are used.

[0559] In one preferred aspect, the GH61 polypeptide with enhanced cellulolytic activity further comprises HX(1,2)-GPX(3)-[YW]-[AILMV] (SEQ ID NO:29 or SEQ ID NO:30). In another preferred aspect, the isolated GH61 polypeptide with enhanced cellulolytic activity further comprises [EQ]-XYX(2)-CX-[EHQN]-[FILV]-X-[ILV] (SEQ ID NO:31). In yet another preferred aspect, the GH61 polypeptide with enhanced cellulolytic activity further comprises HX(1,2)-GPX(3)-[YW]-[AILMV] (SEQ ID NO:32 or SEQ ID NO:33) and [EQ]-XYX(2)-CX-[EHQN]-[FILV]-X-[ILV] (SEQ ID NO:34).

[0560] In the second aspect, the GH61 polypeptide with enhanced cellulose-degrading activity has the following motif:

[0561] [ILMV]-Px(4,5)-GxY-[ILMV]-xRx-[EQ]-x(3)-A-[HNQ](SEQ ID NO:35 or SEQ ID NO:36),

[0562] Where x represents any amino acid, x(4,5) represents any amino acid in 4 or 5 consecutive positions, and x(3) represents any amino acid in 3 consecutive positions. The recognized IUPAC single-letter amino acid abbreviations are used in the above motifs.

[0563] Examples of GH61 polypeptides with cellulose-degrading-enhancing activity suitable for the methods of the present invention include, but are not limited to, GH61 polypeptides from the following: *Clostridium perfringens* (WO 2005 / 074647, WO 2008 / 148131, and WO2011 / 035027), *Thermophilic Ascomycota* (WO 2005 / 074656 and WO 2010 / 065830), *Trichoderma reesei* (WO2007 / 089290), *Thermophilic Aspergillus* (WO 2009 / 085935, WO 2009 / 085859, WO 2009 / 085864, WO 2009 / 085868), and *Aspergillus fumigatus* (WO 2010 / 138754); GH61 polypeptides from the following: *Penicillium pineophilum* (WO 2011 / 005867), species of the genus *Thermophilus* (WO 2011 / 039319), species of the genus *Penicillium* (WO 2011 / 041397), and *Thermophilus scabra* (WO 2011 / 041504).

[0564] On the one hand, as described in WO 2008 / 151043, the polypeptide GH61, which has cellulolytic-enhancing activity, can be used in the presence of soluble activated divalent metal cations, such as manganese sulfate.

[0565] On the other hand, the GH61 polypeptide with enhanced cellulose-degrading activity is used in the presence of dioxygen compounds, bicyclic compounds, heterocyclic compounds, nitrogen-containing compounds, quinone compounds, sulfur-containing compounds, or a liquid obtained from pretreated cellulose material (e.g., pretreated corn stalks (PCS)).

[0566] Dioxy compounds can include any suitable compound containing two or more oxygen atoms. In some aspects, a dioxy compound comprises a substituted aryl moiety as described herein. A dioxy compound can include one or more (e.g., several) hydroxyl groups and / or hydroxyl derivatives, and also includes a substituted aryl moiety lacking hydroxyl groups and hydroxyl derivatives. Non-limiting examples of dioxy compounds include catechol or catechin; caffeic acid; 3,4-dihydroxybenzoic acid; 4-tert-butyl-5-methoxy-1,2-benzenediol; pyrogallol; gallic acid; methyl 3,4,5-trihydroxybenzoate; 2,3,4-trihydroxybenzophenone; 2,6-dimethoxyphenol; sinapic acid; 3,5-dihydroxybenzoic acid; 4-chloro-1,2-benzenediol; 4-nitro-1,2-benzenediol; tannic acid; ethyl gallate; methyl glycolate; dihydroxybenzo ... Hydroxyfumaric acid; 2-butyn-1,4-diol; ketone acid; 1,3-propanediol; tartaric acid; 2,4-pentanediol; 3-ethoxy-1,2-propanediol; 2,4,4'-trihydroxybenzophenone; cis-2-buten-1,4-diol; 3,4-dihydroxy-3-cyclobuten-1,2-dione; dihydroxyacetone; acrolein acetal; methyl 4-hydroxybenzoate; 4-hydroxybenzoic acid; and methyl 3,5-dimethoxy-4-hydroxybenzoate; or their salts or solvates.

[0567] Bicyclic compounds may comprise any fused-ring system suitable for substitution as described herein. These compounds may include one or more (e.g., several) additional rings, and are not limited to a specific number of rings unless otherwise stated. In one aspect, the bicyclic compound is a flavonoid. In another aspect, the bicyclic compound is an optionally substituted isoflavone. In yet another aspect, the bicyclic compound is an optionally substituted anthocyanin, such as an optionally substituted anthocyanin or an optionally substituted anthocyanin glycoside, or a derivative thereof. Non-limiting examples of bicyclic compounds include epicatechin; quercetin; myricetin; taxane; calciferol; morin; robinin; naringenin; isorhamnetin; apigenin; cyanidin; cyanidin glycoside; black soybean polyphenols; anthocyanin rhamnoglucoside; or salts or solvates thereof.

[0568] Heterocyclic compounds can be any suitable compound as described herein, such as an optionally substituted aromatic or non-aromatic ring containing a heteroatom. On one hand, a heterocycle is a compound containing an optionally substituted heterocyclic alkyl moiety or an optionally substituted heteroaryl moiety. On the other hand, the optionally substituted heterocyclic alkyl moiety or the optionally substituted heteroaryl moiety is an optionally substituted 5-membered heterocyclic alkyl moiety or an optionally substituted 5-membered heteroaryl moiety. On the other hand, the optionally substituted heterocyclic alkyl or optionally substituted heteroaryl moiety is a moiety selected from the following optionally substituted moiety: pyrazolyl, furanyl, imidazolyl, isoxazolyl, oxadiazolyl, oxazolyl, pyrroleyl, pyridinyl, pyrimidinyl, pyridazinyl, thiazolyl, triazolyl, thiophene, dihydrothiophene-pyrazolyl, thioindyl, carbazoleyl, benzimidazolyl, benzothiophene, benzofuranyl, indolyl, quinolinyl, benzotriazolyl, benzothiazolyl, benzooxazolyl, benzimidazolyl, isoquinolinyl, isoindolyl, acridineyl, benzoisoazolyl, dimethylhydantoin, pyrazinyl, tetrahydrofuranyl, pyrrolinyl, pyrrolidinyl, morpholinyl, indolyl, diazoponyl, azoponyl, thiapaonyl, piperidinyl, and oxoponyl. On the other hand, the optionally substituted heterocyclic alkyl moiety or the optionally substituted heteroaryl moiety is an optionally substituted furanyl group. Non-limiting examples of heterocyclic compounds include (1,2-dihydroxyethyl)-3,4-dihydroxyfuran-2(5H)-one; 4-hydroxy-5-methyl-3-furanone; 5-hydroxy-2(5H)-furanone; [1,2-dihydroxyethyl]furan-2,3,4(5H)-trione; α-hydroxy-γ-butyrolactone; ribonucleic acid γ-lactone; aldohexuronicaldohexuronic acid γ-lactone; gluconate δ-lactone; 4-hydroxycoumarin; dihydrobenzofuran; 5-(hydroxymethyl)furfural; bifurfural; 2(5H)-furanone; 5,6-dihydro-2H-pyran-2-one; and 5,6-dihydro-4-hydroxy-6-methyl-2H-pyran-2-one; or salts or solvates thereof.

[0569] Nitrogen-containing compounds can be any suitable compound having one or more nitrogen atoms. In one aspect, nitrogen-containing compounds comprise an amine, imine, hydroxylamine, or nitride moiety. Non-limiting examples of nitrogen-containing compounds include acetone oxime; violetic acid; pyridine-2-aldehyde oxime; 2-aminophenol; 1,2-phenylenediamine; 2,2,6,6-tetramethyl-1-piperidinyloxy; 5,6,7,8-tetrahydrobiopterin; 6,7-dimethyl-5,6,7,8-tetrahydropterin; and maleic anhydride; or salts or solvates thereof.

[0570] Quinone compounds can be any suitable compound comprising a quinone moiety as described herein. Non-limiting examples of quinone compounds include: 1,4-benzoquinone, 1,4-naphthoquinone, 2-hydroxy-1,4-naphthoquinone, 2,3-dimethoxy-5-methyl-1,4-benzoquinone or coenzyme Q0, 2,3,5,6-tetramethyl-1,4-benzoquinone or duquinone, 1,4-dihydroxyanthraquinone, 3-hydroxy-1-methyl-5,6-dihydroindoledione or adrenaline red, 4-tert-butyl-5-methoxy-1,2-benzoquinone, pyrroloquinolinequinone, or salts or solvates thereof.

[0571] Sulfur-containing compounds can be any suitable compound comprising one or more sulfur atoms. In one aspect, sulfur-containing compounds contain a moiety selected from: thionyl, thioether, sulfinyl, sulfonyl, thioamide, sulfonamide, sulfonic acid, and sulfonate. Non-limiting examples of sulfur-containing compounds include ethanethiol; 2-propanethiol; 2-propen-1-thiol; 2-mercaptoethanesulfonic acid; benzenethiophenol; benzene-1,2-dithiophenol; cysteine; methionine; glutathione; cystine; or salts or solvates thereof.

[0572] On one hand, the effective amount of the compound described above for cellulose materials, as a molar ratio to the glucosyl units of cellulose, is approximately 10. -6 From approximately 10, for example, approximately 10 -6 To approximately 7.5, approximately 10 -6 Approximately 5, approximately 10 -6 From approximately 2.5, approximately 10 -6 About 1, about 10 -5 About 1, about 10 -5 To about 10 -1 Approximately 10 -4 To about 10 -1 Approximately 10 -3 To about 10 -1 or about 10- 3 To about 10 -2 On the other hand, the effective amount of such compound described above is from about 0.1 μM to about 1 M, for example, from about 0.5 μM to about 0.75 M, from about 0.75 μM to about 0.5 M, from about 1 μM to about 0.25 M, from about 1 μM to about 0.1 M, from about 5 μM to about 50 mM, from about 10 μM to about 25 mM, from about 50 μM to about 25 mM, from about 10 μM to about 10 mM, from about 5 μM to about 5 mM, or from about 0.1 mM to about 1 mM.

[0573] The term "liquid" refers to an aqueous, organic, or combination thereof solution phase, and its soluble contents, produced under the conditions described herein by treatment of lignocellulose and / or hemicellulose materials or their monosaccharides (e.g., xylose, arabinose, mannose, etc.) in a slurry. A cellulose-enhancing liquid for GH61 peptides can be produced by treating a lignocellulose or hemicellulose material (or raw material) by applying heat and / or pressure, optionally in the presence of a catalyst (e.g., acid), optionally in the presence of an organic solvent, and optionally in combination with a material that physically breaks down a lignocellulose or hemicellulose material (or raw material), and then separating the solution from the residual solids. The degree of cellulose-enhancing effect obtainable from the combination of the liquid and GH61 peptides during the hydrolysis of cellulose substrates by cellulase preparations is determined by these conditions. The liquid can be separated from the treated material using standard methods in the art, such as filtration, precipitation, or centrifugation.

[0574] On the one hand, the effective amount of liquid for cellulose is approximately 10. -6 cellulose up to approximately 10 g / g, for example, approximately 10 -6 Approximately 7.5g, approximately 10g -6 Approximately 5g, approximately 10 -6 Approximately 2.5g, approximately 10 -6 Approximately 1g, approximately 10 -5 Approximately 1g, approximately 10 -5 To about 10 - 1 g, approximately 10 -4 To about 10 -1 g, approximately 10 -3 To about 10 -1 g, or about 10 -3 To about 10 -2 cellulose g / g.

[0575] In one aspect, the one or more (e.g., several) hemicellulose-degrading enzymes include commercial hemicellulose-degrading enzyme formulations. Examples of commercially available hemicellulose-degrading enzyme formulations suitable for use in this invention include, for example, SHEARZYME. TM (Novozymes) HTec (Novozymes) HTec2 (Novozymes) HTec3 (Novozymes) (Novozymes) (Novozymes) HC (Novozymes) Xylanase (Genetronics) XY (Genetronics Corporation) XC (Genetronics Corporation) TX-200A (AB Enzymes), HSP 6000 xylanase (DSM), DEPOL TM 333P (Biocatalysts Limit, Wales, UK), DEPOL TM 740L (Biocatalyst Ltd, Wales, UK) and DEPOL TM 762P (Biocatalyst Ltd, Wales, UK).

[0576] Examples of xylanases used in the methods of the present invention include, but are not limited to, xylanases derived from the following: Aspergillus echinosporum (GeneSeqP:AAR63790; WO 94 / 21785), Aspergillus fumigatus (WO 2006 / 078256), Penicillium pineophilum (WO2011 / 041405), Penicillium species (WO 2010 / 126772), Clostridium terrestris NRRL 8126 (WO 2009 / 079210), and Trichophaea saccata GH10 (WO 2011 / 057083).

[0577] Examples of β-xylosidases used in the methods of the present invention include, but are not limited to, β-xylosidases from the following: *Streptococcus roughus* (SwissProt accession number Q7SOW4), *Trichoderma reesei* (UniProtKB / TrEMBL accession number Q92458), and *Talaromyces emersonii* (SwissProt accession number Q8X212).

[0578] Examples of acetylated xylan esterases that can be used in the process of this invention include, but are not limited to, acetylated xylan esterases from the following: Aspergillus echinosporum (WO 2010 / 108918), Chaetomium globosum (Uniprot accession number Q2GWX4), Chaetomium spp. (GeneSeqP accession number AAB82124), Pyrophyte DSM 1800 (WO 2009 / 073709), Sarcoptera rubra (WO2005 / 001036), Myceliophtera thermophila (WO 2010 / 014880), Neurospora crassa (UniProt accession number q7s259), Synsporium globosum (Uniprot accession number Q0UHJ1), and Clostridium terrestris NRRL8126 (WO 2009 / 042846).

[0579] Examples of feruloyl esterases (ferulic acidesterases) that can be used in the processes of this invention include, but are not limited to, feruloyl esterases from the following: *Potentilla spp.* DSM1800 (WO2009 / 076122), *Neosartorya fischeri* (UniProt accession number A1D9T4), *Neurospora crassa* (UniProt accession number Q9HGR3), *Penicillium chrysogenum* (WO 2009 / 127729), and *Clostridium thuringiensis* (WO 2010 / 053838 and WO 2010 / 065448).

[0580] Examples of arabinofuranosidases that can be used in the processes of this invention include, but are not limited to, arabinofuranosidases from the following: Aspergillus niger (GeneSeqP accession number AAR94170), *Hypericum spp.* DSM 1800 (WO 2006 / 114094 and WO 2009 / 073383), and *M. giganteus* (WO 2006 / 114094).

[0581] Examples of α-glucuronidases that can be used in the processes of this invention include, but are not limited to, α-glucuronidases from the following: Aspergillus lanceolata (UniProt accession number alcc12), Aspergillus fumigatus (SwissProt accession number Q4WW45), Aspergillus niger (Uniprot accession number Q96WX9), Aspergillus terreus (SwissProt accession number Q0CJP9), *Pseudomonas aeruginosa* (WO 2010 / 014706), *Penicillium chrysogenum* (WO 2009 / 068565), *Emersonia spp.* (UniProt accession number Q8X211), and *Trichoderma reesei* (Uniprot accession number Q99024).

[0582] The enzymatically active polypeptides used in the methods of this invention can be produced by fermentation of the microbial strains described above using a procedure known in the art on a nutrient medium containing suitable carbon and nitrogen sources and inorganic salts (see, for example, Bennett, JW, and LaSure, L. (ed.), More Gene Manipulations in Fungi, Academic Press, California, 1991). Suitable culture media are available from suppliers or can be prepared according to published compositions (e.g., the catalogue of the U.S. Center for Typical Culture Collections). Suitable temperature ranges and other conditions for growth and enzyme production are known in the art (see, for example, Bailey, JE, and Ollis, DF), Biochemical Engineering Fundamentals, McGraw-Hill Book Company, New York, 1986).

[0583] Fermentation can be any method of cultured cells that results in the expression or isolation of an enzyme or protein. Therefore, fermentation can be understood as including shake-flask culture, or small-scale or large-scale fermentation (including continuous fermentation, batch fermentation, fed-batch fermentation, or solid-state fermentation) in a suitable culture medium and under conditions that allow for the expression or isolation of the enzyme in a laboratory or industrial fermenter. The resulting enzymes can be recovered from the fermentation medium and purified using standard procedures.

[0584] Fermentation Fermentable sugars obtained from hydrolyzed cellulose material can be fermented by one or more (e.g., several) fermentative microorganisms capable of directly or indirectly fermenting sugars into a desired fermentation product. "Fermentation" or "fermentation method" refers to any fermentation method or any method that includes a fermentation step. Fermentation methods also include those used in the consumer alcohol industry (e.g., beer and wine), the dairy industry (e.g., fermented dairy products), the leather industry, and the tobacco industry. Fermentation conditions depend on the desired fermentation product and the fermenting organism and can be readily determined by those skilled in the art.

[0585] In the fermentation step, sugars released from the cellulose material as a result of pretreatment and enzymatic hydrolysis are fermented by a fermenting organism (such as yeast) into products, such as ethanol. As mentioned above, hydrolysis (saccharification) and fermentation can be separate or simultaneous.

[0586] In practicing the fermentation steps of this invention, any suitable hydrolyzed cellulose material can be used. Materials are typically selected based on the desired fermentation product (i.e., the substance to be obtained from fermentation) and the method employed, as is well known in the art.

[0587] The term “fermentation medium” can be understood here as a medium prior to the addition of one or more fermenting microorganisms, such as a medium produced by a saccharification process, and a medium used in a simultaneous saccharification and fermentation (SSF) process.

[0588] "Fermentation microorganism" refers to any microorganism, including bacteria and fungi, suitable for the desired fermentation method to produce a fermentation product. Fermentation organisms can be hexose and / or pentose fermentation organisms, or combinations thereof. Both hexose and pentose fermentation organisms are well known in the art. Suitable fermentation microorganisms are capable of fermenting (i.e., converting) sugars (such as glucose, xylose, xylulose, arabinose, maltose, mannose, galactose, and / or oligosaccharides) directly or indirectly into the desired fermentation product.

[0589] Lin et al., 2006, Applied Microbiology and Biotechnology (Appl. Microbiol. Biotechnol.) 69:627-642, described examples of bacterial and fungal fermentation organisms that produce ethanol.

[0590] Examples of fermenting microorganisms capable of fermenting hexoses include bacterial and fungal organisms, such as yeast. Preferred yeasts include strains of the genera *Candida*, *Kluyveromyces*, and *Saccharomyces*, such as *Candida sonorensis*, *Kluyveromyces maculae*, and *Saccharomyces cerevisiae*.

[0591] Examples of fermentative microorganisms capable of fermenting pentoses in their native state include bacteria and fungi, such as certain yeasts. Preferred xylose-fermenting yeasts include strains of the genus *Candida*, preferably *Candida sheatae* or *Candida sonorensis*; and strains of the genus *Pichia*, preferably *Pichia stylosa*, such as *Pichia stylosa* CBS 5773. Preferred pentose-fermenting yeasts include strains of the genus *Pachysolen*, preferably *Pichia tannophilus*. Organisms that cannot ferment pentoses (such as xylose and arabinose) can be genetically modified to ferment pentoses using methods known in the art.

[0592] Examples of bacteria that can efficiently ferment hexoses and pentoses into ethanol include, for example, Bacillus coagulans, Clostridium acetobutyricum, Clostridium thermofibrinolyticum, Clostridium phytofermentans, Bacillus spp., Thermoanaerobacter saccharolyticum, and motile fermentation monoclonal bacteria (Philippidis, 1996, ibid.).

[0593] Other fermenting organisms include strains of the following: *Bacillus* genus, such as *Bacillus coagulans*; *Candida* genus, such as *Candida sanaresis*, *Candida methylsorbitan*, *Candida didensia*, and *Candida parapsilosis*. Candida parapsilosis, Candida naedodendra, Candida bronchoides, Candida wormi, Candida brassicae, Candida pseudotropica, Candida boyidin, Candida utilis, and Candida shurata; Clostridium species, such as Clostridium acetobutol, Clostridium thermophilum, and Clostridium fermentum; Escherichia coli, especially genetically modified strains to improve ethanol production; Bacillus species; Hansenula species, such as Hansenula anomala; Klebsiella species, such as Klebsiella acidophilus; Kluyveromyces species, such as Kluyveromyces marx, Kluyveromyces lactis, Kluyveromyces thermostableis, and Kluyveromyces brittlewallis; Schizosomyces species, such as Schizosomyces pombe; Thermoanaerobacter species, such as Thermoanaerobacter glycolyticus, and Zymomonas species, such as Zymomonas motiformis.

[0594] In one preferred aspect, the yeast is *Bretannomyces*. In a more preferred aspect, the yeast is *Bretannomyces clausenii*. In another preferred aspect, the yeast is *Candida*. In another more preferred aspect, the yeast is *Candida sanaresis*. In another more preferred aspect, the yeast is *Candida boydin*. In another more preferred aspect, the yeast is *Candida bronchiol*. In another more preferred aspect, the yeast is *Candida brassicae*. In another more preferred aspect, the yeast is *Candida didens*. In another more preferred aspect, the yeast is *Candida entomopathogenica*. In another more preferred aspect, the yeast is *Candida pseudotropica*. In another more preferred aspect, the yeast is *Candida schwahatta*. In another more preferred aspect, the yeast is *Candida utilis*. In another preferred aspect, the yeast is *Clavispora*. In another more preferred aspect, the yeast is *Clavisporalusitaniae*. In another more preferred aspect, the yeast is *Clavispora opuntiae*. In another preferred aspect, the yeast is *Kluyveromyces*. In another more preferred aspect, the yeast is *Kluyveromyces brittle-wall*. In another more preferred aspect, the yeast is *Kluyveromyces marx*. In another more preferred aspect, the yeast is *Kluyveromyces thermostableum*. In another preferred aspect, the yeast is *Saccharomyces*. In another more preferred aspect, the yeast is *Saccharomyces tanninophilus*. In another preferred aspect, the yeast is *Pichia*. In another preferred aspect, the yeast is *Pichia stearosa*. In another preferred aspect, the yeast is a specific species within the genus *Saccharomyces*. In another more preferred aspect, the yeast is *Saccharomyces brewer's yeast*. In another more preferred aspect, the yeast is *Saccharomyces distaticus*. In another more preferred aspect, the yeast is *Saccharomyces uvarum*.

[0595] In one preferred aspect, the bacteria are *Bacillus*. In a more preferred aspect, the bacteria are *Bacillus coagulans*. In another preferred aspect, the bacteria are *Clostridium*. In another more preferred aspect, the bacteria are *Clostridium acetobutanol*. In another more preferred aspect, the bacteria are *Clostridium phytofermentans*. In another more preferred aspect, the bacteria are *Clostridium thermofibrinolyticum*. In another more preferred aspect, the bacteria are *Bacillus* species. In another more preferred aspect, the bacteria are *Anaerobic thermophilic bacteria*. In another more preferred aspect, the bacteria are *Anaerobic thermophilic bacteria*. In another preferred aspect, the bacteria are *Fermentomonas*. In another more preferred aspect, the bacteria are *Fermentomonas motiformis*.

[0596] Commercially available yeasts suitable for ethanol production include, for example, BIOFERM. TMAFT and XR (NABC - North American Bioproducts Corporation, Georgia, USA), ETHANOLRED TM Yeast (Fermentis / Lesaffre, USA), FALI TM (Fleischmann's Yeast, USA), FERMIOL TM (DSM Specialties), GERTSTRAND TM (Gert Strand AB, Sweden), and SUPERSTART TM and THERMOSACC TM Fresh yeast (Ethanol Technology, Wisconsin, USA).

[0597] In a preferred aspect, the fermenting microorganisms have been genetically modified to provide the ability to ferment pentoses, such as those utilizing xylose, arabinose, and microorganisms that utilize both xylose and arabinose.

[0598] By cloning heterologous genes into different fermenting microorganisms, organisms capable of converting hexoses and pentoses into ethanol (co-fermentation) have been constructed (Chen and Ho, 1993, Cloning and improving the expression of Pichia stipitisxylose reductase gene in Saccharomyces cerevisiae, Applied Biochemistry and Biotechnology, 39-40:135-147); Huo et al., 1998, Genetically engineered Saccharomyces yeast capable of effectively cofermenting glucose and xylose, Applied and Environmental Microbiology, 64:1852-1859; Kotter and Ciriacy, 1993, Xylosefermentation by Saccharomyces cerevisiae). cerevisiae), Applied Microbiology and Biotechnology (Appl. Microbiol. Biotechnol).38:776-783; Walfridsson et al., 1995, Xylose-metabolizing Saccharomyces cerevisiae strains overexpressing the TKL1 and TAL1 genes encoding the pentose phosphate pathway enzymes transketolase and transaldolase, Applied and Environmental Microbiology 61:4184-4190; Kuyper et al., 2004, Minimal metabolic engineering of Saccharomyces cerevisiae for efficient anaerobic xylose fermentation: a proof of principle, FEMSYeast Research) 4:655-664; Beall et al., 1991, Parametric studies of ethanol production from xylose and other sugars by recombinant Escherichia coli, Biotechnology and Bioengineering (Biotech.Bioeng).38:296-303; Ingram et al., 1998, Metabolic engineering of bacteria for ethanol production, Biotechnology and Bioengineering 58:204-214; Zhang et al., 1995, Metabolic engineering of a pentose metabolism pathway in ethanologenic Zymomonas mobilis, Science 267:240-243; Deanda et al., 1996, Development of anarabinose-fermenting Zymomonas mobilis strain by metabolic pathway engineering, Applied and Environmental Microbiology 62:4465-4470; WO 2003 / 062430, Xylose isomerase).

[0599] In one preferred aspect, the genetically modified fermenting microorganism is *Candida sanguinis*. In another preferred aspect, the genetically modified fermenting microorganism is *Escherichia coli*. In yet another preferred aspect, the genetically modified fermenting microorganism is *Klebsiella oxytoca*. In yet another preferred aspect, the genetically modified fermenting microorganism is *Kluyveromyces martensii*. In yet another preferred aspect, the genetically modified fermenting microorganism is *Saccharomyces cerevisiae*. In yet another preferred aspect, the genetically modified fermenting microorganism is *Fermentomonas motilityis*.

[0600] It is well known in the art that the aforementioned organisms can also be used to produce other substances, as described herein.

[0601] Typically, fermenting microorganisms are added to the degraded cellulose material or hydrolysate, and fermentation is carried out for about 8 to about 96 hours, for example, about 24 to about 60 hours. The temperature is typically between about 26°C and about 60°C, for example, about 32°C or 50°C, and the pH is between about pH 3 and about pH 8, for example, pH 4 to 5, 6 or 7.

[0602] On one hand, yeast and / or another microorganism are applied to the degraded cellulose material and fermentation is carried out for approximately 12 to approximately 96 hours, typically 24-60 hours. On the other hand, the temperature is preferably between approximately 20°C and approximately 60°C, for example, approximately 25°C to approximately 50°C, approximately 32°C to approximately 50°C, or approximately 32°C to approximately 50°C, and the pH is generally from approximately pH 3 to approximately pH 7, for example, approximately pH 4 to approximately pH 7. However, some fermenting organisms, such as bacteria, have higher optimal fermentation temperatures. Yeast or another microorganism is preferably introduced at approximately 10 mg / ml of fermentation broth. 5 Up to 10 12 Preferably from about 10 7 Up to 10 10 Especially about 2×10 8 The amount applied is based on a live cell count. Further guidance on fermentation using yeast can be found, for example, in "The Alcohol Textbook" (edited by K. Jacques, T.P. Lyons, and D.D. Kelsall, Nottingham University Press, United Kingdom, 1999), which is incorporated herein by reference.

[0603] Fermentation stimulants can be used in combination with any of the methods described herein to further improve the fermentation process, and specifically, to improve the performance of fermenting microorganisms, such as increased growth rate and ethanol yield. “Fermentation stimulant” refers to an agent used to stimulate the growth of fermenting microorganisms (particularly yeast). Preferred fermentation stimulants for growth include vitamins and minerals. Examples of vitamins include multivitamins, biotin, pantothenic acid, niacin, meso-inositol, thiamine, pyridoxine, para-aminobenzoic acid, folic acid, riboflavin, and vitamins A, B, C, D, and E. See, for example, Alfenore et al., Improving ethanol production and viability of Saccharomyces cerevisia by a vitamin feeding strategy during fed-batch process, Springer (2002), which is incorporated herein by reference. Examples of minerals include minerals and mineral salts that can supply nutrients including P, K, Mg, S, Ca, Fe, Zn, Mn, and Cu.

[0604] Fermentation products:Fermentation products can be any substance obtained from fermentation. Fermentation products can be, but are not limited to, alcohols (e.g., arabinol, n-butanol, isobutanol, ethanol, glycerol, methanol, ethylene glycol, 1,3-propanediol (propylene glycol), butanediol, glycerol, sorbitol, and xylitol); alkanes (e.g., pentane, hexane, heptane, octane, nonane, decane, undecane, and dodecane); cycloalkanes (e.g., cyclopentane, cyclohexane, cycloheptane, and cyclooctane); alkenes (e.g., pentene, hexene, heptene, and octene); and amino acids (e.g., aspartic acid, glutamic acid, glycine, and lysine). Fermentation products include: acids (serine and threonine); gases (e.g., methane, hydrogen (H2), carbon dioxide (CO2), and carbon monoxide (CO)); isoprene; ketones (e.g., acetone); organic acids (e.g., acetic acid, acetoic acid, adipic acid, ascorbic acid, citric acid, 2,5-diketo-D-gluconic acid, formic acid, fumaric acid, gluconic acid, glucuronic acid, glutaric acid, 3-hydroxypropionic acid, itaconic acid, lactic acid, malic acid, malonic acid, oxalic acid, oxaloacetic acid, propionic acid, succinic acid, and xylic acid); and polyketide compounds. Fermentation products can also be proteins, which are high-value products.

[0605] In a preferred aspect, the fermentation product is an alcohol. It is understood that the term "alcohol" includes substances containing one or more hydroxyl groups. In a more preferred aspect, the alcohol is n-butanol. In another more preferred aspect, the alcohol is isobutanol. In yet another more preferred aspect, the alcohol is ethanol. In yet another more preferred aspect, the alcohol is methanol. In yet another more preferred aspect, the alcohol is arabinitol. In yet another more preferred aspect, the alcohol is butanediol. In yet another more preferred aspect, the alcohol is ethylene glycol. In yet another more preferred aspect, the alcohol is glycerol. In yet another more preferred aspect, the alcohol is glycerol. In yet another more preferred aspect, the alcohol is 1,3-propanediol. In yet another more preferred aspect, the alcohol is sorbitol. In yet another more preferred aspect, the alcohol is xylitol. See, for example, Gon CS, Káo NJ, Dú J., and Tóu GT, 1999, Ethanol production from renewable resources, Advances in Biochemical Engineering / Biotechnology, edited by Schepper T., Springer, Heidelberg-Berlin, 65:207-241; Silveira MM and Jonas R., 2002, The biotechnological production of sorbitol, Applied Microbiology & Biotechnology 59:400-408; Nigam P. and Singer D., 1995, Processes for fermentative production of xylitol-a sugar substitute, Process Biochemistry Biochemistry) 30(2):117-124; Ezeji, TC, Qureshi, N. and Blaschek, HP, 2003, Production of acetone, butanol and ethanol by Clostridium beijerinckii BA101 and in situ recovery by gas stripping, World Journal of Microbiology and Biotechnology 19(6):595-603.

[0606] In another preferred aspect, the fermentation product is an alkane. The alkane may be unbranched or branched. In another more preferred aspect, the alkane is pentane. In another more preferred aspect, the alkane is hexane. In another more preferred aspect, the alkane is heptane. In another more preferred aspect, the alkane is octane. In another more preferred aspect, the alkane is nonane. In another more preferred aspect, the alkane is decane. In another more preferred aspect, the alkane is undecane. In another more preferred aspect, the alkane is dodecane.

[0607] In another preferred aspect, the fermentation product is a cycloalkane. In another more preferred aspect, the cycloalkane is cyclopentane. In another more preferred aspect, the cycloalkane is cyclohexane. In another more preferred aspect, the cycloalkane is cycloheptane. In another more preferred aspect, the cycloalkane is cyclooctane.

[0608] In another preferred aspect, the fermentation product is an olefin. The olefin may be unbranched or branched. In another more preferred aspect, the olefin is pentene. In another more preferred aspect, the olefin is hexene. In another more preferred aspect, the olefin is hepten. In another more preferred aspect, the olefin is octene.

[0609] In another preferred aspect, the fermentation product is an amino acid. In another more preferred aspect, the organic acid is aspartic acid. In another more preferred aspect, the amino acid is glutamic acid. In another more preferred aspect, the amino acid is glycine. In another more preferred aspect, the amino acid is lysine. In another more preferred aspect, the amino acid is serine. In another more preferred aspect, the amino acid is threonine. See, for example, Richard, A. and Margaritis, A., 2004, Empirical modeling of batch fermentation kinetics for poly(glutamic acid) production and other microbial biopolymers, Biotechnology and Bioengineering 87(4):501-515.

[0610] In another preferred aspect, the substance is a gas. In another more preferred aspect, the gas is methane. In another more preferred aspect, the gas is H2. In another more preferred aspect, the gas is CO2. In another more preferred aspect, the gas is CO. See, for example, Kataoka, N., A. Miya, and K. Kiriyama, 1997, Studies on hydrogen production by continuous culture system of hydrogen-producing anaerobic bacteria, Water Science and Technology 36(6-7):41-47; and Gunaseelan VN, Biomass and Bioenergy, Vol. 13(1-2), pp. 83-114, 1997, Anaerobic digestion of biomass for methane production: A review.

[0611] In another preferred aspect, the fermentation product is isoprene.

[0612] In another preferred aspect, the fermentation product is a ketone. It should be understood that the term "ketone" encompasses substances containing one or more ketone moieties. In yet another more preferred aspect, the ketone is acetone. See, for example, Kuresch and Brassek, 2003, above.

[0613] In another preferred aspect, the fermentation product is an organic acid. In another, more preferred aspect, the organic acid is acetic acid. In another, more preferred aspect, the organic acid is acetoic acid. In another, more preferred aspect, the organic acid is adipic acid. In another, more preferred aspect, the organic acid is ascorbic acid. In another, more preferred aspect, the organic acid is citric acid. In another, more preferred aspect, the organic acid is 2,5-diketo-D-gluconic acid. In another, more preferred aspect, the organic acid is formic acid. In another, more preferred aspect, the organic acid is fumaric acid. In another, more preferred aspect, the organic acid is gluconic acid. In another, more preferred aspect, the organic acid is glucuronic acid. In another, more preferred aspect, the organic acid is glutaric acid. In another, more preferred aspect, the organic acid is 3-dihydroxypropionic acid. In another, more preferred aspect, the organic acid is itaconic acid. In another, more preferred aspect, the organic acid is lactic acid. In another, more preferred aspect, the organic acid is malic acid. In another, more preferred aspect, the organic acid is malonic acid. In another preferred aspect, the organic acid is oxalic acid. In another preferred aspect, the organic acid is propionic acid. In another preferred aspect, the organic acid is succinic acid. In another preferred aspect, the organic acid is xyloic acid. See, for example, Chen, R. and Li, YY, 1997, Membrane-mediated extractive fermentation for lactic acid production from cellulosic biomass, Applied Biochemistry and Biotechnology 63-65:435-448.

[0614] In another preferred aspect, the fermentation product is a polyketide compound.

[0615] Recycle One or more fermentation products can optionally be recovered from the fermentation medium using any method known in the art, including but not limited to chromatography, electrophoresis, differential solubility, distillation, or extraction. For example, alcohols can be separated and purified from fermented cellulose material by conventional distillation methods. Ethanol with a purity of up to about 96 vol.% can be obtained, which can be used as, for example, fuel ethanol, drinking ethanol (i.e., neutral drinking ethanol), or industrial ethanol.

[0616] plant

[0617] The present invention also relates to isolated plants, such as transgenic plants, plant parts, or plant cells, that include the polypeptides of the present invention, thereby expressing and producing recyclable ...

Claims

1. An isolated cellobiose hydrolase variant, said variant comprising: Mature polypeptide sequences of SEQ ID NO: 90 or SEQ ID NO: 92, or SEQ ID NO: 90 or SEQ ID NO:

92.

2. A composition comprising the variant as claimed in claim 1.

3. An isolated polynucleotide encoding the variant as described in claim 1.

4. A nucleic acid construct comprising the polynucleotide as described in claim 3.

5. An expression vector comprising the polynucleotide as described in claim 3.

6. A recombinant host cell comprising the polynucleotide as described in claim 3.

7. A method for producing a variant of parental cellobiose hydrolase, the method comprising: The recombinant host cells as described in claim 6 are cultured under conditions suitable for the expression of the variant.

8. The method of claim 7, further comprising recycling the variant.

9. A method for preparing a transgenic plant, plant part or plant cell, wherein the transgenic plant, plant part or plant cell is transformed with the polynucleotide as described in claim 3.

10. A method for producing a variant of claim 1, the method comprising: Transgenic plants, plant parts, or plant cells containing polynucleotides encoding the variant are cultured under conditions that facilitate the generation of the variant.

11. The method of claim 10, further comprising recycling the variant.

12. A hybrid polypeptide, said hybrid polypeptide comprising: Mature polypeptides of SEQ ID NO: 61, SEQ ID NO: 63, SEQ ID NO: 73, or SEQ ID NO: 94, or SEQ ID NO: 61, SEQ ID NO: 63, SEQ ID NO: 73, or SEQ ID NO:

94.

13. A composition comprising the heterozygous polypeptide as described in claim 12.

14. An isolated polynucleotide encoding a heterozygous polypeptide as described in claim 12.

15. A nucleic acid construct comprising the polynucleotide as described in claim 14.

16. An expression vector comprising the polynucleotide as described in claim 14.

17. A recombinant host cell comprising the polynucleotide as described in claim 14.

18. A method for generating a heterozygous polypeptide, the method comprising culturing a recombinant host cell as described in claim 17 under conditions suitable for expression of the heterozygous polypeptide.

19. The method of claim 18, further comprising recovering the hybrid polypeptide.

20. A method for preparing a transgenic plant, plant part or plant cell, wherein the transgenic plant, plant part or plant cell is transformed with the polynucleotide as described in claim 14.

21. A method for generating a hybrid polypeptide, the method comprising: The transgenic plant, plant part or plant cell as described in claim 20 is cultured under conditions that facilitate the production of the hybrid polypeptide.

22. A method for degrading or converting cellulosic materials, the method comprising: The cellulose material is treated with an enzyme composition, wherein the composition comprises a variant as described in claim 1 or a hybrid polypeptide as described in claim 12.

23. The method of claim 22, wherein the cellulose material is pretreated.

24. The method of claim 22, wherein the enzyme composition further comprises one or more enzymes selected from cellulase, GH61 polypeptide with cellulolytic-enhancing activity, hemicellulase, patulin, esterase, laccase, lignin-degrading enzyme, pectinase, peroxidase, protease, and swelling agent.

25. The method of claim 24, wherein the cellulase is one or more enzymes selected from endoglucanase, cellobiase, and β-glucosidase.

26. The method of claim 24, wherein the hemicellulase is one or more enzymes selected from xylanase, acetylxylan esterase, ferulic acid esterase, arabinofuranylase, xylosidase, and glucuronyl glycosidase.

27. The method of any one of claims 22-26, further comprising recovering the degraded cellulose material.

28. The method of claim 27, wherein the degraded cellulose material is sugar.

29. The method of claim 28, wherein the sugar is selected from glucose, xylose, mannose, galactose, and arabinose.

30. A method for producing fermentation products, the method comprising: (a) Saccharifying cellulose material with an enzyme composition, wherein the composition comprises a variant of claim 1 or a hybrid polypeptide as described in claim 12; (b) fermenting the saccharified cellulose material with one or more fermenting microorganisms to produce the fermentation product; and (c) recovering the fermentation product from the fermentation.

31. The method of claim 30, wherein the cellulose material is pretreated.

32. The method of claim 30 or 31, wherein the enzyme composition comprises one or more enzymes selected from cellulase, GH61 polypeptide with cellulolytic-enhancing activity, hemicellulase, patulin, esterase, laccase, lignin-degrading enzyme, pectinase, peroxidase, protease, and swelling agent.

33. The method of claim 32, wherein the cellulase is one or more enzymes selected from endoglucanase, cellobiase, and β-glucosidase.

34. The method of claim 32, wherein the hemicellulase is one or more enzymes selected from xylanase, acetylxylan esterase, ferulic acid esterase, arabinofuranylase, xylosidase, and glucuronyl glycosidase.

35. The method of claim 30, wherein steps (a) and (b) are performed simultaneously in concurrent saccharification and fermentation.

36. The method of claim 30, wherein the fermentation product is an alcohol, organic acid, ketone, amino acid, alkane, cycloalkanes, olefin, isoprene, polyketide, or gas.

37. A method for fermenting cellulose material, the method comprising: The cellulose material is fermented with one or more fermenting microorganisms, wherein the cellulose material is saccharified with an enzyme composition comprising the variant of claim 1 or the hybrid polypeptide of claim 12.

38. The method of claim 37, wherein the fermentation of the cellulose material produces fermentation products.

39. The method of claim 38, further comprising recovering the fermentation product from the fermentation.

40. The method of claim 37, wherein the cellulose material is pretreated before saccharification.

41. The method of claim 37, wherein the enzyme composition further comprises one or more enzymes selected from cellulase, GH61 polypeptide with cellulolytic-enhancing activity, hemicellulase, patulin, esterase, laccase, lignin-degrading enzyme, pectinase, peroxidase, protease, and swelling agent.

42. The method of claim 41, wherein the cellulase is one or more enzymes selected from endoglucanase, cellobiase, and β-glucosidase.

43. The method of claim 41, wherein the hemicellulase is one or more enzymes selected from xylanase, acetylxylan esterase, ferulic acid esterase, arabinofuranylase, xylosidase, and glucuronyl glycosidase.

44. The method of any one of claims 38-43, wherein the fermentation product is an alcohol, organic acid, ketone, amino acid, alkane, cycloalkanes, olefin, isoprene, polyketide compound, or gas.

45. A whole culture medium formulation or cell culture composition comprising the variant of claim 1 or the hybrid polypeptide of claim 12.