Glucose dehydrogenase
By identifying a specific amino acid region and applying defined conditions, novel FAD-GDHs are efficiently screened and modified, addressing the inefficiencies of traditional methods and enhancing their suitability for blood glucose measurement.
Patent Information
- Application Number
- JP2025163992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-13
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-14
AI Technical Summary
Existing methods for identifying FAD-dependent glucose dehydrogenases (FAD-GDHs) are inefficient and rely on trial and error, making it difficult to find novel enzymes suitable for blood glucose measurement applications.
Identification of a specific amino acid region and associated rules for FAD-GDHs, allowing for the efficient screening and modification of novel FAD-GDHs through the use of defined conditions and conserved regions, enabling the development of enzymes with improved characteristics for glucose measurement.
Enables the efficient identification and utilization of novel FAD-GDHs with enhanced properties for blood glucose measurement, overcoming the inefficiencies of traditional methods.
Smart Images

Figure 2026004420000077 
Figure 2026004420000078 
Figure 2026004420000079
Abstract
Description
[Technical Field]
[0001] The present invention relates to glucose dehydrogenase, specifically glucose dehydrogenase characterized by a specific amino acid region, its gene, and methods for screening and modifying glucose dehydrogenase. [Background technology]
[0002] The number of diabetic patients is increasing year by year, and diabetic patients, especially those who are insulin-dependent, need to monitor their blood glucose levels daily to control their blood glucose. In recent years, it has become possible to check the blood glucose levels of diabetic patients using self-monitoring devices that use enzymes to easily and accurately measure blood glucose in real time. Glucose oxidase (EC 1.1.3.4) and PQQ-dependent glucose dehydrogenase (EC 1.1.5.2) (see, for example, Patent Documents 1 to 3) have been developed for glucose sensors (e.g., sensors used in self-monitoring devices), but problems arose with their reactivity to oxygen and to maltose and galactose. To solve these problems, FAD-dependent glucose dehydrogenase (hereinafter abbreviated as "FAD-GDH") was developed (see, for example, Patent Documents 4 to 6 and Non-Patent Documents 1 to 4). There have been many reports of FAD-GDHs that are useful for measuring blood glucose levels (self-monitoring of blood glucose (SMBG) and continuous glucose monitoring (CGM)), but most of them are derived from filamentous fungi and belong to the glucose-methanol-choline (GMC) oxidoreductase family. Their amino acid sequences are highly homologous, and therefore their three-dimensional structures are thought to be similar. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-350588 [Patent Document 2] Japanese Patent Application Laid-Open No. 2001-197888 [Patent Document 3] Special Announcement No. 2001-346587 [License 4] International Publication No. 2004 / 058958 パンフレット [Patent Document 5] International Publication No. 2007 / 139013 パンフレット [License 6] International Publication No. 2006 / 101239 パンフレット [Non-licensed literature]
[0004] [Non-licensed Document 1] Studies on the glucose dehydrogenase of Aspergillus oryzae. I. Induction of its synthesis by p-benzoquinone and hydroquinone, TC Bak, and R. Sato, Biochim. Biophys.Acta, 139, 265-276(1967). [Non-licensed Document 2] Studies on the glucose dehydrogenase of Aspergillus oryzae. II. Purification and physical and chemical properties, TC Bak, Biochim. Biophys. Acta, 139, 277-293 (1967). [Non-licensed Document 3] Studies on the glucose dehydrogenase of Aspergillus oryzae. III. General enzymatic properties, TC Bak, Biochim. Biophys. Acta, 146, 317-327 (1967). [Non-licensed Document 4] Studies on the glucose dehydrogenase of Aspergillus oryzae. IV. Histidyl residue as an active site, TC Bak, and R. Sato, Biochim. Biophys. Acta, 146, 328-335 (1967). Summary of the Invention [Problem to be solved by the invention]
[0005] When aiming to identify / obtain a new FAD-GDH, the usual method is to search for proteins that show homology to the sequence of known glucose dehydrogenases (hereinafter sometimes abbreviated as "GDH"). Conserved amino acid regions characteristic of FAD-GDH (such as FAD-binding sites) are sometimes used, but this usually involves trial and error, making it difficult to efficiently identify novel FAD-GDHs. Therefore, an objective of the present invention is to find a method for efficiently identifying novel FAD-GDHs that can be expected to be used and utilized in applications such as blood glucose measurement, and to promote the use and application of such a method (for example, providing novel FAD-GDHs and their genes, etc.). [Means for solving the problem]
[0006] As a result of extensive research aimed at solving the above problems, we discovered the amino acid region specific to FAD-GDH and the rules that define it. By utilizing this knowledge, we can efficiently identify and obtain novel FAD-GDHs. In fact, we have succeeded in obtaining several novel FAD-GDHs. Based on the above findings and achievements, the following inventions are provided. [1] The following sequence (SEQ ID NO: 1): [ka] (wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, H / N represents H (histidine) or N (asparagine), L / I / M represents L (leucine), I (isoleucine) or M (methionine), V / I represents V (valine) or I (isoleucine), G / A / S represents G (glycine), A (alanine) or S (serine), F / W / Y represents F (phenylalanine), W (tryptophan) or Y (tyrosine), and X represents any amino acid residue), A glucose dehydrogenase comprising an amino acid sequence that is 40% or more identical to the amino acid sequence of any one of SEQ ID NOs: 2 to 4. [2] The sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues, i.e., (the value of the 8th amino acid residue + the value of the 12th amino acid residue + the value of the 13th amino acid residue) is 3.79 or less; The condition that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for the 8th, 11th, 12th, and 13th amino acid residues, i.e., (value of the 8th amino acid residue + value of the 11th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 2.83 or more; The glucose dehydrogenase according to [1], wherein the amino acid region satisfies the following: [3] The sum of the values of the JANJ780101 (Average accessible surface area) listed in the AAindex database for the 3rd and 4th amino acid residues, i.e., (the value of the 3rd amino acid residue + the value of the 4th amino acid residue) is 66 or more. The glucose dehydrogenase according to [2], wherein the amino acid region satisfies the following: [4] The glucose dehydrogenase according to any one of [1] to [3], which has a total length of 530 aa to 630 aa. [5] The glucose dehydrogenase according to any one of [1] to [4], whose amino acid sequence is 50% or more identical to the amino acid sequence of any one of SEQ ID NOs: 2 to 4. [6] The glucose dehydrogenase according to any one of [1] to [4], whose amino acid sequence is 60% or more identical to the amino acid sequence of any one of SEQ ID NOs: 2 to 4. [7] The glucose dehydrogenase according to any one of [1] to [6], which has one or more of the following conserved regions 1 to 10: Conserved region 1 consisting of the following amino acid sequence: [ka] where (F / Y) represents aromatic F or Y, (I / V) represents aromatic I or V, (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (G / A) represents branched aliphatic G or A, (V / A / S / T) represents small V, A, S or T, (S / A / G / C) represents very small S, A, G or C, (V / A / S / T) represents small V, A, S or T, (I / V / L) represents branched aliphatic I, V or L, (N / S) represents polar small N or S, (I / L) represents branched aliphatic I or L, (E / D) represents negative E or D; Conserved region 2 consisting of the following amino acid sequence: [ka] where (N / S / T) represents polar and small N, S, or T, (V / L / A) represents aliphatic V, L, or A, (I / V / L) represents branched aliphatic I, V, or L, (I / V / L) represents branched aliphatic I, V, or L, and (A / P / R) represents A, P, or R; Conserved region 3 consisting of the following amino acid sequence: [ka] where (A / T) represents hydrophobic, small A or T, (G / A) represents very small G or A, (R / K) represents positively charged R or K, (A / L) represents aliphatic A or L, (I / V / L / W) represents branched aliphatic or aromatic I, V, L or W, (G / A) represents very small G or A, (G / T) represents small G or T, (S / T) represents small S or T, (A / S / T) represents small A, S or T, and (I / V / F) represents branched aliphatic or aromatic I, V or F; Conserved region 4 consisting of the following amino acid sequence: [ka] where (A / V / C / E / D) represents small or negatively charged A, V, C, E, or D, (A / S) represents very small A or S, (A / Y) represents A or Y, (R / T / A) represents R, T, or A, (G / A) represents very small G or A, (Y / I / L) represents branched aliphatic or aromatic Y, I, or L, and (W / Y / F / L / A / H / K / Q) represents W, Y, F, L, A, H, K, or Q. Conserved region 5 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (I / V / L) represents branched aliphatic I, V or L, (S / A) represents very small S or A, (S / T / A) represents small S, T or A, (S / A) represents very small S or A, (I / L) represents branched aliphatic I or L, (K / R / A / G / I) represents aliphatic or positively charged (S / T) represents small S or T, (A / G / V / L / K / Q) represents A, G, V, L, K or Q, (I / V / L) represents branched aliphatic I, V or L, (I / V / L / A / H / R) represents aliphatic or positively charged I, V, L, A, H or R, (I / V) represents branched aliphatic I or V, (D / N / S) represents polar and small D, N or S; Conserved region 6 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (D / N) represents small D or N, (L / A / S / N) represents branched aliphatic or small L, A, S or N, (P / A / T) represents small P, A or T, (G / F / T) represents hydrophobic G, F or T, (E / S) represents hydrophilic polar E or S, (L / M) represents hydrophobic L or M, (Q / V) represents Q or V, (D / E) represents negatively charged D or E, and (Q / H) represents polar Q or H; Conserved region 7 consisting of the following amino acid sequence: [ka] however, (S / I / V / L) represents branched aliphatic or very small S, I, V, or L, (V / L / M / A / F) represents hydrophobic V, L, M, A, or F, (L / F) represents hydrophobic L or F, (A / S) represents very small A or S, (N / S / Y) represents N, S, or Y, (I / V / T) represents hydrophobic I, V, or T, and (I / V / L) represents branched aliphatic I, V, or L; Conserved region 8 consisting of the following amino acid sequence: [ka] where (I / M) represents hydrophobic I or M, (L / M) represents hydrophobic L or M, (P / S) represents small P or S, (R / K / E) represents charged R, K or E, (D / E / A / G / S / K) represents D, E, A, G, S or K, (I / L / M / A / N / K) represents I, L, M, A, N or K, (I / V) represents branched aliphatic I or V, and (D / N / S) represents polar and small D, N or S; Conserved region 9 consisting of the following amino acid sequence: [ka] where (Y / H) represents aromatic Y or H, (G / D) represents small G or D, (A / S / T / K) represents A, S, T or K, (L / V) represents branched aliphatic L or V, and (I / V) represents branched aliphatic I or V; Conserved region 10 consisting of the following amino acid sequence: [ka] where (Y / F) represents aromatic Y or F, (A / G) represents very small A or G, (I / L / V) represents branched aliphatic I, L, or V, (A / S) represents very small A or S, (R / K / L) represents branched aliphatic or positively charged R, K, or L, (A / I / T) represents hydrophobic A, I, or T, (A / S) represents very small A or S, (D / E) represents negatively charged D or E, (I / L / V / F / R / Q) represents I, L, V, F, R, or Q, (I / L) represents branched aliphatic I or L, and (K / E / Q) represents polar K, E, or Q. [8] The glucose dehydrogenase according to [7], which has all of conserved regions 1 to 10. [9] The glucose dehydrogenase according to any one of [1] to [8], which consists of an amino acid sequence of any one of SEQ ID NOs: 2 to 4.
[10] An enzyme preparation comprising the glucose dehydrogenase according to any one of [1] to [9] as an active ingredient.
[11] A glucose dehydrogenase gene encoding the glucose dehydrogenase according to any one of [1] to [9].
[12] The glucose dehydrogenase gene according to
[11] , which comprises any one DNA selected from the group consisting of the following (A) to (C): (A) DNA encoding any one of the amino acid sequences of SEQ ID NOs: 2 to 4; (B) DNA consisting of any one of the nucleotide sequences of SEQ ID NOs: 5 to 7; (C) DNA having a nucleotide sequence equivalent to any one of the nucleotide sequences of SEQ ID NOs: 5 to 7 and encoding a protein having glucose dehydrogenase activity.
[13] A recombinant DNA comprising the glucose dehydrogenase gene according to
[11] or
[12] .
[14] A microorganism carrying the recombinant DNA described in
[13] .
[15] A method for preparing glucose dehydrogenase, comprising the following steps (1) to (3): (1) preparing the glucose dehydrogenase gene according to
[11] or
[12] ; (2) expressing the gene; and (3) recovering the expression product.
[16] A method for screening glucose dehydrogenase, comprising the following steps (i) and (ii): (i) searching an amino acid sequence database for an amino acid sequence that shows 20% or more identity to the amino acid sequence of a known glucose dehydrogenase; and (ii) Among the hit amino acid sequences, the following sequence (SEQ ID NO: 1): [ka] (wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, H / N represents H (histidine) or N (asparagine), L / I / M represents L (leucine), I (isoleucine), or M (methionine), V / I represents V (valine) or I (isoleucine), G / A / S represents G (glycine), A (alanine), or S (serine), F / W / Y represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue).
[17] In the step (ii), the amino acid region further comprises: The condition is that the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues, i.e., (the value of the 8th amino acid residue + the value of the 12th amino acid residue + the value of the 13th amino acid residue) is 3.79 or less; The condition that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for the 8th, 11th, 12th, and 13th amino acid residues, i.e., (value of the 8th amino acid residue + value of the 11th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 2.83 or more; The screening method according to
[16] , wherein a substance satisfying the above is selected.
[18] In the step (ii), the amino acid region further comprises: The condition is that the sum of the values of the JANJ780101 (Average accessible surface area) listed in the AAindex database for the 3rd and 4th amino acid residues, i.e., (the value of the 3rd amino acid residue + the value of the 4th amino acid residue) is 56.9 or more. The screening method according to
[17] , wherein a substance satisfying the above is selected.
[19] The screening method according to any one of
[16] to
[18] , further comprising the following step (iii): (iii) selecting from the selected amino acid sequences those having the following conserved regions 1 to 10: Conserved region 1 consisting of the following amino acid sequence: [ka] where (F / Y) represents aromatic F or Y, (I / V) represents aromatic I or V, (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (G / A) represents branched aliphatic G or A, (V / A / S / T) represents small V, A, S or T, (S / A / G / C) represents very small S, A, G or C, (V / A / S / T) represents small V, A, S or T, (I / V / L) represents branched aliphatic I, V or L, (N / S) represents polar small N or S, (I / L) represents branched aliphatic I or L, (E / D) represents negative E or D; Conserved region 2 consisting of the following amino acid sequence: [ka] where (N / S / T) represents polar and small N, S, or T, (V / L / A) represents aliphatic V, L, or A, (I / V / L) represents branched aliphatic I, V, or L, (I / V / L) represents branched aliphatic I, V, or L, and (A / P / R) represents A, P, or R; Conserved region 3 consisting of the following amino acid sequence: [ka] where (A / T) represents hydrophobic, small A or T, (G / A) represents very small G or A, (R / K) represents positively charged R or K, (A / L) represents aliphatic A or L, (I / V / L / W) represents branched aliphatic or aromatic I, V, L or W, (G / A) represents very small G or A, (G / T) represents small G or T, (S / T) represents small S or T, (A / S / T) represents small A, S or T, and (I / V / F) represents branched aliphatic or aromatic I, V or F; Conserved region 4 consisting of the following amino acid sequence: [ka] where (A / V / C / E / D) represents small or negatively charged A, V, C, E, or D, (A / S) represents very small A or S, (A / Y) represents A or Y, (R / T / A) represents R, T, or A, (G / A) represents very small G or A, (Y / I / L) represents branched aliphatic or aromatic Y, I, or L, and (W / Y / F / L / A / H / K / Q) represents W, Y, F, L, A, H, K, or Q. Conserved region 5 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (I / V / L) represents branched aliphatic I, V or L, (S / A) represents very small S or A, (S / T / A) represents small S, T or A, (S / A) represents very small S or A, (I / L) represents branched aliphatic I or L, (K / R / A / G / I) represents aliphatic or positively charged (S / T) represents small S or T, (A / G / V / L / K / Q) represents A, G, V, L, K or Q, (I / V / L) represents branched aliphatic I, V or L, (I / V / L / A / H / R) represents aliphatic or positively charged I, V, L, A, H or R, (I / V) represents branched aliphatic I or V, (D / N / S) represents polar and small D, N or S; Conserved region 6 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (D / N) represents small D or N, (L / A / S / N) represents branched aliphatic or small L, A, S or N, (P / A / T) represents small P, A or T, (G / F / T) represents hydrophobic G, F or T, (E / S) represents hydrophilic polar E or S, (L / M) represents hydrophobic L or M, (Q / V) represents Q or V, (D / E) represents negatively charged D or E, and (Q / H) represents polar Q or H; Conserved region 7 consisting of the following amino acid sequence: [ka] however, (S / I / V / L) represents branched aliphatic or very small S, I, V, or L, (V / L / M / A / F) represents hydrophobic V, L, M, A, or F, (L / F) represents hydrophobic L or F, (A / S) represents very small A or S, (N / S / Y) represents N, S, or Y, (I / V / T) represents hydrophobic I, V, or T, and (I / V / L) represents branched aliphatic I, V, or L; Conserved region 8 consisting of the following amino acid sequence: [ka] where (I / M) represents hydrophobic I or M, (L / M) represents hydrophobic L or M, (P / S) represents small P or S, (R / K / E) represents charged R, K or E, (D / E / A / G / S / K) represents D, E, A, G, S or K, (I / L / M / A / N / K) represents I, L, M, A, N or K, (I / V) represents branched aliphatic I or V, and (D / N / S) represents polar and small D, N or S; Conserved region 9 consisting of the following amino acid sequence: [ka] where (Y / H) represents aromatic Y or H, (G / D) represents small G or D, (A / S / T / K) represents A, S, T or K, (L / V) represents branched aliphatic L or V, and (I / V) represents branched aliphatic I or V; Conserved region 10 consisting of the following amino acid sequence: [ka] where (Y / F) represents aromatic Y or F, (A / G) represents very small A or G, (I / L / V) represents branched aliphatic I, L, or V, (A / S) represents very small A or S, (R / K / L) represents branched aliphatic or positively charged R, K, or L, (A / I / T) represents hydrophobic A, I, or T, (A / S) represents very small A or S, (D / E) represents negatively charged D or E, (I / L / V / F / R / Q) represents I, L, V, F, R, or Q, (I / L) represents branched aliphatic I or L, and (K / E / Q) represents polar K, E, or Q.
[20] The screening method according to any one of
[16] to
[19] , further comprising the following step (iv): (iv) confirming the glucose dehydrogenase activity of the protein consisting of the selected amino acid sequence.
[21] A method for designing an engineered glucose dehydrogenase, comprising the following steps (I) and (II): (I) an amino acid sequence of a known glucose dehydrogenase or an amino acid sequence showing 20% or more identity to the amino acid sequence of a known glucose dehydrogenase, which is the following sequence (SEQ ID NO: 1): [ka] (wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, H / N represents H (histidine) or N (asparagine), L / I / M represents L (leucine), I (isoleucine), or M (methionine), V / I represents V (valine) or I (isoleucine), G / A / S represents G (glycine), A (alanine), or S (serine), F / W / Y represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue); and (II) A step of identifying a region in the prepared amino acid sequence that corresponds to the amino acid region, and then constructing a modified amino acid sequence by modifying the region so that it corresponds to the amino acid region.
[22] In step (II), the amino acid region further comprises: The condition is that the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues, i.e., (the value of the 8th amino acid residue + the value of the 12th amino acid residue + the value of the 13th amino acid residue) is 3.79 or less; The condition that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for the 8th, 11th, 12th, and 13th amino acid residues, i.e., (value of the 8th amino acid residue + value of the 11th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 2.83 or more; The design method described in
[21] is modified to satisfy the following.
[23] In step (II), the amino acid region further comprises: The condition is that the sum of the values of the JANJ780101 (Average accessible surface area) listed in the AAindex database for the 3rd and 4th amino acid residues, i.e., (the value of the 3rd amino acid residue + the value of the 4th amino acid residue) is 56.9 or more. The design method described in
[22] is modified to satisfy the following.
[24] The design method according to any one of
[21] to
[23] , further comprising the following step (III): (III) confirming whether the constructed modified amino acid sequence contains the following conserved regions 1 to 10, and if not, further modifying the sequence to contain them: Conserved region 1 consisting of the following amino acid sequence: [ka] where (F / Y) represents aromatic F or Y, (I / V) represents aromatic I or V, (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (G / A) represents branched aliphatic G or A, (V / A / S / T) represents small V, A, S or T, (S / A / G / C) represents very small S, A, G or C, (V / A / S / T) represents small V, A, S or T, (I / V / L) represents branched aliphatic I, V or L, (N / S) represents polar small N or S, (I / L) represents branched aliphatic I or L, (E / D) represents negative E or D; Conserved region 2 consisting of the following amino acid sequence: [ka] where (N / S / T) represents polar and small N, S, or T, (V / L / A) represents aliphatic V, L, or A, (I / V / L) represents branched aliphatic I, V, or L, (I / V / L) represents branched aliphatic I, V, or L, and (A / P / R) represents A, P, or R; Conserved region 3 consisting of the following amino acid sequence: [ka] where (A / T) represents hydrophobic, small A or T, (G / A) represents very small G or A, (R / K) represents positively charged R or K, (A / L) represents aliphatic A or L, (I / V / L / W) represents branched aliphatic or aromatic I, V, L or W, (G / A) represents very small G or A, (G / T) represents small G or T, (S / T) represents small S or T, (A / S / T) represents small A, S or T, and (I / V / F) represents branched aliphatic or aromatic I, V or F; Conserved region 4 consisting of the following amino acid sequence: [ka] where (A / V / C / E / D) represents small or negatively charged A, V, C, E, or D, (A / S) represents very small A or S, (A / Y) represents A or Y, (R / T / A) represents R, T, or A, (G / A) represents very small G or A, (Y / I / L) represents branched aliphatic or aromatic Y, I, or L, and (W / Y / F / L / A / H / K / Q) represents W, Y, F, L, A, H, K, or Q. Conserved region 5 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (I / V) represents branched aliphatic I or V, (I / V / L) represents branched aliphatic I, V or L, (S / A) represents very small S or A, (S / T / A) represents small S, T or A, (S / A) represents very small S or A, (I / L) represents branched aliphatic I or L, (K / R / A / G / I) represents aliphatic or positively charged (S / T) represents small S or T, (A / G / V / L / K / Q) represents A, G, V, L, K or Q, (I / V / L) represents branched aliphatic I, V or L, (I / V / L / A / H / R) represents aliphatic or positively charged I, V, L, A, H or R, (I / V) represents branched aliphatic I or V, (D / N / S) represents polar and small D, N or S; Conserved region 6 consisting of the following amino acid sequence: [ka] where (I / V) represents branched aliphatic I or V, (D / N) represents small D or N, (L / A / S / N) represents branched aliphatic or small L, A, S or N, (P / A / T) represents small P, A or T, (G / F / T) represents hydrophobic G, F or T, (E / S) represents hydrophilic polar E or S, (L / M) represents hydrophobic L or M, (Q / V) represents Q or V, (D / E) represents negatively charged D or E, and (Q / H) represents polar Q or H; Conserved region 7 consisting of the following amino acid sequence: [ka] however, (S / I / V / L) represents branched aliphatic or very small S, I, V, or L, (V / L / M / A / F) represents hydrophobic V, L, M, A, or F, (L / F) represents hydrophobic L or F, (A / S) represents very small A or S, (N / S / Y) represents N, S, or Y, (I / V / T) represents hydrophobic I, V, or T, and (I / V / L) represents branched aliphatic I, V, or L; Conserved region 8 consisting of the following amino acid sequence: [ka] where (I / M) represents hydrophobic I or M, (L / M) represents hydrophobic L or M, (P / S) represents small P or S, (R / K / E) represents charged R, K or E, (D / E / A / G / S / K) represents D, E, A, G, S or K, (I / L / M / A / N / K) represents I, L, M, A, N or K, (I / V) represents branched aliphatic I or V, and (D / N / S) represents polar and small D, N or S; Conserved region 9 consisting of the following amino acid sequence: [ka] where (Y / H) represents aromatic Y or H, (G / D) represents small G or D, (A / S / T / K) represents A, S, T or K, (L / V) represents branched aliphatic L or V, and (I / V) represents branched aliphatic I or V; Conserved region 10 consisting of the following amino acid sequence: [ka] where (Y / F) represents aromatic Y or F, (A / G) represents very small A or G, (I / L / V) represents branched aliphatic I, L, or V, (A / S) represents very small A or S, (R / K / L) represents branched aliphatic or positively charged R, K, or L, (A / I / T) represents hydrophobic A, I, or T, (A / S) represents very small A or S, (D / E) represents negatively charged D or E, (I / L / V / F / R / Q) represents I, L, V, F, R, or Q, (I / L) represents branched aliphatic I or L, and (K / E / Q) represents polar K, E, or Q.
[25] The design method according to any one of
[21] to
[24] , further comprising the following step (IV): A step of confirming the glucose dehydrogenase activity of the protein consisting of the constructed modified amino acid sequence.
[26] A method for preparing glucose dehydrogenase, comprising the following steps (1) to (3): (1) preparing a glucose dehydrogenase by the screening method according to any one of
[16] to
[20] or the design method according to any one of
[21] to
[25] ; (2) culturing the microorganism expressing the glucose dehydrogenase; and (3) recovering the expression product from the culture. [Brief explanation of the drawings]
[0007] [Figure 1] Measurement results of GDH activity. [Figure 2-1] Identity of amino acid sequences of FAD-GDH derived from various microorganisms. [Figure 2-2] Continued from Figure 2. DETAILED DESCRIPTION OF THE INVENTION
[0008] 1. Terminology As used herein, the term "isolated" is used interchangeably with "purified." In the case of a substance produced without the intervention of human manipulation, the term "isolated" is used to distinguish it from its natural state, i.e., the state that exists in nature, while in the case of a substance produced through the intervention of human manipulation, the term "isolated" is used to distinguish it from a substance that has not undergone an isolation or purification process. In the former case, the artificial isolation process results in an "isolated state" that is different from the natural state, and the isolated substance is clearly and decisively different from the natural product itself. On the other hand, in the latter case, impurities are typically removed or their amounts are reduced by the isolation or purification process, thereby increasing the purity. The purity of the isolated enzyme is not particularly limited. However, if the enzyme is intended to be used in a manner that requires high purity, it is preferable that the isolated enzyme have a high purity.
[0009] In this specification, amino acid sequences are represented according to conventional notation, with the left end representing the amino terminus and the right end representing the carboxy terminus. Each amino acid is represented by a single letter as follows: A: Alanine, C: Cysteine, D: Aspartic acid, E: Glutamic acid, F: Phenylalanine, G: Glycine, H: Histidine, I: Isoleucine, K: Lysine, L: Leucine, M: Methionine, N: Asparagine, P: Proline, Q: Glutamine, R: Arginine, S: Serine, T: Threonine, V: Valine, W: Tryptophan, Y: Tyrosine
[0010] 2. Glucose dehydrogenase (GDH) A first aspect of the present invention provides a GDH. The GDH of the present invention (hereinafter also referred to as "the present enzyme") is characterized by two requirements (the first and second requirements). The first requirement is that it has a unique amino acid region. The unique amino acid region in the present invention is an amino acid region that has been found to be characteristic of GDH and is composed of the following sequence (SEQ ID NO: 1): [ka] wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, (H / N) represents H (histidine) or N (asparagine), (L / I / M) represents L (leucine), I (isoleucine), or M (methionine), (V / I) represents V (valine) or I (isoleucine), (G / A / S) represents G (glycine), A (alanine), or S (serine), (F / W / Y) represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue. Throughout this specification, the rule defining this sequence (SEQ ID NO: 1) will be referred to as "Rule 1" for ease of explanation.
[0011] The amino acid region characterizing this enzyme corresponds to a sequence of 13 consecutive amino acid residues located in the vicinity of positions 140 to 180 of the filamentous fungal FAD-GDH. For example, in the known FAD-GDH Aspergillus oryzae-derived FAD-GDH (SEQ ID NO: 21), this amino acid region extends from the 156th to 168th amino acid residues, and similarly, in the known FAD-GDH Aspergillus terreus-derived FAD-GDH (SEQ ID NO: 22), this amino acid region extends from the 149th to 161st amino acid residues. On the other hand, in the FAD-GDH derived from Aspergillus cristatus (SEQ ID NO: 2) described below, the above amino acid region is from the 149th amino acid residue to the 161st amino acid residue; in the FAD-GDH derived from Aspergillus turcosus (SEQ ID NO: 3), the above amino acid region is from the 149th amino acid residue to the 161st amino acid residue; and in the FAD-GDH derived from Corynascus sepedonium (SEQ ID NO: 4), the above amino acid region is from the 151st amino acid residue to the 163rd amino acid residue.
[0012] The amino acid region that characterizes this enzyme is thought to generally form an extended structure, but near the center, a proline residue (amino acid residue number 6) creates an angle, forming a slight secondary structure (loop-like structure). The first half is somewhat distant from the active center and is thought to be involved in maintaining the structure on the surface of the enzyme. On the other hand, the second half is close to the pathway through which the substrate moves toward the active center and is thought to be involved in the characteristics of the enzyme activity.
[0013] It is preferable that the amino acid region that characterizes the present enzyme satisfies the following rule (for convenience of explanation, referred to as "Rule 2") in addition to Rule 1. Rule 2 consists of the following two conditions. <Rule 2, Condition 1> The condition is that the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues is 3.79 or less, i.e., (value of the 8th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue). <Rule 2, Condition 2> The condition is that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for amino acid residues 8, 11, 12, and 13, i.e., (value of amino acid residue 8 + value of amino acid residue 11 + value of amino acid residue 12 + value of amino acid residue 13) is 2.83 or more.
[0014] Condition 1 of Rule 2 specifies the amino acids that constitute part of the pathway for the substrate toward the active center. It is believed that these amino acids are required to maintain a certain degree of structure. Furthermore, since amino acids with high values for this condition are disadvantageous for the formation and maintenance of an extended structure, this condition is an important indicator of the relevant portion. Condition 1 of Rule 2 preferably satisfies the condition that the sum of the above values is 3.10 to 3.79.
[0015] On the other hand, condition 2 of rule 2 also specifies the amino acids that constitute part of the pathway for the substrate toward the active center. It is believed that these amino acids are required to maintain a certain degree of structure. Furthermore, since amino acids with low numerical values for this condition are disadvantageous for the formation and maintenance of an extended structure, this condition is an important indicator of the relevant part. Condition 2 of rule 2 is preferably a condition where the sum of the numerical values satisfies 2.87 or more, more preferably a condition where the sum satisfies 2.87 to 5.22, and even more preferably a condition where the sum satisfies 2.87 to 3.55.
[0016] The values used for each condition are as follows: [Table 1]
[0017] [Table 2]
[0018] It is more preferable that the amino acid region characterizing the present enzyme satisfies the following rule (for convenience of explanation, referred to as "Rule 3") in addition to Rules 1 and 2. Rule 3 consists of the following conditions. <Conditions for Rule 3> The condition is that the sum of the values of the JANJ780101 (Average accessible surface area) listed in the database AAindex for the 3rd and 4th amino acid residues, i.e., (the value of the 3rd amino acid residue + the value of the 4th amino acid residue) is 66 or more.
[0019] The 3rd and 4th amino acid residues are thought to be present on the protein surface. This condition is an important indicator of their presence on the protein surface. As a condition for rule 3, the sum of the above numerical values is preferably 100 or more, more preferably 116 or more, even more preferably 116 to 172, and particularly preferably 116 to 131.
[0020] The values used for the conditions are as follows: [Table 3]
[0021] The second requirement for characterizing the present enzyme is that it contains an amino acid sequence that is 40% or more identical to any of the amino acid sequences set forth in SEQ ID NOs: 2 to 4. Preferably, the identity of the amino acid sequence here is 65% or more with the amino acid sequence of SEQ ID NO: 2, 85% or more with the amino acid sequence of SEQ ID NO: 3, and 55% or more with the amino acid sequence of SEQ ID NO: 4. Each of the amino acid sequences set forth in SEQ ID NOs: 2 to 4 has been found to have GDH activity through studies by the present inventors. Specifically, the amino acid sequence of SEQ ID NO: 2 corresponds to FAD-GDH derived from Aspergillus cristatus, the amino acid sequence of SEQ ID NO: 3 corresponds to FAD-GDH derived from Aspergillus turcosus, and the amino acid sequence of SEQ ID NO: 4 corresponds to FAD-GDH derived from Corynascus sepedonium.
[0022] The higher the identity with the reference amino acid sequence (any of SEQ ID NOS: 2 to 4), the more preferable it is in terms of having properties similar to those of a protein consisting of any of the amino acid sequences of SEQ ID NOS: 2 to 4. Therefore, the identity is preferably 50% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, 98% or more, or 100% (i.e., the reference amino acid sequence itself) (the higher the percentage of identity, the more preferable it is).
[0023] In the present application, 11 regions, including the unique region of the present invention (SEQ ID NO: 1) and conserved regions 1 to 10 described below, are identified as highly conserved regions in GDH. In regions excluding these regions, differences in amino acid sequence have little effect on GDH activity, allowing for greater variations in amino acid sequence. The identity of the region excluding the above 11 regions with the reference amino acid sequence (any of SEQ ID NOs: 2 to 4) is 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more (the higher the percentage of identity, the more preferable). Furthermore, it is preferable for the region to exhibit 60% or more identity with the amino acid sequence of SEQ ID NO: 2, 80% or more identity with the amino acid sequence of SEQ ID NO: 3, and 40% or more identity with the amino acid sequence of SEQ ID NO: 4.
[0024] The total length of the enzyme is not particularly limited, but is, for example, 530 aa to 660 aa, preferably 540 aa to 620 aa.
[0025] An amino acid sequence that exhibits the above-mentioned identity (at least 40% identity to the reference amino acid sequence) will have some differences from the reference amino acid sequence, except in cases where the identity is 100%. A "partial difference in the amino acid sequence" can occur, for example, by deletion or substitution of one or more amino acids among the amino acids that make up the amino acid sequence, addition or insertion of one or more amino acids to the amino acid sequence, or any combination thereof. The position at which the amino acid sequence differs is not particularly limited, as long as it is a part other than the amino acid region that satisfies the first requirement above. Furthermore, differences in the amino acid sequence may occur at multiple positions (locations).
[0026] A typical example of a "partial difference in the amino acid sequence" is a mutation (change) in the amino acid sequence caused by deletion or substitution of 1 to 50 (preferably 1 to 10, more preferably 1 to 7, even more preferably 1 to 5, and even more preferably 1 to 3) amino acids among the amino acids that make up the amino acid sequence, or addition or insertion of 1 to 50 (preferably 1 to 10, more preferably 1 to 7, even more preferably 1 to 5, and even more preferably 1 to 3) amino acids to the amino acid sequence, or a combination thereof.
[0027] Preferably, the amino acid sequence exhibiting the above identity is obtained by making conservative amino acid substitutions at amino acid residues that are not essential for GDH activity. Here, "conservative amino acid substitution" refers to the substitution of an amino acid residue with an amino acid residue having a side chain with similar properties. Amino acid residues are classified into several families based on their side chains, such as basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Conservative amino acid substitutions are preferably between amino acid residues within the same family. It is preferred that the amino acid residues at the active center are maintained. The positions of the amino acid residues in the active center can be estimated to be the residues corresponding to H at position 502 and H at position 545 in Aspergillus cristatus FAD-GDH (SEQ ID NO: 2), the residues corresponding to H at position 506 and H at position 549 in Aspergillus tarcosus FAD-GDH (SEQ ID NO: 3), or the residues corresponding to H at position 510 and H at position 553 in Corneascus cepedonium FAD-GDH (SEQ ID NO: 4).
[0028] The percent identity between two amino acid sequences or two nucleic acids (hereinafter, the term "two sequences" is used to encompass both) can be determined, for example, by the following procedure. First, the two sequences are aligned to enable optimal comparison (for example, gaps may be introduced into the first sequence to optimize alignment with the second sequence). When a molecule (amino acid residue or nucleotide) at a specific position in the first sequence is the same as a molecule at the corresponding position in the second sequence, the molecules at that position can be said to be identical. The identity between two sequences is a function of the number of identical positions shared by the two sequences (i.e., identity (%) = number of identical positions / total number of positions × 100), and preferably, the number and size of gaps required for optimal alignment are also taken into account.
[0029] Comparison of two sequences and determination of identity can be achieved using a mathematical algorithm. An example of a mathematical algorithm that can be used for sequence comparison is the algorithm described in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-77. Such an algorithm is incorporated into the NBLAST program and XBLAST program (version 2.0) described in Altschul et al. (1990) J. Mol. Biol. 215:403-10. To obtain a nucleotide sequence equivalent to the nucleic acid molecule of the present invention, for example, a BLAST nucleotide search can be performed using the NBLAST program with a score of 100 and a word length of 12. To obtain an amino acid sequence equivalent to the present enzyme, for example, a BLAST polypeptide search can be performed using the XBLAST program with a score of 50 and a word length of 3. To obtain gapped alignments for comparison, Gapped BLAST, as described in Altschul et al. (1997) Amino Acids Research 25(17):3389-3402, can be used. When using BLAST and Gapped BLAST, the default parameters of the corresponding programs (e.g., XBLAST and NBLAST) can be used. For details, see http: / / www.ncbi.nlm.nih.gov. An example of another mathematical algorithm that can be used for sequence comparison is the algorithm described by Myers and Miller (1988) Comput Appl Biosci. 4:11-17. Such an algorithm is incorporated into the ALIGN program, available, for example, on the GENESTREAM network server (IGH Montpellier, France) or the ISREC server. When using the ALIGN program for comparing amino acid sequences, for example, the PAM120 residue mass table can be used, with a gap length penalty of 12 and a gap penalty of 4.
[0030] The identity of two amino acid sequences can be determined using the GAP program in the EMBOSS package, using a Blosum 62 matrix with a gap weight of 10 and a gap length weight of 2. The identity of two nucleotide sequences can be determined using the GAP program in the EMBOSS package (available at http: / / emboss.open-bio.org / ) with a gap weight of 50 and a gap length weight of 3.
[0031] The enzyme may be part of a larger protein (e.g., a fusion protein). Additional sequences in the fusion protein include, for example, sequences useful for purification, such as multiple histidine residues, additional sequences ensuring stability during recombinant production, and cytochrome-like additional sequences that alter reactivity. Furthermore, chemically attached sequences such as polymers, electron acceptors, and metal complexes may also be included.
[0032] Preferably, the enzyme is characterized by requirement 3 in addition to requirements 1 and 2 above. Requirement 3: Has one or more of the following conserved regions 1 to 10.
[0033] Based on the Venn diagram of William Ramsay Taylor, amino acids were classified into the following groups and used to identify amino acid residues that make up the conserved regions. Aromatic:F / W / Y / H Aliphatic:A / I / V / L / G Branched chain aliphatic: I / V / L Extra small size (extra small): G / A / S / C Small size (small): G / A / S / C / T / V / N / P / D Polarity: T / C / S / N / Q / D / E / Y / W / H / K / R Polarity small size (polarity small): S / N / C / T / D Charge: D / E / K / R / H Negative charge: D / E Positive charge: R / K / H Hydrophobicity:A / G / C / T / I / V / L / K / H / Y / W / F / M Hydrophobic and small size (hydrophobic small): A / T / G / C / V Hydrophilic and polar (hydrophilic polar): S / N / Q / D / E / R
[0034] <Save area 1> Conserved region 1 corresponds to the amino acid residues 5 to 25 in A. cristatus GZAAS20.1005-derived FAD-GDH, the amino acid residues 5 to 25 in A. turcosus HMR AF 23-derived FAD-GDH, and the amino acid residues 8 to 28 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 1 is composed of the following sequence (SEQ ID NO: 24): [ka] however, (F / Y) represents aromatic F or Y; (I / V) represents aromatic I or V; (I / V) represents branched chain aliphatic I or V; (I / V) represents branched chain aliphatic I or V; (G / A) represents branched aliphatic G or A; (V / A / S / T) represents a small V, A, S, or T. (S / A / G / C) represents a minimum of S, A, G, or C; (V / A / S / T) represents a small V, A, S, or T. (I / V / L) represents branched aliphatic I, V, or L; (N / S) indicates polarity and small N or S. (I / L) represents branched aliphatic I or L; (E / D) indicates a negative charge of E or D.
[0035] Preferably, the amino acid sequence defining conserved region 1 consists of the following sequence (SEQ ID NO: 25): [ka] however, (F / Y) represents aromatic F or Y; (I / V) represents branched chain aliphatic I or V; (I / V) represents branched chain aliphatic I or V; (I / V) represents branched chain aliphatic I or V; (G / A) represents the minimum of G or A, (T / A) represents a small T or A, (S / C) indicates the minimum S or C, (V / A) represents a small V or A, (I / V / L) represents branched aliphatic I, V, or L; (E / D) indicates a negative charge of E or D.
[0036] More preferably, the amino acid sequence defining conserved region 1 consists of the following sequence (SEQ ID NO: 26): [ka] however, (I / V) represents branched chain aliphatic I or V; (I / V) represents branched chain aliphatic I or V; (T / A) represents a small T or A, (S / C) indicates the minimum S or C, (V / A) represents a small V or A, (I / V / L) represents branched aliphatic I, V, or L; (E / D) indicates a negative charge of E or D.
[0037] More preferably, the amino acid sequence defining conserved region 1 consists of any of the following sequences: YDYVIVGGGTSGLVVANRLSE (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 27) YDYIVVGGGACGLALANRLSD (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 28) FDYIIIGAGTSGLVIANRLSE (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 29)
[0038] <Save area 2> Conserved region 2 corresponds to the amino acid residues 30 to 37 in A. cristatus GZAAS20.1005-derived FAD-GDH, the amino acid residues 30 to 37 in A. turcosus HMR AF 23-derived FAD-GDH, and the amino acid residues 33 to 40 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 2 is composed of the following sequence (SEQ ID NO: 30): [ka] however, (N / S / T) indicates polarity - small N, S or T, (V / L / A) represents aliphatic V, L, or A; (I / V / L) represents branched aliphatic I, V, or L; (I / V / L) represents branched aliphatic I, V, or L; (A / P / R) indicates that A, P or R is.
[0039] Preferably, the amino acid sequence defining conserved region 2 consists of the following sequence (SEQ ID NO: 31): [ka] however, (S / T) represents S or T, (A / V / L) represents A, V, or L; (V / I) represents V or I, (V / I) represents V or I, (A / R) indicates A or R.
[0040] More preferably, the amino acid sequence defining conserved region 2 consists of the following sequence (SEQ ID NO: 32): [ka] however, (V / L) represents V or L, (V / I) represents V or I, (V / I) represents V or I, (A / R) indicates A or R.
[0041] More preferably, the amino acid sequence defining conserved region 2 consists of any of the following sequences: SVVVIEAG (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 33) SVLIIERG (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 34) TVAVIEPG (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 35)
[0042] <Save area 3> Conserved region 3 corresponds to amino acid residues 81 to 93 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 81 to 93 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 83 to 95 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 3 is composed of the following sequence (SEQ ID NO: 36): [ka] however, (A / T) represents a hydrophobic small A or T. (G / A) represents the minimum of G or A, (R / K) represents a positive charge of R or K, (A / L) represents aliphatic A or L, (I / V / L / W) represents branched aliphatic or aromatic I, V, L, or W; (G / A) represents the minimum of G or A, (G / T) represents a small G or T, (S / T) means small S or T, (A / S / T) represents a small A, S, or T. (I / V / F) represents branched aliphatic or aromatic I, V or F.
[0043] Preferably, the amino acid sequence defining conserved region 3 consists of the following sequence (SEQ ID NO: 37): [ka] however, (W / L) represents branched aliphatic or aromatic W or L; (S / T) indicates a small S or T.
[0044] More preferably, the amino acid sequence defining conserved region 3 consists of the following sequence (SEQ ID NO: 38): [ka] however, (S / T) indicates a small S or T.
[0045] More preferably, the amino acid sequence defining conserved region 3 consists of any of the following sequences: AGKALGGTSTING (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 39) AGKALGGTTTING (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 40) AGKAWGGTSTING (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 41)
[0046] <Save area 4> Conserved region 4 corresponds to amino acid residues 208 to 218 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 209 to 219 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 211 to 221 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 4 is composed of the following sequence (SEQ ID NO: 42): [ka] however, (A / V / C / E / D) represents a small or negatively charged A, V, C, E, or D; (A / S) represents the smallest A or S, (A / Y) means that A or Y is (R / T / A) represents R, T or A, (G / A) represents the minimum of G or A, (Y / I / L) represents branched aliphatic or aromatic Y, I, or L; (W / Y / F / L / A / H / K / Q) represents W, Y, F, L, A, H, K or Q.
[0047] Preferably, the amino acid sequence defining conserved region 4 consists of the following sequence (SEQ ID NO: 43): [ka] however, (E / C) represents a small or negative charge of E or C; (A / S) represents the smallest A or S, (Y / L) represents branched aliphatic or aromatic Y or L; (W / Y / H) represents aromatic W, Y or H.
[0048] More preferably, the amino acid sequence defining conserved region 4 consists of the following sequence (SEQ ID NO: 44): [ka] however, (W / Y) represents aromatic W or Y.
[0049] More preferably, the amino acid sequence defining conserved region 4 is composed of any of the following sequences: REDAARAYYWP (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 45) REDAARAYYYP (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 46) RCDSARAYLHP (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 47)
[0050] <Save area 5> Conserved region 5 corresponds to amino acid residues 268 to 289 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 271 to 292 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 272 to 293 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 5 is composed of the following sequence (SEQ ID NO: 48): [ka] however, (I / V) represents branched chain aliphatic I or V; (I / V) represents branched chain aliphatic I or V; (I / V / L) represents branched aliphatic I, V, or L; (S / A) represents the smallest S or A, (S / T / A) represents a small S, T, or A. (S / A) represents the smallest S or A, (I / L) represents branched aliphatic I or L; (K / R / A / G / I) represents aliphatic or positively charged K, R, A, G, or I; (S / T) represents a small S or T, (A / G / V / L / K / Q) means A, G, V, L, or K, Q. (I / V / L) represents branched aliphatic I, V, or L; (I / V / L / A / H / R) represents aliphatic or positively charged I, V, L, A, H, or R; (I / V) represents branched chain aliphatic I or V; (D / N / S) indicates polarity or smallness, D, N or S.
[0051] Preferably, the amino acid sequence defining conserved region 5 consists of the following sequence (SEQ ID NO: 49): [ka] however, (I / V) represents branched chain aliphatic I or V; (T / A) represents a small T or A, (S / A) represents the smallest S or A, (I / L) represents branched aliphatic I or L; (R / I) represents branched aliphatic or positively charged R or I; (T / S) indicates polarity and small T or S. (A / L) represents aliphatic A or L, (L / I / V) represents branched aliphatic L, I, or V; (L / R / A) represents aliphatic or positively charged L, R or A.
[0052] More preferably, the amino acid sequence defining conserved region 5 consists of the following sequence (SEQ ID NO: 50): [ka] however, (T / A) represents a small T or A, (I / L) represents branched aliphatic I or L; (R / I) represents branched aliphatic or positively charged R or I; (T / S) indicates polarity and small T or S. (L / I) represents branched chain aliphatic L or I; (L / R) represents aliphatic or positively charged L or R.
[0053] More preferably, the amino acid sequence defining conserved region 5 consists of any of the following sequences: EVILSTGSIRTPALLELSGVGN (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 51) EVILSAGSLISPAILERSGVGN (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 52) EVVLSAGALRTPLVLEASGVGN (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 53)
[0054] <Save area 6> Conserved region 6 corresponds to amino acid residues 302 to 314 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 305 to 317 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 306 to 318 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 6 is composed of the following sequence (SEQ ID NO: 54): [ka] however, (I / V) represents branched chain aliphatic I or V; (D / N) represents a small D or N. (L / A / S / N) represents branched chain aliphatic or small L, A, S, or N; (P / A / T) represents a small P, A, or T; (G / F / T) represents hydrophobic G, F or T; (E / S) represents hydrophilic polar E or S, (L / M) represents hydrophobic L or M, (Q / V) represents Q or V, (D / E) represents a negative charge of D or E, (Q / H) indicates the polarity Q or H.
[0055] Preferably, the amino acid sequence defining conserved region 6 consists of the following sequence (SEQ ID NO: 55): [ka] however, (V / I) represents branched aliphatic V or I; (D / N) represents a small D or N. (P / T / A) represents a small P, T, or A; (T / G) represents a hydrophobic small T or G; (Q / V) represents Q or V, (D / E) indicates a negative charge of D or E.
[0056] More preferably, the amino acid sequence defining conserved region 6 consists of the following sequence (SEQ ID NO: 56): [ka] however, (V / I) represents branched aliphatic V or I; (D / N) represents a small D or N. (P / T / A) indicates a small P, T or A.
[0057] More preferably, the amino acid sequence defining the conserved region 6 is composed of any of the following sequences: VDLPTVGENLQDQ (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 57) INLATVGENLQDQ (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 58) IDLPGVGENLVEQ (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 59)
[0058] <Save area 7> Conserved region 7 corresponds to amino acid residues 415 to 425 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 418 to 418 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 420 to 430 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 7 is composed of the following sequence (SEQ ID NO: 60): [ka] however, (S / I / V / L) represents branched aliphatic or tiny S, I, V or L; (V / L / M / A / F) represents hydrophobic V, L, M, A or F; (L / F) represents hydrophobic L or F, (A / S) represents the smallest A or S, (N / S / Y) represents N, S or Y; (I / V / T) represents hydrophobic I, V, or T; (I / V / L) represents branched chain aliphatic I, V or L.
[0059] Preferably, the amino acid sequence defining conserved region 7 consists of the following sequence (SEQ ID NO: 61): [ka] however, (L / F) represents hydrophobic L or F, (S / A) represents the smallest S or A, (N / S) represents N or S, (I / L) represents branched chain aliphatic I or L.
[0060] More preferably, the amino acid sequence defining conserved region 7 consists of the following sequence (SEQ ID NO: 62): [ka] however, (S / A) indicates the smallest S or A.
[0061] More preferably, the amino acid sequence defining conserved region 7 is composed of any of the following sequences: LLPFSRGNVHI (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 63) LLPFARGNVHI (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 64) LFPFSRGSVHL (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 65)
[0062] <Save area 8> Conserved region 8 corresponds to amino acid residues 509 to 519 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 513 to 523 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 517 to 527 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 8 is composed of the following sequence (SEQ ID NO: 66): [ka] however, (I / M) represents hydrophobic I or M, (L / M) represents hydrophobic L or M, (P / S) indicates a small P or S. (R / K / E) represents the charge R, K, or E; (D / E / A / G / S / K) means D, E, A, G, S or K, (I / L / M / A / N / K) represents I, L, M, A, N or K, (I / V) represents branched chain aliphatic I or V; (D / N / S) indicates polarity or smallness, D, N or S.
[0063] Preferably, the amino acid sequence defining conserved region 8 consists of the following sequence (SEQ ID NO: 67): [ka] however, (S / D / E) represents a minimal or negative charge of S, D, or E; (M / L) represents hydrophobic M or L, (S / D) indicates polarity, small S or D.
[0064] More preferably, the amino acid sequence defining conserved region 8 consists of the following sequence (SEQ ID NO: 68): [ka] however, (S / D / E) represents a minimal or negative charge of S, D, or E; (S / D) indicates polarity, small S or D.
[0065] More preferably, the amino acid sequence defining the conserved region 8 is composed of any of the following sequences: MMPRSMGGVVS (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 69) MMPREMGGVVD (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 70) MMPRALGGVVD (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 71)
[0066] <Save area 9> Conserved region 9 corresponds to amino acid residues 524 to 536 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 528 to 540 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 532 to 544 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 9 is composed of the following sequence (SEQ ID NO: 72): [ka] however, (Y / H) represents aromatic Y or H; (G / D) represents a small G or D. (A / S / T / K) represents A, S, T or K, (L / V) represents branched chain aliphatic L or V; (I / V) represents I or V of branched chain aliphatic.
[0067] Preferably, the amino acid sequence defining conserved region 9 consists of the following sequence (SEQ ID NO: 73): [ka] however, (H / Y) represents aromatic H or Y; (S / A) represents the smallest S or A, (I / V) represents I or V of branched chain aliphatic.
[0068] More preferably, the amino acid sequence defining conserved region 9 consists of any of the following sequences: VHGTSNVRIVDAS (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 74) VYGTANVRVVDAS (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 75) VYGTANVRVVDAS (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 76)
[0069] <Save area 10> Conserved region 10 corresponds to amino acid residues 551 to 562 in A. cristatus GZAAS20.1005-derived FAD-GDH, amino acid residues 555 to 566 in A. turcosus HMR AF 23-derived FAD-GDH, and amino acid residues 559 to 570 in Corynascus sepedonium NBRC 31363-derived FAD-GDH. The amino acid sequence defining conserved region 10 is composed of the following sequence (SEQ ID NO: 77): [ka] however, (Y / F) represents aromatic Y or F; (A / G) represents the smallest A or G, (I / L / V) represents branched aliphatic I, L, or V; (A / S) represents the smallest A or S, (R / K / L) represents branched aliphatic or positively charged R, K, or L; (A / I / T) represents hydrophobic A, I or T; (A / S) represents the smallest A or S, (D / E) represents a negative charge of D or E, (I / L / V / F / R / Q) represents I, L, V, F, R or Q, (I / L) represents branched aliphatic I or L; (K / E / Q) indicates the polarity K, E or Q.
[0070] Preferably, the amino acid sequence defining the conserved region 10 consists of the following sequence (SEQ ID NO: 78): [ka] however, (A / G) represents the smallest A or G, (I / V / L) represents branched aliphatic I, V, or L; (S / A) represents the smallest S or A, (L / I) represents branched chain aliphatic L or I.
[0071] More preferably, the amino acid sequence defining the conserved region 10 consists of the following sequence (SEQ ID NO: 79): [ka] however, (I / V) represents branched chain aliphatic I or V; (S / A) represents the smallest S or A, (L / I) represents branched chain aliphatic L or I.
[0072] More preferably, the amino acid sequence defining the conserved region 10 is composed of any of the following sequences: YAIAERASDLIK (FAD-GDH from A. cristatus GZAAS20.1005) (SEQ ID NO: 80) YAVAERAADIIK (FAD-GDH from A. turcosus HMR AF 23) (SEQ ID NO: 81) YGLAERASDIIK (FAD-GDH from Corynascus sepedonium NBRC 31363) (SEQ ID NO: 82)
[0073] The number of conserved regions is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. Usually, the more conserved regions there are, the better. Therefore, typically, the more conserved regions there are, the more preferred the enzyme will be, and in the most preferred embodiment, it will have all of conserved regions 1 to 10.
[0074] This enzyme can be easily prepared by genetic engineering techniques. For example, it can be prepared by transforming a suitable host cell (e.g., Escherichia coli) with DNA encoding this enzyme and recovering the protein expressed in the transformant. The recovered protein is then purified as appropriate depending on the purpose. Obtaining this enzyme as a recombinant protein in this way allows for various modifications. For example, by inserting the DNA encoding this enzyme and other appropriate DNA into the same vector and producing a recombinant protein using that vector, it is possible to obtain this enzyme as a recombinant protein to which any peptide or protein is linked. Furthermore, modifications such as the addition of sugar chains and / or lipids, or modifications that cause N- or C-terminal processing, may also be performed. These modifications can simplify the extraction and purification of the recombinant protein, or add biological functions, etc.
[0075] 3. Nucleic acids encoding glucose dehydrogenase (GDH) A second aspect of the present invention provides a nucleic acid related to the present enzyme. That is, a gene encoding the present enzyme is provided. The gene encoding the present enzyme is typically used to prepare the present enzyme. A genetic engineering preparation method using the gene encoding the present enzyme makes it possible to obtain the present enzyme in a more homogeneous state. This method is also suitable for preparing large quantities of the present enzyme. Note that the use of the gene encoding the present enzyme is not limited to preparing the present enzyme. For example, the nucleic acid can also be used as an experimental tool for the purpose of elucidating the mechanism of action of the present enzyme, or as a tool for designing or creating further mutants of the enzyme.
[0076] As used herein, the term "gene encoding the present enzyme" refers to a nucleic acid that, when expressed, gives the present enzyme, and includes not only a nucleic acid having a base sequence corresponding to the amino acid sequence of the present enzyme, but also a nucleic acid in which a sequence that does not encode an amino acid sequence is added to such a nucleic acid. Codon degeneracy is also taken into consideration.
[0077] Examples of gene sequences encoding this enzyme are shown in SEQ ID NO: 5 (sequence encoding FAD-GDH derived from Aspergillus cristatus), SEQ ID NO: 6 (sequence encoding FAD-GDH derived from Aspergillus tarcosus), and SEQ ID NO: 7 (sequence encoding FAD-GDH derived from Corneascus cepedonium).
[0078] The nucleic acids of the present invention can be prepared in an isolated state by using standard genetic engineering techniques, molecular biological techniques, biochemical techniques, etc., with reference to the sequence information disclosed in this specification or the attached sequence listing.
[0079] In another aspect of the present invention, there is provided a nucleic acid (hereinafter also referred to as an "equivalent nucleic acid"; a nucleotide sequence specifying an equivalent nucleic acid is also referred to as an "equivalent nucleotide sequence") that, when compared with the nucleotide sequence of the gene encoding the present enzyme, encodes a protein that has the same function but a different nucleotide sequence in part. Examples of equivalent nucleic acids include DNA that consists of a nucleotide sequence containing one or more nucleotide substitutions, deletions, insertions, additions, or inversions based on the nucleotide sequence of the nucleic acid encoding the present enzyme, and encodes a protein that has the enzymatic activity characteristic of the present enzyme (i.e., GDH activity). Base substitutions or deletions may occur at multiple sites. Here, "multiple" refers to, for example, 2 to 40 nucleotides, preferably 2 to 20 nucleotides, and more preferably 2 to 10 nucleotides, although this varies depending on the position and type of amino acid residues in the three-dimensional structure of the protein encoded by the nucleic acid.
[0080] An equivalent nucleic acid has 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 92% or more, 95% or more, or 99% or more identity to a reference base sequence (any of SEQ ID NOS: 5 to 7) (the higher the percentage of identity, the more preferable it is). Preferably, the equivalent nucleic acid has a portion encoding the amino acid sequence of 11 regions, i.e., the unique region and conserved regions 1 to 10, of the present invention.Therefore, in a preferred embodiment, the equivalent nucleic acid includes a nucleotide sequence encoding the amino acid sequence of the unique region (SEQ ID NO: 1), a nucleotide sequence encoding the amino acid sequence of conserved region 1 (e.g., the amino acid sequence of SEQ ID NO: 24, preferably the amino acid sequence of SEQ ID NO: 25, more preferably the amino acid sequence of SEQ ID NO: 26), a nucleotide sequence encoding the amino acid sequence of conserved region 2 (e.g., the amino acid sequence of SEQ ID NO: 30, preferably the amino acid sequence of SEQ ID NO: 31, more preferably the amino acid sequence of SEQ ID NO: 32), a nucleotide sequence encoding the amino acid sequence of conserved region 3 (e.g., the amino acid sequence of SEQ ID NO: 36, preferably the amino acid sequence of SEQ ID NO: 37, more preferably the amino acid sequence of SEQ ID NO: 38), a nucleotide sequence encoding the amino acid sequence of conserved region 4 (e.g., the amino acid sequence of SEQ ID NO: 42, preferably the amino acid sequence of SEQ ID NO: 43, more preferably the amino acid sequence of SEQ ID NO: 44), a nucleotide sequence encoding the amino acid sequence of conserved region 5 (e.g., the amino acid sequence of SEQ ID NO: 48, preferably the amino acid sequence of SEQ ID NO: 49, more preferably the amino acid sequence of SEQ ID NO: 50), a nucleotide sequence encoding the amino acid sequence of conserved region 6 (e.g., the amino acid sequence of SEQ ID NO: 51, more preferably the amino acid sequence of SEQ ID NO: 52), a nucleotide sequence encoding the amino acid sequence of conserved region 7 (e.g., the amino acid sequence of SEQ ID NO: 53, more preferably the amino acid sequence of SEQ ID NO: 54), a nucleotide sequence encoding the amino acid sequence of conserved region 8 (e.g., the amino acid sequence of SEQ ID NO: 55, more preferably a nucleotide sequence encoding the amino acid sequence of conserved region 6 (for example, the amino acid sequence of SEQ ID NO: 54, preferably the amino acid sequence of SEQ ID NO: 55, more preferably the amino acid sequence of SEQ ID NO: 56), a nucleotide sequence encoding the amino acid sequence of conserved region 7 (for example, the amino acid sequence of SEQ ID NO: 60, preferably the amino acid sequence of SEQ ID NO: 61, more preferably the amino acid sequence of SEQ ID NO: 62), a nucleotide sequence encoding the amino acid sequence of conserved region 8 (for example, the amino acid sequence of SEQ ID NO: 66, preferably the amino acid sequence of SEQ ID NO: 67, more preferably the amino acid sequence of SEQ ID NO: 68), a nucleotide sequence encoding the amino acid sequence of conserved region 9 (for example, the amino acid sequence of SEQ ID NO: 72, preferably the amino acid sequence of SEQ ID NO: 73), and a nucleotide sequence encoding the amino acid sequence of conserved region 10 (for example, the amino acid sequence of SEQ ID NO: 77, preferably the amino acid sequence of SEQ ID NO: 78, more preferably the amino acid sequence of SEQ ID NO: 79).
[0081] Such equivalent nucleic acids can be obtained, for example, by restriction enzyme treatment, treatment with exonuclease or DNA ligase, or by introducing mutations using site-directed mutagenesis (Molecular Cloning, Third Edition, Chapter 13, Cold Spring Harbor Laboratory Press, New York) or random mutagenesis (Molecular Cloning, Third Edition, Chapter 13, Cold Spring Harbor Laboratory Press, New York). Equivalent nucleic acids can also be obtained by other methods, such as ultraviolet irradiation.
[0082] The present invention further provides recombinant DNA containing the gene of the present invention (a gene encoding the present enzyme). The recombinant DNA of the present invention may be provided, for example, in the form of a vector. As used herein, the term "vector" refers to a nucleic acid molecule that can transport a nucleic acid inserted therein into a target such as a cell.
[0083] An appropriate vector is selected depending on the intended use (cloning, protein expression) and the type of host cell. Examples of vectors using E. coli as a host include M13 phage or modified versions thereof, λ phage or modified versions thereof, and pBR322 or modified versions thereof (pB325, pAT153, pUC8, etc.), vectors using yeast as a host include pYepSec1, pMFa, pYES2, etc., vectors using insect cells as a host include pAc and pVL, vectors using mammalian cells as a host include pCDM8 and pMT2PC, and vectors using filamentous fungi as a host include pUC19.
[0084] The vector of the present invention is preferably an expression vector. An "expression vector" refers to a vector that can introduce a nucleic acid inserted therein into a target cell (host cell) and express it in the cell. An expression vector usually contains a promoter sequence necessary for the expression of the inserted nucleic acid, an enhancer sequence that promotes expression, and the like. Expression vectors containing a selection marker can also be used. When such an expression vector is used, the presence or absence (and the degree of introduction) of the expression vector can be confirmed using the selection marker.
[0085] Insertion of the nucleic acid of the present invention into a vector, insertion of a selectable marker gene (if necessary), insertion of a promoter (if necessary), etc. can be carried out using standard recombinant DNA techniques (for example, see Molecular Cloning, Third Edition, 1.84, Cold Spring Harbor Laboratory Press, New York; well-known methods using restriction enzymes and DNA ligases).
[0086] For ease of handling, microorganisms such as Escherichia coli, Saccharomyces cerevisiae, and filamentous fungi (Aspergillus oryzae) are preferred as host cells. However, any host cell capable of replicating recombinant DNA and expressing the gene for this enzyme can be used. Examples of E. coli include E. coli BL21(DE3)pLysS when using a T7 promoter, and E. coli JM109 when not. Examples of filamentous fungi include Aspergillus oryzae RIB40. Examples of budding yeast include SHY2, AH22, and INVSc1 (Invitrogen).
[0087] The present invention further provides a microorganism (i.e., a transformant) carrying the recombinant DNA of the present invention. The microorganism of the present invention can be obtained by transfection or transformation using the above-mentioned vector of the present invention. For example, calcium chloride method (J. Mol. Biol., Vol. 53, p. 159 (1970)), Hanahan method (J. Molecular Biology, Vol. 166, p. 557 (1983)), SEM method (Gene, Vol. 96, p. 23 (1990)), Chung et al.'s method (Proceedings of the National Academy of Sciences of the USA, Vol. 86, p. 2172 (1989)), calcium phosphate coprecipitation method, electroporation (Potter, H. et al., Proc. Natl. Acad. Sci. USA 81, 7161-7165 (1984)), lipofection (Felgner, P. L. et al., Proc. Natl. Acad. Sci. USA 84, 7413-7417 (1984)) etc. The microorganism of the present invention can be used to produce the present enzyme.
[0088] 4. Preparation of glucose dehydrogenase (GDH) 1 Another aspect of the present invention relates to a method for preparing the present enzyme. In the preparation method of the present invention, the present enzyme is prepared by genetic engineering techniques. Specifically, a gene encoding the present enzyme is first prepared (step (1)). Specifically, for example, a nucleic acid encoding any of the amino acid sequences set forth in SEQ ID NOS: 2 to 4 is prepared. Here, a "nucleic acid encoding any of the amino acid sequences set forth in SEQ ID NOS: 2 to 4" is a nucleic acid that, when expressed, results in a polypeptide having the amino acid sequence. This includes not only nucleic acids consisting of a nucleotide sequence corresponding to the amino acid sequence, but also nucleic acids to which an additional sequence (which may or may not encode an amino acid sequence) has been added. Codon degeneracy is also taken into consideration. A "nucleic acid encoding any of the amino acid sequences set forth in SEQ ID NOS: 2 to 4" can be prepared in an isolated state by standard genetic engineering, molecular biological, or biochemical techniques, with reference to the sequence information disclosed in this specification or the attached sequence listing.
[0089] Following step (1), the prepared gene is expressed (step (2)). For example, an expression vector into which the gene has been inserted is first prepared, and a host cell is transformed with this. Next, the transformant is cultured under conditions in which the mutant enzyme, which is the expression product, is produced. The transformant can be cultured according to conventional methods. The carbon source used in the medium may be any assimilable carbon compound, such as glucose, sucrose, lactose, maltose, molasses, or pyruvic acid. The nitrogen source may be any usable nitrogen compound, such as peptone, meat extract, yeast extract, casein hydrolysate, or alkaline extract of soybean meal. Other salts, such as phosphate, carbonate, sulfate, magnesium, calcium, potassium, iron, manganese, or zinc, as well as specific amino acids and specific vitamins, may also be used as needed.
[0090] The culture temperature can be set taking into consideration the growth characteristics of the transformant to be cultured and the production characteristics of the mutant enzyme. It is preferably set within the range of 30°C to 40°C (more preferably around 37°C). The culture time can be set taking into consideration the growth characteristics of the transformant to be cultured and the production characteristics of the mutant enzyme. The pH of the medium is adjusted within a range that allows the transformant to grow and the enzyme to be produced. The pH of the medium is preferably about 6.0 to 9.0 (preferably around pH 7.0).
[0091] Next, the expression product (GDH) is recovered (step (3)). The culture medium containing the bacterial cells after cultivation can be used as an enzyme solution as is or after concentration, removal of impurities, etc., but the expression product is generally recovered first from the culture medium or bacterial cells. If the expression product is a secretory protein, it can be recovered from the culture medium, and if not, it can be recovered from the bacterial cells. When recovering GDH from the culture medium, for example, the culture supernatant is filtered and centrifuged to remove insoluble matter, followed by vacuum concentration, membrane concentration, salting out using ammonium sulfate or sodium sulfate, fractional precipitation using methanol, ethanol, or acetone, dialysis, heat treatment, isoelectric focusing, and various types of chromatography such as gel filtration, adsorption chromatography, ion exchange chromatography, and affinity chromatography (e.g., gel filtration using Sephadex gel (GE Healthcare Biosciences) or the like, DEAE Sepharose CL-6B (GE Healthcare Biosciences), Octyl Sepharose CL-6B (GE Healthcare Biosciences), and CM Sepharose CL-6B (GE Healthcare Biosciences)). A purified GDH product can be obtained by combining these methods. On the other hand, when recovering GDH from bacterial cells, the culture medium is filtered, centrifuged, or the like to harvest the bacterial cells, which are then disrupted mechanically by pressure treatment, ultrasonication, or enzymatically using lysozyme, followed by separation and purification in the same manner as described above.
[0092] The degree of purification of the enzyme is not particularly limited, but it can be purified to a specific activity of, for example, 0.1 to 1000 (U / mg), preferably 1 to 500 (U / mg). The final form may be liquid or solid (including powder).
[0093] The purified enzyme obtained as described above can be provided as a powder by, for example, freeze-drying, vacuum drying, or spray-drying. In this case, the purified enzyme may be dissolved in advance in phosphate buffer, triethanolamine buffer, Tris-HCl buffer, or Good's buffer. Preferably, phosphate buffer or triethanolamine buffer is used. Examples of Good's buffer include PIPES, MES, and MOPS.
[0094] Typically, gene expression and recovery of the expression product (GDH) are performed using an appropriate host-vector system as described above. However, cell-free synthesis systems can also be used. Here, "cell-free synthesis systems (cell-free transcription systems, cell-free transcription / translation systems)" refer to in vitro synthesis of mRNA and proteins encoded by template nucleic acids (DNA and mRNA) using ribosomes and transcription / translation factors derived from living cells (or obtained by genetic engineering techniques), rather than using living cells. Cell-free synthesis systems generally use cell extracts obtained by purifying cell lysates as needed. Cell extracts generally contain ribosomes, various factors such as initiation factors, and various enzymes such as tRNA, all of which are necessary for protein synthesis. When synthesizing proteins, various amino acids, energy sources such as ATP and GTP, and other substances necessary for protein synthesis, such as creatine phosphate, are added to the cell extract. Of course, separately prepared ribosomes, various factors, and / or various enzymes may be supplemented as needed during protein synthesis.
[0095] The development of a transcription / translation system in which each molecule (factor) required for protein synthesis is reconstituted has also been reported (Shimizu, Y. et al.: Nature Biotech., 19, 751-755, 2001). In this synthesis system, the genes for 31 factors that make up the bacterial protein synthesis system, including three initiation factors, three elongation factors, four factors involved in termination, 20 aminoacyl-tRNA synthetases that bind each amino acid to tRNA, and methionyl-tRNA formyltransferase, were amplified from the Escherichia coli genome, and the protein synthesis system was reconstituted in vitro using these genes. Such a reconstituted synthesis system may be used in the present invention.
[0096] The term "cell-free transcription / translation system" is used interchangeably with cell-free protein synthesis system, in vitro translation system, or in vitro transcription / translation system. In an in vitro translation system, RNA is used as a template to synthesize proteins. Examples of template RNA that can be used include total RNA, mRNA, and in vitro transcription products. In contrast, in vitro transcription / translation systems use DNA as a template. The template DNA should contain a ribosome binding domain and preferably contains an appropriate terminator sequence. In an in vitro transcription / translation system, conditions are established in which factors required for each reaction are added so that the transcription and translation reactions proceed sequentially.
[0097] 5. Uses of Glucose Dehydrogenase A further aspect of the present invention relates to uses of the present enzyme. In this aspect, a glucose measurement method using the present enzyme is first provided. In the glucose measurement method of the present invention, the amount of glucose in a sample is measured by utilizing an oxidation-reduction reaction mediated by the present enzyme. The present invention can be applied to various applications in which changes resulting from this reaction can be utilized.
[0098] The present invention can be used, for example, to measure blood glucose levels (self-monitoring of blood glucose (SMBG) and continuous glucose monitoring (CGM)), to measure glucose contained in body fluids other than blood (e.g., tears, saliva, interstitial fluid, urine, etc.), to measure the glucose concentration in foods (condiments, beverages, etc.), etc. The present invention may also be used to examine the degree of fermentation in the production process of fermented foods (e.g., vinegar) or fermented beverages (e.g., beer and sake).
[0099] The present invention also provides a glucose measurement reagent containing the present enzyme. The reagent is used in the glucose measurement method of the present invention. For the purpose of stabilizing the glucose measurement reagent and activating it during use, serum albumin, proteins, surfactants, sugars, sugar alcohols, inorganic salts, etc. may be added.
[0100] The glucose measurement reagent can also be used as a component of a measurement kit. In other words, the present invention also provides a kit (glucose measurement kit) containing the above-mentioned glucose measurement reagent. The kit of the present invention contains the above-mentioned glucose measurement reagent as an essential component. It also contains reaction reagents, buffer solutions, glucose standard solutions, containers, etc. as optional components. The glucose measurement kit of the present invention is usually accompanied by instructions for use.
[0101] A glucose sensor can be constructed using this enzyme. That is, the present invention also provides a glucose sensor containing this enzyme. In a typical structure of the glucose sensor of the present invention, an electrode system including a working electrode and a counter electrode is formed on an insulating substrate, and a reagent layer containing the present enzyme and a mediator is formed thereon. More specifically, the reagent layer is usually coated on the working electrode. The present invention can also be applied to face-to-face glucose sensors in which the working electrode and the counter electrode face each other. Alternatively, a measurement system including a reference electrode may be used. Using such a so-called three-electrode measurement system makes it possible to express the potential of the working electrode relative to the potential of the reference electrode. The materials of each electrode are not particularly limited. Examples of electrode materials for the working electrode and counter electrode include gold (Au), carbon (C), platinum (Pt), and titanium (Ti). Examples of mediators that can be used include ferricyanide compounds (e.g., potassium ferricyanide), metal complexes (e.g., ruthenium complexes, osmium complexes, vanadium complexes), and quinone compounds (e.g., pyrroloquinoline quinone). For details on the structure of glucose sensors and electrochemical measurement methods using glucose sensors, see, for example, "Bioelectrochemistry in Practice - Practical Development of Biosensors and Biobatteries" (published in March 2007, CMC Publishing).
[0102] The present enzyme can also be provided in the form of an enzyme preparation. In addition to the active ingredient (the present enzyme), the enzyme preparation of the present invention may contain excipients, buffers, suspending agents, stabilizers, preservatives, antiseptics, physiological saline, etc. Examples of excipients that can be used include starch, dextrin, maltose, trehalose, lactose, sorbitol, D-mannitol, sucrose, glycerol, sugar esters, and derivatives thereof, as well as D-glucose derivatives. Examples of buffers that can be used include phosphates, citrates, acetates, etc. Examples of stabilizers that can be used include propylene glycol, ascorbic acid, etc. Examples of preservatives that can be used include phenol, benzalkonium chloride, benzyl alcohol, chlorobutanol, methylparaben, etc. Examples of preservatives that can be used include ethanol, benzalkonium chloride, parahydroxybenzoic acid, chlorobutanol, etc. Examples of uses of enzyme preparations include measuring glucose, blood sugar levels, and fermentation levels.
[0103] 6. Screening method for glucose dehydrogenase (GDH) A further aspect of the present invention relates to a method for screening for GDH. The screening method of the present invention is useful as a means for identifying and obtaining novel GDH. The screening method of the present invention comprises the following steps (i) and (ii): (i) searching for amino acid sequences that show 20% or more identity to the amino acid sequence of a known GDH using an amino acid sequence database; (ii) Among the hit amino acid sequences, the following sequence (SEQ ID NO: 1): [ka] (wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, H / N represents H (histidine) or N (asparagine), L / I / M represents L (leucine), I (isoleucine), or M (methionine), V / I represents V (valine) or I (isoleucine), G / A / S represents G (glycine), A (alanine), or S (serine), F / W / Y represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue)
[0104] Step (i) is a step of searching for sequences that satisfy specific conditions, i.e., show a predetermined homology to the amino acid sequence of a known specific GDH that serves as a reference (hereinafter referred to as "reference GDH"). Typically, one GDH is used as the reference GDH, but multiple GDHs may be used in combination, and sequences that show a predetermined homology to all of them may be searched for. In the present invention, amino acid sequences that show 20% or more identity to the amino acid sequence of the reference GDH are searched for. Searches may also be performed with even higher identities (e.g., 25% or more, 30% or more, 31% or more, 32% or more, 33% or more, 34% or more, 35% or more, 36% or more, 37% or more, 38% or more, 39% or more, 40% or more). When aiming to obtain related enzymes (enzymes that are expected to exhibit the same or similar activity, etc.), it is common to search for those with high sequence identity. In other words, conventional search methods are suitable for obtaining related enzymes with high sequence identity. In contrast, the present invention makes it possible to efficiently find related enzymes from among enzymes with low sequence identity, and is distinct from conventional search methods.
[0105] The known GDH, i.e., the reference GDH, is not particularly limited, and examples thereof include FAD-GDH derived from Aspergillus oryzae (e.g., having the amino acid sequence of SEQ ID NO: 22) and FAD-GDH derived from Aspergillus terreus (e.g., having the amino acid sequence of SEQ ID NO: 23).
[0106] The amino acid sequence database to be used is not particularly limited. For example, databases such as Entrez Protein (http: / / www.ncbi.nlm.nih.gov / entrez / query.fcgi?db=Protein), Swiss-Prot (http: / / www.expasy.org / sprot / ), and PIR (http: / / pir.georgetown.edu / ) can be used.
[0107] Preferably, sequences that satisfy the above conditions are searched for from within the GMC oxidoreductase family. In other words, the search targets are limited to amino acid sequences of proteins (enzymes) that belong to the GMC oxidoreductase family, thereby improving search efficiency.
[0108] In step (ii) following step (i), hit amino acid sequences are identified and selected from the hit amino acid sequences that have the amino acid region defined by rule 1. If there is only one hit sequence, it is confirmed whether or not it has the amino acid region defined by rule 1, and if so, it is judged to be promising. Rule 1 has been explained in the first aspect of the present invention (section 2. Glucose dehydrogenase (GDH)), so the above explanation will be used and details will be omitted.
[0109] In a preferred embodiment, the amino acid region defined by Rule 1 also satisfies Rule 2 (conditions 1 and 2). In other words, an amino acid sequence is selected that has an amino acid region that satisfies Rules 1 and 2. Rule 2 has been explained in the first aspect of the present invention (section 2. Glucose dehydrogenase (GDH)), and therefore the above explanation is incorporated herein and its details are omitted.
[0110] In a more preferred embodiment, the amino acid region defined by rule 1 also satisfies rule 3. In other words, an amino acid sequence is selected that has an amino acid region that satisfies rules 1 to 3. Rule 3 is as explained in the first aspect of the present invention (section 2. Glucose dehydrogenase (GDH)), and therefore the above explanation is incorporated herein and its details are omitted.
[0111] Preferably, from the amino acid sequences selected in step (ii), those having the conserved regions 1 to 10 described in the first aspect above are selected (step (iii)). If only one amino acid sequence is selected in step (ii), it is confirmed whether or not it has the conserved regions 1 to 10, and if so, it is judged to be promising.
[0112] Promising amino acid sequences are selected by steps (i) and (ii), or steps (i) to (iii), and it is preferable to confirm whether the selected amino acid sequence actually functions as GDH. That is, preferably, a step (step (iv)) of confirming the GDH activity of a protein consisting of the selected amino acid sequence is carried out. If GDH activity is confirmed, the protein is identified as a promising novel GDH. Furthermore, if the GDH activity is high, the protein is identified as being particularly promising. Incidentally, GDH activity can be confirmed, for example, by using the activity measurement method shown in the Examples below.
[0113] 7. Method for designing modified glucose dehydrogenase (GDH) A further aspect of the present invention relates to a method for designing a modified GDH. The design method of the present invention is useful, for example, as a means for enhancing the activity of known GDHs. It can also be used as a means for converting proteins that show a certain degree of homology to known GDHs and are promising as GDHs into GDHs (i.e., imparting GDH activity). The term "modified GDH" refers to a GDH obtained by partially altering (modifying) an existing amino acid sequence. Therefore, there is a difference between the amino acid sequence before and after modification.
[0114] The design method of the present invention involves the following steps (I) and (II). (I) an amino acid sequence of a known GDH or an amino acid sequence showing 20% or more identity to the amino acid sequence of a known GDH, which is the following sequence (SEQ ID NO: 1): [ka] (wherein the numbers above each amino acid residue in the formula are consecutive numbers from the N-terminus, H / N represents H (histidine) or N (asparagine), L / I / M represents L (leucine), I (isoleucine), or M (methionine), V / I represents V (valine) or I (isoleucine), G / A / S represents G (glycine), A (alanine), or S (serine), F / W / Y represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue), and a step of preparing an amino acid sequence to be modified that does not have an amino acid region consisting of these. (II) identifying a region in the prepared amino acid sequence that corresponds to the amino acid region, and then modifying the region so that it corresponds to the amino acid region, thereby constructing a modified amino acid sequence;
[0115] Step (I) is the step of preparing an amino acid sequence to be modified (referred to as the "target amino acid sequence"). In the present invention, amino acid sequences that do not contain the amino acid region defined by Rule 1 are selected from among known GDH amino acid sequences or amino acid sequences showing 20% or more identity to known GDH amino acid sequences, and these are used as target amino acid sequences. Preferably, the target amino acid sequence does not contain the amino acid region defined by Rule 1, but contains an amino acid sequence similar to the amino acid region (e.g., one that differs from the amino acid sequence defined by Rule 1 at one to four positions, preferably one to three positions, and more preferably one or two positions). The known GDH is not particularly limited. Examples of known GDHs include Aspergillus oryzae-derived FAD-GDH (e.g., one having the amino acid sequence of SEQ ID NO: 21) and Aspergillus terreus-derived FAD-GDH (e.g., one having the amino acid sequence of SEQ ID NO: 22). Amino acid sequences showing 20% or more identity to the amino acid sequences of known GDHs can be searched for and identified by the method described in step (i) of the above aspect (5. Screening method for glucose dehydrogenase (GDH)). Specific examples of amino acid sequences showing 20% or more identity to the amino acid sequences of known GDHs include the amino acid sequence shown in SEQ ID NO: 8 (derived from Aspergillus wentii), the amino acid sequence shown in SEQ ID NO: 9 (derived from Aspergillus wentii), the amino acid sequence shown in SEQ ID NO: 10 (derived from Aspergillus thermomutatus), and the amino acid sequence shown in SEQ ID NO: 11 (derived from Aspergillus cristatus).
[0116] In step (II) following step (I), a region (referred to as the "region to be modified") corresponding to the amino acid region defined by rule 1 is first identified in the prepared amino acid sequence. For example, the region to be modified can be identified by using criteria or conditions that satisfy part of rule 1 (conditions for the first half (amino acid residues 1 to 6), conditions for the second half (amino acid residues 7 to 13), conditions for amino acid residues 1 to 5, conditions for amino acid residues 5 to 9, conditions for amino acid residues 7 to 11, etc.). As described above, the amino acid region defined by rule 1 corresponds to a sequence of 13 consecutive amino acids located near positions 140 to 180 of FAD-GDH derived from filamentous fungi. When identifying the region to be modified, it is advisable to also use the positional information.
[0117] After identifying the region to be modified, the region to be modified is modified so that it corresponds to the amino acid region defined by Rule 1. That is, the amino acid composition of the region to be modified is changed so that it satisfies the conditions of Rule 1. This modification constructs a modified amino acid sequence having the amino acid region defined by Rule 1 (the region to be modified).
[0118] In a preferred embodiment, the amino acid region defined by Rule 1 also satisfies Rule 2 (conditions 1 and 2). In this embodiment, the region to be modified is modified so that it becomes an amino acid region that satisfies Rules 1 and 2. As a result, a modified amino acid sequence having an amino acid region that satisfies Rules 1 and 2 is constructed. Rule 2 is as explained in the first aspect of the present invention (section 2. Glucose dehydrogenase (GDH)), and therefore the above explanation is incorporated herein and its details are omitted.
[0119] In a more preferred embodiment, the amino acid region defined by rule 1 also satisfies rule 3. In this embodiment, the region to be modified is modified so that it becomes an amino acid region that satisfies rules 1 to 3. As a result, a modified amino acid sequence having an amino acid region that satisfies rules 1 to 3 is constructed. Rule 3 is as explained in the first aspect of the present invention (section 2. Glucose dehydrogenase (GDH)), and therefore the above explanation is incorporated herein and its details are omitted.
[0120] Preferably, the constructed modified amino acid sequence is confirmed to have the conserved regions 1 to 10 described in the first aspect above, and if not, is further modified to have them (step (III)). This step may be carried out in parallel with step (II).
[0121] It is preferable to confirm whether the constructed modified amino acid sequence is actually a GDH sequence. That is, preferably, a step (step (IV)) of confirming the GDH activity of the protein consisting of the constructed modified amino acid sequence is carried out. If an improvement in GDH activity is observed, such as when the GDH activity is improved compared to before modification or when the modification results in the exhibiting of GDH activity, the protein consisting of the constructed modified amino acid sequence is identified as an effective modified GDH.
[0122] 8. Preparation of glucose dehydrogenase (GDH) 2 A further aspect of the present invention provides a method for preparing glucose dehydrogenase as an application of the above screening method and design method, which comprises the following steps (1) to (3): (1) preparing a glucose dehydrogenase by the screening method of the present invention or the design method of the present invention; (2) Culturing the microorganism that expresses the glucose dehydrogenase (3) recovering the expression product from the culture
[0123] First, a glucose dehydrogenase is prepared by the screening method or design method of the present invention (step (1)). Next, a microorganism expressing the prepared glucose dehydrogenase is cultured (step (2)). The microorganism here is typically a wild-type strain or a mutant strain thereof that produces the glucose dehydrogenase, or a transformant into which the gene for the glucose dehydrogenase has been introduced, if the glucose dehydrogenase has been prepared by the screening method; or a transformant into which the gene for the glucose dehydrogenase has been introduced, if the glucose dehydrogenase has been prepared by the design method. After step (2), the expression product is recovered from the culture (step (3)). This operation may be performed in accordance with step (3) in the above-mentioned glucose dehydrogenase (GDH) preparation method 1. [Example]
[0124] <Identification of a novel FAD-GDH and the unique region that defines GDH> In order to identify a novel FAD-GDH and to find the specific amino acid region that defines GDH, the following investigations were carried out. 1. Obtaining FAD-GDH using an artificially synthesized gene A comprehensive search using NCBI BLAST and the unique criteria (rules 1–3) shown below identified two genes in the genomes of A. cristatus and A. turcosus that are expected to exhibit FAD-GDH activity (Asp. cristatus GZAAS20.1005, NCBI Sequence ID: ODM22452.1 and Asp. turcosus HMR AF 23, NCBI Sequence ID: RHZ62616.1). Synthetic genes were prepared and DNA was amplified by PCR using these as templates. The amplified DNA was ligated into the yeast expression vector pYES2 and introduced into S. cerevisiae strains by the protoplast method. <Indicators> Rule 1: It has an amino acid region (hereinafter referred to as "GDH-specific region") consisting of the following sequence (SEQ ID NO: 1). [ka] In the formula, the numbers above each amino acid residue are consecutive numbers from the N-terminus, (H / N) represents H (histidine) or N (asparagine), (L / I / M) represents L (leucine), I (isoleucine), or M (methionine), (V / I) represents V (valine) or I (isoleucine), (G / A / S) represents G (glycine), A (alanine), or S (serine), (F / W / Y) represents F (phenylalanine), W (tryptophan), or Y (tyrosine), and X represents any amino acid residue. Condition 1 of Rule 2: For the 8th, 12th, and 13th amino acid residues, the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database, i.e., (value of the 8th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 3.79 or less. Condition 2 of Rule 2: For amino acid residues 8, 11, 12, and 13, the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database, i.e., (value of amino acid residue 8 + value of amino acid residue 11 + value of amino acid residue 12 + value of amino acid residue 13) must be 2.83 or greater. Rule 3: For the third and fourth amino acid residues, the sum of the values of JANJ780101 (Average accessible surface area) listed in the AAindex database, i.e., (the value of the third amino acid residue + the value of the fourth amino acid residue) must be 66 or greater.
[0125] 2. Creation of FAD-GDH-like genes using synthetic genes A comprehensive search using NCBI BLAST identified four genes from A.wentii, A.thermomutatus, and A.cristatus that did not match the unique criteria mentioned above and were therefore not expected to exhibit FAD-GDH activity (Asp.wentii DTO 134E9, NCBI Sequence ID: OJJ29827.1, Asp.wentii DTO 134E9, NCBI Sequence ID: OJJ29704.1, Asp.thermomutatus HMR AF 39, NCBI Sequence ID: XP#026617872.1, and Asp.cristatus GZAAS20.1005, NCBI Sequence ID: ODM23900.1). Synthetic genes were prepared and used as templates for PCR amplification of DNA. The amplified DNA was ligated into the yeast expression vector pYES2 and introduced into S. cerevisiae strains by the protoplast method.
[0126] 3. Obtaining FAD-GDH candidate genes based on genetic information Based on the above criteria, we conducted a comprehensive search of various genome databases. As a result, we identified a gene (CSFG Corse1p7#015851, annotated as glucose oxidase) (SEQ ID NO: 23) expected to exhibit FAD-GDH activity in the genome of Corynascus sepedonium ATCC 9787 in the CSFG database. Primers were designed based on this gene sequence, and DNA was amplified by PCR using the Corynascus sepedonium NBRC 31363 genome as a template. The amplified DNA was ligated into a filamentous fungal expression vector (pUC19 containing an α-amylase modified promoter, an A. oryzae FAD-GDH terminator, and the pyrG gene) and then introduced into Aspergillus oryzae RIB40 strain by the protoplast-PEG method.
[0127] 4. Enzyme activity measurement The measurement principle is as follows: when D-glucose and phenazine methosulfate (PMS) are reacted with GDH, glucose is oxidized, while PMS is reduced. Reduced PMS reduces nitrotetrazorium blue (NTB). When NTB is reduced by reduced PMS, a diformazan dye is formed. Enzyme activity is measured by detecting this diformazan dye at 570 nm. The experimental procedure is as follows: First, 2.6 mL of 50 mmol / L PIPES-NaOH buffer, pH 6.5 (containing 0.5% Triton X-100), 1 mL of 1.0 mmol / L D-glucose, 1 mL of 6.6 mmol / L NTB solution, and 2 mL of 3.0 mmol / L PMS solution were placed in a spectrophotometer cell and preheated at 37°C for 10 minutes. 0.1 mL of FAD-GDH solution was added, mixed, and incubated at 37°C. After the start of the reaction, the absorbances of the reaction solution at 570 nm were measured at 3 and 5 minutes, respectively (A3, A5). Separately, as a blank, the same procedure was repeated using 50 mmol / L PIPES-NaOH buffer, pH 6.5 (containing 0.1% Triron X-100, 0.1% BSA, and 1 mmol / L CaCl2) instead of the FAD-GDH solution, and the absorbances Ab3, Ab5 were measured at 3 and 5 minutes, respectively. The values of A3, A5, Ab3, and Ab5 were substituted into the following equation to calculate the FAD-GDH activity.
number
[0128] 5. Confirmation of GDH activity We measured the GDH activity of samples prepared from the culture medium of six transformed yeast species (two species expected to exhibit FAD-GDH activity and four species not). We also confirmed that transformed yeast without the GDH gene did not exhibit GDH activity. Of the six transformed yeast species, two species reproducibly exhibited GDH activity (Figure 1), indicating that they were FAD-GDH. The other four species did not reproducibly exhibit GDH activity (Figure 1), indicating that they were FAD-GDH-like genes.
[0129] The stability of the two enzymes that demonstrated activity was confirmed, and it was found that both enzymes were significantly more stable than A. oryzae-derived FAD-GDH (Table 4). These two enzymes, with their excellent stability, are highly practical and useful for glucose measurement (particularly blood glucose measurement). Stability was calculated by comparing the residual activity after treatment in 100 mmol / L phosphate buffer at 50°C for 20 minutes with the activity before treatment. [Table 4]
[0130] Next, we examined the FAD-GDH activity of a sample prepared from the culture medium of the A. oryzae RIB40 strain, which had been transformed with a gene derived from the Corynascus sepedonium NBRC31363 strain. We also confirmed that the culture medium of the transformed RIB40 strain, which had not been transformed with the GDH gene, did not exhibit GDH activity. This sample reproducibly exhibited FAD-GDH activity. Furthermore, stability testing revealed that the stability was significantly superior to that of the A. oryzae-derived FAD-GDH (Table 5). This highly stable enzyme is highly practical and useful, for example, for glucose measurement (particularly for blood glucose measurement). Stability was calculated by comparing the residual activity after 20 minutes of treatment in 100 mmol / L phosphate buffer at 50°C with the activity before treatment. [Table 5]
[0131] The sequences of the GDH-specific regions of the three enzymes that were found to exhibit FAD-GDH activity are shown below, along with the full-length enzymes and the sequences of the genes that encode them. (1)A.cristatus GZAAS20.1005(NCBI Sequence ID: ODM22452.1) GDH-specific region: HGYDGPLHVGFNN (SEQ ID NO: 12) Full length enzyme: SEQ ID NO: 2 Gene: SEQ ID NO: 5 (2)A.turcosus HMR AF 23(NCBI Sequence ID: RHZ62616.1) GDH-specific region: HGKAGPLLVGWTY (SEQ ID NO: 13) Full length enzyme: SEQ ID NO: 3 Gene: SEQ ID NO: 6 (3) Corynascus sepedonium NBRC 31363 (originally cloned based on the sequence of Corynascus sepedonium ATCC 9787 CSFG Corse1p7#015851) GDH-specific region: HGFRGPLHVGYTP (SEQ ID NO: 14) Full length enzyme: SEQ ID NO: 4 Gene: SEQ ID NO: 7
[0132] On the other hand, for the four enzymes that did not satisfy the conditions for a GDH-specific region, the sequences of the regions corresponding to the GDH-specific region and the sequences of the full-length enzymes are shown below. (5)A.wentii DTO 134E9(NCBI Sequence ID: OJJ29827.1) Corresponding region: HNSNGPVQVSFKH (SEQ ID NO: 15) Full length enzyme: SEQ ID NO: 8 (6)A.wentii DTO 134E9(NCBI Sequence ID: OJJ29704.1) Corresponding region: HGTKGPLKVGWPS (SEQ ID NO: 16) Full length enzyme: SEQ ID NO: 9 (7)A.thermomutatus HMR AF 39(NCBI Sequence ID: XP#026617872.1) Corresponding region: HGSRGPLKVAFPH (SEQ ID NO: 17) Full length enzyme: SEQ ID NO: 10 (8)A.cristatus GZAAS20.1005(NCBI Sequence ID: ODM23900.1) Corresponding region: HGYEGPLKVGWPP (SEQ ID NO: 18) Full length enzyme: SEQ ID NO: 11
[0133] Furthermore, the sequences of the GDH-specific regions and full-length enzyme sequences of A. oryzae-derived FAD-GDH and A. terreus-derived FAD-GDH, which are known FAD-GDHs, are shown below. (9) A. oryzae GDH-specific region: NGKEGPLKVGWSG (SEQ ID NO: 19) Full length enzyme: SEQ ID NO: 21 (10) A. terreus GDH-specific region: HGHEGPLDVAFTQ (SEQ ID NO: 20) Full length enzyme: SEQ ID NO: 22
[0134] The numerical values for rules 2 and 3 of the GDH-specific region (or the corresponding region) of each enzyme are as follows (Table 6). [Table 6]
[0135] The identity of the amino acid sequence of the novel FAD-GDH to that of known FAD-GDH sequences (FAD-GDH derived from A. oryzae and FAD-GDH derived from A. terreus) is shown in FIG. [Industrial Applicability]
[0136] The glucose dehydrogenase of the present invention is characterized by a newly discovered unique region. The glucose dehydrogenase of the present invention is intended to be used for various purposes (e.g., blood glucose monitors). Meanwhile, the information provided by this application is useful for the discovery and identification of novel GDHs.
[0137] The present invention is not limited to the above-described embodiments and examples. Various modifications within the scope of the claims and within the scope that can be easily conceived by a person skilled in the art are also included in the present invention. The contents of papers, published patent applications, patent publications, and other publications explicitly stated in this specification are incorporated herein by reference in their entirety.
Claims
1. The following sequence (SEQ ID NO: 1): 【Chemistry 1】 (wherein -omitted-) An enzyme preparation used as an FAD-dependent glucose dehydrogenase, comprising a protein having an amino acid sequence that is 95% or more identical to the amino acid sequence of SEQ ID NO: 2 (excluding those containing the amino acid sequence of SEQ ID NO: 2), or an amino acid sequence that is 80% or more identical to the amino acid sequence of any of SEQ ID NOs: 3 to 4.
2. A method for screening for FAD-dependent glucose dehydrogenase, comprising: The method comprises the following steps (i) and (ii): (i) searching an amino acid sequence database for an amino acid sequence that shows 40% or more identity to the amino acid sequence of a known FAD-dependent glucose dehydrogenase; and (ii) Among the hit amino acid sequences, the following sequence (SEQ ID NO: 1): 【Chemistry 1】 a step of selecting a polypeptide having an amino acid region consisting of (-omitted-); In the step (ii), the amino acid region further comprises The condition is that the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues, i.e., (the value of the 8th amino acid residue + the value of the 12th amino acid residue + the value of the 13th amino acid residue) is 3.79 or less; The condition that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for the 8th, 11th, 12th, and 13th amino acid residues, i.e., (value of the 8th amino acid residue + value of the 11th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 2.83 or more; The screening method as described above, wherein an individual satisfying the above is selected.
3. A method for designing a modified FAD-dependent glucose dehydrogenase, comprising: The method includes the following steps (I) and (II): (I) The amino acid sequence of a known FAD-dependent glucose dehydrogenase, which is the following sequence (SEQ ID NO: 1): 【Chemistry 1】 providing an amino acid sequence to be modified that does not have an amino acid region consisting of (-omitted-); and (II) identifying a region in a prepared amino acid sequence that corresponds to the amino acid region, and then modifying the region so that it corresponds to the amino acid region, thereby constructing a modified amino acid sequence; In the step (II), the amino acid region further comprises The condition is that the sum of the values of CRAJ730103 (Normalized frequency of turns) listed in the AAindex database for the 8th, 12th, and 13th amino acid residues, i.e., (the value of the 8th amino acid residue + the value of the 12th amino acid residue + the value of the 13th amino acid residue) is 3.79 or less; The condition that the sum of the values of TANS770109 (Normalized frequency of coil) listed in the AAindex database for the 8th, 11th, 12th, and 13th amino acid residues, i.e., (value of the 8th amino acid residue + value of the 11th amino acid residue + value of the 12th amino acid residue + value of the 13th amino acid residue) is 2.83 or more; The design method further comprises modifying the design so as to satisfy the above.
4. A method for preparing FAD-dependent glucose dehydrogenase, comprising the following steps (1) to (3): (1) preparing an FAD-dependent glucose dehydrogenase by the screening method according to claim 2 or the design method according to claim 3; (2) culturing the microorganism expressing the FAD-dependent glucose dehydrogenase; and (3) recovering the expression product from the culture.
Citation Information
Patent Citations
Glucose dehydrogenase
JP2000350588A
Glucose dehydrogenase having excellent substrate specificity
JP2001197888A
Glucose dehydrogenase excellent in substrate specificity
JP2001346587A
Coenzyme-binding glucose dehydrogenase
WO2004058958A1
Coenzyme-linked glucose dehydrogenase and polynucleotide encoding the same
WO2006101239A1