Method for selecting template enzyme
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KOBE UNIV
- Filing Date
- 2023-03-29
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional methods for selecting enzymes for biosynthetic pathways rely heavily on researcher experience and known information, leading to inefficiencies and inaccuracies in designing effective pathways for producing useful compounds.
A method for selecting enzymes with high catalytic efficiency and substrate specificity using amino acid and nucleic acid sequence information, combined with enzyme activity basic units, to create an enzyme library that can efficiently produce desired metabolites.
Enables the design of efficient biosynthetic pathways and rapid production of desired metabolites by identifying enzymes with high catalytic efficiency and wide substrate specificity, reducing reliance on researcher knowledge.
Smart Images

Figure 00000022_0000 
Figure 00000022_0001 
Figure 00000022_0002
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, an enzyme library containing such an enzyme, and a method for selecting such an enzyme. More specifically, this disclosure relates to a template enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, and a technology for producing a desired metabolite using such a template enzyme. [Background technology]
[0002] Traditionally, the design of biosynthetic pathways for the production of useful compounds, particularly the selection of enzymes for each reaction step, has been based on known information such as published literature, and therefore heavily relies on the researcher's experience. There is a need for the development of enzyme selection methods that do not depend on the knowledge of a single researcher, and for efficient methods for designing target biosynthetic pathways based on these methods. [Overview of the project] [Means for solving the problem]
[0003] This disclosure provides an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, as well as an enzyme library of such enzymes, thereby providing a method for efficiently designing a target biosynthetic pathway and easily producing a target metabolite.
[0004] Therefore, this disclosure provides the following: (Item 1) An enzyme having a catalytic efficiency above a certain level and substrate specificity above a certain level. (Item 2) The enzyme is the enzyme described in item 2a, which has a catalytic efficiency equal to or greater than the standard catalytic efficiency of the enzyme. The enzyme described in any one of the above items is evaluated by a substrate coverage index, the substrate coverage index being the sum of the catalytic efficiencies of each substrate selected from a basic substrate set or their common logarithms. (Item 2b) The enzyme is the enzyme described in any one of the above items, wherein, among multiple enzyme species having the same reaction specificity, the substrate coverage index is within the top approximately 20%. (Item 3) The enzyme is the enzyme described in any one of the above items, having substrate specificity greater than or equal to the reference substrate specificity of the enzyme. (Item 3a) The enzyme is an enzyme described in any one of the above items, having approximately three or more substrate specificities. (Item 4) A composition comprising an enzyme described in any one of the above items, for use in systematic synthesis of substances. (Item 5) A method for selecting a predetermined enzyme as a template enzyme from amino acid sequence information only, nucleic acid sequence information only, or a combination thereof, a) A method comprising the step of associating an amino acid sequence, nucleic acid sequence, or combination thereof of an enzyme candidate for the template enzyme with an enzyme activity base unit. (Item 6) The method according to any one of the above items, wherein the enzyme activity base unit directly or indirectly specifies catalytic efficiency and / or substrate specificity. (Item 7) The method according to any one of the above items, wherein the basic unit of enzyme activity is identified by an element selected from the group consisting of reaction specificity, coenzyme, substrate species, reaction temperature, reaction pH, molecular crowding during the reaction, and reaction pressure. (Item 8) The method according to any one of the above items, further comprising the step of modifying the aforementioned enzyme activity base unit to an activity not previously known. (Item 9) The enzyme is an enzyme, composition, or method described in any one of the above items, selected by the method of this disclosure. (Item A1) An enzyme library comprising an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, wherein the enzyme library comprises two or more different enzymes. (Item A2) The enzyme has a catalytic efficiency equal to or higher than the reference catalytic efficiency of the enzyme, and is an enzyme library according to any one of the above items. (Item A3) The enzyme has a substrate specificity equal to or higher than the reference substrate specificity of the enzyme, and is an enzyme library according to any one of the above items. (Item A4) An enzyme library according to any one of the above items, comprising three or more different kinds of the enzymes. (Item A5) The enzyme is an enzyme library according to any one of the above items, including oxidoreductase, transferase, hydrolase, lyase, isomerase, ligase, translocase. (Item A6) The enzyme is an enzyme library according to any one of the above items, including alcohol dehydrogenase (ADH), decarboxylase, dehydratase, oxygenase, dehydrogenase, oxidase, aminotransferase, transaldolase, esterase, peptidase, glycosidase, carboxylase, hydrolase, racemase, mutase, intramolecular lyase, carboxylase, amino acid transporter, sugar transporter, organic acid transporter, hydrogenase, halogenase, nitrogenase, methyltransferase, phosphorylase, phosphatase, kinase, sulfatase, nuclease, ATPase, and metal transporter. (Item A7) An enzyme library according to any one of the above items for use in systematic substance synthesis. (Item A8) The enzyme is an enzyme library according to any one of the above items selected by the method of the present disclosure. (Item B1) A method for providing an enzyme or a combination of enzymes for producing a desired metabolite, a) identifying the raw material of the metabolite and the necessary enzymes; b) selecting the enzyme from an enzyme library and comprising the method. (Item B2) The method according to any one of the above items, wherein the enzyme library is the enzyme library according to any one of the above items. (Item C1) A method for producing a desired metabolite, comprising: a) identifying the raw material of the metabolite and the necessary enzymes; b) selecting the enzyme from an enzyme library; c) placing the selected enzyme and the raw material under conditions where the reaction for metabolite production proceeds and including. (Item C2) The method according to any one of the above items, wherein the enzyme library is the enzyme library according to any one of the above items. (Item C3) The tetrahydroisoquinoline analogs of the following formula produced by the method according to any one of the above items. JPEG2023152952000001.jpg8989 (Item C4) The tetrahydroisoquinoline analogs of the following formula. JPEG2023152952000002.jpg8989 (Item D1) A method for selecting a template enzyme only from sequence information, comprising: reducing the sequence information of a plurality of enzymes that are candidates for the template enzyme to a dimension number less than the dimension number of the basic unit indicating the sequence information; classifying each enzyme based on the component unit information corresponding to the reduced dimension number; selecting the enzyme classified into a specific group as the template enzyme; and including. (Item D2) The method according to any one of the above items, wherein the reduction includes vectorizing the sequence information using a one-hot vector conversion table. (Item D3) The classification is a method according to any one of the above items, including principal component analysis (PCA). (Item D4) Furthermore, the method according to any one of the above items, further comprising the step of measuring the activity of enzymes classified into the specific group. (Item D5) The selection process is the method described in any one of the above items, in which an enzyme located near the median is selected as a template enzyme. (Item E1) A method for preparing a template enzyme library, A step of selecting three or more template enzymes from multiple candidate enzymes using the method described in any one of the above items, A step of providing an enzyme library containing three or more of the above-mentioned template enzymes as a template enzyme library. Methods that include...
[0005] In this disclosure, one or more of the above features may be provided in combinations other than those explicitly stated. Further embodiments and advantages of this disclosure will be apparent to those skilled in the art, by reading and understanding the detailed description below as necessary.
[0006] Furthermore, any other features and notable effects of this disclosure will become clear to those skilled in the art by referring to the following sections on embodiments of the invention and the drawings. [Effects of the Invention]
[0007] This disclosure provides a method for selecting enzymes with a predetermined level of reaction specificity. By using this method, it is possible to provide template enzymes that carry out the desired enzymatic reaction with high accuracy, and efficient production of the target compound can be expected. [Brief explanation of the drawing]
[0008] [Figure 1]Figure 1 is a schematic diagram of the metabolic pathway of a useful compound. The biosynthetic pathway from the hub compound to the useful compound utilizes enzymes with common reaction specificity. [Figure 2] Figure 2 is a flowchart of a method for selecting a template enzyme according to one embodiment of this disclosure. Sequence information can be vectorized and its dimensions reduced by principal component analysis. [Figure 3] Figure 3 shows the sequence distribution (left) and phylogenetic tree analysis results (right) of the template enzyme selection results when alcohol dehydrogenase (ADH) is used in one embodiment of the present disclosure. [Figure 4] Figure 4 shows the results of measuring the enzyme activity of a selected template enzyme when alcohol dehydrogenase (ADH) is used in one embodiment of this disclosure. [Figure 5] Figure 5 shows the results of analysis of α-keto acid decarboxylase by principal component analysis in one embodiment of the present disclosure. [Figure 6] Figure 6 shows the results of structural analysis of α-keto acid decarboxylase and various substrates using MOE (Molecular Operating Environment) in one embodiment of this disclosure. [Figure 7] Figure 7 shows the structures of 24 substrates of α-keto acid decarboxylase in one embodiment of the present disclosure. [Figure 8] Figure 8 shows the structures of various products obtained by reacting α-keto acid decarboxylase with 24 different substrates in one embodiment of the present disclosure. [Figure 9] Figure 9 shows the results of activity evaluation of a selected template enzyme when using α-keto acid decarboxylase in one embodiment of the present disclosure. [Figure 10] Figure 10 is a schematic diagram of plasmid DNA incorporating the LlKdcA gene and the norcoclaurine synthase gene. [Figure 11] Figure 11 shows the qualitative analysis results of THIQ1-6 by LC-MS / MS in one embodiment of the present disclosure. [Figure 12] Figure 12 shows the qualitative analysis results of THIQ7-12 by LC-MS / MS in one embodiment of the present disclosure. [Figure 13] Figure 13 shows the qualitative analysis results of THIQ13-18 by LC-MS / MS in one embodiment of the present disclosure. [Figure 14] Figure 14 shows the qualitative analysis results of THIQ19-24 by LC-MS / MS in one embodiment of the present disclosure. [Figure 15] Figure 15 shows the structures of tetrahydroisoquinoline analogs (THIQ1-24) obtained using LlKdcA and norcoclaurine synthase derived from Coptis japonica in one embodiment of the present disclosure. [Modes for carrying out the invention]
[0009] The present disclosure is described below in best form. Throughout this specification, singular expressions should be understood to include the concept of their plural form unless otherwise specified. Accordingly, singular articles (e.g., "a," "an," "the" in English) should be understood to include the concept of their plural form unless otherwise specified. Furthermore, terms used herein should be understood to have the meaning commonly used in the art unless otherwise specified. Accordingly, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. In case of any conflict, this specification (including definitions) shall prevail.
[0010] The following provides definitions of terms used specifically in this specification and / or basic technical concepts as appropriate.
[0011] In this specification, "approximately" means ±10% of the following number.
[0012] In this specification, "enzyme" is construed in a broad sense and is a general term for macromolecular compounds centered on proteins that catalyze any reaction, and which mediate or accelerate chemical changes in vivo without themselves changing or decomposing. The substance that reacts with an enzyme is called a "substrate", and in this specification, "substrate" is construed in the broadest sense. Each enzyme has determined what substance serves as a substrate and what it is changed into, and this is called the specificity of the enzyme. In the present disclosure, such specificity can be utilized. Also, in this specification, "enzyme species" may indicate the origin of such an enzyme.
[0013] In this specification, "catalytic efficiency" means k cat / K m (where K m is the Michaelis constant, and k catは refers to the turnover number indicating the number of enzyme reactions performed by one enzyme per unit time.) It is expressed as the ratio of, and this value reflects the result of comparing the number of times a certain enzyme reacts (converts the corresponding substrate) with the number of times that enzyme forms a complex with that substrate. Therefore, the more efficient an enzyme is, the larger the value of k cat / K m .
[0014] In this specification, "reference catalytic efficiency" refers to the minimum criterion for ensuring the desired catalytic efficiency, and among a group of enzyme species having the same reaction specificity, for an enzyme that occupies a predetermined upper position for the catalytic efficiency (k cat / K m ) with respect to a plurality of substrate species, it refers to the lowest value of the catalytic efficiency of the plurality of substrate species.
[0015] Specifically, the reference catalytic efficiency is for enzyme species having the same reaction specificity (for example, among about 30 or more, the catalytic efficiency (k cat / K m) and for enzymes that rank in the top approximately 20% for each of the multiple substrate species, it can also refer to the lowest catalytic efficiency value for each of the multiple substrate species. For example, assuming there are 30 types of enzymes for the substrates lactide, glycerate, and malate, enzyme 1 and enzyme 2 refer to the lowest catalytic efficiency when all three substrates rank within the top 6.
[0016] In this specification, the "substrate coverage index" indicates the proportion of potential substrate candidates that can actually be used as substrates, and is expressed as the sum of the catalytic efficiencies of each substrate selected from the basic substrate set or their common logarithms.
[0017] In this specification, a "basic substrate set" refers to a representative set of substrates for an enzyme that catalyzes a particular reaction. The types of substrates can be appropriately selected depending on the type of reaction catalyzed by the enzyme. For reduction reactions, a set can be created that includes acetaldehyde, methylglyoxal, etc., while for oxidation reactions, a set can include ethanol, benzyl alcohol, etc. For example, alcohol dehydrogenase catalyzes the reduction of carbonyls in the R1-C(=O)-R2 (ketone) or RC(=O)-H (aldehyde) structures, so any substrate can be selected, but the "basic" condition is that the substrate's side chain contains one of the functional groups or substituents listed as alkyl, aryl, carboxyl, or alkyl alcohol. Therefore, specifically, a "basic substrate set" for reduction reactions can be a carbonyl compound having one of alkyl, aryl, carboxyl, or alkyl alcohol as a side chain. Specifically, isobutyraldehyde (representative alkyl aldehyde), cyclohexanone (representative cyclic ketone), phenylacetaldehyde (representative aryl aldehyde), pyruvate (representative carboxyl side chain), and 4-hydroxy-2-butanone (representative alkyl alcohol side chain) can be used as the basic substrate set. In oxidation reactions, these catalyze the oxidation reaction of alcohols that are R1-CH(-OH)-R2 (secondary alcohol) or R-CH2-OH (primary alcohol). Therefore, the "basic substrate set" for oxidation reactions can be alcohol compounds having one of the following side chains: alkyl, aryl, carboxyl, or alkyl alcohol.
[0018] In one embodiment, for example, in the case of an α-keto acid decarboxylation reaction, a set including pyruvate, phenylglyoxylic acid, etc. can be prepared. For example, any α-keto acid decarboxylase can be selected because it catalyzes the elimination reaction of the carboxyl group of the substrate R1-C(=O)-COOH, but the "basic" condition is that the substrate side chain contains one of the functional groups or substituents of alkyl, aryl, carboxyl, or alkyl alcohol. Therefore, specifically, the "basic substrate set" for an α-keto acid decarboxylation reaction can be said to be an α-keto acid compound having one of alkyl, aryl, carboxyl, or alkyl alcohol as a side chain.
[0019] In this specification, "substrate specificity" means that, in the cleavage of a substrate catalyzed by an enzyme, the enzyme does not catalyze the cleavage of substances other than the substrate, or the degree of catalysis is sufficiently weak.
[0020] In this specification, "reference substrate specificity" refers to the number of substrates that an enzyme can react with, or the proportion of possible substrates, when comparing a predetermined number (e.g., about 30) or more enzyme species from a group of enzymes with the same reaction specificity, and the number of substrates that an enzyme can react with is within the top of a certain number (e.g., about 20%). Specifically, for example, if there are 30 types of enzymes for substrates such as lactide (substrate 1), glycerate (substrate 2), malate (substrate 3), substrate 4, substrate 5, and substrate 6, and the 5th ranked enzyme reacts with 5 substrates, the 6th ranked enzyme reacts with 4 substrates, and the 7th ranked enzyme reacts with 3 substrates, then the 6th ranked enzyme, which corresponds to 20%, is "4" because it reacts with 4 substrates, and if the survey included 30 types of substrates, this can be calculated as 13.3%. Typically, reference substrate specificity can be expressed as a real number rather than a percentage, for example, as "3 or more types." For example, a template enzyme can be defined as having a catalytic efficiency factor (Log) for each substrate. 10The top 20% can be calculated by summing the catalytic efficiency values. Specifically, for example, if there are 30 types of enzymes, and in a study with 30 types of substrates, 20 of them satisfy the standard substrate specificity, then the top 20% (= enzyme species 4) can be calculated.
[0021] In this specification, "enzyme activity base unit" refers to a base unit that defines enzyme activity. Typically, it directly or indirectly specifies catalytic efficiency and / or substrate specificity. It can be associated with the amino acid sequence, nucleic acid sequence, or combination thereof of a candidate enzyme for a template enzyme. The enzyme activity base unit can be identified by elements selected from the group consisting of reaction specificity, coenzymes, substrate type, reaction temperature, reaction pH, molecular crowding during the reaction, and reaction pressure.
[0022] In this specification, "reaction specificity" refers to one of the specific properties of an enzyme, specifically the property that an enzyme exhibits catalytic activity only for certain reactions.
[0023] In this specification, "coenzyme" refers to a compound essential for the activity of certain enzymes. Examples of coenzymes include "flavin adenine dinucleotide (FAD)," "thiamine pyrophosphate (ThPP)," and "flavin mononucleotide (FMN)."
[0024] In this specification, "substrate species" is synonymous with "substrate," and refers to a compound whose reaction is catalyzed by an enzyme. The substrate binds to a specific site (active site) on the surface of the enzyme molecule, forming an enzyme-substrate complex, and is converted into the product.
[0025] In this specification, "reaction temperature" refers to the temperature at which an enzyme carries out a reaction, and varies depending on the type of enzyme. For example, reaction temperatures can range from approximately 0 to approximately 100°C.
[0026] In this specification, "reaction pH" refers to the pH at which an enzyme carries out a reaction, and it varies depending on the type of enzyme. For example, reaction pH can range from approximately 2 to approximately 12.
[0027] In this specification, "molecular crowding during reaction" refers to the concentration of the reaction solution, which mainly consists of organic substances such as proteins, nucleic acids, sugars, lipids, and amino acids, and inorganic substances such as salts, and varies depending on the type of enzyme. For example, reaction solution concentrations can range from approximately 10 g / L to approximately 400 g / L.
[0028] In this specification, "reaction pressure" refers to the pressure at which an enzyme carries out a reaction, and it varies depending on the type of enzyme. For example, reaction pressures can range from approximately 0.1 to approximately 200 megapascals.
[0029] In this specification, “principal component analysis” refers to a statistical data analysis technique that summarizes many quantitative explanatory variables into fewer indicators or composite variables. In this disclosure, it can be used to reduce the sequence information of multiple enzymes that are candidates for a template enzyme to a number of dimensions less than the number of dimensions of the basic units that represent the sequence information. For example, first, the full-length sequence of the target sequence is obtained from publicly available databases such as the National Center for Biotechnology Information (NCBI), Uniprot (Uniprot Consortium), and KEGG Bioinformatics (Institute for Chemical Research, Kyoto University). Then, the amino acid sequence is vectorized using a programming language, for example, according to a numerical table consisting of 21 rows x 5 columns. A two-dimensional plot is obtained by principal component analysis of the vectors, and the principal component analysis here can follow general analytical methods.
[0030] (Preferred embodiment) Preferred embodiments of the Disclosure are described below. The embodiments provided below are provided for a better understanding of the Disclosure, and the scope of the Disclosure should not be limited to the descriptions below. It will be apparent that those skilled in the art can make appropriate modifications within the scope of the Disclosure, taking into consideration the descriptions herein. Furthermore, the embodiments of the Disclosure below can be used individually or in combination.
[0031] (Template enzymes and their selection methods) In one aspect of this disclosure, an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity is provided. In this disclosure, an enzyme having a catalytic efficiency and / or substrate specificity of a standard or higher is provided, and such an enzyme having a predetermined or higher catalytic efficiency and / or substrate specificity is referred to in this disclosure as a template enzyme. Therefore, in one embodiment of this disclosure, an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity is provided. In another aspect of this disclosure, an enzyme library containing such an enzyme is provided, comprising two or more different enzymes.
[0032] Designing biosynthetic pathways for the production of useful compounds, particularly the selection of enzymes for each reaction step, can be explored based on known information such as published literature, but this process heavily relies on the researcher's experience. Furthermore, information on the enzymatic activity of only a small fraction of the vast number of enzymes is available, making it difficult to perform highly accurate screening using conventional enzyme selection methods and biosynthetic pathway design methods based on knowledge and experience. Therefore, there is a need for the development of enzyme selection methods that do not rely on the knowledge of a single researcher, and for efficient methods of designing target biosynthetic pathways based on these methods.
[0033] When the biosynthetic pathways of many useful compounds (e.g., general-purpose compounds, high-performance compounds, secondary metabolites, unnatural compounds, etc.) are investigated, it is found that enzymes with common reaction specificity are used in each, as shown in the schematic diagram of the metabolic pathway in Figure 1. In one embodiment of this disclosure, by collecting various enzymes with such common reaction specificity, especially enzymes with high catalytic efficiency and broad substrate specificity, as template enzymes, it is possible to rapidly design the biosynthetic pathway of the target. In one embodiment of this disclosure, as a method for selecting template enzymes, classification can be performed using only the amino acid sequence of the enzyme. For example, as shown in this embodiment, when the activity of enzymes belonging to the group presumed to be template enzymes was measured individually, using alcohol dehydrogenase (ADH) enzyme as an example, it was shown that they have broad substrate specificity and high catalytic efficiency.
[0034] The template enzyme selection method of this disclosure allows for the grouping of enzymes by first converting the amino acid sequence into numerical matrix data and then analyzing the matrix information. Therefore, in one aspect of this disclosure, a method is provided for selecting a template enzyme solely from sequence information, comprising the steps of: reducing the sequence information of a plurality of candidate enzymes for the template enzyme to a number of dimensions less than the number of dimensions of the basic unit representing the sequence information; classifying each enzyme based on component unit information corresponding to the reduced number of dimensions; and selecting the enzymes classified into a specific group as template enzymes.
[0035] Regarding the numerical and matrix representation of amino acid sequences, it is known that amino acids can be converted into a form usable for calculations based on conversion rules such as AAindex, which are based on the chemical structure and properties of amino acids (for example, Kawashima et al., Nucleic Acids Res. 2008 Jan;36(Database issue):D202-5). In one embodiment of this disclosure, the method of this disclosure is used. By converting the data into a simpler vector consisting of 0s and 1s, the computation speed is significantly improved, making it possible to compute a vast amount of sequence information at once. In one embodiment, principal component analysis can be performed as a matrix analysis method to indicate the group to which the template enzyme belongs from the overall picture of the sequence distribution.
[0036] In one embodiment, as a strategy for selecting a template enzyme, the range of the plot population may be arbitrarily determined from the sequence distribution. Next, enzymes with high sequence conservation among the population can be selected as template enzyme candidates, and catalytic efficiency data for each substrate can be obtained. Subsequently, the population of template enzymes can be identified based on the obtained data. Therefore, in one embodiment, the method of this disclosure allows for the selection of template enzymes from the sequence distribution in a catalytic efficiency data-driven manner.
[0037] In other embodiments, if sufficient enzyme data has been accumulated, feature selection becomes possible from the correlation between sequences and enzyme data, and template enzymes can be selected from the sequence distribution by using regression methods.
[0038] In one embodiment of this disclosure, during the principal component analysis process, five columns of 0s and 1s can be assigned to each amino acid so that the variance is maximized in the calculation of the variance-covariance matrix. Therefore, the sequence distribution can be changed according to the vectorization.
[0039] In one embodiment, the vectorization and principal component analysis in the method for selecting template enzymes according to this disclosure can be calculated based on the amino acid sequence information of the enzyme group that performs the target reaction, such as by the EC number. Subsequently, the enzymes of the group estimated to be template enzymes are expressed and the individual enzyme activity is measured.
[0040] Figure 2 shows an example of the flow chart for selecting a template enzyme according to this disclosure. First, the sequence information of an arbitrary enzyme is vectorized, and the number of dimensions is reduced by principal component analysis. For example, when amino acids are used as the sequence information, the number of dimensions is reduced to fewer than the 21 types of dimensions (20 types of amino acids + gap sequence(-)) (e.g., PC1 to PC5). In other embodiments, when nucleic acid sequences are used, they can be converted into amino acid sequences according to the host codon table, and then the number of dimensions can be reduced in the same way.
[0041] In one embodiment, the transformation of sequence information can be performed using, for example, one-hot vectorization, but is not limited to this, as long as it can transform the sequence information of the enzyme group performing the target reaction into a format that can be written to a program. In another embodiment, dimensionality reduction can be performed using principal component analysis (PCA). Any principal component analysis that can reduce the number of dimensions of the sequence information to a smaller number of dimensions is acceptable. In other embodiments, dimensionality reduction can be performed using k-means clustering or self-organizing maps, which are representative unsupervised learning methods. k-means clustering is a method that classifies unknown data into the cluster closest to an arbitrary representative point, and self-organizing maps are a method that displays multidimensional data on a two-dimensional plane while preserving the phase relationships. Both differ from PCA only in their algorithms and can be used for dimensionality reduction. For more information on k-means clustering and self-organizing maps, see Forgy, EW (1965). Biometrics, International Biometric Society, Wiley-Blackwell; MacQueen, JB (1967). Proceedings of 5th Berkeley Symposium on Mathematical Statistics and Probability, University of California Press; Anderberg, MR (1973). Cluster analysis for application, Academic Press; Hastie, T et al. "Fundamentals of Statistical Learning - Data Mining, Inference, and Prediction," Kyoritsu Shuppan, 2014; Teubon Kohonen, (2001). Self-organizing Maps, Springer; Teuvo Kohonen, "Self-organizing Maps," Springer Japan, 2016, etc., and relevant parts (maybe all) of these are incorporated herein by reference.
[0042] In one embodiment of this disclosure, enzymes from a group presumed to be template enzymes can be expressed, and their individual enzyme activity can be measured to confirm whether they can function as template enzymes. In this case, any method capable of appropriately measuring enzyme activity is acceptable, and any known method can be appropriately selected and used.
[0043] In one embodiment, the step of selecting an enzyme classified into a specific group as a template enzyme can be performed by, for example, estimating an enzyme whose sequence is located near the median after dimensionality reduction by principal component analysis as the template enzyme. However, the specific group is not limited to sequences located near the median. For example, in one embodiment, the selection of a template enzyme can be performed by arbitrarily defining the range of the plotted sequence distribution population using a machine learning method such as the K-means method, and then selecting an enzyme with high sequence conservation among the populations as a candidate template enzyme.
[0044] In one embodiment of the present disclosure, the selection of a template enzyme of the present disclosure is performed by measuring the activity of the candidate template enzyme to confirm whether the candidate enzyme can serve as a template enzyme. Such activity measurement can be performed, for example, by obtaining data on reaction specificity and substrate specificity at room temperature, atmospheric pressure, near neutral (pH 7.0), and with an appropriate reaction composition. Accordingly, in other aspects of the present disclosure, a method is provided for selecting a predetermined enzyme as a template enzyme from amino acid sequence information alone, nucleic acid sequence information alone, or a combination thereof, the method comprising the step of associating the amino acid sequence, nucleic acid sequence, or combination thereof of the candidate enzyme for the template enzyme with the basic unit of enzyme activity.
[0045] In one embodiment, the enzyme activity base unit can directly or indirectly specify catalytic efficiency and / or substrate specificity. For example, the enzyme activity base unit can be specified by elements selected from the group consisting of reaction specificity, coenzyme, substrate type, reaction temperature, reaction pH, molecular crowding during the reaction, and reaction pressure.
[0046] Furthermore, in one embodiment of this disclosure, the basic enzyme activity unit can be modified to have unprecedented activity. In this way, with respect to the template enzyme of this disclosure, it is possible to create a group of artificial enzymes including even more highly active enzymes and novel active enzymes, depending on the useful compound to be produced using that enzyme. In addition, by preparing as many template enzymes as needed, classified and selected by reaction specificity, it becomes possible to quickly realize highly efficient reaction pathways according to needs.
[0047] In one embodiment of the present disclosure, the template enzyme of the present disclosure can have a catalytic efficiency greater than or equal to the standard catalytic efficiency of the enzyme. In one embodiment, for example, in the case of an alcohol dehydrogenase template enzyme, among 30 or more enzyme species having the same reaction specificity, a group of enzymes whose catalytic efficiencies (kcat / Km) for multiple (preferably three or more) substrate species occupy the top 20% can be cited, and in this case, the template enzyme can be evaluated as having a standard catalytic efficiency. That is, in this case, by comparing the catalytic efficiency data of 30 template enzyme candidates, the probability of obtaining a group of template enzymes (or the probability of finding a group of template enzymes) can be 80% or more.
[0048] In one embodiment of the present disclosure, the template enzyme of the present disclosure is evaluated by a substrate coverage index, which can be the sum of the catalytic efficiencies of each substrate selected from a basic substrate set or their common logarithms. In another embodiment, the template enzyme of the present disclosure can be selected from a plurality of enzyme species having the same reaction specificity, with the substrate coverage index being within the top approximately 20%. For example, in one embodiment of the present disclosure, for a plurality of enzyme species, the catalytic efficiency (catalytic efficiency factor = Log) expressed as a common logarithm for each substrate is used. 10 The top 20% of enzymes, ranked based on the total score obtained by adding the catalytic efficiency values, can be designated as template enzymes (for example, if there are 30 enzymes, 6 enzymes can be designated as template enzymes). A higher total score suggests a greater number of substrate species and higher catalytic efficiency for each.
[0049] In one embodiment, in the case of an alcohol dehydrogenase-catalyzed alcohol dehydrogenase reaction, the alcohol dehydrogenase catalyzes the reduction reaction of the carbonyl group of the substrate R1-C(=O)-R2 (ketone) or RC(=O)-H (aldehyde). A "basic" condition for this catalysis is that the side chain of the substrate contains one of the following functional groups or substituents: "alkyl, aryl, carboxyl, alkyl alcohol".
[0050] Therefore, for example, the "basic substrate set" for reduction reactions can be described as carbonyl compounds having one of the following side chains: alkyl, aryl, carboxyl, or alkyl alcohol. Specifically, examples include isobutyraldehyde (a representative example of alkyl aldehydes), cyclohexanone (a representative example of cyclic ketones), phenylacetaldehyde (a representative example of aryl aldehydes), pyruvate (a representative example of carboxyl side chains), and 4-hydroxy-2-butanone (a representative example of alkyl alcohol side chains).
[0051] In oxidation reactions, it catalyzes the oxidation of the substrate alcohol, which is either R1-CH(-OH)-R2 (secondary alcohol) or R-CH2-OH (primary alcohol). Therefore, the "basic substrate set" for oxidation reactions can be described as an alcohol compound having one of the following side chains: alkyl, aryl, carboxyl, or alkyl alcohol.
[0052] In one embodiment of the present disclosure, the template enzyme of the present disclosure can have substrate specificity greater than that of the reference substrate specificity of the enzyme, for example, it can have about three or more substrate specificities. As described above, in one embodiment of the present disclosure, the template enzyme of the present disclosure has a catalytic efficiency (catalytic efficiency factor = Log) expressed on a common logarithmic scale for each substrate for a certain number of enzyme species. 10 By adding the catalytic efficiency value, the top 20% of the enzymes can be selected based on the total score obtained. Therefore, in one embodiment, the reference substrate specificity can be defined as the number of substrates that an enzyme can react with is within the top approximately 20% when comparing about 30 or more enzyme species with the same reaction specificity.
[0053] In one embodiment of the present disclosure, the template enzyme of the present disclosure may be used for systematic synthesis of substances. As described above, the template enzyme of the present disclosure is an enzyme that can commonly catalyze a hub compound in the biosynthesis pathway of a particular useful compound. Therefore, by using the template enzyme of the present disclosure, any compound can be produced in high volume from a common precursor (hub compound) in a short period of time.
[0054] The template enzymes selected as described above can be organized into libraries according to their reaction specificity. Therefore, in one aspect of this disclosure, a method for preparing a template enzyme library is provided, comprising the steps of: selecting three or more template enzymes from a plurality of candidate enzymes using the template enzyme selection method of this disclosure; and providing an enzyme library containing the three or more template enzymes as a template enzyme library.
[0055] In one embodiment, the library obtained in this manner may be an enzyme library containing enzymes having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, and comprising two or more different enzymes, preferably three or more of the said enzymes.
[0056] In one embodiment of the present disclosure, such template enzymes may include, but are not limited to, oxidoreductases, transferases, hydrolases, lyases, isomerases, ligases, and translocases. In other embodiments, the template enzymes of the present disclosure may include, but are not limited to, alcohol dehydrogenases (ADHs), decarboxylases, dehydratases, oxygenases, dehydrogenases, oxidases, aminotransferases, transaldolases, esterases, peptidases, glycosidases, carboxylyases, hydrolyases, racemases, mutases, intramolecular lyases, carboxylases, amino acid transporters, sugar transporters, organic acid transporters, hydrogenases, halogenases, nitrogenases, methyltransferases, phosphorylases, phosphatases, kinases, sulfatases, nucleases, ATPases, and metal transporters.
[0057] In one aspect of this disclosure, a method for providing an enzyme or combination of enzymes for producing a desired metabolite, comprising the steps of: a) identifying the raw materials and required enzymes for the metabolite; A method is provided which includes the step of (b) selecting the enzyme from an enzyme library. In one embodiment, such a library may be one of those described elsewhere in the Disclosure. By using such a library, template enzymes can be classified according to their reaction specificity, so that a desired metabolite can be rapidly produced simply by selecting the necessary enzyme according to the desired metabolite. Accordingly, in one aspect of the Disclosure, a method is provided for producing a desired metabolite, which includes the steps of (a) identifying the raw materials for the metabolite and the necessary enzyme, (b) selecting the enzyme from an enzyme library, and (c) arranging the selected enzyme and the raw materials under conditions that allow the reaction for metabolite production to proceed.
[0058] (General technology) The molecular biological, biochemical, and microbiological methods used herein are well-known and commonly used in their respective fields, as seen, for example, Sambrook J. et al. (1989). Molecular Cloning: A Laboratory Manual, Cold Spring Harbor and its 3rd Ed. (2001); Ausubel, FM (1987). Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience; Ausubel, FM (1989). Short Protocols in Molecular Biology: A Compendium of Methods. from Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience; Innis, MA(1990).PCR Protocols: A Guide to Methods and Applications, Academic Press; Ausubel, FM(1992).Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Greene Pub. Associates; Ausubel, FM (1995).Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Greene Pub. Associates; Innis, MA et al. (1995). PCR Strategies, Academic Press; Ausubel, FM (1999). Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Wiley, and annual updates; Sninsky, JJ et al.(1999). These methods are described in publications such as "PCR Applications: Protocols for Functional Genomics," Academic Press, and "Experimental Methods for Gene Transfer & Expression Analysis," Yodosha, 1997. Relevant parts (or possibly all) of these publications are referenced herein.
[0059] For DNA synthesis techniques and nucleic acid chemistry to create artificially synthesized genes, gene synthesis and fragment synthesis services such as GeneArt, GenScript, and Integrated DNA Technologies (IDT) can be used. Other resources include, for example, Gait, MJ (1985). Oligonucleotide Synthesis: A Practical Approach, IRL Press; Gait, MJ (1990). Oligonucleotide Synthesis: A Practical Approach, IRL This information is found in publications such as Press; Eckstein, F. (1991). Oligonucleotides and Analogues: A Practical Approach, IRL Press; Adams, RL et al. (1992). The Biochemistry of the Nucleic Acids, Chapman & Hall; Shabarova, Z. et al. (1994). Advanced Organic Chemistry of Nucleic Acids, Weinheim; Blackburn, GM et al. (1996). Nucleic Acids in Chemistry and Biology, Oxford University Press; Hermanson, GT (1996). Bioconjugate Techniques, Academic Press, and relevant parts of these publications are referenced herein.
[0060] In this specification, "or" is used when "at least one" of the items listed in the text can be adopted. The same applies to "or else". In this specification, when it is specified that "within the range" of "two values", that range includes the two values themselves. References such as scientific literature, patents, and patent applications cited herein are incorporated herein by reference to the same extent as they are specifically described herein.
[0061] The present disclosure has been described above with reference to preferred embodiments for ease of understanding. The present disclosure will now be described based on examples, but the above description and the following examples are provided for illustrative purposes only and not to limit the present disclosure. Accordingly, the scope of the present disclosure is not limited to the embodiments or examples specifically described herein, but is limited only by the claims. [Examples]
[0062] (Example 1: Template enzyme for alcohol dehydrogenase) As a validation experiment, we searched for template enzymes using alcohol dehydrogenase (ADH). Approximately 2400 ADH amino acid sequences were vectorized, and principal component analysis was performed.
[0063] The full-length sequences were obtained from publicly available databases such as the National Center for Biotechnology Information (NCBI), Uniprot (Uniprot Consortium), and KEGG Bioinformatics (Institute for Chemical Research, Kyoto University). The amino acid sequences were vectorized based on a 21x5 numerical table using programming languages such as Python. Two-dimensional plots were obtained by principal component analysis of the multidimensional vectors. The principal component analysis followed general analytical methods. The results are shown in Figure 3.
[0064] The left panel of Figure 3 shows the results of principal component analysis of 2445 alcohol dehydrogenase sequences. Based on the sequence distribution, they were arbitrarily classified into 6 groups. The right panel of Figure 3 shows the results of phylogenetic tree analysis of alcohol dehydrogenases. The names of each enzyme in the phylogenetic tree correspond to the group names in the left panel. Specifically, Group 1 corresponds to enzymes 22 and 23, Group 2 to enzymes 1 and 24, Group 3 to enzymes 3, 4, 20, 21, and 25-28, Group 4 to enzymes 15-17, 19, 29, and 30, Group 5 to enzymes 6 and 31, and Group 6 to enzymes 32 and 33. Asterisks indicate that enzyme structural information is available.
[0065] Next, eight sequences were selected from the template enzyme group estimated by calculation, and 14 sequences were selected from the group of enzymes estimated not to be template enzymes. Each of these enzymes was then expressed in E. coli. Purified enzymes were then obtained, and their activity was measured against 40 different substrates (Figure 4).
[0066] In Figure 4, for various aldehyde, ketone, and alcohol substrates, data indicating a catalytic efficiency of 0.1 mM-1·s-1 or higher for each enzyme is shown as "active," while data below this level is shown as "inactive." The enzymes evaluated and their origins are as follows. G1-1, Gordonia sputi (H5U1X9); G1-2, Rhodococcus sp. (A0A260B926); G2-1, Arabidopsis thaliana (Q96533); G2-2, Escherichia coli (P25437); G3a-1, Halomonas sp. (A0A1R4HSH9);G3a-2,Trabulsiella guamensis (A0A084ZMM2);G3y-1,Agrococcus casei (A0A1R4G3L4);G3y-2,Tatumella ptyseos (A0A085JKH3);Ahr,Escherichia coli (P27250);YahK,Escherichia coli (P75691);ADH6,Saccharomyces cerevisiae (Q04894);ADH7,Saccharomyces cerevisiae (P25377);G4-1,Schizosaccharomyces pombe (P00332);G4-2,Caenorhabditis elegans (O45687);ADH1,Saccharomyces cerevisiae (P00330);ADH2,Saccharomyces cerevisiae (P00331);ADH3,Saccharomyces cerevisiae (P07246);ADH5,Saccharomyces cerevisiae (P38113);G5-1,Zymomonas mobilis subsp. Mobilis (P20368); G5-2, Escherichia coli (P39451); G6-1, Corynebacterium glutamicum (Q8NLX9); G6-2, Komagataeibacter saccharivorans (A0A347WDM1).
[0067] The results showed that enzymes belonging to the template enzyme group were active against approximately 50-70% of substrate compounds, while enzymes belonging to other groups showed activity against at most about 40% of substrate compounds. The calculations revealed that the enzymes belonging to the template enzyme group possessed broad substrate specificity and exhibited characteristics of a template enzyme.
[0068] Enzyme activity was measured as follows: E. coli BL21(DE3) strain (Merck Millipore) was transformed using a pET plasmid (Merck Millipore) into which the alcohol dehydrogenase gene was inserted. The E. coli was cultured in 5 ml of LB medium (tryptone 1 g / l, yeast extract 0.5 g / l, NaCl 1 g / l) at 37°C. When the OD600 reached 0.5-1.0, 5 μl of 1 M isopropyl-β-thiogalactopyranoside (IPTG) was added, and the culture was then incubated overnight at 20°C. The E. coli was collected by centrifugation at 10000 × g for 10 minutes at 4°C. The cells were suspended in 50 μl of binding buffer (20 mM Tris-HCl (pH 8.0), 100 mM NaCl, 5 mM Imidazole). After disrupting the bacterial cells using a bead shocker (2500 rpm, ON / OFF 30 sec, 3 cycles), the supernatant was collected as crude enzyme solution after centrifugation at 15000 × g, 4°C, 20 minutes. The crude enzyme solution was subjected to 500 μl of TALON Metal Affinity Resin (TaKaRa), and contaminating proteins present in the column were removed with 4 ml of washing buffer 1 (20 mM Tris-HCl (pH 8.0), 100 mM NaCl, 10 mM Imidazole) and 1 ml of washing buffer 2 (20 mM Tris-HCl (pH 8.0), 100 mM NaCl, 20 mM Imidazole). Alcohol dehydrogenase was eluted with elution buffer (20 mM Tris-HCl (pH 8.0), 100 mM NaCl, 500 mM Imidazole). The alcohol dehydrogenase fraction was collected, and 70% glycerol was added to achieve a final concentration of 50%. The mixture was then stored at -30°C until use. Protein quantification was performed using the Bradford method with Bio-Rad protein assay reagent (Bio-Rad Corporation) with bovine serum albumin as the standard.
[0069] To investigate the reducing activity of alcohol dehydrogenase, the reaction reagents (10 μl of 50 mM MOPS buffer (pH 7.0), 3 μl of 10 mM NADPH / NADH) and purified alcohol dehydrogenase were pre-incubated (37 / 30°C, 10 min), and then 10 μl of substrates of various concentrations (5 different concentrations) were added to initiate the reaction. The reaction was carried out at 37 / 30°C, and the reduction activity associated with the oxidation of NADPH or NADH was observed. 340nm / A 370nm The decrease (Tecan, Infinite 200, microplate reader) was measured over time. The formula for calculating specific activity is as follows: units / mg = (Δ A340nm ( / min) / (6.2 × mg enzyme / ml reaction solution) or units / mg = (Δ A370nm (k) / (2.8 × mg enzyme / ml reaction solution). One unit was defined as the amount of enzyme that oxidizes 1 μmol of NADH or NADPH per minute under conditions of 37 / 30°C. Catalytic efficiency (k) cat / K m The specific activity was calculated based on the specific activity at any given concentration of the aldehyde substrate.
[0070] Regarding the oxidative activity of alcohol dehydrogenase, the reaction reagents used were 10 μl of 50 mM MOPS buffer (pH 7.0) and 10 mM NADP. + / NAD + After pre-incubating (3 μl) of purified alcohol dehydrogenase with NADP+ or NAD+ at 37 / 30°C for 10 min, 10 μl of substrates of various concentrations (5 different concentrations) were added to start the reaction. The reaction was carried out at 37 / 30°C, and the reaction was accompanied by the reduction of NADP+ or NAD+. 340nm / A 370nm The increase in (Tecan, Infinite 200, microplate reader) was measured over time. Specific activity and catalytic efficiency were calculated using the same method as for reduction reactions.
[0071] (Example 2: Template enzyme for α-keto acid decarboxylase) This study aims to search for template enzymes for α-keto acid decarboxylase (KDC). The full-length sequences of approximately 17,000 KDC amino acids are obtained from publicly available databases such as the National Center for Biotechnology Information (NCBI), Uniprot (Uniprot Consortium), and KEGG Bioinformatics (Institute for Chemical Research, Kyoto University). The KDC amino acid sequences are vectorized using a 21x5 numerical table based on programming language processing such as Python. A two-dimensional sequence distribution map is obtained using principal component analysis of the multidimensional vectors. Next, the range of the plot population is arbitrarily determined from the sequence distribution using machine learning techniques such as the K-means method. Enzymes with high sequence conservation within the population are selected as template enzyme candidates, and catalytic efficiency data for 20 types of α-keto acids is obtained. Based on the obtained data, enzymes with activity for approximately 50-80% of substrate compounds are identified. Characteristic sequences of the enzymes are extracted computationally and subjected to machine learning to identify the template enzyme group.
[0072] (Example 3: Production of target substance using an alcohol dehydrogenase template enzyme) The template enzyme for alcohol dehydrogenase is not limited to specific bacterial strains, but can be used for the production of various alcohols in hosts such as Escherichia coli, Bacillus subtilis, Pseudomonas, Actinomycetes, Methylobacterium, yeast, hydrogen bacteria, cyanobacteria, and microalgae.
[0073] To create a bacterial strain that expresses the alcohol dehydrogenase gene, the alcohol dehydrogenase template gene is inserted into the genome of the host microorganism. Alternatively, the alcohol dehydrogenase gene is inserted into a plasmid that can be replicated in the host microorganism, and the host microorganism is transformed using the created plasmid.
[0074] By adding a precursor to the culture medium of a host microorganism that possesses the alcohol dehydrogenase gene, the target substance can be produced. For example, in the production of phenethyl alcohol, a fragrance material, and isobutanol, which is used as fuel, production is possible by adding the precursors phenylacetaldehyde and isobutyraldehyde to the culture medium of a host microorganism that possesses the alcohol dehydrogenase gene. Alternatively, in the fermentation production of the target substance by a host microorganism, by introducing or enhancing the expression level of genes involved in the biosynthesis pathway of the precursors phenylacetaldehyde and isobutyraldehyde, phenethyl alcohol and isobutanol can be produced from the growth carbon source (glucose, CO2, methanol, etc.) of the host microorganism that possesses these genes.
[0075] (Example 4: Searching for template enzymes using k-means clustering or self-assembly maps) The acquisition of the full-length sequence is performed in the same manner as in Example 1. Using a programming language such as Python, the amino acid sequence is vectorized based on a numerical table consisting of 21 rows x 5 columns. In k-means clustering of multidimensional vectors, the number of clusters is arbitrarily set to obtain a two-dimensional plot. The k-means clustering method follows general analytical methods. Alternatively, a self-organizing map of multidimensional vectors is used to obtain a two-dimensional plot by providing a set of prepared input vectors at once for training. The self-organizing map follows general analytical methods. Template enzyme candidates are arbitrarily selected from the range of the obtained plot population. Enzyme activity is measured in the same manner as in Example 1. Among the enzyme candidates that showed activity, the characteristic sequences of the enzymes are extracted by calculation and subjected to machine learning to identify the template enzyme group.
[0076] (Example 5: Template enzyme for α-keto acid decarboxylase) We used α-keto acid decarboxylase (KdcA) to search for template enzymes.
[0077] Selection of α-keto acid decarboxylase Searching the protein database UniProt for "decarboxylases" possessing the three characteristic domains of α-keto acid decarboxylase—"TPP_Enzyme_N," "TPP_Enzyme_M," and "TPP_Enzyme_C"—resulted in 17,728 sequences. Furthermore, these were subdivided into "pyruvate decarbonase (EC4.1.1.1)", "benzoyl formate decarbonase (EC4.1.1.7)", "oxalyl CoA decarbonase (EC4.1.1.8)", "phenylpyruvate decarbonase (EC4.1.1.43)", "branched-chain 2-oxo acid decarbonase (EC4.1.1.72)", "indolepyruvate decarbonase (EC4.1.1.74)", "5-guanidino-2-oxopentanoate decarbonase (EC4.1.1.75)", "sulfopyruvate decarbonase (EC4.1.1.79)", "4-hydroxyphenylpyruvate decarbonase (EC4.1.1.80)", and "phosphonopyruvate decarbonase (EC4.1.1.82)". These were further classified by principal component analysis (Figure 5). Furthermore, we constructed the three-dimensional structures of various α-keto acid decarboxylases using alphafold2, and performed docking with 24 different substrates (α-keto acids, Keto1-24) using the molecular simulation software MOE (Molecular Operating Environment). We then calculated and compared the affinity of Keto1-24 (Figure 6).
[0078] Figure 7 shows the structures of Keto1-24. In silico analysis of these enzymes suggests that the enzymes in the vicinity of ScARO10, ScPDC1, and LlKdcA may be a group of enzymes that possess template enzyme properties.
[0079] Purification of α-keto acid decarboxylase Synthetic genes optimized for E. coli were prepared for various α-keto acid decarboxylases. These genes were incorporated into the pACYCDuet-1 vector, and plasmid DNA was obtained for expression as His-tag fusion proteins. This plasmid DNA was introduced into E. coli BL21(DE3), and expression was induced by IPTG. After extracting the proteins from the cultured E. coli, the target α-keto acid decarboxylase was obtained using affinity purification resin. The α-keto acid decarboxylases used are listed below. • ScARO10 (UniProt ID: Q06408, derived from Saccharomyces cerevisiae) • ScPDC1 (UniProt ID: P06169, derived from Saccharomyces cerevisiae) ·Zmpdc (UniProt ID:P06672, derived from Zymomonas mobilis subsp. mobilis) • PpmdlC (UniProt ID: P20906, derived from Pseudomonas putida) ·Ofoxc (UniProt ID:P40149, derived from Oxalobacter formigenes) ·LlKdcA (UniProt ID:Q6QBS4, derived from Lactococcus lactis) ·AbipdC (UniProt ID:P51852, derived from Azospirillum brasilense)
[0080] Sample and reagent preparation Each enzyme was tested for concentration using the Bradford method and its purity confirmed by SDS-PAGE. Finally, the concentration was adjusted to 0.01–1 μM using reaction buffer (50 mM MES, 0.5 mM TPP, 1 mM MgCl2, 100 mM KCl).
[0081] Keto1-24 were adjusted to 100 mM with distilled water or DMSO and stored at -80°C. They were then adjusted to 2 mM with reaction buffer before use. DMB (1,2-Diamino-4,5-methylenedioxybenzene) was adjusted to 100 mg / mL with distilled water and stored at -80°C. It was then adjusted to 10 mM with 1 M 2-ME / 30 mM Na2S2O5 / 1N HCl before use. ABAO (2-Aminobenzamidoxime) was then adjusted to 100 mM with 100 mM AcOH / AcONa buffer (pH 4.5) before use.
[0082] Activity evaluation of α-keto acid decarboxylase (DMB assay) 25 μL of the prepared enzyme was mixed with 25 μL of each α-keto acid and reacted at 30°C for 30 minutes. Figure 8 shows the products (aldehydes) Ald1-24 obtained when Keto1-24 were reacted as substrates. 50 μL of the prepared DMB was added to 50 μL of this enzyme reaction solution and reacted at 80°C for 60 minutes. The absorption spectrum of each DMB-treated sample was measured using a plate reader. The amount of substrate remaining was quantified from the decrease in absorbance (310, 330, or 350 nm), and the enzyme activity (product μM / enzyme μM·reaction time min) was evaluated. As a result of evaluating the various α-keto acid decarboxylases, it was confirmed that LlKdcA had the highest function as a template enzyme (Figure 9).
[0083] Qualitative analysis of α-keto acid decarboxylase products (ABAO assay) 50 μL of the prepared ABAO was added to 50 μL of the above enzyme reaction solution and reacted at room temperature for 10 minutes. The molecular ions and fragment ions of each ABAO-treated sample were detected using LC-MS / MS (ionization method: ESI, scanning method: MRM(+)). As a result, the molecular ions [M+H] + From characteristic fragment ions [M-H2NO] + (m / z 130+R) and [M-OR] + The presence of (m / z 146) was confirmed in each sample. A sample without α-keto acid was used as a negative control.
[0084] In vivo synthesis of tetrahydroisoquinolines (THIQ1-24) Plasmid DNA was obtained by incorporating the above-mentioned LlKdcA gene and the norcocrawlin synthase gene (UniProt ID: A2A1A0, derived from Coptis japonica, a synthetic gene optimized for E. coli was prepared) into the pCDFDuet-1 vector (Figure 10). After introducing this plasmid DNA into E. coli BL21 (DE3), pre-culture was performed in LB medium. This inoculum was inoculated into TB medium and cultured for 5 hours at 37°C and 700 rpm. Then, 1 mM α-keto acid, 1 mM dopamine, 10 mM ascorbic acid, and 300 μM IPTG were added to the culture medium, and cultured for 20 hours at 25°C and 700 rpm. The obtained culture supernatant was diluted 100-fold with MilliQ, and the molecular ions and fragment ions of each tetrahydroisoquinoline were detected using LC-MS / MS (ionization method: ESI, scanning method: MRM(+)). As a result, the molecular ions [M+H] of THIQ1-10 and 17-24, excluding THIQ11-16, + From characteristic fragment ions [M-H3N] + (m / z 148+R) and [MR] + The detection of (m / z 164) was confirmed for each (Figures 11-14). A sample without α-keto acid was used as a negative control. The structures of THIQ1-24 are shown in Figure 15. A literature search revealed that THIQ1 (salsolinol) and THIQ19 (norcoclaurine) are natural products, THIQ2-10, 17, 18, 20, 21, 23, and 24 are non-natural products, and THIQ22 is a novel compound.
[0085] (Note) As described above, while the present disclosure has been illustrated using preferred embodiments thereof, it is understood that the scope of the present disclosure should be interpreted solely by the claims. Patents, patent applications and other documents cited herein should be incorporated by reference in the same way that their contents are specifically described herein. This application claims priority over Japanese Patent Application No. 2022-57114, filed with the Japan Patent Office on March 30, 2022, the contents of which are incorporated by reference in the same way that their entirety constitutes the content of this application. [Industrial applicability]
[0086] This disclosure provides enzymes that have a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, as well as an enzyme library of such enzymes. By using such enzymes, it is possible to efficiently design a target biosynthetic pathway and easily produce the target metabolite, and thus applications are expected in industrial fields such as material production.
Claims
1. An enzyme having a catalytic efficiency above a certain level and substrate specificity above a certain level.
2. The enzyme according to claim 1, wherein the enzyme has a catalytic efficiency equal to or greater than the standard catalytic efficiency of the enzyme.
3. The enzyme according to claim 1 or 2, wherein the enzyme is evaluated by a substrate coverage index, the substrate coverage index being the sum of the catalytic efficiencies of each substrate selected from a basic substrate set or their common logarithms.
4. The enzyme according to claim 3, wherein the substrate coverage index of a plurality of enzyme species having the same reaction specificity is within the top approximately 20%.
5. The enzyme according to claim 1 or 2, wherein the enzyme has substrate specificity greater than or equal to the reference substrate specificity of the enzyme.
6. The enzyme according to claim 1 or 2, wherein the enzyme has about three or more substrate specificities.
7. A composition comprising the enzyme described in claim 1 for use in systematic synthesis of substances.
8. A method for selecting a predetermined enzyme as a template enzyme from amino acid sequence information only, nucleic acid sequence information only, or a combination thereof, a) A method comprising the step of associating the amino acid sequence, nucleic acid sequence, or combination thereof of an enzyme candidate for the template enzyme with an enzyme activity base unit.
9. The method according to claim 8, wherein the enzyme activity base unit directly or indirectly specifies catalytic efficiency and / or substrate specificity.
10. The method according to claim 8 or 9, wherein the basic unit of enzyme activity is identified by an element selected from the group consisting of reaction specificity, coenzyme, substrate species, reaction temperature, reaction pH, molecular crowding during the reaction, and reaction pressure.
11. The method according to claim 8 or 9, further comprising the step of modifying the enzyme activity base unit to an activity not previously known.
12. An enzyme library comprising an enzyme having a predetermined or higher catalytic efficiency and a predetermined or higher substrate specificity, wherein the enzyme library comprises two or more different enzymes.
13. The enzyme library according to claim 12, wherein the enzyme has a catalytic efficiency equal to or greater than the standard catalytic efficiency of the enzyme.
14. The enzyme library according to claim 12 or 13, wherein the enzyme has substrate specificity greater than or equal to the reference substrate specificity of the enzyme.
15. The enzyme library according to claim 12 or 13, comprising three or more different enzymes.
16. The enzyme library according to claim 12 or 13, wherein the enzyme comprises oxidoreductase, transferase, hydrolase, lyase, isomerase, ligase, and translocase.
17. The enzyme library according to claim 12 or 13, wherein the enzyme comprises alcohol dehydrogenase (ADH), decarboxylase, dehydratase, oxygenase, dehydrogenase, oxidase, aminotransferase, transaldolase, esterase, peptidase, glycosidase, carboxylyase, hydrolyase, racemase, mutase, intramolecular lyase, carboxylase, amino acid transporter, sugar transporter, organic acid transporter, hydrogenase, halogenase, nitrogenase, methyltransferase, phosphorylase, phosphatase, kinase, sulfatase, nuclease, ATPase, and metal transporter.
18. An enzyme library according to claim 12 or 13 for use in systematic synthesis of substances.
19. A method for providing an enzyme or combination of enzymes for producing a desired metabolite, a) A step of identifying the raw materials and necessary enzymes for the metabolites, b) A step of selecting the enzyme from the enzyme library. Methods that include...
20. The method according to claim 19, wherein the enzyme library is the enzyme library described in claim 12 or 13.
21. A method for producing a desired metabolite, a) A step of identifying the raw materials and necessary enzymes for the metabolites, b) A step of selecting the enzyme from the enzyme library, c) A step of placing the selected enzyme and the raw materials under conditions that allow the reaction for the production of the metabolites to proceed. Methods that include...
22. The method according to claim 21, wherein the enzyme library is the enzyme library described in claim 12 or 13.
23. A method for selecting a template enzyme based solely on sequence information, A step of reducing the sequence information of multiple candidate enzymes for the template enzyme to a number of dimensions less than the number of dimensions of the basic unit representing the sequence information, A process of classifying each enzyme based on component unit information corresponding to the reduced number of dimensions, The process of selecting an enzyme classified into a specific group as a template enzyme, Methods that include...
24. The method according to claim 23, wherein the reduction includes vectorizing the sequence information using a one-hot vector transformation table.
25. The method according to claim 23 or 24, wherein the classification includes principal component analysis (PCA).
26. The method according to claim 23 or 24, further comprising the step of measuring the activity of enzymes classified into the specific group.
27. The method according to claim 23 or 24, wherein the selection step involves selecting an enzyme located near the median as a template enzyme.
28. A method for preparing a template enzyme library, A step of selecting three or more template enzymes from a plurality of candidate enzymes for template enzymes by the method of claim 23 or 24, A step of providing an enzyme library containing three or more of the above-mentioned template enzymes as a template enzyme library. Methods that include...
29. The enzyme according to claim 1 or 2, wherein the enzyme is selected by the method described in claim 23 or 24.
30. The composition according to claim 7, wherein the enzyme is selected by the method described in claim 23 or 24.
31. The method according to claim 8 or 9, wherein the enzyme is selected by the method according to claim 23 or 24.
32. The enzyme library according to claim 12 or 13, wherein the enzyme is selected by the method described in claim 23 or 24.