Method for predicting enzyme catalysis efficiency and evaluating prediction credibility
By introducing a deep learning model of evidence of bilinear attention mechanism in the enzyme catalytic efficiency prediction model, the problems of insufficient understanding of data limitation and enzyme catalytic mechanism in the existing technology are solved, and more accurate and reliable prediction of enzyme catalytic efficiency is achieved, and the efficiency of enzyme directed evolution is improved.
Patent Information
- Application Number
- CN202510244085.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art present challenges in predicting enzyme catalytic efficiency and evaluating predictive credibility, including limitations of the dataset and the lack of understanding of key local interactions of enzyme catalytic substrates by the model.
An evidence-based deep learning model using a bilinear attention mechanism is used to predict the kcat/Km value by learning pairwise local interactions between enzymes and substrates, and provide uncertainty in the predicted value to improve the credibility and interpretability of the model.
The predictive performance of enzyme catalytic efficiency is improved, the interpretability of the model is enhanced, the false positives of experimental verification are reduced, and the efficiency of directed evolution of high-active enzymes is improved.
Smart Images

Figure CN120148686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting enzyme catalytic efficiency and evaluating the credibility of the prediction, belonging to the fields of bioinformatics and theoretical calculation of enzyme catalysis. An evidence deep learning method is adopted to guide the screening of highly active enzymes and promote the development of enzyme directed evolution. Background Art
[0002] Enzyme catalytic efficiency (k cat / K m ) is a key parameter for guiding the screening of highly active enzymes. At present, the measurement of enzyme kinetic parameters mainly relies on experiments, but the process is time-consuming, expensive and requires a large amount of human input, resulting in a very small database of experimentally measured kinetic parameter values. To bridge the huge gap between the enzyme sequence and the experimentally measured k cat / K m data volume, researchers have attempted to develop machine learning and deep learning-based methods to accelerate the process of obtaining enzyme kinetic parameters. For example, Heckmann et al. used an integrated machine learning model of linear elastic network, decision tree and complex neural network to learn various input features such as metabolic network, protein structure, and substance concentration, and achieved rapid prediction of k cat in the Escherichia coli metabolic network, demonstrating the feasibility of predicting k cat by machine learning models. However, the input features of this model are too complex and difficult to be used for predicting k cat of other species. Kroll et al. proposed two organism-independent machine learning and deep learning models for rapidly inferring the Km and k cat values of wild-type enzymes across species. Li et al. directly used recurrent neural network and graph neural network to learn the information of amino acid sequence and molecular graph of substrate respectively, and achieved high-throughput prediction of k cat . In addition, Yu et al. developed a unified framework based on pre-trained language model, and achieved high-throughput prediction of k cat , K m and k cat / K m of three enzyme kinetic parameters.
[0003] The above deep learning technologies have shown their fast and accurate prediction of k cat / K mHowever, there are still two challenges. The first challenge is that the dataset of enzyme activity parameters used in current work mainly comes from the BRENDA and SABIO-RK databases. Not all entries in these databases have complete annotations. After strict data cleaning, the amount of data available for model training is very limited. The lack of data for different types of enzymes makes it difficult for machine learning models to accurately fit the distribution of this type of data. Therefore, it is very difficult for the model to make accurate predictions when dealing with new data beyond the distribution range. The second challenge is that the essence of enzyme catalysis is the interaction between the key residues in the enzyme active pocket and the important substructures of the substrate molecules. The current enzyme-catalyzed kinetic parameter prediction model globally represents the enzymatic reaction by combining or adding the feature vectors of the enzyme and the substrate, and does not explicitly learn the key local interactions that perform the catalytic function. These limitations restrict the prediction accuracy and generalization of the model, and thus hinder its practical application in enzyme directed evolution. Summary of the Invention
[0004] (I) Objectives of the Invention
[0005] The objective of the present invention is to provide a method for predicting enzyme catalytic efficiency and evaluating the credibility of the prediction, and constructing an evidence deep learning model with a bilinear attention mechanism for predicting k cat / K m value, and at the same time providing the uncertainty of the predicted value to increase the credibility of the deep learning model for predicting k cat / K m , reduce the false positives of experimental verification, and improve the efficiency of directed evolution of highly active enzymes. The bilinear attention mechanism is adopted to focus on learning and visualizing the key local interactions between the enzyme and the substrate, increasing the interpretability of the deep learning model for predicting k cat / K m , and helping to understand the mechanism analysis of enzyme-catalyzed substrates.
[0006] (II) Technical Solutions
[0007] To achieve the above objectives and solve the above technical problems, the following solutions are adopted in the present invention:
[0008] A method for predicting enzyme catalytic efficiency and evaluating the credibility of the prediction, specifically including the following steps:
[0009] Step 1: Construct an evidence deep learning model with a bilinear attention mechanism;
[0010] The evidence deep learning model with a bilinear attention mechanism includes a bilinear attention module A and an evidence deep learning module B;
[0011] Among them, the bilinear attention module A is described as follows:
[0012] 1-1-1 is the initial feature vector H of the substrated , 1-2-1 is H d The product of the transpose and the learnable weight matrix represented by the substrate;
[0013] 1-3-1 is the Hadamard product of 1-2-1 and the learnable weight matrix;
[0014] 1-1-2 is the initial eigenvector H of the enzyme p , 1-2-2 is the product of the transpose of the learnable weight matrix represented by the enzyme and H p ;
[0015] 2-2 is the characteristic matrix of the interaction strength of the substrate-enzyme substructure pair obtained by dot-multiplying 1-3-1 and 1-2-2;
[0016] 2-1 is the product of the transpose of the learnable weight matrix represented by the substrate and 1-1-1, and 2-3 is the product of the learnable weight matrix represented by the enzyme and the transpose of 1-1-2;
[0017] 3 is the joint representation matrix of the substrate-enzyme interaction obtained by dot-multiplying 2-1, 2-2 and 2-3;
[0018] 4 is the joint representation matrix of the substrate-enzyme interaction after the SUM pooling operation on 3;
[0019] Evidence deep learning module B: The joint representation matrix of the substrate-enzyme interaction after the SUM pooling operation is used as the input of evidence deep learning module B. After learning the normal-inverse gamma distribution through the evidence deep learning layer, four parameters γ, v, α and β are output, which are 5-1, 5-1, 5-3 and 5-4 respectively;
[0020] Based on the four parameters β, v, α and β, the predicted value 6-1 and the uncertainty 6-2 of the predicted value are output;
[0021] Step 2: Use the bilinear attention network module constructed in Step 1 to capture the pairwise local interaction between the enzyme and the substrate;
[0022] Use the bilinear attention mechanism to capture the pairwise local interaction between the enzyme and the substrate; The interaction is divided into two parts:
[0023] One part is to use bilinear interaction to capture pairwise attention weights;
[0024] The other part is to use the bilinear pooling layer to extract the joint enzyme-substrate representation;
[0025] The hidden representations of the enzyme and the substrate are respectively and Where M and N represent the number of residues in the enzyme and the number of atoms in the substrate, respectively. These hidden representations are used to construct a bilinear interaction map to obtain a single-headed pairwise interaction moment.
[0026]
[0027] Where is a learnable weight matrix for substrate and enzyme representations, is a learnable weight vector, is a fixed all-ones vector, and ° represents the Hadamard product; the elements in I represent the interaction strengths of the respective substrate-enzyme substructure pairs and are mapped to potential binding sites and molecular substructures;
[0028] Element I i,j can also be written as
[0029]
[0030] Where is the i-th row of H d , is the j-th row of H p , representing the representation of the i-th atom of the substrate and the j-th residue of the enzyme;
[0031] Regarding a bilinear interaction as the first mapped representation and to a common feature space with weight matrices U and V, and then learning an interaction on the Hadamard product and the weight of vector q;
[0032] To obtain the joint representation a bilinear pooling layer is introduced on the interaction map I, and the k-th element of f′ is calculated as
[0033]
[0034] Where U k and V k represent the k-th columns of the weight matrices U and V;
[0035] In addition, by adding SUM pooling on the joint representation vector, a compact feature map is obtained:
[0036] f = Sumpool(f′, s)
[0037] Where the Sumpool(·) function is a one-dimensional non-overlapping sum pooling operation with a stride of s; reducing to
[0038] Expand a single bilinear interaction into a multi - head form by calculating multiple bilinear interaction mappings;
[0039] The final joint representation vector is the sum of the individual heads;
[0040] Step 3: Implement the prediction of k cat / K m value through the evidence layer of evidence - based deep learning, and give a credibility assessment of the predicted value.
[0041] In regression, the present invention gives a dataset of paired training instances where the assumed labels are drawn from an independent and identically - distributed Gaussian distribution θ = {μ, σ 2} with unknown mean and variance. Assuming that the mean comes from a Gaussian distribution and the variance comes from an inverse - gamma distribution, we seek to make probabilistic estimates of the mean and variance. Thus, the joint high - order evidence distribution is represented as a normal - inverse - gamma distribution. Specifically, the normal - inverse - gamma distribution, p(θ|m), which is also the conjugate prior of the Gaussian prior, is parameterized by m = {γ, v, α, β} and is represented as the distribution over θ = {μ, σ 2}. Therefore, in the present invention, the last layer of the evidence layer is modified to output these normal - inverse - gamma hyperparameters. Thus, for each target task, the network has 4 outputs. As Figure 1 shown in B, given the normal - inverse - gamma distribution, the prediction and uncertainty are described according to the distribution moments:
[0042]
[0043] where is the predicted value and Var[μ] represents the uncertainty.
[0044] Use a bi - objective loss consisting of two loss terms to train the evidence model to maximize the model fitness according to the negative log - likelihood and minimize the wrong evidence:
[0045]
[0046] where is the negative log - likelihood and is the evidence regularizer. The regularization coefficient λ controls the strength of the uncertainty relative to the model fitness.
[0047] (III) Effective benefits
[0048] 1. The present invention develops a method for predicting the enzyme catalytic efficiency k cat / K m to evaluate the catalytic efficiency of enzymes. The Pearson correlation coefficient in independent testing on the test set reached 0.778, which is higher than the current state - of - the - art kcat / K m The prediction model has been improved by approximately 3.56%.
[0049] 2. The k developed in the present invention cat / K m The prediction method adopts a bilinear attention mechanism to focus on learning the key local interactions between the enzyme and the substrate. The learning of key local interactions not only improves the prediction performance of k cat / K m but more importantly, visualizes the key residues and substrate atoms during the enzyme-catalyzed substrate process, providing interpretable insights into the enzyme-catalyzed reaction.
[0050] 3. The present invention adopts the method of evidence deep learning for predicting k cat / K m , and evaluates the uncertainty of the predicted value while outputting the k cat / K m value, that is, evaluates the credibility of the prediction. Considering the k cat / K m predicted value and uncertainty to evaluate the catalytic efficiency of the enzyme helps to reduce the false positives in enzyme screening, improve the accuracy of high-activity enzyme screening, reduce the workload of experimental screening, and accelerate the process of directed evolution of high-activity enzymes. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Architecture diagram of an evidence deep learning model with a bilinear attention mechanism;
[0052] Figure 2 Test performance of the method of the present invention on an out-of-domain test set. DETAILED DESCRIPTION OF THE INVENTION
[0053] The present invention will be further explained and described below in conjunction with the drawings and embodiments.
[0054] The present invention proposes an evidence deep learning and bilinear attention network model for k cat / K m prediction. The present invention is a deep learning method that focuses on learning the local interactions between the enzyme and the substrate through a bilinear attention network, and can effectively evaluate the uncertainty of the k cat / K m predicted value through an evidence deep learning layer. The results show that the present invention achieves the current best k cat / K mPredictive performance, and at the same time can provide an uncertainty assessment for the prediction results. This uncertainty information will provide a more confident prediction, which can be used to guide the more accurate screening of highly active target enzymes, thereby accelerating the directed evolution process of highly active enzymes. In addition, the present invention visualizes key residues and substrate atoms, provides interpretable insights into enzyme-substrate interactions, and thus plays a certain auxiliary role in the mechanism study of enzyme-catalyzed substrates.
[0055] A method for predicting enzyme catalytic efficiency and evaluating the credibility of the prediction, specifically comprising the following steps:
[0056] Step 1: Construct an evidence deep learning model with a bilinear attention mechanism
[0057] The architecture of the evidence deep learning model with a bilinear attention mechanism of the present invention is as Figure 1 shown. The evidence deep learning model with a bilinear attention mechanism of the present invention includes a bilinear attention module A and an evidence deep learning module B;
[0058] Among them, the bilinear attention module A is described as follows:
[0059] 1-1-1 is the initial feature vector H of the substrate d , 1-2-1 is the product of the transpose of H d and the learnable weight matrix of the substrate representation;
[0060] 1-3-1 is the Hadamard product of 1-2-1 and the learnable weight matrix;
[0061] 1-1-2 is the initial feature vector H of the enzyme p , 1-2-2 is the product of the transpose of the learnable weight matrix of the enzyme representation and H p ;
[0062] 2-2 is the feature matrix of the interaction strength of the substrate-enzyme substructure pair obtained by dot-multiplying 1-3-1 and 1-2-2;
[0063] 2-1 is the product of the transpose of the learnable weight matrix of the substrate representation and 1-1-1, and 2-3 is the product of the learnable weight matrix of the enzyme representation and the transpose of 1-1-2;
[0064] 3 is the joint characterization matrix of substrate-enzyme interaction obtained by dot-multiplying 2-1, 2-2 and 2-3;
[0065] 4 is the joint characterization matrix of substrate-enzyme interaction after SUM pooling operation on 3;
[0066] Evidence Deep Learning Module B: Evidence Deep Learning Module B takes the joint representation matrix of substrate-enzyme interactions after SUM pooling operation as input, that is, 4 is the input of Evidence Deep Learning Module B. After learning the normal-inverse gamma distribution through the evidence deep learning layer, it outputs four parameters γ, v, α, and β, which are 5-1, 5-1, 5-3, and 5-4 respectively;
[0067] Based on these four parameters, it outputs the predicted value 6-1 and the uncertainty of the predicted value 6-2.
[0068] Step 2: Use the bilinear attention network module to capture the pairwise local interactions between enzymes and substrates
[0069] Use the bilinear attention network module to capture the pairwise local interactions between enzymes and substrates. As Figure 1 shown in A, it consists of two layers: a bilinear interaction graph to capture pairwise attention weights, and a bilinear pooling layer on the interaction graph to extract the joint enzyme-substrate representation.
[0070] The hidden representations of the enzyme and substrate are respectively and where M and N represent the number of residues in the enzyme and the number of atoms in the substrate. These hidden representations are used to construct a bilinear interaction mapping, thereby obtaining a single-headed pairwise interaction moment
[0071]
[0072] where is a learnable weight matrix for substrate and enzyme representations, is a learnable weight vector, is a fixed all-one vector, and ° represents the Hadamard product. The elements in I represent the interaction strengths of the respective substrate-enzyme substructure pairs and are mapped to potential binding sites and molecular substructures.
[0073] To intuitively understand the bilinear interaction, an element I in the above formula i,j can also be written as
[0074]
[0075] where is the i-th row of H d , is the j-th row of H p , representing the representations of the i-th atom of the substrate and the j-th residue of the enzyme. Therefore, a bilinear interaction can be regarded as the first mapped representation and to a common feature space with weight matrices U and V, and then learn an interaction on the Hadamard product and the weights of vector q. In this way, the contribution of pairwise interactions to the prediction results of substructures provides interpretability.
[0076] To obtain the joint representation We introduce a bilinear pooling layer on the interaction mapping I. Specifically, the k-th element of f′ is calculated as
[0077]
[0078] where, U k and V k represent the k-th columns of the weight matrices U and V.
[0079] In addition, by adding SUM pooling on the joint representation vector, a compact feature map is obtained:
[0080] f = Sumpool(f′, s)
[0081] where the Sumpool(·) function is a one-dimensional non-overlapping sum pooling operation with a stride of s. Reduce the dimension to In addition, by calculating multiple bilinear interaction mappings, we can extend the single bilinear interaction to the multi-head form. The final joint representation vector is the sum of individual heads. Since the weight matrices U and V are shared, each additional head only adds a new weight vector q.
[0082] Therefore, using the novel bilinear attention mechanism, the model can explicitly learn the pairwise local interactions between drugs and proteins. To predict k cat / K m , we input the joint representation f into the evidence layer, which is a fully connected classification layer.
[0083] Step 3, through the evidence layer of evidence deep learning, we can achieve the prediction of the k cat / K m value and give a credibility assessment of the predicted value.
[0084] The specific implementation process is as follows:
[0085] In regression, a dataset of pairwise training instances is given where, it is assumed that the labels follow an independent and identically distributed Gaussian distribution θ = {μ, σ 2} Extraction. Assuming that the mean comes from a Gaussian distribution and the variance comes from an inverse gamma distribution, a probabilistic estimate of the mean and variance is sought. Therefore, the joint high-order evidence distribution is represented as a normal-inverse gamma distribution. Specifically, the normal-inverse gamma distribution, p(θ|m), is also the conjugate prior of the Gaussian prior, is parameterized by m = {γ, v, α, β}, and is represented as the distribution at θ = {μ, σ 2} Thus, in the present invention, the last layer of the evidence layer is modified to output these normal-inverse gamma hyperparameters. Therefore, for each target task, the network has 4 outputs. As Figure 1 shown in Module B, given the normal-inverse gamma distribution, the prediction and uncertainty are described according to the distribution moments:
[0086]
[0087] where is the predicted value, and Var[μ] represents the uncertainty.
[0088] A dual-objective loss consisting of two loss terms is used to train the evidence model to maximize the model fitness according to the negative log-likelihood and minimize the false evidence:
[0089]
[0090] where is the negative log-likelihood, is the evidence regularizer. The regularization coefficient λ controls the strength of the uncertainty relative to the model fitness.
[0091] Example 1
[0092] To verify the accuracy of the software of the present invention in practical applications, the present invention organized an out-of-domain test set based on the published literature, with a data volume of 806. The numerical distribution of this test set is significantly different from that of the training set, and each piece of data has no repetition with the training set. The present invention independently tested the accuracy of the software of the present invention on this out-of-domain test set,
[0093] As Figure 2 shown, the abscissa is the true value of k cat / K m in the data set, that is, the experimentally measured k cat / K m , which is transformed by the logarithm to the base 10. The bar graph above is the distribution of the experimental values of k cat / K m . The ordinate is the value of k cat / K m predicted by the present invention, which is also transformed by the logarithm to the base 10. The right surface bar graph is k cat / Km Distribution of predicted values. The Pearson correlation coefficient between the experimental values and the predicted values is shown in the upper left corner, reaching 0.663. The closer this index is to 1, the stronger the positive correlation between the experimental values and the predicted values. The diagonal represents perfect prediction. The closer the predicted points are to the diagonal, the more accurate the prediction. The legend indicating the point density is shown in the lower right corner. The closer the color is to white, the higher the point density.
[0094] From Figure 2 it can be seen that the Pearson correlation coefficient of the independent test of the present invention on the out-of-domain test set reaches 0.663, indicating that the predicted k cat / K m and the experimental k cat / K m values have a good positive correlation. It proves that the present invention has good application potential in actually predicting the k cat / K m value to evaluate the catalytic efficiency of enzymes.
[0095] Example 2
[0096] To verify the guiding role of the deep learning method of the present invention in enzyme directed evolution, further, the present invention collected directed evolution data sets of two sesquiterpene synthases and two monoterpene synthases according to the published literature: (1) Directed evolution data of recombinant epi-zizaene synthase, with a data volume of 24; (2) Directed evolution data of germacrene A synthase, with a data volume of 8; (3) Directed evolution data of myrcene synthase, with a data volume of 8; (4) Directed evolution data of borneol synthase, with a data volume of 44. In this directed evolution data set, the k cat / K m value of the mutant enzyme is compared with that of the wild-type enzyme. If the k cat / K m value of the mutant enzyme is higher than that of the wild-type, the label is marked as 1, otherwise it is marked as 0. By predicting the directed evolution data of the present invention and sorting according to the k cat / K m predicted value and uncertainty, the results with higher k cat / K m predicted value and lower uncertainty are ranked higher. The hit rate of the software prediction is further evaluated according to the sorting results.
[0097] Table 1 shows the hit rates predicted by the deep learning method of the present invention. The recombinant caryophyllene synthase and germacrene A synthase both have relatively high hit rates when ranked in the top 20%, which are better than the hit rates in the top 100%. The hit rate of myrcene synthase reaches 100% when ranked in the top 20%, which is better than the hit rate in the top 100%. However, the hit rate of borneol synthase reaches 44.4% when ranked in the top 20%. Therefore, in practical applications, the present invention considers the results in the top 20% for experimental verification, which will greatly reduce the false positive rate of experimental verification and improve the directed evolution efficiency of highly active enzymes.
[0098] According to the prediction value plus uncertainty ranking method, the predicted hit rates of the present invention for the directed evolution data of two sesquiterpene synthases and monoterpene synthase are shown in Table 1.
[0099] Table 1 Comparison table of predicted hit rates of directed evolution data of two sesquiterpene synthases and monoterpene synthase
[0100]
[0101] The above content is a further detailed description of the present invention in combination with specific embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for predicting enzyme catalytic efficiency and evaluating the reliability of the prediction, characterized in that: The specific steps include: Step 1: Build an evidence-based deep learning model with a bilinear attention mechanism The evidence deep learning model of the bilinear attention mechanism includes a bilinear attention module A and an evidence deep learning module B; The bilinear attention module A is expressed as follows: 1-1-1 is the initial characteristic vector H of the substrate d , 1-2-1 is H d The product of the learnable weight matrix represented by the substrate after transposition; 1-3-1 is the Hadamard product of 1-2-1 and the learnable weight matrix; 1-1-2 is the initial characteristic vector H of the enzyme p , 1-2-2 is the transpose of the learnable weight matrix represented by the enzyme and H p The product of 2-2 is the characteristic matrix of the interaction strength between substrate and enzyme substructure obtained by multiplying 1-3-1 and 1-2-2; 2-1 is the product of the transpose of the learnable weight matrix represented by the substrate and 1-1-1, and 2-3 is the product of the learnable weight matrix represented by the enzyme and the transpose of 1-1-2; 3 is the joint representation matrix of substrate-enzyme interaction obtained by dot-multiplication of 2-1, 2-2 and 2-3; 4 is the joint representation matrix of substrate-enzyme interaction after SUM pooling operation in 3; Evidence deep learning module B: The joint representation matrix of substrate-enzyme interaction after SUM pooling operation is used as the input of evidence deep learning module B. After the evidence deep learning layer learns the normal-inverse gamma distribution, it outputs four parameters γ, v, α and β, which are 5-1, 5-1, 5-3 and 5-4 respectively; Output the predicted value 6-1 and the uncertainty of the predicted value 6-2 based on the four parameters; Step 2: Use the bilinear attention network module constructed in step 1 to capture the pairwise local interactions between enzymes and substrates; An evidence-based deep learning model using a bilinear attention mechanism to capture pairwise local interactions between enzymes and substrates; There are two levels of interaction: One part is to use bilinear interactions to capture pairwise attention weights; The other part is to use bilinear pooling layers to extract joint enzyme-substrate representations; The enzyme and substrate hidden representations are and Where M and N represent the number of residues in the enzyme and the number of atoms in the substrate; These hidden representations are used to construct a bilinear interaction map, which leads to the single-headed pairwise interaction matrix in are learnable weight matrices for substrate and enzyme representation, is a learnable weight vector, is a fixed all-one vector, and ° represents the Hadamard product; the elements in I represent the interaction strength of the respective substrate-enzyme substructure pairs and are mapped to potential binding sites and molecular substructures; Element I i,j It can also be written as in Yes H d The i-th row of Yes H p The jth row of represents the representation of the i-th atom of the substrate and the j-th residue of the enzyme; Consider a bilinear interaction as the first mapping representation and to a common feature space with weight matrices U and V, and then learn an interaction on the weights of the Hadamard product and vector q; To obtain the joint representation A bilinear pooling layer is introduced on the interaction map I, f ′ The kth element of is calculated as Among them, U k and V k represents the k-th column of the weight matrices U and V; In addition, by adding SUM pooling on the joint representation vector, a compact feature map is obtained: f=Sum(f ′ ,s) The Sumpool(·) function is a one-dimensional non-overlapping and pooling operation with a step size of s; Dimensionality reduction to By computing multiple bilinear interaction maps, a single bilinear interaction is extended to a multi-head form; The final joint representation vector is the sum of the individual heads; Step 3: Implement k through the evidence layer of evidence deep learning cat / K m The prediction of the value is given and the credibility of the prediction is evaluated.
2. A method for predicting enzyme catalytic efficiency and evaluating prediction reliability according to claim 1, characterized in that: The specific implementation steps of step 3 are as follows: In regression, given a dataset of paired training examples Here, the labels are assumed to be drawn from an independent and identically distributed Gaussian distribution θ = {μ,σ 2 } Extract; Assuming the mean comes from a Gaussian distribution and the variance comes from an inverse gamma distribution, we seek to make probabilistic estimates of the mean and variance; The joint high-order evidence distribution is represented as a normal-inverse gamma distribution; the normal-inverse gamma distribution, p(θ|m), is the conjugate prior of the Gaussian prior, parameterized by m = {γ, v, α, β} and represented as 2 } distribution; The last layer of the evidence layer outputs normal-inverse gamma hyperparameters. Each target task has 4 output parameters γ, v, α, and β. Given a normal-inverse gamma distribution, the prediction and uncertainty are described according to the distribution moments: in is the predicted value, Var[μ] represents the uncertainty; Use a dual-objective loss consisting of two loss terms Train the evidence model to maximize the model fit according to the negative log-likelihood and minimize the false evidence: in is the negative log-likelihood, is the evidence regularizer, and the regularization coefficient λ controls the strength of the uncertainty relative to the model fit.