Enzyme function prediction method based on structure perception and semantic features
By constructing the contact map of the enzyme and fusion of cross-modal features, the problems of unstable sequence coding and lack of structural data in enzyme function prediction are solved, and high-precision enzyme function and pH prediction are achieved.
Patent Information
- Application Number
- CN202510456516.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The existing enzyme function prediction methods lack stable and efficient sequence coding methods, rely on traditional manual feature engineering to calculate high cost, and the prediction accuracy is low when structural data is lacking. The performance of deep learning models is limited when the enzyme classification level increases, and the interaction relationship between feature perspectives is not fully explored.
The ESMFold model is used to predict the contact graph of enzymes to construct a structural contact graph matrix, combined with the ESM2 and ProtT5 models to extract the semantic features of the enzyme, and information fusion is carried out through graph neural networks and cross-modal attention mechanisms, and initial residual connections and identity mapping are introduced to improve model performance.
It significantly improves the accuracy and robustness of enzyme function prediction, can effectively process complex structural data, and improves the predictive ability of enzyme function and active pH.
Smart Images

Figure CN120340592A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics and relates to a method for predicting enzyme function based on structure perception and semantic features. Background Art
[0002] Bioinformatics integrates biology, computer science, and statistics, focusing on the processing and analysis of biological data to solve information processing problems. Through the application of algorithms, statistical models, and databases, this discipline can efficiently analyze massive biological data, reveal the complex mechanisms of the life system, and promote the development of biological research, biotechnology, and medicine. Proteomics is an important branch of it, which studies the structure, function, and interactions of proteins and predicts their properties. Enzymes, as special proteins, can efficiently catalyze biochemical reactions under appropriate conditions, promoting cell growth and metabolism. Metabolism is the core of life activities, and enzymes play a crucial role in it, ensuring the homeostasis of cell functions. With their excellent catalytic ability, enzymes have broad application prospects in fields such as pharmaceutical research and development, textile industry, food processing, and industrial manufacturing.
[0003] In the pharmaceutical field, enzymes play an important role in the development of new drugs and the treatment of diseases. For example, hyaluronidase is widely used in ophthalmic surgery and the treatment of arthritis. It can degrade hyaluronic acid, increase tissue permeability, and thus improve the drug absorption efficiency. In ophthalmic surgery, this enzyme can be used to assist drug diffusion, make local anesthetics take effect faster, and reduce surgical trauma. In addition, thrombolytic enzymes (such as tissue-type plasminogen activator, tPA) are crucial in the treatment of acute stroke and myocardial infarction. This enzyme can activate the fibrinolytic system, promote thrombolysis, restore blood vessel patency, and reduce the risk of disability and death caused by thrombus obstruction in patients. At the same time, for some genetic enzyme deficiencies, such as Gaucher's disease, patients can supplement β-glucocerebrosidase through enzyme replacement therapy to improve abnormal lipid metabolism in the body, thereby alleviating disease symptoms and improving the quality of life.
[0004] In the textile industry, the application of enzymes is also crucial, running through multiple links such as fiber processing, fabric modification, and final product optimization, significantly improving production efficiency and product quality. For example, cellulase plays a key role in the bio-polishing treatment of denim. It can decompose the fine fibers on the fabric surface, make the cloth softer, and give it a more natural washed effect, replacing traditional chemical methods and reducing environmental pollution. In addition, laccase also has a wide range of applications in the fabric bleaching process. This enzyme can efficiently degrade dyes, improve dyeing uniformity, and at the same time reduce the dependence on chemical reagents during the bleaching process, making textile production more environmentally friendly.
[0005] Enzyme function annotation is one of the important tasks of enzymomics and proteomics, providing a basis for the study of enzyme-catalyzed reactions. According to the type of catalytic reaction, enzymes can be divided into seven categories: oxidoreductases, transferases, hydrolases, lyases, isomerases, synthetases, and transposases. At present, enzyme function is mainly classified by EC number, which is assigned by the Enzyme Commission and contains four digital codes. The core task of enzyme function annotation is to match the correct EC number for a specific protein sequence.
[0006] In the study of enzyme function prediction, although biochemical experimental verification is still regarded as the most reliable means, with the rapid development of DNA sequencing technology, high-throughput sequencing has greatly accelerated the discovery of protein sequences. However, compared with the scale of data growth, the progress of functional annotation has lagged behind significantly. In recent years, the widespread application of metagenomics in enzymology research has led to the accumulation of a large number of different types of enzyme data resources in enzyme and protein databases, including tertiary structure information, evolutionary relationships and gene expression patterns determined by biochemical experiments. Although enzyme function prediction methods based on these biological characteristics have made important breakthroughs, their application scope is still mainly limited to specific enzymes for which complete data can be obtained. In the UniProt protein database, there are currently about 240 million protein sequence records, however, only 0.23% of proteins have additional information other than the sequence itself, which makes enzymes that rely only on sequence data and lack prior knowledge face great challenges in function prediction.
[0007] To solve this problem, improve the accuracy of enzyme function prediction and reduce the cost, researchers have developed a variety of enzyme function annotation tools. Although the current prediction methods based on intelligent computing have made important progress, there are still many challenges that need further optimization and breakthroughs.
[0008] Challenge 1: Lack of stable and efficient enzyme sequence encoding methods. Since enzyme sequences cannot be directly input into the prediction model, they must first undergo a specific encoding conversion. However, at this stage, there is still a lack of a unified and efficient sequence embedding method, which leads researchers to still rely on traditional manual feature engineering, such as one-hot encoding, functional domain encoding, and evolutionary information encoding. These methods are not only computationally expensive, but also have certain limitations when dealing with newly discovered enzyme sequences.
[0009] Challenge 2: The function of an enzyme not only depends on its amino acid sequence, but its three-dimensional structure also plays an important role. Although high-throughput sequencing technology has significantly increased the speed of obtaining enzyme sequences, the analysis of enzyme three-dimensional structures is still extremely difficult. Traditional structural analysis technology, due to its high cost and long experimental cycle, limits large-scale enzyme structure research. Therefore, in the absence of structural data, using only sequence information for effective structural perception to improve prediction accuracy has become a research problem that needs to be broken through.
[0010] Challenge 3: There is still much room for improvement in the accuracy of existing enzyme function prediction models. As the enzyme classification hierarchy increases, the data distribution among enzyme classes is severely imbalanced. Especially when predicting the fourth-level EC numbers, the performance of deep learning models is limited. In addition, current prediction methods mostly adopt simple feature splicing or weighted fusion strategies, failing to fully explore and utilize the interaction relationships between different feature perspectives, thus limiting the overall prediction ability of the models.
[0011] In summary, enzyme function prediction still faces many challenges. How to improve the accuracy and robustness of prediction remains an important research direction in the current bioinformatics field. Summary of the Invention
[0012] To solve the above technical problems and fully explore and rationally utilize the features at the enzyme sequence level for function prediction has become an important direction to improve the ability of enzyme prediction without prior knowledge. The present invention proposes an enzyme function prediction method based on structure awareness and semantic features.
[0013] To overcome the limitations of the prior art, the present invention proposes an enzyme function prediction method based on the fusion of structure awareness and semantic features (SSEFNET). This method uses the ESMFold model to predict the contact map of enzymes and constructs a structural contact map matrix based on the Euclidean distance between amino acids. This matrix is used as an adjacency matrix to construct an undirected graph of enzymes, thereby providing structure awareness information for function prediction. In addition, the present invention uses the ESM2 model to generate feature representations for each amino acid, combines the structural information and sequence information, and constructs a node feature matrix to further enhance the structure awareness ability of the model. To improve the performance of the graph neural network, the present invention proposes an initial residual connection and identity mapping method to alleviate the over-smoothing problem and enhance the model's ability to capture high-order neighborhood information, especially showing superiority when dealing with complex structure data. In addition to structural information, the sequence information of enzymes is also crucial. Therefore, the present invention introduces the ProtT5 model to extract the semantic features of enzyme sequences, and constructs a more comprehensive and accurate enzyme function prediction method by fusing the structural information predicted by ESMFold and the sequence information extracted by ProtT5.
[0014] The technical solution of the present invention is as follows:
[0015] The enzyme function prediction method based on structure awareness and semantic features includes five steps: initial multi-view feature extraction, key structure awareness, key semantic information extraction, multi-view information fusion, and function classification. The steps are as follows:
[0016] Step 1: Input the enzyme sequences into the protein pre-trained models ESM2, ProtT5, and the ESMFold protein structure prediction model respectively to extract multi-view feature information. Among them, ESM2 and ProtT5 extract the ESM semantic feature matrix and the Prt semantic feature matrix respectively to characterize the sequence information of the enzyme, while ESMFold predicts the three-dimensional structure of the enzyme and generates a PDB file.
[0017] Step 2: Input the ESM semantic feature matrix and the PDB file obtained in Step 1 into the graph structure constructor of the key structure perception module to construct the graph structure of the enzyme. Subsequently, learn and extract the key structure information of the enzyme through the key structure perception mechanism.
[0018] Step 3: Input the Prt semantic feature matrix obtained in Step 1 into the key semantic information module to learn the key semantic information of the enzyme.
[0019] Step 4: Input the key structure information and the key semantic information obtained in Steps 2 and 3 into the multi-view information fusion module to obtain the multi-view fusion features.
[0020] Step 5: Input the multi-view fusion features obtained in Step 4 into the function classifier to train and save the enzyme function prediction model.
[0021] Step 6: Input the enzyme sequence to be predicted into the enzyme function prediction model saved in Step 5 to predict the function of the enzyme and map it to the corresponding EC number.
[0022] The specific implementation process of Step 1 is as follows: Input the enzyme sequence into the protein large language model ESM2 to extract the ESM semantic information of the enzyme sequence, and each sequence is encoded as a tensor of Length*1280 dimensions. Input the enzyme sequence into the protein large language model ProtT5 to extract the Prt semantic information of the enzyme sequence, and each sequence is encoded as a tensor of 1*1024 dimensions. Input the enzyme sequence into the protein three-dimensional structure prediction model ESMFold to extract the three-dimensional structure of the enzyme, and each sequence is encoded as a PDB file.
[0023] The specific implementation process of Step 2 is as follows: The key structure perception module consists of a graph structure constructor (GraphFeaturizer) and a key structure feature learning mechanism (KSA) to obtain the key structure feature Z. keySince the PDB file cannot be directly input into the neural network for training, based on this module, the present invention proposes a graph structure construction and feature extraction method to extract key structural information and improve the accuracy and robustness of the model. Graph Featurizer constructs a graph structure using the PDB file generated in Step 1 and the ESM semantic information. The PDB file provides a contact map, which is used to represent the edge information in the graph structure, and the ESM features are used as node features. The contact map is a two-dimensional matrix that describes the spatial contact relationship between amino acid residues in a protein molecule. The matrix element represents the Euclidean distance between residue pairs. If it is less than a set threshold (such as ), the corresponding position is marked as 1, indicating that there is contact between the two; otherwise, it is marked as 0. This representation method presents the complex three-dimensional structural information of the protein in the form of a simple two-dimensional matrix, providing effective support for protein folding, interaction, and function prediction. The specific formula is as follows:
[0024]
[0025] After obtaining the complete graph structure through Graph Featurizer, a 5-layer graph convolutional neural network is used to learn key structural features. To alleviate the over-smoothing problem that may occur in deep networks with multiple layers of GCNs, the present invention introduces an initial residual connection and an identity mapping mechanism to enhance the model's ability to capture high-order neighborhood information, and controls the contribution of the identity mapping through a hyperparameter scaling factor. The updated formula of the improved GCN is as follows:
[0026]
[0027] where σ(·) is the Swish activation function, is the adjacency matrix plus the self-loop I, is 's degree matrix, γ controls the proportion of the initial residual, and controls the proportion of the identity mapping.
[0028] The specific implementation process of Step 3 is as follows: The key semantic feature information module consists of a self-attention mechanism and multiple layers of convolutions with residual connections to obtain the key semantic feature A key . Denote the Prt semantic feature matrix obtained in Step 1 as Prt = [p1, p2,..., p 1024 , where p i represents the i-th extracted feature. Next, use the self-attention mechanism to weight and screen these features to automatically identify and extract the features that are most critical to the enzyme function. The core of the self-attention mechanism is to calculate the correlation matrix A prt, and its calculation formula is:
[0029]
[0030] Among them, Q Prt and K Prt are the query and key matrices of the Prt feature vector respectively, and d k is the dimension of the key vector. In this way, the network can dynamically select the most representative features for subsequent processing according to the correlation between different features. To further improve the performance of the model, the key semantic information extraction module uses a multi-layer convolutional neural network to extract deep features. The specific formula is:
[0031]
[0032] Among them, A Prt l is the output feature map of the l-th layer, W l is the convolutional kernel, * represents the convolution operation, b l is the bias term, and Swish is the activation function. Through layer-by-layer convolution, the model can extract more abstract feature representations. In addition, to avoid the problem of gradient disappearance or gradient explosion in the deep network, the key semantic information extraction module also uses residual connections. The specific formula is:
[0033] A Prt l = A Prt (l-1) + γA Prt 1 (5)
[0034] Among them, γA Prt 1 is the input of the skip connection, ensuring the effective transmission of deep information.
[0035] The specific implementation process of step 4 is as follows: The present invention proposes a multi-view information fusion module to achieve the fusion of semantic information and structural information. The semantic information is based on the amino acid sequence, while the structural information focuses on the three-dimensional structure of the enzyme. These two types of information are based on different theoretical bases. Therefore, directly fusing them by simple summation may lead to information loss or mismatch, thus affecting the prediction accuracy. Therefore, in order to more effectively combine Z key obtained in step 2 and A key obtained in step 3, the multi-view information fusion module uses cross-channel attention to combine the two. The specific formula is as follows:
[0036] M out = Softmax(W Q Z key·W k A key T )·A key (6)
[0037] where W Q and W k are the learned weight matrices that transform Z key and A key into the same dimensional space respectively, and M out is the fused tensor.
[0038] The specific implementation process of step 5 is as follows: The present invention proposes a function predictor. The M out obtained in step 4 is input into a multi-layer perceptron with a hidden layer to construct a classifier. The predicted probability of each label is estimated as follows:
[0039]
[0040] where W represents the weight of the fully connected layer, and f is the non-linear activation function Swish adopted. In addition, the Sigmoid function is used to convert the output value of the network into the probability value of each enzyme function. The cross-entropy loss function has been widely verified to achieve good results in protein function and enzyme function prediction tasks. Considering the class imbalance problem existing in enzyme data, this paper proposes to adopt an improved form of the cross-entropy loss function - Focal loss. Focal loss shows better adaptability in dealing with class imbalance, and its definition is as follows:
[0041]
[0042] where N represents the number of samples, K represents the number of EC number categories, p ic represents the probability that the i-th enzyme sequence has the function of class c, i = 1, 2,..., N, y ic ∈ {0, 1}. γ is a hyperparameter used to adjust the difficulty of samples, and its value is 2. By reducing the loss weight of easy-to-classify samples, γ can effectively alleviate the class imbalance problem and improve the performance of the model in dealing with imbalanced data.
[0043] The specific implementation process of step 6 is as follows: The enzyme sequence to be predicted is input into the enzyme function prediction model saved in step 5 to predict the function of the enzyme and map it to the corresponding EC number.
[0044] The advantages of the present invention include the following points:
[0045] (1) Key Structure Perception Module: Utilize the three-dimensional structure information predicted by ESMFold to construct the contact map of the enzyme molecule, providing an accurate structural representation for the subsequent graph neural network model. In this module, an initial residual connection and an identity mapping mechanism are introduced to effectively alleviate the over-smoothing problem commonly found in traditional graph neural networks, thereby enhancing the model's ability to capture high-order neighbor information. Additionally, this improvement enhances the robustness of the model, enabling it to better retain key information when processing complex enzyme structure data and improving the overall prediction accuracy.
[0046] (2) Key Semantic Information Extraction Module: To further improve the accuracy of enzyme function prediction, the ProtT5 protein large language model is used to semantically encode the amino acid sequence of the enzyme. This module can capture the deep semantic information in the Prt features, extract more expressive key semantic features, provide strong support for enzyme function prediction, and enhance the overall prediction ability of the model.
[0047] (3) Multi-View Information Fusion Module: Based on the cross-channel attention mechanism, establish associations between features of different modalities to achieve the efficient fusion of structural information and semantic features. Through multi-level information interaction, this module further enhances the model's understanding and prediction ability of enzyme functions.
[0048] (4) Comprehensive Experimental Evaluation: In the experimental section, the proposed method was evaluated in detail. The experimental results show that this method can not only accurately predict the function of enzymes but also effectively predict the optimal pH value of enzyme activity. Compared with existing related methods, this method shows significant advantages in terms of accuracy, robustness, and generalization ability, verifying its potential and value in practical applications. Brief Description of the Drawings
[0049] Figure 1 is the algorithm framework diagram of the present invention.
[0050] Figure 2 is the structure diagram of the graph structure constructor of the present invention.
[0051] Figure 3 is the performance comparison diagram of the present invention and other existing enzyme function prediction methods on the New-392 dataset.
[0052] Figure 4 is the performance comparison diagram of the present invention and other existing enzyme function prediction methods on the Price-149 dataset.
[0053] Figure 5 is the performance comparison diagram of the present invention and other existing pH value prediction methods on the optimal pH value test set. Detailed Description of the Invention
[0054] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0055] The present invention realizes an enzyme function prediction method based on structural perception and semantic feature fusion, and its algorithm framework is as Figure 1 shown.
[0056] Figure 1 is the algorithm framework diagram of the present invention. Among them, Step 1: Multi-view feature extraction module; Step 2: Key semantic information extraction module, key structure perception module; Step 4: Multi-view information fusion module; Step 5: Function classifier.
[0057] The invention mainly includes the following five core modules: (1) Initial multi-view feature extraction module: Use ESMFold to predict the three-dimensional structure of the enzyme, and use ESM2 and ProtT5 to extract the semantic information of the enzyme sequence respectively, providing comprehensive basic features for subsequent modeling to ensure that the model can fully capture the structural and sequence characteristics of the enzyme. (2) Key structure perception module: Based on the three-dimensional structure information predicted by ESMFold, construct a contact map, which is used to characterize the spatial topological relationship of the enzyme, and combine the semantic information extracted by ESM2 as node features to construct a graph structure. The graph structure constructor is as Figure 2 shown. Through a multi-layer graph convolutional network (GCN) for feature propagation, aggregating neighborhood information layer by layer, extracting the key structural features of the enzyme from local to global, effectively capturing high-order neighbor relationships, ensuring that the model can deeply understand the spatial topological structure of the enzyme, and at the same time enhancing the perception ability of the key regions that determine the enzyme function, thereby optimizing the accuracy and stability of function prediction. (3) Key semantic information extraction module: Use ProtT5 for deep semantic encoding of the enzyme sequence, and use the self-attention mechanism to identify important features in the sequence. At the same time, combine multi-layer convolution and residual connections to extract deep information to enhance the feature representation ability, ensuring that the model can accurately learn the function-determining information in the enzyme sequence. (4) Multi-view feature fusion module: Based on the cross-channel attention mechanism, establish the association between structural information and sequence semantic features, realize the efficient fusion of multi-modal information, thereby enhancing the model's comprehensive understanding and prediction ability of enzyme functions. (5) Function classifier: Based on the fusion features extracted by the foregoing modules, construct a classification model to accurately predict the function of the enzyme, and ensure that the model has good generalization ability and stability on different types of enzymes. Through the collaborative action of the above five modules, the SSEFNET model framework fully integrates the sequence information and three-dimensional structure information of the enzyme, constructs a function prediction system that can both capture sequence semantic features and understand three-dimensional structure information, and significantly improves the accuracy, robustness and generalization ability of enzyme function prediction.
[0058] Example 1
[0059] The benchmark training set 22W-Enzyme used in this invention comes from Swiss-Prot in Uniprot, which is currently the protein database with the largest amount of information and the widest range, containing a large number of protein sequences. The benchmark dataset 22W-Enzyme is constructed through the following data processing steps: First, all the enzymes recorded in the swiss-prot database before 2022 are obtained. The functional annotations of these enzymes all come from the manual annotations of experimental personnel. Only enzymes with a single function are considered, so enzyme sequences with multiple EC numbers are excluded. To reduce sequence redundancy, the cd-hit tool is used to exclude enzyme sequences with a similarity threshold of 50%. Finally, 227,362 enzyme sequences are obtained as the benchmark dataset 22W-Enzyme, with a total of 4,956 different EC numbers. The specific data distribution is shown in Table 1. To evaluate the performance of the SSEFNET method proposed in this invention, this section compares it with six representative enzyme classification methods, including CLEAN, DeepECtransformer, DeepEC, ECPred, ProteInfer, and GCMEFP. All methods are experimentally verified on the independent test set New-392, and the data distribution of New-392 is shown in Table 2.
[0060] Table 1 Sample distribution of enzyme functions in the benchmark dataset 22W-Enzyme
[0061]
[0062] Table 2 N EW -392 Dataset sample distribution of enzyme functions
[0063]
[0064] The results of this invention and the comparative algorithms on the independent test set New-392 are shown in Table 3. Figure 3 Visualization shows the differences between each method. As shown in the table, this invention has achieved the best results on the benchmark dataset among all inventions.
[0065] Table 3. Performance comparison between this invention and existing methods on New-392
[0066]
[0067] As can be seen from the results in the table, SSEFNET performs outstandingly in the four evaluation metrics of precision, recall, F1-score, and AUROC. In terms of the AUROC metric, SSEFNET scored 0.769, outperforming other methods. GCMEFP (0.761), CLEAN (0.753), and DeepECtransformer (0.572) followed closely, but still lower than SSEFNET. In terms of precision, SSEFNET scored 0.594, far higher than other methods, where GCMEFP and CLEAN were 0.572 and 0.561 respectively, while DeepEC, ECPred, and ProteInfer had lower scores, especially DeepEC and ECPred, which were only 0.142 and 0.117, with poor classification performance. In terms of recall and F1-score performance, SSEFNET continued to lead, with scores of 0.534 and 0.543 respectively, followed by GCMEFP and CLEAN. Overall, SSEFNET performs well in all evaluation metrics, can effectively combine enzyme sequence and structure information, and achieve high-precision enzyme function prediction.
[0068] Example 2
[0069] To evaluate the performance of the proposed method SSEFNET in this chapter from more dimensions, the independent test set Price-149 was used for verification. The six comparison methods selected were the same as those in Section 5.3.3. The specific data are shown in Table 4. Figure 4 .
[0070] According to the experimental results, the performance of each model varies significantly. SSEFNET performs the best among all metrics. Especially in AUROC, its score is 0.749, which is significantly better than other models. Although its precision is 0.584, recall is 0.491, and F1 value is 0.503, it still demonstrates the best overall performance. GCMEFP and CLEAN follow closely, with similar precision, recall, and F1 values, which are 0.512, 0.423, 0.442 and 0.531, 0.434, 0.452 respectively. But in AUROC, GCMEFP scores 0.690 and CLEAN scores 0.717, both of which are better than models such as DeepECtransformer, DeepEC, ECpred, and Proteinfer. DeepECtransformer shows relatively weak performance, with precision of 0.503, recall of 0.429, F1 value of 0.425, and AUROC only 0.566, and its overall effect is not as good as the previous methods. DeepEC performs particularly poorly in precision (0.238) and recall (0.200), with F1 value of 0.211 and AUROC of 0.519, resulting in a relatively poor overall effect. ECpred scores extremely low in recall (0.020) and F1 value (0.038), with precision of 0.333 and AUROC of 0.510, and its performance is almost at the lowest level among all models. Proteinfer also shows poor performance, with precision of 0.243, recall of 0.138, F1 value of 0.166, and AUROC of 0.569, and its overall effect is significantly inferior to other models. In summary, SSEFNET performs the best in all evaluation metrics, while other models have varying degrees of performance deficiencies.
[0071] Table 4. Performance Comparison between the Invention and Existing Methods on Price-149
[0072]
[0073] Example 3
[0074] Since the pH value of the environment where the enzyme is located is crucial for its function, the present invention also predicts the optimal pH value of the enzyme. To train the model, a new dataset was constructed from the Brenda database (released in January 2023), which includes 4,110 proteins with a sequence similarity lower than 25%. The dataset was divided into a training set (Brenda-train, 3,297 enzymes) and an independent test set (Brenda-test, 813 enzymes) according to the collection time, with a ratio of 4:1. SSEFNET obtained a precision of 0.851, a recall of 0.868, an F1 of 0.849, and an AUPR of 0.9321. In contrast, the two latest methods, EpHod and EpHod_SVR, performed poorly, and the precision, recall, and F1 value of SSEFNET exceeded those of the second-best method (EpHod) by 0.6%, 13%, and 7%, respectively. The performance comparison graph is as shown in Figure 5 shown.
Claims
1. A method for predicting enzyme function based on structure perception and semantic features, characterized in that The steps are as follows: Step 1: Input the enzyme sequences into the protein pre-trained models ESM2, ProtT5, and the ESMFold protein structure prediction model respectively to extract multi-view feature information; among them, ESM2 and ProtT5 extract the ESM semantic feature matrix and Prt semantic feature matrix respectively to characterize the sequence information of the enzyme, while ESMFold predicts the three-dimensional structure of the enzyme and generates a PDB file; Step 2: Input the ESM semantic feature matrix and PDB file obtained in Step 1 into the graph structure constructor of the key structure perception module to construct the graph structure of the enzyme; subsequently, learn and extract the key structure information of the enzyme through the key structure perception mechanism; Step 3: Input the Prt semantic feature matrix obtained in Step 1 into the key semantic information module to learn the key semantic information of the enzyme; Step 4: Input the key structure information and key semantic information obtained in Steps 2 and 3 into the multi-view information fusion module to obtain multi-view fusion features; Step 5: Input the multi-view fusion features obtained in Step 4 into the function classifier to train and save the enzyme function prediction model; Step 6: Input the enzyme sequence to be predicted into the enzyme function prediction model saved in Step 5 to predict the function of the enzyme and map it to the corresponding EC number.
2. The method for predicting enzyme function based on structure perception and semantic features according to claim 1, wherein The implementation process of Step 1 is as follows: Input the enzyme sequence into the protein large language model ESM2 to extract the ESM semantic information of the enzyme sequence, and each sequence is encoded as a tensor of Length*1280 dimensions; input the enzyme sequence into the protein large language model ProtT5 to extract the Prt semantic information of the enzyme sequence, and each sequence is encoded as a tensor of 1*1024 dimensions; input the enzyme sequence into the protein three-dimensional structure prediction model ESMFold to extract the three-dimensional structure of the enzyme, and each sequence is encoded as a PDB file.
3. The enzyme function prediction method based on structure perception and semantic features according to claim 1, wherein The implementation process of Step 2 is as follows: The key structure perception module consists of a graph featurizer and a key structure feature learning mechanism (KSA) to obtain the key structure feature Z key ; Since PDB files cannot be directly input into a neural network for training, a graph structure construction and feature extraction method is proposed based on this module to extract key structure information and improve the accuracy and robustness of the model; The Graph Featurizer constructs a graph structure using the PDB file generated in step 1 and the ESM semantic information, where the PDB file provides a contact map to represent the edge information in the graph structure, and the ESM features are used as node features; The contact map is a two-dimensional matrix that describes the spatial contact relationship between amino acid residues in a protein molecule; the matrix element represents the Euclidean distance between residue pairs, and if it is less than the set threshold, the corresponding position is marked as 1, indicating that there is contact between the two; otherwise, it is marked as 0, and the formula is as shown in Equation (1): After obtaining the complete graph structure through Graph Featurizer, use a 5-layer graph convolutional neural network for key structure feature learning; introduce the initial residual connection and identity mapping mechanism, and the updated formula of the improved GCN is as shown in Equation (2): where, σ(!) is the Swish activation function, is the adjacency matrix plus the self-loop I, is the degree matrix of, γ controls the proportion of the initial residual, controls the proportion of the identity mapping.
4. The method for predicting enzyme function based on structure perception and semantic features according to claim 1, wherein The implementation process of Step 3 is as follows: The key semantic feature information module is composed of a self-attention mechanism and multiple layers of convolutions with residual connections, and obtains the key semantic feature A key ; Denote the Prt semantic feature matrix obtained in step 1 as Prt = [p1, p2, …, p 1024 , where p i represents the i-th extracted feature; Next, use the self-attention mechanism (Self-Attention Mechanism) to weight and screen these features to automatically identify and extract the features that are most critical to the enzyme function; The core of the self-attention mechanism is to calculate the correlation matrix A prt of the input features, and its calculation formula is: Among them, Q Prt and K Prt are the query and key matrices of the Prt feature vector respectively, and d k is the dimension of the key vector; the key semantic information extraction module uses a multi-layer convolutional neural network to extract deep features; the specific formula is shown in Equation (4): A Prt l = Swish(W l *A Prt (l-1) +b l ) (4) Among them, A Prt l is the output feature map of the l-th layer, W l is the convolution kernel, * represents the convolution operation, b l is the bias term, and Swish is the activation function; through layer-by-layer convolution, the model can extract more abstract feature representations; the key semantic information extraction module also adopts residual connections; the specific formula is: A Prt l = A Prt (l-1) + γA Prt 1 (5) where γA Prt 1 is the input of the skip connection, ensuring the effective transmission of deep information.
5. The method for predicting enzyme function based on structure perception and semantic features according to claim 1, characterized in that The implementation process of Step 4 is as follows: The multi-view information fusion module realizes the fusion of semantic information and structural information; the semantic information is based on the amino acid sequence, while the structural information focuses on the three-dimensional structure of the enzyme; combining Z obtained in step 2 key and A obtained in step 3 key , the multi-view information fusion module uses cross-channel attention to combine the two; the specific formula is shown in Equation (6): M out = Softmax(W Q Z key ·W k A key T )·A key (6) Where W Q and W k are the learned weight matrices that transform Z key and A key to the same dimensional space, and M out is the fused tensor.
6. The method for predicting enzyme function based on structure perception and semantic features according to claim 1, wherein The implementation process of Step 5 is as follows: Input the M obtained in step 4 out into a multi-layer perceptron with a hidden layer to construct a classifier; the predicted probability estimate for each label is shown in Equation (7): Among them, W represents the weight of the fully connected layer, and f is the non-linear activation function Swish adopted; adopt the improved form of the cross-entropy loss function - Focal loss; Focal loss shows better adaptability in dealing with class imbalance, and its definition is as follows: Among them, N represents the number of samples, K represents the number of EC number categories, and p ic represents the probability that the i-th enzyme sequence has the function of category c, i = 1, 2, ..., N, y ic ∈ {0, 1}; γ is a hyperparameter used to adjust the difficulty of samples, and its value is 2; by reducing the loss weight of easily classified samples, γ can effectively alleviate the class imbalance problem and improve the performance of the model when dealing with imbalanced data.
7. The method for predicting enzyme function based on structure perception and semantic features according to claim 1, wherein The implementation process of Step 6 is as follows: Input the enzyme sequence to be predicted into the enzyme function prediction model saved in Step 5 to predict the function of the enzyme and map it to the corresponding EC number.
Citation Information
Cited By
Protein interaction site prediction method and system
CN120877854A