Multi-level enzyme function prediction method based on multi-view semantics

CN118888033BActive Publication Date: 2026-09-25JIANGNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410912384.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-09-25
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

[0009]2)其次,不同类别的酶数据分布极为不平衡,这导致现有深度学习模型容易过拟合

Benefits of technology

[0047](1)我们提出了一种融合多视图学习技术的多级酶功能预测方法,该方法能够直接预测最后一级即第四级酶功能,实现非酶的蛋白质序列和酶中不同水平标记的同步预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118888033B_ABST
    Figure CN118888033B_ABST
Patent Text Reader

Abstract

The application belongs to the field of biological information, and particularly relates to a multi-level enzyme function prediction method based on multi-view semantics, which comprises technologies such as deep learning, convolutional neural network and multi-head self-attention mechanism. The method comprises five steps of original semantic feature extraction of a protein sequence, a deep semantic feature module, a broad semantic feature module, a part-of-speech attention enhancer and multi-view adaptive network fusion prediction. The method firstly inputs the protein sequence into a protein pre-training model for preprocessing to obtain original semantic information. Then the original semantic information is input into the deep semantic feature module and the broad semantic feature module proposed in the method to obtain deep semantic features and broad semantic features. The obtained original, deep and broad semantic features are input into the part-of-speech attention enhancer and the multi-view adaptive fusion network of the method to obtain the final enzyme function prediction result. A large number of experiments prove that the application can efficiently and accurately predict enzyme functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics, specifically relating to a multi-level enzyme function prediction method based on multi-view semantics. Background Technology

[0002] Bioinformatics is a new discipline formed by the integration of computer science and life sciences, utilizing computer science tools to analyze, retrieve, and store biological information. As one of the cutting-edge fields in natural and life sciences, bioinformatics reveals the laws governing genetic language and the information structure of the genome. Proteomics, as an important topic within bioinformatics, is playing an increasingly vital role in protein research.

[0003] Enzymes, as a special class of proteins, play an indispensable role in the biological field. As biocatalysts, enzymes accelerate the rate of chemical reactions, enabling organisms to carry out complex metabolic processes under mild conditions. The highly efficient catalytic action of enzymes allows organisms to complete a large number of chemical transformations in a short time, thus maintaining the normal functioning of life. Enzymes participate in many key biological processes within organisms, including food digestion, energy production, signal transduction, cell division, and repair. Furthermore, enzymes also play an important role in drug development, bioengineering, and environmental protection.

[0004] With the deepening understanding of DNA sequence knowledge, an increasing number of new proteins are being discovered, making information mining and analysis of protein sequences crucial. Dysfunction of certain enzymes can lead to serious metabolic diseases; for example, a deficiency in lactate dehydrogenase can cause lactic acidosis due to lactate buildup. In fact, almost all life processes rely on the catalytic action of enzymes. In the human body, thousands of different enzymes exist, distributed in specific locations within different cells, working together to participate in biological processes such as immunity, growth, and metabolism. In-depth research on enzymes not only helps in understanding the complex metabolic networks within organisms but also provides an important foundation for the development of new drugs and biotechnologies. Enzyme research has significant theoretical and applied value, playing a vital role in advancing biological sciences and solving major problems facing humanity. Therefore, research on proteins and enzymes in bioinformatics is of great significance, revealing the laws of genetic language, the information structure of the genome, and the complex metabolic networks within organisms, providing an important foundation for drug development and the advancement of biotechnology.

[0005] Enzyme functional annotation is a crucial component of numerous complex protein annotation tasks, providing a key starting point for generating and testing enzyme-catalyzed reactions. Based on the chemical reactions they catalyze, enzymes can be classified into seven classes: oxidoreductases, transferases, hydrolases, lyases, isomerases, synthases, and transposases. Currently, enzyme function is primarily classified through their EC (Enzyme Code) numbers, which are four-part numerical codes assigned by the enzyme committee. Therefore, the main task of enzyme annotation is to assign an EC number to a given protein sequence.

[0006] In the field of enzyme function prediction, the most rigorous method for enzyme function identification is biological experimental methods. However, these methods require significant time and cost. Thanks to the tremendous advancements in amino acid sequencing technology, it is now possible to rapidly sequence large numbers of enzyme sequences, while the number of enzymes with unknown functions is increasing dramatically. To efficiently and accurately predict enzyme function, computational methods have been introduced into the field of enzyme function prediction to help biologists determine the specific functions of enzymes.

[0007] To achieve rapid and low-cost functional annotation of enzymes, many enzyme function annotation tools have been developed. Although current intelligent prediction methods for enzyme function have made significant progress, these existing methods still have some shortcomings. For example:

[0008] 1) First, enzyme function prediction lacks a well-designed method with stable predictive performance for newly discovered proteins. Specifically, encoding quality significantly impacts the performance of downstream algorithms. However, the current lack of effective and universal protein sequence embedding methods forces researchers to spend considerable time on manual feature engineering to encode sequences, such as genetic information encoding, functional domain encoding, and evolutionary information (PSSM) encoding. Therefore, it is necessary to develop a new, universal sequence embedding method.

[0009] 2) Secondly, the distribution of enzyme data across different categories is extremely unbalanced, making existing deep learning models prone to overfitting. In classifier learning, classifiers are typically trained directly using encoded primary features obtained through feature engineering. This strategy is too coarse and doesn't fully utilize the extracted features.

[0010] 3) Finally, although some methods have improved model performance by leveraging multi-view learning to mine enzyme data, existing multi-view enzyme function prediction models treat each view equally. Secondly, the interaction information hidden in multi-view data is particularly important for enzyme function prediction. However, existing multi-view enzyme function prediction methods ignore the importance of this information.

[0011] Therefore, enzyme function prediction still faces significant challenges. Summary of the Invention

[0012] To address the shortcomings of existing technologies, this invention proposes a multi-level enzyme function prediction method based on multi-view semantics. The innovations of this method are mainly reflected in two aspects: protein sequence feature extraction and model construction. In the feature extraction stage, the original semantic features of the protein are initially extracted using the pre-trained protein model ESM2_3B. In the model construction stage, a novel multi-view enzyme function prediction framework is proposed. This framework can extract both deep and broad semantic features of the protein from its original semantic features. For each extracted feature, a multi-head attention mechanism is used to capture the long-range dependencies between part-of-speech features and obtain preliminary enzyme function predictions. Finally, to extract the optimal common representation among the semantic information from multiple views, this method designs a multi-view adaptive fusion network to achieve multi-semantic information fusion and adaptive weighted classification.

[0013] The technical solution of the present invention is as follows:

[0014] A multi-level enzyme function prediction method based on multi-view semantics includes five steps: extraction of original semantic features of protein sequences, deep semantic feature extraction module, broad semantic feature extraction module, part-of-speech attention enhancer, and multi-view adaptive network fusion prediction. The steps are as follows:

[0015] Step 1: Input the protein sequence into the protein pre-training model ESM2_3B for preprocessing to obtain and save the original semantic information after feature extraction.

[0016] Step 2: Input the raw semantic information obtained in Step 1 into the deep semantic extraction feature module to obtain the deep semantic features of the enzyme.

[0017] Step 3: Input the raw semantic information obtained in Step 1 into the broad semantic extraction feature module to obtain the broad semantic features of the enzyme.

[0018] Step 4: Input the original, deep, and broad semantic features obtained in Steps 1, 2, and 3 into the part-of-speech attention enhancer to obtain the original, deep, and broad semantic features after part-of-speech attention enhancement.

[0019] Step 5: Input the original, deep, and broad semantic features enhanced by part-of-speech attention into the multi-view adaptive fusion network, train and save the enzyme function prediction model.

[0020] Step 6: After processing the protein sequence to be predicted in Step 1, input it into the enzyme function prediction model saved in Step 5.

[0021] A multi-level enzyme function prediction method based on multi-view semantics, the specific implementation process of step 1 is as follows: input the protein sequence into the protein pre-training model ESM2_3B, extract the original semantic information of the protein sequence, and finally convert each sample into a 2560-dimensional vector, and save each preprocessed sample vector.

[0022] The specific implementation process of step 2 is as follows: The deep semantic extraction module consists of four downsampling encoders and four upsampling decoders. The downsampling encoder consists of two convolutional layers with a width of 2, a batch normalization layer, and a max-pooling layer, with a ReLU activation function applied after each batch normalization layer. The upsampling decoder consists of an upsampling convolutional layer, feature concatenation, a batch normalization layer, and two convolutional layers with a width of 2, with a ReLU activation function applied after each batch normalization layer. Inputting the 2560-dimensional original semantic features obtained in step 1 into the downsampling encoder will yield 1280-dimensional, 640-dimensional, 320-dimensional, and 160-dimensional semantic features. We then input the 160-dimensional semantic features into the upsampling decoder to obtain 320-dimensional semantic features. We then concatenate the channels of these 320-dimensional semantic features, perform convolution and upsampling on the concatenated semantic features, and obtain 640-dimensional semantic features. This 640-dimensional semantic feature is then concatenated and convolved with the original 640-dimensional semantic features, followed by another upsampling. After four upsampling decoding steps, we obtain a 2560-dimensional deep semantic feature with the same size as the original semantic features.

[0023] The specific implementation process of step 3 is as follows: The broad semantic extraction module consists of two Bi-LSTMs and a residual connection. Bi-LSTM is a variant of a recurrent neural network (RNN). The model is divided into two independent LSTMs. The input sequence is fed into the two LSTM neural networks in both forward and reverse order for feature extraction. The word vector formed by concatenating the two output vectors (i.e., the extracted feature vectors) is used as the final feature representation of the word. It can simultaneously consider the forward and backward contextual information of the sequence. To avoid network degradation and to better utilize local features, we introduce a residual mechanism to fuse the two Bi-LSTM layers. The 2560-dimensional original semantic features obtained in step 1 are input into the broad semantic extraction module to obtain a 2560-dimensional broad semantic feature with the same size as the original semantic features.

[0024] The specific implementation process of step 4 is as follows: This invention proposes a novel part-of-speech attention enhancement mechanism to enhance three types of semantic information. The main goal of this method is to improve the accuracy of enzyme function prediction by capturing the part-of-speech dependencies of enzyme semantic information. In the self-attention mechanism, each input element interacts with all other elements to calculate its attention distribution. Specifically, the self-attention mechanism can be expressed as the following formula:

[0025]

[0026] Where N, A, and V represent query, key, and value, respectively, all of which are linear transformations of the input elements. k It is the dimension of the key vector. This is a scaling factor used to prevent the dot product from becoming too large. The softmax function ensures that all attention weights are between 0 and 1, and that their sum is 1. In the part-of-speech attention enhancer, we take deep, broad, and raw semantic information as input and compute their attention distribution through a multi-head self-attention mechanism. In the multi-head self-attention mechanism, we divide the input semantic information into multiple parts, each with its own N, A, and V. In this way, each head can learn a different attention distribution, thereby capturing different dependencies in the enzyme's semantic information. Specifically, the multi-head self-attention mechanism can be expressed as the following formula:

[0027] POSA(N,A,V)=Concat(head1,…,head n W o (2)

[0028] The output of each head i It can be represented as:

[0029]

[0030] in These are the query, key, and value transformation matrices for the i-th head. The `Concat` function concatenates all heads to form the final output. Because different enzyme functions require different numbers of part-of-speech neighbors, we use a multi-head approach to further extend multi-head self-attention. Specifically, each head uses a different kernel size `k`. Note that the number of heads is represented by `h`. To avoid adjusting the kernel size `k`, our strategy is to choose a single head with a fixed `k=3` (`h=1`), or to use kernel sizes `k1`, `k2`, ..., `k` with a fixed sequence. n In addition to h=1, we also set h=2, 4, etc. Details are as follows:

[0031] When h = 2, k = 3, 5.

[0032] When h = 4, k = 3, 5, 7, 9.

[0033] When h > 1, the kernel size k is set in ascending order. Diversity is introduced into branches with different k values ​​to generate better long-range dependency information. Due to the overlap effect caused by the similarity between different heads, we use adaptive weighted fusion of adaptive joint semantic long-range information instead of simple summation. In this way, our part-of-speech attention enhancer can understand the semantic information of enzymes from multiple perspectives, thus predicting enzyme functions more accurately. Inputting the 2560-dimensional original, deep, and broad semantic features obtained in steps 1, 2, and 3 into the next part-of-speech attention enhancement module yields the original, deep, and broad semantic features after part-of-speech enhancement.

[0034] The specific implementation process of step 5 is as follows: This invention proposes a multi-view adaptive fusion network, which constructs a classifier through a multilayer perceptron with a hidden layer based on the comprehensive feature representation. The predicted probability estimate for each label is as follows:

[0035]

[0036] Where W represents the weights of the fully connected layer. f is the ReLU nonlinear activation function. The sigmoid function is used to convert the output values ​​into probabilities. This invention employs cross-entropy loss because it has been shown to be suitable for protein function prediction [9,18,19]. Therefore, the binary classification loss function for enzyme or non-enzyme annotation tasks is defined as follows:

[0037]

[0038] Where N is the total number of samples, y i p is the category to which the i-th sample belongs. i This is the predicted value for whether the i-th protein sequence sample is an enzyme. The multi-class loss function for EC number function prediction is defined as follows:

[0039]

[0040] Where N represents the number of samples, K represents the number of EC categories, and p ic This represents the probability that the i-th enzyme sequence has class c function. y ic ∈{0,1}.

[0041] The three views extracted from different theories need to be combined to make a comprehensive decision. The following formula defines the comprehensive decision-making process from multiple perspectives:

[0042]

[0043] Where M is the number of views, w v ∈R is the weight of the v-th view. This is the preliminary prediction result for the v-th view. This is the comprehensive prediction result. Finally, we use the cross-entropy loss function (Equation 7) to optimize the comprehensive prediction result and save the trained model, that is:

[0044]

[0045] Step 6: After processing the protein sequence to be predicted in Step 1, input it into the enzyme function prediction model saved in Step 5.

[0046] The advantages of this invention include the following:

[0047] (1) We propose a multi-level enzyme function prediction method that integrates multi-view learning technology. This method can directly predict the function of the last level, i.e. the fourth level enzyme, and realize the synchronous prediction of non-enzyme protein sequences and different levels of labels in enzymes.

[0048] (2) Based on the latest protein language embedding methods and Bi-LSTM, a novel multi-view enzyme function prediction framework is proposed. This framework can extract deep and broad features of protein semantics. For each extracted feature, a multi-head attention mechanism is used to capture the long-range dependencies between part-of-speech features and obtain preliminary enzyme function predictions.

[0049] (3) An adaptive decision-making method based on multi-view learning mechanism is proposed, which combines the deep semantic features, broad semantic features and original semantic features of enzyme sequence to achieve optimal enzyme function prediction.

[0050] (4) To effectively evaluate the performance of our method, we conducted experiments on enzyme and isozyme datasets. We demonstrate that our method outperforms existing methods on both datasets. Attached Figure Description

[0051] Figure 1 This is the overall flowchart of the present invention.

[0052] Figure 2 This is a diagram of the algorithm framework of the present invention.

[0053] Figure 3 This is a structural diagram of the deep semantic extraction module of the present invention.

[0054] Figure 4 This is a structural diagram of the breadth semantic extraction module of the present invention.

[0055] Figure 5This is a comparison chart of the performance of the present invention with other existing enzyme function prediction methods on a benchmark dataset.

[0056] Figure 6 This is a comparison chart of the performance of the present invention with other existing enzyme function prediction methods on isoenzyme datasets. Detailed Implementation

[0057] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0058] like Figures 1-4 As shown, this invention implements a multi-level enzyme function prediction method based on multi-view semantics, and its algorithm architecture is as follows. Figure 2 As shown in the figure, this invention mainly comprises three modules: a raw semantic feature extraction module, a multi-view feature extractor module, and a multi-view adaptive prediction module. Specifically, unlike manually designed features, this invention first designs a new feature extraction module. This module introduces the latest protein large language model ESM2_3B to automatically extract primary features of the enzyme's amino acid sequence. Then, a set of multi-view deep feature extraction methods are designed to extract multiple sets of different view information from deep semantic perspective, broad semantic perspective, and raw semantic perspective. This invention constructs a deep semantic information extraction module (DSEM) and a broad semantic feature extraction module (BSEM) to obtain preliminary predictions of multiple features. The specific structure of DSEM is shown in the figure. Figure 3 As shown, the specific structure of BSEM is as follows: Figure 4As shown in the diagram. First, in the deep semantic information extraction module, a multi-layer convolutional encoder is used to cascade and concatenate multiple convolutional layers, and then a multi-layer convolutional decoder is used for decoding to ensure the scale invariance of features. Due to the small convolutional scale, local deep semantic information can be obtained. Second, since enzyme function is not only determined by the local regions of its sequence, but may also be affected by distant parts of the sequence, this phenomenon indicates that some functionally related regions of the enzyme sequence may be far apart in one-dimensional sequences, but are close together in three-dimensional structures, interacting and jointly determining the enzyme's activity and function. Local patterns cannot capture sufficient long-range dependencies between enzyme sequences. Therefore, introducing Bi-LSTM to construct a broad semantic feature extraction module can not only capture sufficient local information, but also capture long-range dependencies in the enzyme sequence, thereby obtaining broad semantic information. This invention mines primary, deep, and broad information of enzyme sequences. In enzyme sequences, long-range dependencies between residues affect enzyme function. Therefore, to better integrate these three types of information and achieve better enzyme function prediction, this invention proposes a new part-of-speech attention enhancement mechanism to achieve the fusion of these three types of information. The main objective of this method is to improve the accuracy of enzyme function prediction by capturing the part-of-speech dependencies of enzyme semantic information. Finally, this invention proposes an adaptive decision-making module that integrates a multi-view learning mechanism to achieve dynamic joint decision-making on information from different views, thereby realizing enzyme function prediction.

[0059] Example 1

[0060] To fully validate the effectiveness of the proposed method, this paper compares it with six advanced enzyme function prediction models (MVDINET, HECNet, DEEPre, ECPred, Clean, and ABLE). The benchmark dataset used in this invention comes from Swiss-Prot in Uniprot, which is currently the world's largest and most comprehensive protein database, containing a large number of protein sequences. Based on annotations, the proteins in the database are divided into enzymes and non-enzymes, with enzymes consisting of monofunctional and multifunctional enzymes. Specific data for enzymes and non-enzymes in the benchmark dataset are shown in Table 1, and specific data for different enzyme types are shown in Table 2.

[0061] Table 1. Specific data for enzymes and non-enzymes.

[0062]

[0063] Table 2. Specific data for different types of enzymes

[0064]

[0065] The results of the present invention and the comparative invention on the benchmark dataset are shown in Table 3. Figure 5The differences between the various methods are visualized. As shown in the table, this invention achieved the best results on the benchmark dataset among all inventions.

[0066] Table 3 Performance comparison between the present invention and existing multi-level enzyme function prediction methods

[0067]

[0068] Compared to other inventions mentioned above, this invention fully mines the semantic information of enzymes while utilizing the latest protein large language model, resulting in significantly higher performance. Specifically, this invention achieves an F1 score of 95.9% in the third layer, while MVDINET, HECNet, and DEEPre achieve F1 scores of 91.8%, 79.4%, and 68.4%, respectively. Similarly, in the fourth layer prediction, this invention achieves an F1 score of 94.6%, while MVDINET and HECNet achieve F1 scores of 90.5% and 81.90%, respectively. These results demonstrate that this invention outperforms existing methods.

[0069] Example 2

[0070] This section explores the robustness of the proposed invention in the face of isozymes. Isozymes are enzymes that catalyze the same reaction in organisms but differ in structure, charge, solubility, or other biochemical properties. These differences are usually due to being encoded by different genes or different splicing forms of the same gene. Predicting the enzyme function of isozymes is crucial for understanding physiological processes and disease mechanisms in organisms. Experimentally validated isozyme sequences were selected from the Swiss-Prot database as the isozyme dataset for this invention. Isozyme sequences shorter than 50 or longer than or equal to 1000, as well as those belonging to multiple categories, were removed, resulting in 2027 isozyme sequences. The invention was compared with pre-trained models from DeepEC, DETECT_2, and HECNet on the isozyme dataset, and the results are shown in Table 4. Figure 6 The visualizations illustrate the differences between the various methods. For levels 1-4, the accuracy of this invention is higher than the other three methods. Particularly at level 1, the F1 score of this invention is significantly higher than the other three methods. Although the advantage of this invention gradually decreases from levels 2-4, it still outperforms the other three methods by an average of 9%, 11%, and 15%, respectively. Based on the experimental results, this invention can predict different types of isoenzymes and is superior to state-of-the-art methods.

[0071] Table 4 Performance comparison of the present invention and existing methods on isozyme datasets

[0072]

Claims

1. A multi-level enzyme function prediction method based on multi-view semantics, characterized in that, The steps are as follows: Step 1: Input the protein sequence into the protein pre-training model ESM2_3B for preprocessing, and obtain and save the original semantic information after feature extraction; Step 2: Input the raw semantic information obtained in Step 1 into the deep semantic extraction feature module to obtain the deep semantic features of the enzyme; Step 3: Input the raw semantic information obtained in Step 1 into the broad semantic extraction feature module to obtain the broad semantic features of the enzyme; Step 4: Input the original, deep, and broad semantic features obtained in Steps 1, 2, and 3 into the part-of-speech attention enhancer to obtain the original, deep, and broad semantic features after part-of-speech attention enhancement; Step 5: Input the original, deep, and broad semantic features enhanced by part-of-speech attention into the multi-view adaptive fusion network, train and save the enzyme function prediction model; The implementation process of step 5 is as follows: By constructing a classifier using a multilayer perceptron with one hidden layer, the predicted probability of each label is estimated as follows: (4) in The weights of the fully connected layer are denoted by f; f is the ReLU nonlinear activation function; the sigmoid function is used to convert the output value into a probability; cross-entropy loss is used, and the binary classification loss function for enzyme or non-enzyme annotation tasks is defined as follows: (5) in, It is the total number of samples. It is the first The category to which each sample belongs. It is the first Whether a protein sequence sample is a predicted value of an enzyme; the multi-class loss function for EC number function prediction is defined as follows: (6) in, Represents the number of samples. Represents the number of EC category numbers. Representing the Each enzyme sequence has a category The probability of function , , ; By extracting three views from different theories, a comprehensive decision-making process from multiple perspectives is defined, namely: (7) in, M It is the number of views. It is the weight of the v-th view. This is the preliminary prediction result for the v-th view. This is the comprehensive prediction result; finally, the cross-entropy loss function is used to optimize the comprehensive prediction result and save the trained model, that is: (8) Step 6: After processing the protein sequence to be predicted in Step 1, input it into the enzyme function prediction model saved in Step 5.

2. The multi-level enzyme function prediction method based on multi-view semantics as described in claim 1, characterized in that, The deep semantic feature extraction module in step 2 includes four downsampling encoders and four upsampling decoders. The downsampling encoder consists of a convolutional layer with two kernels of width 2, a batch normalization layer, and a max pooling layer, and uses the non-linear activation function ReLU after each batch normalization layer. The upsampling decoder consists of an upsampling convolutional layer, feature concatenation, a batch normalization layer, and two convolutional layers with a kernel size of 2. The non-linear activation function ReLU is used after each batch normalization layer.

3. The multi-level enzyme function prediction method based on multi-view semantics as described in claim 1 or 2, characterized in that, The specific operations in step 2 are as follows: The 2560-dimensional original semantic features obtained in step 1 are input into the downsampling encoder, resulting in 1280-dimensional, 640-dimensional, 320-dimensional, and 160-dimensional semantic features. The 160-dimensional semantic features are then input into the upsampling decoder to obtain 320-dimensional semantic features. The previously obtained 320-dimensional semantic features are then concatenated, and the concatenated semantic features are then convolved and upsampled to obtain 640-dimensional semantic features. These 640-dimensional semantic features are then concatenated and convolved with the previous 640-dimensional semantic features, and then upsampled. After four upsampling decodings, a 2560-dimensional deep semantic feature with the same size as the original semantic features is obtained.

4. The multi-level enzyme function prediction method based on multi-view semantics as described in claim 1 or 2, characterized in that, The breadth semantic feature extraction module in step 3 consists of two Bi-LSTMs and a residual connection. Bi-LSTM is a variant of a recurrent neural network. The model is divided into two independent LSTMs. The input sequence is input into the two LSTM neural networks in forward and reverse order, respectively, for feature extraction. The word vector formed by concatenating the two output vectors is used as the final feature expression of the word.

5. The multi-level enzyme function prediction method based on multi-view semantics as described in claim 1, characterized in that, The implementation process of step 4 is as follows: The self-attention mechanism can be expressed by the following formula: (1) in, , and These represent query, key, and value, respectively, all of which are linear transformations of the input elements; It is the dimension of the key vector. It is a scaling factor; in the multi-head self-attention mechanism, the semantic information of the input is divided into multiple parts, each of which has its own N, A and V; The bullish self-attention mechanism can be expressed by the following formula: (2) Output of each head Represented as: (3) in, , , They are the first i The query, key, and value transformation matrix for each header; the Concat function joins all the headers together to form the final output; Each header uses a different kernel size. k Note that the number of heads is expressed as h The strategy adopted is to select those with fixed... k A single head of 3, or using a kernel size with a fixed sequence. , , , .Apart from h = 1, also set h =2, 4, details are as follows: when h = 2, k = 3, 5. when h = 4, k = 3, 5, 7, 9. when h When the value is greater than 1, the kernel size k is set in ascending order; diversity is introduced into branches with different k values ​​to produce better long-range dependency information.

6. The multi-level enzyme function prediction method based on multi-view semantics as described in claim 1, characterized in that, The implementation process of step 1 is as follows: The protein sequence is input into the protein pre-training model ESM2_3B to extract the original semantic information of the protein sequence. Finally, each sample is converted into a 2560-dimensional vector, and each preprocessed sample vector is saved.

Citation Information

Patent Citations

  • Multistage enzyme function prediction method based on multi-view deep interactive learning

    CN116312851A