A hybrid deep learning strategy for protein melting temperature prediction

By hybridizing deep learning strategies and combining sequence and structural feature extraction modules, the problem of the existing technology failing to effectively utilize the local information of amino acid residues is solved, and more accurate protein melting temperature prediction is achieved.

CN119446287BActive Publication Date: 2025-09-16ANQING NORMAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411436748.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-09-16
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Existing methods fail to effectively utilize local information at the amino acid residue level when predicting protein melting temperature, resulting in insufficient prediction accuracy.

Method used

A hybrid deep learning strategy is adopted, combining the sequence feature extraction module and the structural feature extraction module. The feature information of proteins is extracted through the ProtT5-XL-UniRef50 and ESMFold models. Deep learning models such as multi-head attention network, multi-scale convolutional neural network, bidirectional long short-term memory network and graph attention convolutional network are used to comprehensively consider the structural and sequence characteristics of proteins.

Benefits of technology

The prediction accuracy of protein melting temperature has been improved, especially the utilization of the spatial conformation of amino acid residues, which has achieved more accurate temperature prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446287B_ABST
    Figure CN119446287B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of protein prediction technology, and discloses a method for predicting protein melting temperature using a hybrid deep learning strategy, comprising the following steps: Step 1, constructing a protein melting temperature prediction model using a deep learning strategy, wherein the prediction model includes a sequence feature extraction module, a structural feature extraction module, and a forward network module; Step 2, inputting the primary sequence of the protein to be tested into the prediction model; Step 3, outputting the protein melting temperature and protein category using the prediction model. The prediction method proposed in the present invention has a prediction model that comprehensively considers the structural and sequence features of the protein, particularly the spatial conformation of the amino acid residues in the protein structure. Therefore, the prediction model proposed in the present invention can more accurately predict the protein melting temperature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protein performance prediction, and in particular to a protein melting temperature prediction method using a hybrid deep learning strategy. Background Art

[0002] The melting temperature (Melting temperature) of a protein is an important indicator for assessing its thermal stability. It refers to the temperature at which a protein partially or completely loses its three-dimensional structure and dissociates into amino acids under high temperature conditions. Different proteins have different melting temperatures, depending on their structure and amino acid composition. Generally, the dissociation temperature of proteins ranges from 50°C to 90°C. The melting temperature directly reflects the folding state and thermal stability of a protein and is of great significance for understanding protein structure and function, optimizing protein design, and improving industrial production efficiency.

[0003] Computational methods for predicting protein melting temperatures can quickly obtain this metric. However, existing methods mostly rely on global protein-level input data, failing to effectively utilize local information at the level of the amino acid residues that make up the protein. In particular, protein structure determines function, and protein structure is determined by the spatial conformation of its amino acid residues.

[0004] A protein is made up of multiple amino acids linked together by condensation (removing H+ and OH-). The primary structure of a protein refers to the sequence of residues that make up the protein. Protein structure refers to the three-dimensional conformation of a protein molecule in its native folded state.

[0005] Currently, computational models for representing proteins mostly use features extracted at the protein level, such as amino acid content, dipeptide content, physicochemical properties, and structural composition. These computational models primarily rely on random forests, support vector machines, and gradient boosting trees, focusing primarily on input feature selection or dimensionality reduction.

[0006] Starting from the constituent amino acids of proteins, we can make full use of sequence feature information and amino acid local structure information to better represent the characteristics of proteins;

[0007] Computational methods for predicting protein melting temperatures fall into two main categories: regression problems for predicting specific melting temperature values ​​and classification problems for identifying thermophilic proteins. Classification problems can be categorized into three types: a three-class classification problem for the temperature limits of hyperthermophilic proteins (>65°C), thermophilic proteins (55°C-65°C), and mesophilic proteins (<55°C); and a two-class classification problem for thermophilic versus mesophilic proteins. Summary of the Invention

[0008] In order to solve the technical problems raised in the background technology, the present invention provides a protein melting temperature prediction method based on a hybrid deep learning strategy.

[0009] The present invention is implemented by the following technical solution: a method for predicting protein melting temperature using a hybrid deep learning strategy, comprising the following steps:

[0010] Step 1: constructing a protein melting temperature prediction model based on a deep learning strategy, wherein the prediction model includes a sequence feature extraction module, a structural feature extraction module, and a forward network module;

[0011] Step 2: inputting the primary sequence of the protein to be tested into the prediction model;

[0012] Step 3: Use the prediction model to output protein melting temperature and protein category.

[0013] Specifically, constructing the prediction model in step 1 includes the following steps:

[0014] 1.3. Obtaining training dataset

[0015] 1.4, Feature extraction;

[0016] 1.3. Build a deep learning model.

[0017] Specifically, the specific operation of step 1.1 is as follows: the training and test data are from the Meltome Atlas dataset, proteins with a sequence length of less than 1024 are extracted from the Meltome Atlas dataset, and sequences with a sequence similarity greater than 30% are removed using CD-HIT software. These sequences are then divided into a training set, a validation set, and a test set at an 8:1:1 ratio using a fixed random seed.

[0018] Specifically, the steps 1.2 are as follows:

[0019] 1.21. Feature extraction using the protein language pre-trained large model ProtT5-XL-UniRef50 and the ESMFold model:

[0020] 1.22. Normalize the protein melting temperature;

[0021] Specifically, the steps 1.21 are as follows:

[0022] The primary sequence of the protein is input into the model ProtT5-XL-UniRef50, and the output data is the feature representation vector X of the i-th residue in the corresponding sequence i , and X i The dimension is 1024;

[0023] ESMFold is used to predict the three-dimensional structure of proteins, obtain the three-dimensional coordinates of atoms in each amino acid residue in the sequence, and then calculate the C α -C α The Euclidean distance of atoms, if the Euclidean distance is less than The two amino acids are considered to be in contact, and a residue contact map matrix is ​​obtained, the size of which is L*L, where L is the sequence length;

[0024] Where: Node e in row i and column j of the residue contact map matrix ij =1, indicating that there is a contact relationship between the i-th residue and the j-th residue in the sequence, e ij =0 means that there is no contact relationship between the i-th residue and the j-th residue in the sequence;

[0025] Specifically, the steps 1.22 are as follows:

[0026] The protein melting temperature was normalized using formula (1), where Tm i is the true melting temperature of the i-th protein, Max is the highest melting temperature value in the current data set, and Min is the lowest melting temperature;

[0027] In the three-classification, the protein temperature category C is divided according to formula (2), with 0, 1, and 2 representing the numerical temperature category respectively;

[0028] In the binary classification, proteins with a melting temperature ≥55°C were classified as thermophilic proteins, otherwise they were mesophilic proteins;

[0029]

[0030]

[0031] Then the feature representation vector X of the i-th residue in the corresponding sequence is i , the contact map matrix is ​​input into the sequence feature extraction module and the structural feature extraction module respectively.

[0032] Specifically, the deep learning model in step 1.3 includes a position encoding module, a sequence feature extraction module, a structural feature extraction module, and an output module.

[0033] Specifically, in the position encoding module, a sine function or a cosine function is used for encoding, and the encoding is shown in formula (3) and formula (4);

[0034]

[0035] Where i represents the position of the amino acid in the protein sequence, 0≤i<L, D represents the feature dimension D=1024, 2j and 2j+1 represent the feature component positions of the residue at the current position, 0≤j<D / 2;

[0036] Then, the residue representation input to the dual-channel network module is added with position information. The specific operation is as shown in formula (5):

[0037]

[0038] Specifically, the sequence feature extraction module consists of a multi-head attention network, a multi-scale convolutional neural network, a bidirectional long short-term memory network, and a forward attention network. The input data of the sequence feature extraction module is

[0039] The multi-head attention network uses 8 attention heads; in the multi-head attention network, its inputs each pass through a layer of forward neural network to obtain Q i , K i 、V i , as shown in formula (6), the input data is 1024-dimensional and the output data is 128-dimensional; and an attention head i Calculation is as in formula (7) and formula (8), where d k The value is 128; the "Concat" operation in formula (9) means that the data is merged in the last dimension, and the 8 attention heads are spliced ​​in the last dimension, and then output after passing through a layer of forward neural network. The output dimension is 1024, and the length of the protein sequence remains unchanged;

[0040]

[0041]

[0042] Head i =Score(Q i ,K i ,V i )V i (8)

[0043] MA(Q,K,V)=Concat(Head1,Head2,...,Head8)W o (9)

[0044] The multi-scale convolutional neural network uses one-dimensional convolution and multiple convolutional neural networks with different convolution kernel sizes to perform convolution operations in parallel; the input data comes from the output data of the multi-head attention network, the input data is 1024-dimensional, the output data is 128-dimensional, and the convolution kernel sizes are [3, 5, 7, 9] respectively; the default activation function of the convolution operation is ReLU; the outputs of multiple convolution units are merged, as shown in formula (11): the output of the multi-scale convolutional neural network is 512-dimensional;

[0045] Conv i =f(W*x j:j+k-1 +b)i∈{1,2,3,4},j∈{0,1...,1023},k∈{3,5,7,9}(10)

[0046] Conv=Concat(Conv1,Conv2,Conv3,Conv4)(11)

[0047] The bidirectional long short-term memory network is used to focus on the global information of proteins and capture the long-range dependencies of protein sequences. It is composed of a bidirectional long short-term memory network. The unidirectional LSTM model is described in the formula as shown in (12).

[0048]

[0049]

[0050] Among them, σ is the activation function using the Sigmoid function; ⊙ represents the matrix bitwise multiplication; x t is the network input at time t; i t 、f t 、、o t 、c t and h t Respectively represent the input gate, forget gate, output gate, internal memory unit and output at time t; h t-1 is the output of the previous moment; c t-1 is the output of the internal memory unit at the previous moment; the rest are the learnable parameters of the neural network;

[0051] The unidirectional LSTM network has a 1024-dimensional input and a 256-dimensional output. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (13), and the output is 512-dimensional.

[0052] The forward attention network is used to convert the two-dimensional matrix into a one-dimensional vector, h t is the output of the previous layer network, representing the tth amino acid residue in the sequence;

[0053] Through a layer of forward neural network as shown in formula (14), the attention weight of the current feature representation of residue t in the sequence is obtained by the formula; the feature representations of all residues in the sequence are weighted and summed, as shown in formula (16), to achieve the conversion of two-dimensional feature representation to one-dimensional; here, L is 1024; the input data dimension is L*512, and the output dimension is 1*512;

[0054] e t =h t W t (14)

[0055]

[0056] Specifically, the structural feature extraction module includes a graph attention convolutional network, a graph diffusion network, and a graph average pooling network, wherein:

[0057] The spatial structure of proteins in the graph attention convolutional network is represented by a graph, where nodes are amino acid residues and the edges between nodes represent the contact between two residues. The node representation is the sequence feature information of formula (5), which is input Input1; the edge information is described by the contact map matrix, which is input Input2.

[0058] The network nodes are connected to the forward network layer, as shown in formula (17), with an input dimension of 1024 and an output dimension of 256;

[0059] For the neighbor node j of the current node i, data is concatenated, then multiplied by the learnable vector a (dimension 256*2, 1), compressed into a scalar, and then passed through the LeakyReLU activation function, as shown in formula (19);

[0060] Calculate the attention score α between the current node i and the neighbor node j ij , the calculation formula uses the Softmax function, such as formula (20), where N(i) represents the set of neighbor nodes of node i, and then updates the information of node i, such as formula (21), the activation function σ uses the LeakyReLU function, and the output dimension is 256;

[0061] h' i =Wh i (17)

[0062] h' ij =Concat(h i ,h j ) (18)

[0063] e i,j =LeakyReLU(a T h' ij ) (19)

[0064]

[0065] The graph diffusion network obtains the adjacent node information of the adjacent nodes of a node through diffusion convolution;

[0066] The original protein-protein contact matrix A has a dimension of 1024*1024, I is the identity matrix with a diagonal of 1, and the dimension is the same as A. D is the degree matrix of matrix A, D -1 is the normalized matrix of D, and the matrix A is transformed by formula (22-23) to become H (0) That is, the node information output by the graph attention convolutional network, with a dimension of 256; H (k) is the node feature after k iterations, α is a learnable balance parameter with an initial value of 0.9; W is the parameter matrix;

[0067] The diffusion operation of node information acquisition is realized through formula (24) and formula (25). The operation is repeated 5 times, expecting to obtain neighboring point information at 5 levels; the output node dimension of this network module is still 256.

[0068]

[0069] H (k+1) =ReLU(WH') (25)

[0070] The output node information and the original adjacency matrix of the graph diffusion network are sent back to the graph attention convolutional network unit, and the above operation is repeated iteratively again; at the same time, the node information is sent to the graph average pooling network; the results of repeated iterations are also sent to the graph average pooling network.

[0071] The graph average pooling network is used to adopt a graph global pooling strategy for the node information in the graph. The current feature of the i-th residue in the sequence is represented as x i , take the mean of all such features in the sequence, as shown in formula (26), and compress the L*256 matrix into a 1*256 vector through graph global pooling;

[0072]

[0073] The two output results of the graph average pooling unit are concatenated to obtain a 1*512 dimension vector.

[0074] It is necessary to perform feature concatenation on the output F1 of the sequence feature extraction network and the output F2 of the structure feature extraction network, as shown in formula (27), with an output dimension of 1*1024;

[0075] F=Concat(F1,F2)(27)

[0076] Then connect two layers of feedforward neural networks, as shown in formula (28) and formula (29), respectively. The first feedforward neural network uses tanh as activation function, and the output dimension is 256; the second one uses Sigmoid function or Softmax function, and the output dimension is 1 or 3;

[0077] F′=tanh(W*F+b) (28)

[0078]

[0079] In the output module:

[0080] The error loss between the model prediction value and the true value is described by formula (30) or formula (31), where (30) is the melting temperature prediction loss and formula (31) is the two-class or three-class prediction loss;

[0081]

[0082] In the above formula, N is the number of training samples and i is the i-th protein.

[0083] Compared with the prior art, the present invention has the following beneficial effects:

[0084] In the prior art, when using computational methods to predict protein melting temperatures, most of the input data are global features at the protein level, and local information at the level of amino acid residues that make up the protein is not effectively utilized. Compared with the prior art, the prediction method proposed in the present invention has a prediction model that comprehensively considers the structural characteristics and sequence characteristics of the protein, especially the spatial conformation of amino acid residues in the protein structure. Therefore, the prediction model proposed in the present invention can more accurately predict the protein melting temperature. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 This is a flowchart of the prediction method proposed by the present invention;

[0086] Figure 2 This is a schematic diagram of the overall structure of the prediction model proposed in the present invention;

[0087] Figure 3 This is a diagram of a network module for extracting sequence features in the prediction model proposed by the present invention;

[0088] Figure 4 This is a diagram of the network module for extracting structural features in the prediction model proposed in this invention. DETAILED DESCRIPTION

[0089] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0090] Example:

[0091] Reference Figures 1-4 ,This scheme proposes a protein melting temperature prediction method based on a hybrid deep learning strategy,,which includes the following steps:

[0092] Step 1: constructing a protein melting temperature prediction model based on a deep learning strategy, wherein the prediction model includes a sequence feature extraction module, a structural feature extraction module, and a forward network module;

[0093] Step 2: inputting the primary sequence of the protein to be tested into the prediction model;

[0094] Step 3: Use the prediction model to output protein melting temperature and protein category.

[0095] Specifically, constructing the prediction model in step 1 includes the following steps:

[0096] 1.1. Obtaining training dataset

[0097] 1.2, Feature extraction;

[0098] 1.3. Build a deep learning model.

[0099] Specifically, the specific operation of step 1.1 is as follows: the training and test data are from the Meltome Atlas dataset, proteins with a sequence length of less than 1024 are extracted from the Meltome Atlas dataset, and sequences with a sequence similarity greater than 30% are removed using CD-HIT software. These sequences are then divided into a training set, a validation set, and a test set at an 8:1:1 ratio using a fixed random seed.

[0100] Specifically, the steps 1.2 are as follows:

[0101] 1.21. Feature extraction using the protein language pre-trained large model ProtT5-XL-UniRef50 and the ESMFold model:

[0102] 1.22. Normalize the protein melting temperature;

[0103] Specifically, the steps 1.21 are as follows:

[0104] The primary sequence of the protein is input into the model ProtT5-XL-UniRef50, and the output is the feature representation vector X of the i-th residue in the corresponding sequence i , X i The dimension is 1024;

[0105] ESMFold was used to predict the three-dimensional structure of proteins, obtain the three-dimensional coordinates of atoms in each amino acid residue in the sequence, and calculate the C α -C α The Euclidean distance of atoms, if the Euclidean distance is less than The two amino acids are considered to be in contact, and a residue contact map matrix is ​​obtained, the size of which is L*L, where L is the sequence length;

[0106] Where: Node e in row i and column j of the residue contact map matrix ij =1, indicating that there is a contact relationship between the i-th residue and the j-th residue in the sequence, e ij =0, indicating that there is no contact relationship between the i-th residue and the j-th residue in the sequence;

[0107] Specifically, considering the convenience of data batch input during deep learning model training, all sequence lengths are aligned to 1024, and those less than 1024 are padded with zeros. The residue features are 1024*1024, and the contact map matrix is ​​1024*1024.

[0108] Specifically, the steps 1.22 are as follows:

[0109] Normalize the protein melting temperature using formula (1), where Tm i is the true melting temperature of the i-th protein, Max is the highest melting temperature in the current dataset, and Min is the lowest melting temperature. In the three-class classification, protein temperature category C is divided according to formula (2), with 0, 1, and 2 representing the numerical temperature category, respectively. In the two-class classification (0 and 1 represent numerical categories), proteins with a melting temperature ≥ 55°C are classified as thermophilic proteins, otherwise they are mesophilic proteins.

[0110]

[0111] The feature representation vector X of the i-th residue in the corresponding sequence i , the contact map matrix is ​​input into the sequence feature extraction module and the structural feature extraction module respectively.

[0112] Specifically, the deep learning model in step 1.3 includes a position encoding module, a sequence feature extraction module, a structural feature extraction module, and an output module;

[0113] Specifically, in the position coding, the arrangement of amino acids in the protein starts from the N-terminal amino group and ends at the C-terminal carboxyl group. In order to enable the model to utilize the order information of amino acids in the protein sequence, the sequence features are positionally coded. The coding is performed using a sine function or a cosine function, as shown in formula (3) and formula (4).

[0114]

[0115] Where i represents the position of the amino acid in the protein sequence, 0≤i<L, D represents the feature dimension D=1024, 2j and 2j+1 represent the feature component positions of the residue at the current position, 0≤j<D / 2;

[0116] For the residue representation input to the dual-channel network module, position information is added. The specific operation is as follows:

[0117]

[0118] Specifically, in the sequence feature extraction module;

[0119] The sequence feature extraction module consists of a multi-head attention network, a multi-scale convolutional neural network, a bidirectional long short-term memory network (BLSTM), and a forward attention network. The specific structure is as follows: Figure 2 As shown. And the input of the sequence feature extraction module is the input in formula (5)

[0120] About the multi-head attention network;

[0121] Eight attention heads are used; in the multi-head attention network, each input passes through a layer of forward neural network to obtain Q i , K i 、V i , as shown in formula (6), the input data is 1024-dimensional and the output data is 128-dimensional;

[0122] And an attention head i Calculation is as in formula (7) and formula (8), where d k value 128;

[0123] The "Concat" operation in formula (9) indicates that the data is merged in the last dimension. The meaning of this operation is consistent in the present invention. The eight attention heads are concatenated in the last dimension and then output after a layer of forward neural network (MLP). The output dimension is 1024, and the length of the protein sequence remains unchanged.

[0124]

[0125]

[0126] Head i =Score(Q i ,K i ,V i )V i (8)

[0127] MA(Q,K,V)=Concat(Head1,Head2,...,Head8)W o (9)

[0128] Regarding the multi-scale convolutional neural network, it adopts one-dimensional convolution, and multiple convolutional neural networks with different convolution kernel sizes perform convolution operations in parallel; the input data comes from the output data of the multi-head attention network, the input data is 1024-dimensional, the output data is 128-dimensional, and the convolution kernel size (the number of residues participating in the convolution each time) is [3, 5, 7, 9] respectively, aiming to fully extract the local characteristics of the sequence and fit the characteristics of the secondary structure composed of a small number of continuous residues; the default activation function of the convolution operation is ReLU; the output of multiple convolution devices is merged, as shown in formula (11): the output of this network layer is 512-dimensional.

[0129] Conv i =f(W*x j:j+k-1 +b)i∈{1,2,3,4},j∈{0,1...,1023},k∈{3,5,7,9} (10)

[0130] Conv=Concat(Conv1,Conv2,Conv3,Conv4) (11)

[0131] About a layer of bidirectional long short-term memory network (BLSTM); the second part focuses on the global information of proteins, captures the long-range dependencies of protein sequences, and is composed of a bidirectional long short-term memory network; the unidirectional LSTM model is formulated as shown in (12).

[0132]

[0133] Among them, σ is the activation function using the Sigmoid function; ⊙ represents the matrix bitwise multiplication; x t is the network input at time t; i t 、f t 、、o t 、c t and h t Respectively represent the input gate, forget gate, output gate, internal memory unit and output at time t; h t-1 is the output of the previous moment; c t-1 is the output of the internal memory unit at the previous moment; the rest are the learnable parameters of the neural network.

[0134] The unidirectional LSTM network has a 1024-dimensional input and a 256-dimensional output. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (13), and the output is 512-dimensional.

[0135] The forward attention network is used to convert the two-dimensional matrix into a one-dimensional vector, h t is the output of the previous layer network, representing the tth amino acid residue in the sequence;

[0136] Through a layer of forward neural network as shown in formula (14), the attention weight of the current feature representation of residue t in the sequence is obtained by the formula; the feature representations of all residues in the sequence are weighted and summed, as shown in formula (16), to realize the conversion of two-dimensional feature representation to one-dimensional; the input data dimension is L*512, and the output dimension is 1*512.

[0137] e t =h t W t (14)

[0138]

[0139] Specifically, the structural feature extraction module includes a graph attention convolutional network unit, a graph diffusion network unit, and a graph average pooling unit, wherein:

[0140] The spatial structure of proteins in the graph attention convolutional network unit is represented by a graph, where the nodes are amino acid residues and the edges between the nodes represent the contact between two residues (C α The distance between atoms is less than ); the node representation uses the sequence feature information of formula (5), that is, input Input1; the edge information is described by the contact graph matrix, that is, input Input2;

[0141] The network node is connected to the forward network layer, as shown in formula (17), with an input dimension of 1024 and an output dimension of 256.

[0142] For the neighbor node j of the current node i, data is concatenated, then multiplied by the learnable vector a (dimension 256*2, 1), compressed into a scalar, and then passed through the LeakyReLU activation function, as shown in formula (19);

[0143] Calculate the attention score α between the current node i and the neighbor node j ij , the calculation formula uses the Softmax function, such as formula (20), where N(i) represents the set of neighbor nodes of node i, and then updates the information of node i, such as formula (21), where the activation function σ uses the LeakyReLU function, and the output dimension is 256;

[0144] h' i =Wh i (17)

[0145] h' ij =Concat(h i ,hj ) (18)

[0146] e i,j =LeakyReLU(a T h' ij ) (19)

[0147]

[0148] In the diffusion network unit of the graph, the adjacent node information of the adjacent nodes of the node is obtained through diffusion convolution (in the graph, the information of node 1, adjacent nodes 2 and 3 is obtained after one diffusion convolution, and the information of adjacent node 2, adjacent node 4, and adjacent node 3, adjacent node 5, is obtained).

[0149] The original protein-protein contact matrix A has a dimension of 1024*1024, I is the identity matrix with a diagonal of 1, and the dimension is the same as A. D is the degree matrix of matrix A, D -1 is the normalized matrix of D, and the matrix A is transformed by formula (22-23) to become H (0) That is, the node information output by the graph attention convolutional network, with a dimension of 256; H (k) is the node feature after k iterations, α is a learnable balance parameter with an initial value of 0.9; W is the parameter matrix;

[0150] The diffusion operation of node information acquisition is implemented through formula (24-25). This operation is repeated 5 times, hoping to obtain neighboring point information at 5 levels; the output node dimension of this network module is still 256.

[0151]

[0152] H (k+1) =ReLU(WH') (25)

[0153] The output node information and the original adjacency matrix (i.e., Contact Map) of the graph diffusion network unit are sent back to the graph attention convolutional network unit, and the above operation is repeated again. At the same time, the node information is also sent to the graph average pooling network module. The results of the above repeated operation are also sent to the graph average pooling unit network module.

[0154] The graph average pooling unit is used to apply a graph global pooling strategy to the node information in the graph. The current feature of the i-th residue in the sequence is represented as x i , take the mean of all features of this type in the sequence, as shown in formula (26). Through graph global pooling, the L*256 matrix is ​​compressed into a 1*256 vector;

[0155]

[0156] The output results of the two graph average pooling units are concatenated to obtain a 1*512 dimension vector.

[0157] It is necessary to perform feature concatenation on the output F1 of the sequence feature extraction network and the output F2 of the structure feature extraction network, as shown in formula (27), with an output dimension of 1*1024.

[0158] F=Concat(F1,F2)(27)

[0159] Then connect two layers of feedforward neural networks, as shown in formulas (28) and (29), respectively. The first feedforward neural network uses tanh as the activation function, and the output dimension is 256; the second one uses Sigmoid function (regression and two-classification) or Softmax function (three-classification), and the output dimension is 1 or 3.

[0160] F′=tanh(W*F+b) (28)

[0161]

[0162] Specifically, in the output module:

[0163] The error loss between the model prediction value and the true value is described by (30) or (31), where (30) is the melting temperature prediction loss and (31) is the two-class or three-class prediction loss.

[0164]

[0165] In the above formula, N is the number of training samples and i is the i-th protein.

[0166] The above embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and replacements made by technicians in this field on the basis of the present invention fall within the scope of protection required by the present invention.

Claims

1. A protein melting temperature prediction method based on a hybrid deep learning strategy, characterized in that: The steps include: Step 1: Construct a protein melting temperature prediction model based on a deep learning strategy, wherein the prediction model includes a sequence feature extraction module and a structural feature extraction module; Step 2: inputting the primary sequence of the protein to be tested into the prediction model; Step 3: Use the prediction model to output protein melting temperature and protein category; Constructing the prediction model in step 1 includes the following steps: Step 1.1, obtain the training data set; Step 1.2, feature extraction; Step 1.3: Build a deep learning model. The specific operations of step 1.2 are as follows: Step 1.21: Use the protein language pre-trained large model ProtT5-XL-UniRef50 and the ESMFold model for feature extraction: The specific operations of step 1.21 are as follows: The primary sequence of the protein is input into the model ProtT5-XL-UniRef50, and the output data is the feature representation vector X of the i-th residue in the corresponding sequence i , and X i The dimension is 1024; ESMFold is used to predict the three-dimensional structure of proteins, obtain the three-dimensional coordinates of atoms in each amino acid residue in the sequence, and then calculate the distance between residues. The Euclidean distance of atoms. If the Euclidean distance is less than 10Å, the two amino acids are considered to be in contact, and the residue contact map matrix is ​​obtained, with a size of , and L is the sequence length; Where: the node in row i and column j in the residue contact map matrix , indicating that there is a contact relationship between the i-th residue and the j-th residue in the sequence, Indicates that there is no contact relationship between the i-th residue and the j-th residue in the sequence; Step 1.22, normalize the protein melting temperature; The specific operations of step 1.22 are as follows: The protein melting temperature was normalized using formula (1), where: is the true melting temperature of the kth protein, is the highest melting temperature value in the current data set, is the minimum melting temperature; In the three-classification, the protein temperature category C is divided according to formula (2), with 0, 1, and 2 representing the numerical temperature category respectively; In the binary classification, proteins with a melting temperature ≥55°C were classified as thermophilic proteins, otherwise they were mesophilic proteins; (1) (2) Then the feature representation vector X of the i-th residue in the corresponding sequence is i , the contact map matrix is ​​input into the sequence feature extraction module and the structural feature extraction module respectively; The deep learning model in step 1.3 includes a position encoding module, a sequence feature extraction module, a structural feature extraction module, and an output module; In the position coding module, sine function or cosine function is used for coding, and the coding is shown in formula (3) and formula (4); (3) (4) in, Indicates the position of the amino acid in the protein sequence, 0≤ <L, D means feature dimension D=1024, and Indicates the characteristic component position of the residue at the current position, 0≤ <D / 2; Then, the residue representation input to the dual-channel network module is added with position information. The specific operation is as shown in formula (5): (5) The sequence feature extraction module consists of a multi-head attention network, a multi-scale convolutional neural network, a bidirectional long short-term memory network, and a forward attention network. The input data of the sequence feature extraction module is ; The multi-head attention network uses 8 attention heads; in the multi-head attention network, its inputs each pass through a layer of forward neural network to obtain Q i , K i 、V i , as shown in formula (6), the input data is 1024-dimensional and the output data is 128-dimensional; And an attention head i Calculation is as in formula (7) and formula (8), where d k value 128; The "Concat" operation in formula (9) indicates that the data is merged in the last dimension. The eight attention heads are concatenated in the last dimension and then output after passing through a layer of feedforward neural network. The output dimension is 1024 and the length of the protein sequence remains unchanged. (6) (7) (8) (9) The multi-scale convolutional neural network uses one-dimensional convolution and multiple convolutional neural networks with different convolution kernel sizes to perform convolution operations in parallel; the input data comes from the output data of the multi-head attention network, the input data is 1024-dimensional, the output data is 128-dimensional, and the convolution kernel sizes are [3, 5, 7, 9] respectively; the default activation function of the convolution operation is ReLU; the outputs of multiple convolution units are merged, as shown in formula (11): the output of the multi-scale convolutional neural network is 512-dimensional; (10) (11) The bidirectional long short-term memory network is used to focus on the global information of proteins and capture the long-range dependencies of protein sequences. It is composed of a bidirectional long short-term memory network LSTM. The unidirectional LSTM model is described in the formula as shown in (12). (13) Among them, σ is the activation function using the Sigmoid function; ⊙ represents the matrix bitwise multiplication; is the network input at time t; 、 、 、 and Represent the input gate, forget gate, output gate, internal memory unit and output at time t respectively; is the output of the previous moment; is the output of the internal memory unit at the previous moment; the rest are the learnable parameters of the neural network; The unidirectional LSTM network has a 1024-dimensional input and a 256-dimensional output. When the forward LSTM and backward LSTM data converge, a merge operation is performed at the lowest dimension, as shown in formula (13), and the output is 512-dimensional. The forward attention network is used to convert the two-dimensional matrix into a one-dimensional vector. is the output of the previous layer network, representing the tth amino acid residue in the sequence; Through a layer of forward neural network as shown in formula (14), the attention weight of the current feature representation of residue t in the sequence is obtained by the formula; the feature representations of all residues in the sequence are weighted and summed, as shown in formula (16), to achieve the conversion of two-dimensional feature representation to one-dimensional; here, L is 1024; the input data dimension is L*512, and the output dimension is 1*512; (14) (15) (16) The structural feature extraction module includes a graph attention convolutional network, a graph diffusion network, and a graph average pooling network, wherein: The spatial structure of proteins in the graph attention convolutional network is represented by a graph, where nodes are amino acid residues and the edges between nodes represent the contact between two residues. The node representation is the sequence feature information of formula (5), which is input Input1; the edge information is described by the contact map matrix, which is input Input2. The network nodes are connected to the forward network layer, as shown in formula (17), with an input dimension of 1024 and an output dimension of 256; For the neighboring node p of the current node z, data is concatenated and then multiplied with the learnable vector a to compress it into a scalar with a dimension of 256*2,1, and then passed through the LeakyReLU activation function, as shown in formula (19); Calculate the attention score between the current node z and the neighbor node p , the calculation formula uses the Softmax function, such as formula (20), where Represents the set of neighbor nodes of node z, and then updates the information of node z, as shown in formula (21), the activation function Use LeakyReLU function with output dimension 256; (17) (18) (19) (20) (21) The graph diffusion network obtains the adjacent node information of the adjacent nodes of a node through diffusion convolution; The original protein-protein contact matrix A has a dimension of 1024*1024, I is the identity matrix with a diagonal of 1, and the dimension is the same as A. F is the degree matrix of matrix A, and F is the degree matrix of matrix A. -1 is the normalized matrix of F, and the matrix A is transformed by formula (22-23) to become , That is, the node information output by the graph attention convolutional network, with a dimension of 256; is the node feature after iteration v times, is a learnable balance parameter with an initial value of 0.9; W is a parameter matrix; The diffusion operation of node information acquisition is realized by formula (24) and formula (25). This operation is repeated 5 times, expecting to obtain the neighboring point information of 5 levels. The output node dimension of this network module is still 256. (22) (23) (24) (25) The output node information of the graph diffusion network and the original protein contact matrix are sent back to the graph attention convolutional network unit, and the above operation is repeated again. At the same time, the node information is also sent to the graph average pooling network. The results of repeated iterations are also sent to the graph average pooling network. The graph average pooling network is used to adopt a graph global pooling strategy for the node information in the graph. The current feature of the i-th residue in the sequence is represented by X i , take the mean of all features in the sequence, as shown in formula (26), and compress the L*256 matrix into a 1*256 vector through graph global pooling; (26) Perform feature concatenation on the two output results of the graph average pooling unit to obtain a 1*512 dimension vector; It is necessary to perform feature concatenation on the output F1 of the sequence feature extraction network and the output F2 of the structure feature extraction network, as shown in formula (27), with an output dimension of 1*1024; (27) Then connect two layers of feedforward neural networks, as shown in formula (28) and formula (29), respectively. The first feedforward neural network uses tanh as activation function, and the output dimension is 256; the second one uses Sigmoid function or Softmax function, and the output dimension is 1 or 3; (28) (29) In the output module: The error loss between the model prediction value and the true value is described by formula (30) or formula (31), where (30) is the melting temperature prediction loss and formula (31) is the two-class or three-class prediction loss; (30) ((31) In the above formula, N is the number of training samples and i is the i-th protein.

2. The protein melting temperature prediction method using a hybrid deep learning strategy according to claim 1, wherein: The specific operation of step 1.1 is as follows: the training and test data are from the Meltome Atlas dataset. Proteins with a sequence length of less than 1024 are extracted from the Meltome Atlas dataset, and sequences with a sequence similarity greater than 30% are removed using CD-HIT software. These sequences are then divided into a training set, a validation set, and a test set in an 8:1:1 ratio using a fixed random seed.

Citation Information

Patent Citations

  • Protein sequence classification method based on hierarchical attention network

    CN111402953A

  • Protein function prediction method and device, storage medium and electronic equipment

    CN116343903A

  • Protein thermal stability prediction method, device, equipment and medium

    CN117672381A

  • Protein solubility prediction method

    CN118696380A