Anti-parasitic drug target affinity prediction method based on multi-modal feature fusion
Through the multimodal feature fusion method, combining the sequence and structural information of drugs and proteins, comparative learning strategies are used to strengthen information interaction, and an anti-parasitic drug target affinity prediction model is constructed, which solves the problem of insufficient information in the existing technology and achieves more efficient drug-target affinity prediction.
Patent Information
- Application Number
- CN202510646772.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
The existing drug-target affinity prediction methods are mainly modeled through single modal feature, resulting in limited information and it is difficult to accurately predict the interaction between anti-parasitic drugs and proteins. In addition, traditional experimental screening methods are time-consuming and costly, making it difficult to meet the needs of large-scale drug screening.
The multimodal feature fusion method is adopted, and a multi-scale convolutional neural network, graph convolutional neural network and adaptive gating network are combined with the sequence and structural characteristics of drugs and proteins, and the information interaction is strengthened using a comparison learning strategy to construct an anti-parasitic drug target affinity prediction model.
It improves the accuracy of drug-target affinity prediction and the generalization ability of the model, enhances the characteristic representation of drug-protein interactions, and improves the efficiency and accuracy of anti-parasitic drug discovery.
Smart Images

Figure CN120496626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of antiparasitic drug target affinity prediction, and in particular to an antiparasitic drug target affinity prediction method based on multimodal feature fusion. Background Art
[0002] Parasitic diseases, primarily caused by infections with protozoa and helminths, result in millions of deaths each year, pose a serious threat to global public health, and have a significant impact on socioeconomic development. Currently, medication remains the primary approach for treating and preventing parasitic diseases. However, with the increasing prevalence of drug resistance in parasites, the therapeutic efficacy of existing drugs is gradually declining, and some drugs are also associated with significant side effects, further limiting their clinical application. Therefore, the development of safe and effective antiparasitic drugs has become an urgent issue that needs to be addressed.
[0003] In the field of drug research and development, drug repositioning has attracted widespread attention due to its advantages such as short development cycle and low cost. It provides new ideas for the development of antiparasitic drugs by exploring the potential activity of existing clinical drugs against new targets. However, although traditional experimental screening methods can obtain relatively accurate affinity evaluation results, they are time-consuming and costly, making it difficult to meet the needs of large-scale drug screening. With the development of computer-assisted drug design, drug-target affinity prediction methods based on machine learning and deep learning have gradually become important tools for virtual drug screening. Early drug-target affinity prediction methods mainly modeled a single modality of drugs or proteins, resulting in limited drug or protein feature information learned by the model, thereby affecting the accuracy of the model's prediction.
[0004] In recent years, as drug and protein data have become increasingly diverse, more and more models are moving beyond single modal features and are attempting to integrate multiple modal features to improve their predictive performance. However, existing multimodal approaches still have limitations in cross-modal information interaction, leading to information imbalances between modalities and making it difficult to fully exploit the complementary information between modalities. Furthermore, many models still use relatively simple operations to model the feature interactions between drugs and proteins, making it difficult to fully exploit the interaction between the two, thus hindering the discovery of antiparasitic drugs. Summary of the Invention
[0005] The present invention aims to address the deficiencies of the above-mentioned prior art and proposes a method for predicting the affinity of antiparasitic drug targets based on multimodal feature fusion, in order to learn the rich feature representations of drugs and proteins and balance the information between different modalities of drugs and proteins, while focusing on the key information of drug features and protein features to enhance the interaction characteristics between drugs and proteins, thereby achieving more accurate affinity prediction between antiparasitic drugs and proteins.
[0006] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0007] The present invention is characterized by a method for predicting the affinity of antiparasitic drug targets based on multimodal feature fusion, which is carried out according to the following steps:
[0008] Step 1: Obtain drug sample dataset ,make Any drug sample in is denoted as d; Get protein sample data set ;make Any protein sample in is denoted as p;
[0009] Step 2: Construct a sequence feature extractor of a multi-scale convolutional neural network, including: input module, multi-scale convolution module and output module, and Drug samples in d and The protein sample p in is processed and the corresponding drug sequence features are obtained and protein sequence features ;
[0010] Step 3: Construct a structural feature extractor of the graph convolutional neural network, including: M layers of graph convolution layers and the first graph pooling layer, and process d to obtain the updated drug structural features ;
[0011] Step 4: Construct a structural feature extractor of the graph sampling and aggregation network, including: T-layer graph sampling layer, T-layer graph aggregation layer and second graph pooling layer, and process p to obtain the updated protein structural features ;
[0012] Step 5: Construct the feature fusion module of the adaptive gating network and Processing is performed to obtain the interaction characteristics y of drug sample d-protein sample p dp ;
[0013] Step 6: Build the prediction layer and perform dp Perform linear processing to obtain the predicted affinity value between d and p ;
[0014] Step 7: Construct the predicted total loss function L, which includes: the contrast loss of d, the contrast loss of p, and the prediction loss between d and p;
[0015] Step 8: Train the antiparasitic drug target affinity prediction network through backpropagation and gradient descent, and calculate the total loss function To update the parameters of the entire network until the maximum number of iterations is reached or Until it converges to the minimum value, the optimal antiparasitic drug target affinity prediction model is obtained, which is used to predict the affinity of the input drug samples and protein samples.
[0016] The antiparasitic drug target affinity prediction method based on multimodal feature fusion described in the present invention is also characterized in that step 2 is performed as follows:
[0017] Step 2.1: The input module processes d using a pre-trained chemical language model to obtain the initial drug feature vector of d. , where m1 represents the feature dimension of d;
[0018] The input module processes p using a pre-trained protein language model to obtain the initial protein feature vector of p , where m2 represents the feature dimension of p;
[0019] Step 2.2, the multi-scale convolution module consists of N layers of convolutional layers, batch normalization layers and linear normalization layers, and and Processing is performed to obtain the drug convolution feature vector of the Nth layer. and protein convolution feature vector ;
[0020] Step 2.3: The output module uses global pooling operation to and Processing is performed to obtain the final drug sequence characteristics and final protein sequence features .
[0021] Furthermore, step 2.2 is performed as follows:
[0022] Step 2.2.1, when n=1, initialize the drug feature vector of the n-1th layer , protein feature vector of the n-1th layer ;
[0023] Step 2.2.2: Use formula (1) and formula (2) to obtain the drug feature vector of the nth layer and protein feature vector :
[0024] (1)
[0025] (2)
[0026] In formula (1) and formula (2), σ is the ReLU activation function, Represents the nth convolutional layer respectively The weight parameters and bias parameters of ;
[0027] Step 2.2.3, the batch normalization layer and linear normalization layer of the nth layer are and Processing is performed to obtain the drug convolution feature vector of the nth layer. and protein convolution feature vector ;
[0028] Step 2.2.4, Assign to ,Will Assign to , after assigning n+1 to n, return to step 2.2.2 and execute sequentially until n>N, thus obtaining the drug convolution feature vector of the Nth layer and protein convolution feature vector .
[0029] Furthermore, step 3 is performed as follows:
[0030] Step 3.1: Represent d as a drug graph consisting of nodes and edges ,in, Drug diagram The feature matrix of The number of nodes in the drug graph, R1 represents the characteristic dimension of the drug graph node; Drug diagram The adjacency matrix of
[0031] Step 3.2: When m=1, initialize the drug graph features of the m-1th layer The graph convolution layer of the mth layer uses formula (3) to obtain the drug graph features of the mth layer , thus obtaining the drug graph features of the Mth layer ;
[0032] (3)
[0033] In formula (3), is the adjacency matrix with self-loop added, is the degree matrix, is the weight parameter to be learned in the m-th graph convolution layer, and σ represents the ReLU activation function;
[0034] Step 3.3: The first graph pooling layer is Processing to obtain updated drug structural characteristics .
[0035] Furthermore, step 4 is performed as follows:
[0036] Step 4.1: Represent p as a protein graph consisting of nodes and edges ,in, Representing protein graphs The feature matrix of The number of nodes, R2 represents the characteristic dimension of the protein graph node, Representing protein graphs The adjacency matrix of
[0037] Step 4.2: Protein Map Perform feature update to obtain updated protein structure features .
[0038] Furthermore, step 4.2 is performed as follows:
[0039] Step 4.2.1. When t=1, initialize the protein graph features of the t-1 layer ,initialization Any node Features at the t-1th graph aggregation layer ;in, express Any node characteristics;
[0040] Step 4.2.2, the t-1th layer graph sampling layer uses formula (4) to Any node The feature sampling of a neighbor node u is performed, and the features of the sampled neighbor nodes are aggregated to obtain the sampling node of the t-1 layer. Neighbor node characteristics :
[0041] (4)
[0042] In formula (4), express The set of neighbor nodes of Represents the sampling node of the t-1th layer graph sampling layer The characteristics of neighbor node u, Represents an aggregate function;
[0043] Step 4.2.3: The graph aggregation layer of the tth layer uses formula (5) to obtain the protein graph features of the tth layer , thus obtaining the protein graph features of the Tth layer :
[0044] (5)
[0045] In formula (5), is the weight parameter to be learned for the t-th graph aggregation layer, and Cat represents the concatenation operation;
[0046] Step 4.2.3, the second pooling layer Processing to obtain updated protein structural features .
[0047] Furthermore, step 5 is performed as follows:
[0048] Step 5.1: Use formula (6) to obtain the joint representation Z of drug sample d and protein sample p dp :
[0049] (6)
[0050] Step 5.2: Calculate Z using formula (7) dp Weight score :
[0051] (7)
[0052] In formula (7), Represent the weight parameters of the first and second linear layers in the adaptive gating network, Represent the bias parameters of the first and second linear layers respectively;
[0053] Step 5.3: Use formula (8) to obtain Z dp Attention score ;
[0054] (8)
[0055] In formula (8), sigmoid represents the activation function;
[0056] Step 5.4: Use formula (9) to obtain the interaction characteristics y between drug sample d and protein sample p dp :
[0057] (9)
[0058] In formula (9), They represent the weight parameters and bias parameters of the third linear layer in the adaptive gating network respectively.
[0059] Furthermore, step 7 is performed as follows:
[0060] Step 7.1, based on 、 、 and , construct the contrast loss function of drug sample d and the contrast loss of protein sample p ;
[0061] Step 7.1.1: Use Equation (10) to construct the contrast loss of d :
[0062] (10)
[0063] In formula (10), represents cosine similarity, τ is a parameter used to adjust the feature distribution, represents any drug sample other than d in D, and \ represents the removal operation; express The drug sequence characteristics, express Updated drug structure characteristics, exp represents exponential function, log represents logarithmic function;
[0064] Step 7.1.2: Use formula (11) to construct the contrast loss function of p :
[0065] (11)
[0066] In formula (11), Indicates any protein sample other than p in P, express The protein sequence characteristics of express updated protein structural features;
[0067] Step 7.2: Use Equation (12) to construct the prediction loss between d and p :
[0068] (12)
[0069] In formula (12), represents the true affinity value between d and p proteins;
[0070] Step 7.3: Use formula (13) to construct the total loss function L of the antiparasitic drug target affinity prediction network, which is composed of a sequence feature extractor of a multi-scale convolutional neural network, a structural feature extractor of a graph convolutional neural network, a structural feature extractor of a graph sampling and aggregation network, a feature fusion module of an adaptive gating network, and a prediction layer:
[0071] (13).
[0072] The electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the antiparasitic drug target affinity prediction method, and the processor is configured to execute the program stored in the memory.
[0073] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the antiparasitic drug target affinity prediction method when the computer program is run by a processor.
[0074] Compared with the prior art, the present invention has the following beneficial effects:
[0075] 1. The present invention designs a method for predicting antiparasitic drug target affinity based on multimodal feature fusion. By making full use of the sequence and structural information of drugs and proteins, the accuracy of drug-target affinity prediction is improved, effectively solving the problem of insufficient drug-target training data, and improving the accuracy of antiparasitic drug-target affinity prediction with limited sample data.
[0076] 2. By using a pre-trained model and designing a multi-scale convolutional neural network to encode drug sequences and protein sequences, the present invention can obtain high-quality drug sequence features and protein sequence features, which is beneficial to improving the prediction performance of downstream tasks and enhancing the generalization ability of the model.
[0077] 3. The present invention designs a convolutional neural network and a sampling and aggregation network to effectively obtain the structural information of drugs and proteins, enrich the feature representation of drugs and proteins, enable the model to capture a variety of biological information, and thus improve the predictive performance of the model.
[0078] 4. The present invention introduces a comparative learning strategy to strengthen the information interaction between different modalities of drugs and proteins, learn the consistency between modalities, enhance the robustness of the model in feature representation, and ensure that the model can maintain stable performance in complex biological scenarios.
[0079] 5. The present invention constructs an adaptive gating network, which enables the model to focus on the key features of different modalities, enhances the representation ability of drug-protein interactions, and further improves the predictive performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is the overall framework diagram of the method of the present invention;
[0081] Figure 2 This is a structural diagram of the multi-scale convolutional neural network of the present invention;
[0082] Figure 3 This is a structural diagram of the adaptive gating network of the present invention. DETAILED DESCRIPTION
[0083] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0084] In this example, a method for predicting the affinity of antiparasitic drug targets based on multimodal feature fusion is proposed. Figure 1 , including the following steps:
[0085] Step 1: Obtain drug sample dataset ,make Any drug sample in is denoted as d; Get protein sample data set ;make Any protein sample in is denoted as p;
[0086] Step 2: Construct a sequence feature extractor of a multi-scale convolutional neural network, including: input module, multi-scale convolution module and output module, and Drug samples in d and The protein sample p in is processed and the corresponding drug sequence features are obtained and protein sequence features ;
[0087] Step 2.1: The input module uses the pre-trained chemical language model to perform initial encoding on the smiles sequence of d to obtain the initial drug feature vector of d. , where m1 represents the feature dimension of d; in the specific example, m1=384;
[0088] The input module uses the pre-trained protein language model to initially encode the amino acid sequence of p and obtain the initial protein feature vector of p. , where m2 represents the feature dimension of p; in the specific example, m2=1280.
[0089] Step 2.2, the multi-scale convolution module consists of N layers of convolutional layers, batch normalization layers and linear normalization layers, and and Processing is performed to obtain the drug convolution feature vector of the Nth layer. and protein convolution feature vector ;
[0090] Step 2.2.1, when n=1, initialize the drug feature vector of the n-1th layer , protein feature vector of the n-1th layer ;like Figure 2 As shown, the multi-scale convolution module in this embodiment uses three layers of convolution operations, and each layer uses convolution kernels of sizes 3, 5, and 7 respectively.
[0091] Step 2.2.2: Use formula (1) and formula (2) to obtain the drug feature vector of the nth layer and protein feature vector :
[0092] (1)
[0093] (2)
[0094] In formula (1) and formula (2), σ is the ReLU activation function, Represents the nth convolutional layer respectively The weight parameters and bias parameters of ;
[0095] Step 2.2.3, the batch normalization layer and linear normalization layer of the nth layer are and Processing is performed to obtain the drug convolution feature vector of the nth layer. and protein convolution feature vector ;
[0096] Step 2.2.4, Assign to ,Will Assign to , after assigning n+1 to n, return to step 2.2.2 and execute sequentially until n>N, thus obtaining the drug convolution feature vector of the Nth layer and protein convolution feature vector .
[0097] Step 2.3: The output module uses global pooling operation to and Processing is performed to obtain the final drug sequence characteristics and final protein sequence features ;
[0098] Step 3: Construct a structural feature extractor of the graph convolutional neural network, including: M layers of graph convolution layers and the first graph pooling layer, and process d to obtain the updated drug structural features , in this specific example, M=3:
[0099] Step 3.1: Represent d as a drug graph consisting of nodes and edges ,in, Drug diagram The feature matrix of The number of nodes in the drug graph, R1 represents the characteristic dimension of the drug graph node; Drug diagram The adjacency matrix of
[0100] Step 3.2: When m=1, initialize the drug graph features of the m-1th layer The graph convolution layer of the mth layer uses formula (3) to obtain the drug graph features of the mth layer , thus obtaining the drug graph features of the Mth layer ;
[0101] (3)
[0102] In formula (3), is the adjacency matrix with self-loops added, is the degree matrix, through the drug graph The adjacent node counts of each node in are obtained, is the weight parameter to be learned in the m-th graph convolution layer, and σ represents the ReLU activation function;
[0103] Step 3.3, the first image pooling layer Processing to obtain updated drug structural characteristics .
[0104] Step 4: Construct a structural feature extractor of the graph sampling and aggregation network, including: T-layer graph sampling layer, T-layer graph aggregation layer and second graph pooling layer, and process p to obtain the updated protein structural features , in the specific example, T=3:
[0105] Step 4.1: Represent p as a protein graph consisting of nodes and edges ,in, Representing protein graphs The feature matrix of The number of nodes, R2 represents the characteristic dimension of the protein graph node, Representing protein graphs The adjacency matrix of .
[0106] Step 4.2: Protein Map Perform feature update to obtain updated protein structure features :
[0107] Step 4.2.1. When t=1, initialize the protein graph features of the t-1 layer ,initialization Any node Features at the t-1th graph aggregation layer ;in, express Any node characteristics;
[0108] Step 4.2.2, the t-1th layer graph sampling layer uses formula (4) to Any node The feature sampling of a neighbor node u is performed, and the features of the sampled neighbor nodes are aggregated to obtain the sampling node of the t-1 layer. Neighbor node characteristics :
[0109] (4)
[0110] In formula (4), express The set of neighbor nodes of Represents the sampling node of the t-1th layer graph sampling layer The characteristics of neighbor node u, Represents an aggregate function.
[0111] Step 4.2.3: The graph aggregation layer of the tth layer uses formula (5) to obtain the protein graph features of the tth layer , thus obtaining the protein graph features of the Tth layer :
[0112] (5)
[0113] In formula (5), is the weight parameter to be learned for the t-th graph aggregation layer, and Cat represents the concatenation operation;
[0114] Step 4.2.3, the second pooling layer Processing to obtain updated protein structural features .
[0115] Step 5: Construct the feature fusion module of the adaptive gating network and Processing is performed to obtain the interaction characteristics y of drug sample d-protein sample p dp ;like Figure 3As shown in the figure, based on the multiple modal features of drugs and proteins extracted above, namely sequence features and structural features, each feature is mapped into a 128-dimensional feature vector; this module enhances the representation of drug-protein interaction features by focusing on the key information between the different modal features of drugs and proteins;
[0116] Step 5.1: Use formula (6) to obtain the joint representation Z of drug sample d and protein sample p dp :
[0117] (6)
[0118] Step 5.2: Calculate Z using formula (7) dp Weight score :
[0119] (7)
[0120] In formula (7), Represent the weight parameters of the first and second linear layers in the adaptive gating network, They represent the bias parameters of the first and second linear layers respectively.
[0121] Step 5.3: Use formula (8) to obtain Z dp Attention score ;
[0122] (8)
[0123] In formula (8), sigmoid represents the activation function.
[0124] Step 5.4: Use formula (9) to obtain the interaction characteristics y between drug sample d and protein sample p dp :
[0125] (9)
[0126] In formula (9), They represent the weight parameters and bias parameters of the third linear layer in the adaptive gating network respectively.
[0127] Step 6: Build the prediction layer and perform dp Perform linear processing to obtain the predicted affinity value between d and p .
[0128] Step 7: Construct the predicted loss function, including: contrast loss of d, contrast loss of p, and prediction loss between d and p;
[0129] Step 7.1, based on 、 、 and , construct the contrast loss function of drug sample d and the contrast loss of protein sample p , strengthen the information interaction between different modal features of drugs and proteins, and learn the consistency between modalities;
[0130] Step 7.1.1: Use Equation (10) to construct the contrast loss of d :
[0131] (10)
[0132] In formula (10), represents cosine similarity, τ is a parameter used to adjust the feature distribution, represents any drug sample other than d in D, and \ represents the removal operation; express The drug sequence characteristics, express Updated drug structure features, exp represents exponential function, and log represents logarithmic function.
[0133] Step 7.1.2: Use formula (11) to construct the contrast loss function of p :
[0134] (11)
[0135] In formula (11), Indicates any protein sample other than p in P, express The protein sequence characteristics of express Updated protein structural features.
[0136] Step 7.2: Use Equation (12) to construct the prediction loss between d and p :
[0137] (12)
[0138] In formula (12), represents the true affinity value between d and p proteins;
[0139] Step 7.3: Use formula (13) to construct the total loss function L of the antiparasitic drug target affinity prediction network, which is composed of a sequence feature extractor of a multi-scale convolutional neural network, a structural feature extractor of a graph convolutional neural network, a structural feature extractor of a graph sampling and aggregation network, a feature fusion module of an adaptive gating network, and a prediction layer:
[0140] (13)
[0141] Step 8: Train the antiparasitic drug target affinity prediction network through backpropagation and gradient descent, and calculate the total loss function To update the parameters of the entire network until the maximum number of iterations is reached or Until it converges to the minimum value, the optimal antiparasitic drug target affinity prediction model is obtained, which is used to predict the affinity of the input drug samples and protein samples.
[0142] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0143] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
Claims
1. A method for predicting antiparasitic drug target affinity based on multimodal feature fusion, characterized in that: The steps are as follows: Step 1: Obtain drug sample dataset ,make Any drug sample in is denoted as d; Get protein sample data set ;make Any protein sample in is denoted as p; Step 2: Construct a sequence feature extractor of a multi-scale convolutional neural network, including: input module, multi-scale convolution module and output module, and Drug samples in d and The protein sample p in is processed and the corresponding drug sequence features are obtained and protein sequence features ; Step 3: Construct a structural feature extractor of the graph convolutional neural network, including: M layers of graph convolution layers and the first graph pooling layer, and process d to obtain the updated drug structural features ; Step 4: Construct a structural feature extractor of the graph sampling and aggregation network, including: T-layer graph sampling layer, T-layer graph aggregation layer and second graph pooling layer, and process p to obtain the updated protein structural features ; Step 5: Construct the feature fusion module of the adaptive gating network and Processing is performed to obtain the interaction characteristics y of drug sample d-protein sample p dp ; Step 6: Build the prediction layer and perform dp Perform linear processing to obtain the predicted affinity value between d and p ; Step 7: Construct the predicted total loss function L, which includes: the contrast loss of d, the contrast loss of p, and the prediction loss between d and p; Step 8: Train the antiparasitic drug target affinity prediction network through backpropagation and gradient descent, and calculate the total loss function To update the parameters of the entire network until the maximum number of iterations is reached or Until it converges to the minimum value, the optimal antiparasitic drug target affinity prediction model is obtained, which is used to predict the affinity of the input drug samples and protein samples.
2. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 1, characterized in that: Step 2 is performed as follows: Step 2.1: The input module processes d using a pre-trained chemical language model to obtain the initial drug feature vector of d. , where m1 represents the feature dimension of d; The input module processes p using a pre-trained protein language model to obtain the initial protein feature vector of p , where m2 represents the feature dimension of p; Step 2.2, the multi-scale convolution module consists of N layers of convolutional layers, batch normalization layers and linear normalization layers, and and Processing is performed to obtain the drug convolution feature vector of the Nth layer. and protein convolution feature vector ; Step 2.3: The output module uses global pooling operation to and Processing is performed to obtain the final drug sequence characteristics and final protein sequence features .
3. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 2, characterized in that: Step 2.2 is performed as follows: Step 2.2.1, when n=1, initialize the drug feature vector of the n-1th layer , protein feature vector of the n-1th layer ; Step 2.2.2: Use formula (1) and formula (2) to obtain the drug feature vector of the nth layer and protein feature vector : (1) (2) In formula (1) and formula (2), σ is the ReLU activation function, Represents the nth convolutional layer respectively The weight parameters and bias parameters of ; Step 2.2.3, the batch normalization layer and linear normalization layer of the nth layer are and Processing is performed to obtain the drug convolution feature vector of the nth layer. and protein convolution feature vector ; Step 2.2.4, Assign to ,Will Assign to , after assigning n+1 to n, return to step 2.2.2 and execute sequentially until n>N, thus obtaining the drug convolution feature vector of the Nth layer and protein convolution feature vector .
4. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 3, characterized in that: Step 3 is performed as follows: Step 3.1: Represent d as a drug graph consisting of nodes and edges ,in, Drug diagram The feature matrix of The number of nodes in the drug graph, R1 represents the characteristic dimension of the drug graph node; Drug diagram The adjacency matrix of Step 3.2: When m=1, initialize the drug graph features of the m-1th layer The graph convolution layer of the mth layer uses formula (3) to obtain the drug graph features of the mth layer , thus obtaining the drug graph features of the Mth layer ; (3) In formula (3), is the adjacency matrix with self-loop added, is the degree matrix, is the weight parameter to be learned in the m-th graph convolution layer, and σ represents the ReLU activation function; Step 3.3: The first graph pooling layer is Processing to obtain updated drug structural characteristics .
5. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 4, characterized in that: Step 4 is performed as follows: Step 4.1: Represent p as a protein graph consisting of nodes and edges ,in, Representing protein graphs The feature matrix of The number of nodes, R2 represents the characteristic dimension of the protein graph node, Representing protein graphs The adjacency matrix of Step 4.2: Protein Map Perform feature update to obtain updated protein structure features .
6. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 5, characterized in that: Step 4.2 is performed as follows: Step 4.2.
1. When t=1, initialize the protein graph features of the t-1 layer ,initialization Any node Features at the t-1th graph aggregation layer ;in, express Any node characteristics; Step 4.2.2, the t-1th layer graph sampling layer uses formula (4) to Any node The feature sampling of a neighbor node u is performed, and the features of the sampled neighbor nodes are aggregated to obtain the sampling node of the t-1 layer. Neighbor node characteristics : (4) In formula (4), express The set of neighbor nodes of Represents the sampling node of the t-1th layer graph sampling layer The characteristics of neighbor node u, Represents an aggregate function; Step 4.2.3: The graph aggregation layer of the tth layer uses formula (5) to obtain the protein graph features of the tth layer , thus obtaining the protein graph features of the Tth layer : (5) In formula (5), is the weight parameter to be learned for the t-th graph aggregation layer, and Cat represents the concatenation operation; Step 4.2.3, the second pooling layer Processing to obtain updated protein structural features .
7. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 6, characterized in that: Step 5 is performed as follows: Step 5.1: Use formula (6) to obtain the joint representation Z of drug sample d and protein sample p dp : (6) Step 5.2: Calculate Z using formula (7) dp Weight score : (7) In formula (7), Represent the weight parameters of the first and second linear layers in the adaptive gating network, Represent the bias parameters of the first and second linear layers respectively; Step 5.3: Use formula (8) to obtain Z dp Attention score ; (8) In formula (8), sigmoid represents the activation function; Step 5.4: Use formula (9) to obtain the interaction characteristics y between drug sample d and protein sample p dp : (9) In formula (9), They represent the weight parameters and bias parameters of the third linear layer in the adaptive gating network respectively.
8. The method for predicting antiparasitic drug target affinity based on multimodal feature fusion according to claim 7, characterized in that: Step 7 is performed as follows: Step 7.1, based on 、 、 and , construct the contrast loss function of drug sample d and the contrast loss of protein sample p ; Step 7.1.1: Use Equation (10) to construct the contrast loss of d : (10) In formula (10), represents cosine similarity, τ is a parameter used to adjust the feature distribution, represents any drug sample other than d in D, and \ represents the removal operation; express The drug sequence characteristics, express Updated drug structure characteristics, exp represents exponential function, log represents logarithmic function; Step 7.1.2: Use formula (11) to construct the contrast loss function of p : (11) In formula (11), Indicates any protein sample other than p in P, express The protein sequence characteristics of express updated protein structural features; Step 7.2: Use Equation (12) to construct the prediction loss between d and p : (12) In formula (12), represents the true affinity value between d and p proteins; Step 7.3: Use formula (13) to construct the total loss function L of the antiparasitic drug target affinity prediction network, which is composed of a sequence feature extractor of a multi-scale convolutional neural network, a structural feature extractor of a graph convolutional neural network, a structural feature extractor of a graph sampling and aggregation network, a feature fusion module of an adaptive gating network, and a prediction layer: (13)。 9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the anti-parasitic drug target affinity prediction method according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the antiparasitic drug target affinity prediction method according to any one of claims 1 to 8 are performed.
Citation Information
Cited By
Molecular-protein affinity prediction method, system, equipment and medium
CN121709080A
A method, system, device and medium for predicting molecular-protein affinity
CN121709080B