Drug target affinity prediction method and model

By extracting molecular-scale feature of drugs and target proteins and constructing isomerographic networks, combined with cross-scale feature fusion method, the problem of existing models ignoring the two-dimensional structure and network-scale features of target proteins is solved, and the accuracy and robustness of drug target affinity prediction is improved.

CN119964680AInactive Publication Date: 2025-05-09EAST CHINA UNIV OF SCI & TECH

Patent Information

Application Number
CN202510020261.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing drug target affinity prediction models ignore the two-dimensional structural characteristics of the target protein and the network-scale characteristics in multi-source biological networks, resulting in limitations in feature learning.

Method used

By extracting the molecular-scale feature of drugs and target proteins, a isomerographic network of drug-target interaction relationships is constructed, and a cross-scale feature fusion method is used to fuse molecular-scale and network-scale features to generate drug-target joint feature representations to predict affinity.

Benefits of technology

It improves the affinity prediction performance of drug targets, enhances the robustness and generalization ability of the model, and can more accurately predict the binding affinity of drugs and target proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964680A_ABST
    Figure CN119964680A_ABST
Patent Text Reader

Abstract

The invention provides a drug target affinity prediction method and model, and the prediction method comprises the steps: carrying out the feature extraction of a drug and a target protein, and obtaining a drug molecular scale feature representation d and a target protein molecular scale feature representation p; the method comprises the following steps: constructing a drug-target interaction relationship heterogeneous graph network according to drug and target protein molecular scale feature representation, the drug-target interaction relationship heterogeneous graph network comprising network scale features; and fusing the drug molecule scale feature representation d, the target protein molecule scale feature representation p and the network scale feature by using a cross-scale feature fusion method, and generating a drug-target combined feature representation for predicting a final affinity score. According to the method, the model shows superiority in improving the drug target affinity prediction performance, and meanwhile, the robustness and generalization ability are reflected on different data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a drug target affinity prediction method and model. Background Art

[0002] Drug development is a time-consuming and costly process. According to statistics, it usually takes 10 to 15 years and billions of dollars to develop a new drug, but more than 90% of candidate drugs fail in clinical trials and are ultimately not approved for marketing. How to reduce costs and increase success rates has become a key issue that needs to be addressed in the field of drug development.

[0003] Drug-Target Affinity (DTA) prediction is one of the core links in drug development. It uses computational methods to evaluate the tightness of the binding between drugs and target proteins. This process not only helps researchers to deeply analyze the mechanism of action of drugs, but also can significantly reduce research and development costs and improve the success rate of new drug development. Therefore, accurate prediction of drug-target affinity plays an important role in the drug development process and has attracted widespread attention in academia and industry in recent years.

[0004] After searching the literature on the prior art, it was found that Nguyen et al. mentioned a drug-target affinity prediction model based on graph neural networks in "GraphDTA: predicting drug-target binding affinity with graph neural networks". Different variant models of graph neural networks were used to extract drug topological map features to predict the binding affinity between drugs and target proteins. However, this model represents the target protein as a sequence as the input of the model, ignoring the two-dimensional structural characteristics of the target protein. In addition, the target protein is a complex biological macromolecule, usually composed of hundreds to thousands of amino acids, with complex sequence and structural characteristics. Relying only on a single modal feature may lead to the loss of key information.

[0005] In addition, He et al. proposed a node-adaptive hybrid graph neural network for interpretable drug-target binding affinity prediction in NHGNN-DTA: a node-adaptive hybrid graph neural network for interpretable drug-target binding affinity prediction. It can adaptively obtain the characteristic representation of drugs and proteins and allow information to interact at the graph level, effectively combining the advantages of sequence-based and graph-based methods. However, this method only considers the molecular scale features limited to a single compound, ignoring the network scale feature information contained in the multi-source biological network, resulting in limitations in the learning of biological molecular features. Summary of the invention

[0006] In order to overcome at least one aspect of the above-mentioned defects and problems in the prior art, the present invention provides a method for predicting drug target affinity, comprising:

[0007] Extract features of drugs and target proteins to obtain drug molecular scale feature representation d and target protein molecular scale feature representation p;

[0008] According to the molecular scale feature representation of drugs and target proteins, a heterogeneous graph network of drug-target interaction relationship is constructed, and the heterogeneous graph network of drug-target interaction relationship includes network scale feature;

[0009] A cross-scale feature fusion method is used to fuse the drug molecule scale feature representation d, the target protein molecule scale feature representation p and the network scale features to generate a drug-target joint feature representation to predict the final affinity score.

[0010] Furthermore, in the extraction of drug molecular-scale features, the drug SMILES sequence is converted into a molecular graph, and feature extraction is performed through a multi-layer graph convolutional neural network.

[0011] Furthermore, in the molecular scale feature extraction of the target protein, the target protein is represented as two modes, namely, amino acid sequence and residue graph, for feature extraction, thereby enriching the molecular scale feature representation of the target protein.

[0012] Furthermore, in the drug molecule scale feature extraction, the drug molecule graph is input into a three-layer graph convolutional neural network module to learn the feature representation of the drug molecule graph. The graph convolution formula is as follows:

[0013]

[0014] Among them, A d is the drug molecule adjacency matrix, I is the identity matrix, which is used to add self-connection, and Z is the degree matrix, which means that its diagonal elements are the degrees of the nodes. represents the first layer node feature matrix, W (l) is the learnable weight matrix of the first layer, and σ(·) represents the ReLU activation function.

[0015] Furthermore, through the graph convolution operation, the features of each node and its neighboring nodes are weighted and aggregated according to the graph structure, and the updated node representation is generated using nonlinear transformation, thereby effectively capturing local and global graph features. After the last layer of graph convolutional neural network, a readout block including a global average pooling layer and a fully connected layer is designed to obtain the final drug molecule-scale feature representation d, which is calculated as follows:

[0016]

[0017] in, represents the last layer of drug molecule node embedding, N d is the number of atoms in the drug molecule.

[0018] Furthermore, in the molecular scale feature extraction of the target protein, the target protein is represented as two modes, amino acid sequence and residue graph, for feature extraction, and a neural network module including 1D-CNN and BiLSTM is used to extract features from the sequence mode;

[0019] The target protein sequence is input into a two-layer 1D-CNN network to extract local features. The process expression is as follows:

[0020]

[0021] The local features extracted by 1D-CNN are used as the input of the BiLSTM recurrent neural network. The calculation process of BiLSTM includes the calculation of the forward and reverse LSTM layers. The process is as follows:

[0022]

[0023] Among them, H CNN represents the local features extracted by 1D-CNN, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the reverse LSTM at time step t, [·;·] represents the vector concatenation operation, and p seq As the final output, it contains the forward and backward dependency feature information of the target protein sequence.

[0024] Furthermore, for the target protein residue graph, a three-layer graph convolutional neural network layer is used to extract the modal feature information of the target protein graph, and a readout module including a global average pooling layer and a fully connected layer is used to obtain the target protein graph feature representation P graph , the calculation process is as follows:

[0025]

[0026] in, represents the graph convolution operation, X p Represents the target protein residue feature matrix, A p represents the target protein residue adjacency matrix;

[0027] After obtaining the feature vectors of the target protein sequence and graph respectively, the sequence feature vector p is concatenated using the cascade method. seq and graph feature vector p graph Connect them together to get the complete feature vector p of a single target protein, and the calculation process is as follows:

[0028] p=[p seq ∥p graph ].

[0029] Furthermore, the cross-scale feature fusion method includes:

[0030] The drug molecule scale feature vector d and the target protein molecule scale feature vector p are stacked into a matrix:

[0031] X DTI =[d1,…,d M ,p1,…,p N ] T

[0032] M is the total number of drugs, N is the total number of target proteins, and X DTI It is regarded as the initial feature matrix of the drug-target interaction heterogeneous graph network, and a two-layer graph convolutional neural network is used for feature extraction, fusing the molecular scale and drug-target interaction heterogeneous graph network scale features. The calculation process is as follows:

[0033]

[0034] Furthermore, after feature fusion, the fusion drugs are obtained respectively:

[0035] Drug characterization based on network-scale features of heterogeneous graphs of target interaction relationships and target protein feature vector The formula is as follows:

[0036]

[0037] After integrating the network scale features, the drug feature vector is cascaded. and target protein feature vector Combined to form a joint feature representation of drug-target, which is used as the input of the subsequent fully connected layer to finally obtain the affinity prediction value The formula for the full connection process is as follows:

[0038]

[0039] The present invention also provides a drug target affinity prediction model, which is obtained by training a neural network using the above method.

[0040] In summary, the technical solution of the present invention demonstrates superiority in improving the prediction performance of drug target affinity, and also demonstrates robustness and generalization ability on different data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Attached Figure 1 A flow chart of a method for predicting drug target affinity of the present invention;

[0042] Attached Figure 2 Experimental performance comparison of datasets DAVIS and KIBA. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] Example 1, a method for predicting drug target affinity.

[0046] This embodiment proposes a method for predicting drug target affinity, comprising:

[0047] Extract features of drugs and target proteins to obtain drug molecular scale feature representation d and target protein molecular scale feature representation p;

[0048] According to the molecular scale feature representation of drugs and target proteins, a heterogeneous graph network of drug-target interaction relationship is constructed, and the heterogeneous graph network of drug-target interaction relationship includes network scale feature;

[0049] A cross-scale feature fusion method is used to fuse the drug molecule scale feature representation d, the target protein molecule scale feature representation p and the network scale features to generate a drug-target joint feature representation to predict the final affinity score.

[0050] Preferably, in drug molecular scale feature extraction, the drug SMILES sequence is converted into a molecular graph, and feature extraction is performed through a multi-layer graph convolutional neural network.

[0051] Preferably, in the molecular scale feature extraction of the target protein, the target protein is represented as two modes, namely, an amino acid sequence and a residue graph, for feature extraction, thereby enriching the molecular scale feature representation of the target protein.

[0052] Preferably, in the drug molecule scale feature extraction, the drug molecule graph is input into a three-layer graph convolutional neural network module to learn the feature representation of the drug molecule graph. The graph convolution formula is as follows:

[0053]

[0054] Among them, A d is the drug molecule adjacency matrix, I is the identity matrix, which is used to add self-connection, and Z is the degree matrix, which means that its diagonal elements are the degrees of the nodes. represents the first layer node feature matrix, W (l) is the learnable weight matrix of the first layer, and σ(·) represents the ReLU activation function.

[0055] Preferably, through the graph convolution operation, the features of each node and the features of its neighboring nodes are weighted and aggregated according to the graph structure, and the updated node representation is generated by nonlinear transformation, so as to effectively capture local and global graph features. After the last layer of graph convolutional neural network, a readout block including a global average pooling layer and a fully connected layer is designed to obtain the final drug molecule scale feature representation d, which is calculated as follows:

[0056]

[0057] in, represents the last layer of drug molecule node embedding, N d is the number of atoms in the drug molecule.

[0058] Preferably, in the molecular scale feature extraction of the target protein, the target protein is represented as two modes, namely, an amino acid sequence and a residue graph, for feature extraction, and a neural network module including 1D-CNN and BiLSTM is used to extract features from the sequence mode;

[0059] The target protein sequence is input into a two-layer 1D-CNN network to extract local features. The process expression is as follows:

[0060]

[0061] The local features extracted by 1D-CNN are used as the input of the BiLSTM recurrent neural network. The calculation process of BiLSTM includes the calculation of the forward and reverse LSTM layers. The process is as follows:

[0062]

[0063] Among them, H CNN represents the local features extracted by 1D-CNN, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the reverse LSTM at time step t, [·;·] represents the vector concatenation operation, and p seq As the final output, it contains the forward and backward dependency feature information of the target protein sequence.

[0064] Preferably, for the target protein residue graph, a three-layer graph convolutional neural network layer is used to extract the target protein graph modal feature information, and a readout module including a global average pooling layer and a fully connected layer is used to obtain the target protein graph feature representation P graph , the calculation process is as follows:

[0065]

[0066] in, represents the graph convolution operation, X p Represents the target protein residue feature matrix, A p represents the target protein residue adjacency matrix;

[0067] After obtaining the feature vectors of the target protein sequence and graph respectively, the sequence feature vector p is concatenated using the cascade method. seq and graph feature vector p graph Connect them together to get the complete feature vector p of a single target protein, and the calculation process is as follows:

[0068] p=[p seq ∥p graph ].

[0069] Preferably, in order to comprehensively extract network scale features, the present invention constructs a heterogeneous graph network based on drug-target interaction (DTI). Specifically, drug molecules and target proteins are respectively used as two types of nodes in the heterogeneous graph. By analyzing the affinity data in the training data set, drug-target pairs with affinity values ​​greater than the set threshold K are screened out, and connecting edges are established in the heterogeneous graph for these qualified drug-target pairs to reflect the strong affinity interaction relationship between the drug and the target. By constructing a DTI heterogeneous graph network, the complex relationship between drug molecules and target proteins can be more comprehensively revealed, and effective data support and theoretical basis can be provided for further drug discovery and precision medicine.

[0070] Preferably, the cross-scale feature fusion method includes:

[0071] The drug molecule scale feature vector d and the target protein molecule scale feature vector p are stacked into a matrix:

[0072] X DTI =[d1,…,d M ,p1,…,p N ] T

[0073] M is the total number of drugs, N is the total number of target proteins, and X DTI It is regarded as the initial feature matrix of the drug-target interaction heterogeneous graph network, and a two-layer graph convolutional neural network is used for feature extraction, fusing the molecular scale and drug-target interaction heterogeneous graph network scale features. The calculation process is as follows:

[0074]

[0075] Preferably, after feature fusion, the fusion drugs are obtained:

[0076] Drug characterization based on network-scale features of heterogeneous graphs of target interaction relationships and target protein feature vector The formula is as follows:

[0077]

[0078] After integrating the network scale features, the drug feature vector is cascaded. and target protein feature vector Combined to form a joint feature representation of drug-target, which is used as the input of the subsequent fully connected layer to finally obtain the affinity prediction value The formula for the full connection process is as follows:

[0079]

[0080] Example 2, a drug target affinity prediction model.

[0081] This example proposes a drug target affinity prediction model, which is obtained by training a neural network using the method in Example 1. It should be noted that the parts not mentioned in this example are the same as those in Example 1.

[0082] Specifically, this embodiment is a drug-target affinity prediction model based on multimodal cross-scale feature fusion (Multimodal Cross Scale Feature Fusion, MCSF_DTA). In view of the one-sidedness of target protein feature learning caused by relying only on a single modal feature, the MCSF_DTA model fuses the target protein sequence and graph features of different modalities to enhance the feature representation of a single target protein molecule. In view of the fact that existing methods ignore the network scale features contained in multi-source biological networks, by constructing a heterogeneous graph network of drug-target interaction (DTI), a cross-scale feature fusion method is used to achieve feature fusion at the molecular scale and network scale, further enriching the feature representation of drugs and target proteins.

[0083] The overall framework of the MCSF_DTA model of the present invention is as follows Figure 1 The model adopts a dual encoder architecture to extract features of drugs and target proteins respectively, and finally outputs affinity values ​​through the affinity prediction module.

[0084] In the drug branch, the present invention converts the drug molecule graph G d The input contains a three-layer graph convolutional neural network module to learn the feature representation of drug molecule graphs. The graph convolution formula is as follows:

[0085]

[0086] Among them, A d is the drug molecule adjacency matrix, I is the identity matrix, which is used to add self-connection, and Z is the degree matrix, which means that its diagonal elements are the degrees of the nodes. represents the first layer node feature matrix, W (l) is the learnable weight matrix of the first layer, and σ(·) represents the ReLU activation function. Through the graph convolution operation, the features of each node are weighted and aggregated with the features of its neighboring nodes according to the graph structure, and the updated node representation is generated using nonlinear transformation, thereby effectively capturing local and global graph features. In order to obtain the final drug feature representation d, a readout block containing a global average pooling layer and a fully connected layer is designed after the last layer of the graph convolutional neural network. The calculation formula is as follows:

[0087]

[0088] in, represents the last layer of drug molecule node embedding, N d is the number of atoms in the drug molecule.

[0089] In the target protein branch, the present invention represents the target protein as two modes, amino acid sequence and residue graph, as the input of MCSF_DTA. The target protein sequence can be regarded as a time series, and the present invention uses a combination of 1D-CNN and BiLSTM to extract its features. 1D-CNN can effectively extract local patterns and features in sequence data by sliding the convolution kernel in one-dimensional space. BiLSTM is a deep learning model that can simultaneously capture the forward and backward dependencies of sequence data, and is widely used in tasks such as natural language processing and time series analysis. Specifically, the target protein sequence is input into a network containing two layers of 1D-CNN to extract local features. The 1D-CNN implementation process expression is as follows:

[0090]

[0091] The local features extracted by 1D-CNN are used as the input of BiLSTM recurrent neural network, and the BiLSTM temporal dependency modeling capability is used to achieve richer feature expression. The calculation process of BiLSTM includes the calculation of two LSTM layers, forward and reverse. The specific process is as follows:

[0092]

[0093] Among them, H CNN represents the local features extracted by 1D-CNN, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the reverse LSTM at time step t, [·;·] represents the vector concatenation operation, and p seq As the final output, it contains the forward and backward dependency feature information of the target protein sequence.

[0094] For the target protein residue graph, the present invention extracts the target protein graph modal feature information through a three-layer graph convolutional neural network layer, and uses a readout module including a global average pooling layer and a fully connected layer to obtain the target protein graph feature representation P graph , the calculation process is as follows:

[0095]

[0096]

[0097] in, represents the graph convolution operation, X p Represents the target protein residue feature matrix, Ap Represents the target protein residue adjacency matrix.

[0098] After obtaining the feature vectors of the target protein sequence and the graph respectively, the present invention uses a cascade method to convert the sequence feature vector p seq and graph feature vector p graph Connect them together to get the complete feature vector p of a single target protein, and the calculation process is as follows:

[0099] p=[p seq ∥p graph ]

[0100] The feature representation of a single target protein molecule is enhanced by representing the target protein as a graph and extracting features in different modes of sequence.

[0101] In the affinity prediction module, the present invention constructs a drug-target interaction (DTI) heterogeneous graph network G DTI , using the cross-scale feature fusion method to transfer features in the DTI heterogeneous graph network, integrating the molecular scale and G DTI The network scale feature further enriches the feature representation of drugs and target proteins, thereby improving the model prediction performance. Specifically, the drug feature vector d and the target protein feature vector p are stacked into a matrix:

[0102] X DTI =[d1,…,d M ,p1,…,p N ] T

[0103] M is the total number of drugs, N is the total number of target proteins, and X DTI is regarded as the initial feature matrix of the DTI heterogeneous graph network, and a two-layer graph convolutional neural network is used to perform G DTI Feature extraction is performed to fuse the molecular scale and DTI heterogeneous graph network scale features. The calculation process is as follows:

[0104]

[0105] After feature fusion, drug features integrating DTI heterogeneous graph network scale features are obtained. and target protein feature vector The formula is as follows:

[0106]

[0107] After integrating the network scale features, the drug feature vector is cascaded. and target protein feature vector Combined to form a joint feature representation of drug-target, which is used as the input of the subsequent fully connected layer to finally obtain the affinity prediction value The formula for the full connection process is as follows:

[0108]

[0109] Experimental verification.

[0110] To verify the effect of the present invention, comparative experiments were conducted on two drug target datasets, DAVIS and KIBA datasets. In previous drug-target affinity prediction studies, these two datasets were considered as benchmark datasets.

[0111] In the experiment, the DAVIS and KIBA datasets were divided into training and test sets according to the data partitioning method of previous work. The MCSF_DTA model was trained on the training set, and the consistency index CI and mean square error MSE were output as evaluation indicators after each iteration in the test set to measure the model performance. The model was implemented based on Pytorch 1.8.0 and PyTorchGeometric 2.0.4, and the operating environment was NVIDIARTX 3090GPU.

[0112] In order to evaluate the performance of the MCSF_DTA model (the model in Example 2 of the present invention), the MCSF_DTA model was compared with other advanced drug-target affinity prediction models, including the following 8 DTA prediction models:

[0113] SimBoost: Utilizes the similarity information between drugs and targets and uses the gradient boosting method to construct a similar feature network between drugs and targets for affinity prediction.

[0114] DeepDTA: Drugs and target proteins are represented as sequences, and two 1D-CNN modules are used to extract sequence features of drugs and target proteins respectively to predict affinity.

[0115] GraphDTA: Represent drugs as graph structures and target proteins as sequences, and use GNN and 1D-CNN modules to extract features and predict affinity, respectively.

[0116] AttentionDTA: A DTA prediction model based on attention mechanism is proposed, which uses the attention mechanism to measure the importance of different subsequences in drugs and target proteins.

[0117] DeepGLSTM: Improves the drug molecule graph representation by constructing a multi-channel power graph, effectively extracts the features of neighbor nodes of different orders, and enhances the feature representation of drug molecules.

[0118] MGraphDTA: A dense connection strategy is used to implement deep GNN, and multiple CNN layers are combined to build MGNN and MCNN modules for affinity prediction.

[0119] HiSIF-DTA: By constructing a protein hierarchical semantic graph, the protein interaction network is used to enrich the protein feature representation for DTA prediction.

[0120] SISDTA: A DTA prediction model based on structural similarity is proposed, which measures drug similarity through the inclusion relationship of molecular substructures and constructs a drug similarity network for DTA prediction.

[0121] Figure 2 The average MSE and CI scores of the MCSF_DTA model and 8 benchmark models on the DAVIS and KIBA datasets were recorded. The difference in CI values ​​of the MCSF_DTA model proposed in the present invention on the DAVIS and KIBA datasets is controlled within 1%, indicating that MCSF_DTA exhibits good generalization ability in the DTA tasks of different datasets. The comparative experimental results show that the MCSF_DTA model outperforms other benchmark models in performance. Compared with the SimBoost model, the MSE of MCSF_DTA on the two datasets was reduced by 12.8% and 10.3%, respectively, and the CI was increased by 4% and 7%, respectively. This shows that compared with traditional machine learning methods, deep learning methods have shown significant advantages in DTA tasks due to their powerful data processing and feature extraction capabilities. Compared with the DeepDTA model based on sequence representation and the GraphDTA model based on graph representation, the performance of MCSF_DTA on both datasets is significantly improved, indicating that different modalities have their own unique feature information, and the prediction performance can be improved by fusing multimodal features. Compared with the HiSIF-DTA model that integrates the high-order semantics of the protein interaction network, the MSE of MCSF_DTA on the two data sets was reduced by 3.7% and 0.2%, and the CI was increased by 0.6% and 0.2%, respectively, which further illustrates the effectiveness of integrating molecular scale and network scale features in improving the prediction performance of DTA. Compared with the SISDTA model with the best effect in the benchmark model, the MSE values ​​of the MCSF_DTA model proposed in the present invention on the two data sets were reduced by 1.5% and 0.3%, and the CI values ​​were increased by 0.5% and 0.4%, respectively, which fully demonstrated the superiority of the MCSF_DTA model in improving the prediction performance of drug target affinity, and at the same time demonstrated its robustness and generalization ability on different data sets.

[0122] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, all equivalent changes made according to the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting drug target affinity, characterized in that: include: Extract features of drugs and target proteins to obtain drug molecular scale feature representation d and target protein molecular scale feature representation p; According to the molecular scale feature representation of drugs and target proteins, a heterogeneous graph network of drug-target interaction relationship is constructed, and the heterogeneous graph network of drug-target interaction relationship includes network scale feature; A cross-scale feature fusion method is used to fuse the drug molecule scale feature representation d, the target protein molecule scale feature representation p and the network scale features to generate a drug-target joint feature representation to predict the final affinity score.

2. The method for predicting drug target affinity according to claim 1, characterized in that: In drug molecular scale feature extraction, the drug SMILES sequence is converted into a molecular graph, and feature extraction is performed through a multi-layer graph convolutional neural network.

3. The method for predicting drug target affinity according to claim 1, characterized in that: In the molecular scale feature extraction of the target protein, the target protein is represented as two modes, amino acid sequence and residue graph, for feature extraction, thereby enriching the molecular scale feature representation of the target protein.

4. The method for predicting drug target affinity according to claim 2, characterized in that: In drug molecule scale feature extraction, the drug molecule graph is input into a three-layer graph convolutional neural network module to learn the feature representation of the drug molecule graph. The graph convolution formula is as follows: Among them, A d is the drug molecule adjacency matrix, I is the identity matrix, which is used to add self-connection, and Z is the degree matrix, which means that its diagonal elements are the degrees of the nodes. represents the first layer node feature matrix, W (l) is the learnable weight matrix of the first layer, and σ(·) represents the ReLU activation function.

5. The method for predicting drug target affinity according to claim 4, characterized in that: Through the graph convolution operation, the features of each node and its neighboring nodes are weighted and aggregated according to the graph structure, and the updated node representation is generated using nonlinear transformation, thereby effectively capturing local and global graph features. After the last layer of graph convolutional neural network, a readout block including a global average pooling layer and a fully connected layer is designed to obtain the final drug molecule scale feature representation d. The calculation formula is as follows: in, represents the last layer of drug molecule node embedding, N d is the number of atoms in the drug molecule.

6. The method for predicting drug target affinity according to claim 3, characterized in that: In the molecular scale feature extraction of target proteins, the target proteins are represented as two modes, amino acid sequence and residue graph, for feature extraction. The neural network modules including 1D-CNN and BiLSTM are used to extract features from the sequence mode. The target protein sequence is input into a two-layer 1D-CNN network to extract local features. The process expression is as follows: The local features extracted by 1D-CNN are used as the input of the BiLSTM recurrent neural network. The calculation process of BiLSTM includes the calculation of the forward and reverse LSTM layers. The process is as follows: Among them, H CNN represents the local features extracted by 1D-CNN, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the reverse LSTM at time step t, [·;·] represents the vector concatenation operation, and p seq As the final output, it contains the forward and backward dependency feature information of the target protein sequence.

7. The method for predicting drug target affinity according to claim 3, characterized in that: For the target protein residue graph, a three-layer graph convolutional neural network layer is used to extract the modal feature information of the target protein graph, and a readout module including a global average pooling layer and a fully connected layer is used to obtain the target protein graph feature representation P graph , the calculation process is as follows: in, represents the graph convolution operation, X p Represents the target protein residue feature matrix, A p represents the target protein residue adjacency matrix; After obtaining the feature vectors of the target protein sequence and graph respectively, the sequence feature vector p is concatenated using the cascade method. seq and graph feature vector p graph Connect them together to get the complete feature vector p of a single target protein, and the calculation process is as follows: p=[p seq ∥p graph ]。 8. The method for predicting drug target affinity according to claim 1, characterized in that: Cross-scale feature fusion methods include: The drug molecule scale feature vector d and the target protein molecule scale feature vector p are stacked into a matrix: X DTI =[d1,…,d M ,p1,…,p N ] T M is the total number of drugs, N is the total number of target proteins, and X DTI It is regarded as the initial feature matrix of the drug-target interaction heterogeneous graph network, and a two-layer graph convolutional neural network is used for feature extraction, fusing the molecular scale and drug-target interaction heterogeneous graph network scale features. The calculation process is as follows:

9. The method for predicting drug target affinity according to claim 8, characterized in that: After feature fusion, the fusion drugs are obtained respectively - Drug characterization based on network-scale features of heterogeneous graphs of target interaction relationships and target protein feature vector The formula is as follows: After integrating the network scale features, the drug feature vector is cascaded. and target protein feature vector Combined to form a joint feature representation of drug-target, which is used as the input of the subsequent fully connected layer to finally obtain the affinity prediction value The formula for the full connection process is as follows:

10. A drug target affinity prediction model, characterized in that: The prediction model is obtained by training a neural network using the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Drug target affinity prediction method based on deep learning

    CN110689965A

  • Method for predicting binding affinity of drug molecule and target protein

    CN113936735A

  • Protein multilevel semantic aggregation characterization method for drug-target affinity prediction

    CN117393036A

  • Drug target binding affinity prediction method and system

    CN117594116A

  • Drug target binding affinity prediction method based on drug bimodal characteristics

    CN118298908A

Cited By

  • Drug-target binding affinity prediction model training method and prediction method based on comparative learning

    CN120954561A