A drug-target interaction prediction method integrating multi-layer drug structure information

Through deep learning models, the molecular structure information of drug is preprocessed and feature extracted, combined with the molecular completion map convolutional neural network and the Transformer network, the high cost and low efficiency problems of drug-target interaction relationship verification are solved, and more efficient drug research and development and drug reuse are achieved.

CN114067905BActive Publication Date: 2025-06-06DALIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111313022.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-06-06
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

The prior art has high cost and low efficiency problems in the verification of drug-target interaction relationships, and similarity-based methods ignore deep feature information between drugs and targets, making it difficult for deep learning methods to learn complex omics data.

Method used

The deep learning model is used to preprocess the molecular structure information of the drug, extract the characteristic information of the drug and target, and combine the molecular completion map convolution neural network and the Transformer network to construct a drug-target interaction prediction model.

Benefits of technology

It improves the prediction accuracy and efficiency of drug-target interaction relationships, reduces verification costs, shortens drug development cycles, and provides an important foundation for drug reuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067905B_ABST
    Figure CN114067905B_ABST
Patent Text Reader

Abstract

The present invention provides a method for predicting drug-target interactions by integrating multi-layer drug structure information. First, the drug and target information in the pharmacomics database is preprocessed to extract the drug and target information with interactions; secondly, the drug SMILES molecular fingerprint is represented as a molecular graph structure, and the molecular completion graph convolutional neural network and the Transformer network are used to extract drug feature information; then, the target sequence information is processed using a convolutional neural network, and the target feature information is extracted; finally, the extracted drug feature information and target feature information are sent to a classification model for training, the model is saved, and the relationship between the drug and the target is predicted. The present invention effectively extracts the feature information in the drug molecular structure, has a higher accuracy rate when predicting the drug-target relationship, improves the efficiency and accuracy of drug-target relationship verification, effectively shortens the drug research and development cycle, and greatly reduces the cost of new drug research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical artificial intelligence and natural language processing technology, and in particular to a drug-target interaction prediction method integrating multi-layer drug structure information. Background Art

[0002] The development of new drugs is an expensive and time-consuming process. It is well known that the total average development cost of a new drug ranges from 2 to 3 billion US dollars, and the total development time takes 13-15 years. Therefore, the field of drug development urgently needs a fast and efficient development method to improve the efficiency of drug development and reduce the cost of drug development. A large number of studies have shown that drug repositioning methods can effectively shorten the drug development cycle and greatly reduce the development cost of new drugs. Drug-target interactions play a key role in drug repositioning research. The development of the Human Genome Project has enabled the rapid accumulation of data on drug compounds, targets, and interactions, providing data accumulation for the prediction of drug-target interactions. However, there are still a large number of interactions between drugs and targets that have not been discovered and verified.

[0003] At present, the verification of drug-target interaction relationship mainly relies on large-scale biological or chemical experiments. In the verification process, a lot of manpower, material and financial resources are required, and there is a great deal of contingency in the verification through experiments, which increases the cost of drug-target interaction relationship verification. In order to reduce the cost of drug-target interaction relationship verification, more and more computational methods are used to predict drug-target interaction relationship, but these methods all have certain defects. For example, similarity-based methods often ignore the deep feature information between drugs and targets. Although deep learning methods can obtain more feature information of drugs and targets, it is difficult to learn the complex relationship between different entities for complex omics data, and lacks practical guidance for drug reuse. The drug molecular structure diagram contains various atoms and chemical bonds that constitute the drug, which is a concentrated reflection of the chemical properties and efficacy of the drug, and has an important impact on the prediction of drug-target interaction. However, the current methods based on graph neural networks only focus on the interaction relationship in complex networks and ignore the molecular structure properties of the drug itself. Summary of the invention

[0004] In response to the above-mentioned problems in the prior art, the present application proposes a deep learning model based on drug molecular structure information to automatically predict the interaction relationship between drugs and targets, which improves verification efficiency and reduces verification costs.

[0005] To achieve the above objectives, the technical solution of the present application is: a drug-target interaction prediction method integrating multi-layer drug structure information, comprising:

[0006] Step 1: Preprocess the drug and target information in the pharmacomics database, extract the information of drugs and targets with interactions, and construct drug-target interaction data;

[0007] Step 2: Represent the drug SMILES molecular fingerprint as a molecular graph structure, and use the molecular completion graph convolutional neural network and Transformer network to extract drug feature information;

[0008] Step 3: Embed the target sequence information and process it using a convolutional neural network to extract target feature information;

[0009] Step 4: sending the extracted drug characteristic information and target characteristic information into a classification model for training, and then saving the model;

[0010] Step 5: Load the model, input the drug and target information to be predicted, predict the relationship between the drug and the target, and output the prediction result.

[0011] Furthermore, step 1 specifically includes:

[0012] Step 1.1: Screen the drug information and target information from the pharmacomics database and delete the drug information and target information without interaction relationship;

[0013] Step 1.2: Integrate the drugs and targets with interaction relationships and construct them into the form of <drug number, target number, label>, and mark the label as 1;

[0014] Step 1.3: Obtain the SMILES molecular fingerprint corresponding to the drug and the sequence information corresponding to the target from the pharmacomics database as the specific representation information of the drug and the target respectively;

[0015] Step 1.4: Randomly construct unknown drug-target relationships as negative examples in a ratio of 1:2 between positive examples and negative examples, and mark the negative example labels as 0.

[0016] Furthermore, step 2 specifically includes:

[0017] Step 2.1: By calling the RDKit function library in the Python library, the SMILES molecular fingerprint of each drug is represented as a graph, where the vertices and edges of the graph represent the atoms and chemical bonds of the drug, respectively. Each drug molecule is represented by a feature matrix and an adjacency matrix. Each row of the feature matrix corresponds to the attribute of each atom. Each drug is represented as Where N represents the type of drug, represents the feature matrix of the drug, represents the adjacency matrix of drugs, D irepresents the number of atoms of the ith drug, and C represents the number of characteristic channels of the atoms;

[0018] Step 2.2: See Figure 2 , the molecular completion graph convolutional neural network and Transformer network are used to extract drug feature information.

[0019] Furthermore, step 2.2 specifically includes:

[0020] Step 2.2.1: The molecular completion graph convolutional neural network MCGCN takes the graph G obtained in step 2.1 as input. MCGCN adds a supplementary graph to the original drug molecule graph to ensure that the adjacency matrix and feature matrix of each drug molecule are of the same size, where the original graph and the supplementary graph are independent of each other; after completion, the drug molecule graph is represented as:

[0021]

[0022] in, Represents the connection matrix between the original graph G and the completed graph G′ of the i-th drug; Respectively represent the adjacency matrix and feature matrix after completion; through the completion operation, all drug molecules are represented as a graph G with the same number of nodes. new ; MCGCN contains two hidden layers, and the drug is represented in each hidden layer using formula (2):

[0023]

[0024] in, is the adjacency matrix with self-attention added, yes The weight matrix of and θ (l) is the convolution signal and filter parameter of the lth layer; after each hidden layer, σ(·) is used to represent the activation function, and σ(·) is set to ReLU(·) = max(0,·); at the end of MCGCN, maximum pooling is used to reduce the dimension of the data;

[0025] Step 2.2.2: Transformer network with the output vector of each hidden layer in MCGCN As input, the Encoder part of the Transformer network is used to extract features; in the Transformer network, different multi-head attention modules are used to process the vector information from different hidden layers in MCGCN. For the vector from the first hidden layer in MCGCN, Use a multi-head attention module with 6 heads for processing; for the vector from the second hidden layer of MCGCN A multi-head attention module with 4 heads is used for processing; in the multi-head attention module, formula (3) is used to extract features:

[0026] MultiHead j (Q,K,V)=Concat(head 1 ,...,head i ) (3)

[0027] The feature vectors processed by the multi-head attention module are concatenated using formula (4) and sent to the original layer normalization part of the Transformer network. Finally, the layer normalization output vector is sent to the fully connected feedforward neural network, and the output of the fully connected feedforward neural network is used as the final feature vector M of the drug. alldrug ;

[0028] AllMultiHead=Concat(MultiHead 1 ,...,MultiHead j ) (4).

[0029] Furthermore, the step 3 specifically includes:

[0030] Step 3.1: Randomly initialize a lookup table corresponding to all amino acids appearing in the target sequence, with a size of 26×20; correspond each amino acid in the target sequence to the lookup table, and construct the embedding matrix M of the target sequence tar ; The embedding matrix M tar The length of is the maximum length in the target sequence, which is set to 2500, and the width is consistent with the width of the lookup table. During the model training process, the embedding vector is continuously optimized, so the relevant information in the lookup table will continue to change as the model is optimized.

[0031] Step 3.2: See Figure 3 , use convolutional neural network to extract feature information from the target sequence, and embed the embedding matrix M obtained in step 3.1 tar As the input of the convolutional neural network; empty labels are automatically filled for target sequences that are shorter than the length of the embedding matrix.

[0032] Furthermore, the step 3.2 specifically includes:

[0033] Step 3.2.1: The embedding matrix M obtained in step 3.1 tarThe input is a convolution layer with kernel sizes of 10, 15, and 20 and a step size of 1 for feature extraction. The extracted feature vector is sent to the ELU activation function for optimization. The ELU activation function is defined as follows:

[0034]

[0035] Step 3.2.2: The vector optimized in the ELU activation function is sent to the global maximum pooling layer to extract the most important local features. After the global maximum pooling layer, the vector dimension obtained is 128;

[0036] Step 3.3.3: Concatenate the output vectors of each maximum pooling layer to obtain a concatenated vector with a dimension of 384, and input the concatenated vector into a fully connected neural network to obtain a vector with a dimension of 128 as the final feature vector M of the target. alltar .

[0037] Furthermore, the step 4 specifically includes:

[0038] Step 4.1: Add the drug feature vector M obtained in step 2 alldrug And the target feature vector M obtained in step 3 alltar Splice and get the final vector representation M of the input data all , and the labels corresponding to the original drug-target relationships are used as the labels of the final vector;

[0039] Step 4.2: Represent the final vector obtained in step 4.1 as M all The labels are input into the fully connected neural network to train the model. In order to obtain the best model effect, the L2 norm optimized binary cross entropy function is used to optimize the model, and the best model model_best is saved:

[0040]

[0041]

[0042] Furthermore, the step 5 specifically includes:

[0043] Load the model model_best in step 4.2, input the drug-target information in the validation data into the model, determine whether there is an interaction relationship between the drug and the target, and output the corresponding evaluation index;

[0044] Due to the adoption of the above technical scheme, the present invention can achieve the following technical effects: the present invention adopts a deep learning model, utilizes the information of drugs and targets in the medical database, combines the structural characteristics of drugs and targets, and automatically predicts the information of drug-target interaction through the model. It effectively extracts the characteristic information in the molecular structure of the drug, has a higher accuracy rate when predicting the drug-target relationship, and has robustness, improves the efficiency and accuracy of drug-target relationship verification, effectively shortens the cycle of drug research and development, greatly reduces the cost of new drug research and development, and provides an important foundation and guarantee for new drug research and development and drug reuse. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flow chart of a drug-target interaction prediction method integrating multi-layer drug structure information;

[0046] Figure 2 Extraction flow chart for drug feature information;

[0047] Figure 3 This is the flow chart for extracting target feature information. DETAILED DESCRIPTION

[0048] The embodiments of the present invention are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0049] Example 1

[0050] The present invention is described in detail below in conjunction with embodiments so that those skilled in the art can implement the invention according to the description.

[0051] In this embodiment, Windows system is used as the development environment, Pycharm is used as the development platform, Python is used as the development language, and the drug-target interaction prediction method integrating multi-layer drug structure information of the present invention is used to predict the drug-target interaction relationship and the potential therapeutic drugs for COVID-19.

[0052] In this embodiment, a drug-target interaction prediction method integrating multi-layer drug structure information includes the following steps:

[0053] Step 1: Given a target, the Delta target of COVID-19, find drugs that interact with it in the PubChem database, a total of 54, and set the data label to 1;

[0054] Step 2: Randomly select 108 drugs that have no interaction with Delta targets in the PubChem database to construct negative examples, and set the data label to 0;

[0055] Step 3: Obtain the chemical structure of the drugs involved in step 1 and step 2 and the sequence structure of the Delta target from the PubChem database;

[0056] Step 4: Convert the chemical structure of the drug into a molecular structure diagram using the RDKit toolkit in the Python library, and save the molecular structure diagram as a file in hkl format, with each drug saved as one file;

[0057] Step 5: Take the drug hkl format file and the sequence structure of the Delta target as input, load the saved model, and obtain the evaluation index and prediction score of the interaction relationship between the drug and the Delta target. The evaluation index includes accuracy (ACC), F1 value and AUC;

[0058]

[0059]

[0060]

[0061]

[0062] TP: True positive, the number of positive classes correctly predicted as positive; FP: False positive, the number of negative classes incorrectly predicted as positive; FN: False negative, the number of positive classes incorrectly predicted as negative; TN: True negative, the number of negative classes correctly predicted as negative. AUC is expressed as the area under the ROC curve;

[0063] Step 6: Sort the predicted scores in step 5 in descending order to obtain the top 5 drug information.

[0064] According to the above steps, the present invention compares the drug-target relationship prediction effect with the Deep DTA model, the Deep DTI model, the Deep Conv-DTI model, and the ML-DTI model. As can be seen from Table 1, the method proposed in this invention is significantly better than other methods in terms of AUC, F1 value, and prediction accuracy.

[0065] Table 1 Comparison of prediction results of different models for drug-target relationship

[0066]

[0067] At the same time, the method of the present invention is used to predict potential therapeutic drugs for COVID-19 and Delta targets. In the experimental results, four drugs including Tramadol in the top five drugs have been used in the clinical treatment of COVID-19 or have literature support for the inhibitory effect on COVID-19, as shown in Table 2. Tramadol, Amitriptyline and Dextromethorphan all have a close interaction relationship with the Delta target. Among them, Dexamethasone and Dextromethorphan are widely used in the clinical treatment of COVID-19 and successfully alleviate the complications of COVID-19. Tramadol can protect COVID-19 patients from the complications of the disease by increasing antioxidant enzymes, superoxide dismutase and glutathione peroxidase while reducing the effects of malondialdehyde. Studies have shown that after treating cells with different concentrations of Amitriptyline, the probability of cell infection is reduced by 90%, which also provides a basis for the use of Amitriptyline in the treatment of COVID-19.

[0068] Table 2 The top five therapeutic drugs related to COVID-19 recommended by the present invention

[0069]

[0070] The foregoing description of specific exemplary embodiments of the present invention is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present invention to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present invention and various different selections and changes. The scope of the present invention is intended to be limited by the claims and their equivalents.

Claims

1. A drug-target interaction prediction method integrating multi-layer drug structure information, It is characterized in that include: Step 1: Preprocess the drug and target information in the pharmacomics database, extract the information of drugs and targets with interactions, and construct drug-target interaction data; Step 2: Represent the drug SMILES molecular fingerprint as a molecular graph structure, and use the molecular completion graph convolutional neural network and Transformer network to extract drug feature information; Step 3: Embed the target sequence information and process it using a convolutional neural network to extract target feature information; Step 4: sending the extracted drug characteristic information and target characteristic information into a classification model for training, and then saving the model; Step 5: Load the model, input the drug and target information to be predicted, predict the relationship between the drug and the target, and output the prediction result.

2. According to claim 1, a method for predicting drug-target interactions by integrating multi-layer drug structure information, It is characterized in that Step 1 specifically includes: Step 1.1: Screen the drug information and target information from the pharmacomics database and delete the drug information and target information without interaction relationship; Step 1.2: Integrate the drugs and targets with interaction relationships and construct them into the form of <drug number, target number, label>, and mark the label as 1; Step 1.3: Obtain the SMILES molecular fingerprint corresponding to the drug and the sequence information corresponding to the target from the pharmacomics database as the specific representation information of the drug and the target respectively; Step 1.4: Randomly construct unknown drug-target relationships as negative examples in a ratio of 1:2 between positive examples and negative examples, and mark the negative example labels as 0.

3. According to claim 1, a method for predicting drug-target interactions by integrating multi-layer drug structure information, It is characterized in that Step 2 specifically includes: Step 2.1: By calling the RDKit function library in the Python library, the SMILES molecular fingerprint of each drug is represented as a graph, where the vertices and edges of the graph represent the atoms and chemical bonds of the drug, respectively. Each drug molecule is represented by a feature matrix and an adjacency matrix. Each row of the feature matrix corresponds to the attribute of each atom. Each drug is represented as Where N represents the type of drug, represents the feature matrix of the drug, represents the adjacency matrix of drugs, D i represents the number of atoms of the ith drug, and C represents the number of characteristic channels of the atoms; Step 2.2: Use the molecular completion graph convolutional neural network and Transformer network to extract drug feature information.

4. According to claim 3, a method for predicting drug-target interaction by integrating multi-layer drug structure information, It is characterized in that Step 2.2 specifically includes: Step 2.2.1: The molecular completion graph convolutional neural network MCGCN takes the graph G obtained in step 2.1 as input. MCGCN adds a supplementary graph to the original drug molecule graph to ensure that the adjacency matrix and feature matrix of each drug molecule are of the same size, where the original graph and the supplementary graph are independent of each other; after completion, the drug molecule graph is represented as: in, Represents the connection matrix between the original graph G and the completed graph G′ of the i-th drug; Respectively represent the adjacency matrix and feature matrix after completion; through the completion operation, all drug molecules are represented as a graph G with the same number of nodes. new ; MCGCN contains two hidden layers, and the drug is represented in each hidden layer using formula (2): in, is the adjacency matrix with self-attention added, yes The weight matrix of and θ (l) is the convolution signal and filter parameter of the lth layer; after each hidden layer, σ(·) is used to represent the activation function, and σ(·) is set to ReLU(·) = max(0,·); at the end of MCGCN, maximum pooling is used to reduce the dimension of the data; Step 2.2.2: Transformer network with the output vector of each hidden layer in MCGCN As input, the Encoder part of the Transformer network is used to extract features; in the Transformer network, different multi-head attention modules are used to process the vector information from different hidden layers in MCGCN. For the vector from the first hidden layer in MCGCN, Use a multi-head attention module with 6 heads for processing; for the vector from the second hidden layer of MCGCN A multi-head attention module with 4 heads is used for processing; in the multi-head attention module, formula (3) is used to extract features: MultiHead j (Q,K,V)=Concat(head 1 ,…,head i ) (3) The feature vectors processed by the multi-head attention module are concatenated using formula (4) and sent to the original layer normalization part of the Transformer network. Finally, the layer normalization output vector is sent to the fully connected feedforward neural network, and the output of the fully connected feedforward neural network is used as the final feature vector M of the drug. alldrug ; AllMultiHead=Concat(MultiHead 1 ,…,MultiHead j ) (4)。 5. According to claim 1, a method for predicting drug-target interaction by integrating multi-layer drug structure information, It is characterized in that The step 3 specifically includes: Step 3.1: Randomly initialize a lookup table corresponding to all amino acids appearing in the target sequence; correspond each amino acid in the target sequence to the lookup table to construct the embedding matrix M of the target sequence tar ; The embedding matrix M tar The length of is the maximum length in the target sequence, and the width is consistent with the width of the lookup table; Step 3.2: Use a convolutional neural network to extract the feature information in the target sequence and embed the matrix M obtained in step 3.1 tar As the input of the convolutional neural network; empty labels are automatically filled for target sequences that are shorter than the length of the embedding matrix.

6. A drug-target interaction prediction method integrating multi-layer drug structure information according to claim 5, It is characterized in that The step 3.2 specifically includes: Step 3.2.1: The embedding matrix M obtained in step 3.1 tar The input is a convolution layer with kernel sizes of 10, 15, and 20 and a step size of 1 for feature extraction. The extracted feature vector is sent to the ELU activation function for optimization. The ELU activation function is defined as follows: Step 3.2.2: The vector optimized in the ELU activation function is fed into the global max pooling layer to extract the most important local features; Step 3.3.3: Concatenate the output vectors of each maximum pooling layer and input the concatenated vector into a fully connected neural network as the final feature vector M of the target. alltar .

7. According to claim 1, a method for predicting drug-target interaction by integrating multi-layer drug structure information, It is characterized in that The step 4 specifically includes: Step 4.1: Add the drug feature vector M obtained in step 2 alldrug And the target feature vector M obtained in step 3 alltar Splice and get the final vector representation M of the input data all , and the labels corresponding to the original drug-target relationships are used as the labels of the final vector; Step 4.2: Represent the final vector obtained in step 4.1 as M all The labels are input into the fully connected neural network to train the model; the model is optimized using the binary cross entropy function optimized by the L2 norm, and the best model model_best is saved:

8. A drug-target interaction prediction method integrating multi-layer drug structure information according to claim 7, It is characterized in that The step 5 specifically includes: Load the model model_best in step 4.2, input the drug-target information in the validation data into the model, determine whether there is an interaction relationship between the drug and the target, and output the corresponding evaluation index.

Citation Information

Patent Citations

  • Drug small molecule-protein target reaction prediction method based on multi-dimensional information

    CN112331273A

  • Drug recommendation system for regulating and controlling disease targets based on three-channel deep learning, computer equipment and storage medium

    CN112652358A