A deep learning approach for predicting drug-target interactions based on biological substructures

By extracting the substructure features of drugs and targets, using convolutional neural networks for feature learning, and constructing a multi-layer perceptron for interaction prediction, the accuracy and interpretability problems of drug-target interaction prediction in existing technologies are solved, achieving more efficient drug development.

CN115312125BActive Publication Date: 2025-09-09EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210207457.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-09-09
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

Existing drug-target interaction prediction methods have large granularity and lack biological theoretical support, resulting in large differences in prediction effects and lack of interpretability. In addition, existing methods rely on data statistics and lack pharmacological basis.

Method used

A deep learning method based on biological substructure is adopted to extract the substructure features of drugs and targets through BCM and CFM methods, convolutional neural network is used for feature learning, and a multi-layer perceptron is constructed for interaction prediction.

Benefits of technology

It improves the accuracy of drug-target interaction prediction, reduces noise, enhances the interpretability and pharmacological characteristics of the prediction, and is significantly better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312125B_ABST
    Figure CN115312125B_ABST
Patent Text Reader

Abstract

This paper proposes a deep learning method for predicting drug-target interactions based on biological substructures. The method first extracts the functional substructures of the drug and target, where the drug substructures include molecular branches, common substructures, and retrosynthetic fragments. The amino acid sequence of the target is converted into a species sequence based on its chemical properties, and then non-overlapping k-grams are used for segmentation to obtain the target substructure. The substructure features are then learned based on a convolutional neural network. Experiments show that the method can effectively capture the functional characteristics of drug-target interactions, outperforming existing technologies on datasets of different sizes and distributions, and is reasonable and universal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and deep learning, and in particular to a deep learning method for predicting drug-target interactions based on biological substructures. Background Art

[0002] Drug-target interaction (DTI) specifically refers to the process by which specific intracellular biomacromolecules (called "targets," most of which are proteins such as enzymes and ion channels) bind to drug molecules with appropriate chemical properties and affinities. It is the inherent mechanism of drug treatment for disease, and therefore, drug-target interactions are fundamental to drug discovery and development. Currently, interaction analysis methods based on in vitro experiments face the challenges of long time periods and high costs. In recent years, with the development and achievements of deep learning in various fields of bioinformatics, it has also demonstrated tremendous research potential in the task of drug-target interaction prediction. Using deep learning methods to predict drug-target interactions can accelerate drug discovery and reduce research costs.

[0003] Deep learning-based drug-target interaction prediction methods typically consist of two steps: feature extraction and classification prediction. Since drug molecules are typically represented by SMILES (Simplified Molecular Input Line Entry System) sequences, molecular fingerprints, and molecular graphs, while targets are represented by amino acid sequences, feature extraction often leverages language models and graph neural networks, focusing on biological information, topological structure, and physicochemical properties. In the feature extraction stage, approaches based on language models often take a drug and target sequence representation as input, and encode the drug and target sequences using language models such as word2vec, Doc2vec, and transformers to obtain semantic feature representations of the drug and target. Some studies use the drug's molecular graph as input and employ graph neural networks such as GNNs and GCNs to extract the drug's topological structural features as its feature representation. In classification prediction, the above drug and target features are concatenated and used as input to a classifier. Interaction features are then learned through a multi-layer fully connected network to predict whether an interaction will occur. These methods still have the following limitations:

[0004] 1) The granularity of extracting drug molecular features is mostly at the atomic level or using the entire molecular information and target information as its input. This processing method has no theoretical basis for pharmacological reactions and may also introduce noise. The effects on data sets with different distributions vary greatly and make the predictions lack interpretability.

[0005] 2) Some studies use technologies such as natural language processing to extract drug substructures for prediction. Most of these methods rely on data statistics and lack theoretical support in the biological field. Summary of the Invention

[0006] The purpose of the present invention is to alleviate the limitations of existing drug-target interaction prediction technologies. Inspired by pharmacophore and pharmacological knowledge, namely that drug groups and substructures in drug molecules often make important contributions to drug activity, and the length and type of drug molecule side chains will also affect the affinity of the molecule, thereby affecting the interaction between the drug and the target, a deep learning method for drug-target interaction prediction based on biological substructures is proposed. This method takes drug molecules, targets and their interaction data as research objects, extracts a variety of drug substructures including molecular side chains, common substructures and retrosynthetic fragments, and at the same time extracts target substructures with the help of natural language processing technology. On this basis, a deep learning model is constructed to model drug-target interactions, learn more comprehensive pharmacodynamic properties, obtain high-level, multi-dimensional pharmacological characteristics, improve the accuracy of prediction, and accelerate drug development.

[0007] The specific technical solution for achieving the purpose of the present invention is:

[0008] A deep learning method for predicting drug-target interactions based on biological substructures includes the following steps:

[0009] Step a: Input a drug's normalized SMILES sequence D and target amino acid sequence T;

[0010] Step b: Extract substructures from the drug's normalized SMILES sequence D and the target amino acid sequence T, respectively, including:

[0011] 1) For the normalized SMILES sequence D of the drug, the substructure of the drug is extracted using the BCM method, including:

[0012] 1.1) First, extract the drug's side chains from the standardized SMILES sequence D of the drug. According to the definition rules of SMILES, the side chains are enclosed in brackets (()), and the remaining part is the main chain.

[0013] 1.2) Then extract the common substructures in the main chain based on string matching;

[0014] 1.3) Finally, the main chain is cleaved according to the retrosynthetic fragmentation rules of RECAP (Retrosynthetic Combinatorial Analysis Procedure) to obtain retrosynthetic fragments;

[0015] 1.4) Integrate the above three substructures, namely, the branch chain, the common substructure and the retrosynthetic fragment, as the substructure set F of the standardized SMILES sequence D D .

[0016] 2) For the target amino acid sequence T, the CFM method is used to extract substructures. The CFM method first divides amino acids into 8 categories according to chemical structure or properties, and obtains the category sequence T through category feature mapping. C ; Then use non-overlapping k-gram sequences to convert T C Functional substructure set F that is cut into targets T .

[0017] Step c: Construct a collaborative feature learning module, which consists of two parts: input representation and feature learning. Specifically, it includes:

[0018] 1) In the input representation, the substructure set F of the normalized SMILES sequence D of the drug is D and the functional substructure set F of the target T Perform initial encoding representation, including:

[0019] 1.1) Substructure set F of the normalized SMILES sequence D of the drug D First, label encoding is used to encode the drug substructure set F D Encode and obtain the initial representation of the drug I D ; Then I D Transformed into drug embedding representation E D ∈R max_drug_frag_length*embed_size , where max_drug_frag_length represents the size of the largest drug substructure set, and embed_size represents the embedding dimension;

[0020] 1.2) Functional substructure set F of the target T To represent, first use label encoding to encode the target substructure set F T Encode and obtain the initial representation of the target I T ; Then I T Converted into its embedding representation E T ∈R max _target_frag_length*embed_size , where max_target_frag_length represents the size of the maximum target substructure set, and embed_size represents the embedding dimension.

[0021] 2) Feature learning is divided into drug feature learning and target feature learning, including:

[0022] 2.1) In drug feature learning, E DAs the initial input, it is sent to the convolutional neural network to learn drug features. The convolutional neural network consists of multiple convolution blocks and a final pooling layer. Each convolution block consists of a convolution layer, an activation layer exponential linear unit (ELU), and a batch normalization layer (BatchNorm). Finally, the pooling layer is used to perform feature dimensionality reduction to obtain the final drug representation V D ;

[0023] 2.2) In target feature learning, E T As the initial input, it is sent to the convolutional neural network to learn the target features. The convolutional neural network consists of multiple convolution blocks and a final pooling layer. Each convolution block consists of a convolution layer, an activation layer exponential linear unit (ELU) and a batch normalization layer (BatchNorm). Finally, the pooling layer is connected to obtain the final target representation V T .

[0024] Step d: Build a predictor, including:

[0025] 1) First, assemble the final drug representation V in step c D and target expression V T , the drug-target interaction representation V is obtained;

[0026] 2) V is then fed into a multilayer perceptron for interaction learning. The multilayer perceptron is a fully connected network consisting of multiple fully connected layers and a final sigmoid activation layer. Except for the last layer, each fully connected layer is connected to a rectified linear unit and a dropout layer to prevent overfitting. Finally, the interaction prediction probability is obtained. A value greater than 0.5 indicates that the two are predicted to interact, and a value less than 0.5 indicates that they will not interact.

[0027] The main work and innovations of the present invention include: for drugs, substructure extraction is performed on the standardized SMILES sequence of the drug based on domain knowledge, and the semantic features of the drug are extracted using a convolutional neural network with the substructure as the basic granularity; for targets, the amino acids in the target amino acid sequence are classified according to their chemical properties, and type mapping is performed to obtain the corresponding amino acid type sequence, and then the sequence is segmented with the help of natural language processing methods, and semantic features are extracted using a convolutional neural network with the substructure as the basic granularity; finally, the drug and protein features are spliced ​​as the input of the interaction predictor, and feature learning is performed through a multi-layer fully connected network to predict whether the two will interact.

[0028] Compared with existing technologies, the present invention demonstrates its innovativeness and effectiveness in the following aspects: It proposes a branch chain mining (BCM)-based substructure extraction method that can break down drugs into multiple substructure types, including branches, common substructures, and retrosynthetic fragments. The BCM method first extracts branches according to the SMILES generation rule, then extracts common substructures within the drug, such as benzene rings and hydroxyl groups, and finally uses RECAP to break down the drug backbone into retrosynthetic fragments. It also proposes a category feature mapping (CFM) method that classifies the amino acid sequence of the target protein into eight categories based on the chemical properties of the amino acids, then uses non-overlapping k-grams to segment it into substructures. It then uses a convolutional neural network to learn the synergistic interactions between drug substructures and target substructures, respectively. Finally, the features of both are fused as input to a classifier for interaction prediction. Experimental results show that the present invention significantly outperforms the best existing methods on multiple public datasets, demonstrating that the method can effectively remove noise and extract substructure information of interest for drug-target interactions, achieving excellent results. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic diagram of the process of the present invention;

[0030] Figure 2 yes Figure 1 Detailed flowchart of the multi-layer perceptron in

[15] . DETAILED DESCRIPTION

[0031] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0032] Example 1

[0033] See Figure 1 , the present invention is described in detail according to the following steps:

[0034] Step a: drug-target pair, input the normalized SMILES sequence D of the drug and the target amino acid sequence T;

[0035] Step b: Substructure extraction, which is divided into drug substructure extraction and target substructure extraction, specifically including:

[0036] 1) Drug substructure extraction, using the BCM method, as follows:

[0037] 1.1) First, extract the branches of the drug standardized SMILES sequence D. According to the SMILES generation rule, using sequences enclosed in parentheses (()) to represent branches, traverse the entire drug standardized SMILES sequence D, use a stack to match each parenthesis (()), extract the subsequences within the parentheses (i.e., branches), and concatenate the remaining subsequences to form the main chain.

[0038] 1.2) Detect common substructures in the backbone, such as benzene rings, hydroxyl groups, aldehydes, and carbonyl groups. Since these substructures have corresponding SMILES sequences, for example, a benzene ring can be represented as C1=CC=CC=C1, we can directly detect the presence of these substructures in the backbone through string matching;

[0039] 1.3) Extract the retrosynthetic fragments from the backbone. Fragment the molecule according to the RECAP method, directly using the RDKit tool;

[0040] 1.4) The above three substructures are used as the substructure set F of the drug normalized SMILES sequence D D .

[0041] The above four steps can obtain the substructure set of the drug. For example, the normalized SMILES sequence D of aspirin is represented as CC(=O)OC1=CC=CC=C1C(=O)O. The drug substructure set extracted using BCM is =O (substructure #1), =O (substructure #2), CCOC1=CC=CC=C1CO (substructure #3), and C1=CC=CC=C1 (substructure #4). Substructure #1 and substructure #2 are side chains, substructure #3 is the main chain, and substructure #4 is a common substructure benzene ring. Since the main chain, i.e., substructure #3, cannot be further cleaved by RECAP to obtain retrosynthetic fragments, the substructure set of aspirin does not contain retrosynthetic fragments.

[0042] 2) Target substructure extraction, using the CFM method, as follows:

[0043] 2.1) Amino acids are divided into 8 categories according to their chemical properties. Each category corresponds to a category label, as shown in Table 1. Then, the amino acids in the target amino acid sequence T are mapped to the corresponding category labels to obtain the category sequence T C ;

[0044] Then use the non-overlapping k-gram to T C Divide into substructures of length k and obtain the substructure set F T .

[0045] Table 1

[0046]

[0047] For example, the target sequence T is MKMSRLCLSVALLVLLGTLAASTPGCDTSNQAKAQRPDFCLEPPYTGPCK……ENL, and the type sequence T is obtained after type mapping according to Table 1. C EGEFGAEAFAAAAAAAAFAAAFFHAECFFDDAGADGHCBEACHHBFAH……CDA, and finally use non-overlapping k-gram to segment the sequence, setting k=3, to obtain the substructure set F T Expressed as EGE FGAEAF AAA AAA AAF AAA FFH AEC FFD DAG ADG HCB EAC … A .

[0048] Step c: Collaborative feature learning. This module consists of two parts: input representation and feature learning. Specifically, it includes:

[0049] 1) In the input representation, the substructure set F of the normalized SMILES sequence D of the drug is D and the functional substructure set F of the target T Perform initial encoding representation, including:

[0050] 1.1) Substructure set F of the normalized SMILES sequence D of the drug D First, establish the drug substructure dictionary D d , and sorted by substructure length, for a drug substructure set F D , using label encoding to represent, that is, each substructure uses its d In addition, since the number of substructures generated by each drug is different, the maximum substructure set size is used as its fixed length, and if it is insufficient, it is filled with 0 to obtain the initial representation of the drug I D ∈R max_drug_frag_length , where max_drug_frag_length represents the size of the largest drug substructure set; then the ID is converted into the embedding representation E of the drug D ∈R max_drug_frag_length*embed_size , embed_size represents the embedding dimension;

[0051] 1.2) Functional substructure set F of the target T Establish the target function substructure dictionary Dt , and sorted by substructure length, for a target substructure set F T , using label encoding to represent, that is, each substructure uses its t In addition, since the number of substructures generated by each target is different, the maximum substructure set size is used as its fixed length, and if it is insufficient, it is filled with 0 to obtain the initial representation of the target I T ∈R max_target_frag_length , where max_target_frag_length represents the size of the largest target substructure set; then I T Converted into its embedding representation E T ∈R max_target_frag_length*embed_size , embed_size represents the embedding dimension.

[0052] 2) Feature learning is divided into drug feature learning and target feature learning, including:

[0053] 2.1) In drug feature learning, E D As the initial input, it is sent to the convolutional neural network to learn drug features. The convolutional neural network consists of multiple convolution blocks and a final pooling layer. Each convolution block consists of a convolution layer, an activation layer exponential linear unit (ELU), and a batch normalization layer (BatchNorm). Finally, the pooling layer is used to perform feature dimensionality reduction to obtain the final drug representation V D , as follows:

[0054] 2.1.1) Convolutional layer, representing drug embedding E D After the one-dimensional convolution operation, the feature matrix of the same dimension as the original one is obtained. The convolution operation can be expressed as formula 2, where E D represents drug embedding representation, f represents convolution operation, W j and b j Represent the weight matrix and bias of the j-th convolutional layer respectively;

[0055] E D ∈R max_drug_frag_length*embed_size (1)

[0056] f j =f(E D *W j +b j ) (2)

[0057] 2.1.2) Batch normalization, for f j Normalized, the calculation is as follows, where is the batch set, m represents the size of the batch set, x s express The sth value in and is the mean and variance within the batch, ∈ is the bias, is x s After normalization, γ and β are learning parameters, and the final output is y s ;

[0058]

[0059] 2.1.3) After the normalization function, ELU is selected as the activation function, which can be expressed as formula (6), where x represents the value of the i-th neuron and α is a parameter;

[0060]

[0061] 2.1.4) After passing through N convolutional blocks, a pooling layer is added to reduce the dimension of the feature map and increase the network training speed.

[0062] 2.2) In target feature learning, E is used as the initial input of the target feature and sent to the convolutional neural network to learn the target feature. The convolutional neural network consists of N convolutional blocks and a final pooling layer. Each convolutional block consists of a convolutional layer, an activation layer exponential linear unit (ELU), and a batch normalization layer (BatchNorm).

[0063] Finally, it is connected to the pooling layer to obtain the final target representation V.

[0064] Step d: Send to the predictor to predict whether the drug and target will interact, including:

[0065] 1) First, the final drug representation V and target representation V in step c are concatenated to obtain the drug-target interaction representation V;

[0066] 2) Then V is sent to the multi-layer perceptron for interaction learning. The multi-layer perceptron is a fully connected network, as shown in the figure

[0067] As shown in Figure 2, the network consists of multiple fully connected layers and a final sigmoid activation layer. Except for the last layer, each fully connected layer is connected to a rectified linear unit and a dropout layer to prevent overfitting. Finally, the interaction probability is obtained.

[0068] A value greater than 0.5 indicates that the two are predicted to interact, while a value less than 0.5 indicates that they do not.

[0069] In this embodiment, the network parameters are optimized by the Adam optimizer, and the binary cross entropy is used as the loss function, which can be expressed as:

[0070]

[0071] Among them, M is the number of samples participating in training, p is the probability predicted by sample i, and y represents the true label of sample i.

[0072] Example 2

[0073] In this embodiment, Pytorch is used to implement the present invention, and experiments are conducted on four public datasets, namely C.elegans, Human, DAVIS, and BindingDB datasets. The data distribution is shown in Table 2.

[0074] Table 2

[0075]

[0076]

[0077] 1) The C. elegans and Human datasets contain the most detailed reactivity of known and clinical kinase inhibitors across the entire kinome and fully cover the human protein kinome. Their positive samples come from the DrugBank and Matador databases, two of the most confident biochemical databases, and contain 1,717 drugs and 1,857 targets, with a balanced positive and negative sample base.

[0078] 2) Davis MI, Hunt JP, Herrgard S et al. integrated various resources into a systematic screening framework, guided by the principle that “similar compounds may interact with proteins similar to corresponding known target proteins”.

[0079] Inspiration, we assume that proteins that are different from each known / predicted target of a compound are unlikely to be targeted by the compound. On the other hand, compounds that are dissimilar to any known / predicted compound targeting a protein are also unlikely to target this protein. We then screen out highly confident negative samples and construct the DAVIS dataset, which contains 64 drugs, 379 targets, and a serious imbalance between positive and negative samples.

[0080] 3) The BindingDB dataset comes from the BindingDB database. Its data is proposed from literature research and focuses on collecting protein information that serves as drug targets or candidate targets. The dataset contains 7,134 drugs, 1,251 targets, and an imbalance of positive and negative samples.

[0081] In the experiment, the number of convolution blocks N was set to 2, the optimizer learning rate was set to 1e-5, the batch size was 32, a maximum of 100 iterations were allowed, and early stopping was set. To prevent the model from overfitting, the loss rate was set to 0.5. In the experimental evaluation, AUC, Precision, and Recall were used as indicators to measure model performance. The present invention divided the model into training set, validation set, and test set in a ratio of 8:1:1. The model with the best AUC in the validation set was selected as the best model, and then a comparative evaluation was performed on the test set.

[0082] Table 3

[0083]

[0084]

[0085] Table 4

[0086]

[0087] Table 5

[0088]

[0089] Table 6

[0090]

[0091] The present invention provides a deep learning method for predicting drug-target interactions based on biological substructures, and conducts experiments on four public datasets to compare performance with other machine learning methods. Tables 3 and 4 are the comparative experimental results of the present invention on two balanced positive and negative sample datasets, C.elegans and Human, respectively. Tables 5 and 6 are the experimental results of the present invention on two uneven positive and negative sample datasets, DAVIS and BindingDB, respectively. From the above results, it can be seen that the present invention can achieve better prediction results regardless of whether the dataset distribution is balanced or uneven, indicating that the present invention is effective in modeling the drug-target interaction prediction task based on the idea of ​​substructure, which can explicitly enhance the pharmacological characteristics of drugs and targets, and also reflects the versatility of the present invention. In addition, compared with shallow machine learning methods or deep learning methods, its prediction performance can achieve leading results, indicating that the present invention uses convolutional neural networks to mine the synergistic effects between drug and target substructures, which has positive significance for predicting interactions.

[0092] The above embodiments are only intended to further illustrate the present invention and are not intended to limit the present invention. Any equivalent implementation of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A deep learning method for predicting drug-target interactions based on biological substructures, characterized by: The steps of this method are as follows: Step a: Input a drug's normalized SMILES sequence D and target amino acid sequence T; Step b: Extract substructures from the drug's normalized SMILES sequence D and the target amino acid sequence T, respectively, including: 1) For the normalized SMILES sequence D of the drug, the substructure of the drug is extracted using the BCM method, including: 1.1) First, extract the drug's side chains from the standardized SMILES sequence D of the drug. According to the definition rules of SMILES, the side chains are enclosed in brackets (()), and the remaining part is the main chain. 1.2) Then extract the common substructures in the main chain based on string matching; 1.3) Finally, the main chain is cleaved according to the retrosynthetic fragmentation rules of RECAP to obtain retrosynthetic fragments; 1.4) Integrate the branches, common substructures and retrosynthetic fragments as the substructure set F of the standardized SMILES sequence D D ; 2) For the target amino acid sequence T, the CFM method is used to extract substructures. The CFM method first divides amino acids into 8 categories according to chemical structure or properties, and obtains the category sequence T through category feature mapping. C ; Then use non-overlapping k-gram sequences to convert T C Functional substructure set F that is cut into targets T ; Step c: Construct a collaborative feature learning module, which consists of two parts: input representation and feature learning. Specifically, it includes: 1) In the input representation, the substructure set F of the normalized SMILES sequence D of the drug is D and the functional substructure set F of the target T Perform initial encoding representation, including: 1.1) Substructure set F of the normalized SMILES sequence D of the drug D First, label encoding is used to encode the drug substructure set F D Encode and obtain the initial representation of the drug I D ; Then I D Transformed into drug embedding representation E D ∈R max _drug_frag_length*embed_size , where max_drug_frag_length represents the size of the largest drug substructure set, and embed_size represents the embedding dimension; 1.2) Functional substructure set F of the target T To represent, first use label encoding to encode the target substructure set F T Encode and obtain the initial representation of the target I T ; Then I T Converted into its embedding representation E T ∈R max_target_frag_length*embed_size , where max_target_frag_length represents the size of the largest target substructure set, and embed_size represents the embedding dimension; 2) Feature learning is divided into drug feature learning and target feature learning, including: 2.1) In drug feature learning, E D As the initial input, it is sent to the convolutional neural network to learn drug features. The convolutional neural network consists of multiple convolution blocks and a final pooling layer. Each convolution block consists of a convolution layer, an activation layer exponential linear unit, and a batch normalization layer. Finally, the pooling layer is used to perform feature dimensionality reduction to obtain the final drug representation V D ; 2.2) In target feature learning, E T As the initial input, it is sent to the convolutional neural network to learn the target features. The convolutional neural network consists of multiple convolution blocks and a final pooling layer. Each convolution block consists of a convolution layer, an activation layer exponential linear unit, and a batch normalization layer. Finally, the pooling layer is connected to obtain the final target representation V T ; Step d: Build a predictor, including: 1) First, assemble the final drug representation V in step c D and target expression V T , the drug-target interaction representation V is obtained; 2) V is then fed into a multilayer perceptron for interaction learning. The multilayer perceptron is a fully connected network consisting of multiple fully connected layers and a final sigmoid activation layer. Except for the last layer, each fully connected layer is connected to a rectified linear unit and a dropout layer to prevent overfitting. Finally, the interaction prediction probability is obtained. A value greater than 0.5 indicates that the two are predicted to interact, and a value less than 0.5 indicates that they will not interact.