A method for predicting the binding affinity between drug molecules and target proteins
By constructing a multi-level graph structure of drug molecules and extracting feature embedding representations using neural network technology, the problem of high cost and low efficiency of predicting binding affinity between drug molecules and target proteins in the prior art is solved, and a high-accuracy prediction effect is achieved.
Patent Information
- Application Number
- CN202111291440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-02
AI Technical Summary
The existing methods for predicting binding affinity of drug molecules to target proteins have problems such as high cost, low efficiency and inability to identify binding affinity of new targets.
Using a multi-level graph division and neural network method, by obtaining the SMILES sequence of drug molecules and the amino acid sequence of target proteins, atom-based drug atom structure diagram and sub-structure-based drug sub-structure structure diagram are constructed, and feature embedding representations are extracted using graph convolution neural network and natural language processing technology, and splicing and prediction are combined with attention mechanisms.
High accuracy of predicting binding affinity of drug molecules to target proteins is achieved, reducing dependence on expert experience in drugs and biology fields, and being able to adaptively represent the characteristics of learning drugs and target proteins.
Smart Images

Figure CN113936735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of biological and pharmaceutical informatics and artificial intelligence, and in particular to a method for predicting the binding affinity between a drug molecule and a target protein based on multi-level graph partitioning and a neural network. Background Art
[0002] Drug discovery is the process of identifying new candidate drug compounds with potential therapeutic effects. In this process, predicting drug molecule-target protein interaction (DTI) research is an essential step. Proteins are important drug targets. Drugs play an important role in the human body by interacting with various targets, which can enhance or inhibit their functions and play a regulatory role to achieve the purpose of treating a certain disease. Therefore, identifying DTIs can help understand the mechanism of action of drugs and play a vital role in the discovery of new targets and drug repositioning. In the past few decades, high-throughput screening (HTS) experiments have greatly accelerated the identification of drug molecule-target protein interactions. However, HTS experiments are costly and laborious, and cannot meet the needs of revealing the interactions of millions of existing drug compounds and thousands of target proteins. Therefore, it is very necessary to develop efficient computational methods to make full use of the heterogeneous biological data of known drug molecule-target protein interactions to clarify the mechanism of action of drugs in the human body.
[0003] The rapid increase of DTI data in public databases, such as ChEMBL, DrugBank, and SuperTarget, has enabled large-scale identification of drug molecule-target protein interactions. Methods for calculating the binding affinity between drug molecules and target proteins can be mainly divided into three categories, namely docking-based, similarity search-based, and feature-based methods.
[0004] For docking-based methods, the three-dimensional structure of the target protein is used to simulate the binding position and orientation of the drug molecule and the target protein by considering various transformations and rotations of the target protein ligand to obtain different binding conformations. These methods predict effective compound-ligand binding by designing scoring functions to minimize binding free energy. The efficacy of docking methods depends on protein 3D structural information, while the 3D structure of many target proteins remains unknown, such as GPCRs. In addition, the simulation of the docking process is time-consuming and can only be used when the prediction scale is small.
[0005] Similarity search-based methods assume that small molecule compounds with similar structures or physicochemical properties can act on target proteins with the same or similar properties. Similarity search-based methods have been widely used in recent years due to the rapid increase in drug compound information and target protein annotation information in public databases. However, they are only suitable for predicting the binding of proteins similar to known target proteins, but cannot identify the binding affinity of new targets.
[0006] Unlike docking-based and similarity search-based methods, feature-based methods use various types of features extracted from drug compounds and target proteins, and mainly use machine learning models to predict the binding affinity of drug molecules and target proteins. Feature-based methods can be roughly divided into two categories. The first category uses collaborative matrix decomposition technology. This method decomposes the known drug molecule-target protein relationship matrix into two low-dimensional feature matrices representing the drug molecule and the target protein, respectively. Based on the drug molecule and target protein feature matrices, the similarity matrix of the drug molecule and the target protein can be estimated by taking the inner product of the feature vectors. Given the drug molecule-target protein relationship matrix and the similarity matrix of the drug molecule and the target protein, the potential drug molecule-target protein interaction can be inferred.
[0007] The second feature-based approach extracts feature descriptors of drug compounds and target proteins respectively, and models the prediction of drug molecule-target protein interaction as a binary classification (whether there is interaction) or regression problem (output is the predicted value of binding affinity). Molecular fingerprints are usually used as descriptors of drug substructures, while composition, transition and distribution (CTD) are usually used as protein descriptors.
[0008] In recent years, feature-based methods have been more widely used because they have almost no restrictions on the input information source. However, their performance depends largely on the initial feature representation of drug molecules and target proteins. In the existing initial feature descriptors of drug molecules and target proteins, molecular structure information is often missing, resulting in unsatisfactory prediction results.
[0009] Due to the great success of deep neural networks (DNN) in automatic feature learning of sequence data for image recognition and natural language processing, some deep learning models have also been proposed to predict the binding affinity between drug molecules and target proteins. By inputting the original drug molecules and target protein sequence data, DNN can extract useful information for prediction. Although deep learning methods have made progress in recent years, there is still a lot of room to improve the feature embedding representation of drug molecules and target proteins to enhance the prediction of the binding affinity between drug molecules and target proteins. Summary of the invention
[0010] The purpose of the present invention is to provide a method for predicting the binding affinity between a drug molecule and a target protein.
[0011] The method for predicting the binding affinity between a drug molecule and a target protein provided by the present invention helps research experts in the field of medicine to study the interaction between a drug molecule and a target protein and the strength of the binding affinity, and provides a theoretical and practical basis for the design of new drugs and the new use of old drugs in the downstream. For example, propranolol is a classic drug for the treatment of coronary heart disease and hypertension, and has recently been found to be useful for the treatment of osteoporosis and melanoma; cimetidine is a revolutionary drug for the treatment of peptic gastric ulcers, and has recently been used to treat chronic obstructive pulmonary disease, HIV virus infection, etc. Due to the great success of deep neural networks (DNNs) in automatic feature learning of sequence data for image recognition and natural language processing, the feature embedding representation of drug molecules can be well extracted using the attention mechanism and multi-layer graph convolutional neural networks; the feature embedding representation of the amino acid sequence of the target protein can be extracted using the language model in natural language processing and the one-dimensional convolutional neural network. The learned feature embedding representation of drug molecules integrates the knowledge of the drug field and the structural characteristics of the compound, and the feature embedding representation of the amino acid sequence of the target protein integrates the knowledge of the amino acid field and the protein structure information, which can make a more accurate prediction of the binding affinity between drug molecules and target proteins.
[0012] To achieve the above object, the present invention provides the following technical solution, comprising the following steps:
[0013] (1) Obtain the SMILES sequence of the drug molecule and the amino acid sequence of the target protein;
[0014] (2) For the SMILES sequence of drug molecules, it is represented as an atom-based drug atomic structure diagram and a substructure-based drug substructure structure diagram;
[0015] (3) Representation learning is performed on the drug atomic structure graph and the drug substructure graph respectively, so as to obtain the feature embedding representation of the drug atomic structure graph and the feature embedding representation of the drug substructure graph;
[0016] (4) For amino acid sequences, the feature embedding representation of amino acids is pre-trained using the language model in natural language processing, and then features are extracted to obtain the feature embedding representation of the amino acid sequence;
[0017] (5) concatenating the feature embedding representations of drug molecules and amino acids to obtain a concatenated embedding representation;
[0018] (6) Based on the splicing embedding representation, the predicted value of the binding affinity between the drug molecule and the target protein is obtained.
[0019] In step (1), the step of obtaining the SMILES sequence of the drug molecule and the amino acid sequence of the target protein comprises: obtaining the SMILES simplified molecular linear input canonical sequence of the drug molecule; and obtaining the amino acid sequence of the target protein.
[0020] In step (2), the drug molecule SMILES sequence is represented as an atom-based drug atomic structure diagram and a substructure-based drug substructure structure diagram, including: constructing an atom-based drug atomic structure diagram based on the drug molecule SMILES sequence; dividing the drug substructures based on the drug molecule SMILES sequence, and constructing a substructure-based drug substructure structure diagram.
[0021] Among them, based on the SMILES sequence of drug molecules, the drug substructures are divided, and the substructure-based drug substructure structure diagram is constructed, including: obtaining the drug atomic structure diagram of the drug molecule; numbering the atomic nodes in the drug atomic structure diagram; initializing the drug substructure set C as an empty set; constructing the set V1 as a set of all chemical bonds; constructing the set V2 as a set of all simple rings; if the chemical bond in V1 does not belong to any simple ring, add it to the drug substructure set C; loop through all the rings in V2, and merge the rings with more than or equal to 3 common atoms in V2 into new rings, until all the rings in V2 do not have three or more common atoms; add all the rings in V2 to the drug substructure set C; form the final drug substructure set C; based on the substructure set of drug molecules, construct a substructure-based drug substructure structure diagram.
[0022] In step (3), the drug atomic structure graph and the drug substructure structure graph are represented and learned respectively, so as to obtain the feature embedding representation of the drug atomic structure graph and the feature embedding representation of the drug substructure structure graph, including: using a deep learning neural network and an attention mechanism to extract the feature embedding representation of each atomic node in the drug atomic structure graph; performing a maximum pooling operation on the feature embedding representation of each atomic node in the drug atomic structure graph to obtain the feature embedding representation of the drug atomic structure graph; using a deep learning neural network and an attention mechanism to extract the feature embedding representation of each substructure node in the drug substructure structure graph; performing a maximum pooling operation on the feature embedding representation of each substructure node in the drug substructure structure graph to obtain the feature embedding representation of the drug substructure structure graph.
[0023] Among them, using a deep learning neural network and an attention mechanism to extract a feature embedding representation of each atomic node in a drug atomic structure graph includes: using an attention mechanism to extract the relative importance weight of each atomic node in the drug atomic structure graph; using a graph convolutional neural network of a deep learning neural network to extract an adjacency relationship representation between adjacent atomic nodes of the drug atomic structure graph and an initial feature representation of the atomic nodes; training the adjacency relationship representation and the initial feature representation, and using the trained embedding representation as the feature embedding representation of the atomic node.
[0024] The adjacency relationship representation and the initial feature representation are trained, and the trained embedding representation is used as the feature embedding representation of the atomic node, including: using the trained embedding representation as the initial feature representation of the atomic node, and continuously and cyclically executing the feature embedding representation of the atomic node of the drug atomic structure diagram extracted by using the graph convolutional neural network training; when the feature embedding representation of the atomic node is extracted cyclically for a specified number of times, the best training result is used as the feature embedding representation of the atomic node of the drug atomic structure diagram.
[0025] Among them, using a deep learning neural network and an attention mechanism to extract a feature embedding representation of each substructure node of a drug substructure structure graph includes: using an attention mechanism to extract the relative importance weight of each substructure node in the drug substructure structure graph; using a graph convolutional neural network of a deep learning neural network to extract an adjacency relationship representation between adjacent substructure nodes of the drug substructure structure graph and an initial feature representation of the substructure node; training the adjacency relationship representation and the initial feature representation, and using the trained embedding representation as the feature embedding representation of the substructure node.
[0026] Training the adjacency relationship representation and the initial feature representation, and using the trained embedding representation as the feature embedding representation of the substructure node includes: using the trained embedding representation as the initial feature representation of the substructure node, and continuously and cyclically executing the feature embedding representation of the substructure node of the drug substructure graph extracted by using the graph convolutional neural network training; after the feature embedding representation of the substructure node is cyclically extracted for a specified number of times, using the best training result as the feature embedding representation of the substructure node of the drug substructure graph.
[0027] In step (4), for the amino acid sequence, the feature embedding representation of the amino acid is pre-trained using the language model in natural language processing, and then feature extraction is performed on it to obtain the feature embedding representation of the amino acid sequence, including: unsupervised pre-training of the amino acid sequence of the target protein using the language model in natural language processing to obtain the initial feature representation of each amino acid; extracting the feature embedding representation of multiple amino acids using a one-dimensional convolutional network of deep learning; and performing a maximum pooling operation on the feature embedding representations of the multiple amino acids to obtain the feature embedding representation of the amino acid sequence of the target protein.
[0028] Among them, the amino acid sequence of the target protein is unsupervisedly pre-trained using the language model in natural language processing to obtain the initial feature representation of each amino acid, including: treating each amino acid in the target protein sequence as a word in the natural language processing text sequence and dividing the amino acid words; constructing a co-occurrence matrix of amino acid words; and training the regression method based on the least squares principle to obtain the initial feature embedding representation of each amino acid.
[0029] Among them, using a one-dimensional convolutional network of deep learning to extract feature embedding representations of multiple amino acids includes: inputting the initial feature embedding representation of each amino acid into the one-dimensional convolutional network for cyclic training to obtain feature embedding representations of multiple amino acids; when the feature embedding representations of multiple amino acids are cyclically extracted for a specified number of times, the last training result is used as the final feature embedding representation of multiple amino acids.
[0030] In step (5), the feature embedding representations of the drug molecule and the amino acid sequence are spliced to obtain the spliced embedded representation, which includes: finishing and splicing the feature embedding representation of the drug atomic structure graph, the feature embedding representation of the drug substructure graph, and the feature embedding representation of the amino acid sequence to obtain the spliced embedded representation.
[0031] In step (6), obtaining the binding affinity value between the drug molecule and the target protein based on the spliced embedded representation includes: inputting the spliced embedded representation into a multi-layer fully connected neural network to obtain a predicted value of the binding affinity between the drug molecule and the target protein.
[0032] The above technical scheme realizes the prediction of the binding affinity between drug molecules and target proteins with high prediction accuracy; at the same time, the entire technical scheme can adaptively realize the representation learning of drug molecule SMILES sequence and target protein sequence respectively, and can automatically learn their implicit feature representation without relying on a large amount of expert experience in the field of drugs and biology. The present invention reconstructs the drug molecule SMILES sequence into a drug atomic structure graph and a drug substructure structure graph, and uses a graph convolutional neural network to extract the node feature embedding representation of the drug molecule structure graph, thereby converting two-dimensional data into one-dimensional data. In order to utilize the rich structural information from drug molecules, the present invention models each drug molecule as a drug atomic structure graph and a drug substructure structure graph. At the same time, the present invention proposes an algorithm for segmenting substructures and extracting their features. Experimental results show that the representation of drug molecule structure graphs in two different dimensions helps to significantly improve the prediction ability of the binding affinity between drug molecules and target proteins. In order to make full use of the target protein sequence information, we use the word embedding method of the language model in natural language processing to pre-train the initial feature embedding representation of amino acids through a large corpus, and can learn the potential semantic correlation between amino acids. In addition, the present application further adopts a one-dimensional convolutional neural network to learn the high-level abstract feature embedding representation of proteins. At the same time, the present invention adds an attention mechanism before the graph convolutional neural network for drug molecules to identify important nodes in the drug atomic structure graph and the drug substructure structure graph and their interactions in the structure graph, which provides useful hints for drug discovery.
[0033] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow chart of the method for predicting the binding affinity between the drug molecule and the target protein of the present invention;
[0035] Figure 2 Indicates the different meanings of different chemical bonds in drug molecules in drug atomic structure diagrams;
[0036] Figure 3 This is an example diagram of the division of drug molecule substructures;
[0037] Figure 4 It is a flow chart of the algorithm for cutting the substructure of drug molecules;
[0038] Figure 5 It is a schematic diagram of attention propagation of drug atom / substructure structure diagram;
[0039] Figure 6 It is the GCN training flow chart of drug atom / substructure structure graph;
[0040] Figure 7 It is a model structure diagram of the EmbedDTI of the present invention;
[0041] Figure 8 It is the distribution diagram of the true binding affinity and the binding affinity predicted by EmbedDTI on the test set of the Davis dataset;
[0042] Fig. 9 This is the distribution of the true binding affinity and the binding affinity predicted by EmbedDTI on the test set of the KIBA dataset. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention is described more clearly and completely below in conjunction with the accompanying drawings.
[0044] like Figure 1 As shown, Figure 1 The flowchart of the method for predicting the binding affinity between drug molecules and target proteins based on multi-level graph partitioning and neural network of the present invention includes the following steps:
[0045] Step S11: Obtain the SMILES sequence of the drug molecule and the amino acid sequence of the target protein.
[0046] It costs approximately $2.6 billion to develop a new drug, and FDA approval can take up to 17 years. Finding new uses for approved drugs can avoid the expensive and lengthy drug development process. In order to effectively repurpose drugs, it is useful to know which proteins are targeted by which drugs. High-throughput screening experiments are often used to examine the binding affinity of drugs to their target proteins; however, these experiments are expensive and time-consuming, and exhaustive searches are not feasible because there are millions of drug-like compounds and hundreds of potential target proteins. Therefore, there is a strong motivation for the present invention to build computational models that can estimate the interaction strength of new drug molecule-target protein pairs based on previous drug molecule-target protein experiments.
[0047] SMILES (Simplified molecular input line entry specification), or simplified molecular linear input specification, is a specification that uses ASCII strings to clearly describe molecular structures. The SMILES sequence is a linear representation symbol for drug molecules, which is used to express the structure of a compound in a single line of text. It can represent information such as the atomic types of drug molecules and the connection relationship between atoms. Since SMILES uses a string of characters to describe a three-dimensional chemical structure, it must convert the chemical structure into a spanning tree, which uses a vertical priority traversal tree algorithm. During the conversion, hydrogen is removed first, and then the ring is opened. When representing, the atoms at the ends of the removed bonds are marked with numbers, and the side chains are written in parentheses. The open source chemical information software RDKit can be used to convert the SMILES sequence of drug molecules into a structural diagram of the drug molecules.
[0048] Each protein sequence is composed of more than 20 kinds of amino acid combinations. The amino acid sequence contains information about the types of amino acids and the positional relationship between amino acids, and is also the primary amino acid sequence representation of the protein.
[0049] Step S12: For the drug molecule SMILES sequence, it is represented as an atom-based drug atomic structure diagram and a substructure-based drug substructure structure diagram. Specifically, it includes the following steps:
[0050] The atomic structure diagram of the drug molecule can be converted from the SMILES string. Compounds are usually represented as graph-structured data in computers, where the vertices and edges of the graph correspond to atoms and chemical bonds, respectively, which is consistent with the atomic structure diagram of the drug molecule. The functions provided by the open source chemical information software RDKit can convert the SMILES sequence of the drug molecule into the atomic structure diagram of the drug molecule and number each atom in the diagram.
[0051] The atomic structure graph of drug molecules can represent the structural information between atoms at short distances, but ignores the substructures in the compound, which play an important role in determining the properties and reactions of the compound. For example, a single atom in a benzene ring can understand the information of its neighboring atoms, but it is difficult to understand the structure of the entire benzene ring as a whole and the role it plays in the drug molecule. Therefore, we define substructures and convert the original atom-based drug atomic structure graph into a higher-level drug substructure structure graph, where the nodes and edges in the drug substructure structure graph correspond to substructures and connections between substructures, respectively.
[0052] A major limitation of the atomic structure graph of drug molecules is that it treats all edges equally and extracts information from a single vertex, whereas atoms and their associated edges usually work in pairs. Figure 2 For example, the chemical bond represented by S22 is very important to the entire molecule. Therefore, the independent chemical bonds in the drug molecule play a key role in the structure and chemical properties of the entire molecule. However, if separated from the benzene ring substructure S23, the chemical bond represented by S21 is meaningless in terms of structure and chemical properties.
[0053] Therefore, the present invention proposes a method for segmenting drug molecules and obtains a complete substructure set of the drug to ensure that all compounds in the database can be composed of substructures in the set.
[0054] like Figure 3 As shown, we cut the entire drug molecule into substructures. Figure 3 The left figure is an atom-based drug atomic structure diagram, in which the substructures of the drug molecules are marked with dotted ovals, as shown in S1-S15; the right figure is a substructure-based drug substructure structure diagram, in which each substructure is represented by a node in the diagram; S1-S15 on the left corresponds one-to-one to S1-S15 on the right. Therefore, the substructure consists of a cyclic substructure or a pair of atoms connected by chemical bonds that do not belong to the ring. In this way, the drug molecule compound can be regarded as a topological graph connected by substructures. The substructure cutting algorithm is Figure 4 The main steps include:
[0055] Step S41: obtaining a drug atomic structure diagram of a drug molecule;
[0056] Step S42: numbering the atomic nodes in the drug atomic structure graph;
[0057] Step S43: Initialize the drug substructure set C to an empty set;
[0058] Step S44: constructing a set V1 as a set of all chemical bonds;
[0059] Step S45: constructing a set V2 as a set of all simple rings;
[0060] Step S46: If the chemical bond in V1 does not belong to any simple ring, add it to the drug substructure set C;
[0061] Step S47: loop through all the rings in V2, and merge the rings in V2 with more than or equal to 3 common atoms into new rings, until all the rings in V2 do not have three or more common atoms;
[0062] Step S48: adding all the rings in V2 to the drug substructure set C;
[0063] Step S49: forming the final drug substructure set C.
[0064] The drug atomic structure diagram is obtained by the Chem.MolFromSmiles function in RDKit. V1 and V2 represent the set of independent chemical bonds and the set of simple rings, respectively. Independent chemical bonds are extracted from the GetBonds function of RDKit, while simple ring structures are extracted from the Chem.GetSymmSSSR function of RDKit. Finally, we established a substructure set consisting of chemical bonds that do not belong to any ring and independent rings that share less than 3 atoms with other rings.
[0065] Step S13: Representation learning is performed on the drug atomic structure graph and the drug substructure structure graph respectively, so as to obtain the feature embedding representation of the drug atomic structure graph and the feature embedding representation of the drug substructure structure graph. Specifically, learning the feature representation of the drug atomic structure graph includes the following steps:
[0066] (1) Using deep learning neural networks and attention mechanisms to extract feature embedding representations of each atomic node in the drug atomic structure graph;
[0067] Convolutional neural networks (CNNs) have not only achieved great success in computer vision and natural language processing, but also have shown good performance in various graph-related learning tasks. The difference is that the nodes in the graph are located in a non-Euclidean space, and the introduction of graph convolutional neural networks (GCNs) aims to capture the local correlation of node signals on the graph. Since drug molecular compounds can be represented in the form of graphs, GCNs are used in the present invention to learn the feature embedding representation of drug atomic structure graphs.
[0068] For a graph G = (V, E), V is the set of all nodes (atoms in drug molecules) in the graph, and E is the set of all edges (chemical bonds in drug molecules) in the graph. Each node i has its characteristic x i , all node features x iA feature matrix Wherein N represents the number of nodes, and d represents the number of features of each node, that is, the dimension of each node feature vector. The present invention utilizes the open source chemical information software RDKit to extract the initial features of each atomic node. The initial features of each atomic node are represented as a one-hot feature vector, which contains eight kinds of information (atomic symbol, degree of atom in drug molecule, total number of explicit and implicit hydrogen atoms connected to atom, number of implicitly connected hydrogen atoms, total valence of explicit and implicit atoms, charge number of atom, whether atom belongs to aromatic family, and whether atom is in ring), so as to obtain a 101-dimensional one-hot initial feature vector representation for each atomic node.
[0069] The connection relationship between the drug atomic nodes in the graph forms an N×N dimensional adjacency matrix A. The initial feature matrix of the atomic node is and the adjacency matrix is the input of GCN. The propagation between GCN layers can be expressed as Indicates that is the matrix obtained by adding the adjacency matrix A to the self-connection edges, yes The degree matrix, H (l) represents the atomic node feature embedding representation matrix of the l-th layer GCN, σ() is an activation function (such as ReLU function). For the input layer GCN, H (0) is equal to the initial feature matrix X.
[0070] In addition, before the adjacency matrix and initial feature matrix of the drug atomic structure graph are input, we add an attention mechanism matrix W to help learn the relative importance weight of each atomic node in the structure graph. (0) Equal to W×X. Figure 5 The propagation process of the attention mechanism is illustrated.
[0071] (2) performing a maximum pooling operation on the feature embedding representation of each atomic node to obtain a feature embedding representation of the drug atomic structure graph;
[0072] The GCN model learns the atomic node feature embedding matrix of drug molecules to represent the output Where F represents the number of filters in the graph convolution layer. In order to obtain the feature embedding representation of the drug atomic structure graph, the present invention adds a maximum pooling layer after the GCN layer. Similar to the pooling operation in traditional CNN, maximum pooling is a reasonable reduction of the graph.
[0073] The feature embedding representation of the atomic nodes of the drug atomic structure diagram is extracted by continuously executing the graph convolutional neural network training. After the feature embedding representation of the atomic nodes is extracted cyclically for a specified number of times, the best training result is used as the feature embedding representation of the atomic nodes of the drug atomic structure diagram.
[0074] (3) Convert the feature embedding representation of the drug atomic structure graph into a 128-dimensional feature embedding representation vector of the drug atomic structure graph.
[0075] The feature embedding representation of the drug atomic structure graph is input into two fully connected layer neural networks for training, and a 128-dimensional feature embedding representation vector of the drug atomic structure graph is obtained as an output. Figure 6 The GCN learning process of the drug atomic structure graph is shown.
[0076] Learning the feature representation of drug molecules from drug substructure graphs includes the following steps:
[0077] (1) Using deep learning neural networks and attention mechanisms to extract feature embedding representations of each substructure node in the drug substructure graph;
[0078] The present invention also uses GCN to learn the features of drug substructure graphs. Similar to the initial feature vector representation of the atomic nodes extracted from the drug atomic structure graph, the present invention uses the open source chemical information software RDKit to extract the features of each substructure node. Each substructure node feature is represented as a one-hot feature vector, which contains five types of information based on graph theory (the number of atoms, the number of edges connecting to other substructures, the number of explicit and implicit hydrogen atoms, whether it contains a ring, and whether it contains chemical bonds that do not belong to a simple ring), thereby obtaining a 35-dimensional one-hot vector initial feature representation for each substructure node.
[0079] The process of training and learning the feature embedding representation of each substructure node in the drug substructure graph is consistent with the process of obtaining the feature embedding representation of each atomic node in the drug atomic structure graph.
[0080] (2) performing a maximum pooling operation on the feature embedding representation of each substructure node to obtain a feature embedding representation of the drug substructure graph.
[0081] The feature embedding representation of the drug substructure graph is consistent with the feature embedding representation of the drug atomic structure graph.
[0082] (3) Convert the feature embedding representation of the drug substructure graph into a 128-dimensional feature embedding representation vector of the drug substructure graph.
[0083] The feature embedding representation of the drug substructure diagram is input into two fully connected layer neural networks for training, and a 128-dimensional feature embedding representation vector of the drug substructure diagram is obtained as an output.
[0084] Step S14: For the amino acid sequence, the feature embedding representation of the amino acid is pre-trained using the language model in natural language processing, and then feature extraction is performed on it, so as to obtain the feature embedding representation of the amino acid sequence. Specifically, the following steps are included:
[0085] (1) using the language model in natural language processing to perform unsupervised pre-training on the amino acid sequence of the target protein to obtain an initial feature embedding representation of each amino acid;
[0086] In the present invention, the input features of proteins are extracted from amino acid sequences. In order to obtain a good representation of amino acid sequences, we pre-trained the large protein database UniRef50 using the word embedding technology of the language model in natural language processing, and obtained the initial feature embedding vector of amino acids. The UniProt reference cluster UniRef provides a set of sequence clusters from the UniProt knowledge base (including isoforms) and selected UniParc records to obtain complete coverage of the sequence space at multiple resolutions while hiding redundant sequences (but not their description information). Unlike UniParc, sequence fragments in UniRef are merged: the UniRef100 database merges identical sequences and subfragments with 11 or more residues from any organism into one UniRef entry, showing the sequence of representative proteins, all merged accession number entries and links to the corresponding UniProtKB and UniParc records. UniRef90 was constructed by clustering UniRef100 sequences with 11 or more residues using the MMseqs2 algorithm, such that each cluster was constructed from sequences with at least 90% sequence identity and 80% overlap with the longest sequence of the cluster (also known as the seed sequence). Similarly, UniRef50 was constructed by clustering UniRef90 seed sequences that had at least 50% sequence identity and 80% overlap with the longest sequence in the cluster. UniRef90 and UniRef50 reduce the database size by approximately 58% and 79%, respectively, thereby providing faster sequence similarity searches. The UniRef50 database is used in the present invention as a pre-trained corpus, comprising 48,524,161 amino acid sequences.
[0087] The present invention uses the GloVe model to obtain the initial feature embedding representation of amino acids. GloVe is an unsupervised model that can learn fixed-length feature vector representations from variable-length texts. It is based on the global word-word co-occurrence statistics of the corpus. In the present invention, each amino acid is regarded as a word in a natural language processing language model, and the initial feature embedding representation of each amino acid is obtained. i .
[0088] (2) Using a one-dimensional convolutional network with deep learning to extract feature embedding representations of multiple amino acids;
[0089] The initial feature embedding representation of all amino acids is e i The initial feature embedding matrix E of the entire amino acid sequence is constructed. The initial feature embedding matrix E of the amino acid sequence is used as the input of a deep convolutional neural network (CNN) for further feature representation learning. The present invention adopts a one-dimensional CNN model (i.e., TextCNN) to extract local sequence features through convolution kernels operating near amino acids, so that the amino acid feature embedding representation obtained by training can contain the structure and chemical property information of the amino acid. Through the convolution operation of multiple convolution kernels, the feature embedding representation of multiple amino acids is obtained.
[0090] (3) performing a maximum pooling operation on the feature embedding representations of the multiple amino acids to obtain a feature embedding representation of the amino acid sequence of the target protein.
[0091] Convolution operations with convolution kernels of different sizes will result in different feature embedding representations of multiple amino acids. After the feature embedding representations of multiple amino acids are aggregated, a maximum pooling operation is performed on them to obtain a feature embedding representation of the amino acid sequence of the target protein.
[0092] (4) Converting the feature embedding representation of the amino acid sequence of the target protein into a 128-dimensional target protein feature embedding representation vector.
[0093] The feature embedding representation of the amino acid sequence of the target protein is input into a fully connected layer neural network for training, and a 128-dimensional target protein feature embedding representation vector is output.
[0094] Step S15: performing splicing based on the feature embedding representations of the drug molecule and the amino acid sequence to obtain a spliced embedding representation;
[0095] The 128-dimensional drug atomic structure graph feature embedding representation vector, the 128-dimensional drug substructure graph feature embedding representation vector and the 128-dimensional target protein feature embedding representation vector are concatenated end to end to obtain a 384-dimensional concatenated embedding representation vector.
[0096] Step S16: based on the spliced embedded representation, obtaining the binding affinity value between the drug molecule and the target protein;
[0097] The concatenated embedding representation vector is input into a three-layer fully connected neural network to obtain the binding affinity value between the drug molecule and the target protein. The three-layer fully connected neural network has 1024, 512, and 1 hidden units respectively.
[0098] The present invention proposes an EmbedDTI model to predict the binding affinity value between drug molecules and target proteins. The model extracts features of drug molecules and target proteins and implicitly embeds them to finally complete the prediction of the binding affinity between drug molecules and target proteins. The whole process is steps S11-S16 described above. The model structure diagram is shown in Figure 7 As shown. A one-dimensional SMILES sequence of a drug molecule and a one-dimensional amino acid sequence of a target protein are used as inputs of the EmbedDTI model. For the amino acid sequence of the target protein, the present invention first uses a language model of natural language processing to pre-train a single amino acid in the corpus to obtain an initial feature embedding representation of each amino acid; then, the amino acid sequence is represented as an initial feature matrix consisting of an initial feature embedding representation vector of each amino acid, which is input into a three-layer TextCNN network for feature extraction, and then a 128-dimensional target protein feature embedding representation is obtained through a fully connected layer. For the SMILES sequence of the drug molecule, the open source chemical information software RDKit is first used to represent it as an atom-based drug atomic structure diagram and a substructure-based drug substructure structure diagram. For the drug atomic structure diagram, the adjacency matrix and initial feature matrix between atoms are obtained, and the attention mechanism is used to help learn the relative importance weight of each atomic node; the adjacency matrix and initial feature matrix of the drug atomic structure diagram are input into the three-layer graph convolutional neural network for feature extraction to obtain the feature embedding representation of each atomic node; the feature embedding representation of the atomic nodes is fused through the maximum pooling operation to obtain the feature embedding representation of the drug atomic structure diagram, and finally the standard 128-dimensional feature embedding representation of the drug atomic structure diagram is obtained through two fully connected layers. For the drug substructure diagram, the 128-dimensional feature embedding representation of the standard drug substructure diagram is obtained by training and compared with the drug atomic structure diagram. Figure 1 The difference lies in the adjacency matrix and initial feature matrix of the input graph convolutional neural network. Finally, the 128-dimensional feature embedding representation of the drug atomic structure graph, the 128-dimensional feature embedding representation of the drug substructure graph, and the 128-dimensional feature embedding representation of the target protein are concatenated head to tail, and the obtained concatenated embedding representation is passed through three fully connected layers to obtain the predicted binding affinity value of the drug molecule and the target protein.
[0099] Experimental datasets: In this paper, the EmbedDTI model is evaluated on two benchmark datasets, namely Kinase Davis and KIBA datasets. Binding affinity provides specific information about the interaction between drug molecules and target protein pairs. It can be expressed by the half maximal inhibitory concentration (IC 50 ), dissociation constant (K d ), inhibition constant (K i ) and the binding constant (K a ) and other indicators. 50 Represents the concentration of a drug or inhibitor required to inhibit a specified biological process (or a component of a process, such as an enzyme, receptor, cell, etc.) by half. i It reflects the inhibitory strength of the inhibitor on the target protein. The smaller the value, the stronger the inhibitory ability. d Reflects the affinity of the drug compound for the target protein. The smaller the value, the stronger the affinity. In some cases, it is equivalent to K i . K a It's K d Therefore, K a The larger the value of , the stronger the binding affinity. Following the practice of previous studies, the present invention uses logarithmic transformation of K d ,Right now As model output, the Davis dataset collects clinically relevant kinase protein families and related inhibitors and their respective dissociation constants K d The KIBA dataset is a more general dataset and much larger than Davis. The Davis dataset contains 30,056 drug-target protein pair interactions, covering 442 target proteins and 68 drug compound molecules. In the Davis dataset, only K d It is used to measure the biological activity of kinase inhibitors; KIBA binds K i , K d and IC 50 The KIBA scores of protein families and related inhibitors were obtained. There are 229 proteins and 2111 drug compounds involved in the KIBA dataset. Table 1 summarizes the number of target proteins, drug molecules, and drug molecule-target protein pair interactions in the two datasets.
[0100] Table 1 Summary of Davis and KIBA datasets
[0101] Dataset Number of drug molecules Number of target proteins Number of drug-target interactions Davis 68 442 30,056 KIBA 2,111 229 118,254
[0102] Experimental setup: The present invention evaluates the performance of EmbedDTI on two benchmark sets, the Davis dataset and the KIBA dataset. For each dataset, the present invention divides it into 6 equal parts, one part is used as an independent test set, and the remaining five parts are used for training. The present invention performs five-fold cross-validation in the training set to search for the best hyperparameters. For each hyperparameter, a grid search is used to narrow the search range to the neighborhood of the optimal parameter, and then a refined search is performed. In the process of feature extraction, we used three filter convolution layers of different sizes for the target protein; the GCN used to learn the drug atomic structure graph and the drug substructure structure graph also contains three graph convolution layers. The parameters of the model training process are shown in Table 2.
[0103] Table 2 Parameter settings of EmbedDTI
[0104] parameter Setting value Batch size 512 Learning rate 0.0005 Epoch 1500 Dropout 0.2 Optimizer Adam Number of filters in three-layer CNN 1000,256,32 Filter size of three-layer CNN 8,8,3 Input dimension of three-layer GCN N, N, 2N Output dimension of three-layer GCN N, 2N, 4N The number of hidden units in the three fully connected layers after concatenation 1024,512,1
[0105] * Where N represents the dimension of the input feature vector
[0106] Experimental analysis: Since the present invention regards DTI as a regression problem of predicting the binding affinity between drug molecules and target protein pairs, we use mean square error (MSE) as the loss function. MSE measures the difference between the predicted value (P) and the true value (Y) of the target variable. The smaller the MSE, the closer the predicted value is to the true value, and vice versa. Where N represents the number of samples. In addition, another indicator used to evaluate the performance is the consistency index (CI), which is used to calculate the difference between the predicted value of the model and the true value. where b x is relative to the true maximum binding affinity δ x The predicted binding affinity of y is relative to the actual smaller binding affinity δ y The predicted binding affinity of , h(x) is a step equation, Z is a normalization constant used to map values to the interval [0,1]. The CI metric measures whether the predicted affinity values of two randomly selected drug-target pairs maintain a similar relative order in the real dataset. The larger the CI value, the better the result.
[0107] To evaluate the performance of EmbedDTI, we compare it with five state-of-the-art models listed below.
[0108] KronRLS: It uses the Smith-Waterman algorithm to calculate the similarity between proteins and the PubChem structural clustering service to calculate the similarity between drug compounds. It then uses a kernel-based approach to calculate the Kronecker product and integrates multiple heterogeneous information sources within a least squares regression (RLS) framework.
[0109] SimBoost: It has the same representation of proteins and drug compounds as KronRLS. It builds features for drugs, targets, and drug-target pairs, and extracts feature vectors of drug-target pairs through feature engineering to train gradient boosting machines to predict binding affinity.
[0110] DeepDTA: It encodes the original one-dimensional protein sequence and drug analysis SMILES sequence. The encoded vector is passed through two independent CNN modules to obtain the corresponding representation vector, which is then concatenated and output through a fully connected layer to predict the binding affinity.
[0111] WideDTA: It adds protein domain and membrane information, as well as the largest common substructure words, based on DeepDTA. Together with the original information, there are four parts to train the model together.
[0112] GraphDTA: It uses TextCNN to learn features of one-dimensional protein sequences. For drug molecule SMILES sequences, it uses four models: GCN, GAT, GIN, and GAT_GCN to obtain the representation vector of the SMILES sequence.
[0113] Furthermore, we perform an ablation study on EmbedDTI by comparing three variants of EmbedDTI, namely, EmbedDTI_noPre, EmbedDTI_noClq, and EmbedDTI_noAttn.
[0114] (1) EmbedDTI_noPre: No GloVe pre-training was performed on the amino acid sequence of the target protein;
[0115] (2) EmbedDTI_noClq: There is no drug substructure diagram representation, that is, for the drug molecule sequence, it is only represented as an atom-based drug atomic structure diagram;
[0116] (3) EmbedDTI_noAttn: No attention module is added to the GCN. That is, the initial feature representation matrix is directly input into the GCN without considering the relative importance weight of the nodes in the graph.
[0117] Table 3 shows the MSE and CI scores of EmbedDTI on the independent Davis test dataset compared with 5 baseline models. It can be seen that EmbedDTI achieves the lowest MSE and highest CI, with a 9.5% reduction in MSE and a 2.5% improvement in CI compared to the state-of-the-art method GraphDTA. The performance improvement can be attributed to the following three factors.
[0118] First, we use graphs to represent compounds, which preserves more structural information than raw sequence-based methods. In addition, we represent compounds with two graph structures, preserving more structural and functional information at the atomic and substructure levels, rather than using only one atom-based drug atomic structure graph in most existing methods such as GraphDTA.
[0119] Secondly, the attention mechanism before GCN helps to learn the relative importance weights of nodes (atoms or substructures). By outputting the attention score of each node, we can observe the focus nodes that the model pays attention to.
[0120] Third, pre-training improves the representation of the amino acid sequence of the target protein by introducing some prior background knowledge, which also improves the overall performance of EmbedDTI. The predicted binding affinity and the true binding affinity are plotted in Figure 8 It can be observed that most of the points are close to the line x=y, indicating that the EmbedDTI model is more accurate in predicting the binding affinity between drug molecules and target proteins on the Davis dataset.
[0121] Table 4 MSE and CI scores of EmbedDTI and other 5 baseline models on the KIBA dataset. Although the data scale of KIBA is much larger than that of Davis, the performance of these models follows the same trend as on the Davis dataset. The drug representation based on two structural graphs of atoms and substructures greatly improves the performance, with an MSE improvement of 0.268 over WideDTA, an improvement of 0.058 over GraphDTA, and a CI improvement of 0.012 over GraphDTA. The predicted binding affinities and true binding affinities are plotted in Fig. 9 It can be observed that most of the points are close to the line x=y, indicating that the EmbedDTI model is more accurate in predicting the binding affinity between drug molecules and target protein pairs on the KIBA dataset.
[0122] Table 3 Comparison results with the baseline model on the Davis dataset
[0123]
[0124] Table 4 Comparison results with the baseline model on the KIBA dataset
[0125]
[0126] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for predicting the binding affinity between a drug molecule and a target protein, characterized in that: The following steps are involved: Obtain the SMILES sequence of the drug molecule and the amino acid sequence of the target protein; For the drug molecule SMILES sequence, it is represented as an atom-based drug atomic structure diagram and a substructure-based drug substructure structure diagram; this step includes: constructing an atom-based drug atomic structure diagram based on the drug molecule SMILES sequence; dividing the drug substructure based on the drug molecule SMILES sequence, and constructing a substructure-based drug substructure structure diagram; Representation learning is performed on the drug atomic structure graph and the drug substructure structure graph respectively, so as to obtain the feature embedding representation of the drug atomic structure graph and the feature embedding representation of the drug substructure structure graph; For amino acid sequences, the language model in natural language processing is used to pre-train the feature embedding representation of amino acids, and then feature extraction is performed on them to obtain the feature embedding representation of the amino acid sequence; Concatenate the feature embedding representations of drug molecules and amino acids to obtain a concatenated embedding feature representation; Based on the splicing embedding representation, the predicted value of the binding affinity between the drug molecule and the target protein is obtained; Based on the SMILES sequence of drug molecules, the drug substructures are divided and a substructure-based drug substructure diagram is constructed, including the following steps: Obtaining a drug atomic structure diagram of the drug molecule; Numbering the atomic nodes in the drug atomic structure diagram; Initialize the drug substructure set C to an empty set; Construct set V1 as the set of all chemical bonds; Construct set V2 as the set of all simple rings; If the chemical bond in V1 does not belong to any simple ring, add it to the drug substructure set C; Loop through all the rings in V2, merge the rings in V2 with more than 3 common atoms into new rings, until all the rings in V2 do not have three or more common atoms; Add all the rings in V2 to the drug substructure set C; Form the final drug substructure set C; Based on the substructure set of drug molecules, a substructure-based drug substructure diagram is constructed.
2. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that: Obtaining the SMILES sequence of the drug molecule and the amino acid sequence of the target protein includes the following steps: Obtaining the SMILES simplified molecular linear input specification sequence of the drug molecule; Obtain the amino acid sequence of the target protein.
3. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that: Representation learning is performed on the drug atomic structure graph and the drug substructure structure graph respectively, so as to obtain the feature embedding representation of the drug atomic structure graph and the feature embedding representation of the drug substructure structure graph, including the following steps: The attention mechanism is used to extract the relative importance weight of each atomic node in the atomic structure graph of the drug; A graph convolutional neural network of a deep learning neural network is used to extract adjacency relationship representations between adjacent atomic nodes of the drug atomic structure graph and initial feature representations of the atomic nodes; The trained embedding representation is used as the initial feature representation of the atomic node, and the feature embedding representation of the atomic node of the drug atomic structure graph is extracted by using the graph convolutional neural network training in a continuous loop; When the feature embedding representation of the atomic node is extracted cyclically for a specified number of times, the best training result is used as the feature embedding representation of the atomic node of the drug atomic structure graph; Performing a maximum pooling operation on the feature embedding representation of each atomic node in the drug atomic structure graph to obtain a feature embedding representation of the drug atomic structure graph; The attention mechanism is used to extract the relative importance weight of each substructure node in the drug substructure graph; A graph convolutional neural network of a deep learning neural network is used to extract adjacency relationship representations between adjacent substructure nodes of the drug substructure structure graph and initial feature representations of the substructure nodes; The trained embedding representation is used as the initial feature representation of the substructure node, and the feature embedding representation of the substructure node of the drug substructure graph is extracted by using the graph convolutional neural network training in a continuous loop; After the feature embedding representation of the substructure node is extracted cyclically for a specified number of times, the best training result is used as the feature embedding representation of the substructure node of the drug substructure graph; A maximum pooling operation is performed on the feature embedding representation of each substructure node in the drug substructure graph to obtain a feature embedding representation of the drug substructure graph.
4. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that: For the amino acid sequence, the feature embedding representation of the amino acid is pre-trained using the language model in natural language processing, and then feature extraction is performed on it to obtain the feature embedding representation of the amino acid sequence, including the following steps: Use the language model in natural language processing to perform unsupervised pre-training on the amino acid sequence of the target protein to obtain the initial feature representation of each amino acid; A one-dimensional convolutional network based on deep learning is used to extract feature embedding representations of multiple amino acids; The feature embedding representations of multiple amino acids are subjected to a maximum pooling operation to obtain a feature embedding representation of the amino acid sequence of the target protein.
5. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 4, characterized in that: The language model in natural language processing is used to perform unsupervised pre-training on the amino acid sequence of the target protein to obtain the initial feature representation of each amino acid, including the following steps: Treat each amino acid in the target protein sequence as a word in the natural language processing text sequence and perform amino acid word segmentation; Construct a co-occurrence matrix of amino acid words; The regression method based on the least squares principle is trained to obtain the initial feature embedding representation of each amino acid.
6. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 4, characterized in that: The feature embedding representation of multiple amino acids is extracted using a one-dimensional convolutional network of deep learning, including the following steps: The initial feature embedding representation of each amino acid is input into a one-dimensional convolutional network for cyclic training to obtain feature embedding representations of multiple amino acids; After the feature embedding representations of multiple amino acids are extracted cyclically for a specified number of times, the last training result is used as the final feature embedding representation of multiple amino acids.
7. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that: The feature embedding representations of drug molecules and amino acid sequences are concatenated to obtain a concatenated embedding representation, including the following steps: The feature embedding representation of the drug atomic structure graph, the feature embedding representation of the drug substructure graph, and the feature embedding representation of the amino acid sequence are concatenated head to tail to obtain a concatenated embedding representation.
8. The method for predicting the binding affinity between a drug molecule and a target protein according to claim 1, characterized in that: Based on the splicing embedding representation, the binding affinity value of the drug molecule to the target protein is obtained, including the following steps: The spliced embedding representation is input into a multi-layer fully connected neural network to obtain the predicted value of the binding affinity of the drug molecule to the target protein.
Citation Information
Patent Citations
Transform and graph neural network-combined drug target interaction prediction method
CN116417093A
RNA drug binding prediction method based on subgraph matching and domain adaptation
CN118116508A
Drug target binding affinity prediction method based on drug bimodal characteristics
CN118298908A