DTA prediction method based on multi-modal neural network model
By adopting a multimodal neural network model in drug-target affinity prediction, combining GCN and CNN to extract features, and fusion of features through attention mechanisms, the problem of insufficient feature extraction in the existing technology is solved, and the prediction accuracy and model generalization ability are improved.
Patent Information
- Application Number
- CN202510066521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The prior art has problems such as strong data dependence, limited prediction accuracy, and insufficient model generalization ability in drug-target affinity prediction, making it difficult to effectively extract the comprehensive characteristics of drug molecules and proteins.
The DTA prediction method based on the multimodal neural network model is adopted to extract the characteristics of the drug and protein sequences through GCN and CNN respectively, and the characteristics are fused through the attention mechanism to form a more complete protein feature representation.
The accuracy of drug-target affinity prediction and the generalization ability of the model are improved, and efficient prediction of the interaction intensity or binding ability of the drug and target protein is achieved.
Smart Images

Figure CN119993272A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug discovery, and in particular to a DTA prediction method based on a multimodal neural network model. Background Art
[0002] Drug discovery is a long and complex process, including target screening, lead compound discovery, preclinical research, and clinical trials. It often takes a large number of potential compounds to find a molecule that effectively binds to the target protein. Although the traditional drug development process has made great contributions to the development of medicine, its high cost, low success rate, and long R&D cycle have become the key bottlenecks that limit the efficiency of new drug development. According to statistics, the process from new drug development to final market launch takes an average of more than 10 years, and the total cost can be as high as tens of billions of RMB. DTA prediction is a key link in drug discovery. Its purpose is to predict the interaction strength or binding ability between drugs and target proteins through computational models or algorithms, also known as binding affinity prediction. This work helps to understand the mechanism of drug action, improve drug design, and accelerate the drug discovery process.
[0003] Relying on the progress of bioinformatics, high-tech technologies such as computational modeling, machine learning and deep learning have been introduced into drug-target affinity prediction, which has greatly reduced experimental costs and improved the ability of high-throughput prediction through computer methods. Early computational methods, such as molecular docking and collaborative filtering technology, have improved the efficiency of affinity prediction to a certain extent, but there are still problems such as strong data dependence, limited prediction accuracy, and insufficient model generalization ability. In recent years, deep learning technology has brought new opportunities for drug-target affinity prediction tasks and has shown great potential in this field. The modeling method based on neural networks can effectively explore the potential interaction relationship between drug and target data and show excellent performance on multiple standard data sets.
[0004] Mainstream biological sequence feature extraction methods include MLP, CNN, LSTM, etc. On the one hand, the application of these methods alone may not be sufficient to extract the comprehensive features of drug molecules and proteins. Among them, 1D convolution and 2D convolution mainly rely on fixed-size convolution kernels. Although too small convolution kernels can capture local patterns of biological sequences, they ignore global context information. Although 2D convolution can expand the receptive field by stacking convolution layers to obtain global information, longer biological sequences increase the computational cost. LSTM controls the transmission and forgetting of information through its gate mechanism, so it can retain local features or short-term dependency information of the sequence. Although LSTM can retain longer dependency information through certain designs, its long-distance dependency ability is still limited. In contrast, the fully connected structure of MLP means that all features of the input biological sequence will participate in the calculation together, so although the global features of the sequence can be calculated, its local features cannot be taken into account. On the other hand, there is currently no recognized and effective protein sequence graph conversion method that can convert protein sequences from sequence form to graph structure while retaining their global and local patterns. At present, the commonly used graph conversion method is to represent it according to the graph structure mode: nodes represent amino acids, and edges represent the interactions between amino acids. However, the sequence information of proteins is linear, and the order of amino acids is crucial. GCN is mainly used for graph structure data, focusing on the topological relationship between nodes rather than directly modeling the order information of nodes. Therefore, it is very necessary to design a DTA prediction method based on a multimodal neural network model. Summary of the invention
[0005] In order to overcome the deficiencies of the prior art, an object of the present invention is to provide a DTA prediction method based on a multimodal neural network model.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] The present invention provides a DTA prediction method based on a multimodal neural network model, comprising:
[0008] Step 100: Collect drug data and target protein data, and pre-process them;
[0009] Step 200: constructing an MDNN-DTA prediction model;
[0010] Step 300: Generate a data set based on the collected data, and train the MDNN-DTA prediction model based on the data set to obtain a trained MDNN-DTA prediction model;
[0011] Step 400: Input the data to be predicted into the trained MDNN-DTA prediction model to obtain the prediction result.
[0012] Preferably, in step 100, drug data and target protein data are collected and preprocessed, specifically:
[0013] Obtain drug molecules, where the drug molecules are in SMILES format, extract nodes, edge features, and adjacency matrices from the SMILES of the drug molecules based on TorchDrug and RDKit tools, and construct each drug molecule sequence into a graph structure, where nodes represent atoms and edges represent chemical bonds between atoms;
[0014] For the target protein data, that is, the target protein sequence, an integer dictionary of the FASTA sequence is constructed, each character is mapped to an integer, and the target protein sequence is represented as an integer sequence.
[0015] Preferably, the maximum length of the target protein sequence is fixed at 1200.
[0016] Preferably, in step 200, an MDNN-DTA prediction model is constructed, specifically:
[0017] The MDNN-DTA prediction model includes a drug feature extraction module, a protein overall feature extraction module, a protein feature fusion module, and an affinity prediction module. The protein overall feature extraction module is connected to the protein feature fusion module, and the drug feature extraction module and the protein feature fusion module are connected to the affinity prediction module.
[0018] Preferably, the drug feature extraction module is a three-layer GCN stack, which is used to extract the sequence features of drug molecules and incorporate the edge features of drug molecules into the graph message passing process of GCN to enrich the feature representation of nodes and edges in the drug graph. The formula is:
[0019]
[0020] In the formula, H (l) represents the feature output of the node at layer l, E represents the edge feature of the drug graph, is a learnable weight parameter.
[0021] Preferably, the protein overall feature extraction module includes a protein feature extraction module and an ESM module, the protein feature extraction module includes a global feature extractor and a local feature extractor, the global feature extractor is composed of an affine block, an LR layer and a residual connection, and the global feature extractor is used to extract the global features of the entire sequence through MLP without disrupting the order of the protein amino acid sequence and transmit it to subsequent operations, and its formula is:
[0022] X out =X in+[FC1(AF(X in ))] (2) (2)
[0023] AF(X)=Diag(α)X+β (3)
[0024] Where AF represents an Affine block for linear transformation of input features, Diag creates a diagonal matrix, α and β are trainable weight vectors, FC1(·) includes a fully connected layer and a ReLU layer, and [·]2 means looping twice;
[0025] The local feature extractor integrates the CNN method to obtain the local pattern features of proteins through a convolution receptive field of appropriate size, and its formula is:
[0026] X CNN =X out +Att se (CNN (3) (X out )) (4)
[0027] Where, X CNN is the output of the local feature extractor. The CNN contains three convolutional blocks, each of which contains a 1D convolutional layer, a batch normalization layer, and a ReLU layer. out is the output of the global feature extractor, Att se Represents a SE-Block that dynamically learns the importance of each channel through two main steps of squeezing and excitation;
[0028] The ESM module is an ESM-1V pre-trained model, and the features extracted from it are used as supplementary features of the target protein sequence.
[0029] Preferably, the protein feature fusion module is used to fuse the global features and local features extracted by the protein feature extraction module and the supplementary features extracted by the ESM module to form a more complete protein feature, and its formula is:
[0030] X f =X CNN ×W att +X e ×(1-W att ) (5)
[0031] W att =Att(FC2(X CNN )+FC2(X E )) (6)
[0032] Where, X fis the output of the protein feature fusion module, i.e., the final feature vector of the protein sequence, X CNN is the output of the protein feature extraction module, X E It is the supplementary feature extracted by the ESM module, and Att() is the self-attention mechanism, which is used to calculate the attention weight.
[0033] Preferably, the affinity prediction module splices the fused protein features and drug molecule features, inputs the spliced features into multiple fully connected layers, and obtains affinity scores.
[0034] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0035] The present invention provides a DTA prediction method based on a multimodal neural network model, the method comprising: collecting drug data and target protein data, preprocessing them, constructing an MDNN-DTA prediction model, generating a data set based on the collected data, training the MDNN-DTA prediction model based on the data set, obtaining a trained MDNN-DTA prediction model, inputting the data to be predicted into the trained MDNN-DTA prediction model, and obtaining a prediction result. The present invention has the following advantages:
[0036] 1. According to the different biochemical characteristics of drug molecules and target proteins, a MDNN-DTA prediction model based on multimodal neural network is designed and proposed. The model uses GCN and CNN methods to extract the characteristic representations of drug and protein sequences respectively, so as to achieve efficient prediction of DTA;
[0037] 2. A protein feature extraction (PFE) block is proposed, which aims to follow the linear sequence law of the target protein to deeply explore its global and local features. At the same time, a pre-training model is introduced to supplement the attribute information of the sequence at the molecular level, thereby generating a more detailed, robust and multi-dimensional protein feature representation;
[0038] 3. By introducing the protein feature fusion (PFF) block based on the attention mechanism, the model can efficiently integrate the multi-scale features of protein sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0040] Figure 1 A flow chart of a method provided by an embodiment of the present invention;
[0041] Figure 2 This is a schematic diagram of the MDNN-DTA prediction model structure;
[0042] Figure 3 It is a schematic diagram of the structure of the protein feature extraction module (PFE);
[0043] Figure 4 Schematic diagram of the protein feature fusion module (PFF) structure. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] The purpose of the present invention is to provide a DTA prediction method based on a multimodal neural network model. A feature extraction module is constructed using GCN and CNN to capture the sequence features of drug molecules and target proteins, respectively. The multi-scale features of protein sequences are integrated through an attention-based feature fusion block (PFF) to further improve the accuracy of the model.
[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0047] Figure 1 A flow chart of a method provided by an embodiment of the present invention, such as Figure 1 As shown, the present invention provides a DTA prediction method based on a multimodal neural network model, comprising:
[0048] Step 100: Collect drug data and target protein data, and pre-process them;
[0049] Step 200: constructing an MDNN-DTA prediction model;
[0050] Step 300: Generate a data set based on the collected data, and train the MDNN-DTA prediction model based on the data set to obtain a trained MDNN-DTA prediction model;
[0051] Step 400: Input the data to be predicted into the trained MDNN-DTA prediction model to obtain the prediction result.
[0052] In step 100, drug data and target protein data are collected and preprocessed, specifically:
[0053] Obtain drug molecules, where drug molecules use SMILES format, which is a specification that uses ASCII strings to concisely describe molecular structures. Based on TorchDrug and RDKit tools, nodes, edge features, and adjacency matrices are extracted from the SMILES of drug molecules, and each drug molecule sequence is constructed into a graph structure, where nodes represent atoms and edges represent chemical bonds between atoms.
[0054] For the target protein data, that is, the target protein sequence, an integer dictionary of the FASTA sequence is constructed to map each character to an integer. For example, the MLK3 protein sequence "VQIARGM" can be encoded as [22, 17, 9, 1, 18, 7, 13] according to the protein dictionary {'V': 22, 'Q': 17, 'I': 9, 'A': 1, 'R': 18, 'G': 7, 'M': 13}. This encoding method can represent the protein sequence as an integer sequence. For the convenience of training, the maximum length of the protein sequence is fixed to 1200, so that the maximum length can cover at least 80% of the proteins.
[0055] In step 200, the MDNN-DTA prediction model is constructed, specifically:
[0056] like Figure 2 As shown, the MDNN-DTA prediction model includes a drug feature extraction module, a protein overall feature extraction module, a protein feature fusion module, and an affinity prediction module. The protein overall feature extraction module is connected to the protein feature fusion module, and the drug feature extraction module and the protein feature fusion module are connected to the affinity prediction module.
[0057] The drug feature extraction module is a three-layer GCN stack, which is used to extract the sequence features of drug molecules and incorporate the edge features of drug molecules into the GCN graph message passing process to enrich the feature representation of nodes and edges in the drug graph. The formula is:
[0058]
[0059] In the formula, H (l) represents the feature output of the node at layer l, E represents the edge feature of the drug graph, is a learnable weight parameter.
[0060] The protein overall feature extraction module includes a protein feature extraction module (PFE) and an ESM module, such as Figure 3As shown in FIG. 1 , the protein feature extraction module includes a global feature extractor and a local feature extractor. The global feature extractor is composed of an affine block, an LR layer (a linear layer and a ReLU activation layer) and a residual connection. The global feature extractor is used to extract the global features of the entire sequence through MLP without disrupting the order of the protein amino acid sequence and transmit it to subsequent operations. The formula is:
[0061] X out =X in +[FC1(AF(X in ))] (2) (2)
[0062] AF(X)=Diag(α)X+β (3)
[0063] Where AF represents an Affine block for linear transformation of input features, Diag creates a diagonal matrix, α and β are trainable weight vectors, FC1(·) includes a fully connected layer and a ReLU layer, and [·]2 means looping twice;
[0064] The local feature extractor combines the advantages of the CNN method and obtains the local pattern features of proteins through a convolution receptive field of appropriate size. The application of an SE attention mechanism enables the extractor to focus on the importance of each channel, thereby generating a feature vector that is more in line with the natural amino acid sequence. The formula is:
[0065] X CNN =X out +Att se (CNN (3) (X out )) (4)
[0066] Where, X CNN is the output of the local feature extractor. The CNN contains three convolutional blocks, each of which contains a 1D convolutional layer, a batch normalization layer, and a ReLU layer. out is the output of the global feature extractor, Att se Represents a SE-Block that dynamically learns the importance of each channel through two main steps of squeezing and excitation;
[0067] The ESM module is an ESM-1V pre-trained model, and the features extracted by it are used as supplementary features of the target protein sequence. For protein sequences longer than 1024 amino acids, random numbers are used to reduce the length of the protein sequence to 1024 tokens to obtain a sample sequence.
[0068] Based on the above steps, two parts of protein sequence features are obtained, namely, the sequence features extracted by the PFE block, including the global and local features of the protein, and the biological characteristics supplemented by the ESM pre-trained model. In order to integrate these two parts of features together to form a more complete protein feature and enhance the representation ability of the model, a protein feature fusion module (PFF) is constructed, such as Figure 4 As shown in the figure, the attention weight is obtained by the self-attention formula. The two parts of the features are linearly processed, multiplied by the attention weights and added to obtain the final fused target protein feature vector. In short, the protein feature fusion module (PFF) is used to fuse the global features and local features extracted by the protein feature extraction module and the supplementary features extracted by the ESM module to form a more complete protein feature. The formula is:
[0069] X f =X CNN ×W att +X e ×(1-W att ) (5)
[0070] W att =Att(FC2(X CNN )+FC2(X E )) (6)
[0071] Where, X f is the output of the protein feature fusion module, i.e., the final feature vector of the protein sequence, X CNN is the output of the protein feature extraction module, X E It is the supplementary feature extracted by the ESM module, and Att() is the self-attention mechanism, which is used to calculate the attention weight.
[0072] After the drug and target protein sequences have undergone their respective feature extraction operations, the features of the two need to be fused to predict affinity. The affinity prediction module splices the fused protein features and drug molecule features, and inputs the spliced features into multiple fully connected layers to obtain affinity scores. The present invention provides an embodiment for using three fully connected layers, with the number of neurons being 1024, 512 and 128 respectively, and each linear layer is connected by a batchnormalization layer and a ReLU activation function.
[0073] The present invention compares the model recorded in this application with other models on two public datasets, Davis and KIBA, and determines the effectiveness of each component through a series of ablation experiments. On the Davis dataset, our MSE and CI reach 0.209 and 0.901 respectively, and on the KIBA dataset, our MSE and CI reach 0.135 and 0.894 respectively. These experimental results highlight the advantages and superiority of the MDNN-DTA method.
[0074] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0075] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A DTA prediction method based on a multimodal neural network model, characterized in that: include: Step 100: Collect drug data and target protein data, and pre-process them; Step 200: constructing an MDNN-DTA prediction model; Step 300: Generate a data set based on the collected data, and train the MDNN-DTA prediction model based on the data set to obtain a trained MDNN-DTA prediction model; Step 400: Input the data to be predicted into the trained MDNN-DTA prediction model to obtain the prediction result.
2. The method according to claim 1, characterized in that In step 100, drug data and target protein data are collected and preprocessed, specifically: Obtain drug molecules, where the drug molecules are in SMILES format, extract nodes, edge features, and adjacency matrices from the SMILES of the drug molecules based on TorchDrug and RDKit tools, and construct each drug molecule sequence into a graph structure, where nodes represent atoms and edges represent chemical bonds between atoms; For the target protein data, that is, the target protein sequence, an integer dictionary of the FASTA sequence is constructed, each character is mapped to an integer, and the target protein sequence is represented as an integer sequence.
3. The method according to claim 2, characterized in that The maximum length of the target protein sequence is fixed at 1200.
4. The method according to claim 1, characterized in that: In step 200, the MDNN-DTA prediction model is constructed, specifically: The MDNN-DTA prediction model includes a drug feature extraction module, a protein overall feature extraction module, a protein feature fusion module, and an affinity prediction module. The protein overall feature extraction module is connected to the protein feature fusion module, and the drug feature extraction module and the protein feature fusion module are connected to the affinity prediction module.
5. The method according to claim 4, characterized in that The drug feature extraction module is a three-layer GCN stack, which is used to extract the sequence features of drug molecules and incorporate the edge features of drug molecules into the GCN graph message passing process to enrich the feature representation of nodes and edges in the drug graph. The formula is: In the formula, H (l) represents the feature output of the node at layer l, E represents the edge feature of the drug graph, is a learnable weight parameter.
6. The method according to claim 5, characterized in that The protein overall feature extraction module includes a protein feature extraction module and an ESM module. The protein feature extraction module includes a global feature extractor and a local feature extractor. The global feature extractor is composed of an affine block, an LR layer and a residual connection. The global feature extractor is used to extract the global features of the entire sequence through MLP without disrupting the order of the protein amino acid sequence and transmit it to subsequent operations. The formula is: X out =X in +[FC1(AF(X in ))] (2) (2) AF(X)=Diag(α)X+β (3) Where AF represents an Affine block for linear transformation of input features, Diag creates a diagonal matrix, α and β are trainable weight vectors, FC1(·) includes a fully connected layer and a ReLU layer, and [·]2 means looping twice; The local feature extractor integrates the CNN method to obtain the local pattern features of proteins through a convolution receptive field of appropriate size, and its formula is: X CNN =X out +To se (CNN (3) (X out )) (4) Where, X CNN is the output of the local feature extractor. The CNN contains three convolutional blocks, each of which contains a 1D convolutional layer, a batch normalization layer, and a ReLU layer. out is the output of the global feature extractor, Att se Represents a SE-Block that dynamically learns the importance of each channel through two main steps of squeezing and excitation; The ESM module is an ESM-1V pre-trained model, and the features extracted from it are used as supplementary features of the target protein sequence.
7. The method according to claim 6, characterized in that The protein feature fusion module is used to fuse the global features and local features extracted by the protein feature extraction module and the supplementary features extracted by the ESM module to form a more complete protein feature. The formula is: X f =X CNN ×W att +X e ×(1-W att ) (5) W att =At(FC2(X CNN )+FC2(X E )) (6) Where, X f is the output of the protein feature fusion module, i.e., the final feature vector of the protein sequence, X CNN is the output of the protein feature extraction module, X E It is the supplementary feature extracted by the ESM module, and Att() is the self-attention mechanism, which is used to calculate the attention weight.
8. The method according to claim 7, characterized in that The affinity prediction module splices the fused protein features and drug molecule features, inputs the spliced features into multiple fully connected layers, and obtains affinity scores.
Citation Information
Patent Citations
Pretraining model-based drug target affinity prediction method
CN116994644A
Method and system for predicting affinity between drug and target
US20220284990A1