Multi-modal data fusion drug-target affinity prediction system based on Graph Transform

By employing a multimodal data fusion method combining Graph Transformer and 1D-CNN, the problem of insufficient accuracy and generalization ability in drug target affinity prediction was solved, achieving more accurate prediction of drug-target interaction strength and improving the efficiency and effectiveness of new drug discovery.

CN120853664APending Publication Date: 2025-10-28TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510800044.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing methods for predicting drug target affinity are insufficient in terms of accuracy and generalization ability, especially in regression tasks where it is difficult to accurately quantify the interaction strength between molecules and targets, which affects the drug discovery and optimization process.

Method used

A multimodal data fusion method combining Graph Transformer and one-dimensional convolutional neural network (1D-CNN) is adopted. High affinity samples are generated through GraphSMOTE, graph structure features of drugs and proteins are extracted using Graph Transformer, and text features are fused through cross-attention mechanism. Finally, affinity prediction is performed through multi-layer fully connected neural network.

Benefits of technology

It significantly improves the accuracy and generalization ability of drug-target affinity prediction, simplifies the feature extraction process, and improves the efficiency of model training and inference, providing a reliable technical means for high-throughput virtual screening and new drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853664A_ABST
    Figure CN120853664A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion drug-target affinity prediction system based on Graph Transform. The system comprises a data preprocessing module, a graph representation module, a text representation module and an affinity prediction module. The data preprocessing module is responsible for analyzing and processing the SMILES character string of the compound and the protein ID, and extracting the structural information of the compound and the three-dimensional structural data of the protein. And the affinity prediction module performs multi-modal fusion on the graph features and the text features, processes fusion feature vectors through a feedforward neural network comprising three full-connection layers, and outputs a prediction result of drug-target affinity. According to the method, multi-modal information of a graph structure and a sequence text is combined, the accuracy of cross-domain drug-target affinity prediction can be effectively improved, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0002] This invention belongs to the field of AI drug design. It uses 1D-CNN to extract textual features of drugs and proteins, and graph structural features are extracted by Graph Transformer to fully explore the rich information of compounds and proteins. To more accurately characterize protein information, this patent constructs protein graphs and their features based on the three-dimensional structure of proteins to more accurately capture intermolecular interactions, thereby achieving effective prediction of binding affinity. Background Technology

[0004] Drug-target affinity (DTA) is a key indicator in the drug discovery stage. It refers to the binding strength between a small molecule drug and its target protein. Affinity reflects the stability of the binding between the drug molecule and the target, and is usually expressed using the binding free energy (GFE). ), dissociation constant ( ), suppression constant ( ) and half-inhibitory concentration ( The higher the affinity, the tighter the drug binds to the target, which means better efficacy. Affinity is influenced by various intermolecular interactions, including hydrogen bonds, hydrophobic interactions, van der Waals forces, and ionic bonds. Molecular rigidity, binding site matching, and the surrounding molecular environment also affect binding stability. Previous studies often treated affinity prediction as a binary classification problem, using a threshold to determine whether a molecule-target pair possessed high affinity, commonly used in large-scale screening and hit rate prediction. Current research, however, directly predicts the specific numerical value of affinity, treating the prediction as a regression task, thus more accurately quantifying the interaction strength between the molecule and the target, often used for tasks such as optimizing molecular design and precise drug screening. Binary classification methods are more suitable for the initial screening stage, while regression methods provide more accurate affinity predictions, aiding in subsequent drug optimization. Summary of the Invention

[0006] This invention discloses a multimodal data fusion drug-target affinity prediction system based on Graph Transformer. Utilizing textual and image information of drugs and proteins, it extracts drug and protein features using Graph Transformer and attention mechanisms, and achieves accurate affinity prediction through a prediction module. Its excellent predictive performance has significant practical implications for candidate drug discovery and downstream pharmaceutical tasks.

[0007] To achieve the above objectives, the following technical solution is adopted:

[0008] This invention discloses a multimodal data fusion drug-target affinity prediction system based on Graph Transformer, including a data preprocessing module, a graph representation module, a text representation module, a feature fusion module, and an affinity prediction module.

[0009] Data preprocessing module: GraphSMOTE is used to generate new high-affinity samples on the graph structure of drug-target pairs, enabling the model to better learn minority class features.

[0010] The core idea of ​​GraphSMOTE is to interpolate in the expressive embedding space obtained by the GNN-based feature extractor to generate minority class nodes. At the same time, an edge generator is introduced to predict the connection relationship between the synthesized node and other nodes, thereby constructing an enhanced and balanced graph structure. This oversampled graph data augmentation method solves the class imbalance problem. GraphSMOTE consists of four parts: (1) a GNN-based feature extractor, which is used to learn node representations and fully preserve node attributes and graph topology information to assist in the generation of synthesized nodes; (2) a synthesized node generator, which generates synthesized minority class nodes in the latent space to alleviate the class imbalance problem; (3) an edge generator, which predicts and generates the connection relationship of the synthesized nodes to ensure that the nodes are reasonably integrated into the original graph structure, thereby constructing an enhanced graph with a balanced class distribution; and (4) a GNN-based classifier, which uses the enhanced graph to perform node classification and improves the ability to identify minority classes.

[0011] In terms of feature extraction, theoretically any type of GNN can be used. In this invention, GraphSAGE was chosen as the backbone model structure of GraphSMOTE, primarily because it can efficiently learn node representations from various topologies and possesses good generalization ability, adapting to new graph structures. In this part, the message passing and feature fusion process can be represented as:

[0012]

[0013] The attribute feature matrix of the input node. Representing node attributes , It is the th in the adjacency matrix List, These are the parameters of the weight matrix. It is the ReLU activation function. The calculated node Embedded features.

[0014] After obtaining the representation of each node in the embedding space constructed by the feature extractor, oversampling can be performed on this basis. The SMOTE algorithm is used here to generate synthetic samples through random interpolation, enhancing the effect of traditional oversampling. The basic idea of ​​SMOTE is to select target minority class samples in the embedding space and interpolate between their nearest neighbors to generate synthetic minority class nodes. The SMOTE algorithm calculation formula is as follows:

[0015]

[0016] In the formula, it is assumed that It is marked as For minority class nodes, the first step is to find those that are related to... The nearest labeled node of the same category, for The nearest neighbors within the same class are calculated in the embedding space using Euclidean distance. The formula for generating composite nodes using nearest neighbors is as follows:

[0017]

[0018] in, It is a random variable that follows a uniform distribution in the range [0, 1]. Because and Since they belong to the same category and are very similar, the resulting composite node... They also belong to the same category. Using this method, labeled synthetic nodes can be obtained. For each minority class, hyperparameters and oversampling scales are used to control the number of samples generated for each class, thus making the class size distribution more balanced.

[0019] The above operations have now generated synthetic nodes to balance the class distribution. However, these generated nodes differ from the original graph. They are isolated; there are no edges connecting them yet. This necessitates an edge generator to compute the existence of edges between nodes. GNNs need to learn how to simultaneously extract and propagate features, and this edge generator can provide relational information for the synthesized samples, thus facilitating the training of the GNN classifier. The edge generator is trained on real nodes and existing edges and used to predict the neighbor information of these synthesized nodes. The specific implementation formula is as follows:

[0020]

[0021] in, For the predicted nodes and Relationship information This is a parameter matrix used to capture the interactions between nodes. The loss function used to train the edge generator is:

[0022]

[0023] In the formula for Predicted connections between nodes The initial adjacency matrix consists of nodes and edges. This loss function prompts the edge generator to efficiently reconstruct the adjacency matrix using node representations, making it more naturally integrated into the original graph structure. Then, edge reconstruction is used, placing the predicted edges of the synthesized nodes into the augmented adjacency matrix, with a threshold set... Generate composite nodes The edge:

[0024]

[0025] in, By to The oversampled adjacency matrix obtained by inserting new nodes and edges is finally passed to the classifier.

[0026] The GNN classifier is calculated as shown in the formula. Through connection The augmented node representation is set up by embedding real nodes and synthesized nodes. These are the parameters of the weight matrix. It is the ReLU activation function. It is the updated node. The embedding features are then used to calculate the node distribution of node class labels using a formula.

[0027]

[0028]

[0029] in, This represents the node representation matrix of the second GNN block. Represents the weight parameters. It is a node The probability distribution on the class labels is determined, and finally, the classifier module is optimized using the following cross-entropy loss function:

[0030]

[0031] in By merging synthetic nodes to Enhanced tag set, It is a node The true category label, and It is an indicator function, when the node The true category is The value is 1 if the condition is met, and 0 otherwise.

[0032] The graph representation module is used for extracting graph structure features from drug compounds and target proteins, fully mining molecular-level topological information and establishing high-dimensional feature representations. First, drug compounds and target proteins are respectively processed through feature embedding layers, converting molecular structures and protein sequences into high-dimensional vector representations. Then, a Graph Transformer module is introduced, utilizing its multi-head self-attention mechanism to construct long-range dependencies within the molecule, thereby capturing global topological features. Finally, a max-pooling layer is used to reduce the dimensionality of the features, unifying the feature dimensions for scale alignment with text features while preserving key representative features.

[0033] Specifically, given Layer node features , To calculate the number of nodes. To the node The formula for multi-head attention is as follows:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] in the formula For the number of attention heads, the query vector Based on nodes Feature matrix The key vector is obtained by linear transformation. Sum value vector Then use nodes Feature matrix The linear transformation yields the following: , , , , , For different trainable parameters. The attention score is calculated using a scale dot product scaling function. Let be the dot product of two vectors. The dimension of the attention vector is the same as the dimension of the key vector. Then, the attention score is converted into attention weights using the Softmax function. .

[0040] After calculating the multi-head attention weights, for the remote node To the source node Perform information aggregation (including its own node) to obtain the source node. In the The feature information in the formula is as follows:

[0041]

[0042] in, for The splicing operation of individual attention heads. This represents the number of remote nodes. Compared to traditional graph neural network formulas, a multi-head attention matrix replaces the normalized adjacency matrix as the message passing transition matrix. Each layer aggregates neighbor features, gradually expanding the receptive field of nodes, transitioning from local relationships to global ones.

[0043] After calculating the features using Graph Transformer, LayerNorm and ReLU activation functions are introduced to generate higher-quality feature representations. Then, Maxpool is used to scale the feature dimensions to ensure that the feature dimensions of compounds and proteins are consistent, which facilitates subsequent feature concatenation. The specific formula is as follows:

[0044]

[0045]

[0046] Where parameters , , , The learning process can be adjusted. After the above steps, graph structure features are obtained for further affinity prediction.

[0047] The Graph Transformer, as the core module, is used to extract deep structural features from drug and target maps. Following a GNN-based message passing mechanism, the Graph Transformer layer utilizes an attention mechanism to aggregate the feature representations of neighboring atoms, thereby dynamically updating the atom embeddings. By stacking multiple Graph Transformer layers, the model can effectively capture a wider range of domain structural information, enabling atoms to obtain richer representations in a broader local environment, thus enhancing the expressive power of molecular features.

[0048] In the Layer and first In each attention head, for atoms Input feature representation The features are updated through calculation. For example, the following formula:

[0049]

[0050] in, The transformation matrix is... For nodes In the The attention head receives the nodes Attention score. By... Attention score for each channel The attention score is obtained by adding them together. Then, the Softmax activation function is used for calculation:

[0051]

[0052] In the formula Represents the dot product between two vectors, where 1 is... A 1-dimensional vector, where each element has a value of 1. Defined as:

[0053]

[0054] In the formula Represents the element-wise product of two vectors. Let be the transformation matrix.

[0055] atom In the The layer's output representation is obtained by updating the different attention heads. The representations are concatenated and then subjected to a linear transformation:

[0056]

[0057] in, This represents the vector concatenation operation. To focus on the number of heads, Let be the transformation matrix.

[0058] After deep feature extraction through multiple layers of Graph Transformer, the structural information of the drug and target is fully captured. To ensure scale consistency across different modalities, the model employs max pooling to align the extracted molecular features with the feature dimensions of the text modality output. This process not only standardizes the feature scale but also effectively reduces interference from redundant features, laying a spatial foundation for subsequent multimodal feature fusion.

[0059] The text representation module is used to process both the SMILES molecular representation and the protein amino acid sequence to capture their sequence features and construct high-dimensional feature representations suitable for drug-target affinity prediction. For the SMILES branch, the character-level SMILES representation is first mapped to a high-dimensional vector space through a feature embedding layer to preserve the structural information of the molecular sequence. Then, a one-dimensional convolutional neural network (1D-CNN) is used to extract local patterns to identify key substructures in the SMILES sequence, and max pooling is used to reduce the data dimensionality to retain the most significant sequence features. Finally, a linear transformation is used for nonlinear mapping to further enhance feature representation capabilities, ultimately generating a text representation for downstream prediction tasks. Similarly, the protein sequence branch's representation module also extracts features from the amino acid sequence based on this approach.

[0060] For text feature embedding, 64 integer labels are used to represent each symbol in SMILES, while 25 categories are used for protein sequences. To ensure that the convolution operation can effectively slide across the entire text and extract features, the maximum length of SMILES feature embeddings is uniformly set to 85, and the maximum length of protein sequences is set to 1200. Sequences exceeding the maximum length are truncated; sequences shorter than the maximum length are padded with zeros. This strategy aims to ensure consistency in model evaluation.

[0061] In the text representation module, a multi-layer one-dimensional convolutional neural network is used to perform convolution operations on text features. The basic principle of this method is to extract key information from the sequence by sliding the convolution kernel (filter) with a specific stride. The calculation formula is shown below:

[0062]

[0063] in, yes The first sequence in the layer Output a character vector. yes The first sequence in the layer A character vector representation, It is The first convolution kernel in the layer Each weight, For the first Layer bias terms, This represents the number of characters covered by the convolution kernel. After the convolution calculation is completed, the feature dimensions are scaled using max pooling, and finally, a linear layer is used to align the scale of the extracted features with the scale of the graph features.

[0064] Feature Fusion Module: Feature fusion strategies based on cross-attention mechanisms have gradually become a hot topic in the field of deep learning. Cross-attention mechanisms allow the model to assign different weights to each part of the input data, enabling the model to extract more relevant and crucial information. In affinity prediction, cross-attention modules are often used to fuse multi-scale features of the graph and the text information. Introducing an attention mechanism allows for the measurement of the importance of each residue and compound molecule during the binding process using attention scores, resulting in a comprehensive and reasonable feature embedding. The module integrates two labeled dot-product attention modules, collectively called scaled dot-product attention layers. Multi-head scaled dot-product attention can also be achieved by stacking multiple scaled dot-product attention layers in parallel. The scaled dot-product attention layer converts the query and a set of key-value pairs into a weighted sum of values, with the weight of each value calculated by a compatibility function between the query and its corresponding key. This module will convert the query ( ), key ( ), value ( The feature matrix is ​​used as input, and the feature matrix is ​​processed... The input matrix is ​​obtained by performing a linear mapping. ), input matrix (dimension is) ), input matrix (dimension is) ):

[0065]

[0066] in Assuming This represents a piece of text. This indicates the number of words appearing in the text. The feature embedding dimension of a word. , , These are their respective weight matrices. After calculating the three values, the dot product of the query key and all keys is calculated, and each result is divided by... Then, the Softmax activation function is used to obtain the weight values:

[0067]

[0068] In the affinity prediction task, the cross-attention module helps the model learn the relationship between independent one-dimensional text sequences and three-dimensional graph structures through the interaction of two encoded embedding modalities. This allows the model to obtain more comprehensive and effective information from relevant protein modalities, thus promoting affinity prediction.

[0069] Affinity prediction module: Predicts the affinity between molecular and protein features, fuses the obtained features, and performs prediction using a multi-layer fully connected neural network. The multi-layer fully connected neural network consists of three fully connected layers with ReLU activation functions. A Dropout layer is added after the fully connected layers to reduce the risk of overfitting and improve the model's generalization ability. This patent treats affinity prediction as a regression task, therefore using mean squared error (MSE) as the model loss function.

[0070]

[0071] in, The number of drug-target bindings. The affinity value predicted by the model. This represents the true affinity value. The model is trained using the Adam optimizer as the optimization function.

[0072] The beneficial effects of this invention are:

[0073] This invention fuses the graph structure features of compounds and proteins with text sequence features using Graph Transformer technology and a one-dimensional convolutional neural network. This approach not only deeply mines spatial interaction information in three-dimensional structures but also captures local features at the sequence level, thereby significantly improving the accuracy and generalization ability of drug-target affinity prediction. At the same time, the modular end-to-end learning framework simplifies the feature extraction process and improves the efficiency of model training and inference, providing a reliable and scalable technical means for high-throughput virtual screening and new drug discovery. Attached Figure Description

[0075] Figure 1 This describes the overall network architecture of a multimodal data fusion drug-target affinity prediction system based on Graph Transformer.

[0076] Figure 2 This is the GraphSMOTE data preprocessing module.

[0077] Figure 3 This refers to the network architecture of the Graph Transformer layer.

[0078] Figure 4 This refers to the multi-head attention mechanism in Graph Transformer.

[0079] Figure 5 This is a network architecture for cross-attention.

[0080] Figure 6 The network architecture for the prediction module.

[0081] Figure 7 Multi-head attention matrix for drug molecules

[0082] Figure 8 Distribution of drug molecules of interest

[0083] Figure 9 Image showing molecular docking results

[0084] Figure 10 The top 20 high-affinity drugs screened based on the GPR55 protein Detailed Implementation

[0086] The present invention will be described in detail below with reference to specific embodiments.

[0087] A multimodal data fusion drug-target affinity prediction system based on Graph Transformer, the overall network architecture of which is as follows: Figure 1 Shown, including:

[0088] A1. Data Preprocessing Module: GraphSMOTE is used to generate new high-affinity samples on the graph structure of drug-target pairs, enabling the model to better learn the features of the minority class.

[0089] The GraphSMOTE network structure involved in step A1 is as follows: Figure 2 As shown. The core idea is to interpolate in the expressive embedding space obtained by the GNN-based feature extractor to generate minority class nodes. At the same time, an edge generator is introduced to predict the connection relationship between the synthetic node and other nodes, thereby constructing an enhanced and balanced graph structure. This oversampled graph data augmentation method solves the class imbalance problem. GraphSMOTE consists of four parts: (1) a GNN-based feature extractor, which is used to learn node representations and fully preserve node attributes and graph topology information to assist in the generation of synthetic nodes; (2) a synthetic node generator, which generates synthetic minority class nodes in the latent space to alleviate the class imbalance problem; (3) an edge generator, which predicts and generates the connection relationship of synthetic nodes to ensure that the nodes are reasonably integrated into the original graph structure, thereby constructing an enhanced graph with a balanced class distribution; (4) a GNN-based classifier, which uses the enhanced graph to perform node classification and improves the ability to identify minority classes.

[0090] The specific steps of step A1 are as follows:

[0091] The input graph data is fed into the feature extractor. Theoretically, any type of GNN can be used as the feature extractor. In this invention, GraphSAGE is chosen as the backbone model structure of GraphSMOTE, primarily because it can efficiently learn node representations from various topologies and possesses good generalization ability, adapting to new graph structures. In this part, the message passing and feature fusion process can be represented as:

[0092]

[0093] The attribute feature matrix of the input node. Representing node attributes , It is the th in the adjacency matrix List, These are the parameters of the weight matrix. It is the ReLU activation function. The calculated node Embedded features.

[0094] After obtaining the representation of each node in the embedding space constructed by the feature extractor, oversampling can be performed on this basis. The SMOTE algorithm is used here to generate synthetic samples through random interpolation, enhancing the effect of traditional oversampling. The basic idea of ​​SMOTE is to select target minority class samples in the embedding space and interpolate between their nearest neighbors to generate synthetic minority class nodes. The SMOTE algorithm calculation formula is as follows:

[0095]

[0096] In the formula, it is assumed that It is marked as For minority class nodes, the first step is to find those that are related to... The nearest labeled node of the same category, for The nearest neighbors within the same class are calculated in the embedding space using Euclidean distance. The formula for generating composite nodes using nearest neighbors is as follows:

[0097]

[0098] in, It is a random variable that follows a uniform distribution in the range [0, 1]. Because and Since they belong to the same category and are very similar, the resulting composite node... They also belong to the same category. Using this method, labeled synthetic nodes can be obtained. For each minority class, hyperparameters and oversampling scales are used to control the number of samples generated for each class, thus making the class size distribution more balanced.

[0099] The above operations have now generated synthetic nodes to balance the class distribution. However, these generated nodes differ from the original graph. They are isolated; there are no edges connecting them yet. This necessitates an edge generator to compute the existence of edges between nodes. GNNs need to learn how to simultaneously extract and propagate features, and this edge generator can provide relational information for the synthesized samples, thus facilitating the training of the GNN classifier. The edge generator is trained on real nodes and existing edges and used to predict the neighbor information of these synthesized nodes. The specific implementation formula is as follows:

[0100]

[0101] in, For the predicted nodes and Relationship information This is a parameter matrix used to capture the interactions between nodes. The loss function used to train the edge generator is:

[0102]

[0103] In the formula for Predicted connections between nodes The initial adjacency matrix consists of nodes and edges. This loss function prompts the edge generator to efficiently reconstruct the adjacency matrix using node representations, making it more naturally integrated into the original graph structure. Then, edge reconstruction is used, placing the predicted edges of the synthesized nodes into the augmented adjacency matrix, with a threshold set... Generate composite nodes The edge:

[0104]

[0105] in, By to The oversampled adjacency matrix obtained by inserting new nodes and edges is finally passed to the classifier.

[0106] The GNN classifier is calculated as shown in the formula:

[0107]

[0108]

[0109] in, Through connection The augmented node representation is set up by embedding real nodes and synthesized nodes. These are the parameters of the weight matrix. It is the ReLU activation function. It is the updated node. The embedding features are then used to calculate the node distribution of node class labels using a formula. This represents the node representation matrix of the second GNN block. Represents the weight parameters. It is a node The probability distribution on the class labels is determined, and finally, the classifier module is optimized using the following cross-entropy loss function:

[0110]

[0111] in By merging synthetic nodes to Enhanced tag set, It is a node The true category label, and It is an indicator function, when the node The true category is The value is 1 if the condition is met, and 0 otherwise.

[0112] A2. The graph characterization module is used to extract graph structural features of drug compounds and target proteins, fully explore molecular-level topological information, and establish high-dimensional feature representations;

[0113] The graph representation modules involved in step A2 are as follows: Figure 3 As shown, the Graph Transformer network structure is as follows: Figure 4 As shown, this method is used for graph structure feature extraction of drug compounds and target proteins, fully mining molecular-level topological information and establishing high-dimensional feature representations. The input atomic features are first processed by addition and normalization after a residual connection, then fed into a feedforward network to extract local features. Subsequently, they are added and normalized again, and then fed into a multi-head self-attention module to capture global dependencies between atoms in the graph. The final output is atomic-level embedded features for subsequent affinity prediction tasks.

[0114] The specific steps are as follows:

[0115] Given Layer node features , To calculate the number of nodes. To the node The formula for multi-head attention is as follows:

[0116]

[0117]

[0118]

[0119]

[0120]

[0121] in the formula For the number of attention heads, the query vector Based on nodes Feature matrix The key vector is obtained by linear transformation. Sum value vector Then use nodes Feature matrix The linear transformation yields the following: , , , , , For different trainable parameters. The attention score is calculated using a scale dot product scaling function. Let be the dot product of two vectors. The dimension of the attention vector is the same as the dimension of the key vector. Then, the attention score is converted into attention weights using the Softmax function. .

[0122] After calculating the multi-head attention weights, for the remote node To the source node Perform information aggregation (including its own node) to obtain the source node. In the The feature information in the formula is as follows:

[0123]

[0124] in, for The splicing operation of individual attention heads. This represents the number of remote nodes. Compared to traditional graph neural network formulas, a multi-head attention matrix replaces the normalized adjacency matrix as the message passing transition matrix. Each layer aggregates neighbor features, gradually expanding the receptive field of nodes, transitioning from local relationships to global ones.

[0125] After calculating the features using Graph Transformer, LayerNorm and ReLU activation functions are introduced to generate higher-quality feature representations. Then, Maxpool is used to scale the feature dimensions to ensure that the feature dimensions of compounds and proteins are consistent, which facilitates subsequent feature concatenation. The specific formula is as follows:

[0126]

[0127]

[0128] Where parameters , , , The learning process can be adjusted. After the above steps, graph structure features are obtained for further affinity prediction.

[0129] refer to Figure 4 The Graph Transformer, as the core module, is used to extract deep structural features from drug and target maps. The left side features a basic scaled dot-product attention mechanism. After linear transformations, matrix multiplications, and scaling of the query, key, and value, attention weights are obtained through Softmax, and then multiplied by the value vector to obtain the output. The right side features a multi-head attention mechanism, mapping the input to multiple attention heads. Scaled dot-product attention is calculated independently in each head, and the results from each head are concatenated and merged to model the complex relationships between the central node and its neighbors. Following a GNN-based message passing mechanism, the Graph Transformer layer uses the attention mechanism to aggregate the feature representations of adjacent atoms, thereby dynamically updating the atom embeddings. By stacking multiple Graph Transformer layers, the model can effectively capture a wider range of domain structural information, allowing atoms to obtain richer representations in a broader local environment, thus improving the expressive power of molecular features.

[0130] In the Layer and first In each attention head, for atoms Input feature representation The features are updated through calculation. For example, the following formula:

[0131]

[0132] in, The transformation matrix is... For nodes In the The attention head receives the nodes Attention score. By... Attention score for each channel The attention score is obtained by adding them together. Then, the Softmax activation function is used for calculation:

[0133]

[0134] In the formula Represents the dot product between two vectors, where 1 is... A 1-dimensional vector, where each element has a value of 1. Defined as:

[0135]

[0136] In the formula Represents the element-wise product of two vectors. Let be the transformation matrix.

[0137] atom In the The layer's output representation is obtained by updating the different attention heads. The representations are concatenated and then subjected to a linear transformation:

[0138]

[0139] in, This represents the vector concatenation operation. To focus on the number of heads, Let be the transformation matrix.

[0140] After deep feature extraction through multiple layers of Graph Transformer, the structural information of the drug and target is fully captured. To ensure scale consistency across different modalities, the model employs max pooling to align the extracted molecular features with the feature dimensions of the text modality output. This process not only standardizes the feature scale but also effectively reduces interference from redundant features, laying a spatial foundation for subsequent multimodal feature fusion.

[0141] A3. The text representation module is used to process SMILES molecular representations and protein amino acid sequences to capture their sequence features and construct high-dimensional feature representations suitable for drug-target affinity prediction.

[0142] The specific steps are as follows:

[0143] For the SMILES branch, the character-level SMILES representation is first mapped to a high-dimensional vector space through a feature embedding layer to preserve the structural information of the molecular sequence. Then, a one-dimensional convolutional neural network (1D-CNN) is used to extract local patterns to identify key substructures in the SMILES sequence, and max pooling is used to reduce the data dimensionality to retain the most significant sequence features. Finally, a linear transformation is used for non-linear mapping to further enhance feature representation capabilities, ultimately generating a text representation for downstream prediction tasks. Similarly, the representation module for the protein sequence branch also extracts features from the amino acid sequence based on this approach.

[0144] For text feature embedding, 64 integer labels are used to represent each symbol in SMILES, while 25 categories are used for protein sequences. To ensure that the convolution operation can effectively slide across the entire text and extract features, the maximum length of SMILES feature embeddings is uniformly set to 85, and the maximum length of protein sequences is set to 1200. Sequences exceeding the maximum length are truncated; sequences shorter than the maximum length are padded with zeros. This strategy aims to ensure consistency in model evaluation.

[0145] In the text representation module, a multi-layer one-dimensional convolutional neural network is used to perform convolution operations on text features. The basic principle of this method is to extract key information from the sequence by sliding the convolution kernel (filter) with a specific stride. The calculation formula is shown below:

[0146]

[0147] in, yes The first sequence in the layer Output a character vector. yes The first sequence in the layer A character vector representation, It is The first convolution kernel in the layer Each weight, For the first Layer bias terms, This represents the number of characters covered by the convolution kernel. After the convolution calculation is completed, the feature dimensions are scaled using max pooling, and finally, a linear layer is used to align the scale of the extracted features with the scale of the graph features.

[0148] A4. The feature fusion module plays a crucial bridging role in this system. It integrates multi-source heterogeneous information from drug molecule maps, SMILES sequences, protein structure maps, and amino acid sequences, deeply fusing graph structure with textual semantics to generate a unified feature representation. This provides comprehensive and accurate input for downstream affinity prediction and is key to model performance and generalization ability. Regarding the fusion method, the cross-attention mechanism allows the model to assign different weights to each part of the input data, enabling it to extract more relevant and critical information.

[0149] Cross-attention modules are often used to fuse multi-scale features of graphs and textual information. The network architecture is as follows: Figure 5 As shown. Introducing an attention mechanism allows for the measurement of the importance of each residue and compound molecule during the binding process using attention scores, resulting in a comprehensive and reasonable feature embedding. The module integrates two labeled dot-product attention modules, collectively called the scaled dot-product attention layer. Multi-head scaled dot-product attention can also be achieved by parallel stacking of multiple scaled dot-product attention layers. The scaled dot-product attention layer converts the query and a set of key-value pairs into a weighted sum of values, where the weight of each value is calculated by a compatibility function between the query and its corresponding key. This module will convert the query ( ), key ( ), value ( The feature matrix is ​​used as input, and the feature matrix is ​​processed... The input matrix is ​​obtained by performing a linear mapping. ), input matrix (dimension is) ), input matrix (dimension is) ):

[0150]

[0151] in Assuming This represents a piece of text. This indicates the number of words appearing in the text. The feature embedding dimension of a word. , , These are their respective weight matrices. After calculating the three values, the dot product of the query key and all keys is calculated, and each result is divided by... Then, the Softmax activation function is used to obtain the weight values:

[0152]

[0153] In the affinity prediction task, the cross-attention module helps the model learn the relationship between independent one-dimensional text sequences and three-dimensional graph structures through the interaction of two encoded embedding modalities. This allows the model to obtain more comprehensive and effective information from relevant protein modalities, thus promoting affinity prediction.

[0154] The A5 affinity prediction module predicts the affinity between molecular and protein features, fuses the obtained features together, and performs predictions through a multi-layer fully connected neural network.

[0155] Prediction module network architecture such as Figure 6 As shown, the multilayer fully connected neural network consists of three fully connected layers with ReLU activation functions. A Dropout layer is added after the fully connected layers to reduce the risk of overfitting and improve the model's generalization ability. This patent treats affinity prediction as a regression task, therefore, the mean squared error (MSE) is used as the model loss function.

[0156]

[0157] in, The number of drug-target bindings. The affinity value predicted by the model. This represents the true affinity value. The model is trained using the Adam optimizer as the optimization function.

[0158] The attention mechanism in GraphTransDTA can be used to analyze which molecular sites in drugs are more likely to bind to large protein molecules. We input a small molecule compound with PubChemCID 20744853 and a protein kinase with UniprotID Q2M2I8 into a trained model to generate an attention matrix, demonstrating the importance of drug molecules and proteins in drug-target binding and providing some interpretability for biochemistry. Figure 7 This diagram illustrates the multi-head attention matrix of a molecular compound, based on a 4-layer graph transformer. The calculation results of the four attention heads in the first two layers are extracted. The color differences in the diagram reflect the attention level of each node in the molecular and protein graphs. Figure 7 In the diagram, the x-axis and y-axis represent node numbers, respectively. It can be seen that the attention of the first layer (ad) in the small molecule is mainly distributed at nodes 0, 9, 14, 16, 18, 19, and 26, while the second layer (eh) is mainly at nodes 0, 8, 14, and 16. A summary of these observations shows... Figure 8 Pay attention to the distribution map.

[0159] To further illustrate this, we used AutoDock4 to perform molecular docking between compound 20744853 and Q2M2I8 protein kinase. Based on the docking energy, we selected the four docking results with the lowest energies and visualized them, as shown below. Figure 9 As shown. Lower energy indicates that the binding of the drug molecule to the pocket protein is spontaneous and relatively stable, requiring no additional energy stimulation. At a docking energy of -4.11 kcal / mol, node 26 of the compound is connected to the TYR-305 residue, and node 14 is connected to the CYS-319 residue. At a docking energy of -3.95 kcal / mol, node 16 is connected to the ASN-279 residue. At a docking energy of -3.84 kcal / mol, node 26 is connected to the LEU-142 residue. At a docking energy of -2.97 kcal / mol, node 8 is connected to ILE-111, and node 26 is connected to LYS-88. Therefore, it can be seen that... Figure 8 The nodes in the model are relatively close together, and we only selected the top four docking energies here. Some nodes may bind to residues at other energies. Therefore, the attention mechanism in the model can provide some site references for molecular docking, which is beneficial for the research and development of related drugs.

[0160] Figure 10 These are the top 20 high-affinity drugs screened by the Graph Transformer-based multimodal data fusion drug-target affinity prediction system of this invention. Actual affinity values ​​need to be verified through clinical trials; the system results are for reference only.

[0161] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A drug-target affinity prediction system based on Graph Transformer multimodal data fusion, characterized in that, It includes a data preprocessing module, a graph representation module, a text representation module, a feature fusion module, and an affinity prediction module. Data preprocessing module: preprocesses the input data, which includes: drug compound molecular graphs, SMILES text, target protein graphs and amino acid sequences. GraphSMOTE is used to generate new high-affinity samples on the graph structure of drug-target pairs, enabling the model to better learn minority class features. Graph characterization module: used to extract features from the graph structure of preprocessed drug compounds and target proteins, fully explore molecular-level topological information, and establish high-dimensional feature representations; Text Representation Module: Used to process preprocessed SMILES molecular representations and protein amino acid sequences to capture their sequence features and construct high-dimensional feature representations suitable for drug-target affinity prediction; Feature fusion module: This module fuses the outputs of the graph representation module and the text representation module using a cross-attention method. The cross-attention mechanism enables the model to assign different weights to each part of the input data, allowing the model to extract more relevant and critical information. Affinity prediction module: Predicts the affinity between molecular features and protein features, fuses the obtained features together, and performs prediction through a multi-layer fully connected neural network.

2. The system according to claim 1, characterized in that, The data preprocessing module includes four parts: (1) a GNN-based feature extractor for learning node representations, fully preserving node attributes and graph topology information to assist in the generation of synthetic nodes; (2) a synthetic node generator for generating synthetic minority class nodes in the latent space to alleviate class imbalance; (3) an edge generator for predicting and generating the link relationships of synthetic nodes to ensure that nodes are reasonably integrated into the original graph structure, thereby constructing an enhanced graph with a balanced class distribution; and (4) a GNN-based classifier for performing node classification using the enhanced graph to improve the ability to identify minority classes. For feature extraction, GraphSAGE was chosen as the backbone model structure of GraphSMOTE. It can efficiently learn node representations from various topologies and has good generalization ability, adapting to new graph structures. In this part, the message passing and feature fusion process is represented as follows: ; The attribute feature matrix of the input node. Representing node attributes , It is the th in the adjacency matrix List, These are the parameters of the weight matrix. It is the ReLU activation function. The calculated node Embedded features.

3. The system according to claim 2, characterized in that, After obtaining the representation of each node in the embedding space constructed by the feature extractor, oversampling is performed on this basis. The SMOTE algorithm is used to generate synthetic samples through random interpolation to enhance the effect of traditional oversampling. The basic idea of ​​SMOTE is to select target minority class samples in the embedding space and interpolate between their nearest neighbors to generate synthetic minority class nodes. The SMOTE algorithm calculation formula is as follows: In the formula, it is assumed that... It is marked as For minority class nodes, the first step is to find those that are related to... The nearest labeled node of the same category, for The nearest neighbors within the same class are calculated using Euclidean distance in the embedding space; the formula for generating composite nodes using nearest neighbors is as follows: ;in, It is a random variable that follows a uniform distribution in the range [0, 1]. Because and Since they belong to the same category and are very similar, the resulting composite node... They also belong to the same category; by calculating in this way, the labeled synthetic nodes can be obtained; for each minority class, hyperparameters and oversampling scales are used to control the number of samples generated for each class, so that the class size distribution can be more balanced.

4. The system according to claim 3, characterized in that, An edge generator is used to compute the existence of edges between nodes; the GNN needs to learn how to extract and propagate features simultaneously, and this edge generator can provide relational information for the synthesized samples, thereby facilitating the training of the GNN classifier; the edge generator is trained on real nodes and existing edges and used to predict the neighbor information of these synthesized nodes; the specific implementation formula is as follows: ;in, For the predicted nodes and Relationship information The parameter matrix is ​​used to capture the interactions between nodes; the loss function used to train the edge generator is: In the formula for Predicted connections between nodes The initial adjacency matrix consists of nodes and edges. This loss function prompts the edge generator to effectively reconstruct the adjacency matrix using node representations, making it more naturally integrated into the original graph structure. Then, edge reconstruction is used, placing the predicted edges of the synthesized nodes into the augmented adjacency matrix, with a threshold set. Generate composite nodes The edge: ;in, By to The oversampled adjacency matrix obtained by inserting new nodes and edges is finally passed to the GNN classifier.

5. The system according to claim 4, characterized in that, The GNN classifier is calculated as shown in the formula. Through connection The augmented node representation is set up by embedding real nodes and synthesized nodes. These are the parameters of the weight matrix. It is the ReLU activation function. It is the updated node. The embedding features are then used to calculate the node distribution of the node class label using a formula; ; ;in, This represents the node representation matrix of the second GNN block. Represents the weight parameters. It is a node The probability distribution on the class labels is determined, and finally, the classifier module is optimized using the following cross-entropy loss function: ;in By merging synthetic nodes to Enhanced tag set, It is a node The true category label, and It is an indicator function, when the node The true category is The value is 1 if the condition is met, and 0 otherwise.

6. The system according to claim 1, characterized in that, The graph characterization module is used to extract graph structure features of drug compounds and target proteins, fully explore molecular-level topological information, and establish high-dimensional feature representations. First, the drug compound and target protein are respectively processed through a feature embedding layer to convert the molecular structure and protein sequence into high-dimensional vector representations. Then, a Graph Transformer module is introduced, which uses the multi-head self-attention mechanism to construct long-range dependencies within the molecule, thereby capturing global topological features. Finally, the features are reduced in dimensionality through a max pooling layer to unify the feature dimension, making it easier to scale with text features while retaining key representative features.

7. The system according to claim 6, characterized in that, Given Layer node features , To calculate the number of nodes. To the node The formula for multi-head attention is as follows: ; ; ; ; ;In the formula For the number of attention heads, the query vector Based on nodes Feature matrix The key vector is obtained by linear transformation. Sum value vector Then use nodes Feature matrix The linear transformation yields the following: , , , , , For different trainable parameters; The attention score is calculated using a scale dot product scaling function. Let be the inner product of two vectors. The dimension of the attention vector is the same as the dimension of the key vector; then, the attention score is converted into attention weights using the Softmax function. After calculating the multi-head attention weights, for the remote node To the source node Perform information aggregation to obtain the source node. In the The feature information in the formula is as follows: ;in, for The splicing operation of individual attention heads. The number of remote nodes; compared with the traditional graph neural network formula, the multi-head attention matrix replaces the normalized adjacency matrix as the transition matrix for message passing; each layer gradually expands the information receptive field of the nodes through neighbor feature aggregation, transitioning from local relations to global relations; After the features are calculated by Graph Transformer, LayerNorm and ReLU activation functions are introduced to generate higher quality feature representations. Then, Maxpool is used to scale the feature dimensions to ensure that the feature dimensions of compounds and proteins are consistent, which facilitates subsequent feature splicing. The Graph Transformer, as the core module, is used to extract deep structural features from drug and target maps. Following the message passing mechanism based on GNN, the Graph Transformer layer uses an attention mechanism to aggregate the feature representations of adjacent atoms, thereby dynamically updating the atom embeddings. By stacking multiple Graph Transformer layers, the model can effectively capture a wider range of domain structural information, enabling atoms to obtain richer representations in a broader local environment, thereby improving the expressive power of molecular features. After deep feature extraction through multiple layers of Graph Transformer, the structural information of the drug and target is fully captured. To ensure the scale consistency of features across different modalities, the model employs max pooling to adjust the extracted molecular features to align with the feature dimensions of the text modality output. This process not only standardizes the feature scale but also effectively reduces the interference of redundant features, laying a spatial foundation for subsequent multimodal feature fusion.

8. The system according to claim 1, characterized in that, The text representation module is used to process SMILES molecular representations and protein amino acid sequences to capture their sequence features and construct high-dimensional feature representations suitable for drug-target affinity prediction. For the SMILES branch, the character-level SMILES representation is first mapped to a high-dimensional vector space through a feature embedding layer to preserve the structural information of the molecular sequence. Then, a one-dimensional convolutional neural network (1D-CNN) is used to extract local patterns to identify key substructures in the SMILES sequence, and max pooling is used to reduce the data dimensionality to preserve the most significant sequence features. Finally, a linear transformation is used for nonlinear mapping to further enhance the feature expression capability, ultimately generating a text representation for downstream prediction tasks; the representation module of the protein sequence side branch also extracts features from the amino acid sequence based on this. For text feature embedding, 64 categories of integer labels are used to represent each symbol in SMILES, while 25 categories are used for protein sequences. To ensure that the convolution operation can effectively slide across the entire text and extract features, the maximum length of SMILES feature embedding is uniformly set to 85, and the maximum length of protein sequences is set to 1200. For sequences exceeding the maximum length, truncation is performed; for sequences shorter than the maximum length, zero padding is used. This strategy aims to ensure the consistency of model evaluation. In the text representation module, a multi-layer one-dimensional convolutional neural network is used to perform convolution operations on text features. The basic principle of this method is to extract key information from the sequence by sliding the convolution kernel (filter) with a specific stride on the sequence. After the convolution calculation is completed, the feature dimension is scaled by max pooling, and finally, a linear layer is used to align the scale of the extracted features with the scale of the graph features.

9. The system according to claim 1, characterized in that, The feature fusion module employs a cross-attention mechanism, enabling the model to assign different weights to each part of the input data so that the model can extract more relevant and critical information. The cross-attention module integrates two labeled dot product attention modules, collectively called the scaled dot product attention layer. Multi-head scaled dot product attention can also be achieved by stacking multiple scaled dot product attention layers in parallel. The scaled dot product attention layer can convert the query and a set of key-value pairs into a weighted sum of values. The weight of each value is calculated by a compatibility function between the query and the corresponding key. This module will query ( ), key ( ), value ( The feature matrix is ​​used as input, and the feature matrix is ​​processed... The input matrix is ​​obtained by performing a linear mapping. ), input matrix (dimension is) ), input matrix (dimension is) ): ; in Assuming This represents a piece of text. This indicates the number of words appearing in the text. The feature embedding dimension of a word; , , These are their respective weight matrices; After calculating the three values, calculate the dot product of the query key and all keys, and divide each result by . Then, the Softmax activation function is used to obtain the weight values: In the affinity prediction task, the cross-attention module helps the model learn the relationship between independent one-dimensional text sequences and three-dimensional graph structures through the interaction of two encoded embedding modalities. This allows the model to obtain more comprehensive and effective information from relevant protein modalities, thus promoting affinity prediction.

10. The system according to claim 1, characterized in that, The affinity prediction module predicts the affinity between molecular features and protein features, fuses the obtained features together, and performs prediction through a multi-layer fully connected neural network. The multilayer fully connected neural network consists of three fully connected layers with ReLU activation functions. A Dropout layer is added after the fully connected layers to reduce the risk of overfitting and improve the model's generalization ability. Since affinity prediction is considered a regression task, mean squared error (MSE) is used as the model loss function. ; in, The number of drug-target bindings. The affinity value predicted by the model. The affinity is the true value; the model is trained using the Adam optimizer as the optimization function.

Citation Information

Cited By

  • DTI prediction method based on LLM and RBMO

    CN121506234A

  • Molecular-protein affinity prediction method, system, equipment and medium

    CN121709080A