Compound-target binding affinity prediction method based on multi-modal feature fusion

Through the deep learning method of multimodal feature fusion, the problem of compound-target binding information integration was solved, high-precision compound-target affinity prediction was achieved, and the efficiency and accuracy of drug development were improved.

CN120613003APending Publication Date: 2025-09-09GUANGXI UNIVERSITY OF TECHNOLOGY +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510668101.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively integrating the multimodal information of compounds, lack the ability to model the synergistic/antagonistic effects of multiple targets, cannot accurately characterize the complex pharmacological mechanism of "one drug, multiple targets", and have insufficient predictive generalization capabilities for novel targets or rare structural compounds.

Method used

A multimodal feature fusion method is adopted to extract the multi-scale structural information of compounds and targets through convolutional neural networks and graph convolutional networks. Combined with bidirectional long short-term memory networks and graph attention networks, a compound-target binding affinity prediction model is constructed, and affinity prediction is performed using deep learning methods.

Benefits of technology

It achieves high-precision prediction of compound-target binding affinity, improves the success rate of drug development and reduces clinical research costs, and supports the screening of active ingredients in traditional Chinese medicines and the prediction of drug-target interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613003A_ABST
    Figure CN120613003A_ABST
Patent Text Reader

Abstract

The invention discloses a compound-target binding affinity prediction method based on multi-modal feature fusion, and relates to the technical field of compound activity prediction, and the technical key point is that the method comprises the steps of data acquisition, multi-modal feature characterization, multi-modal feature extraction and fusion, and affinity prediction to predict the compound activity. The problems of long time consumption, high cost, low efficiency and the like of compound activity prediction are solved through assistance of an artificial intelligence algorithm, and information can be processed in parallel in a self-adaptive and self-learning manner by referring to a multi-layered structure of a human brain and a layer-by-layer analysis processing mechanism of neuron information interaction. The intensity information of the interaction between the compound-target pair is provided by combining the affinity, the affinity between the compound and the target is predicted through a deep learning method, and the specific biological activity and key action target of the compound are analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of compound activity prediction, and in particular to a compound-target binding affinity prediction method based on multimodal feature fusion. Background Art

[0002] During drug development, analyzing the interactions between candidate compounds and biological targets is a core step in discovering lead compounds and optimizing drug activity. Modern drug molecules encompass a wide range of types, including small molecule compounds, biomacromolecule preparations, and natural product extracts. Their chemical space is highly complex: small molecule compounds exhibit stereoisomerism and diverse functional group combinations; protein drugs exhibit multi-domain synergistic effects; and natural active ingredients often possess unique topological structures, such as polycyclic fusions and dense chiral centers. These compounds can regulate disease-related signaling pathways by specifically binding to a variety of biological targets, such as enzymes, receptors, and ion channels. However, they also carry the potential risk of off-target effects leading to toxic side effects. Therefore, constructing an accurate compound-target interaction prediction system is of great significance for improving the success rate of drug development and reducing the cost of clinical research.

[0003] The current mainstream bioactivity evaluation systems can be divided into experimentally verified biological methods (including biophysical detection techniques such as surface plasmon resonance and isothermal titration calorimetry) and computationally based bioinformatics methods (such as molecular docking, free energy perturbation, and quantitative structure-activity relationship modeling). Although the former can provide high-confidence binding kinetic parameters, it has inherent defects such as strong equipment dependence, large sample consumption, and low throughput. Although the latter can achieve large-scale virtual screening, it still faces challenges in dealing with the following key issues: (1) It is difficult to effectively integrate multimodal information such as the sequence characteristics, three-dimensional conformation of the compound and the sequence conservation and dynamic conformational changes of the target protein; (2) It lacks the ability to model the synergistic / antagonistic effects of multiple targets and cannot accurately characterize the complex pharmacological mechanism of "one drug, multiple targets"; (3) The predictive generalization ability of novel targets or rare structural compounds is insufficient, and it is heavily dependent on the coverage of known activity data.

[0004] In recent years, deep learning technology has shown breakthrough potential in the field of drug discovery. Compared with traditional machine learning models, deep neural networks can automatically mine nonlinear relationships between compounds and targets through hierarchical feature extraction. However, existing deep learning methods still have the following limitations: (1) Single-modal feature representation cannot fully utilize the multi-scale structural information of compounds (such as the grammatical rules of SMILES sequences and the topological constraints of molecular graphs); (2) Protein representation is mostly limited to amino acid sequence analysis, ignoring the decisive role of secondary structure motifs and tertiary conformational space in the formation of binding pockets; (3) Multimodal feature fusion strategies are simple and fail to establish cross-modal correlation mapping between chemical space and biological space. Therefore, it is urgent to build a new multimodal fusion framework that can achieve high-precision compound-target binding affinity prediction by synergistically optimizing molecular sequence semantic understanding, spatial conformation analysis, and cross-modal correlation analysis, providing intelligent solutions for major needs such as anti-tumor drug design, antibiotic development, and immunomodulator discovery. Summary of the Invention

[0005] The purpose of the present invention is to solve the above problems and provide a compound-target binding affinity prediction method based on multimodal feature fusion.

[0006] In order to achieve the above object, the technical solution of the present invention is as follows:

[0007] The present invention provides a compound-target binding affinity prediction method based on multimodal feature fusion, comprising the following steps:

[0008] Step S1, data acquisition: obtaining the SMILES sequence of the compound and the amino acid sequence of the target protein;

[0009] Step S2, multimodal feature characterization: numerically represent the SMILES sequence of the compound and convert it into a molecular graph structure, numerically represent the amino acid sequence of the target protein, extract the secondary structure features of the target protein, and construct the adjacency matrix of the tertiary structure through the residue contact map;

[0010] Step S3, multimodal feature extraction and fusion: Convolutional neural networks and gated recurrent units are used to extract local and long-term sequence dependency features from the compound's SMILES sequence; graph convolutional networks and graph attention networks are used to extract spatial topological features from the molecular graph structure, and these features are concatenated into a 128-dimensional feature vector through global maximum pooling and average pooling; bidirectional long short-term memory networks are used to extract sequence context features from the amino acid sequence of the target protein; graph convolutional networks and graph attention networks are used to extract structural features from the secondary and tertiary structures; the four feature vectors of the compound and target protein are concatenated into a 512-dimensional fusion feature vector;

[0011] Step S4, affinity prediction: The fused feature vector is input into a regression model consisting of three fully connected layers, and the binding affinity value between the compound and the target protein is output through optimization of the mean square error loss function.

[0012] The present invention is further configured as follows: the node feature matrix of the molecular graph structure is generated by one-hot encoding, the edge index matrix is ​​generated based on the chemical bond connection relationship, and the molecular graph data structure is constructed by the PyTorch Geometric library.

[0013] The present invention is further configured as follows: the secondary structure characteristics of the target protein are predicted by the NetSurfP model, the residue contact map of the tertiary structure is constructed by the PconsC4 model, and the adjacency matrix elements are based on the Cβ atomic distance threshold between residues. determination.

[0014] The present invention is further configured as follows: the graph attention network model adopts a multi-head attention mechanism with 10 heads, and calculates the inter-node attention coefficient through the Leaky ReLU activation function, and the formula is as follows:

[0015]

[0016] e ij =LeakyReLU(α T [Wh i ||Wh i ])

[0017] Among them, α is a learnable parameter vector, || represents the vector concatenation operation, and e ij is the unnormalized attention score between node i and its neighbor node j, N i represents all first-order neighbor nodes of the central node i, h i represents the feature vector of node i, w is the learnable weight, α ij is the attention weight of node i to node j in layer l, and LeakyReLU is the activation function;

[0018] For each node i, the features of its neighbor j can be expressed as:

[0019]

[0020] Among them, || means concatenating the results of K groups of attention mechanisms. represents the normalized attention coefficient of the jth node obtained by the kth group of attention mechanism, w k is the linear transformation weight of the kth group, and σ is the activation function.

[0021] The present invention is further configured as follows: the bidirectional long short-term memory network model extracts contextual features of the amino acid sequence through forward and reverse LSTM layers, the hidden layer dimension is 64, and the output layer is mapped to a 128-dimensional feature vector through full connection.

[0022] The present invention is further configured as follows: the fully connected layer parameters of the regression model are set to:

[0023] First layer: input 512 dimensions, output 256 dimensions, activation function is ReLU;

[0024] Second layer: input 256 dimensions, output 128 dimensions, activation function is ReLU;

[0025] The third layer: input 128 dimensions, output 1 dimension;

[0026] A Dropout layer is added after each layer with a dropout rate of 0.2, an AdamW optimizer, and a learning rate of 0.005.

[0027] The present invention is further configured as follows: the performance evaluation indicators of the method include mean square error (MSE), consistency index (CI) and regression mean

[0028]

[0029] in is the predicted value of the i-th sample, y i is the true value of the i-th sample;

[0030]

[0031] Among them, b i , b j Represent the high prediction value and low prediction value respectively, δ i and δ j Represent the high true value and the low true value respectively, Z represents the normalization constant, and f(x) is the step function, whose formula is as follows:

[0032]

[0033] in, is the squared correlation coefficient between the true and predicted values ​​in the test set, is the squared correlation coefficient between the true and predicted affinity values ​​in the test set with an intercept of 0.

[0034] The present invention is further configured as follows: the method is applied to the screening of active ingredients of traditional Chinese medicines or the prediction of drug-target interactions, and supports the prediction of compound-target affinity for DAVIS and KIBA data sets.

[0035] Compared with existing technologies, this solution offers the following benefits: It uses artificial intelligence algorithms to address the time-consuming, costly, and inefficient nature of compound activity prediction. By drawing on the multi-layered structure of the human brain and its layer-by-layer analysis and processing mechanism for neuronal information interaction, it enables adaptive, self-learning parallel processing of information. Binding affinity provides information on the strength of compound-target interactions, enabling deep learning methods to predict compound-target affinity and evaluate the specific biological activity of a compound and its key target. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flow chart of the compound-target binding affinity prediction model in the embodiment of the present invention;

[0037] Figure 2 is a statistical analysis of the data set in the embodiment of the present invention;

[0038] Figure 3 This is the process of characterizing the molecular structure of the compound in the embodiment of the present invention;

[0039] Figure 4 is the secondary structure of the protein obtained according to NetSurfP in the embodiment of the present invention;

[0040] Figure 5 This is the process of characterizing the secondary and tertiary structures of proteins in the embodiments of the present invention;

[0041] Figure 6 It is the training result of the MGBDTA model in the example of the present invention. DETAILED DESCRIPTION

[0042] In order to enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the embodiments of the present invention and the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0043] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the embodiments.

[0044] Example:

[0045] A method for predicting compound target binding affinity based on multimodal feature fusion, such as Figure 1 , including the following steps:

[0046] 1. Dataset introduction: The DAVIS dataset and KIBA dataset are used to evaluate the performance of the prediction model. These two datasets are often used as benchmark datasets for compound-target affinity prediction. Figure 2 As shown in the figure, in the DAVIS dataset, the average SMILES sequence length of compounds is 54, with the lengths of SMILES sequences mainly concentrated in the range of 40 to 60. The average sequence length of kinase protein amino acid sequences is 743, with the lengths of kinase protein amino acid sequences mainly concentrated in the range of 400 to 1000. In the KIBA dataset, the average SMILES sequence length of compounds is 50, with the lengths of most SMILES sequences distributed around 50. The average sequence length of kinase protein amino acid sequences is 726, with the lengths of most kinase protein amino acid sequences ranging from 1000.

[0047] 2. Compound molecular characterization: Compounds are represented by their SMILES sequences and molecular graph structures.

[0048] 2.1. Representation of compound molecular sequences: SMILES sequence data involves a total of 24 atoms and symbols. A character-to-index mapping strategy is used to convert the characters in the SMILES sequence into numerical indices ranging from 1 to 24 as the input data of the model. Since the input sequence must have a fixed length when processing sequence data using a convolutional neural network, the SMILES sequence length is set to 54 based on the average length of the SMILES sequence in S1. If the numerical sequence length exceeds 54, it is truncated. For sequences with a length less than 54, zeros are added to padded to 54.

[0049] 2.2. Compound molecular graph structure representation: The SMILES sequence of the compound is converted into a two-dimensional molecular graph format using the RDKit chemical informatics software library. An undirected graph is constructed with atoms as nodes and chemical bonds as edges. The digital representation of the graph is the node feature matrix X∈R N×M and the adjacency matrix A∈R N×N (N represents the number of nodes, M represents the feature dimension of the node), such as Figure 3 The graph data representation process for the compound quercetin is shown in Figure 1. Five atomic attributes are used to describe atomic features, as shown in Table 1. Each attribute is uniquely encoded to form a 78-dimensional feature vector for the atom. The feature matrix is ​​composed of these eigenvectors.

[0050] Table 1 Atomic feature categories

[0051]

[0052]

[0053] 3. Protein molecule characterization: Protein molecules are characterized by their amino acid sequence, secondary structure and tertiary structure.

[0054] 3.1. Protein Amino Acid Sequence Representation: The length of protein amino acid sequences varies greatly. To conserve computing power, we fixed the protein sequence length to 1000 based on the average length of amino acid sequences in S1. Sequences longer than 1000 were truncated, and sequences less than 1000 were padded with the letter "Z." Because one-hot encoding of long sequences produces sparse vectors, which can lead to excessive dimensionality, each amino acid was mapped to an integer and embedded into a 128-dimensional feature vector using an embedding layer.

[0055] 3.2. Protein secondary structure representation: NetSurfP model is used to predict the secondary structure of protein, such as Figure 4 As shown, the secondary structure and solvent accessibility of each amino acid residue can be used as its property characteristics.

[0056] 3.3. Protein tertiary structure representation: The tertiary structure of a protein is represented by an amino acid residue contact map, which is a two-dimensional representation of the three-dimensional structure of a protein. The amino acid sequences of the protein are aligned using HHbilts, and the amino acid residue contact map is obtained by setting the threshold to 0.5 using PconsC4, and a contact map is generated, M∈R L×L , L represents the sequence length. Define the physicochemical properties of amino acids to construct a node feature matrix, and introduce the secondary structure features of amino acid sequences into the node feature matrix, as shown in Table 2. The process of extracting secondary and tertiary structure features of proteins is as follows Figure 5 shown.

[0057] Table 2 Amino acid residue characteristics

[0058]

[0059]

[0060] 4. Model design: A graph convolutional neural network combined with a graph attention neural network is used to extract the structural features of the compound molecules and the target protein. A bidirectional long short-term memory network is used to extract the amino acid sequence features of the target protein. A convolutional neural network combined with a gated recurrent unit is used to extract the SMILES sequence features of the compound. Feature fusion is performed in series, and input into three fully connected layers to predict the compound-target protein binding affinity, such as Figure 1 shown.

[0061] 4.1. Compound SMILES sequence feature extraction: A convolutional neural network combined with a gated recurrent unit was used to extract SMILES sequence features.

[0062] 4.2. Compound molecular graph structure feature extraction: The SMILES sequence is converted into a molecular graph structure using the Rdkit chemical informatics library. The node feature matrix and edge index matrix are obtained through the Pytorch Geometric library and input into the graph convolutional network combined with the graph neural network module of the graph attention network to extract the compound molecular graph structure features. After passing through the graph convolutional network and graph attention network layers, a global maximum pooling and global average pooling layer are input to extract molecular features. These two feature tensors are concatenated in the feature dimension to enrich the representation of the input features. Each layer of the network is activated by the ReLu activation function. After the pooling layer, the input is converted into 128-dimensional features in the linear layer.

[0063] The propagation of compound molecules in the graph convolutional network is shown in the following formula:

[0064]

[0065] in, I N is the identity matrix, for The degree matrix of H (l) is the input matrix of the lth layer, W (l) is the trainable weight matrix for each layer l, and σ is the activation function.

[0066] After passing through the graph convolutional network layer, it is passed to the graph attention network layer, and the softmax function is used to normalize the attention weights:

[0067]

[0068] Where α is a learnable parameter vector, || represents the vector concatenation operation, is the unnormalized attention score between node i and its neighbor node j, N i represents all first-order neighbor nodes of the central node i, Represents the feature vector of node i in the previous layer, W (l) is the learnable weight of layer l, is the attention weight of node i to node j in layer l, and Leaky ReLU is the activation function.

[0069] This paper uses a multi-head attention mechanism such as Figure 5 As shown, the characteristics of node i and node j can be expressed as:

[0070]

[0071] Among them, || means concatenating the results of K groups of attention mechanisms, represents the normalized attention coefficient of the jth node obtained by the kth group of attention mechanism in the lth layer, is the linear transformation weight of the kth group in the lth layer, and σ is the activation function.

[0072] 4.3. Protein sequence feature extraction: Each amino acid is mapped to an integer, and each integer is embedded into a 128-dimensional feature vector through the embedding layer. Subsequently, a two-layer bidirectional long-short-term attention network is applied to capture the dependencies between characters in the protein sequence, extract protein sequence features, and output 128-dimensional features through a fully connected layer.

[0073] 4.4. Extraction of protein secondary and tertiary structure features: The amino acid sequences of proteins are aligned using the HHbilts model, and then converted into amino acid residue contact maps using the CCMPred model as a representation of the protein tertiary structure. The secondary structural features of the amino acid sequences are extracted using the NetSurfP model and incorporated into the node feature matrix of the graph data. The secondary structural features are then passed into the graph convolutional network combined with the graph attention network model to extract the multi-level structural features of the protein.

[0074] 5. Model training: The four feature vectors are concatenated to form a 512-dimensional feature vector and input into three fully connected layers to predict the binding affinity value. A ReLu activation function and a Dropout layer are added after each fully connected layer. The mean square error is used as the loss function. The hyperparameter settings are shown in Table 3. The training results are shown in Figure 6 As shown in Figure 2, seven ablation experiments are used to verify the optimality of the model feature combination.

[0075] Table 3. Optimal hyperparameter settings for the model

[0076]

[0077] Table 3 Optimal hyperparameter settings for the model

[0078]

[0079] 5.1 Model Evaluation: Mean Square Error (MSE), Concordance Index (CI) and Regression toward the Mean ( ) evaluates the performance of the constructed compound-protein affinity prediction model, with values ​​ranging from 0 to 1. MSE is a commonly used metric to measure the difference between predicted and true values; a smaller MSE indicates better model prediction performance. CI measures whether the predicted and true values ​​have the same ranking; a higher CI indicates better model prediction performance. The closeness between the predicted results and the actual results can be evaluated. The larger the model, the better the performance. The specific formula is as follows:

[0080]

[0081] Among them, y i Describe the true value of the i-th sample, Represents the predicted value of the i-th sample.

[0082]

[0083] Among them, b i , b j Represent the high prediction value and low prediction value respectively, δ i and δ j Represent the high true value and the low true value respectively, Z represents the normalization constant, and f(x) is the step function, whose formula is as follows:

[0084]

[0085] Among them, r 2 is the squared correlation coefficient between the true and predicted values ​​in the test set, is the squared correlation coefficient between the true and predicted affinity values ​​in the test set with an intercept of 0.

[0086] 5.2 Comparison with baseline models: The comparison results of the MGBDTA prediction model and six baseline models on the DAVIS and KIBA datasets are shown in Table 3. The KronRLS model and SimBoost model are machine learning models, and DeepDTA, GraphDTA, FusionDTA, and KCDTA are deep learning models. The results show that in the DAVIS dataset, the performance evaluation indicators of the MGBDTA prediction model, MSE, CI, were 0.221, 0.894, and 0.71, respectively, among which CI and The MSE, CI, and PR of the MGBDTA model are all better than the other 6 benchmark models. The MSE is slightly higher than that of the KCDTA model, but better than the other 5 models. They are 0.174, 0.884, and 0.750 respectively. The best performing model in terms of MSE and CI is GraphDTA. The MSE is reduced by 0.025 compared with the MGBDTA model, and the CI is improved by 0.007 compared with the MGBDTA model. Compared with the present model, the MGBDTA model is 0.076 lower. It performs best in MSE and CI, and is slightly lower than the optimal baseline model, but has the best overall performance.

[0087] Table 3 Experimental results of baseline models on DAVIS and KIBA datasets

[0088]

[0089] 5.3 Ablation Experiments: To verify the impact of different feature representations on model performance and the optimality of the proposed model feature combination, models with different feature combinations were tested on the DAVIS and KIBA datasets. The comparison results of models with different feature combinations are shown in Table 4. The proposed MGBDTA model performed best in all evaluation metrics. On the DAVIS and KIBA datasets, the MSEs were 0.222 and 0.174, respectively, which were 0.039-0.135 and 0.010-0.063 lower than those of the other six models, respectively. The CIs were 0.894 and 0.862, respectively, which were 0.015-0.053 and 0.003-0.025 higher than those of the other six models, respectively. The performance of the M6 ​​model, which does not consider compound sequence features, and the M5 model, which does not consider protein structure, are significantly lower than that of the MGBDTA model, indicating that compound sequence features and protein structure features have a significant contribution to the prediction model. Compared with model M6, model M3 does not consider protein sequence feature representation, and the prediction performance of M3 is significantly lower, indicating that protein sequence feature representation helps improve model performance. Compared with model M5, model M1 does not consider compound graph structure features, and the prediction performance of M1 is lower than that of M5, indicating that compound graph structure feature representation can help improve model performance.

[0090] Table 4 Ablation experiment results of DAVIS and KIBA dataset models

[0091]

[0092] The above specific embodiments are merely explanations of the present invention and are not limitations of the present invention. After reading this specification, those skilled in the art may make non-creative modifications to the embodiments as needed. However, as long as they are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. A compound-target binding affinity prediction method based on multimodal feature fusion, characterized in that: The following steps are involved: Step S1, data acquisition: obtaining the SMILES sequence of the compound and the amino acid sequence of the target protein; Step S2, multimodal feature characterization: numerically represent the SMILES sequence of the compound and convert it into a molecular graph structure using the RDKit cheminformatics software library, numerically represent the amino acid sequence of the target protein, extract the secondary structure features of the target protein using NetSurfP, and construct an amino acid residue contact map as an adjacency matrix of the tertiary structure using CCMPred; Step S3, multimodal feature extraction and fusion: A convolutional neural network combined with a gated recurrent unit is used to extract local and long-term dependency features of the compound's SMILES sequence; a graph convolutional network and a graph attention network are used to extract spatial topological features of the molecular graph structure, and the dimension is reduced to a 128-dimensional feature vector through global maximum pooling and average pooling; a bidirectional long short-term memory network is used to extract sequence context features of the target protein's amino acid sequence; Graph convolutional networks and graph attention networks are used to extract structural features of secondary and tertiary structures; the four types of feature vectors of the compound and target protein are spliced ​​into a 512-dimensional fusion feature vector; Step S4, affinity prediction: The fused feature vector is input into the three-layer fully connected layer for affinity prediction, and the binding affinity value between the compound and the target protein is output through mean square error loss function optimization.

2. The method for predicting compound-target binding affinity based on multimodal feature fusion according to claim 1, wherein: The node feature matrix of the molecular graph structure is generated by one-hot encoding, the edge index matrix is ​​generated based on the chemical bond connection relationship, and the molecular graph data structure is constructed using the PyTorch Geometric library.

3. The compound-target binding affinity prediction method based on multimodal feature fusion according to claim 1, characterized in that: The secondary structure features of the target protein were predicted by the NetSurfP model, the residue contact map of the tertiary structure was constructed by the PconsC4 model, and the adjacency matrix elements were based on the Cβ atomic distance threshold between residues. determination.

4. The method for predicting compound-target binding affinity based on multimodal feature fusion according to claim 1, wherein: The graph attention network model adopts a multi-head attention mechanism with 10 heads. The inter-node attention coefficient is calculated by the Leaky ReLU activation function. The formula is as follows: Use the softmax function to normalize the attention weights: e ij =LeakyReLU(α T [Wh i ||Wh i ]) α is a learnable parameter vector, || represents the vector concatenation operation, e ij is the unnormalized attention score between node i and its neighbor node j, N i represents all first-order neighbor nodes of the central node i, h i represents the feature vector of node i, W is the learnable weight, α ij is the attention weight of node i to node j, and Leaky ReLU is the activation function; Using the multi-head attention mechanism, for each node i, the features of its neighbor j can be expressed as: Among them, || means concatenating the results of K groups of attention mechanisms. represents the normalized attention coefficient of the jth node obtained by the kth group of attention mechanism, w k is the linear transformation weight of the kth group, and σ is the activation function.

5. The compound-target binding affinity prediction method based on multimodal feature fusion according to claim 1, characterized in that: The bidirectional long short-term memory network model extracts contextual features of the amino acid sequence through forward and reverse LSTM layers, the hidden layer dimension is 64, and the output layer is mapped to a 128-dimensional feature vector through full connection.

6. The method for predicting compound-target binding affinity based on multimodal feature fusion according to claim 1, wherein: The fully connected layer parameters of the regression model are set as: First layer: input 512 dimensions, output 256 dimensions, activation function is ReLU; Second layer: input 256 dimensions, output 128 dimensions, activation function is ReLU; The third layer: input 128 dimensions, output 1 dimension; A Dropout layer is added after each layer with a dropout rate of 0.2, an AdamW optimizer, and a learning rate of 0.

005.

7. The method for predicting compound-target binding affinity based on multimodal feature fusion according to claim 1, wherein: The performance evaluation indicators of the method include mean square error (MSE), consistency index (CI) and regression mean in is the predicted value of the i-th sample, y i is the true value of the i-th sample: Among them, b i , b j Represent the high prediction value and low prediction value respectively, δ i and δ j Represent the high true value and the low true value respectively, Z represents the normalization constant, and f(x) is the step function, whose formula is as follows: in, is the squared correlation coefficient between the true and predicted values ​​in the test set, is the squared correlation coefficient between the true and predicted affinity values ​​in the test set with an intercept of 0.

8. The method for predicting compound-target binding affinity based on multimodal feature fusion according to claim 1, wherein: The method is applied to the screening of active ingredients and key targets of traditional Chinese medicines or the prediction of drug-target interactions.

Citation Information

Cited By

  • Protein post-translational modification prediction method based on multi-modal deep learning

    CN121096427A