Drug-target interaction prediction method based on multistage collaborative fusion enhancement network
By combining a multi-level collaborative fusion enhancement network with attention mechanisms and data augmentation algorithms, the problems of global information integration and data scarcity in drug-target interaction prediction are solved, achieving higher prediction accuracy and model stability.
Patent Information
- Application Number
- CN202510794428.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies neglect the integration of global information about drugs and targets in drug-target interaction prediction, and the scarcity of labeled data leads to insufficient utilization of negative samples, resulting in problems such as model overfitting.
We employ a multi-level collaborative fusion enhancement network approach, combining attention mechanisms and data augmentation algorithms. Through feature extraction, multi-level collaborative fusion, and synthesis enhancement modules, we generate new positive samples, expand the dataset, reduce the impact of label noise, and improve prediction accuracy.
By effectively integrating the multi-level features of drugs and targets and their interaction features, the accuracy of drug-target interaction prediction is improved, the problems of data scarcity and model overfitting are solved, and the prediction performance is enhanced.
Smart Images

Figure CN120853733A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drug-target interaction prediction technology, and in particular to a drug-target interaction prediction method based on a multi-level synergistic fusion enhancement network. Background Technology
[0002] Drug development has a long history and is crucial for improving human health and medical standards. However, traditional drug development faces challenges such as long development cycles, high investment, and high risks, especially in response to emerging diseases. Therefore, shortening development cycles and reducing costs have become critical issues that urgently need to be addressed. With the rapid development of big data and computing technologies, high-throughput data in the biomedical field is accumulating rapidly, and computer technology is being effectively applied to drug-target interaction (DTI) prediction. This method can screen potential drug target pairs from massive amounts of data, optimize drug screening and retargeting processes, and thus significantly improve development efficiency.
[0003] However, most existing studies focus only on extracting the independent features of drugs and targets, neglecting the integration of drug-target interaction features and global information. Therefore, accurately and deeply extracting feature representations of drugs and targets is crucial for drug-target interaction prediction research. Furthermore, the limited availability of labeled data is also a major challenge in this field. Due to the scarcity of labeled data, the selection and processing of negative samples are particularly critical. Although traditional methods such as random negative sampling, fixed-proportion negative sampling, and simple data augmentation alleviate the data shortage problem to some extent, these methods may ignore potential negative samples, leading to insufficient utilization of negative samples and potentially causing problems such as model overfitting. Summary of the Invention
[0004] In view of this, the present invention provides a drug-target interaction prediction method based on a multi-level collaborative fusion enhancement network, which applies attention mechanism and data augmentation algorithm to drug-target interaction analysis, solves the shortcomings of existing research in drug-target interaction prediction, and can effectively improve the accuracy of drug-target interaction prediction.
[0005] Therefore, the present invention provides the following technical solution: A drug-target interaction prediction method based on a multi-level collaborative fusion enhancement network includes: Preprocessing of drug and target data; A drug-target interaction prediction model based on a multi-level collaborative fusion enhancement network was constructed, which includes a feature extraction module, a multi-level collaborative fusion module, a synthesis enhancement module, and a prediction module. The feature extraction module extracts local and global features from the drug and target, respectively. The fusion characteristics of drug targets are obtained through a multi-level collaborative fusion module; The synthesis enhancement module generates new positive samples based on the synthesis enhancement algorithm, expanding the dataset and reducing the impact of label noise. The prediction module predicts the final outcome of drug-target interactions through a fully connected layer.
[0006] A further improvement to the technical solution of this invention lies in: preprocessing the drug and target data, including: Drugs were represented using SMILES and converted into molecular graphs using RDKit to extract topological information about the drugs. ; Convert the SMILES string into a drug feature matrix using Smi2Vec. ; For the target sequence, word2vec is used to embed it into a real-valued vector representation. Then, Prot2Vec is used to divide the target sequence into overlapping subsequences of 3 characters each, and a target feature matrix is generated. .
[0007] A further improvement of the technical solution of this invention lies in that: the feature extraction module includes: global target feature extraction, global drug topology feature extraction, and drug-target pair local interaction feature extraction, and the feature extraction process includes: First, a graph attention network is used to extract drug-specific topological features. ; The drug's topological structure is extracted as independent features through a T-layer graph attention network, where T is the total number of layers in the graph attention network; after T convolutions, the aggregate representation of the drug molecule is replaced by... : (1) in, This indicates the number of atoms in a drug molecule diagram. Indicates the first in the figure The feature vectors obtained by each atom after passing through a T-layer graph attention network; Secondly, semantic features of the target are extracted using a multi-scale one-dimensional convolutional network. ; Sequence features of the target Feature extraction using multi-scale one-dimensional convolution Then, filters with multiple different convolutional kernels are used to further extract representative features of semantic information: (2) in, The number of different convolution kernels, and For learnable parameters, It is a non-linear function, obtained by splicing different... Obtain global target feature representation
[0008] Finally, the interaction feature representation of the drug and target is obtained using the drug feature decoder and target feature encoder. and ; In the local interaction feature extraction part of the drug-target pair, the target feature matrix Drug feature matrix As input information for this part; for drug orientation, the drug feature matrix and target feature matrix The data are input into the drug feature decoder and the target feature encoder, respectively, to obtain the local interaction feature representation of the drug. : (3) (4) in, This represents the sigmoid activation function. This represents element-wise multiplication. This represents a one-dimensional convolution operation. and All are learnable parameters; This is a characteristic representation of the drug at the next higher level. This represents the characteristics of the drug after this change. This is a representation of the target features; For the target direction, the target feature matrix Drug feature matrix The data are input into the target feature decoder and the drug feature encoder, respectively; then, following the steps above, the final interactive feature representation of the target is obtained.
[0009] (5) (6) in, and All are learnable parameters; This represents the characteristics of the target at the next higher level. This represents the target characteristics after this change. This describes the characteristics of the drug.
[0010] A further improvement to the technical solution of this invention lies in that: the multi-level collaborative fusion module consists of a local-global multimodal fusion submodule and a cross-guided attention submodule; the fusion process is as follows: (1) In the local-global multimodal fusion submodule, the importance of each residue in the binding process is evaluated through the attention mechanism, which integrates local and global features and can fuse rich biological features from the drug and the target to obtain a comprehensive feature representation of the drug. X d Comprehensive feature representation of the target X p : (7) in, Represents the global atomic diagram features of a drug. Representing local and local interaction features, Represents the first attention scaling factor, and three sets of learnable parameters. These are used to map local and global features to query, key, and value spaces, respectively. Then, using the scaling dot product formula in the attention mechanism, they are fused to obtain the final multimodal feature representation of the drug. ; (8) in, This represents the global feature input of the target. This represents the local feature input of the target. For learnable parameters, It is the second attention scaling factor, and the final fused feature representation of the target is calculated through the local-global multimodal fusion function. ; (2) In the cross-modal guided attention submodule, the comprehensive features of the drug and the target are fused through an interactive cross-modal guided attention mechanism to obtain the drug-target interaction embedding vector. This effectively extracts the interaction characteristics between the drug and the target; Following the local-global multimodal fusion submodule, the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network learns a more comprehensive and in-depth representation of drug features. With target feature representation Then, a cross-modal mutual guidance attention mechanism is introduced for further processing; First, for drug-oriented approaches, target features guide the learning of drug features; given the learned drug representation... Target representation The affinity matrix was calculated.
[0011] (9) in, These are weight parameters; Then, the entire affinity matrix As a feature input, the drug-target interaction prediction model based on multi-level collaborative fusion enhancement network can learn to automatically generate attention maps of drugs and targets, thereby obtaining richer contextual information and improving the overall performance of the drug-target interaction prediction model based on multi-level collaborative fusion enhancement network. (10) (11) (12) (13) in, and All are learnable parameters. To represent the drug after incorporating target attention-guided information. To represent the target after incorporating drug attention guidance information. This represents the attention probability distribution of the drug region. This represents the attention probability distribution of the target region; (14) in, This is the attention-weighted drug representation vector. This is the target representation vector after attention weighting.
[0012] A further improvement of the technical solution of this invention lies in the following: the synthesis enhancement module is a drug-target pair synthesis enhancement strategy based on synthetic minority class oversampling technology, aiming to solve the impact of limited data and label noise on the drug-target interaction prediction model based on multi-level collaborative fusion enhancement network; the SMOTE-based drug-target pair synthesis enhancement strategy generates new positive sample drug-target pairs by learning the distribution of existing positive sample drug-target pairs, as shown in the following formula: (15) in, Represents the feature vector of the original sample. Indicates from positive samples of A neighbor sample vector randomly selected from the nearest neighbors. This represents a random number between 0 and 1, used to control new samples. exist and The interpolation ratio between them.
[0013] A further improvement to the technical solution of this invention lies in the following: the prediction module predicts the final result of drug-target interaction through a fully connected layer, specifically including: integrating multi-level collaborative fusion features through a fully connected layer to finally predict the prediction score of drug-target interaction, and determining whether the drug and target interact based on the prediction score; the prediction label is calculated using the sigmoid function and backpropagated using the cross-entropy loss function, including an L2 regularization term, to optimize the parameters of the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network. (16) (17) in, This represents the sigmoid activation function. This is the final predicted label probability value. and These are learnable weights and bias parameters; the loss function uses cross-entropy loss, where... Indicates the first The true label of each sample This represents the label probability predicted by the model. It is the set of all learnable parameters in the entire end-to-end model. yes The coefficients of the regularization term are used to prevent the model from overfitting.
[0014] A further improvement of the technical solution of the present invention is that the threshold for the drug target interaction prediction score is set to 0.5. If the score is greater than 0.5, it is considered that there is an interaction; otherwise, there is no interaction.
[0015] Advantages and positive effects of the present invention: 1. This invention proposes a novel Multi-Level Collaborative Fusion Enhancement Network (MCFAM), in which the multi-level collaborative fusion module consists of a local-global multimodal fusion submodule and a cross-guided attention submodule. It aims to more effectively extract global interaction information between drugs and targets, as well as their respective local substructure features. These extracted features will be effectively fused, which can effectively integrate the multi-level features of drugs and targets and their interaction features, thereby enhancing the accuracy of prediction.
[0016] 2. In order to avoid the model focusing too much on overlapping information and to reduce the negative impact on downstream tasks, this invention introduces learnable parameters.
[0017] 3. This invention designs a synthesis enhancement module based on the SMOTE algorithm. By selecting similar positive samples and generating new positive samples in the feature space, it can make full use of negative samples while ensuring that the proportion of positive samples is not too small. This effectively increases the diversity and quality of training data, thereby reducing the bias of the model in practical applications and effectively solving the problem of data scarcity.
[0018] 4. The drug-target interaction prediction method based on multi-level collaborative fusion enhancement network provided by this invention can achieve state-of-the-art (SOTA) results on multiple datasets. Attached Figure Description To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the drug-target interaction prediction method based on a multi-level collaborative fusion enhancement network provided by the present invention. Figure 2 A schematic diagram of the overall structure of the drug-target interaction prediction model based on a multi-level collaborative fusion enhancement network provided by this invention; Figure 3 This is a detailed diagram of the multi-level collaborative fusion module provided by the present invention; Figure 4 This is a detailed diagram of the synthesis enhancement module provided by the present invention. Detailed Implementation
[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0021] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0022] like Figure 1 As shown, a drug-target interaction prediction method based on a multi-level collaborative fusion enhancement network specifically includes the following: S1. Preprocess drug and target data; Specifically, drugs are represented using SMILES, and RDKit is used to convert them into molecular graphs to extract the topological information of the drugs. Then, the SMILES string is converted into a target feature matrix using Smi2Vec. For the target sequence, word2vec is used to embed it as a real-valued vector representation. Then, Prot2Vec is used to divide the target sequence into overlapping subsequences of 3 characters each, and a drug feature matrix is generated. .
[0023] S2. Construct a drug-target interaction prediction model based on a multi-level collaborative fusion enhancement network, including a feature extraction module, a multi-level collaborative fusion module, a synthesis enhancement module, and a prediction module; such as... Figure 2 As shown.
[0024] S3, the feature extraction module extracts local and global features from the drug and target, respectively.
[0025] Specifically, the feature extraction module includes: global target feature extraction, global drug topology feature extraction, and drug-target pair local interaction feature extraction.
[0026] First, a graph attention network (GAT) is used to extract drug-independent topological features. G drug ; The drug's topological structure is extracted as independent features using a T-layer Gaussian Artificial Arithmetic (GAT), where T is the total number of GAT layers. After T convolutions, the aggregate representation of the drug molecule is replaced with... : (1) in, This indicates the number of atoms in a drug molecule diagram. Indicates the first in the figure The feature vectors obtained by passing each atom through a T-layer GAT.
[0027] Secondly, semantic features of the target are extracted using a multi-scale one-dimensional convolutional network. ; Sequence features of the target Feature extraction using multi-scale one-dimensional convolution Then, filters with multiple different convolutional kernels are used to further extract representative features of semantic information: (2) in, The number of different convolution kernels, and For learnable parameters, It is a non-linear function, obtained by splicing different... Obtain global target feature representation ; Finally, the interaction feature representation of the drug and target is obtained using the drug feature decoder (DFD) and target feature encoder (TFE). and .
[0028] In the local interaction feature extraction part of the drug-target pair, the target feature matrix Drug feature matrix This section serves as the input information for two directions: drug orientation and target orientation. For the drug orientation, the drug feature matrix... and target feature matrix The data are input into the drug feature decoder (DFD) and the target feature encoder (TFE), respectively, to obtain the local interaction feature representation of the drug. : (3) (4) in, This represents the sigmoid activation function. This represents element-wise multiplication. This represents a one-dimensional convolution operation. and All of these are learnable parameters. This is a characteristic representation of the drug at the next higher level. This represents the characteristics of the drug after this change. This is a representation of the target characteristics.
[0029] For the target direction, the target feature matrix Drug feature matrix The data are input into the target feature decoder (TFE) and the drug feature encoder (DFD), respectively. Then, following the steps described above, the final result is obtained. .
[0030] (5) (6) in, This represents the sigmoid activation function. This represents element-wise multiplication. This represents a one-dimensional convolution operation. and All of these are learnable parameters. This represents the characteristics of the target at the next higher level. This represents the target characteristics after this change. This describes the characteristics of the drug.
[0031] S4. A multi-level collaborative fusion module extracts and fuses the global structural features and local substructural features of the drug and target. During this integration process, the multi-level features of the target and drug, as well as their interaction features, are effectively combined. By introducing learnable parameters, the model avoids overemphasizing overlapping information, thereby mitigating the negative impact on downstream tasks.
[0032] Specifically, the multi-level collaborative fusion module consists of a local-global multimodal fusion sub-module and a cross-guided attention sub-module, such as... Figure 3 As shown.
[0033] (1) In the local-global multimodal fusion submodule, the importance of each residue in the binding process is evaluated by the over-attention mechanism, which integrates local and global features and can fuse rich biological features from drugs and targets to obtain comprehensive feature representations of drugs and targets respectively. This improves the ability of drug-target interaction prediction models based on multi-level collaborative fusion enhancement networks to handle hidden biases and shortcut learning problems during training.
[0034] (7) in, Represents the global atomic diagram features of a drug. Representing local and local interaction features, three sets of learnable parameters. These are used to map local and global features to query, key, and value spaces, respectively, utilizing the scaling dot product formula in the attention mechanism (where...). (Representing the first attention scaling factor), the final drug multimodal feature representation is obtained by fusion. ; (8) in, This represents the global feature input of the target. This represents the local feature input of the target. For learnable parameters, It is the second attention scaling factor, and the final fused feature representation of the target is calculated through the local-global multimodal fusion function. ; (2) In the cross-modal guided attention submodule, the comprehensive features of the drug and the target are fused through an interactive cross-modal guided attention mechanism to obtain the drug-target interaction embedding vector. It effectively extracts the interaction characteristics between drugs and targets.
[0035] Following the local-global multimodal fusion submodule, the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network learns a more comprehensive and in-depth representation of drug and target features. Next, a cross-modal mutual guidance attention mechanism is introduced for further processing. Unlike traditional methods that simply splice together drug and target features, the cross-modal mutual guidance attention mechanism effectively captures the interaction information between the drug and target through bidirectional attention calculation, thereby avoiding information loss, hidden bias, and shortcut learning problems.
[0036] First, for drug-oriented approaches, target features guide the learning of drug features. Given the learned drug representation... Target representation The affinity matrix was calculated. .
[0037] (9) in, These are the weighting parameters. Then, the entire affinity matrix is... As a feature input, allowing the model to learn to automatically generate attention maps of drugs and targets can obtain richer contextual information, thereby improving the overall model performance; (10) (11) (12) (13) in, and For learnable parameters, To represent the drug after incorporating target attention-guided information. To represent the target after incorporating drug attention guidance information. This represents the attention probability distribution of the drug region. This represents the attention probability distribution of the target region; (14) in, This is the attention-weighted drug representation vector. The target representation vector after attention weighting; S5. The synthesis enhancement module generates new positive samples based on the synthesis enhancement algorithm, expanding the dataset and reducing the impact of label noise. Specifically, such as Figure 4 As shown, to fully utilize negative samples and further extract comprehensive feature representations, a drug-target pair synthesis enhancement strategy based on Synthetic Minority Oversampling Technique (SMOTE) is proposed, aiming to address the impact of limited data and label noise on target prediction models. This strategy generates new positive drug-target pairs by learning the distribution of existing positive samples, thereby expanding the learnable data of the drug-target interaction prediction model based on a multi-level collaborative fusion enhancement network and solving the problem of increased label noise after enhancing negative samples.
[0038] (15) Among them, among them, Represents the feature vector of the original sample. Indicates from positive samples of A neighbor sample vector randomly selected from the nearest neighbors. This represents a random number between 0 and 1, used to control new samples. exist and The interpolation ratio between them.
[0039] S6. The prediction module uses a fully connected layer to predict samples.
[0040] The prediction module predicts the final result of drug-target interaction through a fully connected layer. Specifically, it integrates multi-level collaborative fusion features through the fully connected layer to predict the prediction score of drug-target interaction, and determines whether the drug and target interact based on the prediction score. The prediction label is calculated using the sigmoid function and backpropagated using the cross-entropy loss function, which includes an L2 regularization term to optimize the parameters of the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network.
[0041] (16) (17) in, This represents the sigmoid activation function. This is the final predicted label probability value. and These are learnable weights and bias parameters. The loss function uses cross-entropy loss, where... Indicates the first The true label of each sample This represents the label probability predicted by the model. It is the set of all learnable parameters in the entire end-to-end model. These are the coefficients of the L2 regularization term, used to prevent the model from overfitting.
[0042] The threshold for the predicted score of drug-target interaction is set to 0.5. If the score is greater than 0.5, an interaction is considered to exist; otherwise, there is no interaction.
[0043] The drug-target interaction prediction method (MCFAM) based on multi-level collaborative fusion enhancement network provided by this invention can achieve state-of-the-art results on multiple datasets, as shown in Table 1. Table 1. Detailed prediction results of each model on each dataset.
[0044] As shown in Table 1, the drug-target interaction prediction method (MCFAM) based on multi-level collaborative fusion enhancement network provided by this invention outperforms existing prediction methods (DeepConv-DTI, HyperAttentionDTI, IIFDTI, MCANet, etc.) in terms of AUC, AUPR, Precision, and Recall on the datasets DrugBank, BindingDB, GPCR, and CYP. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting drug-target interactions based on a multi-level collaborative fusion enhancement network, characterized in that, include: Preprocessing of drug and target data; A drug-target interaction prediction model based on a multi-level collaborative fusion enhancement network was constructed, which includes a feature extraction module, a multi-level collaborative fusion module, a synthesis enhancement module, and a prediction module. The feature extraction module extracts local and global features from the drug and target, respectively. The fusion characteristics of drug targets are obtained through a multi-level collaborative fusion module; The synthesis enhancement module generates new positive samples based on the synthesis enhancement algorithm, expanding the dataset and reducing the impact of label noise. The prediction module predicts the final outcome of drug-target interactions through a fully connected layer.
2. The drug target interaction prediction method according to claim 1, characterized in that, Preprocessing of drug and target data includes: Drugs were represented using SMILES and converted into molecular graphs using RDKit to extract topological information about the drugs. ; Convert the SMILES string into a drug feature matrix using Smi2Vec. ; For the target sequence, word2vec is used to embed it into a real-valued vector representation. Then, Prot2Vec is used to divide the target sequence into overlapping subsequences of 3 characters each, and a target feature matrix is generated. .
3. The drug target interaction prediction method according to claim 1, characterized in that, The feature extraction module includes: global target feature extraction, global drug topology feature extraction, and drug-target pair local interaction feature extraction. The feature extraction process includes: First, a graph attention network is used to extract drug-specific topological features. ; The drug's topological structure is extracted as independent features through a T-layer graph attention network, where T is the total number of layers in the graph attention network; after T convolutions, the aggregate representation of the drug molecule is replaced by... : (1) in, This indicates the number of atoms in a drug molecule diagram. Indicates the first in the figure The feature vectors obtained by each atom after passing through a T-layer graph attention network; Secondly, semantic features of the target are extracted using a multi-scale one-dimensional convolutional network. ; Sequence features of the target Feature extraction using multi-scale one-dimensional convolution Then, filters with multiple different convolutional kernels are used to further extract representative features of semantic information: (2) in, The number of different convolution kernels, and All are learnable parameters. It is a non-linear function, obtained by splicing different... Obtain global target feature representation Finally, the interaction feature representation of the drug and target is obtained using the drug feature decoder and target feature encoder. and ; In the local interaction feature extraction part of the drug-target pair, the target feature matrix Drug feature matrix As input information for this part; for drug orientation, the drug feature matrix and target feature matrix The data are input into the drug feature decoder and the target feature encoder, respectively, to obtain the local interaction feature representation of the drug. : (3) (4) in, This represents the sigmoid activation function. This represents element-wise multiplication. This represents a one-dimensional convolution operation. and All are learnable parameters; This is a characteristic representation of the drug at the next higher level. This represents the characteristics of the drug after this change. This is a representation of the target features; For the target direction, the target feature matrix Drug feature matrix The data are input into the target feature decoder and the drug feature encoder, respectively; then, following the steps above, the final interactive feature representation of the target is obtained. (5) (6) in, and All are learnable parameters; This represents the characteristics of the target at the next higher level. This represents the target characteristics after this change. This describes the characteristics of the drug.
4. The drug target interaction prediction method according to claim 1, characterized in that, The multi-level collaborative fusion module consists of a local-global multimodal fusion submodule and a cross-guided attention submodule; the fusion process is as follows: (1) In the local-global multimodal fusion submodule, the importance of each residue in the binding process is evaluated through the attention mechanism, which integrates local and global features and can fuse rich biological features from the drug and the target to obtain a comprehensive feature representation of the drug. X d Comprehensive feature representation of the target X p : (7) in, Represents the global atomic diagram features of a drug. Representing local and local interaction features, Represents the first attention scaling factor, and three sets of learnable parameters. These are used to map local and global features to query, key, and value spaces, respectively. Then, using the scaling dot product formula in the attention mechanism, they are fused to obtain the final multimodal feature representation of the drug. ; (8) in, This represents the global feature input of the target. This represents the local feature input of the target. For learnable parameters, It is the second attention scaling factor, and the final fused feature representation of the target is calculated through the local-global multimodal fusion function. ; (2) In the cross-modal guided attention submodule, the comprehensive features of the drug and the target are fused through an interactive cross-modal guided attention mechanism to obtain the drug-target interaction embedding vector. This effectively extracts the interaction characteristics between the drug and the target; Following the local-global multimodal fusion submodule, the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network learns a more comprehensive and in-depth representation of drug features. With target feature representation Then, a cross-modal mutual guidance attention mechanism is introduced for further processing; First, for drug-oriented approaches, target features guide the learning of drug features; given the learned drug representation... Target representation The affinity matrix was calculated. (9) in, These are weight parameters; Then, the entire affinity matrix As a feature input, the drug-target interaction prediction model based on multi-level collaborative fusion enhancement network can learn to automatically generate attention maps of drugs and targets, thereby obtaining richer contextual information and improving the overall performance of the drug-target interaction prediction model based on multi-level collaborative fusion enhancement network. (10) (11) (12) (13) in, and All are learnable parameters. To represent the drug after incorporating target attention-guided information. To represent the target after incorporating drug attention guidance information. This represents the attention probability distribution of the drug region. This represents the attention probability distribution of the target region; (14) in, This is the attention-weighted drug representation vector. This is the attention-weighted target representation vector.
5. The drug target interaction prediction method according to claim 1, characterized in that, The synthesis enhancement module is a drug-target pair synthesis enhancement strategy based on synthetic minority class oversampling technology, aiming to address the impact of limited data and label noise on drug-target interaction prediction models based on multi-level collaborative fusion enhancement networks. The SMOTE-based drug-target pair synthesis enhancement strategy generates new positive sample drug-target pairs by learning the distribution of existing positive sample drug-target pairs, as shown in the following equation: (15) in, Represents the feature vector of the original sample. Indicates from positive samples of A neighbor sample vector randomly selected from the nearest neighbors. This represents a random number between 0 and 1, used to control new samples. exist and The interpolation ratio between them.
6. The drug target interaction prediction method according to claim 1, characterized in that, The prediction module predicts the final result of drug-target interaction through a fully connected layer. Specifically, it integrates multi-level collaborative fusion features through the fully connected layer to predict the final drug-target interaction prediction score, and determines whether the drug and target interact based on the prediction score. The prediction label is calculated using the sigmoid function and backpropagated using the cross-entropy loss function, which includes an L2 regularization term, to optimize the parameters of the drug-target interaction prediction model based on the multi-level collaborative fusion enhancement network. (16) (17) in, This represents the sigmoid activation function. This is the final predicted label probability value. and These are learnable weights and bias parameters; the loss function uses cross-entropy loss, where... Indicates the first The true label of each sample This represents the label probability predicted by the model. It is the set of all learnable parameters in the entire end-to-end model. yes The coefficients of the regularization term are used to prevent the model from overfitting.
7. The drug target interaction prediction method according to claim 6, characterized in that, The threshold for the drug target interaction prediction score is set to 0.
5. If the score is greater than 0.5, an interaction is considered to exist; otherwise, there is no interaction.