A drug-target interaction prediction method and system based on robust multi-core ensemble method
Through the drug target interaction prediction method based on the robust multi-core integration method, the challenges of data scarcity, noise interference and multi-view information fusion in the prior art are solved, and high-precision and strong robust drug target interaction prediction are achieved, and the coverage of the drug target database is expanded.
Patent Information
- Application Number
- CN202411710162.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Existing drug target interaction prediction methods face challenges such as scarcity of data, noise interference and multi-view information fusion, resulting in insufficient prediction accuracy and robustness.
Using a drug target interaction prediction method based on a robust multi-core integration method, we use multi-core learning and ensemble learning strategies to optimize model parameters to improve prediction accuracy and robustness by constructing objective functions, including loss functions, integrated learning terms and regular terms.
It realizes high accuracy, strong robustness and drug target interaction prediction suitable for multiple data structures, which can effectively reduce the interference of noise data, improve the flexibility and adaptability of prediction, and expand the coverage of drug target database.
Smart Images

Figure CN119207547B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer biology, and in particular relates to a drug target interaction prediction method based on a robust multi-core integration method and a system thereof. Background Art
[0002] Drug-Target Interaction (DTI) refers to the specific interaction between drug molecules and molecular targets in biological systems. These interactions may occur through a variety of mechanisms, such as binding to receptors, inhibiting enzyme activity, or regulating signaling pathways. Understanding drug-target interactions is crucial because it helps us understand the mechanism of action of drugs, how drugs bind to target molecules, and how drugs exert their therapeutic effects. By identifying interactions between drugs and targets, researchers can reveal the mechanism of action of drugs and design more efficient and precise targeted treatment strategies.
[0003] However, known drug-target interaction data are still relatively scarce, and many potential drug-target interactions have not yet been discovered, which poses a huge challenge to drug development and disease treatment. To address this problem, researchers have proposed a variety of computational drug-target interaction prediction methods in recent years. These methods can be roughly divided into model-based methods and deep learning-based methods.
[0004] Model-based methods mainly predict drug-target interactions by leveraging prior knowledge and information about drugs and targets. These methods use known drug-target interaction data, combined with similarities between drugs and targets or other relevant information, to perform mathematical modeling to infer potential drug-target interactions. For example, using matrix decomposition methods, graph regularization techniques, and graph convolutional networks, these methods have alleviated the challenges of data scarcity and multi-view learning to a certain extent.
[0005] Deep learning methods encode the representations of drugs and targets, extract potential features from them, and combine these representations for interactive prediction. Deep learning methods usually rely on multi-layer neural networks, graph neural networks (GCNs), convolutional neural networks (CNNs), attention mechanisms and other technologies for modeling. These methods can effectively capture complex nonlinear relationships and have strong learning capabilities.
[0006] Although existing drug-target interaction prediction methods have achieved certain results, there are still many problems to be solved. For example, there are still challenges in how to deal with data noise, optimize computational complexity, and ensure model robustness; there is noisy data in the interaction matrix, and traditional loss functions (such as L2 loss) often perform poorly when dealing with these noisy data. In addition, multi-view learning and the fusion of multiple data structures remain an important challenge. Summary of the invention
[0007] The purpose of the present invention is to propose a new drug-target interaction prediction method to overcome the multiple problems such as data scarcity, noise interference, and multi-view information fusion in the existing drug-target interaction prediction tasks in order to overcome these limitations.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A drug-target interaction prediction method based on a robust multi-core ensemble method, the method comprising:
[0010] Construct the objective function, including the loss function of the predicted interaction matrix, the ensemble learning term and the regularization term;
[0011] Training the model using known interaction matrices, drug similarity kernel matrix sets, and target similarity kernel matrix sets as training data to optimize the objective function;
[0012] The drug similarity core matrix set includes the drugs to be tested, and the target similarity core matrix set includes the targets to be tested;
[0013] After training, the model outputs a predicted interaction matrix containing the confidence of the interaction between the drug to be tested and the target to be tested.
[0014] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, the drug similarity kernel matrix set and the target similarity kernel matrix set are obtained based on the known interaction matrix through multiple kernel functions.
[0015] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, each drug similarity kernel matrix also includes drug side effect similarity data, drug chemical structure similarity data, and substructure similarity data;
[0016] Each target similarity kernel matrix also includes protein sequence similarity data, protein-protein interaction data, and gene ontology functional annotation data;
[0017] And the kernel function includes Gaussian similarity kernel, cosine similarity kernel and correlation coefficient kernel;
[0018] The drug similarity kernel matrix set is:
[0019] (4)
[0020] The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix;
[0021] The fourth to sixth items are the similarity of drug side effects, drug chemical structure similarity, and substructure similarity;
[0022] The target similarity kernel matrix set is:
[0023] (5)
[0024] The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix;
[0025] The fourth to sixth items are protein sequence similarity, protein-protein interaction, and gene ontology functional annotation, respectively.
[0026] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, in the model, a weighting coefficient is assigned to each kernel matrix and the optimal similar kernel matrix is obtained through training:
[0027] (6)
[0028] (7)
[0029] and They are the weighted coefficients of the i-th drug similarity kernel matrix and the j-th target similarity kernel matrix updated through training, respectively.
[0030] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, the ensemble learning item is constructed in the following manner:
[0031] It is assumed that there are four latent structures in the data: drug data structure, target data structure, drug-target pair data structure, and low-rank structure;
[0032] Construct the objective functions of the four structures respectively;
[0033] By weighted combination of four objective functions, an integrated learning item for adaptive fusion of different structures is obtained.
[0034] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, the objective functions of the four structures are:
[0035] , drug data structure objective function;
[0036] , target data structure objective function;
[0037] , drug-target pair data structure objective function;
[0038] , low-rank structure objective function;
[0039] Respectively represent drug data, target data, drug-target pair data, and parameters of low-rank matrix, which are updated through training;
[0040] is the predicted interaction matrix, F is the Frobenius norm;
[0041] The ensemble learning term obtained by weighted combination of four objective functions is as follows:
[0042] (15)
[0043] They are the weighted terms of the four objective functions, which are updated through training.
[0044] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, the loss function includes the robust loss function , used to update the model parameters according to the error of the predicted interaction matrix compared with the known interaction matrix;
[0045] The robust loss function combines Loss of precision and Robustness of loss.
[0046] In the above-mentioned drug-target interaction prediction method based on the robust multi-core ensemble method, the regularization term includes a regularization term on the learning parameters required by the model;
[0047] The parameters that the model needs to learn include parameter A corresponding to drug data, parameter B corresponding to target data, parameter α and low-rank matrix (U, V) of drug-target pair data, regularization term of weighted parameter w of four structural objective functions, and weight coefficient of drug similarity kernel matrix. Nuclear Target Similarity Matrix The weight coefficient of .
[0048] In the above-mentioned drug-target interaction prediction method based on robust multi-core ensemble method, the robust loss function The definition is as follows:
[0049] (10)
[0050] in is the prediction error, is a hyperparameter that controls the degree of smoothing;
[0051] This method model uses an alternating optimization method to train the model.
[0052] A drug-target interaction prediction system comprises a processor for performing drug-target interaction prediction by executing the drug-target interaction prediction method.
[0053] The advantages of the present invention are:
[0054] 1. This solution implements a DTI (drug-target interaction) prediction method that is highly accurate, robust, and applicable to a variety of data structures, providing more powerful technical support for drug development and precision medicine;
[0055] 2. This model uses multi-kernel learning, multi-perspective information fusion and integrated learning strategies, combined with a robust loss function, to achieve high accuracy in the reconstruction of the drug-target interaction matrix. It can expand the existing drug target database by predicting unknown drug-target interactions.
[0056] 3. This model assumes that the data conforms to four different structures, including drug-target pair structure, drug structure, target structure and low-rank structure. It optimizes the combined weights of these structures through integrated learning, effectively reduces the interference caused by noisy data, and can adaptively fuse information of different structures to improve the flexibility and adaptability of prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A graphical display of the drug-target interaction prediction model based on the robust multi-core integration method proposed in the present invention;
[0058] Figure 2 is the docking pose of the predicted interaction between the drug (DrugBankID: DB00734) and the target protein (UniprotID: P41595);
[0059] Figure 3 is the docking pose of the predicted interaction between the drug (DrugBankID: DB00307) and the target protein (UniprotID: P13631). DETAILED DESCRIPTION
[0060] The present invention provides a drug-target interaction prediction method based on a robust multi-core ensemble method, which proposes a new method to overcome these limitations by introducing the following key technologies:
[0061] Designed a new The loss function combines Loss of precision and The robustness of the loss effectively reduces the interference caused by noise data (i.e., undiscovered interactions), thereby improving the reconstruction accuracy of the interaction matrix.
[0062] Multi-perspective integration of multi-kernel learning: Multi-kernel learning is used to automatically select the feature kernels most relevant to the drug target prediction task and assign weights according to their importance in the task, thereby achieving effective fusion of multi-perspective information.
[0063] Data structure assumptions and ensemble learning: It is assumed that the data conforms to four different structures, including drug-target pair structure, drug structure, target structure, and low-rank structure, and the weights of these structural combinations are optimized through ensemble learning, effectively improving the accuracy of prediction and the generalization ability of the model.
[0064] like Figure 1 As shown, the specific method is that the input of the model includes the following aspects:
[0065] (1) Known interaction matrix (Known interaction matrix), which records the known interaction relationships between drugs and targets. This data can be obtained from an experimentally verified drug-target interaction database, such as DrugBank.
[0066] (2) Protein sequence similarity , based on the standardized Smith-Waterman score, helps the model understand the evolutionary relationships and functional similarities between proteins. This data can be obtained through tools such as BLAST (Basic Local Alignment Search Tool).
[0067] (3) Known protein-protein interactions , providing possible physical or functional connections between proteins, and this data can be obtained from databases such as STRING.
[0068] (4) Gene Ontology (GO) functional annotation of proteins , providing a detailed description of protein function. The GO database was established by the Gene Ontology Consortium, which classifies and summarizes all gene-related research results in the world and uniformly defines and describes gene and protein functions.
[0069] (5) Similarity of drug side effects , provides information on possible side effects of drugs, which can be obtained from drug databases such as SIDER.
[0070] (6) Chemical structure similarity of drugs , provides chemical structure information of drug molecules, which can be obtained through cheminformatics tools and databases.
[0071] (7) Substructural similarity of drugs , which provides information on specific substructures in drug molecules. This data can be obtained from chemical structure databases such as ChEMBL.
[0072] In addition, the drug core matrix and target core matrix are based on the known interaction matrix Specifically, the following three commonly used kernel functions are used to measure the similarity between drugs and the similarity between targets to obtain the drug kernel matrix and target kernel matrix respectively:
[0073] Gaussian Similarity Kernel:
[0074] (1)
[0075] Cosine similarity kernel:
[0076] (2)
[0077] Correlation coefficient kernel:
[0078] (3)
[0079] in, is the covariance, is the variance, represents two drugs whose similarity is measured, Represents two targets to be measured by similarity, and the kernel function calculates the similarity between drugs and targets by extracting the features of drugs and targets from the known interaction matrix.
[0080] Finally, the drug similarity core matrix set and target similarity core matrix set are:
[0081] (4)
[0082] The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix;
[0083] The fourth to sixth items are the similarity of drug side effects, similarity of drug chemical structures, and similarity of substructures, respectively.
[0084] (5)
[0085] The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix;
[0086] The fourth to sixth items are protein sequence similarity, protein-protein interaction, and gene ontology functional annotation, respectively.
[0087] in, and are the number of drug and target core matrices, respectively.
[0088] This method uses multi-kernel learning (MKL) to fuse information from different sources (drugs and targets). Specifically, a weighting coefficient is assigned to each kernel matrix, and the contributions of different kernels are combined to obtain the optimal drug similarity matrix and target similarity matrix. The final drug and target similarity kernel matrix is calculated using the following linear weighting method:
[0089] (6)
[0090] (7)
[0091] in, and are the weighted coefficients of the i-th drug similarity kernel matrix and the j-th target similarity kernel matrix updated through training, respectively. These two weighted coefficients satisfy the constraints:
[0092] (8)
[0093] (9)
[0094] This proposal proposes a new loss, which effectively reconstructs the DTI interaction matrix and improves robustness. The loss function may not perform well when facing outliers. The loss combined loss and -loss (corretropy-inducedloss) to improve robustness to outliers. The loss function is defined as follows:
[0095] (10)
[0096] in is the prediction error, is a hyperparameter that controls the degree of smoothing. For positive errors (the predicted value is smaller than the actual value but larger than the observed value), C-loss provides robustness; for negative errors, The loss penalizes it to avoid underestimating the interaction.
[0097] This scheme assumes that there are four potential structures in the data: drug data structure, target data structure, drug-target pair data structure and low-rank structure. Specifically, the objective functions of the four structures are:
[0098] (11)
[0099] (12)
[0100] (13)
[0101] (14)
[0102] in, They represent drug data, target data, drug-target pair data, and low-rank matrix representations respectively. By weighted combination of these loss functions and introduction of regularization terms, adaptive fusion of different structures is finally achieved:
[0103] (15)
[0104] Furthermore, this scheme sets regularization on the learned parameters to avoid overfitting:
[0105] (16)
[0106] The final objective function can be expressed as:
[0107] (17)
[0108] After constructing the above objective function, the interaction matrix Y and the drug similarity kernel matrix set are known. Similar to the target core matrix set The model of the objective function is trained for the training data, where the drug similarity kernel matrix set contains the drug to be tested, the target similarity kernel matrix set contains the target to be tested, and the known interaction matrix can contain the drug-target interaction matrix to be tested, but it is displayed as 0. After the training is completed, the model outputs a predicted interaction matrix containing the confidence of the interaction between the drug to be tested and the target to be tested. The 0 of the drug-target interaction matrix to be tested will become a continuous value of 0-1, indicating the confidence of the interaction.
[0109] Furthermore, this scheme trains the model through an alternating optimization algorithm. The optimization process of the objective function is non-convex, so we adopt an alternating optimization method, fixing other parameters in each step and updating only one parameter group until convergence.
[0110] The above-mentioned model proposed in this scheme utilizes multi-kernel learning, multi-view information fusion and integrated learning strategy, and combines robust loss function to achieve high accuracy in the reconstruction of drug-target interaction matrix. This embodiment takes DrugBank database as an example. After training on database version 3.0, 17 of the top 50 prediction results were verified in the new version (6.0).
[0111] Table 1. Drug-target interactions predicted by this validated approach.
[0112]
[0113]
[0114] The top-ranked DTI predictions, especially rank 1 (DrugBankID: DB00734, UniProtID: P41595) and rank 3 (DrugBankID: DB00307, UniProtID: P13631), are not supported by any known experimental or clinical evidence in the existing literature. To further investigate the validity of these predictions, computational docking studies were performed to evaluate the biological activities of the DTI predictions of rank 1 and 3 made by DTI-RME. Atomistic molecular dynamics (MD) simulations were used to simulate docking interactions, and MD simulations were performed using the AmberTools package, and the AMBERff19SB and AM1-BCC force fields were applied to the top-scoring compound-protein complexes.
[0115] The simulation results of ranking 1 are as follows Figure 2 As shown, the interaction with residue ASP-135 suggests that this amino acid may play a key role in hydrogen bonding or electrostatic interactions, potentially stabilizing the ligand. Residues such as GLU, PHE, and THR form hydrogen bonds (indicated by green dashed lines) and hydrophobic contacts (marked by curved green regions) with DB00734, further supporting the potential binding of the drug.
[0116] The simulation results for ranking 3 are as follows Figure 3 As shown, we can see that there is Stacking, which is a non-covalent interaction between aromatic rings that is usually strong, even in the absence of hydrogen bonding, plays an important role in the stabilization of the ligand-receptor complex.
[0117] From the above, we can see that the method proposed in this scheme can expand the existing drug target database by predicting unknown drug-target interactions, providing more powerful technical support for drug development and precision medicine.
[0118] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A drug-target interaction prediction method based on a robust multi-core ensemble method, characterized in that: The method includes: Construct the objective function, including the loss function of the predicted interaction matrix, the ensemble learning term and the regularization term; The loss function includes a robust loss function , used to update the model parameters according to the error of the predicted interaction matrix compared with the known interaction matrix; The robust loss function combines Loss of precision and Robustness of loss; The ensemble learning item is constructed as follows: It is assumed that there are four latent structures in the data: drug data structure, target data structure, drug-target pair data structure, and low-rank structure; Construct the objective functions of the four structures respectively; By weighted combination of four objective functions, an integrated learning item for adaptive fusion of different structures is obtained; Training the model using known interaction matrices, drug similarity kernel matrix sets, and target similarity kernel matrix sets as training data to optimize the objective function; The drug similarity kernel matrix set includes the drug to be tested, the target similarity kernel matrix set includes the target to be tested, and the drug similarity kernel matrix set and the target similarity kernel matrix set are obtained based on a known interaction matrix through multiple kernel functions, and the multiple kernel functions include a Gaussian similarity kernel, a cosine similarity kernel and a correlation coefficient kernel; Each drug similarity kernel matrix also includes drug side effect similarity data, drug chemical structure similarity data, and substructure similarity data; Each target similarity kernel matrix also includes protein sequence similarity data, protein-protein interaction data, and gene ontology functional annotation data; The drug similarity kernel matrix set is: (4) The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix; The fourth to sixth items are the similarity of drug side effects, drug chemical structure similarity, and substructure similarity; The target similarity kernel matrix set is: (5) The first to third terms are calculated using the Gaussian similarity kernel, the cosine similarity kernel, and the correlation coefficient kernel according to the known interaction matrix; The fourth to sixth items are protein sequence similarity, protein-protein interaction, and gene ontology functional annotation, respectively; After training, the model outputs a predicted interaction matrix containing the confidence of the interaction between the drug to be tested and the target to be tested; The objective functions of the four structures are: , drug data structure objective function; , target data structure objective function; , drug-target pair data structure objective function; , low-rank structure objective function; Respectively represent drug data, target data, drug-target pair data, and parameters of low-rank matrix, which are updated through training; is the predicted interaction matrix, F is the Frobenius norm; The ensemble learning term obtained by weighted combination of four objective functions is as follows: (15) They are the weighted terms of the four objective functions, which are updated through training.
2. The drug-target interaction prediction method based on the robust multi-core ensemble method according to claim 1, characterized in that: In the model, a weight coefficient is assigned to each kernel matrix and the optimal similar kernel matrix is obtained through training: (6) (7) and They are the weighted coefficients of the i-th drug similarity kernel matrix and the j-th target similarity kernel matrix updated through training, respectively.
3. The drug-target interaction prediction method based on the robust multi-core ensemble method according to claim 2, characterized in that: The regularization term includes a regularization term about the learning parameters required by the model; The parameters that the model needs to learn include parameter A corresponding to drug data, parameter B corresponding to target data, parameter α and low-rank matrix (U, V) of drug-target pair data, regularization term of weighted parameter w of four structural objective functions, and weight coefficient of drug similarity kernel matrix. Nuclear Target Similarity Matrix The weight coefficient of .
4. The drug-target interaction prediction method based on the robust multi-core ensemble method according to claim 3, characterized in that: Robust loss function The definition is as follows: (10) in is the prediction error, is a hyperparameter that controls the degree of smoothing; This method uses an alternating optimization method to train the model.
5. A drug-target interaction prediction system, characterized in that: It comprises a processor for performing drug-target interaction prediction by executing the drug-target interaction prediction method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Drug and side effect identification method based on self-weighted multi-kernel learning
CN111477344A
Drug target association prediction method and device based on heterogeneous network
CN118248206A