General prediction method, device and medium for drug targets based on self-supervised learning

The self-supervised learning method enhances drug-target prediction by leveraging unlabeled data to improve feature extraction, addressing the generalization and mechanism prediction limitations of existing methods, particularly for new drugs and targets.

CN116013428BActive Publication Date: 2025-07-15CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310097306.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-07-15
Estimated Expiration
2043-02-10

Smart Images

  • Figure CN116013428B_ABST
    Figure CN116013428B_ABST
Patent Text Reader

Abstract

The present invention discloses a general drug target prediction method, device and medium based on self-supervised learning. The method includes: using a compound feature extraction module to extract drug feature vectors: splitting the drug molecular structure into a sub-structure sequence, converting each sub-structure into a vector encoding to obtain a sequence vector, and inputting it into an encoder for feature extraction; wherein, using masked language model prediction, molecular descriptor prediction and molecular functional group prediction, and performing self-supervised training on the compound feature extraction module and these three prediction models based on the feature vectors of drug samples to obtain the compound feature extraction module; extracting the feature vectors of the target; based on the feature vectors of the drug and the target, using an automated machine learning model to perform task prediction between the drug and the target. The present invention is applicable to prediction tasks including drug-target interaction, binding affinity and mechanism of action, etc., and the prediction accuracy in each task is better than that of the same type of prediction method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning, and particularly relates to a general drug target prediction method, device and medium based on self-supervised learning. Background Art

[0002] The identification of drug-target interactions is the most critical step in drug discovery and drug development. It can help understand the mechanism of action of drugs at the system level and has important clinical guiding significance for drug repositioning. Traditional experimental methods for measuring drug-target interactions are time-consuming and expensive. Therefore, researchers have proposed various computational methods to predict the potential interactions between drugs and targets. If the interactions between drug small molecules and target proteins can be accurately predicted, efficient screening of compounds can be achieved, a large number of unnecessary biochemical experiments can be reduced, thereby accelerating the process of drug development and reducing the research and development costs. However, the generalization ability of existing computational methods still needs to be further improved. They can achieve good prediction results in known drugs or targets, but perform poorly in predicting unknown drugs or targets. Moreover, currently, the vast majority of computational methods can only be used for binary classification prediction of drug-target interactions or regression prediction of binding affinity, and cannot predict the mechanism of action of the interactions between the two. In fact, the identification of the mechanism of action has important guiding significance in clinical drug use.

[0003] Currently, the most direct and effective method to improve the generalization ability of the model is to increase the training data. However, the existing labeled data is obviously insufficient to train a high-precision drug target prediction model because the known interaction data is scarce, which is also the main reason for the insufficient generalization ability of current methods, especially in the prediction of new drugs and new targets. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a general drug target prediction method, device and medium based on self-supervised learning, which has strong scalability and good prediction performance, aiming at the disadvantages of insufficient generalization ability and inability to predict the mechanism of action in existing drug-target interaction prediction methods.

[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions:

[0006] A general drug target prediction method based on self-supervised learning, comprising:

[0007] (1) Using a compound feature extraction module to extract the feature vector of the drug: splitting the drug molecular structure into a sequence of several substructures, converting each substructure into a vector encoding to obtain a sequence vector, and inputting it into a Transformer encoder for feature extraction to obtain the feature vector of the drug;

[0008] Among them, the pre-training method of the compound feature extraction module is as follows: extract the feature vectors of each drug sample in the drug sample set, and use the extracted drug sample feature vectors to perform masked language model prediction, molecular descriptor prediction, and molecular functional group prediction respectively. Update all the parameters of the compound feature extraction module and these three prediction models by weighted fusion of the losses of these three prediction models and backpropagation;

[0009] (2) Use the protein pre-training model to extract the feature vector of the target;

[0010] (3) Based on the feature vectors of the drug and the target, use the automated machine learning model to perform task prediction between the drug and the target.

[0011] Furthermore, the specific process of step (1) is as follows:

[0012] First, use the RDKit toolkit to split the drug molecular structure into n sub-structure sequences S with a radius of 1:

[0013] S = (x1; x2;...; x n )

[0014] In the formula, x i represents the i-th sub-structure obtained by splitting the drug molecular structure;

[0015] Subsequently, perform vector encoding on each sub-structure and map it to a d-dimensional vector space:

[0016]

[0017] Among them is the d-dimensional vector representation obtained by vector encoding of the i-th sub-structure x i ;

[0018] Finally, input the d-dimensional vector representation set X of the drug into the multi-layer Transformer encoder for feature extraction of multi-head self-attention.

[0019] Furthermore, in the pre-training method of the compound feature extraction module, the loss function of the masked language model is defined as:

[0020]

[0021] In the formula, loss MLM represents the prediction loss of the masked language model, mask represents the set of masked sub-structures of the drug, i represents the sub-structure index in mask, and p(x i ) represents the probability that the prediction output is the real sub-structure x i .

[0022] Further, in the pre-training method of the compound feature extraction module, the loss function of the molecular descriptor prediction model is defined as:

[0023]

[0024] In the formula, loss MDP represents the prediction loss of the molecular descriptor prediction model, n is the number of molecular descriptors of the drug, and y i is the true value of the i-th molecular descriptor of the drug, calculated by RDKit, is the predicted value of the i-th molecular descriptor.

[0025] Further, in the pre-training method of the compound feature extraction module, the loss function of the molecular functional group prediction model is defined as:

[0026]

[0027] In the formula, loss MFGP represents the prediction loss of the molecular functional group prediction model, m is the number of functional groups, and z i is the binary label indicating that the drug contains the i-th functional group, 1 means the drug contains the corresponding functional group, and 0 means it does not. This label is calculated by RDKit, represents the predicted probability that the drug contains the i-th functional group.

[0028] Further, by weighted fusion of the losses of these three prediction models and performing backpropagation, the weighted fusion expression is:

[0029] loss = loss MLM + α·loss MDP + β·loss MFGP

[0030] In the formula, loss is the total loss of weighted fusion, loss MLM represents the prediction loss of the masked language model, loss MDP represents the prediction loss of the molecular descriptor prediction model, loss MFGP represents the prediction loss of the molecular functional group prediction model, and α and β are weighted coefficients.

[0031] Further, the task prediction between the drug and the target includes: whether there is an interaction between the drug and the target, the strength of the interaction between the drug and the target, or whether the interaction between the drug and the target is an activating effect or an inhibitory effect.

[0032] Further, the protein pre-training model uses the existing protein language model ESM-2.

[0033] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor implements the general drug target prediction method described in any one of the above.

[0034] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the general drug target prediction method described in any one of the above is implemented.

[0035] Beneficial effects

[0036] Existing databases store a vast amount of unlabeled data, including drug compounds and target proteins. Therefore, the present invention uses a vast amount of unlabeled data to pre-train a large-scale self-supervised model. By mining the implicit relationships between compound substructures or protein subsequences from the vast amount of data, this model can accurately extract the feature vectors of drug molecules and target proteins. Thus, in various downstream drug target prediction tasks, for unknown drugs or targets, good prediction effects can also be obtained based on the learned substructure and subsequence information, effectively improving the generalization ability of the downstream task prediction model, and having strong scalability, and can be applied to multiple drug target-related prediction tasks including drug-target interaction, binding affinity, and mechanism of action. Description of the drawings

[0037] Figure 1 It is the overall architecture diagram of the method described in the embodiments of the present application. Detailed implementation manners

[0038] The embodiments of the present invention will be described in detail below. Based on the technical solutions of the present invention, detailed implementation manners and specific operation processes are given, and the technical solutions of the present invention are further explained and illustrated.

[0039] This embodiment provides a general drug target prediction method based on self-supervised learning. Referring to Figure 1 as shown, the method includes the following steps:

[0040] I. Pre-training of drug compounds

[0041] The input of the pre-training model is the SMILES string of the compound. RDKit is used to split the compound into a sequence of substructures with a radius of 1. Then, the substructures of the compound are encoded into feature vectors and input into the Transformer encoder to extract the implicit relationships and features between the substructures. Finally, the extracted feature vectors are respectively used for masked language model prediction, molecular descriptor prediction, and molecular functional group prediction, and the prediction losses of these models are weighted and fused, and all model parameters are updated through backpropagation.

[0042] More specifically, for an input drug compound, assuming its SMILES string is "CCCON", all substructures with a radius of 1 are extracted using the Morgan algorithm of RDKit, obtaining a substructure sequence ("CC", "CCC", "CCO", "CON", "ON"). Then these substructures are encoded into learnable Embedding vectors, and each substructure has a corresponding Embedding vector, with the same substructures sharing the same Embedding vector. Then the encoded sequence of Embedding vectors is input into the Transformer encoder to calculate the self-attention between substructures and perform feature extraction. The calculation of self-attention is as follows:

[0043]

[0044] where Q, K, and V are linear transformations of the input sequence, all with a dimension of d. The Transformer encoder incorporates multiple self-attention mechanisms and stacks multiple identical modules to improve the model's expressive power.

[0045] Next, the feature vectors extracted by the Transformer encoder are used for masked language model prediction, molecular descriptor prediction, and molecular functional group prediction respectively. These three models are all simple neural network models. Among them, the masked language model is a multi-classification prediction problem, where part of the substructure sequence in the input is randomly masked and the masked part is predicted based on the context information of the unmasked substructures. Molecular descriptor prediction is a regression task aimed at predicting the true values of all molecular descriptors of the input compound. Molecular functional group prediction is a multi-label classification problem aimed at predicting which functional groups the input compound contains. Finally, the prediction losses of these three models are weighted and fused, and backpropagation is performed to update all the parameters of the model, including the Embedding vectors, the Transformer encoder, and the parameters of the three prediction models. Through continuous iterative training until the model converges, a trained compound pre-training model is obtained. In this embodiment, the masked language model masks 15% of the substructures for each compound as prediction labels, while molecular descriptor prediction uses 123 molecular descriptors as prediction true values, and molecular functional group prediction uses 60 functional groups as prediction labels.

[0046] II. Pre-training of target proteins

[0047] Regarding the pre-training part of the target protein, in this embodiment, the protein model ESM-2 trained by the Meta AI research team is directly used. The input of this model is the protein sequence. The Transformer encoder is also used for self-attention calculation and feature extraction between amino acids. The prediction model is only the masked language model. ESM-2 trained multiple protein language models of different scales using more than 100 million protein sequences. In this embodiment, a model with 650 million parameters is used as the feature extraction model for the target.

[0048] III. Prediction of Downstream Tasks

[0049] The pre-trained compound and protein models have learned rich semantic information between substructures and subsequences, and can extract accurate compound and protein feature vectors, which can be widely applied to downstream prediction tasks related to drug targets. The present invention mainly involves the prediction of drug-target interaction, binding affinity, and mechanism of action. First, the compound and protein pre-training models are used to extract the feature vectors of drugs and targets respectively. Then, the feature vectors of the two are concatenated as the input of the automatic machine learning model AutoGluon. AutoGluon improves the accuracy and stability of the model by fusing multiple models that do not require hyperparameter search. Finally, the predictions of drug-target interaction, binding affinity, and mechanism of action are made respectively. Among them, the drug-target interaction prediction is a binary classification problem, that is, predicting whether a given drug-target pair has an interaction. The predicted labels are 1 and 0. 1 indicates that the corresponding drug-target pair has a known interaction, and 0 indicates that there is no interaction. The binding affinity prediction is used to evaluate the strength of their interaction. The predicted label is a continuous value after log transformation, indicating the binding affinity between the corresponding drug-target pair. The prediction of the mechanism of action is mainly used to determine whether the interaction between the drug and the target is an activation effect or an inhibitory effect. The present invention divides the prediction of the mechanism of action into two models. One model is used to predict whether a given drug-target pair has an activation effect, and the other model is used to predict whether there is an inhibitory effect between the two. Both models are binary classification predictions.

[0050] IV. Experimental Verification

[0051] To verify the effectiveness of using the present invention [hereinafter referred to as GFDTI] for drug target prediction and its performance superiority compared with other methods, this section evaluates the performance of GFDTI through extensive experiments. The following comparative experiments were conducted on 6 datasets in three prediction tasks of drug-target interaction, binding affinity, and mechanism of action. Each prediction task contains 2 different datasets, and each comparative experiment was conducted with three settings: warm start, drug cold start, and target cold start. The warm start setting means that both the drugs and targets in the test set appear in the training set. Drug cold start means that the drugs in the test set do not appear in the training set, and target cold start means that the targets in the test set do not appear in the training set. For each prediction task, some corresponding classical models were selected as the baseline models for experimental comparison. To ensure the fairness of the experimental comparison, the same random seed was used for cross-validation of all datasets. Each dataset was divided into a training set and a test set. Each method was trained on the same training set, and the test results of the obtained models were tested on the test set. In addition, AUC and AUPR were used as evaluation indicators for drug-target interaction prediction and mechanism of action prediction, while mean squared error MSE and concordance index CI were used as evaluation indicators for binding affinity prediction.

[0052] The experimental results of each prediction task are shown in Tables 1, 2, and 3 respectively.

[0053] Table 1 Performance comparison of GFDTI and other baseline models in drug-target interaction prediction

[0054]

[0055]

[0056] Table 2 Performance comparison of GFDTI and other baseline models in binding affinity prediction

[0057]

[0058] Table 3 Performance comparison of GFDTI and other baseline models in mechanism of action prediction

[0059]

[0060] As shown in Table 1, in the drug-target interaction prediction task, GFDTI achieved the best prediction performance under various experimental settings on all datasets. Especially on the yamanishi08 dataset with a small data scale, the prediction performance of GFDTI significantly outperformed other baseline models, indicating that the pre-trained model extracted accurate implicit features from a large amount of unlabeled data, so that only a small amount of labeled data was required to train an accurate model in downstream tasks. Additionally, it can be seen that the performance of other baseline models decreased significantly under the two cold-start experimental settings, while GFDTI still maintained a high prediction performance, suggesting that the substructure and subsequence information learned through pre-training can be effectively applied to the prediction of unknown drugs and targets. However, on the hetionet dataset with a large data scale, the performance advantage of GFDTI was not obvious, and other baseline models could also train accurate models when there was sufficient data volume.

[0061] From the results in Table 2, it can be seen that in the binding affinity prediction task, GFDTI also achieved the optimal prediction performance under various experimental settings on all datasets. Similarly, the performance advantage was more significant on the davis dataset with a small data scale, and the performance advantage was not obvious on the kiba dataset with a large data scale. Under the cold-start experimental setting, the prediction performance of all models decreased significantly, but GFDTI still maintained its performance advantage compared with other baseline models.

[0062] In the mechanism of action prediction task, as can be seen from Table 3, the prediction performance of GFDTI was significantly ahead of another baseline model under various experimental settings on all datasets. Consistent with the conclusions drawn from the previous two prediction tasks, GFDTI also had a more obvious performance advantage on the activator dataset with a small data volume. Under the drug cold-start experimental setting, the prediction performance of GFDTI on the two datasets was almost the same as that in the hot-start experimental setting. Under the target cold-start experimental setting, the prediction performance of both methods decreased significantly, but the prediction performance of GFDTI was still significantly ahead compared with another baseline model.

[0063] The above experimental results show that the pre-training of drugs and targets based on self-supervised learning proposed by the present invention can effectively improve the performance of downstream prediction tasks. Especially in downstream prediction tasks with insufficient labeled data, the pre-trained model can significantly improve the prediction performance. At the same time, for the prediction of unknown drugs and targets, GFDTI can also effectively improve the generalization ability and prediction performance of the model. This further demonstrates that the GFDTI method has learned rich implicit features and correlation relationships between drug substructures and target protein subsequences from a large amount of unlabeled data. Even for prediction tasks with insufficient data or for predicting unknown drugs and targets, GFDTI can still make accurate predictions relying on the implicit features learned during pre-training. In addition, GFDTI has achieved the best prediction performance in the above three prediction tasks, reflecting its strong scalability and being applicable to downstream prediction tasks related to drug targets.

[0064] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements based on this. Without departing from the patent concept of the present application, these transformations or improvements should fall within the scope of protection required by the present application.

Claims

1. A general prediction method for drug targets based on self-supervised learning, characterized in that, Including: (1) Using a compound feature extraction module to extract the feature vector of a drug: splitting the drug molecular structure into a sequence of several sub-structures, converting each sub-structure into a vector encoding to obtain a sequence vector, and inputting it into a Transformer encoder for feature extraction to obtain the feature vector of the drug; Among them, the pre-training method of the compound feature extraction module is: extracting the feature vectors of each drug sample in the drug sample set, respectively performing masked language model prediction, molecular descriptor prediction, and molecular functional group prediction using the extracted drug sample feature vectors, and updating all parameters of the compound feature extraction module and these three prediction models by weighted fusion of the losses of these three prediction models and performing backpropagation; (2) Using a protein pre-trained model to extract the feature vector of the target; (3) Based on the feature vectors of the drug and the target, using an automatic machine learning model to perform task prediction between the drug and the target.

2. The general prediction method for drug targets according to claim 1, wherein The specific process of step (1) is as follows: First, use the RDKit toolkit to split the drug molecular structure into a sequence S of n sub-structures with a radius of 1: S = (x1; x2;...; x n ) where x i represents the i-th sub-structure obtained by splitting the drug molecular structure; Subsequently, perform vector encoding on each sub-structure and map it to a d-dimensional vector space: wherein is the d-dimensional vector representation obtained by vector encoding of the i-th sub-structure x i ; Finally, input the d-dimensional vector representation set X of the drug into a multi-layer Transformer encoder for feature extraction of multi-head self-attention.

3. The general prediction method for drug targets according to claim 1, characterized in that In the pre-training method of the compound feature extraction module, the loss function of the masked language model is defined as: where loss MLM represents the prediction loss of the masked language model, mask represents the set of substructures where the drug is masked, i represents the substructure index in mask, p(x i ) represents the probability that the prediction output is the true substructure x i .

4. The general prediction method for drug targets according to claim 1, characterized in that In the pre-training method of the compound feature extraction module, the loss function of the molecular descriptor prediction model is defined as: where loss MDP represents the prediction loss of the molecular descriptor prediction model, n is the number of molecular descriptors of the drug, and y i is the true value of the i-th molecular descriptor of the drug, calculated by RDKit, is the predicted value of the i-th molecular descriptor.

5. The general prediction method for drug targets according to claim 1, characterized in that, In the pre-training method of the compound feature extraction module, the loss function of the molecular functional group prediction model is defined as: where loss MFGP represents the prediction loss of the molecular functional group prediction model, m is the number of functional groups, and z i is the binary label indicating that the drug contains the i-th functional group, where 1 means the drug contains the corresponding functional group and 0 means it does not. This label is calculated by RDKit, represents the predicted probability that the drug contains the i-th functional group.

6. The general prediction method for drug targets according to claim 1, characterized in that The weighted fusion of the losses of these three prediction models and performing backpropagation, the weighted fusion expression is: loss = loss MLM + α·loss MDP + β·loss MFGP where loss is the total loss of weighted fusion, loss MLM represents the prediction loss of the masked language model, loss MDP represents the prediction loss of the molecular descriptor prediction model, loss MFGP represents the prediction loss of the molecular functional group prediction model, and α and β are weighting coefficients.

7. The general prediction method for drug targets according to claim 1, characterized in that The task prediction between the drug and the target includes: whether there is an interaction between the drug and the target, the strength of the interaction between the drug and the target, or whether the interaction between the drug and the target is an activation effect or an inhibitory effect.

8. The general prediction method for drug targets according to claim 1, characterized in that The protein pre-trained model uses the existing protein language model ESM-2.

9. An electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the computer program is executed by the processor, the processor is enabled to implement the general drug-target prediction method described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the general drug-target prediction method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Drug and target interaction prediction method and device, equipment and storage medium

    CN113160894A

  • Drug small molecule property prediction method, device and equipment based on graph neural network

    CN113707236A