Drug-Target Protein Association Prediction Method and System Based on Multi-Source Data Embedding

Through the multi-core learning algorithm and noise-decreasing autoencoder, the problem of high time and cost in drug development is solved, and efficient prediction of drug-target protein association is achieved, and prediction accuracy and stability are improved.

CN119920305BActive Publication Date: 2025-06-24HUNAN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510396484.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-24
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The traditional drug development process is time-consuming, costly and low success rate. The existing drug-target protein association prediction methods perform poorly on the imbalanced dataset and are susceptible to noise data.

Method used

Multi-core learning algorithm is used to integrate multiple similarity networks, combine noise reduction and dimensionality reduction of feature data, and multi-layer perceptrons are used to predict drug-target protein associations.

Benefits of technology

Improve the accuracy and stability of drug-target protein association prediction, reduce noise interference, and simplify operational procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920305B_ABST
    Figure CN119920305B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting drug-target protein associations based on multi-source data embedding. The method of multi-source data embedding consists of a multi-kernel learning algorithm and a denoising autoencoder. Based on the multi-kernel learning method, the robustness of the model is improved by integrating multiple kernel functions, the over-reliance on a single kernel function is reduced, and weights are assigned to different similarity networks, and finally a comprehensive similarity matrix kernel is obtained. Based on the denoising autoencoder method, by introducing noise into the input data and attempting to restore the original data from the noisy data, the functions of removing noise and restoring data are realized. The present invention is simple, and only needs to accurately predict the potential associations in the drug-disease association network according to the association data of drugs and target proteins. Moreover, the effectiveness of this method is proved by experiments, which has certain biological significance and has theoretical significance and reference value for drug repositioning and the promotion of the drug R & D process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computational biology and relates to a drug-target protein association prediction method and system based on multi-source data embedding. Background Art

[0002] Traditional drug development usually requires a long process of drug discovery, preclinical research, clinical research, and finally drug marketing. This process usually takes 10-15 years, plus more than $1.5 billion in funding, and the overall development difficulty is relatively high. However, with such huge consumption, the probability of successful development is less than 10%. Obviously, such R&D consumption and success probability are unacceptable. How to find some new ways to reduce the time and capital investment costs in the traditional drug development process and increase the success probability of drug development is a question worthy of in-depth thinking. Drug repositioning, as a new R&D strategy, uses computer-assisted calculations to predict biological data including biological associations such as drug-target proteins, reduce initial screening costs, and obtain more related information in the drug development process, which can effectively solve the above problems.

[0003] As time goes by and science and technology develop, more and more biology-related data are available for researchers to use. For example, drug-related databases such as DrugBank and SIDER record the chemical structure, biological activity, pharmacology, toxicology and other information of drugs. Disease-related databases such as MalaCards provide detailed descriptions of the name, definition, pathogenesis, symptoms, clinical characteristics and other aspects of the disease. Protein-related databases such as HPRD and UniProt collect information on the structure, function, interaction and expression pattern of proteins in different tissues and cell types. The development of various biological information-related databases has created unlimited possibilities for association prediction based on computer calculation methods. The computational method uses the chemical fingerprint information of drugs, the amino acid sequence information of proteins, and the association data of various aspects such as drugs, diseases, proteins, and side effects as a medium to mine the deep information in the data, and finally calculate the possibility of association between target proteins, and then further verify the association through biological wet experiments. Computational biology methods conduct in-depth research on various biological factors at different levels, which greatly promotes the theoretical understanding of the pathogenic mechanism of complex diseases in organisms, and improves the speed of drug development and the efficiency of successful drug development.

[0004] Analyzing the complex biological association network formed by biomolecules such as drugs, diseases, and target proteins is an important research content in bioinformatics. With the development of computers, using computational methods such as machine learning and deep learning to capture the deep information between data is crucial for the association prediction between the two. The machine learning-based method improves the performance of the model by obtaining the global structural information between drugs and diseases. Although this type of method takes into account the global structural information and effectively improves the prediction performance of the model, this type of method often performs poorly on unbalanced data sets. The deep learning-based method improves the prediction performance by deeply fusing multi-source drug and protein data. However, this type of method usually requires sufficient data for effective training, and at the same time involves a large number of hyperparameter optimizations, and iterative training is required to obtain good results; at the same time, outliers and noise data in the data may lead to a decline in the performance of the deep learning-based method.

[0005] However, the more data is collected, the more association information it contains. How to retain the effective information during the data information fusion process, reduce the impact of noise data on the results, and further improve the prediction performance of the model is a question worthy of consideration.

[0006] Therefore, it is necessary to design a drug-target protein association prediction method based on multi-source data embedding. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a drug-target protein association prediction method and system based on multi-source data embedding. This method fuses multiple similarity matrices based on the multi-kernel learning algorithm to obtain a comprehensive similarity matrix, and uses a denoising autoencoder to denoise and reduce the dimension of the feature data. It can accurately identify potential drug-target protein associations only based on the association matrix of drugs and target proteins, which has certain biological significance.

[0008] The technical solution provided by the present invention is as follows:

[0009] On the one hand, a drug-target protein association prediction method based on multi-source data embedding includes the following steps:

[0010] Step 1), based on the data sources of drugs and target proteins, calculate the similarity between drugs and between target proteins, and construct similarity networks of drugs and target proteins in each data source respectively;

[0011] Step 2), use the multi-kernel learning algorithm MKL to integrate the similarity networks of drugs and target proteins respectively, and obtain the comprehensive similarity matrices SD and SP of drugs and target proteins;

[0012] Step 3), use each row of data in the comprehensive similarity matrices SD and SP obtained in Step 2) as the initial feature representations of drugs and target proteins, and introduce noise into the initial feature data through a denoising autoencoder; and restore the initial feature data with noise introduced;

[0013] , ;

[0014] wherein, represents randomly generated noise, represents the mean of the distribution, represents the variance; represents the feature data of the original input, represents the feature data after adding noise;

[0015] The encoder and decoder used for restoration are as follows:

[0016] ;

[0017] ;

[0018] wherein, and respectively represent the outputs of the l-th layer of the encoder and decoder, and respectively represent the weight matrix and bias matrix of the l-th layer of the encoder, and σ(*) represents the non-linear activation function; represents the reconstructed data finally output by the decoder, and respectively represent the weight matrix and bias matrix of the l-th layer of the decoder;

[0019] Restoration is used to remove noise and restore data. The data after denoising has redundant or irrelevant information removed and retains the core features of the data, making it more suitable for subsequent modeling tasks. This enhances the model's ability to learn potential patterns in the data and reduces the interference of noise on prediction accuracy.

[0020] Step 4), splice the restored feature data of drugs and target proteins processed in Step 3) to obtain the feature vectors of drug-target protein pairs, and use a multi-layer perceptron to model the associations between the feature vectors of drug-protein pairs, obtaining a drug-target protein association prediction model based on the multi-layer perceptron to predict the association relationship between drugs and target proteins.

[0021] Furthermore, the prediction score of the association relationship between drugs and target proteins is calculated by the following formula:

[0022] ;

[0023] Among them, is the feature vector of the input drug-target protein pair, , , and , , are the corresponding weight and bias vectors determined in the drug-target protein association prediction model based on the multi-layer perceptron, The function calculation formula is , where e represents the natural exponent, and the Relu function calculation formula is , is the calculation variable of the input function.

[0024] The parameters in the drug-target protein association prediction model based on the multi-layer perceptron are usually optimized using the backpropagation algorithm and the gradient descent algorithm;

[0025] Furthermore, the construction process of the similarity networks of drugs and target proteins in each data source is as follows:

[0026] Step 1.1: Construction of the drug similarity network;

[0027] Obtain drug data sources, including at least the chemical structures of drugs, drug-drug associations, drug-disease associations, and drug-drug side effect data; for each data source, calculate the similarity between drugs using the Jaccard similarity coefficient, and each data source generates an independent drug similarity network, represented in matrix form, where each element represents the similarity score between a pair of drugs;

[0028] The Jaccard similarity coefficient obtains the similarity score by calculating the ratio of the intersection to the union of two drugs on a specific data source.

[0029] Step 1.2: Construction of the target protein similarity network;

[0030] Obtain target protein data sources, including at least the sequence information of target proteins, target protein-target protein interactions, and target protein-disease associations; for each data source, calculate the similarity between target proteins using the Jaccard similarity coefficient, and each data source generates an independent target protein similarity network;

[0031] The Jaccard similarity coefficient obtains the similarity score by calculating the ratio of the intersection to the union of two target proteins on a specific data source.

[0032] Furthermore, the multi-kernel learning algorithm MKL is used to integrate the similarity networks of drugs and target proteins respectively, and weight values are assigned to each similarity network, so as to obtain the comprehensive similarity matrices SD and SP of drugs and target proteins:

[0033] In the process of integrating the similarity network by using the multi-kernel learning algorithm MKL, the following objective function is used to minimize the distance between the optimal kernel of the integrated similarity network and the ideal optimal kernel:

[0034] Objective function: ,

[0035] ,

[0036] , ,

[0037] ;

[0038] Among them, represents the i-th similarity network, represents the weight vector of the similarity network of, is the regularization parameter, is the ideal optimal kernel of, β is a regularization parameter used to balance the weight of the similarity matrix fitting error and the weight regularization term; represents a drug or a target protein, represents a drug, represents a target protein; represents the drug-target protein association matrix; and respectively represent the comprehensive similarity matrices generated after integrating multiple similarity networks, * is the ideal optimal kernel matrix of; represents the number of similarity networks; represents the square of the Frobenius norm, represents the difference between the integrated comprehensive similarity matrix SX and the ideal optimal kernel matrix * ; represents the regularization term for the weight vector used to constrain the sparsity of the full-time and avoid overfitting;

[0039] To avoid overfitting during training and improve the generalization ability of the model, making the system performance more robust, an L2 penalty is imposed on the weight vector to limit the growth of its magnitude.

[0040] Weight solution formula:

[0041] ,

[0042] ,

[0043] ;

[0044] Where: ,

[0045] ,

[0046] ;

[0047] Among them, trace represents the trace of the matrix;

[0048] The transformation formula is:

[0049]

[0050] ;

[0051] represents the drug or target protein, and the weights of each similarity network of the drug and target protein are obtained according to the above formula.

[0052] The significance of multi-source data fusion is that since the similarity between drugs and target proteins can be described from multiple perspectives (such as chemical structure, sequence information, functional association, etc.), a single data source may not be able to fully reflect their similarity. Therefore, by integrating the similarity networks generated from multiple data sources, the complex relationships between drugs and target proteins can be captured more comprehensively, thereby improving the accuracy of subsequent prediction models.

[0053] The Multiple Kernel Learning (MKL) algorithm is used to integrate multiple similarity networks to generate a comprehensive similarity matrix kernel. Each similarity network can be regarded as a kernel function. The MKL algorithm assigns different weights to each kernel function to reflect its importance in the comprehensive similarity matrix. By using an optimization algorithm (such as gradient descent or quadratic programming) to determine the optimal weights of each kernel function, the comprehensive similarity matrix can better reflect the true relationship between drugs and target proteins. Finally, the MKL algorithm weights and fuses multiple similarity networks to generate a comprehensive similarity matrix kernel for subsequent drug-target protein association prediction models.

[0054] By using the multi-core learning algorithm, the model no longer depends on a single similarity network, but can comprehensively consider the information from multiple data sources. This can not only improve the prediction accuracy of the model, but also enhance the robustness of the model to noise and abnormal data, and avoid prediction errors caused by the deviation of a single data source.

[0055] Furthermore, the comprehensive similarity matrix SD of drugs and the comprehensive similarity matrix SP of target proteins are calculated by the following formula:

[0056] ;

[0057] ;

[0058] where respectively represent the numbers of drugs and target proteins, N D and N P are the numbers of similarity networks of drugs and target proteins respectively, and R represents the real number field.

[0059] Furthermore, the Softplus function and the RMSProp algorithm are used to optimize the mean square error formula in the dimensionality reduction process. The mean square error calculation formula is as follows:

[0060] ;

[0061] where L represents the loss function of the denoising autoencoder, represents the reconstructed data finally output by the decoder, represents the original input data input to the denoising autoencoder;

[0062] Reduce the loss of data information in the dimensionality reduction process;

[0063] Furthermore, the drug-target protein association prediction model based on the multi-layer perceptron adopts the Adam algorithm and sets the learning rate to 0.001 to optimize the binary cross loss:

[0064] ;

[0065] where and respectively represent the positive and negative sample labels of the drug-target protein pair, Y is the true label, with a value of 0 or 1, N represents the number of drug-target protein pairs, represents the prediction score of the association relationship between drugs and target proteins, represents the loss function of the drug-target protein association prediction model based on the multi-layer perceptron.

[0066] In a second aspect, a system for predicting the association between a drug and a target protein based on the above multi-source data embedding includes:

[0067] A similarity network construction module that calculates the similarity between drugs and between target proteins based on the data sources of drugs and target proteins, and constructs similarity networks of drugs and target proteins in each data source respectively;

[0068] A similarity matrix acquisition module that uses the multi-kernel learning algorithm MKL to integrate the similarity networks of drugs and target proteins respectively, and obtains the comprehensive similarity matrices SD and SP of drugs and target proteins;

[0069] An initial feature reduction module that takes each row of data in the obtained comprehensive similarity matrices SD and SP as the initial feature representation of drugs and target proteins, introduces noise into the initial feature data through a denoising autoencoder; and restores the initial feature data with noise introduced;

[0070] An association prediction module that splices the processed and restored feature data of drugs and target proteins to obtain the feature vectors of drug-target protein pairs, and uses a multi-layer perceptron to model the association between the feature vectors of drug-target protein pairs, and obtains a drug-target protein association prediction model based on the multi-layer perceptron to predict the association relationship between drugs and target proteins.

[0071] In a third aspect, a computer-readable storage medium stores a computer program, and the computer program is called by a processor to implement:

[0072] The steps of the above method for predicting the association between a drug and a target protein based on multi-source data embedding.

[0073] In a fourth aspect, an electronic device includes:

[0074] One or more processors;

[0075] A storage device for storing one or more programs;

[0076] The one or more processors execute the one or more programs to implement the steps of the above method for predicting the association between a drug and a target protein based on multi-source data embedding.

[0077] Advantageous effects:

[0078] The present invention provides a method (MKLDAE) and system for predicting drug-target protein associations based on multi-source data embedding. This method takes into account the differences in the amount of information contained in different association matrices. The sparser the association matrix, the less information it contains, and vice versa. A multi-kernel learning approach is used to dynamically assign fusion weights to the association matrices, so as to retain as much information as possible in the fused similarity matrix. To reduce the impact of noise in the data on the prediction results, a denoising autoencoder is used to denoise and reduce the dimensionality of the feature data, improving the stability of the model. This prediction system has a simple structure and is easy to operate;

[0079] Compared with existing methods for predicting drug-target protein associations, the MKLDAE method described in the present invention has the following advantages:

[0080] 1) Using the multi-kernel learning algorithm, dynamically assign corresponding weights during the fusion process of the similarity matrix, so that the amount of information contained in the fused similarity matrix is maximized.

[0081] 2) Using a denoising autoencoder to reduce the dimensionality and denoise the fused feature data, reducing the impact of noisy data on performance and improving the stability of the model.

[0082] The present invention is easy to implement. By using the association data of drugs and target proteins, it can accurately predict the potential associations between drugs and target proteins. Experimental verification shows that the MKLDAE method described in the present invention can effectively predict the associations between drugs and target proteins. At the same time, by comparing with other methods, the AUROC and AUPR values of MKLDAE have obtained the highest values, and the performance is better. For the specific comparison and analysis of the experimental results graph, please refer to the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 is a schematic flow chart of the method (MKLDAE) for predicting drug-target protein associations based on multi-source data embedding provided by an embodiment of the present invention;

[0084] Figure 2 is an ROC curve graph of the prediction method (MKLDAE) provided by an embodiment of the present invention and other comparison methods;

[0085] Figure 3 is a PR curve graph of the prediction method (MKLDAE) provided by an embodiment of the present invention and other comparison methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] The following will further elaborate on the present invention in detail with reference to the drawings and specific embodiments.

[0087] Example 1:

[0088] A method for predicting drug-target protein associations based on multi-source data embedding, as Figure 1 shown, includes the following steps:

[0089] In Figure 1 , Drug Similarity Networks represents the drug similarity network, Protein SimilarityNetworks represents the protein similarity network, Multiple Kernels Learning represents multi-kernel learning, Fusionsimilarity network represents the fusion similarity network, DAE represents the denoising autoencoder, Loss represents the loss, Decode represents decoding, Encode represents encoding, Drug representation represents the drug representation, Proteinrepresentation represents the target protein representation, Fully Connected Classification represents full connection classification, Concatenating represents concatenation, and Prediction Score represents the prediction score;

[0090] Step 1), based on the data sources of drugs and target proteins, calculate the similarities between drugs and between target proteins, and construct similarity networks of drugs and target proteins in each data source respectively;

[0091] The construction process of the similarity networks of drugs and target proteins in each data source is as follows:

[0092] Step 1.1: Construction of the drug similarity network;

[0093] Obtain drug data sources, including at least the chemical structures of drugs, drug-drug associations, drug-disease associations, and drug-drug side effect data; for each data source, calculate the similarities between drugs using the Jaccard similarity coefficient, and each data source generates an independent drug similarity network, represented in matrix form, where each element represents the similarity score between a pair of drugs;

[0094] The similarity score is obtained by calculating the ratio of the intersection to the union of two drugs on a specific data source.

[0095] Step 1.2: Construction of the target protein similarity network;

[0096] Obtain target protein data sources, including at least the sequence information of target proteins, target protein-target protein interactions, and target protein-disease associations; for each data source, calculate the similarities between target proteins using the Jaccard similarity coefficient, and each data source generates an independent target protein similarity network;

[0097] The similarity score is obtained by calculating the ratio of the intersection to the union of two target proteins on a specific data source.

[0098] Step 2), the multiple kernel learning algorithm MKL is used to integrate the similarity networks of drugs and target proteins respectively, and the comprehensive similarity matrices SD and SP of drugs and target proteins are obtained.

[0099] During the integration of the similarity network using the multiple kernel learning algorithm MKL, the following objective function is used to minimize the distance between the optimal kernel of the integrated similarity network and the ideal optimal kernel:

[0100] Objective function: ,

[0101] ,

[0102] , ,

[0103] ;

[0104] Among them, represents the similarity network of the i-th one, represents the weight vector of the similarity network of, is the regularization parameter, is the ideal optimal kernel of, β is a regularization parameter used to balance the weight of the similarity matrix fitting error and the weight regularization term; represents a drug or a target protein, represents a drug when, represents a target protein when; represents the drug-target protein association matrix; represents integrating multiple the comprehensive similarity matrix generated after the similarity networks of, * is the ideal optimal kernel matrix of; represents the number of similarity networks; represents the square of the Frobenius norm, represents the difference between the integrated comprehensive similarity matrix SX and the ideal optimal kernel matrix * ; represents the regularization term for the weight vector used to constrain the sparsity of the full-time and avoid overfitting;

[0105] Weight solution formula:

[0106] ,

[0107] ,

[0108] ;

[0109] Wherein: ,

[0110] ,

[0111] ;

[0112] where trace represents the trace of the matrix;

[0113] The conversion formula is:

[0114]

[0115] ;

[0116] represents the drug or target protein, and the weights of each similarity network of the drug and target protein are obtained according to the above formula.

[0117] In order to avoid the occurrence of overfitting during training and improve the generalization ability of the model, making the system performance more robust, an L2 penalty is imposed on the weight vector to limit the growth of its authority.

[0118] The significance of multi-source data fusion is that since the similarity between drugs and target proteins can be described from multiple perspectives (such as chemical structure, sequence information, functional association, etc.), a single data source may not be able to fully reflect their similarity. Therefore, by integrating the similarity networks generated from multiple data sources, the complex relationships between drugs and target proteins can be captured more comprehensively, thereby improving the accuracy of subsequent prediction models.

[0119] The Multiple Kernel Learning (MKL) algorithm is used to integrate multiple similarity networks to generate a comprehensive similarity matrix kernel. Each similarity network can be regarded as a kernel function, and the MKL algorithm reflects its importance in the comprehensive similarity matrix by assigning different weights to each kernel function. By using an optimization algorithm (such as the gradient descent method or quadratic programming) to determine the optimal weights of each kernel function, the comprehensive similarity matrix can better reflect the true relationship between drugs and target proteins. Finally, the MKL algorithm weights and fuses multiple similarity networks to generate a comprehensive similarity matrix kernel for subsequent drug-target protein association prediction models.

[0120] By using the multi-core learning algorithm, the model no longer depends on a single similarity network, but can comprehensively consider the information from multiple data sources. This can not only improve the prediction accuracy of the model, but also enhance the robustness of the model to noise and abnormal data, and avoid prediction errors caused by the deviation of a single data source.

[0121] The comprehensive similarity matrix SD of the drug and the comprehensive similarity matrix SP of the target protein are calculated by the following formula:

[0122] ;

[0123] ;

[0124] where respectively represent the number of drugs and target proteins, N D and N P are respectively the number of similarity networks of drugs and target proteins, and R represents the real number field.

[0125] Step 3), take each row of data in the comprehensive similarity matrices SD and SP obtained in step 2) as the initial feature representation of the drug and the target protein, and introduce noise into the initial feature data through a denoising autoencoder; and restore the initial feature data with noise introduced;

[0126] Restoration is used to remove noise and restore data. The denoised data removes redundant or irrelevant information and retains the core features of the data, making it more suitable for subsequent modeling tasks. Thus, it enhances the model's ability to learn potential patterns in the data and reduces the interference of noise on prediction accuracy.

[0127] The process of taking each row of data in the comprehensive similarity matrices SD and SP obtained in step 2) as the initial feature representation of the drug and the target protein and introducing noise into the feature data through a denoising autoencoder is as follows:

[0128] First, the denoising autoencoder randomly generates noise data through a Gaussian normal distribution and adds the randomly generated noise to the initial input features as the initial data for the dimensionality reduction of the denoising autoencoder, as shown in the following formula:

[0129] ;

[0130] ;

[0131] where represents the mean of the distribution, represents the variance; represents the original input feature data, denote the feature data after adding noise, denote the corresponding random noise;

[0132] Subsequently, use the denoising autoencoder to perform data reconstruction, reduce the dimension of through the reconstructed data dimension, and continuously optimize the loss of the denoising autoencoder by reducing the loss between the reconstructed data and the original input data to achieve the optimal result. The L-layer encoder is defined as follows:

[0133] ,

[0134] ,

[0135] …

[0136] ;

[0137] where denotes the output of the l-th layer of the encoder, denotes the original input data of the first layer of the encoder, and denote the weight matrix and bias matrix of the l-th layer of the encoder respectively, denotes the non-linear activation function;

[0138] The L-layer decoder is defined as follows:

[0139] ,

[0140] ,

[0141] …

[0142] .

[0143] where, denotes the output of the l-th layer of the decoder, denotes the reconstructed data finally output by the decoder, and denote the weight matrix and bias matrix of the l-th layer of the decoder respectively, denotes the non-linear activation function.

[0144] Adopt the Softplus function and RMSProp algorithm to optimize the mean square error formula in the dimension reduction process. The mean square error calculation formula is as follows:

[0145] ;

[0146] where, L denotes the loss function of the denoising autoencoder, represents the reconstructed data of the final output of the decoder, and x represents the original input data of the input denoising autoencoder;

[0147] Reduce the loss of data information during the dimensionality reduction process;

[0148] Step 4), splice the restored feature data of the drug and the target protein processed in step 3) to obtain the feature vector of the drug-target protein pair, and use a multi-layer perceptron to model the association between the feature vectors of the drug-target protein pair, and obtain a drug-target protein association prediction model based on the multi-layer perceptron to predict the association relationship between the drug and the target protein.

[0149] The prediction score of the association relationship between the drug and the target protein is calculated by the following formula:

[0150] ;

[0151] where, is the feature vector of the input drug-target protein pair, , , and , , are the corresponding weight and bias vectors determined in the drug-target protein association prediction model based on the multi-layer perceptron, The function calculation formula is , e represents the natural exponent, the Relu function calculation formula is , is the calculation variable of the input function.

[0152] The parameters in the drug-target protein association prediction model based on the multi-layer perceptron are usually optimized using the backpropagation algorithm and the gradient descent algorithm;

[0153] The drug-target protein association prediction model based on the multi-layer perceptron adopts the Adam algorithm and sets the learning rate to 0.001 to optimize the binary cross loss:

[0154] ;

[0155] where, and respectively represent the positive and negative sample labels of the drug-target protein pair, Y is the true label, with a value of 0 or 1, N represents the number of drug-target protein pairs, represents the prediction score of the association relationship between the drug and the target protein, represents the loss function of the drug-target protein association prediction model based on the multi-layer perceptron.

[0156] Effectiveness Verification of Drug-Target Protein Association Prediction Method Based on Multi-Source Data Embedding

[0157] To verify the effectiveness of the method MKLDAE, the MKLDAE method provided by the technical solution example of the present invention was applied to a set of drug and target protein related data sets. It includes 708 drugs, 1512 proteins, 5603 diseases, 4192 drug side effects, and drug-drug, drug-disease, drug-drug side effect, drug-target protein, target protein-disease, target protein-drug association data. The specific data information is shown in Table 1.

[0158] Table 1 Drug and Protein Association Data Information:

[0159] 。

[0160] Among them, drug, target protein, disease, and drug side effect information were collected from the DrugBank (Version 3.0) database, HPRD (Release 9) database, Comparative Toxicogenomics database, and SIDER (Version 2) database respectively.

[0161] 1. Comparative Experiment to Verify the Effectiveness of the Algorithm:

[0162] To further verify the performance of MKLDAE, a comparative experiment was conducted by comparing the model with four drug-target protein association prediction models in recent years: MolTrans, AEFS, DTINet, and DTICNN. The four comparative models used their respective optimal parameters, and the AUROC and AUPR values of the models were evaluated through ten-fold cross-validation. Considering the differences in data information processing among different models, for fairness, MKLDAE, DTINet, and DTICNN used the same data set, AEFS used the optimal data set in its own experiment, and MolTrans used the BIOSNAP data set. The specific comparison results are shown in Table 2, Figure 2 and 3 are the ROC curves and PR curves corresponding to the five models. The AUROC values corresponding to MKLDAE, DTICNN, DTINet, AEFS, and MolTrans are 0.9528, 0.9390, 0.9147, 0.9081, and 0.8886 respectively, and the AUPR values are 0.9595, 0.9445, 0.9323, 0.3696, and 0.9004 respectively. MKLDAE achieved the best result during the comparative experiment, demonstrating its superiority over other methods in the process of drug-target protein association prediction.

[0163] Among them, MKLDAE is the method proposed in this solution, MolTrans is the method disclosed in the literature "Molecular InteractionTransformer for drug–target interaction prediction", AEFS is the method disclosed in the literature "Autoencoder-based drug–target interaction prediction by preserving the consistency of chemical properties and functions of drugs"; DTINet is the method disclosed in the literature "A network integration approach for drug-target interaction prediction and computational drug repositioning from heterogeneous information"; DTICNN is the method disclosed in the literature "A learning-target based method for drug-interaction prediction based on feature representation learning and deep neural network".

[0164] Table 2 The AUROC and AUPR values corresponding to the five methods in the comparative experiment and the data set sources:

[0165] 。

[0166] 2. Case analysis experiment to verify the effectiveness of the results:

[0167] To further prove the performance of the MKLDAE model in discovering drug-target protein associations, a case analysis experiment was conducted for verification. By deleting the association information between amitriptyline (DB00321: Amitriptyline) and quetiapine (DB01224: Quetiapine) in the drug-target protein association dataset and all target proteins as the test set, and the associations between the remaining drugs and target proteins as the training set to train the model, the accuracy of the top 20 data predicted in the test set was used to evaluate the model. Tables 3 and 4 detail the comparison between the predicted labels and the true labels of amitriptyline and quetiapine with target proteins predicted by the MKLDAE model. It is defaulted that there is an association between the top 20 drugs and target proteins.

[0168] The accuracy rate of the association between amitriptyline and the top 20 target proteins reached 86.7%, and the prediction accuracy rate of quetiapine reached 100%. The case analysis experiment effectively proved that the MKLDAE model has a real effect in predicting the association between drugs and target proteins.

[0169] Table 3 Information on the top 20 associated target proteins predicted for amitriptyline:

[0170] 。

[0171] Table 4 Information on the top 20 associated target proteins predicted for quetiapine:

[0172] 。

[0173] Example 2

[0174] A system for a drug-target protein association prediction method based on the above multi-source data embedding, comprising:

[0175] A similarity network construction module, based on the data sources of drugs and target proteins, calculates the similarities between drugs and between target proteins, and constructs similarity networks of drugs and target proteins in each data source respectively;

[0176] A similarity matrix acquisition module, using the multi-kernel learning algorithm MKL to integrate the similarity networks of drugs and target proteins respectively, and obtaining the comprehensive similarity matrices SD and SP of drugs and target proteins;

[0177] An initial feature reduction module, takes each row of data in the obtained comprehensive similarity matrices SD and SP as the initial feature representations of drugs and target proteins, and introduces noise into the initial feature data through a denoising autoencoder; and reduces the initial feature data with noise introduced;

[0178] An association prediction module, splices the processed and reduced feature data of drugs and target proteins to obtain the feature vectors of drug-protein pairs, and uses a multi-layer perceptron to model the association between the feature vectors of drug-target protein pairs, and obtains a drug-target protein association prediction model based on the multi-layer perceptron to predict the association relationship between drugs and target proteins.

[0179] It should be understood that for the specific implementation processes of each module, please refer to the above method content, which will not be elaborated herein. Moreover, the division of the above functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined, some can be split, and each functional module can be implemented in software, hardware, or a combination of software and hardware. Among them, the software and hardware devices include, but are not limited to, general computer devices, programmable gate arrays, digital signal processors, microprocessors, and their corresponding programming or burning software.

[0180] Embodiment 3

[0181] A computer-readable storage medium stores a computer program, and the computer program is called by a processor to implement:

[0182] The steps of the above drug-target protein association prediction method based on multi-source data embedding.

[0183] For the specific implementation processes of each step, please refer to the elaboration of the foregoing method.

[0184] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the readable storage medium can also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0185] Embodiment 4

[0186] According to an embodiment of the present invention, the present invention further provides an electronic device, including:

[0187] One or more processors;

[0188] A storage device for storing one or more programs,

[0189] When the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0190] In specific use, a user can interact with a server, which is also an electronic device, through an electronic device serving as a terminal device and based on a network to implement functions such as receiving or sending messages. The terminal device is generally various electronic devices equipped with a display device and used based on a human-machine interface, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc. Various specific application software can be installed on the terminal device as needed, including but not limited to web browser software, instant messaging software, social platform software, shopping software, etc.

[0191] The server is a network server for providing various services. The location recognition method provided in this embodiment is generally executed by the server. In actual application, under necessary conditions, the terminal device can also directly execute location recognition. Correspondingly, the location recognition device can be set in the server. Similarly, under necessary conditions, the location recognition device can also be set in the terminal device.

[0192] Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing readable storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The present application is a device that, according to the flowchart of the method, device (system), and computer program product of the embodiments of the present application, and / or the instructions executed by the processor in the block diagram, generates functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more blocks of the block diagram.

[0194] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art according to the technical solutions of the present invention, whether modified or replaced, as long as they do not depart from the spirit and scope of the present invention, also belong to the protection scope of the present invention.

Claims

1. A drug-target protein association prediction method based on multi-source data embedding, characterized in that: The following steps are involved: Step 1), based on the data sources of drugs and target proteins, calculate the similarities between drugs and target proteins, and construct similarity networks of drugs and target proteins in each data source; Step 2), the multi-kernel learning algorithm MKL is used to integrate the similarity networks of drugs and target proteins respectively to obtain the comprehensive similarity matrices SD and SP of drugs and target proteins; Step 3), using each row of data in the comprehensive similarity matrix SD and SP obtained in step 2) as the initial feature representation of the drug and the target protein, and introducing noise into the initial feature data through a denoising autoencoder, and restoring the initial feature data with the introduced noise; , ; in, represents randomly generated noise, represents the mean of the distribution, represents variance; Represents the feature data of the original input, express Feature data after adding noise; The encoder and decoder used for restoration are as follows: ; ; in, and denote the output of the encoder and decoder layer l, respectively. and Represent the weight matrix and bias matrix of the encoder layer l, respectively. represents a nonlinear activation function; Represents the reconstructed data finally output by the decoder, and Respectively represent the weight matrix and bias matrix of the decoder layer l; Step 4), the restored drug and target protein feature data processed in step 3) are spliced ​​to obtain the feature vector of the drug-protein pair, and the association between the feature vectors of the drug-protein pair is modeled using a multi-layer perceptron to obtain a drug-target protein association prediction model based on a multi-layer perceptron to predict the association relationship between the drug and the target protein; The multi-core learning algorithm MKL is used to integrate the similarity networks of drugs and target proteins respectively, and weights are assigned to each similarity network to obtain the comprehensive similarity matrices SD and SP of drugs and target proteins: In the process of integrating the similarity network using the multi-core learning algorithm MKL, the objective function corresponding to the following formula is used to minimize the distance between the optimal kernel of the integrated similarity network and the ideal optimal kernel: Objective function: , , , , ; in, express The similarity network of the i-th, express The weight vector of the similarity network, β is a regularization parameter used to balance the similarity matrix fitting error and the weight of the weight regularization term; Indicates drug or target protein, Indicates drugs, indicates the target protein; It represents the association matrix of drug-target protein; Indicates the integration of multiple The comprehensive similarity matrix generated after the similarity network is * yes The ideal optimal kernel matrix of express The number of similarity networks; represents the square of the Frobenius norm, Represents the integrated comprehensive similarity matrix With the ideal optimal kernel matrix * The difference between Represents the weight vector The regularization term is used to constrain the sparsity of weights and avoid overfitting; Weight solution formula: , , ; in: , , ; Among them, trace represents the trace of the matrix; The conversion formula is: ; According to the above formula, the weights of each similarity network of drugs and target proteins are obtained.

2. The method according to claim 1, characterized in that Prediction score of the association between drugs and target proteins Calculated by the following formula: ; in, is the feature vector of the input drug-target protein pair, , , and , , The corresponding weights and bias vectors determined in the drug-target protein association prediction model based on the multi-layer perceptron, The function calculation formula is: , e represents the natural exponent, and the Relu function calculation formula is , is the calculated variable of the input function.

3. The method according to claim 1, characterized in that The construction process of the similarity network of drugs and target proteins in each data source is as follows: Step 1.1: Construction of drug similarity network; Obtain drug data sources, including at least the chemical structure of the drug, drug-drug association, drug-disease association, and drug-drug side effect data; for each data source, use the Jaccard similarity coefficient to calculate the similarity between drugs, and generate an independent drug similarity network for each data source, which is expressed in a matrix form, in which each element represents a similarity score between a pair of drugs; Step 1.2: Construction of target protein similarity network; Obtain target protein data sources, including at least target protein sequence information, target protein-target protein interactions, and target protein-disease associations; For each data source, the similarity between target proteins was calculated using the Jaccard similarity coefficient, and each data source generated an independent target protein similarity network.

4. The method according to claim 1, characterized in that: The drug comprehensive similarity matrix SD and the target protein comprehensive similarity matrix SP are calculated by the following formula: , ; in, Represent the number of drugs and target proteins, respectively. N D and N P are the number of similarity networks of drugs and target proteins, respectively, and R represents the real number domain.

5. The method according to claim 4, characterized in that The Softplus function and RMSProp algorithm are used to optimize the mean square error formula in the dimensionality reduction process. The mean square error calculation formula is as follows: ; Where L represents the loss function of the denoising autoencoder, Represents the reconstructed data finally output by the decoder, Represents the raw input data fed into the denoising autoencoder.

6. The method according to claim 1, characterized in that The drug-target protein association prediction model based on multi-layer perceptron uses the Adam algorithm and sets the learning rate to 0.001 to optimize the binary crossover loss: ; in, and They represent the positive and negative sample labels of drug-target protein pairs, respectively. Y is the true label, with a value of 0 or 1. N represents the number of drug-target protein pairs. Represents the prediction score of the association between drugs and target proteins, Represents the loss function of the drug-target protein association prediction model based on multi-layer perceptron.

7. A system for predicting drug-target protein association based on the multi-source data embedding method according to any one of claims 1 to 6, characterized in that: include: The similarity network construction module calculates the similarities between drugs and target proteins based on the data sources of drugs and target proteins, and constructs the similarity network of drugs and target proteins in each data source; The similarity matrix acquisition module uses the multi-kernel learning algorithm MKL to integrate the similarity networks of drugs and target proteins respectively to obtain the comprehensive similarity matrices SD and SP of drugs and target proteins; The initial feature restoration module uses each row of data in the obtained comprehensive similarity matrix SD and SP as the initial feature representation of the drug and target protein, and introduces noise into the initial feature data through a denoising autoencoder; And restore the initial feature data with introduced noise; The association prediction module splices the processed and restored feature data of the drug and target protein to obtain the feature vector of the drug-target protein pair, and uses a multi-layer perceptron to model the association between the feature vectors of the drug-target protein pair, obtains a drug and target protein association prediction model based on a multi-layer perceptron, and predicts the association relationship between the drug and the target protein.

8. A computer-readable storage medium, characterized in that: A computer program is stored, which is called by a processor to implement: The steps of the method according to any one of claims 1 to 6.

9. An electronic device, comprising: one or more processors; A storage device for storing one or more programs; It is characterized in that the one or more processors execute the one or more programs to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Drug target interaction prediction method based on multi-channel graph convolutional network

    CN112863693A

  • TransGAT-based drug-target interaction prediction method

    CN116312808A