Drug-target binding prediction method based on variational coding

By modeling the characteristics of drugs and targets in the potential semantic space based on variational encoding, the high cost and performance attenuation of drug target binding prediction in the prior art are solved, and the prediction effect of high accuracy and robustness is achieved.

CN119580827BActive Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510139658.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-06
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The prior art has high costs, high failure rates, long research cycles, and high requirements for experimental conditions and sample purity in drug target binding prediction, and machine learning-based models attenuate performance when processing large-scale data sets.

Method used

Using a variational encoding method, by collecting data from the drug library and target database, the characteristic representation of the drug and target is constructed, and modeled in the latent semantic space, the drug-target interaction characteristics are reconstructed through up-down sampling paths, the initial prediction model is constructed and combined loss function training is performed.

Benefits of technology

It improves the accuracy and robustness of drug target interaction prediction, can effectively integrate multiple data sources and feature information, accurately model complex nonlinear relationships, and reduces the requirements for experimental conditions and sample purity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580827B_ABST
    Figure CN119580827B_ABST
Patent Text Reader

Abstract

The present invention discloses a drug-target binding prediction method based on variational coding, which belongs to the field of bioinformatics. The present invention realizes efficient extraction of drug and target features, construction of potential feature space of drug and target features, and prediction of their interaction relationship through a two-stage coupled end-to-end neural network. The model is mainly composed of two stages. In the first stage, the potential representation of drug and target features is obtained through variational compression and feature extraction. In the second stage, the drug-target interaction matrix is ​​reconstructed through a deep neural network with up and down sampling. Through this architecture, the chemical characteristics of the drug and the biological characteristics of the target can be effectively integrated to improve the accuracy of interactive relationship prediction. The present invention also provides a novel and efficient solution for the in-depth analysis of biological big data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of bioinformatics, and in particular relates to a drug target binding prediction method based on variational coding. Background Art

[0002] In the modern drug development process, the study of predicting the binding of drugs to targets plays a vital role. Accurate binding analysis can not only deeply understand the function and mechanism of drugs, but also be used to explore their binding pathways, expand the potential application areas of drugs, and even support drug reuse, that is, identifying new indications for existing drugs. Traditional drug-target binding studies mainly rely on experimental methods, but such methods usually face high costs, high failure rates, and long research cycles. In addition, some experiments may not be implemented for ethical reasons. In contrast, artificial intelligence methods provide a new approach for drug-target binding prediction. Its advantage is that it can integrate and analyze a large amount of biomolecular information and learn key feature expressions from it. Through technologies such as deep learning, artificial intelligence methods can quickly and efficiently predict potential drug-target pairs, thereby making up for the shortcomings of traditional experimental methods and providing strong technical support and guarantee for the drug development process.

[0003] Prediction of drug-target interaction (DTI) is one of the key steps in drug development. In existing studies, biological wet experimental methods are widely used for preliminary screening and verification of DTI. These methods include high-throughput screening (HTS), affinity mass spectrometry (AP-MS), and yeast two-hybrid experiments, which can provide direct experimental verification. However, these methods are usually time-consuming and labor-intensive, and require a lot of resource investment. In addition, biological wet experiments have high requirements for sample purity and experimental conditions, and are limited by experimental techniques and equipment. Therefore, computational analysis to assist drug repurposing has received widespread attention. Prediction models for drug-target binding are generally divided into machine learning (ML)-based methods and deep learning (DL)-based strategies. Machine learning-based models widely use technologies such as matrix decomposition, random forests, and support vector machines to infer drug-target binding by mining known data. These methods can achieve certain prediction results when dealing with small-scale data sets or simple drug-target relationships. However, in the face of large-scale data sets and complex biomolecular features, the performance of these models is severely attenuated, and it is often difficult to accurately summarize and characterize the complex properties of biomolecules.

[0004] With the accelerated development of deep learning, a variety of deep learning-based models have emerged in recent years, which can be roughly classified into models based on attention mechanisms, multimodal models, hybrid network models, and models based on network representation learning. Models based on attention mechanisms use self-attention and cross-attention layers to capture the complex background information between drugs and targets. The advantage of such models is that they can effectively extract important relationships between features and improve prediction accuracy. However, these models usually require a large amount of training data and have high computational complexity, making the training process time-consuming. Multimodal models improve drug-target binding prediction by integrating information from drug-drug and target-target interaction networks. The advantage of these models is that they can use multiple similarity measures to enrich feature representations and improve the generalization ability of model predictions. However, they also have high requirements for data quality and diversity, and the feature engineering process is complex, and it may be difficult to determine the best similarity measure. Hybrid network models combine convolutional neural networks and long short-term memory networks to achieve a balance between sequence and local feature learning. Such models improve expressive power by combining feature extraction methods, but may face high complexity in implementation, and precise adjustment of model details such as hyperparameters is crucial to achieve the desired effect. Models based on network representation learning use heterogeneous data sources to learn low-dimensional feature representations. These models are generally able to capture complex biological relationships well, however, they are sensitive to the way the input network is constructed and require efficient feature integration and pattern recognition techniques. Their effectiveness may be limited when heterogeneous data sources are not available. Although these models have shown significant potential in improving the accuracy and robustness of drug-target binding predictions, challenges such as data dependence, complexity, and computational resource requirements still need to be overcome. Summary of the invention

[0005] The purpose of the present invention is to propose a drug target binding prediction method based on variational coding in response to the problems of the prior art.

[0006] To achieve the above object, the present invention provides a method for predicting drug target binding based on variational coding, comprising the following steps:

[0007] Collect drug fingerprints and target descriptors from existing drug libraries and target databases, calculate drug fingerprint similarity and target similarity, and construct drug features and target features as data sets respectively;

[0008] Based on variational embedding, the dataset is modeled in a joint latent semantic space of drugs and targets to obtain low-dimensional features of drugs and targets in the latent space;

[0009] splicing low-dimensional features of drugs and targets in latent space as joint latent representation, extracting the joint latent representation, reconstructing drug-target interaction features through up- and down-sampling paths, and building an initial prediction model based on the drug-target interaction features;

[0010] The initial prediction model is trained based on the joint loss function to obtain a prediction model, and the probability of interaction between the drug and the target is predicted by the prediction model.

[0011] Furthermore, the process of constructing drug features and target features respectively as data sets includes:

[0012] Collect drug molecular fingerprints, construct fingerprint vectors through Morgan fingerprints, evaluate the similarity data between different drug molecular fingerprints based on Tanimoto similarity, and construct drug similarity vectors; collect target descriptors, construct target molecular descriptor feature vectors, use cosine similarity to evaluate the similarity data between different targets, and construct target similarity vectors; concatenate fingerprint vectors and drug similarity vectors, target molecular descriptor feature vectors and target similarity vectors, respectively, clean and standardize them, and obtain drug features and target features as data sets.

[0013] Furthermore, the expression for evaluating the similarity between different drug molecular fingerprints is:

[0014] ;

[0015] in, Represents drug molecule and drug molecules The closer the similarity is to 1, the more similar the two molecular fingerprints are; and Represent drug molecules and drug molecules The fingerprint vector, and Represents molecules and molecules The square norm of the fingerprint vector.

[0016] Furthermore, the expression for evaluating the similarity between different targets is:

[0017] ;

[0018] in, Indicates target and target The cosine similarity of and Represent the target molecule descriptor feature vectors and The second norm of .

[0019] Furthermore, a method for modeling the dataset in a potential semantic space of a drug-target combination to obtain low-dimensional features of the drug and the target in the potential space includes:

[0020] Encode the drug features and target features in the data set, and obtain the mean and variance of the drug features and target features mapped to the latent space respectively;

[0021] Sampling using a reparameterization technique based on the mean and variance to generate a drug latent feature vector and a target latent feature vector;

[0022] The drug latent feature vector and the target latent feature vector are respectively input into multiple fully connected layers, each containing a nonlinear activation function, for processing to obtain low-dimensional features of the drug and target in the latent space.

[0023] Furthermore, the low-dimensional features of the drug and target in the latent space are expressed as:

[0024] ;

[0025] ;

[0026] in, and are the weight matrices of drugs and targets, respectively, and are the bias vectors for drug and target, respectively; and are the drug latent feature vector and target latent feature vector, respectively, using ReLu as the activation function.

[0027] Furthermore, the method for reconstructing drug-target interaction characteristics through up- and down-sampling paths includes:

[0028] The joint latent representation is downsampled to extract key features, which are then restored to the original size of the feature map layer by layer through deconvolution operations, and the feature map is output through the encoding-decoding model.

[0029] Furthermore, the expression of the joint loss function is:

[0030] ;

[0031] in, The expression is:

[0032] ;

[0033] The expression is:

[0034] ;

[0035] The expression is:

[0036] ;

[0037] The expression is:

[0038] ;

[0039] In the formula, represents the upsampled output, represents the joint latent representation, and They represent the mean and variance of drug features mapped to the latent space, and They represent the mean and variance of the target feature mapped to the latent space, represents the true adjacency matrix, represents the reconstructed adjacency matrix, which is obtained by normalizing the upsampled output through the activation function; and Represent drug molecules and target , Represents the number of effective drug-target pairs.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention innovatively proposes a new method for characterizing biological information of drug targets and predicting binding properties. First, high-dimensional vector representations of drugs and targets are extracted through feature engineering, which lays the foundation for subsequent calculations. In the first stage, a variational autoencoder is used to compress features to capture key information in the latent space and achieve effective information compression and expression. In the second stage, an initial feature map is generated through a specific sampling strategy, and multiple feature channels are combined to enhance the diversity of information, organically combine downsampling and upsampling techniques, and combine multi-scale information through a feature expansion algorithm to effectively reconstruct and interpret the interaction relationship in a complex biological network. Finally, the two stages of the model realize feature sharing and joint training, so that the biomolecular information can be effectively compressed, while improving the model's ability to integrate network information. The entire process of the model emphasizes the effective integration of multiple data sources and feature information, and realizes accurate modeling of complex nonlinear relationships through the innovative architecture of deep learning. The present invention not only improves the accuracy of drug-target interaction prediction through the combination of multi-stage processing and advanced deep learning methods, but also provides a novel and efficient solution for the in-depth analysis of biological big data. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of feature extraction of the chemical structure of the drug and target molecule according to an embodiment of the present invention; wherein: Figure 1 (a) is a schematic diagram of the process of extracting drug features; Figure 1 (b) is a schematic diagram of the process of extracting target features; Figure 1 (c) is a schematic diagram of the process of calculating drug or target similarity;

[0043] Figure 2 The figure is a flow chart of a method for predicting drug-target binding according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0045] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0046] Example 1: Figure 1 and Figure 2 As shown, this embodiment provides a drug target binding prediction method based on variational coding.

[0047] 1. The process of constructing features based on biomolecular ontology.

[0048] The first step is to construct the drug molecular fingerprint.

[0049] The present invention obtains the structural data of drug molecules from a variety of chemical information databases. The chemical characteristics of drugs are realized by calculating their Morgan fingerprints. This fingerprint is generated based on the topological information of the molecular structure to capture the chemical characteristics of the molecule. Given a set of atoms V and a corresponding bond E of a molecule, the Morgan fingerprint can be generated by iterating adjacent atoms, and the fingerprint formula can be expressed as:

[0050] ;

[0051] Among them, R is the maximum radius of the fingerprint, and the internal Neighborhood function is used to extract the neighboring atoms of molecule m within a distance d and their connection methods.

[0052] The topological features and connection relationships of these neighborhood atoms are mapped to specific unique identifiers through hash functions. If the feature has appeared, it is set to true, otherwise it is set to zero for encoding. After generating the fingerprint, the fingerprint can be represented as a binary vector , each element It can be expressed by the following formula:

[0053] ;

[0054] in, It refers to the hash function for different feature bits, which is used to map the extracted features to a binary vector of fixed size. Represents the fingerprint vector The final length is The eigenvector of , as a fingerprint vector, can be expressed as:

[0055] .

[0056] The second step is the similarity calculation based on drug fingerprints.

[0057] After obtaining the drug fingerprint features, in order to further explore its potential feature relationship, the similarity between drug features is calculated. The similarity between molecular fingerprints is evaluated by calculating the Tanimoto similarity. The specific formula is as follows:

[0058] ;

[0059] in, Represents molecules and molecules The closer the similarity is to 1, the more similar the two molecular fingerprints are. and Respectively represent molecules and molecules The fingerprint vector, and Represents molecules and molecules The square norm of the fingerprint vector.

[0060] The third step is to construct drug characteristics.

[0061] After completing the generation of drug fingerprints and similarity calculation, in order to fully characterize the molecular characteristics of the drug, the present invention splices the fingerprint features generated in the first step with the similarity features calculated in the second step to construct a comprehensive drug feature vector:

[0062] , ;

[0063] in, Indicates drug The final comprehensive feature vector of Indicates drug The fingerprint vector of length is , Indicates drug Similarity vector with other drugs (drug similarity vector), length is .

[0064] The fourth step is target feature construction based on molecular descriptors.

[0065] In order to accurately describe the structure and functional characteristics of the target, the target features are realized by extracting and encoding molecular descriptors. , whose characteristics are described by its sequence information, structural information and functional information. The extracted features can be expressed by the following formula:

[0066] , ;

[0067] in, Indicates target The molecular descriptor feature vector of is a feature extraction function that generates features based on the properties of the target, including amino acid composition, secondary structure distribution, and functional domain information. Its length is .

[0068] Step 5: Calculate the similarity of target features.

[0069] In order to further explore the potential relationship between targets, the similarity of targets is calculated based on their descriptors. Weigh the characteristic vectors of the two target molecular descriptors and The similarity is calculated as follows:

[0070] , ;

[0071] in, and Represent the target molecule descriptor feature vectors and The second norm of .

[0072] Step 6: Construct target features.

[0073] After the target molecular descriptor extraction and similarity calculation are completed, in order to comprehensively construct the feature vector of the target, the feature vector based on the descriptor is Similarity-based features Perform stitching to generate comprehensive target signatures The final feature representation of the target is as follows:

[0074] , ;

[0075] in, Indicates target The final comprehensive feature vector of Indicates target Molecular descriptor feature vector, length , Indicates target Similarity vector with other targets (target similarity vector), length is .

[0076] Step 7: Data preprocessing of drug and target characteristics.

[0077] The molecular feature data collected from the database undergoes cleaning and standardization steps. During the cleaning step, missing values ​​and incomplete records will be removed to ensure data integrity. Standardization is applied to the retained data and processed using the z-score (standard score), the specific formula is as follows:

[0078] ;

[0079] in, is the original characteristic value of the drug or target, is the mean of the feature, is the standard deviation of the feature, It is the characteristic of the drug or target after processing. This process ensures that the characteristic value is within a uniform numerical range. After the above steps, the drug characteristic is obtained. Target characteristics .

[0080] 2. Feature encoding and latent space modeling based on drug-target molecule descriptors.

[0081] The first step is drug and target feature encoding.

[0082] The present invention is based on the characteristics of the drug after pretreatment and target characteristics As input, they are processed through independent encoder networks. , the encoder network maps it to the mean of the latent space and variance Similarly, target features Mapped to and , the specific formula is:

[0083] ;

[0084] ;

[0085] in Represents the encoder network. These means and variances not only summarize the statistics of the input, but also form the basis for the subsequent latent space representation.

[0086] The second step is latent feature vector sampling.

[0087] The mean and variance obtained in the encoding stage are further used to sample the latent feature vector. In order to comply with the requirements of differentiable programming, the present invention applies the reparameterization technique, which can maintain the stability of gradient propagation during the training process. Specifically, starting from the mean and variance, the latent feature vector is generated. and :

[0088] ;

[0089] ;

[0090] in, and are noise vectors sampled from standard normal distribution, represents element-wise multiplication. Such random sampling construction is intended to reflect the distributional assumptions that hold for unknown inputs and helps reduce the risk of overfitting.

[0091] The third step is latent space modeling.

[0092] The drug potential feature vector obtained after the encoding stage and the target latent feature vector It is input into multiple fully connected layers, and nonlinear activation functions are added to each layer to improve the complexity and flexibility of feature expression. Rectified linear unit (ReLu) is selected as the activation function to obtain a modal low-dimensional potential representation of the drug. and the target low-dimensional latent representation :

[0093] ;

[0094] ;

[0095] in, and are the weight matrices of drugs and targets, respectively, and are the bias vectors for drug and target, respectively.

[0096] 3. Adjacency matrix prediction and up- and down-sampling paths.

[0097] The first step is the integration of potential drug target features.

[0098] The low-dimensional latent representation of drugs and the target low-dimensional latent representation Fusion to form a joint latent representation The joint representation is achieved through the concatenation operation of the vectors, and the formula is:

[0099] ;

[0100] By splicing the potential features of drugs and targets, the characteristic information of both is preserved.

[0101] The second step is to build the encoding-decoding model.

[0102] The construction of the encoding-decoding model adopts an up-down sampling strategy based on the U-Net structure to extract the interaction features between drugs and targets in detail. In this structure, the joint potential representation is first A series of convolution and pooling operations are performed to downsample the features layer by layer, thereby extracting their key features. In the downsampling path, for each convolution layer, the output after activation is defined as:

[0103] ;

[0104] in, and It is Layer and The feature map of the layer, is the convolution kernel, represents the convolution operation, is the bias vector. After the downsampling step, the joint potential representation is obtained In the downsampling path, the size of the feature map is reduced by the maximum pooling operation. In the upsampling path, this process restores the original size of the feature map layer by layer through deconvolution operations. In each upsampling layer, the feature map of the corresponding layer of the original downsampling path is combined for skip connection to retain high-resolution features. The generated upsampled output for:

[0105] ;

[0106] in, Represents the number of layers in the downsampling process. This operation ensures that high-level semantic information and low-level detail features are integrated through jump connections, which comprehensively improves the model's expressiveness. Finally, the feature map that is further processed by the encoder-decoder model output generates an accurate drug-target interaction prediction matrix. This matrix is ​​predicted by combining the output of the last layer of the convolutional layer and applying the Sigmoid activation function to standardize the result:

[0107] ;

[0108] in, This step provides quantitative results for the final prediction of the relationship between drugs and targets for the reconstructed adjacency matrix, which can effectively convey the potential interaction information.

[0109] 4. Joint optimization training.

[0110] The first step is to design the model loss function.

[0111] The joint optimization training step of the present invention designs a comprehensive loss function and optimization algorithm to perform end-to-end model parameter updates, thereby improving the accuracy and robustness of drug-target interaction prediction. The loss function of the present invention comprehensively considers various important factors to ensure the training effect of the model. Reconstruction loss during upsampling and downsampling for:

[0112] ;

[0113] The KL divergence is used to measure the difference between the potential distribution of the encoder output and the standard normal distribution. The KL divergence of the joint representation of drug and target features is:

[0114] ;

[0115] ;

[0116] Using binary cross entropy The loss function measures the difference between the predicted adjacency matrix and the true interaction relationship:

[0117] ;

[0118] in, represents the true adjacency matrix, and Respectively represent drugs and target , Represents the number of effective drug-target pairs.

[0119] Combining the above loss terms, the total loss function is constructed:

[0120] ;

[0121] In order to minimize the total loss function, the Adam optimizer (adaptive moment estimation optimizer) is used to learn and update the \(\theta\) model parameters. The specific parameter update is given by the following formula:

[0122] ;

[0123] in, represents the model parameters, is the learning rate, are the parameters of the updated model.

[0124] The second step is to determine the evaluation indicators of the prediction model.

[0125] In the present invention, in order to comprehensively evaluate the performance of the model in the drug-target interaction prediction task, the present invention adopts a variety of performance evaluation indicators. As a key indicator, the F1 score comprehensively considers the precision (Precision) and recall (Recall) of the model, and provides a balanced representation of the classifier when processing unbalanced data sets. In addition, the accuracy (Accuracy, ACC) is used to measure the overall predictive ability of the model. Sensitivity (Sensitivity, Sen) reflects the proportion of correct identification among all actual positive examples. The specificity (Spe) indicates the model's ability to recognize negative examples, which corresponds to the sensitivity. Finally, the precision (Precision, Pre) measures the proportion of samples that are actually positive examples among those predicted by the model as positive examples, thereby expressing the accuracy of the model.

[0126] The third step is the division and selection of experimental data sets.

[0127] In order to verify the universality and robustness of the model, the present invention conducted experiments on three data sets with different ratios of positive to negative samples. In these data sets, positive samples correspond to all known drug-target associations, and negative samples are generated by randomly combining drugs and targets. The ratios of positive to negative samples in the data sets were set to 1:1, 1:10, and 1:100, respectively. 10-fold cross validation was used to evaluate the performance of the model on these data sets. In each 10-fold cross validation, the data was randomly divided into 90% training set and 10% test set in proportion. In order to reduce the fluctuation of results caused by data deviation, the process was repeated 10 times, and the average results were taken to obtain a stable evaluation.

[0128] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A drug target binding prediction method based on variational coding, characterized in that: The following steps are involved: Collect drug fingerprints and target descriptors from existing drug libraries and target databases, calculate drug fingerprint similarity and target similarity, and construct drug features and target features as data sets, specifically: collect drug molecular fingerprints, construct fingerprint vectors through Morgan fingerprints, evaluate the similarity data between different drug molecular fingerprints based on Tanimoto similarity, and construct drug similarity vectors; collect target descriptors, construct target molecular descriptor feature vectors, use cosine similarity to evaluate the similarity data between different targets, and construct target similarity vectors; respectively splice fingerprint vectors and drug similarity vectors, target molecular descriptor feature vectors and target similarity vectors, clean and standardize them, and obtain drug features and target features as data sets; Based on variational embedding, the dataset is modeled in a joint latent semantic space of drugs and targets to obtain low-dimensional features of drugs and targets in the latent space; splicing low-dimensional features of drugs and targets in latent space as joint latent representation, extracting the joint latent representation, reconstructing drug-target interaction features through up- and down-sampling paths, and building an initial prediction model based on the drug-target interaction features; The initial prediction model is trained based on the joint loss function to obtain a prediction model, and the probability of interaction between the drug and the target is predicted by the prediction model.

2. The drug target binding prediction method based on variational coding according to claim 1, characterized in that: The expression for evaluating the similarity between different drug molecular fingerprints is: ; in, Represents drug molecule and drug molecules The closer the similarity is to 1, the more similar the two molecular fingerprints are; and Represent drug molecules and drug molecules The fingerprint vector, and Represents molecules and molecules The square norm of the fingerprint vector.

3. The drug target binding prediction method based on variational coding according to claim 1, characterized in that: The expression for evaluating the similarity between different targets is: ; in, Indicates target and target The cosine similarity of and Represent the target molecule descriptor feature vectors and The second norm of .

4. The drug target binding prediction method based on variational coding according to claim 1, characterized in that: The method of modeling the dataset in the drug-target joint latent semantic space and obtaining low-dimensional features of the drug and the target in the latent space includes: Encode the drug features and target features in the data set, and obtain the mean and variance of the drug features and target features mapped to the latent space respectively; Sampling using a reparameterization technique based on the mean and variance to generate a drug latent feature vector and a target latent feature vector; The drug latent feature vector and the target latent feature vector are respectively input into multiple fully connected layers, each of which contains a nonlinear activation function for processing to obtain the low-dimensional features of the drug and target in the latent space.

5. The drug target binding prediction method based on variational coding according to claim 4, characterized in that: The low-dimensional features of the drug and target in the latent space are expressed as: ; ; in, and are the weight matrices of drugs and targets, respectively, and are the bias vectors for drug and target, respectively; and are the drug latent feature vector and target latent feature vector, respectively, using ReLu as the activation function.

6. The drug target binding prediction method based on variational coding according to claim 1, characterized in that: The method for reconstructing drug-target interaction characteristics through up- and down-sampling paths comprises: The joint latent representation is downsampled to extract key features, which are then restored to the original size of the feature map layer by layer through deconvolution operations, and the feature map is output through the encoding-decoding model.

7. The drug target binding prediction method based on variational coding according to claim 1, characterized in that: The expression of the joint loss function is: ; in, The expression is: ; The expression is: ; The expression is: ; The expression is: ; In the formula, represents the upsampled output, represents the joint latent representation, and They represent the mean and variance of drug features mapped to the latent space, and They represent the mean and variance of the target feature mapped to the latent space, represents the true adjacency matrix, represents the reconstructed adjacency matrix, which is obtained by normalizing the upsampled output through the activation function; and Represent drug molecules and target , Represents the number of effective drug-target pairs.

Citation Information

Patent Citations

  • Drug-target interaction prediction method based on heterograph contrast learning

    CN117594117A

  • Predicting drug-drug interactions based on clinical side effects

    US20150324693A1