A drug-target interaction prediction method and device

The interaction fingerprint information is extracted through the molecular graph model and the protein sequence language model, and the degree of drug-target interaction is predicted by combining the interaction type, which solves the problem of inaccurate prediction in the prior art and achieves more accurate drug-target interaction prediction.

CN114724648BActive Publication Date: 2025-08-19PING AN TECH (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210499668.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-08-19
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The existing drug-target interaction prediction methods are not accurate enough, resulting in the inability to effectively identify new candidate compounds with potential therapeutic effects during drug discovery.

Method used

The first sub-model of the target is used to extract the interaction fingerprint information, the second sub-model of the target is used to predict the degree of interaction, and the feature vector is extracted through the molecular graph model and the protein sequence language model, and accurately predicted with the interaction type.

Benefits of technology

Improve the accuracy and reliability of drug-target interaction prediction, ensuring the reliability and accuracy of the final predicted results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724648B_ABST
    Figure CN114724648B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for predicting drug-target interactions, wherein the method comprises: extracting an interaction fingerprint based on a test drug and a test target using a target first sub-model in a target prediction model to obtain target interaction fingerprint information between the test drug and the test target; and predicting the degree of interaction based on the type of interaction in the interaction fingerprint information using a target second sub-model in the target prediction model to obtain a predicted result of the interaction between the test drug and the test target. The present application first accurately extracts the interaction fingerprint based on the target first sub-model, and then predicts the interaction based on the interaction fingerprint information based on the target second sub-model, thereby making the final prediction result more accurate and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug action prediction, and in particular to a drug-target interaction prediction method and device. Background Art

[0002] Drug action prediction is the process of identifying new candidate compounds with potential therapeutic effects for specific diseases. Predicting drug-target interactions is an essential step in the drug discovery process. The efficacy of a drug molecule depends on its affinity for the target protein or receptor. Drug molecules without any interaction or affinity for the target protein will not provide a therapeutic effect.

[0003] Since the experimental determination of drug-target interactions (DTI) between drug molecules and target proteins is both time-consuming and resource-intensive, it is necessary to develop efficient computational simulation methods. By utilizing known measured drug-target protein interaction data, machine learning methods can be used to train models and learn from such heterogeneous biological data to understand the mechanism of action of drugs in the human body.

[0004] Current interaction prediction methods typically require acquiring a large number of drug molecule and target protein features. These features are then simply spliced together and fed into a model to fit the predicted interaction signature. However, this approach requires a large amount of data and is ineffective, often failing to capture true "interaction information."

[0005] Therefore, a drug-target interaction prediction method is urgently needed to solve the problem of inaccurate interaction prediction in existing interaction prediction methods. Summary of the Invention

[0006] In view of this, the present invention provides a method and device for predicting drug-target interactions, the main purpose of which is to solve the problem of inaccurate interaction prediction in the prior art.

[0007] To solve the above problems, the present application provides a method for predicting drug-target interactions, comprising:

[0008] Based on the drug to be tested and the target to be tested, using the target first sub-model in the target prediction model to extract the interaction fingerprint to obtain the target interaction fingerprint information between the drug to be tested and the target to be tested;

[0009] Based on the interaction type in the interaction fingerprint information, the target second sub-model in the target prediction model is used to predict the degree of interaction to obtain a prediction result of the interaction between the drug to be tested and the target to be tested.

[0010] Optionally, based on the drug to be tested and the target to be tested, using the target first sub-model in the target prediction model to extract the interaction fingerprint to obtain the target interaction fingerprint information between the drug to be tested and the target to be tested, specifically includes:

[0011] Extracting a feature vector of the drug to be tested using the molecular graph model in the target first sub-model to obtain a first feature vector;

[0012] Using the protein sequence language model in the first sub-model of the target, extracting a feature vector of the target to be tested to obtain a second feature vector;

[0013] The target interaction fingerprint information is extracted based on the first feature vector and the second feature vector.

[0014] Optionally, before extracting a feature vector of the drug to be tested using the molecular graph model in the target first sub-model to obtain a first feature vector, the method further includes:

[0015] Preprocessing the drug to be tested to obtain graph data of the drug to be tested based on chemical bond connection information for extracting a first eigenvector;

[0016] The target to be detected is preprocessed to obtain protein sequence data of the target to be detected for extracting a second feature vector.

[0017] Optionally, the step of predicting the degree of interaction based on the interaction type in the target interaction fingerprint information and using the target second sub-model in the target prediction model to obtain a prediction result of the interaction between the test drug and the test target specifically includes:

[0018] Calculating an interaction degree value based on the target second sub-model and the interaction type in the target interaction fingerprint information;

[0019] Obtaining a prediction result of the interaction between the drug to be tested and the target to be tested based on the interaction degree value;

[0020] Among them, the interaction types include any one or more of the following: non-polar interactions, face-to-face π-π interactions, edge-to-edge π-π interactions, hydrogen bonds where the residues are donors, hydrogen bonds where the residues are acceptors, electrostatic interactions where the residues are positively charged, and electrostatic interactions where the residues are negatively charged.

[0021] Optionally, before extracting the interaction fingerprint based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model, the method further includes: performing model pre-training on the initial first sub-model in the initial prediction model to obtain a pre-trained first sub-model, wherein the specific training process includes:

[0022] Preprocessing the sample drugs to obtain graph data of each sample drug based on chemical bond connection information;

[0023] Preprocess the sample targets to obtain protein sequence data of each sample target;

[0024] Based on the sample graph data, the protein sequence data of the sample targets and the interaction fingerprint labels, the initial first sub-model is pre-trained using a back-propagation model training method to obtain the pre-trained first sub-model.

[0025] Optionally, before extracting the interaction fingerprint based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model, the method further includes:

[0026] Based on the target sample drug, the target target, and the interaction information between the target sample drug and the target target, the pre-trained first sub-model and the initial second sub-model in the initial prediction model are trained to obtain the target first sub-model and the target second sub-model to obtain the target prediction model.

[0027] Optionally, the first eigenvector includes any one or more of the following: an atomic node initial eigenvector and an atomic edge initial eigenvector.

[0028] To solve the above problems, the present application provides a device comprising:

[0029] An extraction module is used to extract interaction fingerprints based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model to obtain target interaction fingerprint information between the drug to be tested and the target to be tested;

[0030] The prediction module is used to predict the interaction size based on the mutual interaction type in the interaction fingerprint information and use the target second sub-model in the target prediction model to obtain the prediction result of the interaction between the drug to be tested and the target to be tested.

[0031] To solve the above problems, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned drug-target interaction prediction methods.

[0032] To solve the above problems, the present application provides an electronic device, which includes at least a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the computer program on the memory, it implements the steps of any of the above-mentioned drug-target interaction prediction methods.

[0033] The drug-target interaction prediction method and device in the present application can accurately extract the interaction fingerprint based on the first target sub-model, and then use the second target sub-model to accurately predict the degree of interaction based on the extracted interaction fingerprint information. That is, the second target sub-model is used to accurately predict the degree of interaction based on the interaction type in the interaction fingerprint information, and then it can be accurately determined whether the drug to be tested and the target to be tested have an interaction based on the degree of interaction, which can make the final prediction result more accurate and reliable.

[0034] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0036] Figure 1 This is a flow chart of a method for predicting drug-target interactions according to an embodiment of the present application;

[0037] Figure 2 This is a structural block diagram of the drug-target interaction prediction device in the embodiment of the present application;

[0038] Figure 3 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0039] Various aspects and features of the present application are described herein with reference to the accompanying drawings.

[0040] It should be understood that various modifications may be made to the embodiments of the present application. Therefore, the above description should not be considered as limiting, but merely as an example of an embodiment. Other modifications within the scope and spirit of the present application will occur to those skilled in the art.

[0041] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0042] These and other characteristics of the present application will become apparent from the following description of a preferred form of embodiment given as a non-limiting example with reference to the accompanying drawings.

[0043] It should also be understood that although the present application has been described with reference to certain specific examples, those skilled in the art will readily be able to implement many other equivalent forms of the present application.

[0044] The above and other aspects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings.

[0045] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments described are merely examples of the present application and may be implemented in a variety of ways. Familiar and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details described herein are not intended to be limiting, but rather serve merely as a basis and representative basis for the claims to teach those skilled in the art to variously utilize the present application with substantially any suitable detailed structure.

[0046] This specification may use the phrases "in one embodiment," "in another embodiment," "in yet another embodiment," or "in other embodiments," which may all refer to one or more of the same or different embodiments according to the present application.

[0047] The present invention provides a method for predicting drug-target interactions. Figure 1 As shown, the following steps are included:

[0048] Step S101, based on the drug to be tested and the target to be tested, using the target first sub-model in the target prediction model to extract the interaction fingerprint, and obtain the target interaction fingerprint information between the drug to be tested and the target to be tested;

[0049] In this example, interaction fingerprinting (IFP) provides an alternative method for characterizing drug-target interaction patterns. IFP can reveal the details and specificity of intermolecular interaction forces. IFP is a method that converts three-dimensional (3D) protein-ligand interactions into one-dimensional (1D) bit strings. It can be used to compare differences in the interaction forces between ligands and proteins.

[0050] In this step, the target first sub-model can specifically include a molecular graph model (GNN) and a protein sequence language model (Transformer). During the specific implementation of this step, the drug to be tested can be preprocessed to obtain graph data of the drug to be tested based on chemical bond connectivity information; the target to be tested can also be preprocessed to obtain protein sequence data of the target to be tested; the graph data of the drug to be tested is then input into the molecular graph model (GNN), and the protein sequence data of the target to be tested is input into the protein sequence language model (Transformer), thereby extracting the interaction fingerprint information.

[0051] Step S102 : Based on the interaction type in the interaction fingerprint information, the target second sub-model in the target prediction model is used to predict the degree of interaction, and obtain a prediction result of the interaction between the drug to be tested and the target to be tested.

[0052] During the specific implementation of this step, when predicting the degree of interaction, the target second sub-model can be used to obtain the interaction type from the interaction fingerprint information, and finally, based on the interaction type, the interaction degree value is calculated. Finally, based on the interaction degree value, it is determined whether there is an interaction between the drug to be tested and the target to be tested, and the prediction result of the interaction is obtained.

[0053] The drug-target interaction prediction method in the present application can accurately extract the interaction fingerprint based on the first target sub-model, and then use the second target sub-model to accurately predict the degree of interaction based on the extracted interaction fingerprint information. That is, the second target sub-model is used to accurately predict the degree of interaction based on the interaction type in the interaction fingerprint information, and then it can be accurately determined whether the drug to be tested and the target to be tested have an interaction based on the degree of interaction, which can make the final prediction result more accurate and reliable.

[0054] Based on the above embodiments, another embodiment of the present application provides a method for predicting drug-target interactions, comprising the following steps:

[0055] Step S201: Preprocessing the sample drugs to obtain graph data of each sample drug based on chemical bond connection information; preprocessing the sample targets to obtain protein sequence data of each sample target; and pretraining the initial first sub-model using a back-propagation model training method based on each sample graph data, each protein sequence data of each sample target, and interaction fingerprint labels to obtain the pre-trained first sub-model.

[0056] In the specific implementation process of this step, the initial first sub-model can be pre-trained through large-scale unlabeled data, that is, each sample drug is processed into each sample drug graph data with atoms as nodes and connected by chemical bonds using RDKit, that is, graph data containing chemical bond connection information is obtained. The sample targets are pre-processed to obtain the protein sequence data of each sample target. Then, the graph data of each sample drug and the protein sequence data of each sample target are respectively passed through the molecular graph model GNN in the initial first sub-model and the protein sequence language model Transformer in the initial first sub-model to output their respective latent vectors (both are 768 dimensions) to obtain interaction fingerprints. The specific process of obtaining the interaction fingerprint is: using the molecular graph model GNN to extract the feature vector of the sample drug graph data to obtain the first feature vector; using the protein sequence language model Transformer to extract the feature vector of the protein sequence data of the sample target to obtain the second feature vector; based on the first feature vector and the second feature vector extraction, the target interaction fingerprint information is obtained. After obtaining the interaction fingerprint information, the interaction fingerprint can be used as a label to train the two branch models of the initial first sub-model through error back propagation, that is, training the molecular graph model GNN and the protein sequence language model Transformer, thereby obtaining the pre-trained initial first sub-model.

[0057] In this step, a pre-training approach is used to obtain the initial first sub-model after pre-training. This has the advantage that the general source model obtained through training on a large-scale unlabeled source dataset can be migrated to subsequent target tasks (such as drug-target interaction prediction here, that is, whether the drug can inhibit or activate the protein target). When the prediction model for the target task needs to be trained later, only a small amount of labeled data is needed and the final interaction label is combined to fine-tune the retraining model, which can achieve more accurate prediction of drugs that can act on the target protein, laying the foundation for subsequent rapid and accurate training to obtain the target prediction model.

[0058] Step S202, based on the target sample drug, the target target, and the interaction information between the target sample drug and the target target, the pre-trained first sub-model and the initial second sub-model in the initial prediction model are trained to obtain the target first sub-model and the target second sub-model, thereby obtaining the target prediction model;

[0059] In the specific implementation process of this step, after obtaining the pre-trained initial first sub-model, the pre-trained initial first sub-model and the initial second sub-model can be retrained using small-scale interaction training data to obtain the final target prediction model. The specific model retraining process is:

[0060] (1) Input the target sample drug molecule data and target protein sequence data into the initial first sub-model after pre-training.

[0061] In the specific implementation process, the target sample drug molecules can be preprocessed to obtain the target sample drug graph data. At the same time, the target sample targets can be preprocessed to obtain the protein sequence data of each target sample target. The graph data of each target sample drug is then input into the molecular graph model GNN, and the protein sequence data of each target sample target is input into the protein sequence language model Transformer. The molecular graph model GNN and protein sequence language model Transformer are used to extract the interaction fingerprint information.

[0062] (2) Inputting the obtained interaction fingerprint information into the initial second sub-model, and using the target second sub-model in the target prediction model to predict the degree of interaction, thereby obtaining the prediction result of the interaction between the target sample drug and the target target.

[0063] (3) Based on the prediction results of the interactions and the label data of each interaction, the parameters of the pre-trained initial first sub-model and the initial second sub-model are adjusted to obtain the target first sub-model and the target second sub-model, thereby obtaining the target prediction model.

[0064] In this example, the training dataset includes: input: drug molecule information and target protein information; output: a Boolean value indicating whether an interaction occurs. The labeled datasets used here are: the Human dataset, which contains 369 positive interactions between 1052 drug molecules and 852 proteins; and the C. elegans dataset, which contains 4000 positive interactions between 1434 drug molecules and 2504 proteins.

[0065] In this embodiment, after the 768*2-dimensional vector output by the target first sub-model, that is, the interaction neural fingerprint is obtained, the second stage is entered. Because the obtained interaction neural fingerprint is relatively rich and has indicative information, this interaction neural fingerprint can be used as x, and the final interaction label can be used as Y as training data. Then, a fully connected neural network model is used to train to obtain the correlation between x and Y, that is, to train to obtain the target second sub-model, which can give a good prediction effect.

[0066] Step S203: pre-processing the drug to be tested to obtain graph data of the drug to be tested based on chemical bond connection information for extracting a first eigenvector; pre-processing the target to be tested to obtain protein sequence data of the target to be tested for extracting a second eigenvector;

[0067] During the specific implementation of this step, RDKit may be used to process the drug to be tested, so as to process it into graph data with atoms as nodes and connected by chemical bonds, that is, to obtain graph data containing chemical bond connection information.

[0068] Step S204: Use the molecular graph model in the first sub-model of the target to extract the feature vector of the graph data of the drug to be tested to obtain a first feature vector; use the protein sequence language model in the first sub-model of the target to extract the feature vector of the protein sequence data of the target to be tested to obtain a second feature vector; and extract the target interaction fingerprint information based on the first feature vector and the second feature vector.

[0069] In this step, a molecule (Molecule) can be regarded as a graph. The graph structure data contains nodes (node) and edges (edge), among which the nodes contain entity information (such as atoms in compounds, individuals in social networks), and the edges contain relationship information (relation) between entities (chemical bonds between atoms in graph compounds). Features include the type of each atom (Oxygen, oxygen atom), the properties of the atom itself (Atom Properties), some features of the compound (GlobalProperties), etc., that is, the node initial feature vector. Specifically, the node initial feature vector includes the feature vectors listed in Table 1 below. Consider each atom as a node in the graph, the atomic bond as an edge, and the edge also has corresponding features, the edge initial feature vector. The specific edge initial feature vector includes the feature vectors listed in Table 2 below.

[0070] Table 1:

[0071]

[0072]

[0073] Table 2:

[0074]

[0075] In this step, the input of the graph neural network model (i.e., molecular graph model) is usually a graph structure with node or edge attributes as mentioned above, that is, it includes the adjacency matrix A of the graph and the corresponding attribute information X. Its final output depends on the specific task. For example, node classification outputs the label of the node, graph classification outputs the label of the graph, and link prediction outputs whether the link exists or not. Taking graph classification as an example, the graph neural network model (i.e., molecular graph model) trains the implicit vector representation of each node in the graph based on the graph structure and input node attributes. Its goal is to make the vector representation contain sufficiently powerful expression information so that it can help each node to extract information. Finally, through average pooling and other methods, the information vector representation of the entire graph can be obtained. For example, through the characteristics of atomic nodes and the chemical bond information between atoms, the molecular-level information representation of the entire molecular compound can be extracted.

[0076] Specifically, the learning process of a graph neural network model (also known as a molecular graph model) involves iteratively aggregating and updating the neighbor information of nodes in the graph data. In each iteration, each node updates its own information by aggregating the features of its neighboring nodes and its own features from the previous layer, often performing nonlinear transformations on the aggregated information. By stacking multiple layers of the network, each node can obtain information about neighboring nodes within a certain number of hops.

[0077] The learning process of the graph network model (i.e., molecular graph model) involves two processes: the message passing phase and the readout phase. The message passing phase is the forward propagation phase, which runs T steps in a loop and passes the function M t Get information through function U t Update the node. The equation for this stage is as follows:

[0078]

[0079]

[0080] Among them, e vw represents the feature vector of the edge from node v to w; M t , U t Represents a function that can be set using different models. Take node v as the current node, m represents the message summary that the current node v receives from the neighbor node set Nv, M represents the message function, and the hidden state h of the current node v is integrated. v , neighbor node N v The hidden state h of each w node in w , and the edge e connecting v and w vw information,

[0081] h stands for hidden, which means hidden state; Indicates that the hidden state of the v node at time t+1 is obtained by the previous time t, h v Indicates the hidden state information of the v node itself, Indicates neighbor summary messages.

[0082] The readout stage calculates a feature vector for the representation of the entire graph, using the function R.

[0083]

[0084] Where T represents the total time step number, and the function M t , U t and R can use different model settings; h v Represents the hidden state information of node v at time T.

[0085] In this step, the molecular graph model is used to extract features from the graph data, which makes the extracted features more accurate and reliable, laying the foundation for subsequent accurate training to obtain the target prediction model.

[0086] Step S205 , calculating an interaction degree value based on the target second sub-model according to the interaction type in the target interaction fingerprint information; and obtaining a prediction result of the interaction between the drug to be tested and the target to be tested based on the interaction degree value.

[0087] In this step, the interaction type includes any one or more of the following: non-polar interaction, face-to-face π-π interaction, edge-to-edge π-π interaction, hydrogen bond with a residue as a donor, hydrogen bond with a residue as an acceptor, electrostatic interaction with a residue with a positive charge, and electrostatic interaction with a residue with a negative charge. In the specific implementation of this step, corresponding interaction values can be determined for different interaction types, or interaction values for each combination can be determined based on different combinations of interaction types, thereby calculating the final interaction degree value. In the specific implementation, a critical value can be set according to actual needs, and then the calculated interaction value is compared with the critical value. When the interaction degree value is greater than or equal to the critical value, it is determined that the drug to be tested has an interaction relationship with the target to be tested. When the interaction degree value is less than the critical value, it is determined that the drug to be tested does not have an interaction relationship with the target to be tested.

[0088] In this application, a self-supervised "interaction neural fingerprint" is constructed, that is, a neural network model is used to automatically learn and extract low-dimensional dense representation vectors of molecules and protein targets from structures such as molecules and protein targets or descriptors that retain a large amount of original structural information, that is, the extraction of 768*2 dimensional vectors. It is different from and has a higher degree of information enrichment than the original high-dimensional input vector (molecules and proteins may be features of thousands of dimensions), and is also better than the 7-dimensional traditional interaction fingerprint. The computable interaction neural fingerprint is then used as the intermediate label y, which breaks through the demand for interaction label data to a certain extent. By self-supervised learning of a large number of drug-target interaction fingerprints and training the learned model, better "interaction neural fingerprints" can be extracted. Furthermore, by fine-tuning the retraining model in combination with the final interaction label Y, more accurate predictions of drugs that can act on target proteins can be achieved.

[0089] At present, the success of deep learning in various fields including drug discovery is partly due to the possession of a large amount of labeled training data, because the performance of the model usually increases accordingly with the increase in the quality, diversity and quantity of training data. However, it is often very difficult to collect enough high-quality data to train the model to have good performance, especially in professional fields such as medicine and biochemistry where the cost and risk of sample data labeling are high. The method in this application obtains a model for accurately extracting interaction fingerprint information by using large-scale unlabeled unlabeled data to train the initial first sub-model. Therefore, in the subsequent interaction prediction model training process, a small batch of labeled data can be used for the target task. On the basis of the pre-trained initial first sub-model, further model training is performed to quickly train and obtain the target prediction model. In addition, the method in this application can use the pre-trained initial first sub-model for different target tasks, which improves the model reuse rate and further improves the efficiency of model training. The method in this application solves the problem that the existing machine learning-based methods cannot accurately predict interactions, and at the same time solves the problems of the existing methods such as the long model training process, the large amount of training data, and the inability to migrate and reuse the model.

[0090] Another embodiment of the present application provides a drug-target interaction prediction device, such as Figure 2 Shown, including:

[0091] Extraction module 1, for extracting interaction fingerprints based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model to obtain target interaction fingerprint information between the drug to be tested and the target to be tested;

[0092] Prediction module 2 is used to predict the interaction size based on the interaction type in the interaction fingerprint information using the target second sub-model in the target prediction model to obtain the prediction result of the interaction between the drug to be tested and the target to be tested.

[0093] During the specific implementation of this embodiment, the extraction module is specifically used to: use the molecular graph model in the first sub-model of the target to extract the feature vector of the drug to be tested to obtain a first feature vector; use the protein sequence language model in the first sub-model of the target to extract the feature vector of the target to be tested to obtain a second feature vector; and extract the target interaction fingerprint information based on the first feature vector and the second feature vector.

[0094] During the specific implementation of this embodiment, the drug-target interaction prediction device also includes a preprocessing module, which is used to: preprocess the drug to be tested to obtain graph data of the drug to be tested based on chemical bond connection information for extracting a first eigenvector; and preprocess the target to be tested to obtain protein sequence data of the target to be tested for extracting a second eigenvector.

[0095] During the specific implementation of this embodiment, the prediction module is specifically used to: calculate the interaction degree value based on the target second sub-model according to the interaction type in the target interaction fingerprint information; obtain the predicted result of the interaction between the drug to be tested and the target to be tested based on the interaction degree value; wherein the interaction type includes any one or more of the following: non-polar interaction, face-to-face π-π interaction, edge-to-edge π-π interaction, hydrogen bond with the residue as donor, hydrogen bond with the residue as acceptor, electrostatic interaction with positively charged residue and electrostatic interaction with negatively charged residue.

[0096] During the specific implementation of this embodiment, the drug-target interaction prediction device further includes a training module, which is configured to pre-train an initial first sub-model within the initial prediction model to obtain a pre-trained first sub-model. The pre-processing module is further configured to pre-process sample drugs to obtain graph data for each sample drug based on chemical bond connectivity information; and to pre-process sample targets to obtain protein sequence data for each sample target. The training module is specifically configured to pre-train the initial first sub-model using a back-propagation model training method based on each sample graph data, each sample target protein sequence data, and interaction fingerprint labels to obtain the pre-trained first sub-model.

[0097] During the specific implementation of this embodiment, the training module is also used to: perform model training on the pre-trained first sub-model and the initial second sub-model in the initial prediction model based on the target sample drug, the target target, and the interaction information between the target sample drug and the target target, to obtain the target first sub-model and the target second sub-model, so as to obtain the target prediction model.

[0098] In the specific implementation of this embodiment, the first eigenvector includes any one or more of the following: an atomic node initial eigenvector and an atomic edge initial eigenvector.

[0099] The drug-target interaction prediction device in the present application can accurately extract the interaction fingerprint based on the first target sub-model, and then use the second target sub-model to accurately predict the degree of interaction based on the extracted interaction fingerprint information. That is, the second target sub-model is used to accurately predict the degree of interaction based on the interaction type in the interaction fingerprint information, and then it can accurately determine whether the drug to be tested and the target to be tested have an interaction based on the degree of interaction, which can make the final prediction result more accurate and reliable.

[0100] Another embodiment of the present application provides a storage medium storing a computer program. When the computer program is executed by a processor, the following method steps are implemented:

[0101] Step 1: Based on the drug to be tested and the target to be tested, the target first sub-model in the target prediction model is used to extract the interaction fingerprint to obtain the target interaction fingerprint information between the drug to be tested and the target to be tested;

[0102] Step 2: Based on the interaction type in the interaction fingerprint information, the target second sub-model in the target prediction model is used to predict the degree of interaction to obtain a prediction result of the interaction between the test drug and the test target.

[0103] The specific implementation process of the above method steps can be found in the embodiments of any of the above-mentioned drug-target interaction prediction methods, and this embodiment will not be repeated here.

[0104] The storage medium in the present application can accurately extract the interaction fingerprint based on the target first sub-model, thereby accurately predicting the degree of interaction based on the extracted interaction fingerprint information using the target second sub-model, that is, using the target second sub-model to accurately predict the degree of interaction based on the interaction type in the interaction fingerprint information, and then accurately determine whether the drug to be tested and the target to be tested have an interaction based on the degree of interaction, which can make the final prediction result more accurate and reliable.

[0105] Another embodiment of the present application provides an electronic device, such as Figure 3 As shown, it at least includes a memory 1 and a processor 2. The memory 1 stores a computer program. When the processor 2 executes the computer program on the memory 1, it implements the following method steps:

[0106] Step 1: Based on the drug to be tested and the target to be tested, the target first sub-model in the target prediction model is used to extract the interaction fingerprint to obtain the target interaction fingerprint information between the drug to be tested and the target to be tested;

[0107] Step 2: Based on the interaction type in the interaction fingerprint information, the target second sub-model in the target prediction model is used to predict the degree of interaction to obtain a prediction result of the interaction between the test drug and the test target.

[0108] The specific implementation process of the above method steps can be found in the embodiments of any of the above-mentioned drug-target interaction prediction methods, and this embodiment will not be repeated here.

[0109] The electronic device in the present application can accurately extract the interaction fingerprint based on the target first sub-model, and then use the target second sub-model to accurately predict the degree of interaction based on the extracted interaction fingerprint information. That is, the target second sub-model is used to accurately predict the degree of interaction based on the interaction type in the interaction fingerprint information, and then it can accurately determine whether the drug to be tested and the target to be tested have an interaction based on the degree of interaction, which can make the final prediction result more accurate and reliable.

[0110] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A method for predicting drug-target interactions, characterized in that: include: Extracting feature vectors of the drug to be tested using the molecular graph model in the target first sub-model in the target prediction model to obtain a first feature vector; the first feature vector includes any one or more of the following: an atomic node initial feature vector and an edge initial feature vector; Using the protein sequence language model in the first sub-model of the target, extract the feature vector of the target to be tested to obtain a second feature vector; Extracting and obtaining the target interaction fingerprint information based on the first feature vector and the second feature vector; Calculating an interaction degree value based on the target second sub-model in the target prediction model and the interaction type in the target interaction fingerprint information; Obtaining a prediction result of the interaction between the drug to be tested and the target to be tested based on the interaction degree value; The interaction types include any one or more of the following: non-polar interactions, electrostatic interactions of residues with positive charges, and electrostatic interactions of residues with negative charges.

2. The method according to claim 1, wherein Before extracting a feature vector of the drug to be tested using the molecular graph model in the target first sub-model to obtain a first feature vector, the method further includes: Preprocessing the drug to be tested to obtain graph data of the drug to be tested based on chemical bond connection information for extracting a first eigenvector; The target to be detected is preprocessed to obtain protein sequence data of the target to be detected for extracting a second feature vector.

3. The method according to claim 1, wherein Before extracting the interaction fingerprint based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model, the method further includes: performing model pre-training on the initial first sub-model in the initial prediction model to obtain a pre-trained first sub-model. The specific training process includes: Preprocessing the sample drugs to obtain graph data of each sample drug based on chemical bond connection information; Preprocess the sample targets to obtain protein sequence data of each sample target; Based on the sample graph data, the protein sequence data of the sample targets and the interaction fingerprint labels, the initial first sub-model is pre-trained using a back-propagation model training method to obtain the pre-trained first sub-model.

4. The method according to claim 3, wherein Before extracting the interaction fingerprint based on the drug to be tested and the target to be tested using the target first sub-model in the target prediction model, the method further includes: Based on the target sample drug, the target target, and the interaction information between the target sample drug and the target target, the pre-trained first sub-model and the initial second sub-model in the initial prediction model are trained to obtain the target first sub-model and the target second sub-model to obtain the target prediction model.

5. A drug-target interaction prediction device, characterized in that: include: An extraction module is used to extract feature vectors of the drug to be tested using the molecular graph model in the target first sub-model in the target prediction model to obtain a first feature vector; Using the protein sequence language model in the first sub-model of the target, extract the feature vector of the target to be tested to obtain a second feature vector; extracting the target interaction fingerprint information based on the first feature vector and the second feature vector; the first feature vector includes any one or more of the following: an initial feature vector of a node and an initial feature vector of an edge of an atom; A prediction module is used to calculate an interaction degree value based on the target second sub-model in the target prediction model according to the interaction type in the target interaction fingerprint information; and obtain a prediction result of the interaction between the drug to be tested and the target to be tested based on the interaction degree value; wherein the interaction type includes any one or more of the following: non-polar interaction, electrostatic interaction of residues with positive charge, and electrostatic interaction of residues with negative charge.

6. A storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the drug-target interaction prediction method according to any one of claims 1 to 5 are implemented.

7. An electronic device, characterized in that: The method comprises at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the drug-target interaction prediction method according to any one of claims 1 to 5 when executing the computer program on the memory.

Citation Information

Patent Citations

  • Protein-ligand interaction fingerprint spectrum-based drug target prediction method

    CN107038348A

  • Drug and target interaction prediction method and device, equipment and storage medium

    CN113160894A

  • Drug small molecule property prediction method, device and equipment based on graph neural network

    CN113707236A

  • Drug molecule property prediction method, device and equipment based on comparative learning

    CN114386694A