Drug target interaction prediction method based on joint Q learning optimization meta path
Through joint Q learning, optimized metapaths, combined with higher-order graph convolutional neural networks and attention mechanisms, the problem of insufficient utilization of expert knowledge and node information in the existing technology is solved, and the efficiency and accuracy of drug target interaction prediction is improved.
Patent Information
- Application Number
- CN202510435035.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
AI Technical Summary
The existing meta-path-based drug target interaction prediction methods rely on expert knowledge, are prone to empirical subjectivity, and are not fully utilized, resulting in limited subgraph representation capabilities and reduced prediction performance.
Joint Q learning is used to optimize the metapaths, and by building a heterogeneous information network, using the two agents of drug and target to optimize the metapath collection, combining advanced graph convolutional neural networks and improved Vanilla graph convolutional neural networks to perform node representation learning, and fusing different metapath subgraph features through attention mechanisms to build a fully connected neural network for prediction.
It significantly improves the sub-graph representation ability, improves the adaptability and efficiency of the method, and improves the accuracy and robustness of drug target interaction prediction.
Smart Images

Figure CN120299506A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics and artificial intelligence, and particularly relates to a method for predicting drug-target interactions by optimizing meta-paths based on joint Q-learning. Background Art
[0002] Existing meta-path-based prediction methods mainly rely on expert knowledge to manually specify meta-paths, which are easily limited by the subjectivity and insufficient coverage of expert experience, and are difficult to adapt to dynamic data changes; at the same time, when constructing the meta-path subgraph, the node information in the graph is not fully utilized, ignoring the potentially useful features not covered by the path, resulting in limited subgraph representation ability and reduced prediction performance. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method for predicting drug-target interactions by optimizing meta-paths based on joint Q-learning, which fully utilizes node information, significantly improves the subgraph representation ability, and enhances the adaptability and efficiency of the method.
[0004] To solve the above technical problem, the present invention provides a method for predicting drug-target interactions by optimizing meta-paths based on joint Q-learning, including the following steps:
[0005] 1) Network construction and feature extraction: Integrate biological interaction data from different sources to construct a heterogeneous information network, and learn drug structure features and protein sequence features from drug SMILES and protein sequences;
[0006] 2) Meta-path optimization: Use the joint Q-learning algorithm to jointly optimize the meta-paths to obtain the optimal meta-path set;
[0007] 3) Node representation learning: Use the optimal meta-path set for node representation learning, including constructing a meta-path subgraph, using a high-order graph convolutional neural network and an improved Vanilla graph convolutional neural network to process meta-path subgraphs of different depths, and fusing node representation features learned from different meta-paths through an attention mechanism;
[0008] 4) Feature splicing and calculation: Splice the drug features and target features in the node features, send them into a fully connected neural network to calculate the interaction probability, form a final drug-target interaction prediction model and use it for prediction.
[0009] Further, the heterogeneous information network consists of four types of nodes, namely drugs, targets, diseases, and side effects, and six types of interaction relationships, namely drug-target, drug-drug, drug-disease, disease-side effect, target-target, and target-disease.
[0010] Further, the feature extraction method is as follows:
[0011] Convert the drug SMILES into a drug molecular graph and process it using Rdkit, where each atom is represented as a node and the chemical bond is represented as an edge. Use a graph convolutional neural network to encode the drug molecular structure, obtain the graph-level representation of the drug as the drug structure feature, and summarize the drug structure features output by the GCN through a global max pooling layer to obtain a compact feature representation for each drug.
[0012] For the protein sequence, use the CTD method to convert each protein sequence into a 128-dimensional feature vector, that is, the protein sequence feature; use a multi-layer perceptron (MLP) to process the protein sequence feature to obtain a compact feature representation with the same dimension as the drug structure feature.
[0013] Furthermore, use DEC-POMDP for meta-path optimization. DEC-POMDP is formalized as a tuple <N; S; A; T; R; O; Z; γ>, where N = 2 is the number of agents, S is the state space, A is the action, T is the state transition function, R is the reward function, O is the set of joint observations, Z is the observation function, and γ is the discount factor.
[0014] Furthermore, the optimal meta-path set includes the optimal drug meta-path set and the optimal target meta-path set.
[0015] Furthermore, when learning node representations, the meta-path subgraph includes the drug meta-path subgraph and the target meta-path subgraph; in the meta-path subgraph, when the meta-path depth is less than or equal to three, use the improved Vanilla graph convolutional neural network for information aggregation, and when the meta-path depth exceeds three, use the high-order graph convolutional neural network for information aggregation.
[0016] Furthermore, when the meta-path depth is less than or equal to three, each layer of the graph convolutional layer is defined as follows:
[0017]
[0018] where and are the input and output of the i-th layer respectively, represents the trainable weight matrix of edge type r in the i-th layer, σ is the activation function, and is the symmetric normalized adjacency matrix with self-connections of the meta-path subgraph,
[0019] Furthermore, when the meta-path depth is greater than three, each layer of the graph convolutional layer is defined as follows:
[0020] where the hyperparameter P is a set of integer adjacency matrix powers, represents the adjacency matrix Self - multiply by j times, and Π represents column - by - column concatenation.
[0021] Furthermore, the calculation formula of the fully - connected neural network is: where the input is the overall feature vector of the drug - target pair h drug is the drug - representation feature vector, h target is the target - representation feature vector, represents the concatenation operation, W is the weight matrix, b is the bias term, and σ represents the sigmoid activation function;
[0022] Output ranges from [0, 1], representing the probability value of possible interaction between the drug - target pair. When is greater than 0.5, it indicates that there is an interaction between the drug - target pair; conversely, when is less than or equal to 0.5, it indicates that there is no interaction between the drug - target pair.
[0023] Furthermore, input the prediction model result and the actual result into the cross - entropy loss function for parameter update. The cross - entropy loss function formula is as follows:
[0024]
[0025] where y represents the actual drug - target pair interaction label, which is 0 or 1, is the predicted interaction probability value of the prediction model for the corresponding sample; when y is 1, it is expected that the prediction of the prediction model is closer to 1; when y is 0, it is expected that the prediction of the prediction model is closer to 0; during the training process, by accumulating the losses of all samples and using optimization algorithms such as gradient descent, the parameters of the prediction model are adjusted.
[0026] Advantages of the present invention:
[0027] 1. By modeling the meta - path optimization as a Markov decision process and combining the joint Q - learning algorithm, the present invention can automatically learn and optimize the meta - path selection strategy according to the actual data and downstream task requirements, getting rid of the dependence on manual specification and expert knowledge, and improving the adaptability and efficiency of the method.
[0028] 2. Using two agents of drugs and targets to jointly optimize the meta - path set, constructing diverse meta - path sub - graphs, and combining the high - order graph convolutional neural network and the improved Vanilla graph convolutional neural network, making full use of the deep information of nodes and edges in the graph, significantly improving the sub - graph representation ability.
[0029] 3. By integrating the node features learned from different meta-path subgraphs through the attention mechanism, a more comprehensive graph neural network model is formed, effectively improving the accuracy and robustness of drug-target interaction prediction. Description of the Drawings
[0030] Figure 1 is the method framework diagram of the present invention;
[0031] Figure 2 is the aggregation mode diagram of Vanilla GCN and HOGCN of the present invention;
[0032] Figure 3 is the performance experimental result diagram of ROC-AUC of the present invention on Hetero-A and Hetero-B datasets;
[0033] Figure 4 is the performance experimental result diagram of PR-AUC of the present invention on Hetero-A and Hetero-B datasets;
[0034] Figure 5 is the performance experimental result diagram of Precision of the present invention on Hetero-A and Hetero-B datasets;
[0035] Figure 6 is the performance experimental result diagram of F1 of the present invention on Hetero-A and Hetero-B datasets;
[0036] Figure 7 is the performance experimental result diagram of Recall of the present invention on Hetero-A and Hetero-B datasets. Detailed Embodiments
[0037] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it, but the embodiments given are not intended to limit the present invention.
[0038] Referring to Figure 1 As shown, in an embodiment of the method for predicting drug-target interactions based on jointly optimizing meta-paths of the present invention, biological interaction datasets from different data sources are aggregated and a heterogeneous information network HIN is constructed, which consists of four types of nodes, namely drugs (D), targets (T), diseases (I), and side effects (S), and six types of interactions, namely D-T, D-D, D-I, D-S, T-T, and T-I.
[0039] Given the drug SMILES U = {u1,..., u m} of m drugs and the protein sequences S = {s1,..., sn}. In the feature extraction stage, drug structure features and protein sequence features are learned from drug SMILES and protein sequences. First, the drug SMILES is converted into a drug molecular graph, which is processed using Rdkit, where each atom is represented as a node and the chemical bond is represented as an edge. Next, a graph convolutional neural network (GCN) is used to encode the drug molecular structure to obtain the graph-level representation of the drug as the drug structure feature. Subsequently, the features output by the GCN are aggregated through a global max pooling layer to obtain a compact feature representation for each drug. For protein sequences, the Composition, Transition and Distribution (CTD) method is used to convert each sequence into a 128-dimensional feature vector. Finally, a multi-layer perceptron (MLP) is applied to process the protein sequence features to obtain a compact feature representation with the same dimension as the drug structure features. Thus, the unified dimension of drug structure features and protein sequence features is achieved.
[0040] After completion, meta-path optimization is performed, specifically the meta-path joint optimization combining multi-agent and joint Q-learning. The meta-path joint optimization problem can be regarded as a fully cooperative multi-agent problem, denoted as DEC-POMDP (Distributed Execution of Cooperative Partially Observable Markov Decision Processes). DEC-POMDP is a framework for modeling cooperation problems in multi-agent systems, which allows agents to cooperate in a partially observable environment to achieve a common goal. DEC-POMDP is formalized as a tuple <N; S; A; T; R; O; Z; γ>, where the meaning of each part is as follows:
[0041] Number of agents (N): N = 2 indicates that there are two agents, corresponding to drugs and targets respectively.
[0042] State space (S): The state space s represents the set of meta-paths of the current drug and target:
[0043] s = (s drug , s target ),
[0044] where: Φ drug and Φ target are the sets of meta-paths of drugs and targets respectively, and E φis the encoding of the meta-path φ. The state is processed by L1 normalization to ensure the relative contribution ratio of each meta-path encoding. The present invention assigns a unique ID (from 1 to n, where n is the number of relationship types) to each relationship type on the HIN. Each meta-path is encoded as a vector of length n, where each entity corresponds to the relationship type number in the meta-path. This ensures that all encodings have the same length. For example, on a HIN with 6 relationship types numbered from 1 to 6. If the ID of the meta-path φ is represented as [2, 6, 6, 4], then its encoding E φ will be (0, 1, 0, 1, 0, 2). The encoding is obtained by counting the number of occurrences of the relationships in the ID array. Here, ID = 2 appears once, ID = 4 appears once, and ID = 6 appears twice, corresponding to the values in the second, fourth, and sixth positions of the encoding set.
[0045] Action space (A): The action space A includes all relationships and a special STOP action:
[0046] A drug = {r1, r2,..., r n , STOP}
[0047] A target = {r1, r2,..., r m , STOP}
[0048] When the agent selects the relationship action r, the corresponding set of meta-paths will be extended to include the new relationship. If the STOP action is selected, the meta-path extension process of the agent terminates.
[0049] State transition function (T): The state transition function defines the probability of transitioning from state s to a new state s' under the joint action a. In this context, the transition is deterministic because the new state s' is completely determined by the current state s and the joint action a:
[0050] s' = UpdateMetapathSets(s, a)
[0051] The function to update the set of meta-paths specifically includes:
[0052] Φ′ drug = ExpandMetapaths(Φ drug , a drug )
[0053] Φ' target = ExpandMetapaths(Φ target , a target )
[0054] Then recalculate s'drug . and s' target .
[0055] Reward function (R): The reward function R(s,a) evaluates the quality of the meta-path optimization based on the prediction performance of the downstream model. Formally, the reward is defined as:
[0056] R(s,a) = N(s',a) - N(s,a prev )
[0057] where N(s,a) is an evaluation function based on the performance of the downstream prediction model. The reward encourages the agent to choose actions that can improve the prediction performance.
[0058] Joint observation set (O): Each agent has a certain local observation of the environment. The observation o of the drug agent drug includes the set of meta-paths of the current drug and the expandable relationships, and the observation o of the target agent target includes the set of meta-paths of the current target and the expandable relationships. Formally, the observation can be defined as:
[0059] o drug = ExtractObservation(s drug )
[0060] o target = ExtractObservation(s target )
[0061] These observation functions extract relevant agent information from the global state s.
[0062] Observation function (Z): The observation function Z defines the probability that the agent observes o after choosing the joint action a in state s. In this fully cooperative setting, the observation function is deterministic because the observation is directly obtained from the current state and action:
[0063] Z(o drug s,a) = P(o drug |s,a)
[0064] Z(o targeb s,a) = P(o target |s,a)
[0065] Discount factor (γ): The discount factor γ balances the importance of immediate rewards and future rewards. It is usually set between 0 and 1, for example: γ = 0.9, to ensure that the agent gives priority to actions that can bring long-term improvement in prediction performance.
[0066] Through the optimization of multi-agent and the selection of the joint Q-learning algorithm, the optimal drug meta-path set and the optimal target meta-path set can be obtained; in each training step, the two agents respectively select joint actions according to the current state. After selecting the actions, the system performs state transition according to the actions of drugs and targets and calculates the rewards. The reward function is designed based on the prediction accuracy of the model on the training set and updates its behavior according to the policy. Through such a training process, the reinforcement learning agent can learn the optimal meta-path selection strategy, thereby automatically selecting the most relevant meta-path set and maximizing the performance of drug-target interaction prediction.
[0067] Based on the optimal drug meta-path set and the optimal target meta-path set, the meta-path subgraphs corresponding to drugs and targets can be constructed respectively for node representation learning:
[0068] Among them, the depth of the meta-path subgraph refers to the number of layers used when constructing the graph neural network. In this paper, the depth of the meta-path is measured by the number of edges in the path. For meta-path subgraphs with different depths, different graph neural networks are used for information aggregation. When the meta-path depth is less than or equal to three, an improved Vanilla graph convolutional neural network is used because it performs well in capturing relatively neighboring node relationships. However, when the meta-path depth exceeds three, the relationships between nodes become more complex. Therefore, a high-order graph convolutional neural network is introduced to better capture these deep node associations and handle complex meta-paths involving multiple intermediate nodes.
[0069] The meta-path subgraph of each node is also a heterogeneous network with multiple types of nodes and edges. For each meta-path subgraph, they can be represented as G=(V, E, R), where V represents the set of nodes in the heterogeneous information network, E represents the set of edges in the heterogeneous information network, and R represents the set of edge types. Specifically, there are four types of nodes in the heterogeneous information network (i.e., drugs, proteins, diseases, side effects). Therefore, R includes six types of edges, namely drug-drug interaction, drug-protein interaction, drug-disease association, drug-side effect association, protein-protein interaction, and protein-disease association. The construction method of the meta-path subgraph determines that there is only one fixed edge type between layers of the subgraph. In the node feature initialization stage of the heterogeneous information network, except for drug and protein nodes, a normal distribution initialization method is used to initialize the features of other nodes.
[0070] Based on the Vanilla graph convolutional neural network, this application adds differential consideration of different edge types during the information transmission process. When the meta-path subgraph depth is less than or equal to three, each layer of the graph convolutional layer is defined as follows:
[0071]
[0072] where and are the input and output of the i-th layer respectively, represents the trainable weight matrix of edge type r in the i-th layer, σ is the activation function, and is the symmetric normalized adjacency matrix with self-connections of the meta-path subgraph, where A represents the adjacency matrix of the meta-path subgraph, I represents the identity matrix, and D represents the degree matrix of A + I n .
[0073] When the depth of the meta-path subgraph is greater than three, the relationships between nodes become more complex, and a larger range of information propagation is required to capture this complexity. For this purpose, the high-order graph convolutional neural network Mixhop is adopted, which allows the fusion of neighbor node information of different orders in information aggregation.
[0074] In Mixhop, the formula of the convolutional layer in GCN is replaced by:
[0075]
[0076] where the hyperparameter P is a set of integer adjacency matrix powers, represents the adjacency matrix multiplied by itself j times, and Π represents column-wise concatenation.
[0077] Generally speaking, for the case where the depth of the meta-path subgraph is less than or equal to three, the improved Vanilla graph convolutional neural network is used to aggregate node information. For the case where the depth is greater than three, the high-order graph convolutional neural network is adopted to better capture the complex node relationships, where the relationships between nodes involve multiple intermediate nodes. This strategy allows the model to flexibly select appropriate information aggregation methods in meta-path subgraphs of different depths, thereby improving the expressive power and performance of the model, as shown in Figure 2 .
[0078] For any node v ∈ V belonging to the drug type D D , there are M meta-paths, denoted as P D = {P1, P2,..., P M}. After the aggregation operation, M vector representations corresponding to the meta-paths of the drug node can be obtained, that is Each contains one aspect of the semantic information embedded in the drug node v. Since in the heterogeneous information network, meta-paths are not equally important, the attention mechanism is used to assign different weights to different meta-paths, so that the meta-paths with greater contributions play a greater role in the fusion.
[0079] First, for all drug nodes v ∈ VD The average operation of the meta-path representation vectors is used to summarize each drug meta-path P i ∈P D .
[0080] Among them, W D and b D are learnable parameters. Then, the attention mechanism is used to fuse the meta-path representation vectors of node v as follows:
[0081]
[0082] Among them, is the parameterized attention vector of the drug node type. can be interpreted as the relative importance of meta-path P i for drug type nodes. After calculating the i ∈P D of each meta-path P , the weighted sum of each meta-path representation vector of node v is performed to obtain the final node representation vector h v .
[0083] After completing the above node representation learning method and obtaining the node representation features, a fully connected neural network is used to predict drug-target pairs. The calculation formula is: where the input is the overall feature vector of the drug-target pair h drug is the drug representation feature vector, h target is the target representation feature vector, represents the concatenation operation, W is the weight matrix, b is the bias term, and σ represents the sigmoid activation function. The output ranges from [0,1], indicating the probability value of possible interaction between drug-target pairs. If is greater than 0.5, it is considered that the drug-target pair has an interaction; conversely, if is less than or equal to 0.5, it is considered that there is no interaction.
[0084] The model prediction and the actual result are input into the cross-entropy loss function for parameter update to achieve more accurate prediction of drug-target interactions. The cross-entropy loss function formula is as follows:
[0085]
[0086] Among them, y represents the actual drug-target interaction label (which can be 0 or 1), is the predicted interaction probability value of the model for the corresponding sample. When y is 1, it is hoped that the prediction of the model The closer to 1; when y is 0, it is expected that the prediction of the model is closer to 0.
[0087] The optimization process of this loss function aims to minimize the difference between the model's prediction of interactions and the actual interaction labels, so that the model can more accurately learn the association rules between drug targets. During the training process, by accumulating the losses of all samples and using optimization algorithms such as gradient descent, the model parameters are adjusted to improve its performance in the interaction prediction task.
[0088] The method proposed in the present invention was evaluated on two heterogeneous biological datasets. The heterogeneous dataset A (Hetero-A) is the gold standard dataset for predicting drug-target interactions (DTIs) of heterogeneous biological data. This dataset was collected from public resources and contains 12,015 biological entities, including 708 drugs (D), 1,512 targets (T), 5,603 diseases (I), and 4,192 side effects (S), as well as six types of interactions (a total of 1,895,445 interactions), namely D-T, D-D, D-I, D-S, T-T, and T-I. The heterogeneous dataset B (Hetero-B) is an extension of Hetero-A. It was collected from resources from October 1 to 15, 2021. Compared with Hetero-A, Hetero-B contains more samples for model training and evaluation (15,322 biological entities and 5,126,875 interactions).
[0089] The present invention evaluated the performance of the proposed method and multiple baseline methods on the Hetero-A and Hetero-B datasets through comparative experimental results. The experimental results can be seen Figures 3 - 7。On the Hetero-A dataset, the proposed method significantly outperforms other baseline methods in terms of ROC-AUC and PR-AUC. In particular, compared with HampDTI, it improves by 0.0543 and 0.0381 respectively. The ROC curve is the Receiver Operating Characteristic curve, which is plotted with the True Positive Rate (TPR) on the vertical axis and the False Positive Rate (FPR) on the horizontal axis. AUC is the Area Under the ROC Curve, and its value range is [0, 1]. It performs best in terms of Precision, reaching 0.9554, which is a significant improvement compared to other baseline methods. Precision is the proportion of true positive samples (correctly predicted positive samples) among the samples predicted as positive by the model. Although it is slightly inferior to HampDTI in terms of F1 and Recall, it still achieves relatively high values, 0.9304 and 0.9336 respectively. F1 is the F1 score, which is the harmonic mean of Precision and Recall and is used to comprehensively measure the balance between the two. Recall is the recall rate of the prediction model.
[0090] On the Hetero-B dataset, the proposed method also significantly outperforms other baseline methods in terms of ROC-AUC and PR-AUC, with improvements of 0.0813 and 0.0651. It achieves the highest score in terms of Precision, which is 0.9401, representing a significant improvement compared to other methods. In terms of F1 and Recall, they are 0.9241 and 0.9101 respectively, showing excellent performance compared to other baseline methods, especially having a significant advantage in terms of Recall. The experimental results on the two datasets show that the proposed method exhibits obvious advantages in all evaluation metrics, especially achieving significant improvements in key metrics such as Precision, ROC-AUC, and PR-AUC. This indicates that the proposed method has higher performance and better comprehensive performance in the drug-target interaction prediction task.
[0091] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art based on the present invention are all within the protection scope of the present invention.
Claims
1. A method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning, characterized in that It includes the following steps: 1) Network construction and feature extraction: Integrate biological interaction data from different sources, construct a heterogeneous information network, and learn drug structure features and protein sequence features from drug SMILES and protein sequences; 2) Meta-path optimization: Use the joint Q-learning algorithm to perform joint optimization of meta-paths to obtain an optimal set of meta-paths; 3) Node representation learning: Use the optimal set of meta-paths for node representation learning, including constructing meta-path subgraphs, using high-order graph convolutional neural networks and improved Vanilla graph convolutional neural networks to process meta-path subgraphs of different depths, and fusing node representation features learned from different meta-paths through an attention mechanism; 4) Feature concatenation and calculation: Concatenate the drug features and target features in the node features, send them into a fully connected neural network to calculate the interaction probability, form a final drug-target interaction prediction model and use it for prediction.
2. The method for predicting drug-target interactions based on optimizing meta-paths by combining Q-learning according to claim 1, wherein The heterogeneous information network consists of four types of nodes, namely drugs, targets, diseases, and side effects, and six types of interaction relationships, namely drug-target, drug-drug, drug-disease, disease-side effect, target-target, and target-disease.
3. The method for predicting drug-target interactions based on optimizing meta-paths by combined Q-learning according to claim 2, wherein The feature extraction method is as follows: Convert the drug SMILES into a drug molecular graph and process it using Rdkit, where each atom is represented as a node and the chemical bond is represented as an edge. Use a graph convolutional neural network to encode the drug molecular structure to obtain the graph-level representation of the drug as the drug structure feature, and summarize the drug structure features output by the GCN through a global max pooling layer to obtain a compact feature representation of each drug; For protein sequences, use the CTD method to convert each protein sequence into a 128-dimensional feature vector, that is, the protein sequence feature; use a multi-layer perceptron (MLP) to process the protein sequence feature to obtain a compact feature representation with the same dimension as the drug structure feature.
4. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 1, wherein, Use DEC-POMDP for meta-path optimization. DEC-POMDP is formalized as a tuple <N; S; A; T; R; O; Z; γ>, where N = 2 is the number of agents, S is the state space, A is the action, T is the state transition function, R is the reward function, O is the joint observation set, Z is the observation function, and γ is the discount factor.
5. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 1, characterized in that The optimal set of meta-paths includes the optimal drug meta-path set and the optimal target meta-path set.
6. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 1, wherein During node representation learning, the meta-path subgraph includes the drug meta-path subgraph and the target meta-path subgraph; in the meta-path subgraph, when the meta-path depth is less than or equal to three, use the improved Vanilla graph convolutional neural network for information aggregation, and when the meta-path depth exceeds three, use the high-order graph convolutional neural network for information aggregation.
7. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 6, wherein When the meta-path depth is less than or equal to three, each layer of the graph convolutional layer is defined as follows: where and are the input and output of the $i$-th layer respectively, represents the trainable weight matrix of edge type $r$ in the $i$-th layer, $\sigma$ is the activation function, and is the symmetric normalized adjacency matrix with self-connections of the metapath subgraph, 8. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 6, wherein, When the meta-path depth is greater than three, each layer of the graph convolutional layer is defined as follows: where the hyperparameter P is a set of integer powers of the adjacency matrix, representing the adjacency matrix multiplied by itself j times, and Π represents column-wise concatenation.
9. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 1, wherein The calculation formula of the fully connected neural network is as follows: Among them, the input is the overall feature vector of the drug-target pair h drug is the drug representation feature vector, h target is the target representation feature vector, represents the concatenation operation, W is the weight matrix, b is the bias term, and σ represents the sigmoid activation function; Output ranges from [0, 1], representing the probability value of possible interaction between drug target pairs. When is greater than 0.5, it indicates that there is an interaction between the drug target pair; conversely, when is less than or equal to 0.5, it indicates that there is no interaction between the drug target pair.
10. The method for predicting drug-target interactions based on optimizing meta-paths by joint Q-learning according to claim 1, characterized in that, Input the prediction model result and the actual result into the cross-entropy loss function for parameter update. The formula of the cross-entropy loss function is as follows: Among them, y represents the actual drug target pair interaction label, which is 0 or 1, is the predicted interaction probability value of the prediction model for the corresponding sample; when y is 1, it is hoped that the prediction of the prediction model is closer to 1; when y is 0, it is hoped that the prediction of the prediction model is closer to 0; during the training process, by accumulating the losses of all samples and using optimization algorithms such as gradient descent, the parameters of the prediction model are adjusted.