A Drug Property Prediction Method Based on Reinforcement Learning and Molecular Network Data Augmentation
By adopting a graph enhancement model based on reinforcement learning in drug properties prediction, a new graph data set that generates and combines local structural information, the problems of low prediction accuracy and overfitting in the prior art are solved, and more efficient drug properties prediction is achieved.
Patent Information
- Application Number
- CN202211089773.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-07
AI Technical Summary
Existing drug properties prediction methods have insufficient in capturing the local structural characteristics of compound graph data, resulting in low prediction accuracy and easy overfitting of training on small-scale compound datasets.
A graph enhancement model based on reinforcement learning is adopted, and the edges are added or deleted on the input graph through the generator to generate an enhancement graph, and the generator is trained in combination with the Monte Carlo strategy gradient method to capture local structural information. Combine the generated new graph dataset with the original dataset to enhance the data and improve the prediction accuracy.
By capturing the local structural characteristics of compound graph data, the accuracy of drug properties prediction is significantly improved, overfitting problems are avoided, and optimal performance is achieved on multiple public data sets.
Smart Images

Figure CN116312846B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of reinforcement learning, graph data augmentation, and graph data mining, and mainly relates to a drug property prediction method based on reinforcement learning and molecular network data augmentation. Background Art
[0002] Drug discovery is the process of discovering new candidate drugs, involving knowledge in fields such as medicine, biology, and pharmacology. Researchers generally select targets based on the biochemical mechanisms involved in disease conditions, and then use candidate drugs found in academic, pharmaceutical, or biotech research laboratories to test their interactions with drug targets. Once a drug is confirmed to be effective against the target, the drug is verified by examining its effects on the related disease. There may be up to 50,000 to 100,000 potential candidate drugs for each disease, and all these molecules need to be rigorously screened. With the emergence of interdisciplinary thinking, the application of artificial intelligence in the field of drug discovery has attracted the interest of a large number of researchers. Traditional pharmaceutical companies initially needed to manually screen a large number of candidate drugs and then conduct tests on this basis. However, by introducing artificial intelligence technology in the process of preparing new drug compounds, the properties of compound molecules can be predicted more objectively and accurately, thereby determining the efficacy and safety of the prediction samples for specific diseases, not only accelerating the drug development cycle but also improving research efficiency. As one of the artificial intelligence technologies, graph neural networks (GNNs) have shown good performance in different graph tasks, such as node classification, graph classification, and link prediction. Since the complex relationships between atoms in a molecular formula can be modeled using graph data, graph neural networks play an increasingly important role in the field of drug property prediction. Existing research frameworks generally use a graph data set of compounds as the input to the GNN model, learn graph-based representations, capture structural features such as the order, topology, and geometry of the graph data, and then perform classification or regression tasks to obtain property prediction results. However, such a framework has two limitations: (1) The model only generates node-level or graph-level representations of the original graph data, making it difficult to discover the role of local structures that have an important impact on the prediction of chemical molecule properties; (2) For small-scale compound data sets, due to the small number of training samples, the training is prone to overfitting, the classification accuracy is low, and good property prediction tasks cannot be completed.
[0003] Based on the above problems, the present invention proposes a drug property prediction method based on reinforcement learning and molecular network data augmentation. Specifically, for problem (1), we first propose a graph augmentation model based on reinforcement learning. The model includes a graph generator that adds or deletes edges on the input graph according to the node selection strategy to generate an augmented graph, and then uses the Monte Carlo policy gradient method to train the generator according to the feedback of the trained graph isomorphism network (GIN) model. By studying the influence of local structure information in the molecular graph on the prediction result, the graph generator can combine the local structure information to generate a new graph that maximizes the prediction result under the trained GIN model. Adding the generated new graph to the original molecular dataset realizes the data augmentation effect, thereby avoiding problem (2) and improving the prediction accuracy. In addition, to make the generated graph actually effective, we combine some graph rules. Summary of the Invention
[0004] The present invention aims to overcome the above-mentioned disadvantages of the prior art and provides a drug property prediction method based on reinforcement learning and molecular network data augmentation.
[0005] The present invention is mainly composed of two modules spliced together: one module is a graph augmentation model based on reinforcement learning, which generates an augmented graph by learning the influence of the local structure of the graph on the prediction result of the classification model. It mainly includes the following steps: selecting a starting node, selecting an ending node, adding or deleting edges between the two nodes, calculating the reward of the new graph, training the generator, and constructing a new graph dataset; the other module is a graph classification module, which expands the new graph dataset generated by the previous module into the original dataset, and then inputs the augmented dataset into a graph classification model composed of a three-layer graph isomorphism network and a softmax function for training. The actual drug molecule dataset is input into the trained model to classify the drug properties and realize the property prediction function.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A drug property prediction method based on reinforcement learning and molecular network data augmentation, comprising the following steps:
[0008] S1: Obtain the original compound dataset and input it into the graph augmentation model based on reinforcement learning;
[0009] S2: Input the graph data G n , first aggregate the neighborhood information and learn the node features through a three-layer graph convolutional neural network (GCNs), and then predict the probability p of the starting node through the first multi-layer perceptron (MLP) n,start , and determine the starting node index from the probability distribution, and mark this node as a n,start ;
[0010] S3: "Broadcast" the features of the starting node and all nodes, and use the second MLP to calculate the probability p of the ending node n,end , and determine the ending node index from the probability distribution, and mark this node as a n,end ;
[0011] S4: After determining the starting node and ending node of the good edges, construct a new graph G through the edge addition and deletion strategy n+1 ;
[0012] S5: Calculate the reward R n to evaluate the generated new graph. One part of the reward R n comes from the feedback of the trained classification model g(·), and the other part comes from the graph rule constraints;
[0013] S6: Train the generator with Monte Carlo policy gradient;
[0014] S7: Repeat the process from S2 to S6 until the number of iterations reaches the preset value N, and then store the new graphs generated under the condition that R n is greater than 0 and the original graphs where R n is less than 0 in all iterations as a new graph data set;
[0015] S8: Randomly divide the original graph data set into a training set and a validation set, and then expand the original graph data set with the new graph data set to achieve data augmentation;
[0016] S9: Use all the original training sets and the same number of data extracted from the new graph data set to train the GIN model, and at the same time use the original validation set to verify the performance of the GIN model. When the validation loss no longer shows a downward trend and only fluctuates within a small range, use the early stopping method to stop training;
[0017] S10: After training the model, input the chemical drug molecule data set into the model for classification of drug properties to achieve the drug prediction function.
[0018] Furthermore, in step S1, the present invention downloads a chemical molecule data set or a biochemical data set from the TUDataset of the graph neural network library PyG, and then inputs these compound data in the form of graph data into the graph augmentation model based on reinforcement learning, where the nodes correspond to atoms, the edges correspond to chemical bonds, and the graph labels correspond to the chemical properties of the corresponding compounds.
[0019] Furthermore, the step S2 includes the following steps:
[0020] S2.1: Let the current number of iterations be n (n ≤ N), and input the graph data G nFirst, aggregate domain information and learn node features through a three-layer Graph Isomorphism Network (GIN) to obtain a node feature representation
[0021]
[0022] S2.2: Node feature representation After passing through the first Multi-Layer Perceptron (MLP), a one-dimensional node vector is obtained. Then, this one-dimensional node vector is input into the softmax function, and the resulting one-dimensional vector is denoted as the probability p of the starting node n,start :
[0023]
[0024] S2.3: Take the probability p n,start The index of the maximum value in it is used as the starting node and is marked as a n,start :
[0025] a n,start = argmax(p n,start )(3) Further, the step S3 includes the following steps:
[0026] S3.1: Concatenate the node features and the starting node features in the form of "broadcast", and then input them into the softmax function to obtain a one-dimensional vector result, which is denoted as the probability p of the ending node n,end :
[0027]
[0028] S3.2: Since the starting node and the ending node must be different, the values at the starting node index in the probability p n,end need to be masked. Let mask n,end be a mask with 1s outside the position of a n,start . After masking p n,end , find the index of the maximum value in the probability p n,end as the ending node and mark it as a n,end :
[0029] a n,end = argmax(p n,end · mask n,end )(5)
[0030] Further, in the step S4, after selecting the starting node and the ending node, it is judged whether there is an edge connecting the two nodes in the original graph. If the two nodes are originally connected, then their connecting edge is deleted to generate a new graph; otherwise, an edge is added, and the generated new graph is denoted as G n+1。
[0031] Further, the step S5 includes the following steps:
[0032] S5.1: Calculate the reward R n,f 。Reward R n,f is the feedback of the new graph under the trained classification model g(·), which is divided into the intermediate reward R n,f-mid and the final reward R n,f-final , and the weighted sum of the two, where α represents the weight value of the final reward:
[0033] R n,f = R n,f-mid + α·R n,f-final (6)
[0034] Among them, the intermediate reward R n,f-mid is the predicted probability of obtaining the original graph type class(Gn) when the new graph passes through the classification model for the first time, and it is updated as feedback to the graph generator. Its calculation is shown in formula (7), where g(·) consists of 3-layer graph isomorphism network and a softmax function, and l represents the total number of all possible classes of the classification model g(·);
[0035]
[0036] The final reward is to perform operations S2 to S4 on G n+1 iteratively until the preset number of times m is reached and then stop. Each result is recorded as Rollout i (G n+1 ), i = 1, 2....m; in each loop, when S4 is to add an edge, the reward is calculated according to the intermediate reward formula (7), otherwise the reward value is 0; the average of all rewards in the loop process is obtained to get the final reward R n,f-final :
[0037]
[0038] S5.2: Calculate the rule reward R n,r ——Perform rule constraints on the generated new graph Rollout m (G n+1 ): The first judgment rule is whether the new graph is a biological compound. If the corresponding molecular smiles can be found in the RDkit library, return 0, otherwise return -1; another rule is that the degree of the node cannot exceed the valence of the atom corresponding to the node. If it is satisfied, return 0, otherwise return 1;
[0039] S5.3: The reward R n is expressed as the following formula (9), where β represents the weight value of the rule-based reward.
[0040] Rn = R n,f + β × R n,r (9)
[0041] Furthermore, in step S6, a Monte Carlo policy gradient is used to train the generator, and the loss function in the nth iteration is as shown in formula (10), where L CE represents the cross-entropy loss function.
[0042] Loss g = -R n (L CE (p n,start , a n,start ) + L CE (p n,end , a n,end )) (10)
[0043] Furthermore, in step S7, the processes of S2 to S6 are repeated until the number of iterations reaches the preset value N, and then the new graphs generated under the condition that R t is greater than 0 and the original graphs where R t is less than 0 in all iterations are stored as a new graph data set.
[0044] Furthermore, step S8 includes the following steps:
[0045] S8.1: The original compound data set is randomly divided into a training set and a validation set at a ratio of 8:2;
[0046] S8.2: The feature matrices and adjacency matrices of the original graph data set and the new graph data set are respectively stacked in the tensor depth direction; the labels of the original graph data set and the new graph data set are stacked in the tensor horizontal direction.
[0047] Furthermore, in step S9, all the original training set data and the same number of data in the second half of the augmented training set are taken to further train the 3-layer GIN classification model. At the same time, each time the model is trained, the performance of the classification model is verified using the original validation set; when the validation loss no longer shows a downward trend and only fluctuates within a small range, the early stopping method is used to stop the training;
[0048] Furthermore, in step S10, by applying a drug property prediction method of the present invention, a data set in the field of chemical drugs (such as MUTAG, PTC_MR, or NCI1, etc.) is input into the trained classification model for drug property classification, realizing the drug prediction function, and compared with an un-augmented graph classification model with the same configuration, the prediction accuracy is significantly improved.
[0049] The beneficial effects of the present invention are as follows: The present invention trains a generator based on reinforcement learning, and generates a new graph containing local structure information by studying the influence of the information of local structures in the molecular graph on the prediction results of the classification model; the generated new graph is used to expand the original dataset, achieving the effect of data augmentation, and solving the above problems (1) and (2); compared with directly training a classification model composed of three-layer graph isomorphism networks using the original dataset, this method achieves optimal performance on multiple public datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is the overall method flowchart of the method of the present invention;
[0051] Figure 2 for graph G at the nth iteration n generating a new graph G under the graph augmentation model n+1 process schematic diagram;
[0052] Figure 3 is the flowchart of the graph augmentation module based on reinforcement learning. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following further describes in detail the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0054] The purpose of the present invention is to generate new graphs by strategically adding and deleting edges to the original graph data, expand the original graph dataset, and further train the graph classification model, so that the accuracy of the model in predicting drug properties is significantly improved.
[0055] The whole method includes a graph augmentation module and a graph classification model: the generator of the graph augmentation module determines the starting node and the ending node through a node selection strategy, and adds and deletes edges according to the connection situation of the two nodes to generate a new graph; the new graph is input into the trained graph classification model to obtain classification feedback, and combined with some graph rules to form a reward signal to continuously improve the performance of the generator, so that it can learn the local structure information in the molecular graph and generate a new graph with a different structure from the original data but with a large class score in the classification model; expanding the new graph into the original dataset and further training the classification model will effectively improve the accuracy of drug property prediction.
[0056] Referring to Figure 1 , a method for predicting drug properties based on reinforcement learning and molecular network data augmentation, includes the following steps:
[0057] S1: Obtain the original compound dataset PTC_MR and input it into the graph augmentation model based on reinforcement learning;
[0058] S2: Input graph data G n, first aggregate domain information and learn node features through a three-layer Graph Isomorphism Network (GIN), and then predict the probability p of the starting node through the first Multi-Layer Perceptron (MLP). n,start , and determine the starting node index from the probability distribution, and mark this node as a. n,start ;
[0059] S3: "Broadcast" the features of the concatenated starting node and all nodes, and use the second MLP to calculate the probability p of the ending node. n,end , and determine the ending node index from the probability distribution, and mark this node as a. n,end ;
[0060] S4: After determining the starting node and ending node of the edge, construct a new graph G through the edge addition and deletion strategy. n+1 ;
[0061] S5: Calculate the reward R. n to evaluate the generated new graph, and one part of the reward R. n comes from the feedback of the trained classification model g(·), and the other part comes from the graph rule constraints;
[0062] S6: Train the generator with Monte Carlo policy gradient;
[0063] S7: Repeat the process from S2 to S6 until the number of iterations reaches the preset value N, and then store the new graphs generated under the condition that R n is greater than 0 and the original graphs where R n is less than 0 in all iterations as a new graph data set;
[0064] S8: Randomly divide the original graph data set into a training set and a validation set, and then expand the original graph data set with the new graph data set to achieve data augmentation;
[0065] S9: Train the GIN model with all the original training sets and the same number of data extracted from the new graph data set, and at the same time verify the performance of the GIN model with the original validation set. When the validation loss no longer shows a downward trend and only fluctuates within a small range, use the early stopping method to stop training;
[0066] S10: After training the model, input the PTC_MR data set into the model for classification of drug properties to achieve the drug prediction function.
[0067] Furthermore, in step S1, the present invention downloads the mouse carcinogenicity data set PTC_MR from the TUDataset of the graph neural network library PyG. This data set contains 344 graphs representing compounds and is commonly used for rodent carcinogenicity prediction tasks; input the downloaded original PTC_MR data set into the graph augmentation model based on reinforcement learning, and set the model iteration number N = 200.
[0068] Further, the steps S2 to S6 correspond to Figure 2 in the nth iteration of Figure G in n the process of generating a new graph G under the graph enhancement model, where the step S2 includes the following steps: n+1 The process of generating a new graph G under the graph enhancement model, where the step S2 includes the following steps:
[0069] S2.1: Let the current iteration number be n (n ≤ N), and input the graph data G n is a graph composed of 4 nodes. It first aggregates the neighborhood information and learns the node features through a three-layer graph isomorphism network (GIN) to obtain a node feature representation
[0070]
[0071] S2.2: The node feature representation passes through the first multi-layer perceptron MLP to obtain a one-dimensional node vector, and then this one-dimensional node vector is input into the softmax function. The obtained one-dimensional vector result is expressed as the probability p of the starting node n,start , where Figure 2 the length of p n,start represents the total number of nodes in the graph, and the larger the value, the darker the corresponding color;
[0072]
[0073] S2.3: Take the index of the maximum value in the probability p n,start as the starting node, Figure 2 mark node 1 as a n,start :
[0074] a n,start = argmax(p n,start ) (3) Further, the step S3 includes the following steps:
[0075] S3.1: Concatenate the node feature and the starting node feature in the form of "broadcast", and then input it into the softmax function to obtain a one-dimensional vector result expressed as the probability p of the ending node n,end :
[0076]
[0077] S3.2: Since the starting node and the ending node must be different, the value at the starting node index needs to be masked in the probability p n,end . Let mask n,end be a mask with 1 outside the position of a n,start and mask pn,end Find the probability p after masking n,end Use the index of the maximum value as the end node, and mark node 3 as a on the graph n,end :
[0078] a n,end = argmax(p n,end * mask n,end ) (5)
[0079] Furthermore, in step S4, after selecting the start node and the end node, determine whether there is an edge connecting the two nodes in the original graph - in Figure 2 , there was originally no edge between node 1 and node 3 on the original graph. According to the strategy, an edge needs to be added to form a new graph denoted as G n+1 .
[0080] Furthermore, step S5 includes the following steps:
[0081] S5.1: Calculate the reward R n,f . The reward R n,f is the feedback of the new graph under the trained classification model g(·). It includes an intermediate reward R n,f-mid and a final reward R n,f-final . The two are weighted and summed, where α represents the weight value of the final reward;
[0082] R n,f = R n,f-mid + α·R n,f-final (6)
[0083] Among them, the intermediate reward is the predicted probability of obtaining the original graph type class(G n ) when the new graph first passes through the GIN classification model. It is used as feedback to update the graph generator, and its calculation is shown in formula (7). Among them, g(·) consists of 3 layers of graph isomorphism networks and a softmax function, and l represents the total number of all possible classes of the classification model g(·);
[0084]
[0085] The final reward is obtained by iteratively performing operations from S2 to S4 on G n+1 until the preset number of times m = 4 is reached and then stopping. Each result is recorded as Rollout i (G n+1 ), i = 1, 2....m; In each loop, when S4 is to add an edge, calculate the reward according to the intermediate reward formula (7), otherwise the reward value is 0; Average all the rewards during the loop process to obtain the final reward R n,f-final :
[0086]
[0087] S5.2: Calculate the rule reward R n,r ——Perform rule constraints on the generated new graph Rollout m (G n+1 ) to perform rule constraints: The first judgment rule is whether the new graph is a biological compound. If the corresponding molecular smiles can be found in the RDkit library, return 0; otherwise, return -1. Another rule is that the degree of the node cannot exceed the valence of the atom corresponding to the node. If satisfied, return 0; otherwise, return 1;
[0088] S5.3: Reward R n It is expressed as the following formula (9), where β represents the reward weight based on rules.
[0089] R n = R n,f + β × R n,r (9)
[0090] Furthermore, in the step S6, the Monte Carlo policy gradient is used to train the generator. The loss function in the nth iteration is as shown in formula (10), where L CE represents the cross-entropy loss function.
[0091] Loss g = -R n (L CE (p n,start , a n,start ) + L CE (p n,end , a n,end )) (10)
[0092] Furthermore, referring to Figure 3 , it shows the entire framework process of generating the new graph dataset. In the step S7, the processes of S2 to S6 are repeated until the number of iterations reaches the preset value N = 200. Then, the new graphs generated under the condition that R t is greater than 0 in each iteration and the original graphs where R t is less than 0 in all iterations are stored as the new graph dataset.
[0093] Furthermore, the step S8 includes the following steps:
[0094] S8.1: Randomly divide the original compound dataset into a training set and a validation set at a ratio of 8:2;
[0095] S8.2: Stack the feature matrices and adjacency matrices of the original graph dataset and the new graph dataset respectively in the tensor depth direction; stack the labels of the original graph dataset and the new graph dataset in the tensor horizontal direction.
[0096] Furthermore, in step S9, all the original training set data and the same amount of data in the second half of the augmented training set are taken to further train the 3-layer GIN classification model. At the same time, each time the model is trained, the performance of the classification model is verified using the original validation set. When the validation loss no longer shows a downward trend and only fluctuates within a small range, the early stopping method is used to stop the training.
[0097] Furthermore, in step S10, by applying a drug property prediction method of the present invention, the PTC_MR data set is input into the trained classification model for drug property classification, and the accuracy of predicting rodent carcinogenicity is as high as 79.2±3.6 (%). Comparing the prediction result of the present invention with the non-augmented graph classification model, under the same configuration, the prediction accuracy of the non-augmented graph classification model is 64.6±7.0 (%). The method of the present invention improves the accuracy of the drug property prediction model by 14.6%, and the effect is remarkable.
[0098] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept of the present invention.
Claims
1. A drug property prediction method based on reinforcement learning and molecular network data augmentation, characterized in that, It includes the following steps: S1: Obtain the original compound dataset and input it into the graph augmentation model based on reinforcement learning; S2: Input graph data G n , first aggregate neighborhood information and learn node features through a three-layer graph convolutional neural network GCNs, and then predict the probability p of the starting node through the first multi-layer perceptron MLP n,start , determine the starting node index from the probability distribution, and mark this node as a n,start ; S3: "Broadcast" the features of the starting node and all nodes, and use the second MLP to calculate the probability p of the ending node n,end , determine the ending node index from the probability distribution, and mark this node as a n,end ; S4: After determining the starting node and ending node of the edge, construct a new graph G through the edge addition and deletion strategy n+1 ; S5: Calculate the reward R n to evaluate the generated new graph. One part of the reward R n comes from the feedback of the trained classification model g(·), and the other part comes from the graph rule constraints; S6: Train the generator with Monte Carlo policy gradient; S7: Repeat the process from S2 to S6 until the number of iterations reaches the preset value N, and then store the new graphs generated under the condition that R in each iteration n is greater than 0 and the original graphs where R in all iterations n are all less than 0 as a new graph data set; S8: Randomly divide the original graph dataset into a training set and a validation set, and then augment the original graph dataset with the new graph dataset to achieve data augmentation; S9: Use all the original training set and the same number of data samples drawn from the new graph dataset to train the GIN model, and at the same time use the original validation set to verify the performance of the GIN model. When the validation loss no longer shows a downward trend and only fluctuates within a small range, the early stopping method is used to stop the training; S10: After the model is trained, input the chemical drug molecule dataset into the model for classification of drug properties to achieve the drug prediction function.
2. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, wherein, In step S1, download the chemical molecule dataset or the biochemical dataset from the TUDataset of the graph neural network library PyG, and then input these compound data in the form of graph data into the graph augmentation model based on reinforcement learning, where the nodes correspond to atoms, the edges correspond to chemical bonds, and the graph labels correspond to the chemical properties of the corresponding compounds; Step S2 includes the following steps: S2.1: Let the current iteration number be n, where n ≤ N, and input the graph data G n First, aggregate the neighborhood information and learn the node features through a three-layer Graph Isomorphism Network (GIN) to obtain a node feature representation S2.2: Node Feature Representation After passing through the first multi-layer perceptron (MLP), a one-dimensional node vector is obtained. Then, this one-dimensional node vector is input into the softmax function, and the resulting one-dimensional vector is denoted as the probability p of the starting node n,start :[[]]END]] S2.3: Obtain probability p n,start Take the index of the maximum value in n,start as the starting node and label it as a n,start : a n,start = argmax(p n,start ) (3).
3. A drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, characterized in that Step S3 includes the following steps: S3.1: Concatenate the node features and the starting node features in a "broadcast" form, and then input them into the softmax function to obtain a one-dimensional vector, which is expressed as the probability p of the ending node n,end : S3.2: Since the start node and the end node must be kept different, the value at the start node index needs to be masked in the probability p n,end Let mask n,end be a n,start a mask with 1s outside the position, and mask p n,end After masking, find the index of the maximum value in the probability p n,end as the end node, marked as a n,end : a n,end = argmax(p n,end · mask n,end ) (5).
4. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, wherein In the step S4, after selecting the starting node and the ending node, it is determined whether there is an edge connecting the two nodes in the original graph; if the two nodes are originally connected, their connecting edge is deleted to generate a new graph, otherwise an edge is added, and the generated new graph is denoted as G n+1 .
5. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, wherein Step S5 includes the following steps: S5.1: Calculate the reward R n,f ; The reward R n,f is the feedback of the new graph under the trained classification model g(·), which is divided into the intermediate reward R n,f-mid and the final reward R n,f-final , and the weighted sum of the two, where α represents the weight of the final reward; R n,f = R n,f-mid + α·R n,f-final (6) Among them, the intermediate reward is the predicted probability that the new graph first passes through the classification model to obtain the original graph type class(G n ). This is used as feedback to update the graph generator. Its calculation is as shown in formula (7), where g(·) consists of three layers of graph isomorphism networks and a softmax function, and l represents the total number of all possible classes of the classification model g(·); The final reward is for G n+1 Iterate through operations S2 to S4 until reaching the preset number of times m and then stop. Each result is recorded as Rollout i (G n+1 ), i = 1, 2....m; In each loop, when S4 is adding an edge, calculate the reward according to the intermediate reward formula (7), otherwise the reward value is 0; Average all the rewards during the loop to obtain the final reward R n,f-final : S5.2: Calculate the rule reward R n,r ; Perform a rollout on the generated new graph m (G n+1 ) to perform rule constraints: The first judgment rule is whether the new graph is a biological compound. If the corresponding molecular smiles can be found in the RDkit library, return 0; otherwise, return -1. Another rule is that the degree of a node cannot exceed the valence of the atom corresponding to the node. If satisfied, return 0; otherwise, return 1; S5.3: Reward R n It is expressed as the following formula (9), where β represents the rule-based reward weight; R n = R n,f + β × R n,r (9).
6. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, wherein The step S6 uses Monte Carlo policy gradient to train the generator, and the loss function in the nth iteration is as shown in formula (10), where L CE represents the cross-entropy loss function; Loss g = -R n (L CE (p n,start ,a n,start ) + L CE (p n,end ,a n,end )) (10).
7. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, wherein, In the step S7, the processes from S2 to S6 are repeated until the number of iterations reaches a preset value N, and then the new graphs generated under the condition that R t is greater than 0 and the original graphs where R t is less than 0 in all iterations are stored as a new graph data set.
8. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, characterized in that Step S8 includes the following steps: S8.1: Randomly divide the original compound dataset into a training set and a validation set according to 8:2; S8.2: Stack the feature matrices and adjacency matrices of the original graph dataset and the new graph dataset in the tensor depth direction respectively; stack the labels of the original graph dataset and the new graph dataset in the tensor horizontal direction.
9. The drug property prediction method based on reinforcement learning and molecular network data augmentation according to claim 1, characterized in that In step S9, take all the original training set data and the same number of data samples from the second half of the augmented training set to further train the 3-layer GIN classification model. At the same time, each time the model is trained, use the original validation set to verify the performance of the classification model; when the validation loss no longer shows a downward trend and only fluctuates within a small range, the early stopping method is used to stop the training.
10. A drug property prediction method based on reinforcement learning and molecular network data augmentation as claimed in claim 1, wherein, In step S10, apply a drug property prediction method, input the dataset in the field of chemical drugs into the trained classification model for classification of drug properties to achieve the drug prediction function, and compare it with the un-augmented graph classification model with the same configuration, and the prediction accuracy is significantly improved.