Drug data augmentation method and device, electronic equipment and storage medium

By processing the molecular structure diagrams of drug samples to generate enhanced samples, the overfitting problem caused by small sample drug datasets is solved, the accuracy and generalization ability of machine learning models are improved, and the drug dataset is expanded.

CN115240789BActive Publication Date: 2026-03-03INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

The high cost of drug development results in small sample sizes for drug datasets, leading to overfitting in machine learning models. This reduces the accuracy and generalization ability of the models, hindering their widespread application.

Method used

By processing the molecular structure diagrams of drug samples, including discarding atomic nodes and edges, or splicing them with the diagrams of other drug samples, enhanced samples are generated to expand the drug dataset.

Benefits of technology

It effectively alleviates the overfitting problem on small sample drug datasets, improves the generalization ability and robustness of machine learning models, enhances the accuracy of models, and facilitates their widespread application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240789B_ABST
    Figure CN115240789B_ABST
Patent Text Reader

Abstract

The application provides a drug data enhancement method and device, electronic equipment and storage medium. First, the molecular structure data of the target drug sample is obtained, and the molecular structure data is represented by a graph structure. Then, at least one of operation mode one, operation mode two and operation mode three is used to process the graph structure of the target drug sample to obtain an enhanced sample corresponding to the target drug sample. The method can perturb the graph structure of the target drug sample by deleting and / or adding operations, and the obtained enhanced sample can be used to increase the number and diversity of drug samples, expand the small sample drug data set, effectively alleviate the overfitting problem of the machine learning model on the small sample drug data set, and improve the generalization ability and robustness of the machine learning model. The accuracy of the machine learning model trained by the expanded drug data set is greatly improved, which is helpful for the wide application of the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data augmentation technology, and in particular to a method, apparatus, electronic device, and storage medium for drug data augmentation. Background Technology

[0002] In recent years, with the development of artificial intelligence, especially machine learning technology, it has been applied to solve problems in the field of drug design.

[0003] When applying machine learning techniques in drug design, it is typically necessary to acquire drug data and train a machine learning model based on specific needs. However, due to the high cost of experiments at each stage of drug development, the amount of drug data acquired is often too small, resulting in a small sample drug dataset. This can lead to overfitting in machine learning models trained on small sample datasets, reducing the accuracy of the machine learning model and hindering its widespread application.

[0004] Therefore, there is an urgent need to provide a method for enhancing drug data. Summary of the Invention

[0005] This invention provides a drug data augmentation method, apparatus, electronic device, and storage medium to address the deficiencies in the prior art.

[0006] This invention provides a drug data augmentation method, comprising:

[0007] Molecular structure data of a target drug sample is obtained and represented by a graph structure. The graph structure of the target drug sample includes multiple nodes and edges. The multiple nodes include supernodes corresponding to the molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule. The supernodes and atomic nodes are connected by the edges. The edges are also used to characterize the chemical bonds in the molecule.

[0008] The graph structure of the target drug sample is processed using at least one of the following methods to obtain an enhanced sample corresponding to the target drug sample:

[0009] Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample;

[0010] Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes;

[0011] Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0012] According to a drug data augmentation method provided by the present invention, the ratio of the number of atomic nodes corresponding to carbon atoms in the first number of atomic nodes to the first number is within a first preset range.

[0013] According to a drug data augmentation method provided by the present invention, the ratio of the number of backbone edges in the second number of edges to the second number is within a second preset range, wherein the backbone edges are edges that connect atomic nodes corresponding to two carbon atoms.

[0014] According to a drug data augmentation method provided by the present invention, the third operation mode includes:

[0015] The supernodes in the graph structure of the target drug sample and the graph structure of the other drug samples are merged into one.

[0016] According to a drug data augmentation method provided by the present invention, the label corresponding to the augmented sample obtained after processing the graph structure of the target drug sample using the third operation method is determined based on an OR operation between the label corresponding to the target drug sample and the labels corresponding to other drug samples.

[0017] According to a drug data augmentation method provided by the present invention, the target drug sample includes each drug sample in the drug dataset, and the other drug samples include drug samples in the drug dataset other than the target drug sample;

[0018] Accordingly, the first quantity and the second quantity are determined based on the average number of carbon atoms in the molecules in the drug dataset.

[0019] According to a drug data augmentation method provided by the present invention, the operation mode corresponding to each drug sample in the drug dataset includes operation mode one and operation mode three, and the ratio of the number of augmented samples corresponding to operation mode one to the number of augmented samples corresponding to operation mode three is within a third preset range.

[0020] The present invention also provides a drug data enhancement device, comprising:

[0021] The data acquisition module is used to acquire molecular structure data of the target drug sample and represent the molecular structure data through a graph structure. The graph structure of the target drug sample includes multiple nodes and edges. The multiple nodes include supernodes corresponding to the molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule. The supernodes and atomic nodes are connected by the edges. The edges are also used to characterize the chemical bonds in the molecule.

[0022] The data augmentation module is used to process the graph structure of the target drug sample using at least one of the following methods to obtain an augmented sample corresponding to the target drug sample:

[0023] Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample;

[0024] Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes;

[0025] Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the drug data augmentation method as described above.

[0027] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the drug data augmentation method as described above.

[0028] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the drug data augmentation method as described above.

[0029] The drug data augmentation method, apparatus, electronic device, and storage medium provided by this invention first acquire the molecular structure data of a target drug sample and represent the molecular structure data using a graph structure. Then, at least one of operation methods one, two, and three is used to process the graph structure of the target drug sample to obtain an augmented sample corresponding to the target drug sample. This method performs perturbation on the graph structure of the target drug sample by performing deletion and / or addition operations. The resulting augmented samples can be used to increase the number and diversity of drug samples, thereby expanding small sample drug datasets. This effectively alleviates the overfitting problem of machine learning models on small sample drug datasets and improves the generalization ability and robustness of machine learning models. The accuracy of machine learning models trained on the expanded drug dataset is greatly improved, which is conducive to the widespread application of machine learning models. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on the drawings described below without creative effort.

[0031] Figure 1 This is a flowchart illustrating the drug data augmentation method provided by the present invention;

[0032] Figure 2 This is a schematic diagram of the structure of the enhanced sample obtained using operation mode one in the drug data enhancement method provided by the present invention;

[0033] Figure 3 This is a schematic diagram of the structure of the enhanced sample obtained using operation mode two in the drug data enhancement method provided by the present invention;

[0034] Figure 4 This is a schematic diagram of the structure of the enhanced sample obtained using operation mode three in the drug data enhancement method provided by the present invention;

[0035] Figure 5 This is a schematic diagram of the structure of the enhanced sample obtained simultaneously using operation mode one, operation mode two, and operation mode three in the drug data enhancement method provided by the present invention;

[0036] Figure 6 This is a schematic diagram of the structure of the drug data enhancement device provided by the present invention;

[0037] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0039] When using machine learning techniques to model and analyze drug data, the small sample size of the drug data often leads to overfitting during the training process. Overfitting occurs when a machine learning model captures data features that only exist on the training set but not on the test set. This causes the machine learning model to perform significantly better on the training set than on the test set, reducing its generalization ability and robustness, as well as its accuracy, and hindering its widespread application.

[0040] There are two main approaches to mitigating overfitting in machine learning models caused by small sample datasets: regularization on the model side and augmentation on the data side. Regularization limits the number of parameters that function within the model, restricting its ability to fit features and thus preventing it from fitting noisy features. Augmentation on the data side involves performing augmentation operations on samples within the dataset and then adding the augmented samples back into the dataset.

[0041] Data augmentation methods introduce new noise into the dataset, reducing the impact of the original noise features on the machine learning model and improving the robustness of the machine learning model.

[0042] In the field of computer vision, image samples are often augmented through operations such as cropping, flipping, rotating, and scaling. Images are data with a regular Euclidean structure. This data augmentation method cannot be applied to drug data augmentation. Therefore, this invention provides a drug data augmentation method.

[0043] Figure 1 This is a flowchart illustrating the drug data augmentation method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0044] S1, Obtain the molecular structure data of the target drug sample and represent the molecular structure data through a graph structure; the graph structure of the target drug sample includes multiple nodes and edges, the multiple nodes include supernodes corresponding to the molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule, the supernodes and each atomic node are connected through the edges; the edges are also used to characterize the chemical bonds in the molecule;

[0045] S2, the graph structure of the target drug sample is processed using at least one of the following methods to obtain the enhanced sample corresponding to the target drug sample:

[0046] Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample;

[0047] Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes;

[0048] Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0049] Specifically, the drug data augmentation method provided in this embodiment of the invention is executed by a drug data augmentation device, which can be configured in a server. The server can be a local server or a cloud server. The local server can be a computer, etc., and this embodiment of the invention does not make specific limitations.

[0050] First, execute step S1 to obtain the molecular structure data of the target drug sample. The target drug sample can be any drug sample in the drug dataset, that is, each drug sample in the drug dataset can be used as the target drug sample to determine its corresponding enhancement sample.

[0051] Molecular structure data is used to characterize the internal and overall structure of a target drug sample molecule. It can include molecular information of the target drug sample, atomic information of each atom, and information on the connections between atoms.

[0052] In this embodiment of the invention, the molecular structure data of the target drug sample can be represented by a graph structure. The graph structure of the target drug sample can include multiple nodes and edges. The node types include supernodes and atomic nodes. Supernodes are used to represent molecules of the target drug sample, i.e., molecular nodes, and there is only one supernode. Atomic nodes are used to represent atoms in the molecules of the target drug sample, and there is a one-to-one correspondence between atomic nodes and atoms.

[0053] In the graph structure, the edge lines serve two purposes: firstly, to connect supernodes to atomic nodes, and secondly, to characterize chemical bonds in molecules.

[0054] The atomic information of each atom can be represented in the form of a graph-structured list of atoms, and the connection information between each atom can be represented in the form of a graph-structured adjacency matrix. Each row or column of the adjacency matrix stores the connection information between a node and other nodes except that node.

[0055] Then, step S2 is executed, employing at least one of the following operational methods to process the graph structure of the target drug sample, thereby obtaining the augmented sample corresponding to the target drug sample. It can be understood that the process of determining the augmented sample corresponding to the target drug sample is the process of data augmentation of the target drug sample.

[0056] In this embodiment of the invention, the data augmentation process may involve three operation methods: operation method one, operation method two, and operation method three. Each processing step may employ one operation method to process the graph structure of the target drug sample, resulting in one augmented sample corresponding to the target drug sample. Alternatively, two or three operation methods may be used simultaneously to process the graph structure of the target drug sample, resulting in two or three augmented samples corresponding to the target drug sample. Different processing steps may employ the same or two different operation methods to process the graph structure of the target drug sample, resulting in different augmented samples corresponding to the target drug sample.

[0057] Each augmented sample can be obtained by processing the graph structure of the target drug sample using one operation method, or by processing the graph structure of the target drug sample using two or three operations simultaneously. No specific limitation is made here.

[0058] Operation method one refers to discarding a first number of atomic nodes and the edges connected to them in the graph structure of the target drug sample. The first number can be set as needed, or a discard ratio can be set. The first number is the product of this discard ratio and the total number of atoms in the graph structure of the target drug sample; no specific limitation is made here. The discard ratio can be set as needed, for example, ranging from 10% to 30%, or determined based on the molecular size in the dataset containing the target drug sample; no specific limitation is made here.

[0059] The first number of atomic nodes can be randomly selected from the graph structure of the target drug sample.

[0060] In operation method one, not only do we need to discard some atomic nodes, but we also need to discard the edges connected to each discarded atom. This is a pruning operation.

[0061] Operation method two refers to discarding the second number of edges in the graph structure of the target drug sample, excluding the edges connected to supernodes. Multiple supernode edges can be discarded. In operation method two, supernode edges cannot be discarded; the other edges refer to the edges in the graph structure of the target drug sample excluding the edges connected to supernodes. This is also a type of deletion operation.

[0062] The second quantity can be set as needed, or a boundary discard ratio can be set. The product of this boundary discard ratio and the total number of all boundaries contained in the graph structure of the target drug sample is used as the second quantity. There is no specific limitation here. This boundary discard ratio can be set as needed, for example, the value range can be 10%-30%.

[0063] The second number of borders can be randomly selected from the graph structure of the target drug sample.

[0064] In operation mode two, only some other edge lines need to be discarded, and the atomic nodes connected to both ends of each discarded edge line are not processed.

[0065] Operation method three refers to concatenating the graph structure of the target drug sample with the graph structures of other drug samples. "Other drug samples" refers to drug samples other than the target drug sample. The concatenation method can be set as needed and is not specifically limited here. For example, a shared node can be established between the graph structure of the target drug sample and the graph structures of other drug samples, and the supernodes in the original two graph structures can be deleted, with the shared node becoming the supernode, thus achieving concatenation. Alternatively, the supernodes in the original two graph structures can be directly merged to obtain a single supernode.

[0066] Operation method three involves splicing two graph structures, which can be understood as adding the graph structure of other drug samples to the graph structure of the target drug sample, and is therefore an additive operation.

[0067] After obtaining the augmented samples, the target drug samples can be combined with the augmented samples to train the deep learning model.

[0068] The drug data augmentation method provided in this embodiment of the invention first acquires the molecular structure data of the target drug sample and represents the molecular structure data using a graph structure. Then, it processes the graph structure of the target drug sample using at least one of operation methods one, two, and three to obtain augmented samples corresponding to the target drug sample. This method performs perturbation on the graph structure of the target drug sample by performing deletion and / or addition operations. The resulting augmented samples can be used to increase the number and diversity of drug samples, thereby expanding small sample drug datasets. This effectively alleviates the overfitting problem of machine learning models on small sample drug datasets and improves the generalization ability and robustness of machine learning models. The accuracy of machine learning models trained on the expanded drug dataset is greatly improved, which is conducive to the widespread application of machine learning models.

[0069] Based on the above embodiments, in the drug data augmentation method provided in the embodiments of the present invention, the ratio of the number of atomic nodes corresponding to carbon atoms in the first number of atomic nodes to the first number is within a first preset range.

[0070] Specifically, in this embodiment of the invention, since carbon atoms are typically the backbone atoms in a molecule, in operation mode one, it is necessary to ensure that the ratio of the number of atomic nodes corresponding to carbon atoms in the first number of discarded atomic nodes to the first number is within a first preset range. That is, it is necessary to ensure that the discarded atomic nodes have a certain number of atomic nodes corresponding to carbon atoms. The first preset range can be set as needed, for example, it can be set to 0.2-0.4.

[0071] In this embodiment of the invention, by limiting the number of atomic nodes corresponding to carbon atoms in the first number of discarded atomic nodes, the first operation method can be made more reasonable.

[0072] Based on the above embodiments, in the drug data augmentation method provided in the embodiments of the present invention, the ratio of the number of backbone edges in the second number of edges to the second number is within a second preset range, and the backbone edges are edges that connect atomic nodes corresponding to two carbon atoms.

[0073] Specifically, in this embodiment of the invention, in operation mode two, it is also necessary to ensure that the ratio of the number of backbone edges to the second number of discarded edges is within a second preset range, that is, it is necessary to ensure that a certain number of backbone edges are among the discarded edges. The backbone edge refers to the edge connecting atomic nodes corresponding to two carbon atoms. The second preset range can be set as needed, for example, it can be set to 0.4-0.6.

[0074] In this embodiment of the invention, by limiting the number of backbone edges in the second number of discarded edges, the second operation method can be made more reasonable.

[0075] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention, wherein the third operation mode includes:

[0076] The supernodes in the graph structure of the target drug sample and the graph structure of the other drug samples are merged into one.

[0077] Specifically, in the embodiments of the present invention, the splicing method in operation mode three can be to directly merge the super nodes in two graph structures into one. The merging method can be to directly retain the super node in either graph structure, use the retained super node as the merged super node, discard the super node in the other graph structure, connect all atomic nodes connected to the discarded super node to the retained super node, and at the same time keep the connection relationship between all atomic nodes connected to the discarded super node unchanged.

[0078] In this embodiment of the invention, the splicing of two graph structures is achieved by merging supernodes, which simplifies the operation.

[0079] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention adopts the third operation mode, and the label corresponding to the augmented sample obtained after processing the graph structure of the target drug sample is determined by performing an OR operation on the label corresponding to the target drug sample and the labels corresponding to other drug samples.

[0080] Specifically, in this embodiment of the invention, the purpose of determining augmented samples is to expand the drug dataset, therefore, the labels corresponding to the augmented samples need to be considered. For augmented samples obtained through operation method one and operation method two, since no graph structure of other drug samples was introduced during the processing, the label corresponding to the target sample can be used as the label corresponding to its augmented sample. However, for augmented samples obtained through operation method three, since the graph structure of other drug samples was introduced during the processing, it is necessary to redetermine the labels corresponding to the augmented samples. The determination method can be to perform an OR operation on the label corresponding to the target drug sample and the labels corresponding to other drug samples, and then use the result of the OR operation as the label corresponding to the augmented sample.

[0081] Taking a binary classification label as an example, the binary classification label can take the value 0 or 1. Therefore, during the OR operation, if at least one of the two labels is 1, the label corresponding to the augmented sample is 1. When both labels are 0, the label corresponding to the augmented sample is 0.

[0082] In this embodiment of the invention, a method for determining the labels of augmented samples is provided, so that the augmented samples can be used for supervised training of deep learning models.

[0083] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention includes a target drug sample comprising each drug sample in the drug dataset, and other drug samples comprising drug samples in the drug dataset other than the target drug sample;

[0084] Accordingly, the first quantity is determined based on the average number of carbon atoms in the molecules in the drug dataset.

[0085] Specifically, in this embodiment of the invention, the target drug sample can be any drug sample in the drug dataset, and other drug samples can be any drug samples in the drug dataset other than the target drug sample. Furthermore, the first quantity can be determined by the average number of carbon atoms in the molecules of the drug dataset. The average number of carbon atoms in the molecules of the drug dataset can be the ratio of the total number of carbon atoms in the molecules of all drug samples in the drug dataset to the total number of atoms in the molecules of all drug samples.

[0086] The first quantity can be inversely proportional to the average number of carbon atoms in the molecules of the drug dataset. That is, the larger the average number of carbon atoms in the molecules of the drug dataset, the smaller the first quantity and the smaller the proportion of discarded atomic nodes. Conversely, the larger the first quantity, the larger the proportion of discarded atomic nodes. For example, for drug datasets with an average number of carbon atoms in molecules below 15, a slightly larger proportion of discarded atomic nodes, such as 30%, can be used.

[0087] Furthermore, after obtaining the augmented samples, they can be added to the drug dataset to increase the number and diversity of samples in the drug dataset. This allows the drug dataset to be used to train deep learning models, effectively alleviating the overfitting problem of machine learning models on small sample drug datasets.

[0088] In this embodiment of the invention, when discarding atomic nodes, the average number of carbon atoms in the molecules of the drug dataset is taken into account. This ensures that the average number of carbon atoms in the molecules of the entire expanded drug dataset is not much different from that before expansion, thus guaranteeing the performance of the drug dataset.

[0089] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention includes operation mode one and operation mode three for each drug sample in the drug dataset, and the ratio of the number of augmented samples corresponding to operation mode one to the number of augmented samples corresponding to operation mode three is within a third preset range.

[0090] Specifically, in this embodiment of the invention, when performing data augmentation on each drug sample in the drug dataset to obtain the corresponding augmented sample, at least operation mode one and operation mode three can be used to process the graph structure of different drug samples.

[0091] The ratio between the number of augmented samples corresponding to Operation Mode 1 and the number of augmented samples corresponding to Operation Mode 3 can be guaranteed to be within a third preset range, which can be between 0.25 and 0.5. This ratio can be selected as 1 / 3.

[0092] In this embodiment of the invention, when multiple operation methods are mixed, the ratio between the number of augmented samples corresponding to different operation methods can be set, so as to ensure the balance of augmented samples.

[0093] Based on the above embodiments, the complete process of the drug data augmentation method provided in the embodiments of the present invention may include:

[0094] Step 1: For each drug sample in the drug dataset, transform the molecular structure of each drug sample into a graph structure. The graph structure contains a list of nodes and a list of edges. The node list includes supernodes and atomic nodes, and the edge list includes multiple edges representing the connections between nodes. Supernodes are connected to all atomic nodes in the graph structure. The edge list of this graph structure can be stored in the form of an adjacency matrix, where each row or column stores the connections between a node and other nodes.

[0095] Step 2: Randomly select 10% of the atoms within the molecule of the target drug sample and discard their corresponding atomic nodes. During this process, the edges connected to those atomic nodes are also discarded. Supernodes are not selected or deleted during this process. Adjust the proportion of carbon atom-corresponding atomic nodes in the selected discarded atomic nodes to 0.3. Repeat this process twice, treating all drug samples in the drug dataset as target drug samples.

[0096] like Figure 2 The diagram shown is a schematic representation of the structure of the enhanced sample obtained using operation method one in an embodiment of the present invention. Figure 2 The left-hand diagram represents the target drug sample, consisting of eight nodes numbered 1-8, with node 8 being a supernode. Node 6 on the left is discarded, and the labels of subsequent nodes are adjusted accordingly to obtain the right-hand diagram.

[0097] Figure 2 The list of edges corresponding to the left-hand structure is shown in Table 1. Figure 2 The edge list corresponding to the right side structure is shown in Table 2. In Tables 1 and 2, 0 indicates that there is no connection between two nodes, and 1 indicates that there is a connection between two nodes.

[0098] Table 1 Figure 2 The edge list corresponding to the left side structure

[0099]

[0100] Table 2 Figure 2 The edge list corresponding to the right side structure

[0101]

[0102] Step 3: Randomly select 10% of the intramolecular edges of the target drug sample and discard them. Edges of supernodes connected to supernodes will not be selected or deleted during this process. Among the selected edges to be discarded, adjust the proportion of the backbone edges between carbon atoms to 0.5. Perform the above operation twice on each drug sample in the drug dataset, treating them as target drug samples.

[0103] like Figure 3 The diagram shown is a schematic representation of the enhanced sample obtained using operation method two in an embodiment of the present invention. Figure 3 The left-hand diagram represents the target drug sample, consisting of eight nodes numbered 1-8, with node 8 being a supernode. The boundary between node 4 and node 7 on the left is discarded to obtain the right-hand diagram.

[0104] Figure 3 The list of edges corresponding to the left-hand structure is shown in Table 3. Figure 3 The edge list corresponding to the right side structure is shown in Table 4. In Tables 3 and 4, 0 indicates that two nodes are not connected, and 1 indicates that two nodes are connected.

[0105] Table 3 Figure 3 The edge list corresponding to the left side structure

[0106]

[0107] Table 4 Figure 3 The edge list corresponding to the right side structure

[0108]

[0109] Step 4: For drug samples other than the target drug sample in the drug dataset, randomly select another drug sample and integrate the graph structure of the other drug sample with the graph structure of the target drug sample through supernodes to form a new augmented sample.

[0110] First, create a supernode for the graph structure of the target drug sample and connect all atomic nodes in the graph structure to this supernode. Then, randomly select another drug sample and connect all atomic nodes in the graph structure of that other drug sample to this supernode.

[0111] Because the graph structures of different drug samples are merged, the labels of the augmented samples need to be updated by combining the labels of the two drug samples. For a binary classification task, an OR operation is used to obtain the labels of the augmented samples. If at least one of the two integrated drug samples has a label of 1, the label of the augmented sample is 1; otherwise, it is 0.

[0112] The above operation was performed twice, treating all drug samples in the drug dataset as target drug samples.

[0113] like Figure 4 The diagram shown is a schematic representation of the enhanced sample obtained using operation method three in an embodiment of the present invention. Figure 4 The top left side of the diagram represents the target drug sample, consisting of seven nodes numbered 1-7, with node 7 being a supernode. Figure 4 The upper right diagram shows other drug samples, including eight nodes numbered 1-8, where node 8 is a supernode. Concatenating these nodes yields... Figure 4 The lower middle side view structure is an enhanced sample. Figure 4 In the lower middle side diagram structure, 7 is a supernode.

[0114] Step 5: Instead of using a single additive or subtractive operation on each drug sample within the drug dataset, the diversity of augmented samples is improved by combining different operation methods. A single augmented sample may originate from only one augmentation method, thus not increasing the diversity of that individual sample. However, by employing different operation methods, the overall diversity of the augmented samples is improved. Here, operation method one and operation method three are combined, with the ratio of the number of operations for operation method one to operation method three being 1:3, meaning that 0.25% of the augmented samples come from operation method one. The above operation method is performed 7 times on all drug samples in the drug dataset.

[0115] like Figure 5 The diagram shown is a structural schematic of the enhanced sample obtained simultaneously using operation mode one, operation mode two, and operation mode three in an embodiment of the present invention. Figure 5 The upper-middle graph structure represents the target drug sample, comprising seven nodes numbered 1-7, with node 7 being a supernode. The edge between nodes 3 and 5 in the target drug sample's graph structure is discarded, resulting in... Figure 5 The top left side of the graph represents the first enhanced sample, corresponding to operation mode two; node 4 and its connected edges in the graph structure of the target drug sample are discarded, resulting in... Figure 5 The upper right image shows the second enhanced sample, corresponding to operation mode one; the graph structure of the target drug sample is stitched together with the graph structures of other drug samples to obtain... Figure 5 The lower middle side view structure represents the third enhanced sample, corresponding to operation mode three.

[0116] Step 6: Select one or more steps from Steps 2 to 5 to perform data augmentation on the drug samples in the drug dataset and add the augmented samples to the drug dataset. Multiple operations can be performed on a single drug sample. A machine learning model is then trained on the augmented drug dataset.

[0117] To test the actual performance of the drug data augmentation method provided in this embodiment of the invention, baseline models were trained on both the original drug dataset and the augmented drug dataset, and the differences in the actual effects brought about by data augmentation were compared. Five different drug attribute prediction datasets were selected as drug datasets, namely BACE, BBBP, Tox21, ToxCast, and Clintox from MoleculeNet. Detailed information for each drug attribute prediction dataset is shown in Table 5 below:

[0118] Table 5. Detailed information for each drug attribute prediction dataset.

[0119]

[0120]

[0121] First, the baseline model MG-BERT was selected. The MG-BERT model uses a Transformer encoder as its core, and each drug sample is fed into MG-BERT in the form of a graph structure for encoding. MG-BERT takes the list of atoms in the molecule and the adjacency matrix storing the connections between these atoms as input, and introduces supernodes into the graph structure corresponding to the molecule. After MG-BERT completes the encoding, the output of the supernodes serves as the feature representation of the molecule. The MG-BERT used here has 6 layers, 4 attention heads, a feature dimension of 128, a dropout regularization retention ratio of 0.9, a batch size of 64, a maximum sequence length of 150, and uses the Adam optimizer with a learning rate of 5e-5. The parameters of the baseline model are kept consistent, and the code framework uses Tensorflow. The operating system used is Ubuntu, the CPU is an Intel(R) Xeon(R) Gold6226R CPU@2.90GHz, the GPU is an NVIDIA RTX 2080Ti (11GB), the Python version is 3.7, and the Tensorflow version is 2.4.1.

[0122] Table 6 shows that each drug attribute prediction dataset contains one or more binary classification tasks. For a single binary classification task, AUC (Area Under ROC Curve) can be used as the evaluation metric. However, since some drug attribute prediction datasets contain more than one binary classification task, the average AUC of different binary classification tasks on the same drug attribute prediction dataset, as proposed in MoleculeNet, is used as the performance of MG-BERT on that dataset. Based on the above settings, the experimental results of each operation method are shown in Table 6 below:

[0123] Table 6. Experimental Results of Data Augmentation Methods for Various Drugs

[0124]

[0125]

[0126] As can be seen, different manipulation methods significantly improve upon the baseline model on most datasets. This is due to the introduction of noise into the dataset by each manipulation method, which enhances the robustness of the baseline model during training. However, the effectiveness of different manipulation methods varies considerably across different datasets due to the different types of noise introduced. For example, manipulation method three achieved good enhancement results on both BBBP and ToxCast, while manipulation method two only showed some advantage on the Tox21 dataset. The combination of manipulation methods one, two, and three improved the diversity of the enhanced samples, achieving the best results on the BACE, BBBP, and Clintox datasets, with improvements of 6.5%, 1.9%, and 12.5% ​​respectively compared to the baseline model. This indicates that for small-sample drug datasets, mixing different manipulation methods helps the model learn more accurate task-relevant features.

[0127] like Figure 6 As shown, based on the above embodiments, this embodiment of the invention provides a drug data enhancement device, including:

[0128] The data acquisition module 61 is used to acquire molecular structure data of the target drug sample and represent the molecular structure data through a graph structure. The graph structure of the target drug sample includes multiple nodes and edges. The multiple nodes include supernodes corresponding to the molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule. The supernodes and atomic nodes are connected through the edges. The edges are also used to characterize the chemical bonds in the molecule.

[0129] Data augmentation module 62 is used to process the graph structure of the target drug sample using at least one of the following methods to obtain an augmented sample corresponding to the target drug sample:

[0130] Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample;

[0131] Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes;

[0132] Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0133] Based on the above embodiments, in the drug data enhancement device provided in the embodiments of the present invention, the ratio of the number of atomic nodes corresponding to carbon atoms in the first number of atomic nodes to the first number is within a first preset range.

[0134] Based on the above embodiments, in the drug data augmentation method provided in the embodiments of the present invention, the ratio of the number of backbone edges in the second number of edges to the second number is within a second preset range, and the backbone edges are edges that connect atomic nodes corresponding to two carbon atoms.

[0135] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention, wherein the third operation mode includes:

[0136] The supernodes in the graph structure of the target drug sample and the graph structure of the other drug samples are merged into one.

[0137] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention adopts the third operation mode, and the label corresponding to the augmented sample obtained after processing the graph structure of the target drug sample is determined by performing an OR operation on the label corresponding to the target drug sample and the labels corresponding to other drug samples.

[0138] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention includes a target drug sample comprising each drug sample in the drug dataset, and other drug samples comprising drug samples in the drug dataset other than the target drug sample;

[0139] Accordingly, the first quantity is determined based on the average number of carbon atoms in the molecules in the drug dataset.

[0140] Based on the above embodiments, the drug data augmentation method provided in this embodiment of the invention includes operation mode one and operation mode three for each drug sample in the drug dataset, and the ratio of the number of augmented samples corresponding to operation mode one to the number of augmented samples corresponding to operation mode three is within a third preset range.

[0141] Specifically, the functions of each module in the drug data enhancement device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above method-like embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0142] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communications bus 740. The processor 710 can call logic instructions in the memory 730 to execute the drug data augmentation method provided in the above embodiments. The method includes: acquiring molecular structure data of a target drug sample and representing the molecular structure data through a graph structure; the graph structure of the target drug sample includes multiple nodes and edges, the multiple nodes include supernodes corresponding to molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule, the supernodes and each atomic node are connected through the edges; the edges are also used to characterize chemical bonds in the molecule; the graph structure of the target drug sample is processed using at least one of the following operation methods to obtain an augmented sample corresponding to the target drug sample: operation method one: discarding a first number of atomic nodes and edges connected to the first number of atomic nodes in the graph structure of the target drug sample; operation method two: discarding a second number of edges other than supernode edges in the graph structure of the target drug sample, the supernode edges being edges connected to the supernodes; operation method three: splicing the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0143] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0144] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the drug data augmentation method provided by the above methods. The method includes: acquiring molecular structure data of a target drug sample and representing the molecular structure data through a graph structure; the graph structure of the target drug sample includes multiple nodes and edges, the multiple nodes including supernodes corresponding to molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule, and the supernodes and atomic nodes are connected through the edges. The edge lines are also used to characterize the chemical bonds in the molecule; the graph structure of the target drug sample is processed using at least one of the following methods to obtain an enhanced sample corresponding to the target drug sample: Method 1: Discard a first number of atomic nodes and edge lines connected to the first number of atomic nodes in the graph structure of the target drug sample; Method 2: Discard a second number of edge lines other than supernode edge lines in the graph structure of the target drug sample, where the supernode edge lines are edge lines connected to the supernodes; Method 3: Concatenate the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0145] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the drug data augmentation method provided by the methods described above. The method includes: acquiring molecular structure data of a target drug sample and representing the molecular structure data using a graph structure; the graph structure of the target drug sample includes multiple nodes and edges, the multiple nodes including supernodes corresponding to molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule, the supernodes and atomic nodes being connected through the edges; the edges are also used to characterize the molecules... The chemical bonds; the graph structure of the target drug sample is processed using at least one of the following methods to obtain an enhanced sample corresponding to the target drug sample: Method 1: Discard a first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample; Method 2: Discard a second number of edges other than the supernode edges in the graph structure of the target drug sample, where the supernode edges are the edges connected to the supernodes; Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for enhancing drug data, characterized in that, include: Obtain the molecular structure data of the target drug sample and represent the molecular structure data using a graph structure; The molecular structure data includes the molecular information of the target drug sample, as well as the atomic information of each atom and the connection relationship between the atoms. The graph structure of the target drug sample includes multiple nodes and edge lines. The multiple nodes include supernodes corresponding to molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule. The supernodes are molecular nodes, and the supernodes are connected to each atomic node through the edge lines. The edge lines are also used to characterize the chemical bonds in the molecule. The graph structure of the target drug sample is processed using at least one of the following methods to obtain an enhanced sample corresponding to the target drug sample: Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample; Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes; Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

2. The drug data augmentation method according to claim 1, characterized in that, The ratio of the number of atomic nodes corresponding to carbon atoms in the first number of atomic nodes to the first number is within a first preset range.

3. The drug data augmentation method according to claim 1, characterized in that, The ratio of the number of backbone edges in the second number of edges to the second number is within a second preset range. The backbone edges are edges that connect atomic nodes corresponding to two carbon atoms.

4. The drug data augmentation method according to claim 1, characterized in that, The third operation method includes: The supernodes in the graph structure of the target drug sample and the graph structure of the other drug samples are merged into one.

5. The drug data augmentation method according to claim 1, characterized in that, Using the third operation method, the label corresponding to the enhanced sample obtained after processing the graph structure of the target drug sample is determined by performing an OR operation on the label corresponding to the target drug sample and the labels corresponding to the other drug samples.

6. The drug data augmentation method according to any one of claims 1-5, characterized in that, The target drug sample includes each drug sample in the drug dataset, and the other drug samples include drug samples in the drug dataset other than the target drug sample; Accordingly, the first quantity is determined based on the average number of carbon atoms in the molecules in the drug dataset.

7. The drug data augmentation method according to claim 6, characterized in that, The operation methods corresponding to each drug sample in the drug dataset include operation method one and operation method three, and the ratio of the number of enhanced samples corresponding to operation method one to the number of enhanced samples corresponding to operation method three is within a third preset range.

8. A drug data enhancement device, characterized in that, include: The data acquisition module is used to acquire the molecular structure data of the target drug sample and represent the molecular structure data through a graph structure. The molecular structure data includes the molecular information of the target drug sample, as well as the atomic information of each atom and the connection relationship between the atoms. The graph structure of the target drug sample includes multiple nodes and edge lines. The multiple nodes include supernodes corresponding to molecules of the target drug sample and atomic nodes corresponding to each atom in the molecule. The supernodes are molecular nodes, and the supernodes are connected to each atomic node through the edge lines. The edge lines are also used to characterize the chemical bonds in the molecule. The data augmentation module is used to process the graph structure of the target drug sample using at least one of the following methods to obtain an augmented sample corresponding to the target drug sample: Operation Method 1: Discard the first number of atomic nodes and the edges connected to the first number of atomic nodes in the graph structure of the target drug sample; Operation Method 2: Discard the second number of edges in the graph structure of the target drug sample, excluding the edges of the supernodes, where the supernode edges are the edges connected to the supernodes; Operation Method 3: Combine the graph structure of the target drug sample with the graph structures of other drug samples besides the target drug sample.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the drug data augmentation method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the drug data augmentation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Drug small molecule property prediction method, device and equipment based on self-supervised learning

    CN113707235A

  • Machine learning-based drug recommendation method, device, equipment and medium

    CN113707264A