Drug property prediction method, device, electronic device and storage medium

Through molecular attribute prediction, atomic attribute prediction and molecular comparison learning synchronous training drug pre-training model, the problems of difficult migration and low prediction accuracy of drug pre-training model are solved, and more efficient drug attribute prediction is achieved.

CN115240782BActive Publication Date: 2025-08-12INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210719727.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-23
Publication Date
2025-08-12
Estimated Expiration
2042-06-23

AI Technical Summary

Technical Problem

Among the existing drug pre-training methods, the task difference between the pre-training stage and the fine-tuning stage makes it difficult to transfer the drug pre-training model, the construction process takes a long time and the prediction accuracy is not high, which affects the widespread application of drug prediction models.

Method used

Three pre-training methods: molecular attribute prediction, atomic attribute prediction and molecular comparison learning are used to synchronize the training of drug pre-training models. Through the molecular structure data of drug samples and atomic attribute labels, combined with supervised and unsupervised training, the differences between pre-training and fine-tuning tasks are narrowed, and the model performance and generalization are improved.

Benefits of technology

It reduces the migration difficulty of drug pre-trained models, shortens the construction time, and improves the prediction accuracy of the fine-tuning stage, providing guarantees for the widespread application of drug prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240782B_ABST
    Figure CN115240782B_ABST
Patent Text Reader

Abstract

The present invention provides a drug property prediction method, apparatus, electronic device, and storage medium. The method first obtains target molecular structure data for the drug to be predicted; then, the target molecular structure data is input into a drug property prediction model to obtain drug property information for the drug to be predicted. The drug property prediction model employed is obtained by fine-tuning a pre-trained drug model, which is obtained through simultaneous training of molecular property prediction pre-training, atomic property prediction pre-training, and molecular comparative learning pre-training. This ensures the performance and generalizability of the pre-trained drug model in drug property prediction tasks, reduces the difficulty of migrating the pre-trained drug model, shortens the construction time of the drug prediction model, and improves the prediction accuracy of the drug prediction model obtained during the fine-tuning phase, thus ensuring its widespread application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug property prediction, and in particular to a drug property prediction method, device, electronic equipment and storage medium. Background Art

[0002] With the recent development of artificial intelligence (AI), this technology has been widely applied across all stages of drug development, including target identification, drug design, drug repositioning, and biomedical information analysis. With its powerful data analysis and modeling capabilities, AI offers new solutions to the inefficiencies and uncertainties inherent in traditional drug development methods, while also reducing bias and human intervention in the process.

[0003] Machine learning is the most commonly used technology in the field of artificial intelligence. However, when incorporating machine learning into drug property prediction, a common problem arises: the small sample size problem. This refers to the small number of samples in relevant datasets due to the high experimental costs at each stage of drug development. Furthermore, because drug sample data contains the structure and composition of drug molecules, the data is high in dimensionality. High-dimensional drug feature data and a small sample size pose challenges for parameter optimization in deep learning models.

[0004] Drug pretraining is an important approach to addressing the limited sample size problem in drug-related tasks. Pretraining involves first training a model on a large amount of unlabeled data to complete a predefined self-supervised training task, enhancing the model's drug representation capabilities. After pretraining, the model is then applied to drug property prediction tasks, a phase known as fine-tuning. Essentially, pretraining provides domain-specific parameter initialization for the fine-tuning phase by completing the pretraining task.

[0005] In existing drug pre-training methods, since there are often large differences in the tasks completed in the pre-training stage and the fine-tuning stage, this difference between tasks determines that the drug pre-training model obtained in the pre-training stage needs to undergo a large number of parameter updates to adapt to the fine-tuning tasks when migrating to the fine-tuning stage. This not only increases the difficulty of migrating the drug pre-training model, which in turn causes the construction process of the drug prediction model to take too long, but also reduces the prediction accuracy of the drug prediction model obtained in the fine-tuning stage, which is not conducive to the widespread application of drug prediction models. Summary of the Invention

[0006] The present invention provides a drug property prediction method, device, electronic device and storage medium to solve the defects in the prior art.

[0007] The present invention provides a method for predicting drug properties, comprising:

[0008] Obtain target molecular structure data of the drug to be predicted;

[0009] Inputting the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted;

[0010] The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method:

[0011] Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated;

[0012] performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules;

[0013] Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

[0014] According to a drug property prediction method provided by the present invention, the drug pre-training model is trained based on the following steps:

[0015] Calculating the atomic property prediction loss and the molecular property prediction loss of the initial model based on the molecular property labels of the original molecules, the atomic property labels of the target atoms in the original molecules, and the second type of sample molecular structure data of the defective molecules;

[0016] Calculating the contrastive learning loss of the initial model based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule;

[0017] Based on the atomic property prediction loss, the molecular property prediction loss and the contrastive learning loss, the overall loss of the initial model is determined, and based on the overall loss, the initial model is trained to obtain the drug pre-training model.

[0018] According to a drug property prediction method provided by the present invention, the calculation of the contrastive learning loss of the initial model based on the second-category sample molecular structure data of each defective molecule and the third-category sample molecular structure data of each original molecule includes:

[0019] Constructing positive sample pairs and negative sample pairs based on the defective molecules and the original molecules;

[0020] Inputting each structural data corresponding to the positive sample pair into the initial model to obtain each representation vector corresponding to the positive sample pair output by the initial model;

[0021] Inputting each structural data corresponding to the negative sample pair into the initial model to obtain each representation vector corresponding to the negative sample pair output by the initial model;

[0022] Based on the similarity between each representation vector corresponding to the positive sample pair and the similarity between each representation vector corresponding to the negative sample pair, a contrastive learning loss between the positive sample pair and the negative sample pair is calculated.

[0023] According to a drug property prediction method provided by the present invention, constructing positive sample pairs and negative sample pairs based on the defective molecules and the original molecules includes:

[0024] For a target original molecule, constructing a positive sample pair based on the target original molecule and a defective molecule corresponding to the target original molecule;

[0025] A negative sample pair is constructed based on the target original molecule and a defective molecule corresponding to other original molecules among the original molecules except the target original molecule.

[0026] According to a drug property prediction method provided by the present invention, the target molecular structure data, the first category sample molecular structure data, the second category sample molecular structure data and the third category sample molecular structure data are all represented based on a molecular graph structure, and the molecular graph structure includes molecular-level nodes and atomic-level nodes, and the molecular-level nodes are connected to the atomic-level nodes.

[0027] According to a drug property prediction method provided by the present invention, the molecular property labels of each original molecule include regression task labels, and the atomic property labels of the target atoms include classification task labels.

[0028] According to a drug attribute prediction method provided by the present invention, the initial model includes a model constructed based on a Transformer encoder.

[0029] The present invention also provides a drug property prediction device, comprising:

[0030] A data acquisition module is used to obtain target molecular structure data of the drug to be predicted;

[0031] A prediction module, configured to input the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted;

[0032] The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method:

[0033] Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated;

[0034] performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules;

[0035] Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

[0036] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for predicting drug properties as described above is implemented.

[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for predicting drug properties.

[0038] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-described methods for predicting drug properties.

[0039] The drug property prediction method, device, electronic device, and storage medium provided by the present invention first obtain target molecular structure data for the drug to be predicted; then, the target molecular structure data is input into a drug property prediction model to obtain drug property information for the drug to be predicted. The drug property prediction model employed is obtained by fine-tuning a pre-trained drug model, which is obtained through simultaneous training of molecular property prediction pre-training, atomic property prediction pre-training, and molecular comparative learning pre-training. This ensures the performance and generalizability of the pre-trained drug model in drug property prediction tasks, reduces the difficulty of migrating the pre-trained drug model, shortens the time required to build the drug prediction model, and improves the prediction accuracy of the drug prediction model obtained during the fine-tuning phase, thus ensuring its widespread application. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on the drawings in the following description without any creative work.

[0041] Figure 1 Schematic diagram of the process of drug property prediction method provided by the present invention;

[0042] Figure 2 Schematic diagram of the operation flow of three pre-training methods in the drug property prediction method provided by the present invention;

[0043] Figure 3 Schematic diagram of the structure of the drug property prediction device provided by the present invention;

[0044] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0045] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0046] In existing drug pre-training methods, the tasks completed in the pre-training stage and the fine-tuning stage often differ significantly. This difference in tasks means that the drug pre-training model obtained in the pre-training stage needs to undergo a large number of parameter updates to adapt to the fine-tuning tasks when migrating to the fine-tuning stage. This not only increases the difficulty of migrating the drug pre-training model, which in turn leads to a time-consuming process of building a drug prediction model, but also reduces the prediction accuracy of the drug prediction model obtained in the fine-tuning stage, which is not conducive to the widespread application of drug prediction models. To this end, an embodiment of the present invention provides a drug property prediction method.

[0047] Figure 1 FIG. 1 is a flow chart of a method for predicting drug properties according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0048] S1, obtaining the target molecular structure data of the drug to be predicted;

[0049] S2, inputting the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted;

[0050] The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method:

[0051] Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated;

[0052] performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules;

[0053] Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

[0054] Specifically, the drug property prediction method provided in the embodiment of the present invention is executed by a drug property prediction device, which can be configured in a server. The server can be a local server or a cloud server. The local server can specifically be a computer, etc., and this is not specifically limited in the embodiment of the present invention.

[0055] First, step S1 is executed to obtain the target molecular structure data of the drug to be predicted. The drug to be predicted refers to the drug whose drug properties need to be determined. Drug properties may include the four properties, five flavors, meridians, efficacy, etc., and can be divided into three categories: medicinal properties, medicinal flavors, and meridians, which are not specifically limited here. The target molecular structure data refers to the molecular structure data of the drug to be predicted, which is used to characterize the internal structure and overall structure of the drug molecule to be predicted. It can include molecular information of the drug molecule to be predicted, atomic information of each atom within the molecule, and information on the connection relationships between atoms.

[0056] Then, step S2 is executed to input the target molecular structure data into the drug property prediction model. The drug property prediction model extracts and analyzes the target molecular structure data to obtain and output drug property information of the drug to be predicted. The drug property information is the relevant information used to characterize the drug properties of the drug to be predicted.

[0057] In an embodiment of the present invention, the drug property prediction model used can be obtained by training the first-category sample molecular structure data of drug samples carrying drug property labels on the basis of the drug pre-training model. The training process of the drug pre-training model can be understood as a fine-tuning process. In this process, the first-category sample molecular structure data of the drug sample can be input into the drug pre-training model, and the drug pre-training model outputs the prediction results corresponding to the drug sample. The loss is calculated based on the prediction results and the drug property labels, and the drug pre-training model is trained based on the loss to obtain the drug property prediction model.

[0058] It is understood that the first type of sample molecular structure data refers to the molecular structure data of the drug sample, which is used to characterize the internal structure and overall structure of the drug sample molecule, and may include molecular information of the drug sample molecule, atomic information of each atom within it, and information about the connection relationship between each atom. The drug attribute label refers to the drug attribute of the drug sample, and may include the drug properties, medicinal flavor, and meridians of the drug sample, etc., which are not specifically limited here. The atomic information of each atom can be represented in the form of an atom list, and the connection relationship information between each atom can be represented in the form of an adjacency matrix.

[0059] The drug pre-training model can be obtained by synchronously training the initial model using the following pre-training methods. The initial model can be a neural network model, for example, a model based on a convolutional neural network (CNN) such as ResNet and Inception, a Transformer, a segmentation network, etc. The pre-training methods include molecular property prediction pre-training, atomic property prediction pre-training, and molecular comparative learning pre-training.

[0060] Molecular property prediction pre-training refers to the process of training the initial model by using the second-category sample molecular structure data of the defective molecules of each compound sample as the input of the initial model, and the molecular property labels of the original molecules of each compound sample as the object for calculating the loss with the output of the initial model.

[0061] Each compound sample can be taken from an existing compound database, such as the ZINC database. In an embodiment of the present invention, 7 million compound samples can be selected from the ZINC database. Each compound sample can be represented in molecular form, that is, the original molecule of each compound sample. Here, the original molecule refers to a molecule that has not been defect-treated, which is mainly used to distinguish it from a defective molecule. Correspondingly, the defective molecule can be a molecule obtained by defect-treating the target atom in the original molecule of each compound sample. The defect treatment can include replacing the target atom in the original molecule with a defective atom that has no practical meaning, so that the connection relationship between the defective molecule and the original molecule remains unchanged. The defective atom can be represented as a Mask atom.

[0062] The target atom refers to an atom randomly selected from the original molecule that needs to be defect-treated. The number of target atoms can be one or more, which is related to the random selection method. The random selection method can be to randomly select atoms in the original molecule with a first probability, and set the selected atoms to have a second probability of undergoing defect treatment to obtain defect atoms, the selected atoms to have a third probability of being randomly replaced by other atoms, and the selected atoms to have a fourth probability of still being target atoms. Here, the first probability, the second probability, the third probability and the fourth probability can all be set as needed, wherein the second probability can be greater than 50%, and the first probability, the third probability and the fourth probability can be less than 50%. For example, the first probability can be 15%, the second probability can be 80%, the third probability can be 10%, and the fourth probability can be 10%.

[0063] The random selection method may also be to directly select 15% of the atoms from the original molecule as target atoms for defect treatment to obtain defect atoms.

[0064] The second type of sample molecular structure data refers to the molecular structure data of defective molecules, which is used to characterize the internal and overall structure of the defective molecules. It can include molecular information of the defective molecules, atomic information of each atom within them, and information about the connectivity between atoms. The atomic information of each atom can be represented as an atom list, and the connectivity information between atoms can be represented as an adjacency matrix.

[0065] The molecular attribute label refers to the attributes of the compound sample, which may include topological polar surface area, oil-water partition coefficient, relative molecular mass, number of hydrogen donors, number of hydrogen acceptors, functional groups, number of heteroatoms, and number of rings, etc., and is not specifically limited here.

[0066] Atomic property prediction pre-training refers to the process of training the initial model using the second-category sample molecular structure data of the defective molecules of each compound sample as the input of the initial model, and the atomic property labels of the target atoms in the original molecules of each compound sample as the object for calculating the loss with the output of the initial model.

[0067] Atomic attribute labels refer to the attributes of target atoms in a compound sample, which may include aromaticity, whether they are on a ring, hybridization mode, formal charge number, chirality, degree, number of hydrogen atoms, and atom type, etc., which are not specifically limited here.

[0068] Molecular comparative learning pre-training refers to the process of training the initial model using the second-category sample molecular structure data of each defective molecule and the third-category sample molecular structure data of each original molecule as the input of the initial model, and using the difference information between the representation vectors of each input obtained by the initial model as the object of calculating the loss.

[0069] It should be noted that in the embodiment of the present invention, the initial model is trained synchronously through three pre-training methods: molecular property prediction pre-training, atomic property prediction pre-training, and molecular comparative learning pre-training. That is, in each round of training, the overall loss is calculated by combining the losses calculated by the three pre-training methods, and then the model parameters of the initial model are adjusted according to the overall loss, and the next round of training is carried out until the overall loss converges or the total number of training rounds reaches the preset number.

[0070] As can be seen, both molecular property prediction and atomic property prediction pre-training are supervised training using labels, while molecular comparative learning pre-training is unsupervised training without labels. Adopting a pre-training strategy that combines unsupervised and supervised learning can help drug pre-training models learn more generalizable feature representations, thereby improving their performance and generalization in drug property prediction tasks. On this basis, drug pre-training can be applied to predict the drug properties of the target drug using input molecular structure data, thereby obtaining more accurate drug property information.

[0071] Furthermore, pre-training with atomic property prediction guides the initial model to enhance atomic property knowledge in the atomic features extracted by the drug pre-training model during feature extraction. Pre-training with molecular property prediction narrows the gap between the pre-training task and the fine-tuning drug property prediction task. Pre-training with molecular comparative learning further improves the quality of the molecular features extracted by the drug pre-training model. Furthermore, by simultaneously training the initial model with these three pre-training methods, the resulting drug pre-training model exhibits superior performance and generalization in drug property prediction tasks.

[0072] The drug property prediction method provided in an embodiment of the present invention first obtains target molecular structure data for the drug to be predicted; then, the target molecular structure data is input into a drug property prediction model to obtain drug property information for the drug to be predicted. The drug property prediction model employed is obtained by fine-tuning a pre-trained drug model, which is obtained through simultaneous training of molecular property prediction pre-training, atomic property prediction pre-training, and molecular comparative learning pre-training. This ensures the performance and generalizability of the pre-trained drug model in drug property prediction tasks, reduces the difficulty of migrating the pre-trained drug model, shortens the time required to build the drug prediction model, and improves the prediction accuracy of the drug prediction model obtained during the fine-tuning phase, thus ensuring its widespread application.

[0073] Based on the above embodiment, in the drug property prediction method provided in the embodiment of the present invention, the drug pre-training model is trained based on the following steps:

[0074] Calculating the atomic property prediction loss and the molecular property prediction loss of the initial model based on the molecular property labels of the original molecules, the atomic property labels of the target atoms in the original molecules, and the second type of sample molecular structure data of the defective molecules;

[0075] Calculating the contrastive learning loss of the initial model based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule;

[0076] Based on the atomic property prediction loss, the molecular property prediction loss and the contrastive learning loss, the overall loss of the initial model is determined, and based on the overall loss, the initial model is trained to obtain the drug pre-training model.

[0077] Specifically, in an embodiment of the present invention, in the process of training the initial model to obtain a drug pre-training model, the atomic property prediction loss and the molecular property prediction loss of the initial model can be calculated based on the molecular property labels of each original molecule, the atomic property labels of the target atoms in each original molecule, and the second type of sample molecular structure data of each defective molecule.

[0078] That is, for each original molecule, the second-category sample molecular structure data of the defect molecule corresponding to the original molecule can be input into the initial model. The initial model then performs feature extraction on the second-category sample molecular structure data of each defect molecule to obtain and output a representation vector for each defect molecule and a representation vector for the defect atom corresponding to the target atom in each defect molecule. It will be understood that the representation vector for each defect molecule can be obtained by feature extraction of the molecular information in each second-category sample molecular structure data, and the representation vector for each defect atom can be obtained by feature extraction of the atomic information of each defect atom.

[0079] Then, the atomic property prediction loss is calculated based on the representation vectors of each defect atom and the atomic property labels of the target atoms in each original molecule. During the calculation process, the correspondence between the defect atoms and the target atoms must be maintained. Only the representation vectors of defect atoms with a corresponding relationship can be used to calculate the atomic property prediction loss for their corresponding target atoms. This atomic property prediction loss can be used to characterize the error in the atomic property predictions made by the initial model.

[0080] The molecular property measurement loss is calculated based on the representation vectors of each defect molecule and the molecular property labels of each original molecule. During the calculation process, the correspondence between the defect atoms and the target atoms must be maintained. Only representation vectors of defect molecules with a corresponding relationship can be used to calculate the molecular property prediction loss for their corresponding original molecules. This molecular property prediction loss can be used to characterize the error in the initial model's prediction of molecular properties.

[0081] Then, the contrastive learning loss of the initial model is calculated based on the second-category sample molecular structure data of each defective molecule and the third-category sample molecular structure data of each original molecule. That is, for each original molecule, the second-category sample molecular structure data of the defective molecule corresponding to that original molecule is still used as the input of the initial model. In addition, the third-category sample molecular structure data and the second-category sample molecular structure data of each original molecule can be input into the initial model in sequence. In this way, the contrastive learning loss of the initial model can be calculated by combining the model output corresponding to the second-category sample molecular structure data with the model output corresponding to the third-category sample molecular structure data. This contrastive learning loss can be used to characterize the error generated when the initial model predicts the molecular properties of two different molecules.

[0082] The third type of sample molecular structure data refers to the molecular structure data of the original molecule, which is used to characterize the internal and overall structure of the original molecule. It can include molecular information of the original molecule, atomic information of each atom within it, and information about the connectivity between atoms. The atomic information of each atom can be represented in the form of an atom list, and the connectivity information between atoms can be represented in the form of an adjacency matrix.

[0083] Finally, the overall loss of the initial model is determined by combining the obtained atomic property prediction loss, molecular property prediction loss, and contrastive learning loss. Here, the atomic property prediction loss, molecular property prediction loss, and contrastive learning loss can be directly added together to obtain the overall loss of the initial model.

[0084] Furthermore, the model parameters of the initial model can be adjusted through the overall loss, and the above process can be performed again to train the initial model and obtain a drug pre-training model.

[0085] In an embodiment of the present invention, the second type of sample molecular structure data of each defective molecule can be used to simultaneously obtain the atomic property prediction loss and molecular property prediction loss of the initial model, two parameters for measuring the accuracy of the initial model. On this basis, the third type of sample molecular structure data of each original molecule is introduced to obtain the comparative learning loss. In this way, the time required for the initial model to learn knowledge can be shortened on the basis of increasing the knowledge learned by the initial model, thereby shortening the duration of pre-training.

[0086] Based on the above embodiment, the drug property prediction method provided in the embodiment of the present invention, wherein the comparative learning loss of the initial model is calculated based on the second-category sample molecular structure data of each defective molecule and the third-category sample molecular structure data of each original molecule, comprises:

[0087] Constructing positive sample pairs and negative sample pairs based on the defective molecules and the original molecules;

[0088] Inputting each structural data corresponding to the positive sample pair into the initial model to obtain each representation vector corresponding to the positive sample pair output by the initial model;

[0089] Inputting each structural data corresponding to the negative sample pair into the initial model to obtain each representation vector corresponding to the negative sample pair output by the initial model;

[0090] Based on the similarity between each representation vector corresponding to the positive sample pair and the similarity between each representation vector corresponding to the negative sample pair, a contrastive learning loss between the positive sample pair and the negative sample pair is calculated.

[0091] Specifically, in embodiments of the present invention, when calculating the contrastive learning loss of the initial model, positive and negative sample pairs can be constructed based on each defective molecule and each original molecule. The positive sample pairs can be composed of original molecules and defective molecules that have a corresponding relationship, and the negative sample pairs can be composed of original molecules and defective molecules that do not have a corresponding relationship. Here, multiple positive and negative sample pairs can be constructed.

[0092] Then, each structural data point corresponding to the positive sample pair is input into the initial model, resulting in the output of each representation vector corresponding to the positive sample pair. Here, the structural data points corresponding to the positive sample pair include the structural data points of the second-category sample molecules and the structural data points of the third-category sample molecules. These structural data points are derived from defective molecules and original molecules that have a corresponding relationship. Specifically, the defective molecules are obtained by performing defect treatment on the target atoms in the original molecules. Each structural data point in the positive sample pair corresponds to a representation vector.

[0093] Each structural data point in the negative sample pair is input into the initial model to obtain the representation vectors corresponding to the negative sample pair output by the initial model. Here, each structural data point in the negative sample pair includes the structural data of the second type of sample molecules and the structural data of the third type of sample molecules. These structural data points are derived from defective molecules and original molecules that do not have a corresponding relationship. That is, the defective molecules are not obtained by defect treatment of the target atoms in the original molecules. Each structural data point in the negative sample pair also corresponds to a representation vector.

[0094] Finally, the contrastive learning loss between the positive and negative sample pairs is calculated by combining the similarities between the representation vectors corresponding to the positive sample pairs and the similarities between the representation vectors corresponding to the negative sample pairs.

[0095] The contrastive loss function used to calculate the contrastive learning loss can be:

[0096]

[0097] Among them, τ is the heat factor, which is a hyperparameter, s a,i represents the similarity between the representation vectors of the original molecule a and the corresponding defect molecule i in the positive sample pair, s a,k It represents the similarity between the representation vectors of the original molecule a and the non-corresponding defective molecule k in the negative sample pair.

[0098] In the embodiment of the present invention, the similarity can use cosine similarity as a similarity calculation function, and the calculation method is as follows:

[0099]

[0100] Among them, P a is the representation vector of the original molecule a containing n atoms, It's P a The representation vector of the jth atom in k is the representation vector of the defect molecule k containing n atoms, It's P k The representation vector of the j-th atom in .

[0101] The contrast loss function requires that the numerator of the logarithmic function is as large as possible, and the similarity s between the positive sample pairs a,i As large as possible, and at the same time require the denominator part, that is, the similarity s between the negative sample pairs a,k As small as possible.

[0102] Based on the above embodiments, the drug property prediction method provided in the embodiments of the present invention, wherein the positive sample pairs and the negative sample pairs are constructed based on the defective molecules and the original molecules, comprises:

[0103] For a target original molecule, constructing a positive sample pair based on the target original molecule and a defective molecule corresponding to the target original molecule;

[0104] A negative sample pair is constructed based on the target original molecule and a defective molecule corresponding to other original molecules among the original molecules except the target original molecule.

[0105] Specifically, in the embodiment of the present invention, when constructing positive sample pairs and negative sample pairs, each original molecule can be used as a target original molecule to construct a positive sample pair and a negative sample pair containing the target original molecule.

[0106] For a target original molecule, a positive sample pair is constructed using the target original molecule as the original sample and a defect molecule corresponding to the target original molecule as the positive sample. The defect molecule corresponding to the target original molecule is one of the defect molecules obtained by defect-treating the target atom in the target original molecule.

[0107] A negative sample pair can also be constructed using the target original molecule as the original sample and a defect molecule corresponding to the original molecules other than the target original molecule in each original molecule as the negative sample. The original molecules other than the target original molecule refer to each original molecule other than the target original molecule, and the defect molecule corresponding to the original molecules other than the target original molecule in each original molecule refers to one of the defect molecules obtained by defect-treating the target atom in the other original molecules, i.e., the defect molecule does not correspond to the target original molecule.

[0108] In the embodiment of the present invention, a method for constructing positive sample pairs and negative sample pairs is provided, which can ensure the smooth implementation of molecular comparative learning pre-training and ensure the pre-training effect of molecular comparative learning pre-training.

[0109] On the basis of the above embodiments, in the drug property prediction method provided in the embodiments of the present invention, the target molecular structure data, the first category sample molecular structure data, the second category sample molecular structure data and the third category sample molecular structure data are all represented based on a molecular graph structure, and the molecular graph structure includes molecular-level nodes and atomic-level nodes, and the molecular-level nodes are connected to the atomic-level nodes.

[0110] Specifically, in embodiments of the present invention, a molecular graph structure can be introduced to represent molecular structure data. The molecular graph structure includes molecular-level nodes and atomic-level nodes. A molecular-level node can be a supernode, representing the entire molecule, so there is only one molecular-level node. An atomic-level node represents each atom in a molecule, with each atom corresponding to an atomic-level node. That is, the number of atomic-level nodes is the same as the number of atoms in the molecule. A molecular-level node can be connected to each atomic-level node.

[0111] Taking the original sample [S0]CHOHH as an example, the positive sample can be [S0]CH[M]H[M], and the negative sample can be [S1]NH[M]H[M]. [S0] represents the molecular-level node in the original sample, C, H, O, H, and N represent atomic-level nodes, [S1] represents the molecular-level node in the negative sample, and [M] represents the defect atom.

[0112] In the embodiment of the present invention, due to the introduction of the molecular graph structure, the representation of each molecular structure data can be simplified, thereby reducing the entire training process of the drug property prediction model and the hardware resource consumption of the drug property prediction process.

[0113] On the basis of the above embodiments, in the drug property prediction method provided in the embodiments of the present invention, the molecular property labels of each original molecule include regression task labels, and the atomic property labels of the target atoms include classification task labels.

[0114] Specifically, in an embodiment of the present invention, since the attributes of each atom have multiple categories or a range of values, the classification task label can be used as the atomic attribute label of the target atom. Since the attributes of each original molecule are mostly numerical data, the regression task label can be used as the molecular attribute label of each original molecule. Here, the classification task can include binary classification and multi-classification, and the number of corresponding classification task labels can include 2 or more. For example, aromaticity in the atomic attribute is a binary value, the corresponding atomic attribute label is 0 or 1, and the corresponding task type is a binary classification task.

[0115] Regression tasks involve predicting numerical values. Evaluation metrics are typically measured by numerical differences, primarily using mean squared error and absolute error. Therefore, the regression task labels are the values of the properties of each original molecule. Table 1 lists a total of 30 functional groups in the molecular properties. The presence of each functional group in a molecule corresponds to a binary classification task.

[0116] As shown in Table 1, it is a comparison table of atomic properties, molecular properties and task types.

[0117] Table 1 Comparison table of atomic attributes and task types

[0118]

[0119] On this basis, the atomic attribute prediction loss can be calculated by the multi-classification task loss function, which can be expressed as:

[0120]

[0121] Among them, x is the representation vector of the target atom of d dimensions, K is the number of categories of the atomic attribute label of the target atom corresponding to x, and W i ,W j are all d-dimensional parameter vectors in the initial model, d is the dimension of the representation vector set by the initial model, y i is the i-th type atomic attribute label of the target atom corresponding to x.

[0122] The molecular property prediction loss can be calculated by the regression task loss function, which can be expressed as:

[0123] L mole =(xW-y) 2

[0124] Among them, x is the representation vector of the d-dimensional original molecule, W is the d-dimensional parameter vector in the initial model, and y is the molecular attribute label corresponding to the original molecule.

[0125] On the basis of the above embodiments, in the drug property prediction method provided in the embodiments of the present invention, the initial model includes a model constructed based on a Transformer encoder.

[0126] Specifically, in an embodiment of the present invention, the initial model used may be a model constructed based on a Transformer encoder, and the model structure of the initial model may be the structure of a Transformer encoder.

[0127] The Transformer encoder can be composed of 6 identical layers stacked together, each of which is divided into two sub-layers: a multi-head attention layer and a position-based fully connected feedforward network layer. The multi-head attention layer uses a self-attention mechanism, and its formula is as follows:

[0128] MH(H l )=[head1;head2;...;head h ]W O

[0129] head i =Attention(H l W i Q ,H l W i K ,H l W i V );

[0130] Among them, H l ∈R t×d is the input of the current layer, l is the number of layers, t is the sequence length, and d is the vector dimension specified by the Transformer encoder. The multi-attention layer first linearly maps the input of this layer to h different subspaces, where each subspace mapping uses different learnable parameters. Then, h attention calculation functions are executed in parallel and the output results are concatenated and linearly mapped to form the output of the multi-head attention layer. Here, the query space mapping matrix W of each head is i Q ∈R d×d / h , key space mapping matrix W i K ∈R d×d / h , value space mapping matrix W i V ∈R d×d / h And the final mapping matrix W i O ∈R d×d / h These are all learnable parameters, and the parameters of the mapping matrices of different layers are not shared. The scaled dot product attention function used in this layer is shown in the following formula:

[0131]

[0132] Here, the query Q, key K and value V are all obtained by transforming the input of this layer through the mapping matrix of different spaces. Generate a softer attention score distribution. Combined with the formula, we can see that linear transformation is used in the multi-head attention layer. In order to make the model have nonlinearity and interaction of different dimensions, the position fully connected feedforward network layer is introduced in the Transformer encoder, and its composition is shown in the following formula:

[0133] PFFN(H l )=[FFN(h1 l ) T ;...;FFN(h t l ) T ] T

[0134] FFN(x)=GELU(xW (1) +b (1) )W (2) +b (2)

[0135] GELU(x)=xΦ(x)

[0136] The fully connected position feedforward network layer contains two nonlinear transformations, where the nonlinear function uses the Gaussian Error Linear Unit (GELU) activation function. Here Φ(x) is the cumulative distribution function of the standard Gaussian distribution. W (1) ∈R d×4d , W (2) ∈R 4d×d , b (1) ∈R 4d and b (2) ∈R d are all learnable parameters.

[0137] In order to facilitate the observation of the operation process of the three pre-training methods, namely molecular property prediction pre-training, atomic property prediction pre-training and contrastive learning pre-training, as shown in the following figure: Figure 2 As shown in FIG, the process of inputting the third type sample molecular structure data of each original molecule and the second type sample molecular structure data of the corresponding defective molecule into the initial model is displayed in parallel. Figure 2 It includes two Transformer encoders on the left and right. The input of the left Transformer encoder is the graph structure of the original sample, expressed as [S0]CHOHH, and the input of the right Transformer encoder is the graph structure of the positive sample, expressed as [S0]CH[M]H[M].

[0138] Each Transformer encoder contains a multi-head attention layer using a self-attention mechanism and a position-based fully-connected feedforward network layer. A normalization layer is connected after the multi-head attention layer and the position-based fully-connected feedforward network layer to add and normalize the input with the output of the multi-head attention layer or the position-based fully-connected feedforward network layer, and finally output the result by the normalization layer connected after the position-based fully-connected feedforward network layer.

[0139] Transformer encoder on the left: Each atom in the graph structure of the input original sample is converted into an initial vector x1 to x6. After the initial vectors x1 to x6 pass through the multi-head attention layer, the feature vectors z1 to z6 of each atom are obtained respectively. The feature vectors z1 to z6 and the initial vectors x1 to x6 are added and normalized by the layer normalization layer to obtain the feature vectors z1 to z6 again. The feature vectors z1 to z6 pass through the position fully connected feedforward network layer and the layer normalization layer connected to it, and then the representation vectors are output. The representation vector corresponding to [S] is used to calculate the contrastive learning loss and the molecular knowledge loss.

[0140] Transformer encoder on the right: Each atom in the graph structure of the input positive sample is converted into initial vectors m1 to m6. After the initial vectors m1 to m6 pass through the multi-head attention layer, the feature vectors y1 to y6 of each atom are obtained. The feature vectors y1 to y6 and the initial vectors m1 to m6 are added and normalized by the layer normalization layer to obtain the feature vectors y1 to y6. The feature vectors y1 to y6 pass through the position fully connected feedforward network layer and the connected layer normalization layer to output the respective representation vectors. The representation vector corresponding to [S] is used to calculate the contrastive learning loss and the molecular knowledge loss. The representation vector corresponding to [M] is used to calculate the atomic knowledge loss.

[0141] Figure 2 In [1], X represents the combination vector of initial vectors x1 to x6, Z represents the combination vector of initial vectors z1 to z6, M represents the combination vector of initial vectors m1 to m6, and Y represents the combination vector of initial vectors y1 to y6. Molecular knowledge loss refers to the loss in molecular property prediction, and atomic knowledge loss refers to the loss in atomic property prediction.

[0142] In the embodiment of the present invention, a model constructed based on a Transformer encoder is used as the initial model, which can enable the drug property prediction model to have a good feature extraction function, thereby helping to improve the prediction accuracy of the drug property information of the drug to be predicted.

[0143] Based on the above embodiment, an embodiment of the present invention provides a method for training a drug property prediction model in a drug property prediction method, the training method comprising:

[0144] A) Extract unlabeled molecular structure data from an existing compound database as a drug pre-training dataset, extract the attribute knowledge items of the original molecules in the drug pre-training dataset as molecular attribute labels, extract the attribute knowledge items of each atom in the original molecule as atomic attribute labels, and determine the molecular attribute labels as regression task labels, and the atomic attribute labels as classification task labels.

[0145] B) A graph structure is used to represent the molecules of the compound sample. The graph structure includes molecular-level nodes and atomic-level nodes. Molecular-level nodes represent molecules, and atomic-level nodes represent atoms within the molecule. For the target original molecule as the original sample, some nodes in the graph structure (i.e., nodes corresponding to target atoms) are randomly selected and replaced with mask nodes (i.e., nodes corresponding to defect atoms). The molecule is then encoded using an initial model built based on the Transformer encoder.

[0146] C) Use the representation vector of the Mask node output by the initial model to infer the attribute knowledge of the corresponding target atom, and calculate the atomic attribute prediction loss of the initial model through the atomic attribute label; use the representation vector of the molecular-level node output by the initial model to infer the attribute knowledge of the molecule, and calculate the molecular attribute prediction loss of the initial model through the molecular attribute label.

[0147] D) Defective molecules containing Mask nodes corresponding to the target original molecule are used as positive samples, and defective molecules containing Mask nodes corresponding to other original molecules except the target original molecule during training are used as negative samples. The contrastive learning loss is calculated, and the initial model is updated based on the atomic property prediction loss, molecular property prediction loss, and contrastive learning loss.

[0148] E) completing pre-training steps B to D on the initial model on the drug pre-training dataset consisting of drug samples, and using the drug pre-training model obtained after completing the pre-training task to fine-tune the drug property prediction dataset to obtain a drug property prediction model.

[0149] The training method of the drug property prediction model provided in the embodiment of the present invention can effectively reduce the difficulty of migrating the drug pre-training model from the pre-training task to the drug property prediction task, and improve the prediction effect of the drug property prediction model.

[0150] The above method enables the drug property prediction model provided in the embodiments of the present invention to achieve better results on publicly available drug property prediction datasets than existing baseline drug property prediction models. To this end, four drug property prediction datasets from MoleculeNet were selected, and their detailed information is shown in Table 2 below:

[0151] Table 2 Detailed information of each drug attribute prediction dataset

[0152]

[0153] Each of the aforementioned drug attribute prediction datasets contains one or more classification tasks. Each task uses the ROC-AUC evaluation metric. The average of the initial model's evaluation metrics across all tasks in each drug attribute prediction dataset is used as the final metric for that initial model. The drug attribute prediction datasets were partitioned into training, validation, and test sets in an 8:1:1 ratio. For the Tox21 and ToxCast datasets, the Scaffold partitioning method recommended by MoleculeNet was used, while random partitioning was performed for the other datasets.

[0154] In the drug property prediction model (DPECK) provided in the embodiment of the present invention, the number of encoder layers is 6, the number of attention heads is 4, the dimension of the feature is 128, the dropout regularization retention ratio is 0.9, the batch size in the pre-training stage is 128, the batch size in the fine-tuning stage is 64, the maximum sequence length is 150, the optimizer is Adam, and the learning rate is 5e-5. In all experiments, the parameters of the baseline model remain consistent, and the code framework uses Tensorflow. The operating system used in the experiment is Ubuntu, the CPU is Intel(R) Xeon(R) Gold 6226R CPU@2.90GHz, the GPU is NVIDIA RTX 2080Ti (11GB), Python version 3.7, and Tensorflow version is 2.4.1. In the embodiment of the present invention, 12 baseline drug property prediction models are used for comparison with DPECK. The performance of each model on each data set is shown in Table 3 below:

[0155] Table 3 Performance of each model on each dataset

[0156]

[0157] From the experimental results, it can be seen that the DPECK in the present invention has achieved the best performance on all data sets. Compared with the optimal baseline drug property prediction model, DPECK has achieved an improvement of 0.6-1.4 percentage points on four data sets of different sizes and types, proving the effectiveness of the DPECK for drug property prediction tasks. On the relatively small BACE and BBBP data sets, DPECK achieved improvements of 1.5% and 2.0% respectively compared to MG-BERT, indicating that the atomic property prediction pre-training, molecular property prediction pre-training and molecular comparative learning pre-training in the embodiments of the present invention have significantly improved the performance of DPECK on small sample data sets. Compared with non-pretrained models, the pre-trained model performs better and more stably. This proves that pre-learning drug data representation on large data sets can help improve the performance of the model on small sample drug property prediction tasks.

[0158] Table 4 shows the actual contribution of different components of DPECK to the model's final performance. MG-BERT was used for comparison. MG-BERT and DPECK share the same model architecture, but MG-BERT's pre-training tasks only include a single atom type prediction task. The compound property knowledge enhancement module (based on atomic and molecular property prediction pre-training) and the contrastive learning module (based on molecular contrastive learning pre-training) were removed from DPECK, forming the wo-knowledge model and wo-contrastive model, respectively, to evaluate the performance of each module independently. The experimental results, shown in Table 4, show that combining the compound property knowledge enhancement module and the contrastive learning module significantly improves DPECK compared to MG-BERT, demonstrating the effectiveness of combining these two modules in drug pre-training algorithms for drug property prediction. As shown in Table 4, combining these two modules yields improvements of 1.5%, 2.0%, 1.6%, and 1.3% on the four datasets, respectively, compared to MG-BERT. This demonstrates that contrastive learning, based on enhanced compound property knowledge, can further improve DPECK's performance by distinguishing different samples.

[0159] Table 4 Actual contribution of different parts of DPECK to the final performance of the model

[0160]

[0161]

[0162] like Figure 3 As shown, based on the above embodiment, an embodiment of the present invention provides a drug property prediction device, including:

[0163] The data acquisition module 31 is used to obtain the target molecular structure data of the drug to be predicted;

[0164] Prediction module 32, used to input the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted;

[0165] The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method:

[0166] Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated;

[0167] performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules;

[0168] Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

[0169] Based on the above embodiment, the drug property prediction device provided in the embodiment of the present invention further includes a pre-training module for:

[0170] Calculating the atomic property prediction loss and the molecular property prediction loss of the initial model based on the molecular property labels of the original molecules, the atomic property labels of the target atoms in the original molecules, and the second type of sample molecular structure data of the defective molecules;

[0171] Calculating the contrastive learning loss of the initial model based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule;

[0172] Based on the atomic property prediction loss, the molecular property prediction loss and the contrastive learning loss, the overall loss of the initial model is determined, and based on the overall loss, the initial model is trained to obtain the drug pre-training model.

[0173] Based on the above embodiments, in the drug property prediction device provided in the embodiments of the present invention, the pre-training module is specifically used to:

[0174] Constructing positive sample pairs and negative sample pairs based on the defective molecules and the original molecules;

[0175] Inputting each structural data corresponding to the positive sample pair into the initial model to obtain each representation vector corresponding to the positive sample pair output by the initial model;

[0176] Inputting each structural data corresponding to the negative sample pair into the initial model to obtain each representation vector corresponding to the negative sample pair output by the initial model;

[0177] Based on the similarity between each representation vector corresponding to the positive sample pair and the similarity between each representation vector corresponding to the negative sample pair, a contrastive learning loss between the positive sample pair and the negative sample pair is calculated.

[0178] Based on the above embodiments, in the drug property prediction device provided in the embodiments of the present invention, the pre-training module is specifically used to:

[0179] For a target original molecule, constructing a positive sample pair based on the target original molecule and a defective molecule corresponding to the target original molecule;

[0180] A negative sample pair is constructed based on the target original molecule and a defective molecule corresponding to other original molecules among the original molecules except the target original molecule.

[0181] On the basis of the above embodiments, in the drug property prediction device provided in the embodiments of the present invention, the target molecular structure data, the first category sample molecular structure data, the second category sample molecular structure data and the third category sample molecular structure data are all represented based on a molecular graph structure, and the molecular graph structure includes molecular-level nodes and atomic-level nodes, and the molecular-level nodes are connected to the atomic-level nodes.

[0182] On the basis of the above embodiment, in the drug property prediction device provided in the embodiment of the present invention, the molecular property label of each original molecule includes a regression task label, and the atomic property label of the target atom includes a classification task label.

[0183] On the basis of the above-mentioned embodiment, in the drug property prediction device provided in the embodiment of the present invention, the initial model includes a model constructed based on a Transformer encoder.

[0184] Specifically, the functions of each module in the drug property prediction device provided in the embodiment of the present invention correspond one-to-one to the operating procedures of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further details will be given in the embodiment of the present invention.

[0185] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4As shown, the electronic device may include: a processor (Processor) 410, a communication interface (Communications Interface) 420, a memory (Memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the drug property prediction method provided in the above embodiments, which method includes: obtaining the target molecular structure data of the drug to be predicted; inputting the target molecular structure data into the drug property prediction model to obtain the drug property information of the drug to be predicted; wherein, the drug property prediction model is based on the first type of sample molecular structure data of drug samples carrying drug property labels, and the drug pre-training model is trained; the drug pre-training model is obtained by synchronous training on the basis of the initial model using the following pre-training method: based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of the defective molecules obtained after the target atoms in each original molecule are processed by defects, molecular property prediction pre-training is performed; based on the atomic property labels of the target atoms in each original molecule and the second type of sample molecular structure data of each defective molecule, atomic property prediction pre-training is performed; based on the second type of sample molecular structure data of each defective molecule and the third type of sample molecular structure data of each original molecule, molecular comparison learning pre-training is performed.

[0186] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0187] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the drug property prediction method provided by the above-mentioned methods, the method including: obtaining the target molecular structure data of the drug to be predicted; inputting the target molecular structure data into a drug property prediction model to obtain the drug property information of the drug to be predicted; wherein, the drug property prediction model is obtained by training a drug pre-training model based on the first type of sample molecular structure data of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training on the basis of the initial model using the following pre-training method: molecular property prediction pre-training based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of the defective molecules obtained after the target atoms in each original molecule are processed by defects; atomic property prediction pre-training based on the atomic property labels of the target atoms in each original molecule and the second type of sample molecular structure data of each defective molecule; molecular comparative learning pre-training based on the second type of sample molecular structure data of each defective molecule and the third type of sample molecular structure data of each original molecule.

[0188] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the drug property prediction method provided by the above-mentioned methods, the method comprising: obtaining target molecular structure data of the drug to be predicted; inputting the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted; wherein, the drug property prediction model is obtained by training a drug pre-training model based on the first type of sample molecular structure data of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training on the basis of the initial model using the following pre-training method: molecular property prediction pre-training based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of the defective molecules obtained after defect treatment of the target atoms in each original molecule; atomic property prediction pre-training based on the atomic property labels of the target atoms in each original molecule and the second type of sample molecular structure data of each defective molecule; molecular comparative learning pre-training based on the second type of sample molecular structure data of each defective molecule and the third type of sample molecular structure data of each original molecule.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for predicting drug properties, characterized in that: include: Obtain target molecular structure data of the drug to be predicted; Inputting the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted; The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method: Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated; performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules; Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

2. The drug property prediction method according to claim 1, characterized in that: The drug pre-training model is trained based on the following steps: Calculating the atomic property prediction loss and the molecular property prediction loss of the initial model based on the molecular property labels of the original molecules, the atomic property labels of the target atoms in the original molecules, and the second type of sample molecular structure data of the defective molecules; Calculating the contrastive learning loss of the initial model based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule; Based on the atomic property prediction loss, the molecular property prediction loss and the contrastive learning loss, the overall loss of the initial model is determined, and based on the overall loss, the initial model is trained to obtain the drug pre-training model.

3. The drug property prediction method according to claim 2, characterized in that: The calculating the contrastive learning loss of the initial model based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule includes: Constructing positive sample pairs and negative sample pairs based on the defective molecules and the original molecules; Inputting each structural data corresponding to the positive sample pair into the initial model to obtain each representation vector corresponding to the positive sample pair output by the initial model; Inputting each structural data corresponding to the negative sample pair into the initial model to obtain each representation vector corresponding to the negative sample pair output by the initial model; Based on the similarity between each representation vector corresponding to the positive sample pair and the similarity between each representation vector corresponding to the negative sample pair, a contrastive learning loss between the positive sample pair and the negative sample pair is calculated.

4. The drug property prediction method according to claim 3, characterized in that: The constructing of positive sample pairs and negative sample pairs based on the defective molecules and the original molecules comprises: For a target original molecule, constructing a positive sample pair based on the target original molecule and a defective molecule corresponding to the target original molecule; A negative sample pair is constructed based on the target original molecule and a defective molecule corresponding to other original molecules among the original molecules except the target original molecule.

5. The method for predicting drug properties according to any one of claims 1 to 4, characterized in that: The target molecular structure data, the first type of sample molecular structure data, the second type of sample molecular structure data and the third type of sample molecular structure data are all represented based on a molecular graph structure, and the molecular graph structure includes molecular-level nodes and atomic-level nodes, and the molecular-level nodes are connected to the atomic-level nodes.

6. The method for predicting drug properties according to any one of claims 1 to 4, characterized in that: The molecular attribute labels of the original molecules include regression task labels, and the atomic attribute labels of the target atoms include classification task labels.

7. The method for predicting drug properties according to any one of claims 1 to 4, characterized in that: The initial model includes a model built based on a Transformer encoder.

8. A drug property prediction device, characterized in that: include: A data acquisition module is used to obtain target molecular structure data of the drug to be predicted; A prediction module, configured to input the target molecular structure data into a drug property prediction model to obtain drug property information of the drug to be predicted; The drug property prediction model is obtained by training a drug pre-training model based on the molecular structure data of the first type of drug samples carrying drug property labels; the drug pre-training model is obtained by synchronous training based on the initial model using the following pre-training method: Molecular property prediction pre-training is performed based on the molecular property labels of the original molecules of each compound sample and the second type of sample molecular structure data of defective molecules obtained after the target atoms in each original molecule are defect-treated; performing atomic property prediction pre-training based on the atomic property labels of the target atoms in the original molecules and the second type of sample molecular structure data of the defective molecules; Molecular comparative learning pre-training is performed based on the second type sample molecular structure data of each defective molecule and the third type sample molecular structure data of each original molecule.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the drug property prediction method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting drug properties according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Drug molecule property prediction method, device and equipment based on comparative learning

    CN114386694A

  • Semi-supervised training method for bioactivity prediction of compounds and system thereof

    KR102343523B1