Method and apparatus for predicting drug properties, electronic device, storage medium
By constructing a deep learning model that combines unsupervised and supervised, using drug molecular sequences and molecular maps, the problem of low prediction accuracy in the existing technology is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202310492482.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-05-04
AI Technical Summary
In the prior art, when a model predicting drug properties is obtained using drugs with known drug properties, the trained model has a low accuracy in predicting drug properties of the drug.
By obtaining the drug molecular sequence and molecular map of the drug to be predicted, using unsupervised and supervised training deep learning models, combining an encoder and graph neural network, a drug property prediction model is constructed, and a multiple first sample drug without drug properties tags and a second sample drug data set with drug properties tags are extracted.
The accuracy of drug properties prediction models is improved and the properties of drugs to be predicted can be predicted more accurately.
Smart Images

Figure CN116612828B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of drug property prediction, for example, to a method and device, electronic device, and storage medium for predicting drug properties. Background Art
[0002] Currently, during drug development, researchers often rely on the pharmacological properties of existing drugs to develop new drugs. Therefore, predicting the pharmacological properties of drugs has become a crucial issue. Related technologies typically use machine learning to generate a model for predicting drug properties using drugs with known pharmacological properties. This model is then used to predict the pharmacological properties of the drug to be predicted.
[0003] During the implementation of the embodiments of the present disclosure, it was discovered that the related art has at least the following problems: In the related art, a model for predicting drug properties is obtained using drugs with known drug properties. This model is then used to predict drug properties. However, due to the very small number of drugs with known drug properties, the trained model has a low accuracy rate in predicting the drug properties of the drug to be predicted.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0006] The embodiments of the present disclosure provide a method and apparatus, an electronic device, and a storage medium for predicting drug properties, so as to improve the accuracy of predicting drug properties.
[0007] In some embodiments, the method for predicting drug properties includes: obtaining a drug molecule sequence and a drug molecule graph of the drug to be predicted. And determining the target property type. Inputting the drug molecule sequence and the drug molecule graph into a drug property prediction model corresponding to the target property type to obtain drug property prediction data. The drug property prediction model corresponding to the target property type is obtained based on a sample drug data set corresponding to the target property type; the sample drug data set includes multiple first sample drugs and multiple second sample drugs. The first sample drug is a sample drug that does not carry a drug property label. The second sample drug is a sample drug that carries a drug property label. Based on the drug property prediction data and the target property type, the drug property of the drug to be predicted is determined.
[0008] In some embodiments, a drug property prediction model is obtained by the following method, including: obtaining a first sample drug molecular sequence corresponding to each first sample drug, a first sample drug molecular graph corresponding to each first sample drug, a second sample drug molecular sequence corresponding to each second sample drug, and a second sample drug molecular graph corresponding to each second sample drug. Inputting each first sample drug molecular sequence and each first sample drug molecular graph into a preset first deep learning model for unsupervised training to obtain a first alternative drug property prediction model. Migrating the model parameters of the first alternative drug property prediction model into a preset second deep learning model to obtain a second alternative drug property prediction model. Inputting each second sample drug molecular sequence, each second sample drug molecular graph, and the drug property label of each second sample drug into the second alternative drug property prediction model for supervised training to obtain a drug property prediction model.
[0009] In some embodiments, the preset first deep learning model includes a preset first encoder and a preset first graph neural network. Inputting each first sample drug molecule sequence and each first sample drug molecule graph into the preset first deep learning model for unsupervised training to obtain a first candidate drug property prediction model includes: inputting each first sample drug molecule sequence into the preset first encoder for training to obtain a first molecular formula embedding vector; and inputting each first sample drug molecule graph into the preset first graph neural network for training to obtain a first molecular graph embedding vector. The first candidate drug property prediction model is obtained based on the first molecular formula embedding vector and the first molecular graph embedding vector.
[0010] In some embodiments, inputting each first sample drug molecule sequence into a preset first encoder for training to obtain a first molecular formula embedding vector includes: encoding the atomic properties and atomic positions of each atom in the first sample drug molecule sequence, respectively, to obtain a first sample atomic property matrix corresponding to the first sample drug molecule sequence and a first sample atomic position matrix corresponding to the first sample drug molecule sequence. Utilizing an attention mechanism to obtain the first molecular formula embedding vector based on the first sample atomic property matrix and the first sample atomic position matrix.
[0011] In some embodiments, inputting each first sample drug molecule graph into a preset first graph neural network for training to obtain a first molecular graph embedding vector includes: obtaining node information of each node in each first sample drug molecule graph, and aggregating the node information to obtain the first molecular graph embedding vector.
[0012] In some embodiments, obtaining the first candidate drug property prediction model based on the first molecular formula embedding vector and the first molecular graph embedding vector includes: obtaining a contrast loss value based on the first molecular formula embedding vector and the first molecular graph embedding vector. Obtaining the first candidate drug property prediction model based on the contrast loss value.
[0013] In some embodiments, the second candidate drug property prediction model includes a second encoder, a second graph neural network, and a preset fully connected layer network. Each second sample drug molecular sequence, each second sample drug molecular graph, and each second sample drug drug's drug property label are input into the second candidate drug property prediction model for supervised training to obtain a drug property prediction model, including: determining a training set and a test set from each second sample drug molecular sequence, each second sample drug molecular graph, and each drug property label in a preset ratio. The training set includes multiple third sample drug molecular sequences, multiple third sample drug molecular graphs, and multiple training drug labels. Each third sample drug molecular sequence is input into the second encoder for training to obtain a second molecular formula embedding vector. Each third sample drug molecular graph is input into the second graph neural network for training to obtain a second molecular graph embedding vector. The second molecular formula embedding vector and the second molecular graph embedding vector are input into a preset fully connected layer network for training to obtain a first predicted property label. The drug property prediction model is obtained based on the first predicted property label, the training drug label, and the test set.
[0014] In some embodiments, the apparatus for predicting drug properties includes a processor and a memory storing program instructions, and the processor is configured to execute the above-mentioned method for predicting drug properties when running the program instructions.
[0015] In some embodiments, the electronic device includes an electronic device body, and the device for predicting drug properties is installed in the electronic device body.
[0016] In some embodiments, the storage medium stores program instructions, and when the program instructions are run, the above-mentioned method for predicting drug properties is executed.
[0017] The method and device, electronic device, and storage medium for predicting drug properties provided by the embodiments of the present disclosure can achieve the following technical effects: by inputting the drug molecule sequence and drug molecule graph of the drug to be predicted into the drug property prediction model. Then, the drug property of the drug to be predicted is determined based on the drug property prediction data and the property type of the drug to be predicted. Among them, the drug property prediction model is obtained through a plurality of first sample drugs that do not carry drug property labels and second sample drugs that carry drug property labels. In this way, compared with only using a very small number of drugs with known drug properties to obtain the drug property prediction model. This solution uses the first sample drug that does not carry the drug property label to obtain the drug property prediction model, so that the drug property prediction model can extract richer features of the drug to be predicted, and can more accurately predict the drug properties of the drug to be predicted. Thereby, the accuracy of the drug property prediction model in predicting drug properties is improved.
[0018] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,
[0020] Figure 1 is a schematic diagram of a method for predicting drug properties provided by an embodiment of the present disclosure;
[0021] Figure 2 is a schematic diagram of a preset first deep learning model provided by an embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram of obtaining a third candidate drug property prediction model provided by an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of a method for predicting drug properties provided by an embodiment of the present disclosure;
[0024] Figure 5 Schematic diagram of a device for obtaining a drug property prediction model provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0026] In the description and claims of the embodiments of the present disclosure, as well as in the accompanying drawings, the terms "first," "second," and the like are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe the embodiments of the present disclosure herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.
[0027] Unless otherwise stated, the term "plurality" means two or more.
[0028] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0029] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0030] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0031] The present application is applied to electronic devices. The electronic device obtains the drug molecule sequence of the drug to be predicted, the drug molecule graph of the drug to be predicted, and the property type of the property to be predicted. And determines the target property type. Then the drug molecule sequence and the drug molecule graph are input into the drug property prediction model corresponding to the target property type to obtain drug property prediction data. The drug property of the drug to be predicted is determined based on the drug property prediction data and the target property type. Because the drug property prediction model corresponding to the property type is obtained based on a sample drug data set corresponding to the property type, including multiple first sample drugs without drug property labels and multiple second sample drugs with drug property labels. Compared with using a very small number of drugs with known drug properties to obtain a drug property prediction model, this solution uses the first sample drug without drug property labels to obtain the drug property prediction model, so that the drug property prediction model can extract richer features of the drug to be predicted, so as to accurately predict the drug properties of the drug to be predicted. Thereby improving the accuracy of the drug property prediction model in predicting drug properties.
[0032] Combine Figure 1 As shown, the present disclosure provides a method for predicting drug properties, comprising:
[0033] Step S101: obtaining a drug molecule sequence and a drug molecule graph of a drug to be predicted; and determining a target property type.
[0034] In step S102, the electronic device inputs the drug molecule sequence and the drug molecule graph into the drug property prediction model corresponding to the target property type to obtain drug property prediction data; wherein, the drug property prediction model corresponding to the target property type is obtained based on a sample drug data set corresponding to the target property type; the sample drug data set includes multiple first sample drugs and multiple second sample drugs; the first sample drugs are sample drugs without drug property labels; and the second sample drugs are sample drugs with drug property labels.
[0035] In step S103 , the electronic device determines the drug property of the drug to be predicted according to the drug property prediction data and the target property type.
[0036] The method for predicting drug properties provided by the embodiment of the present disclosure is adopted, and the drug molecular sequence and drug molecular graph of the drug to be predicted are input into the drug property prediction model. Then, the drug property of the drug to be predicted is determined based on the drug property prediction data and the property type of the drug to be predicted. Among them, the drug property prediction model is obtained through a plurality of first sample drugs that do not carry drug property labels and second sample drugs that carry drug property labels. In this way, compared with only using a very small number of drugs with known drug properties to obtain the drug property prediction model, this solution uses the first sample drugs that do not carry drug property labels to obtain the drug property prediction model, so that the drug property prediction model can extract richer features of the drug to be predicted, and can more accurately predict the drug properties of the drug to be predicted. Thereby, the accuracy of the drug property prediction model in predicting drug properties is improved.
[0037] The drug molecule sequence of the drug to be predicted is represented by a Smiles (Simplified Molecular Input Line Entry System) representation of the drug molecule of the drug to be predicted. Smiles are string data. The Smiles representation of the drug to be predicted preserves the structural features of the drug molecule of the drug to be predicted. The Smiles representation enables the reconstruction of the drug molecule graph of the drug to be predicted.
[0038] The drug molecule graph of the drug to be predicted is composed of nodes and edges. The nodes of the drug molecule graph represent atoms of the drug molecule of the drug to be predicted. The edges of the drug molecule graph represent chemical bonds between atoms of the drug molecule of the drug to be predicted.
[0039] The sample drug dataset corresponding to the target property type includes a first dataset and a second dataset; wherein the first dataset includes a ZINC dataset and / or a CHEMBL dataset. The first dataset includes multiple first sample drugs. The second dataset includes second sample drugs. Both the Chembl dataset and the ZINC dataset are datasets of small molecule compounds.
[0040] The target property type is obtained by: receiving a property type determination instruction input by a user, the property type determination instruction including a property type selected by the user. The property type selected by the user is determined as the target property type. Thus, by determining the property type selected by the user as the target property type, and inputting the drug molecular sequence and drug molecular graph of the drug to be predicted into a drug property prediction model corresponding to the target property type, drug property prediction data is obtained. For the drug to be predicted, drug property prediction can be performed for the property type desired by the user.
[0041] Among them, the property types include: whether the drug to be predicted contains compounds with blood-brain barrier permeability, whether the drug to be predicted is a compound that can be used as a human β-secretase 1 inhibitor, the solubility of the drug to be predicted, the atomization energy of the stable and comprehensively accessible organic molecules possessed by the drug to be predicted, the hydration free energy of small molecules in water obtained by experiments and alchemical free energy of the drug to be predicted, the FDA (Food and Drug Administration) approval status of the drug to be predicted and whether it is toxic in clinical trials, the type of toxicity possessed by the drug to be predicted, the type of source of side effects of the drug to be predicted, and the quantum mechanical properties of the drug to be predicted. Among them, quantum mechanical properties include the electronic energy spectrum and excited state energy of small molecules.
[0042] The second sample drug corresponding to the target property type is a sample drug in the data set corresponding to the target property type. The drug property label represents the drug property corresponding to the second sample drug in the target property type.
[0043] In some embodiments, when the target property type is whether the drug to be predicted contains a compound with the ability to penetrate the blood-brain barrier, the second data set is a BBBP data set, and the drug property prediction model corresponding to the target property type is a binary classification model. The drug property label carried by the second sample drug is a first preset value or a second preset value. For example, the first preset value is 1 and the second preset value is 0. When the drug property label carried by the second sample drug is the first preset value, the drug property characterizing the second sample drug is that it contains a compound with the ability to penetrate the blood-brain barrier. When the drug property label carried by the second sample drug is the second preset value, the drug property characterizing the second sample drug is that it does not contain a compound with the ability to penetrate the blood-brain barrier.
[0044] In some embodiments, when the target property type is a compound to be predicted as a human β-secretase 1 inhibitor, the second data set is a BACE data set, and the drug property prediction model corresponding to the target property type is a binary classification model. The drug property label carried by the second sample drug is a first preset value or a second preset value. For example, the first preset value is 1 and the second preset value is 0. When the drug property label carried by the second sample drug is the first preset value, the drug property of the second sample drug is characterized as a compound that can be used as a human β-secretase 1 inhibitor. When the drug property label carried by the second sample drug is the second preset value, the drug property of the second sample drug is characterized as a compound that cannot be used as a human β-secretase 1 inhibitor.
[0045] In some embodiments, when the target property type is the solubility of the drug to be predicted, the second data set is an ESOL data set, and the drug property prediction model corresponding to the target property type is a regression task model. The drug property label carried by the second sample drug is a numerical value. This numerical value is used to characterize the solubility of the second sample drug. For example, when the drug property label carried by the second sample drug is 50, the drug property used to characterize the second sample drug is: the solubility of the second sample drug is 50.
[0046] In some embodiments, when the target property type is the atomization energy of stable and comprehensively accessible organic molecules possessed by the drug to be predicted, the second dataset is a QM7 dataset, and the drug property prediction model corresponding to the target property type is a regression task model. The drug property label carried by the second sample drug is a numerical value. This numerical value is used to represent the amount of atomization energy of stable and comprehensively accessible organic molecules possessed by the second sample drug.
[0047] In some embodiments, when the target property type is the hydration free energy of small molecules in water obtained through experiments and alchemical free energy for the drug to be predicted, the second dataset is a FreeSolv dataset, and the drug property prediction model corresponding to the target property type is a regression task-type model. The drug property label carried by the second sample drug is a numerical value. This numerical value is used to represent the amount of hydration free energy of small molecules in water obtained by experiments and alchemical free energy for the second sample drug.
[0048] In some embodiments, when the target property type is the FDA approval status of the drug to be predicted and whether it exhibits toxicity in clinical trials, the second dataset is the ClinTox dataset, and the drug property prediction model corresponding to the target property type is a multi-label classification model. The drug property label carried by the second sample drug is a two-dimensional array. The value of each element in the array is either a first preset value or a second preset value. Each element position in the array represents a property. One element characterizes the FDA approval status of the drug, for example, FDA-approved or FDA-not approved. Another element characterizes whether the drug exhibits toxicity in clinical trials. Based on the values of each element and the properties corresponding to each element position, multiple sub-properties of the second sample drug are obtained. The sub-properties are combined to obtain the drug property of the second sample drug. For example, if the drug property label is [0, 1], the property corresponding to the element at position 1 is the FDA approval status of the second sample drug. The value of the element at position 1 is 0. Based on the value of this element and the property corresponding to this element position, a sub-property of the second sample drug is obtained: the second sample drug is not FDA-approved. The property corresponding to the element at position 2 is whether the second sample drug exhibits toxicity in clinical trials. The value of the element at position 2 is 1. Based on the value of this element and the property corresponding to its position, another sub-property of the second sample drug is obtained: the second sample drug is toxic in clinical trials. Combining these two sub-properties yields the drug property of the second sample drug. This means that the drug property of the second sample drug is that the second sample drug is not FDA-approved and is toxic in clinical trials. Similarly, if the drug property label is [0,0], the drug property of the second sample drug is that it is not FDA-approved and is non-toxic in clinical trials. If the drug property label is [1,1], the drug property of the second sample drug is that the second sample drug is FDA-approved and is toxic in clinical trials. If the drug property label is [1,0], the drug property of the second sample drug is that the second sample drug is FDA-approved and is non-toxic in clinical trials.
[0049] In some embodiments, when the target property type is the toxicity type of the drug to be predicted, the second data set is the Tox21 data set, and the drug property prediction model corresponding to the target property type is a multi-label classification model. The drug property label carried by the second sample drug is a 12-dimensional array. The position of each element in the array represents a toxicity. The value of each element in the array is a first preset value or a second preset value. Based on the properties corresponding to the values of each element and the positions of each element, multiple sub-properties of the second sample drug are obtained. The sub-properties are merged to obtain the drug properties of the second sample drug.
[0050] In some embodiments, when the target property type is the type of side effect source of the drug to be predicted, the second data set is a SIDER data set, and the drug property prediction model corresponding to the target property type is a multi-label classification model. The drug property label carried by the second sample drug is a 27-dimensional array. The position of each element in the array represents a side effect source. The value of each element in the array is a first preset value or a second preset value. Based on the value of each element and the property corresponding to the position of each element, multiple sub-properties of the second sample drug are obtained. The sub-properties are merged to obtain the drug properties of the second sample drug.
[0051] In some embodiments, when the target property type is a quantum mechanical property of the drug to be predicted, the second dataset is a QM8 dataset, and the drug property prediction model corresponding to the target property type is a regression task model. The drug property labels carried by the second sample drug are a 12-dimensional array. The position of each element in the array represents a quantum mechanical property of the second sample drug, such as the electronic energy spectrum of the second sample drug and the amount of excited state energy of the second sample drug.
[0052] Optionally, determining the drug property of the drug to be predicted according to the drug property prediction data and the target property type includes: when the drug property prediction data is a numerical value, determining the drug property of the drug to be predicted according to the target property type.
[0053] In some embodiments, the target property type is whether the drug to be predicted contains a compound with the ability to penetrate the blood-brain barrier. The drug property prediction data is a numerical value. And the numerical value is less than or equal to the first preset numerical value and the numerical value is greater than or equal to the second preset numerical value. For example, the first preset numerical value is 1 and the second preset numerical value is 0. When the drug property prediction data is greater than or equal to the third preset numerical value, it is determined that the drug property of the drug to be predicted is that the drug to be predicted contains a compound with the ability to penetrate the blood-brain barrier. When the drug property prediction data is less than or equal to the third preset numerical value, it is determined that the drug property of the drug to be predicted is that the drug to be predicted does not contain a compound with the ability to penetrate the blood-brain barrier. The third preset numerical value is less than the first preset numerical value and the third preset numerical value is greater than the second preset numerical value.
[0054] In some embodiments, the target property type is whether the drug to be predicted can be used as a human β-secretase 1 inhibitor. The drug property prediction data is a numerical value. This numerical value is less than or equal to a first preset numerical value and greater than or equal to a second preset numerical value. For example, the first preset numerical value is 1 and the second preset numerical value is 0. If the drug property prediction data is greater than or equal to a third preset numerical value, the drug property of the drug to be predicted is determined to be a compound that can be used as a human β-secretase 1 inhibitor. If the drug property prediction data is less than or equal to the third preset numerical value, the drug property of the drug to be predicted is determined to be a compound that cannot be used as a human β-secretase 1 inhibitor.
[0055] In some embodiments, the target property type is the solubility of the drug to be predicted. The drug property prediction data is a numerical value, and the target property type is combined with the numerical value to obtain the drug property of the drug to be predicted. That is, the drug property of the drug to be predicted is determined to be the solubility of the drug to be predicted being the numerical value. For example, if the drug property prediction data is 48, the solubility of the drug to be predicted is combined with the numerical value 48 to obtain the solubility of the drug to be predicted being 48, and the drug property of the drug to be predicted is the solubility of the drug to be predicted being 48.
[0056] In some embodiments, the target property type is the atomization energy of a stable and comprehensively accessible organic molecule possessed by the drug to be predicted. The drug property prediction data is a numerical value. The target property type is combined with the numerical value to obtain the drug property of the drug to be predicted. That is, the drug property of the drug to be predicted is determined to be the atomization energy of a stable and comprehensively accessible organic molecule possessed by the drug to be predicted having the numerical value.
[0057] In some embodiments, the target property type is the hydration free energy of a small molecule in water obtained by experiments and alchemical free energy for the drug to be predicted. The drug property prediction data is a numerical value. The target property type is combined with the numerical value to obtain the drug property of the drug to be predicted. That is, the drug property of the drug to be predicted is determined to be the numerical value of the hydration free energy of a small molecule in water obtained by experiments and alchemical free energy for the drug to be predicted.
[0058] Optionally, determining the drug property of the drug to be predicted based on the drug property prediction data and the target property type includes: if the drug property prediction data is an array, determining the property corresponding to the position of each element in the array based on the target property type; obtaining multiple candidate properties of the drug to be predicted based on the values of each element in the array and the properties corresponding to the positions of each element; and combining the candidate properties to obtain the drug property of the drug to be predicted.
[0059] In some embodiments, the target property type is the type of side effect source of the drug to be predicted. The drug property prediction data is a 27-dimensional array. The value of each element in the array is less than or equal to a first preset value and the value is greater than or equal to a second preset value. For example, the first preset value is 1 and the second preset value is 0. The property corresponding to the position of each element in the array is determined according to the target property type, that is, the position of each element in the array represents a side effect source. When the value of the element is greater than or equal to a third preset value, one of the alternative properties of the drug to be predicted is determined to be: the drug to be predicted has the side effect source corresponding to the position of the element. When the value of the element is less than the third preset value, one of the alternative properties of the drug to be predicted is determined to be: the drug to be predicted does not have the side effect source corresponding to the position of the element. Combine the alternative properties to obtain the drug property of the drug to be predicted.
[0060] In some embodiments, the target property type is the FDA approval status of the drug to be predicted and whether it has toxicity in clinical trials. The drug property prediction data is a two-dimensional array. The property corresponding to each element position in the array is determined based on the target property type, i.e., each element position in the array represents a property. One element represents the FDA approval status of the drug, for example, FDA-approved or FDA-not approved. Another element represents whether the drug has toxicity in clinical trials. Based on the values of each element in the array and the properties corresponding to each element position, multiple candidate properties of the drug to be predicted are obtained. For example, if the value of the element at the position representing the FDA approval status of the drug is greater than or equal to a third preset value, one of the candidate properties of the drug to be predicted is determined to be: FDA approval of the drug to be predicted. If the value of the element at the position representing the FDA approval status of the drug is less than the third preset value, one of the candidate properties of the drug to be predicted is determined to be: FDA non-approval of the drug to be predicted. If the value of the element at the position representing the toxicity of the drug in clinical trials is greater than or equal to the third preset value, one of the candidate properties of the drug to be predicted is determined to be: toxicity in clinical trials of the drug to be predicted. If the value of the element at the position representing whether the drug was toxic in clinical trials is greater than or equal to a third preset value, one of the candidate properties of the drug to be predicted is determined to be: the drug to be predicted was non-toxic in clinical trials. The candidate properties are combined to obtain the drug property of the drug to be predicted. For example, if the drug property prediction data is [0.1, 0.8] and the third preset value is 0.6, the property corresponding to the position of the element at position 1 is the FDA-approved status of the drug to be predicted. The value of the element at position 1 is 0. If the value 0.1 of the element at the position representing the FDA-approved status of the drug is less than the third preset value 0.6, one of the candidate properties of the drug to be predicted is determined to be: the drug to be predicted was not FDA-approved. If the value 0.8 of the element at the position representing whether the drug was toxic in clinical trials is greater than or equal to the third preset value 0.6, one of the candidate properties of the drug to be predicted is determined to be: the drug to be predicted was toxic in clinical trials. Combining the two candidate properties determines that the drug to be predicted was not FDA-approved and was toxic in clinical trials.
[0061] In some embodiments, the target property type is the toxicity type of the drug to be predicted. The drug property prediction data is a 12-dimensional array. The value of each element in the array is less than or equal to a first preset value and the value is greater than or equal to a second preset value. For example, the first preset value is 1 and the second preset value is 0. The property corresponding to the position of each element in the array is determined according to the target property type, that is, the position of each element in the array represents a toxicity. When the value of the element is greater than or equal to a third preset value, one of the alternative properties of the drug to be predicted is determined to be: the drug to be predicted has the toxicity corresponding to the position of the element. When the value of the element is less than the third preset value, one of the alternative properties of the drug to be predicted is determined to be: the drug to be predicted does not have the toxicity corresponding to the position of the element. Combine the alternative properties to obtain the drug property of the drug to be predicted.
[0062] In some embodiments, the target property type is a quantum mechanical property of the drug to be predicted. The drug property prediction data is a 12-dimensional array. The property corresponding to the position of each element in the array is determined based on the target property type. That is, each element position in the array represents a quantum mechanical property. The value of each element is combined with the property corresponding to the position of each element to obtain the drug property to be predicted.
[0063] Optionally, a drug property prediction model is obtained by the following method, including: obtaining a first sample drug molecular sequence corresponding to each first sample drug, a first sample drug molecular graph corresponding to each first sample drug, a second sample drug molecular sequence corresponding to each second sample drug, and a second sample drug molecular graph corresponding to each second sample drug. Each first sample drug molecular sequence and each first sample drug molecular graph are input into a preset first deep learning model for unsupervised training to obtain a first candidate drug property prediction model. Model parameters of the first candidate drug property prediction model are transferred into a preset second deep learning model to obtain a second candidate drug property prediction model. Each second sample drug molecular sequence, each second sample drug molecular graph, and drug property labels of each second sample drug are input into the second candidate drug property prediction model for supervised training to obtain a drug property prediction model. Thus, by inputting each first sample drug molecular sequence and each first sample drug molecular graph into the preset first deep learning model for unsupervised training, a first candidate drug property prediction model is obtained. The model parameters of the first candidate drug property prediction model are then used to obtain a second candidate drug property prediction model. The second candidate drug property prediction model is then trained using the molecular sequences, molecular graphs, and drug property labels of each second sample drug as input to a supervised training process, yielding a drug property prediction model. Unsupervised training is performed first, allowing for the initial learning of shallow features of the molecular sequences of the first sample drugs and the corresponding molecular graphs of each first sample drug. Supervised training is then performed, leveraging both the first sample drugs without drug property labels and the second sample drugs with drug property labels, enabling the trained drug property prediction model to extract more and richer features.
[0064] The first sample drug molecule sequence corresponding to the first sample drug is a smiles representation of the drug molecule of the first sample drug. The second sample drug molecule sequence corresponding to the second sample drug is a smiles representation of the drug molecule of the second sample drug.
[0065] The nodes in the first sample drug molecule graph corresponding to the first sample drug are used to represent atoms of the drug molecule of the drug to be predicted. The edges in the first sample drug molecule graph are used to represent chemical bonds between atoms of the drug molecule of the first sample drug. The nodes in the second sample drug molecule graph corresponding to the second sample drug are used to represent atoms of the drug molecule of the drug to be predicted. The edges in the second sample drug molecule graph are used to represent chemical bonds between atoms of the drug molecule of the second sample drug.
[0066] Optionally, the preset first deep learning model includes a preset first encoder and a preset first graph neural network. Each first sample drug molecule sequence and each first sample drug molecule graph are input into the preset first deep learning model for unsupervised training to obtain a first candidate drug property prediction model, including: inputting each first sample drug molecule sequence into the preset first encoder for training to obtain a first molecular formula embedding vector; inputting each first sample drug molecule graph into the preset first graph neural network for training to obtain a first molecular graph embedding vector. The first candidate drug property prediction model is obtained based on the first molecular formula embedding vector and the first molecular graph embedding vector. In this way, the first encoder can be trained using the first sample drug molecule sequence, and the first graph neural network can be trained using the first sample drug molecule graph. Unsupervised training of the first deep learning model is achieved. Unsupervised training is performed using the first sample drug that does not carry a drug property label. This enables the first candidate drug property prediction model to more accurately extract the features of more drugs.
[0067] The dimension of the first molecular formula embedding vector is the same as the dimension of the first molecular graph embedding vector.
[0068] The preset first encoder is a Transformer encoder. The preset first graph neural network is a GNN (Graph Neural Network) network. The Transformer encoder is used to embed the first sample drug molecule sequence into a low-dimensional first molecular graph embedding vector using natural language processing technology. The GNN graph neural network is used to obtain the first molecular graph embedding vector based on the first sample drug molecule graph using neighbor aggregation.
[0069] In some embodiments, a GNN network is a network that executes algorithms that use neural networks to learn graph-structured data, extract and discover features and patterns in graph-structured data, and meet the requirements of graph learning tasks such as clustering, classification, prediction, segmentation, and generation. Natural Language Processing (NLP) is a method that uses computer technology to analyze, understand, and process string-type data, such as the sequence of a first sample drug molecule, by extracting features from languages in different fields and performing learning.
[0070] Optionally, each first sample drug molecule sequence is input into a preset first encoder for training to obtain a first molecular formula embedding vector, including encoding the atomic properties and atomic positions of each atom in the first sample drug molecule sequence, respectively, to obtain a first sample atomic property matrix and a first sample atomic position matrix corresponding to the first sample drug molecule sequence. An attention mechanism is then used to obtain the first molecular formula embedding vector based on the first sample atomic property matrix and the first sample atomic position matrix. Thus, by encoding the atomic properties and atomic positions of each atom in the first sample drug molecule sequence of the first sample drug that does not carry a drug property label, a first sample atomic property matrix and a first sample atomic position matrix corresponding to the first sample drug molecule sequence are obtained. The first molecular formula embedding vector is then obtained based on the first sample atomic property matrix and the first sample atomic position matrix using the attention mechanism. This enables the first encoder to accurately extract features of the first sample drug molecule sequence, thereby improving the richness of features of the drug to be predicted obtained from the ultimately trained drug property prediction data.
[0071] Furthermore, the atomic properties and atomic positions of each atom in the first sample drug molecule sequence are encoded respectively to obtain a first sample atomic property matrix corresponding to the first sample drug molecule sequence and a first sample atomic position matrix corresponding to the first sample drug molecule sequence, including: encoding the atomic properties of each atom in the first sample drug molecule sequence to obtain a first sample atomic property matrix corresponding to the first sample drug molecule sequence. Encoding the atomic positions of each atom in the first sample drug molecule sequence to obtain a first sample atomic position matrix corresponding to the first sample drug molecule sequence. The atomic properties of the atom include: one or more of the atomic number, the charge number around the atom, the atomic chirality type, the orbital hybridization type of the atom, the number of hydrogen atoms around the atom, the atomic degree, and the type of chemical bond connected to the atom. The atomic number of the atom is the number of protons in the atomic nucleus. The types of chemical bonds include: single bonds, double bonds, triple bonds, or aromatic bonds.
[0072] In some embodiments, the atomic properties of an atom are the atomic number of the atom and the number of charges surrounding the atom. The electronic device encodes the atomic properties of the carbon atoms in the first sample drug molecule sequence. The atomic number of the carbon atom is 6. The number of charges surrounding the carbon atom is 4. Therefore, the atomic property vector of the carbon atoms in the first sample drug molecule sequence is [6, 4]. After obtaining the atomic property vectors of each atom in the first sample drug molecule sequence, the atomic property vectors of each atom are concatenated to obtain a first sample atomic property matrix corresponding to the first sample drug molecule sequence.
[0073] Furthermore, encoding the atomic position of each atom in the first sample drug molecule sequence to obtain the first sample atomic position matrix corresponding to the first sample drug molecule sequence includes: calculating p k,2i =sin(k / 10000 2i / d ), obtain the first position code of the kth atom in the first sample drug molecule sequence. k,2i is the first position code of the kth atom in the first sample drug molecule sequence. i is the position reference data, which is used to characterize whether the kth atom is in an odd position or an even position in the first sample drug molecule sequence. d is the preset dimension data, which is used to characterize the dimension of the atomic position vector of the kth atom. By calculating p k,2i+1 =cos*(k / 10000 2i / d ) obtain the second position code of the kth atom in the first sample drug molecule sequence. k,2i+1 The second position code of the kth atom in the first sample drug molecule sequence is obtained. The atomic position vector of each atom is obtained based on the first position code and the second position code of each atom. The atomic position vectors of each atom are concatenated to obtain the first sample atomic position matrix corresponding to the first sample drug molecule sequence.
[0074] Furthermore, an attention mechanism is used to obtain a first molecular formula embedding vector based on the first sample atomic property matrix and the first sample atomic position matrix, including: multiplying the first sample matrix by a preset query weight matrix to obtain a query matrix. Multiplying the first sample matrix by a preset bond weight matrix to obtain a bond matrix. Multiplying the first sample matrix by a preset value weight matrix to obtain a value matrix. The first molecular formula embedding vector is obtained based on the query matrix, the bond matrix, and the value matrix. The first sample matrix is a matrix obtained by concatenating the first sample atomic property matrix and the first sample atomic position matrix.
[0075] Further, obtaining the first molecular formula embedding vector according to the query matrix, the key matrix and the value matrix includes: calculating Obtain the first molecular formula embedding vector. Where Attention(Q,K,V) is the first molecular formula embedding vector. is the attention matrix. Q is the query matrix. K is the key matrix. V is the value matrix. d K is the dimension of the K matrix. T is the transpose symbol. Softmax is a preset normalization algorithm used to make the sum of the attention weights of each atom and other atoms equal to 1. The attention matrix is transformed into a standard normal distribution and then normalized using softmax, ensuring that the embedding of each atom in the attention matrix contains information about all other atoms. Multiplying the attention matrix with the value matrix allows attention to be paid to all atoms in the drug molecule sequence, improving the accuracy and richness of the features captured in the drug molecule sequence.
[0076] Optionally, each first sample drug molecule graph is input into a preset first graph neural network for training to obtain a first molecular graph embedding vector, including: obtaining node information of each node in each first sample drug molecule graph. Aggregating each node information to obtain the first molecular graph embedding vector. The first graph neural network has a multi-layer network. In this way, the first sample drug molecule graph of the first sample drug without a drug property label can be used to train the first graph neural network. This allows the drug property prediction data finally trained to extract more node information of each node in the drug molecule graph and obtain the first molecular graph embedding vector based on it. This improves the accuracy of the finally trained drug property prediction data in obtaining the characteristics of the drug to be predicted.
[0077] Further, obtaining the node information of each node in each first sample drug molecule graph includes: calculating Get the node information of the vth node in the fth layer network. is the reference node information of the vth node in the fth layer network. σ is the preset activation function. W (f) It is the network weight matrix preset for the f-th layer network. is the aggregate information of the vth node in the fth layer network. f is the bias matrix preset for the f-th layer network. Perform dimension transformation to achieve Linear transformation of .
[0078] Further, Obtained by the following method: Get the aggregate information of the vth node in the fth layer network. is the aggregate information of the v-th node in the f-th layer network. is the reference node information of the vth node in the j-1th layer network. uv is the embedding representation of the edge connecting the vth node and the uth node. is the reference node information of the u-th node in the j-1 layer network. v is the set of adjacent nodes of the vth node. ∈ is the belonging symbol. u∈N v Representing the u-th node as set N v AGGREGATE is the first preset aggregation function.
[0079] Thus, through the formula and formula The node information of the next layer of the network can be obtained according to the node information of the previous layer of the first graph neural network, thereby realizing the iterative update of the node information of the nodes corresponding to each atom of the molecule of the first sample drug.
[0080] Furthermore, the information of each node is aggregated to obtain the first molecular graph embedding vector, including: calculating Get the first molecular graph embedding vector. For the default. h G is the embedding vector of the first molecular graph. F is the number of layers in the first graph neural network. V is the set of nodes in the first sample drug molecular graph. READOUT is a preset second aggregation function. The preset second aggregation function is used for summing, averaging, or maximizing. In this way, by using READOUT, the node information of the nodes corresponding to each atom in the molecules of the first sample drug can be aggregated to obtain the embedding vector of the first molecular graph.
[0081] Optionally, obtaining a first candidate drug property prediction model based on the first molecular formula embedding vector and the first molecular graph embedding vector includes: obtaining a contrast loss value based on the first molecular formula embedding vector and the first molecular graph embedding vector. Obtaining the first candidate drug property prediction model based on the contrast loss value. In this way, obtaining the contrast loss value based on the first molecular formula embedding vector and the first molecular graph embedding vector enables comparative learning of data from two modalities, the first molecular formula embedding vector and the first molecular graph embedding vector. Compared to learning unimodal data, multimodal learning facilitates the discovery of drug patterns from the first sample drug that does not carry drug property labels, thereby obtaining richer and more accurate features.
[0082] Furthermore, a contrast loss value is obtained based on the first molecular formula embedding vector and the first molecular graph embedding vector, including: obtaining a contrast loss value corresponding to the third alternative drug property prediction model based on the first molecular formula embedding vector and the first molecular graph embedding vector through a preset contrast loss algorithm.
[0083] Furthermore, the contrast loss value corresponding to the third candidate drug property prediction model is obtained according to the first molecular formula embedding vector and the first molecular graph embedding vector by a preset contrast loss algorithm, including: calculating Obtain the comparative loss value corresponding to the third candidate drug property prediction model. Con is the comparative loss value corresponding to the third candidate drug property prediction model. N1 is the total number of the first sample drugs input. sim() is a preset similarity function. For example, the cosine similarity function. x i is the first molecular formula embedding vector of the i-th first sample drug.i y is the first molecular graph embedding vector of the i-th first sample drug. j is the first molecular graph embedding vector of the jth first sample drug. In this way, the first molecular formula embedding vector of the i-th first sample drug and the first molecular graph embedding vector of the i-th first sample drug are used as positive samples, and the first molecular formula embedding vector of the i-th first sample drug and the first molecular graph embedding vector of the j-th first sample drug are used as negative samples for comparative learning. At the same time, the first half of the formula The comparison between the embedding vector of the first molecular formula of the i-th element and the embedding vector of all first molecular graphs is realized. The jth first molecular graph embedding vector is compared with all first molecular formula embedding vectors. Multimodal comparative learning using data from both first molecular formula embedding vectors and first molecular graph embedding vectors facilitates drug discovery from first samples without drug property labels, resulting in richer and more accurate features.
[0084] Furthermore, obtaining a first candidate drug property prediction model based on the contrast loss value includes: if the contrast loss value does not stabilize, optimizing the first deep learning model until the contrast loss value stabilizes. If the contrast loss value stabilizes, pausing optimization of the first deep learning model and determining the first deep learning model as the first drug property prediction model. In this way, training of the first drug property prediction model can be determined to be complete when the contrast loss value stabilizes, i.e., when the contrast loss value converges.
[0085] Furthermore, optimizing the first deep learning model includes: continuing to train the first deep learning model to adjust model parameters of the first deep learning model.
[0086] Furthermore, optimizing the first deep learning model includes: optimizing the first deep learning model by adjusting the model parameters of the first deep learning model. The model parameters of the first deep learning model include: query weight matrix, key weight matrix, value weight matrix, d K 、W (f) and b f wait.
[0087] Furthermore, migrating the model parameters of the first candidate drug property prediction model to a preset second deep learning model to obtain the second candidate drug property prediction model includes: obtaining the model parameters of the first candidate drug property prediction model; and determining the model parameters of the first candidate drug property prediction model as the model parameters of the preset second deep learning model.
[0088] In some embodiments, combined Figure 2As shown, the preset first deep learning model 1 includes a preset first encoder 2 and a preset first graph neural network 3. The electronic device inputs each first sample drug molecule sequence into the preset first encoder for training to obtain a first molecular formula embedding vector. Each first sample drug molecule graph is input into the preset first graph neural network for training to obtain a first molecular graph embedding vector. According to the output of the first deep learning model, a contrast loss value is output based on the first molecular formula embedding vector and the first molecular graph embedding vector. Then, a third alternative drug property prediction model corresponding to the contrast loss value is obtained, and the third alternative drug property prediction model is optimized based on the contrast loss value to obtain a first alternative drug property prediction model.
[0089] Optionally, each second sample drug molecular sequence, each second sample drug molecular graph, and each second sample drug drug's drug property label is input into a second candidate drug property prediction model for supervised training to obtain a drug property prediction model, including: determining a training set and a test set from each second sample drug molecular sequence, each second sample drug molecular graph, and each drug property label in a preset ratio; the training set includes multiple third sample drug molecular sequences, multiple third sample drug molecular graphs, and multiple training drug labels; the test set includes multiple fourth sample drug molecular sequences, multiple fourth sample drug molecular graphs, and multiple test property labels; each third sample drug molecular sequence is input into a second encoder for training to obtain a second molecular formula embedding vector; each third sample drug molecular graph is input into a second graph neural network for training to obtain a second molecular graph embedding vector; the second molecular formula embedding vector and the second molecular graph embedding vector are input into a preset fully connected layer network for training to obtain a first predicted property label; and a drug property prediction model is obtained based on the first predicted property label, the drug property label, and the test set. In this way, the second candidate drug property prediction model can be supervised trained using the second sample drug molecular sequences, the second sample drug molecular graphs, and the drug property labels of the second sample drugs. The drug property prediction model obtained through training can extract the characteristics of the drug and predict the properties of the drug based on the characteristics.
[0090] In some embodiments, the preset ratio is 9:1. Nine-tenths of the sample drug molecule sequences are randomly selected from each second sample drug molecule sequence. The selected second sample drug molecule sequence is determined as the third sample drug molecule sequence. The second sample drug molecule graph corresponding to the selected second sample drug molecule sequence is determined as the third sample drug molecule graph. The drug property label corresponding to the selected second sample drug molecule sequence is determined as the training drug label. A training set is obtained, including multiple third sample drug molecule sequences, multiple third sample drug molecule graphs, and multiple training drug labels. The unselected second sample drug molecule sequence is determined as the fourth sample drug molecule sequence. The second sample drug molecule graph corresponding to the unselected second sample drug molecule sequence is determined as the fourth sample drug molecule graph. The drug property label corresponding to the unselected second sample drug molecule sequence is determined as the test property label. A test set is obtained, including multiple fourth sample drug molecule sequences, multiple fourth sample drug molecule graphs, and multiple test property labels. The ratio of the training set to the test set is the preset ratio, i.e., 9:1.
[0091] In some embodiments, the training method for inputting each third sample drug molecule sequence into the second encoder for training to obtain the second molecular formula embedding vector is the same as the training method for inputting each first sample drug molecule sequence into the preset first encoder for training to obtain the first molecular formula embedding vector, and is not further described here. The training method for inputting each third sample drug molecule graph into the second graph neural network for training to obtain the second molecular graph embedding vector is the same as the training method for inputting each first sample drug molecule graph into the preset first graph neural network for training to obtain the first molecular graph embedding vector, and is not further described here.
[0092] Furthermore, the second molecular formula embedding vector and the second molecular graph embedding vector are input into a preset fully connected layer network for training, including: utilizing the preset fully connected layer network to perform dimensionality reduction on the second molecular formula embedding vector and the second molecular graph embedding vector to obtain a predicted property label. Wherein, if the drug property label is a numerical value, the predicted property label is a numerical value. If the drug property label is an array, the predicted property label is an array. Furthermore, the dimension of the predicted property label is the same as the dimension of the drug property label.
[0093] Furthermore, obtaining a drug property prediction model based on the first predicted property label, the training drug label, and the test set includes: obtaining a label loss value based on the first predicted probability and the training drug label; obtaining a third candidate drug property prediction model based on the label loss value; and evaluating the third candidate drug property prediction model using the test set to obtain the drug property prediction model.
[0094] Combine Figure 3 As shown, Figure 3To obtain a schematic diagram of the third candidate drug property prediction model, the second candidate drug property prediction model 4 includes: a second encoder 5, a second graph neural network 6, and a preset fully connected layer network 7. The model parameters of the second encoder 5 are the model parameters of the first encoder 2. The model parameters of the second graph neural network 6 are the model parameters of the third graph neural network 3. The electronic device inputs each third sample drug molecule sequence into the preset second encoder for training to obtain a second molecular formula embedding vector. Each third sample drug molecule graph is input into the preset second graph neural network for training to obtain a second molecular graph embedding vector. The second molecular formula embedding vector and the second molecular graph embedding vector are then input into the preset fully connected layer network to obtain a first predicted property label. A label loss value is then obtained based on the first predicted property label and the training drug label. The second candidate drug property prediction model is optimized based on the label loss value to obtain a third candidate drug property prediction model. The second encoder is a Transformer encoder. The second graph neural network is a GNN network.
[0095] The fully connected layer network can adapt to classification or regression tasks. When the second dataset is the ESOL, QM7, QM8, or FreeSolv dataset, the fully connected layer network adapts to fit the second molecular formula embedding vector and the second molecular graph embedding vector into a straight line. This results in the trained drug property prediction model being a regression task model.
[0096] When the second dataset is the BBBP dataset, the BACE dataset, the ClinTox dataset, the Tox21 dataset, or the SIDER dataset, the fully connected layer network performs classification based on the second molecular formula embedding vector and the second molecular graph embedding vector after adaptation, so that the trained drug property prediction model is a binary classification model or a multi-label classification model.
[0097] Optionally, obtaining a label loss value based on the first predicted property label and the training drug label includes: when the second dataset is a BBBP dataset, a BACE dataset, a ClinTox dataset, a Tox21 dataset, or a SIDER dataset, determining real data based on the training drug label; determining a predicted probability based on the first predicted property label; and obtaining a label loss value based on the predicted probability and the real data.
[0098] In some embodiments, when the second data set is a BBBP data set or a BACE data set, the drug property prediction model corresponding to the target property type is a binary classification model. The predicted property label is the probability obtained by predicting the drug property label. The training drug label is determined as the real data. The predicted probability is determined based on the first predicted property label. When the second data set is a ClinTox data set, a Tox21 data set, or a SIDER data set, the drug property prediction model corresponding to the target property type is a multi-label classification model. The training drug label and the first predicted property label are both arrays. The training drug label and the first predicted property label are converted into vectors. The vector corresponding to the training drug label is determined as the real data. The vector corresponding to the first predicted property label is determined as the predicted probability.
[0099] Furthermore, the label loss value is obtained based on the predicted probability and the actual probability, including: Obtain the label loss value. Where L is the label loss value. N2 is the number of drugs in the second sample. t is the true value of the t-th second sample drug. t is the predicted probability of the drug in the tth second sample. In this way, the label loss value is obtained by calculating the cross entropy loss value based on the predicted probability and the actual probability.
[0100] Optionally, obtaining a label loss value according to the predicted property label and the drug property label includes: when the second data set is an ESOL data set, a QM7 data set, a QM8 data set or a FreeSolv data set, by calculating Get the label loss value. Where L is the label loss value. N2 is the number of drugs in the second sample. t is the drug property label of the t-th second sample drug. t is the predicted property label of the t-th second sample drug. In this way, the label loss value is obtained by calculating the root mean square error based on the predicted probability and the actual probability.
[0101] Furthermore, a third candidate drug property prediction model is obtained based on the label loss value, including: if the label loss value does not stabilize, optimizing the second candidate drug property prediction model until the label loss value stabilizes. If the label loss value stabilizes, suspending the optimization of the second candidate drug property prediction model, and determining the second candidate drug property prediction model as the third candidate drug property prediction model. In this way, when the label loss value stabilizes, that is, when the label loss value converges, it can be determined that the drug property prediction model training is complete.
[0102] Furthermore, the second candidate drug property prediction model is optimized, including: continuing to train the second candidate drug property prediction model to adjust model parameters of the second candidate drug property prediction model.
[0103] Furthermore, the third candidate drug property prediction model is evaluated using the test set to obtain a drug property prediction model, including: inputting each fourth sample drug molecular sequence and each fourth sample drug molecular graph into the third candidate drug property prediction model to obtain a second predicted property label. A prediction accuracy is obtained based on the test drug label and the second predicted property label. When the accuracy is greater than or equal to the preset accuracy, the third candidate drug property prediction model is determined as the drug property prediction model. When the accuracy is less than the preset accuracy, the second sample drug molecular sequence, each second sample drug molecular graph, and the drug property label of each second sample drug are re-input into the second candidate drug property prediction model for supervised training to obtain a drug property prediction model.
[0104] Further, obtaining a prediction accuracy rate based on the test drug label and the second predicted property label includes: calculating Obtain prediction accuracy. Where Z is the prediction accuracy. X is the number of second predictive property labels equal to the test drug label, i.e., the number of accurately predicted second predictive property labels. Y is the number of test drug labels. Thus, by obtaining prediction accuracy based on the test drug label and the second predictive property label, the prediction accuracy of the third candidate drug property prediction model can be evaluated. This ensures the accuracy of the trained drug property prediction model.
[0105] Combine Figure 4 As shown, the embodiment of the present disclosure provides a method for obtaining a drug property prediction model, comprising:
[0106] In step S201 , the electronic device obtains a first sample drug molecular sequence corresponding to each first sample drug, a first sample drug molecular graph corresponding to each first sample drug, a second sample drug molecular sequence corresponding to each second sample drug, and a second sample drug molecular graph corresponding to each second sample drug.
[0107] In step S202, the electronic device inputs each first sample drug molecule sequence into a preset first encoder for training to obtain a first molecular formula embedding vector; and inputs each first sample drug molecule graph into a preset first graph neural network for training to obtain a first molecular graph embedding vector. The first encoder and the first graph neural network are in a preset first deep learning model.
[0108] In step S203 , the electronic device obtains a contrast loss value according to the first molecular formula embedding vector and the first molecular graph embedding vector.
[0109] Step S204: The electronic device obtains a first candidate drug property prediction model according to the comparison loss value.
[0110] In step S205, the electronic device transfers the model parameters of the first candidate drug property prediction model to a preset second deep learning model to obtain a second candidate drug property prediction model. The second candidate drug property prediction model includes a second encoder, a second graph neural network, and a preset fully connected layer network.
[0111] In step S206, the electronic device determines a training set and a test set from each second sample drug molecular sequence, each second sample drug molecular graph, and each drug property label according to a preset ratio; the training set includes multiple third sample drug molecular sequences, multiple third sample drug molecular graphs, and multiple training drug labels.
[0112] In step S207, the electronic device inputs each third sample drug molecule sequence into the second encoder for training to obtain a second molecular formula embedding vector; and inputs each third sample drug molecule graph into the second graph neural network for training to obtain a second molecular graph embedding vector.
[0113] In step S208 , the electronic device inputs the second molecular formula embedding vector and the second molecular graph embedding vector into a preset fully connected layer network for training to obtain a first predicted property label.
[0114] In step S209 , the electronic device obtains a label loss value based on the first predicted probability and the training drug label.
[0115] In step S210 , the electronic device obtains a third candidate drug property prediction model according to the label loss value.
[0116] In step S211 , the electronic device evaluates the third candidate drug property prediction model using the test set to obtain a drug property prediction model.
[0117] Using the method for obtaining a drug property prediction model provided by the embodiments of the present disclosure, a first candidate drug property prediction model is obtained by performing unsupervised training on a preset first deep learning model using the first sample drug molecular sequence corresponding to the first sample drug and the first sample drug molecular graph corresponding to each first sample drug. A second candidate drug property prediction model is then obtained based on the first candidate drug property prediction model. The second sample drug molecular sequence, the second sample drug molecular graph, and the drug property labels of the second sample drug are then divided into a training set and a test set. The second candidate drug property prediction model is then subjected to supervised training using the training set to obtain a third candidate drug property prediction model. The third candidate drug property prediction model is then evaluated using the test set to obtain a drug property prediction model. In this way, the unsupervised training enables the drug property prediction model obtained at the end of the training to extract more accurate and richer features of the drug molecular sequence and drug molecular graph of the drug to be predicted. The supervised training enables the drug property prediction model obtained at the end of the training to predict the property type of the drug to be predicted based on more accurate and richer features, thereby improving the accuracy of the drug property prediction model in predicting drug properties. At the same time, evaluation through the test set can help users more intuitively understand the properties of the third candidate drug property prediction model.
[0118] Combine Figure 5 As shown, an embodiment of the present disclosure provides a device 8 for predicting drug properties, including a processor 9 and a memory 10. Optionally, the device may also include a communication interface 11 and a bus 12. The processor 9, the communication interface 11, and the memory 10 can communicate with each other through the bus 12. The communication interface 11 can be used for information transmission. The processor 9 can call the logic instructions in the memory 10 to execute the method for predicting drug properties of the above embodiment.
[0119] In addition, the logic instructions in the memory 10 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0120] The memory 10, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods of the embodiments of the present disclosure. The processor 9 executes the program instructions / modules stored in the memory 10 to perform functional applications and data processing, thereby implementing the method for predicting drug properties in the above-mentioned embodiments.
[0121] The memory 10 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 10 may include high-speed random access memory and non-volatile memory.
[0122] The device for predicting drug properties provided by the embodiment of the present disclosure is used, and the drug molecular sequence and drug molecular graph of the drug to be predicted are input into the drug property prediction model. Then, the drug property of the drug to be predicted is determined based on the drug property prediction data and the property type of the drug to be predicted. Among them, the drug property prediction model is obtained through a plurality of first sample drugs that do not carry drug property labels and second sample drugs that carry drug property labels. In this way, compared with only using a very small number of drugs with known drug properties to obtain the drug property prediction model, this solution uses the first sample drugs that do not carry drug property labels to obtain the drug property prediction model, so that the drug property prediction model can extract richer features of the drug to be predicted, and can more accurately predict the drug properties of the drug to be predicted. Thereby, the accuracy of the drug property prediction model in predicting drug properties is improved.
[0123] An embodiment of the present disclosure provides an electronic device, comprising: an electronic device body, and the above-mentioned device for predicting drug properties. The device for predicting drug properties is installed on the electronic device body. The installation relationship described here is not limited to placement inside the electronic device, but also includes installation connections with other components of the electronic device, including but not limited to physical connections, electrical connections or signal transmission connections, etc. It can be understood by those skilled in the art that the device for predicting drug properties can be adapted to a feasible electronic device body, thereby realizing other feasible embodiments.
[0124] In some embodiments, the electronic device is a computer or a server.
[0125] The electronic device provided by the embodiment of the present disclosure is used to input the drug molecule sequence and drug molecule graph of the drug to be predicted into the drug property prediction model. Then, the drug property of the drug to be predicted is determined based on the drug property prediction data and the property type of the drug to be predicted. Among them, the drug property prediction model is obtained through multiple first sample drugs that do not carry drug property labels and second sample drugs that carry drug property labels. In this way, compared with only using a very small number of drugs with known drug properties to obtain the drug property prediction model, this solution uses the first sample drugs that do not carry drug property labels to obtain the drug property prediction model, so that the drug property prediction model can extract richer features of the drug to be predicted and can more accurately predict the drug properties of the drug to be predicted. Thereby, the accuracy of the drug property prediction model in predicting drug properties is improved.
[0126] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned method for predicting drug properties.
[0127] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0128] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient storage medium.
[0129] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.
[0130] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0131] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of the present disclosure may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0132] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for predicting drug properties, characterized in that include: Obtaining the drug molecular sequence and drug molecular map of the drug to be predicted; and determining the target property type; Inputting the drug molecule sequence and the drug molecule graph into a drug property prediction model corresponding to the target property type to obtain drug property prediction data; wherein the drug property prediction model corresponding to the target property type is obtained based on a sample drug dataset corresponding to the target property type; the sample drug dataset includes a plurality of first sample drugs and a plurality of second sample drugs; the first sample drugs are sample drugs without drug property labels; and the second sample drugs are sample drugs with drug property labels; Determine the drug properties of the drug to be predicted based on the drug property prediction data and the target property type; The drug property prediction model is obtained by the following method, including: obtaining a first sample drug molecular sequence corresponding to each first sample drug, a first sample drug molecular graph corresponding to each first sample drug, a second sample drug molecular sequence corresponding to each second sample drug, and a second sample drug molecular graph corresponding to each second sample drug; inputting each first sample drug molecular sequence and each first sample drug molecular graph into a preset first deep learning model for unsupervised training to obtain a first candidate drug property prediction model; migrating model parameters of the first candidate drug property prediction model into a preset second deep learning model to obtain a second candidate drug property prediction model; inputting each second sample drug molecular sequence, each second sample drug molecular graph, and drug property label of each second sample drug into the second candidate drug property prediction model for supervised training to obtain a drug property prediction model; The preset first deep learning model includes a preset first encoder and a preset first graph neural network; each first sample drug molecule sequence and each first sample drug molecule graph are input into the preset first deep learning model for unsupervised training to obtain a first candidate drug property prediction model, including: inputting each first sample drug molecule sequence into the preset first encoder for training to obtain a first molecular formula embedding vector; inputting each first sample drug molecule graph into the preset first graph neural network for training to obtain a first molecular graph embedding vector; and obtaining the first candidate drug property prediction model based on the first molecular formula embedding vector and the first molecular graph embedding vector. The second alternative drug property prediction model includes a second encoder, a second graph neural network and a preset fully connected layer network; each second sample drug molecular sequence, each second sample drug molecular graph and each second sample drug drug property label are input into the second alternative drug property prediction model for supervised training to obtain a drug property prediction model, including: determining a training set and a test set among each second sample drug molecular sequence, each second sample drug molecular graph and each drug property label according to a preset ratio; the training set includes multiple third sample drug molecular sequences, multiple third sample drug molecular graphs and multiple training drug labels; each third sample drug molecular sequence is input into the second encoder for training to obtain a second molecular formula embedding vector; each third sample drug molecular graph is input into the second graph neural network for training to obtain a second molecular graph embedding vector; the second molecular formula embedding vector and the second molecular graph embedding vector are input into the preset fully connected layer network for training to obtain a first predicted property label; and the drug property prediction model is obtained according to the first predicted property label, the training drug label and the test set.
2. The method according to claim 1, characterized in that Inputting each first sample drug molecule sequence into a preset first encoder for training to obtain a first molecular formula embedding vector, including: Encoding the atomic properties and atomic positions of each atom in the first sample drug molecule sequence respectively to obtain a first sample atomic property matrix corresponding to the first sample drug molecule sequence and a first sample atomic position matrix corresponding to the first sample drug molecule sequence; The attention mechanism is used to obtain the first molecular formula embedding vector according to the first sample atomic property matrix and the first sample atomic position matrix.
3. The method according to claim 1, characterized in that Inputting each first sample drug molecule graph into a preset first graph neural network for training to obtain a first molecule graph embedding vector, including: Obtaining node information of each node in each first sample drug molecule graph; Aggregate the information of each node to obtain the first molecular graph embedding vector.
4. The method according to claim 1, wherein Obtaining a first candidate drug property prediction model according to the first molecular formula embedding vector and the first molecular graph embedding vector, including: Obtaining a contrast loss value according to the first molecular formula embedding vector and the first molecular graph embedding vector; A first candidate drug property prediction model is obtained according to the comparative loss value.
5. A device for predicting drug properties, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the method for predicting drug properties according to any one of claims 1 to 4 when running the program instructions.
6. An electronic device, characterized in that: include: Electronic device body; The device for predicting drug properties according to claim 5 is installed in the electronic device body.
7. A storage medium storing program instructions, characterized in that: When the program instructions are executed, the method for predicting drug properties according to any one of claims 1 to 4 is executed.