Molecular property prediction model training method, storage medium, and property prediction device
By splitting the molecular graph into candidate graph hints, extracting target graph hints and fusing them into a global graph hint vector, and training a molecular property prediction model, the problems of data scarcity and task heterogeneity in existing technologies are solved, enabling multi-task molecular property prediction and improving drug development efficiency.
Patent Information
- Application Number
- CN202310799064.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing molecular property prediction methods based on graph neural networks face problems such as data scarcity, task heterogeneity, and lack of subgraph information, making it difficult for the models to generalize and adapt to different molecular property tasks.
By splitting the molecular graph into candidate graph cues, extracting target graph cues using preset rules, calculating importance weights and fusing them into a global graph cue vector, concatenating it with the molecular feature vector, and training a molecular property prediction model, multi-task learning and self-supervised learning are achieved.
It improves the model's generalization ability and efficiency, enabling it to predict multiple molecular properties simultaneously, shortening the drug development cycle, and increasing drug development efficiency.
Smart Images

Figure CN117059199B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of molecular property prediction and digital medical technology, and in particular to a molecular property prediction model training method, a storage medium, a computer device and a molecular property prediction device. BACKGROUND
[0002] Drug development cycle is long, large investment and high risk. Only small molecules that meet certain properties can become clinically useful drugs (toxicity, absorption, metabolic effect, etc.), usually each of these properties needs to be obtained through biochemical experiments, combined with computing technology, constructing an accurate and efficient drug molecular property prediction model can greatly reduce the dependence on experiments, reduce costs and speed up progress. Traditional methods based on molecular fingerprints and descriptors require a lot of professional knowledge for optimization design, and lack of universality and scalability. At present, deep learning models can automatically extract relevant features from raw molecular data and make significant breakthroughs in a large number of tasks.
[0003] Molecular property prediction is an important problem in computational chemistry, which is used in drug discovery, material design and other fields. At present, graph neural network (GNN) is an effective method for molecular representation learning, which can use the structural information of molecules to extract feature vectors and be used for downstream classification or regression tasks. However, due to the scarcity and high cost of labeled data, the generalization ability of GNN in the vast molecular space is limited. In order to solve this problem, a common training method is "pre-training, fine-tuning", that is, first pre-train GNN in a self-supervised manner, and then fine-tune GNN with a small amount of labeled data to adapt to specific molecular property tasks. The existing GNN-based molecular property prediction methods have the following problems: data scarcity, for some new or rare molecular properties, experimental data is often very limited, which makes it difficult for the model to learn effective feature representation and generalization ability; task heterogeneity, different molecular properties may have different complexity and difficulty, which requires different model parameters and structures to adapt; subgraph information loss, existing GNN-based methods usually only focus on node-level or graph-level tasks, ignoring the rich information contained in subgraphs or graph motifs. For example, functional groups (subgraphs often appearing in molecular graphs) usually carry indicative information about molecular properties.
[0004] In order to solve the problem of data scarcity, some studies have proposed methods based on self-supervised learning, which use unlabeled molecular data for pre-training to improve the performance of the model on downstream tasks. However, at present, collecting data and building models for each property separately cannot fully utilize the effective and learnable information between data and tasks, which is increasingly showing its shortcomings. SUMMARY
[0005] Therefore, the application provides a molecular property prediction model training method, a storage medium, a computer device and a molecular property prediction device.
[0006] According to an aspect of the application, a molecular property prediction model training method is provided, which comprises:
[0007] obtaining molecular graphs of a plurality of sample molecules, splitting the molecular graphs of the sample molecules to obtain a plurality of candidate graph prompts, and extracting a preset number of candidate graph prompts as target graph prompts according to a preset candidate graph prompt extraction rule in the candidate graph prompts;
[0008] encoding the target graph prompts into target graph prompt vectors, calculating importance weights of the target graph prompt vectors for the molecular graphs of the sample molecules, and calculating a global graph prompt vector according to the target graph prompt vectors and the importance weights;
[0009] splicing the global graph prompt vector and feature vectors of the sample molecules respectively to obtain molecular vectors of the sample molecules;
[0010] training a molecular property prediction model according to preset molecular property labels corresponding to the sample molecules and the molecular vectors.
[0011] Optionally, the extracting a preset number of candidate graph prompts as target graph prompts according to a preset candidate graph prompt extraction rule comprises:
[0012] determining a number of intra-group graph prompts based on the preset number;
[0013] constructing a plurality of candidate graph prompt combinations for each number of intra-group graph prompts, wherein each candidate graph prompt combination comprises the number of intra-group graph prompts;
[0014] calculating evaluation values of the candidate graph prompts in each candidate graph prompt combination, and calculating a combination evaluation value of the candidate graph prompt combination based on the evaluation values of the candidate graph prompts;
[0015] for each number of intra-group graph prompts, taking a candidate graph prompt combination corresponding to a maximum value in a plurality of combination evaluation values corresponding to the number of intra-group graph prompts as a target graph prompt combination of the number of intra-group graph prompts;
[0016] for each target graph prompt combination, taking a candidate graph prompt with the maximum evaluation value in the target graph prompt combination as a target graph prompt.
[0017] Optionally, the calculating the evaluation value of each candidate graph hint in the candidate graph hint combination and the combination evaluation value of the candidate graph hint combination based on the evaluation value of each candidate graph hint comprises:
[0018] calculating a quality score and a diversity score of each candidate graph hint in the candidate graph hint combination, wherein the quality score of any candidate graph hint is determined based on at least one of the frequency of occurrence of the candidate graph hint in all candidate graph hints corresponding to the sample molecule, the rarity of the candidate graph hint, and the complexity of the candidate graph hint, and the diversity score of any candidate graph hint is determined based on the structural difference and / or information complementarity between the candidate graph hint and other candidate graph hints corresponding to the sample molecule;
[0019] calculating the evaluation value of each candidate graph hint based on the quality score and the diversity score, and calculating the combination evaluation value of the candidate graph hint combination based on the evaluation value of each candidate graph hint in the candidate graph hint combination.
[0020] Optionally, the calculating the importance weight of each target graph hint vector to the molecular graph of each sample molecule comprises:
[0021] calculating the importance weight of each target graph hint vector to the molecular graph of each sample molecule according to a preset attention mechanism, and calculating the average value of the importance weight of each target graph hint vector to the molecular graph of different sample molecules;
[0022] taking the average value of the importance weight of each target graph hint vector as the weighted weight value of each target graph hint vector, and performing weighted fusion calculation on each target graph hint based on the weighted weight value to obtain the global graph hint vector.
[0023] Optionally, the concatenating the global graph hint vector with the feature vector of each sample molecule to obtain the molecular vector of each sample molecule comprises:
[0024] constructing a primary node feature matrix according to the atomic features of each atom in the molecular graph of the sample molecule, and constructing an adjacency matrix according to the force features between adjacent atoms in the molecular graph;
[0025] concatenating the global graph hint vector with the primary node feature matrix to obtain a secondary node feature matrix, and encoding the secondary node feature matrix and the adjacency matrix into a molecular vector.
[0026] Optionally, the training the molecular property prediction model according to the preset molecular property label corresponding to the sample molecule and the molecular vector comprises:
[0027] According to a plurality of predicted properties, a plurality of predicted property task nodes are constructed, a task relationship graph is constructed based on the correlation between the predicted property task nodes, and a task embedding vector of each predicted property task node is calculated;
[0028] For each predicted property task node, the degree of dependence between the predicted property task node and the related nodes is calculated according to the task embedding vector of the predicted property task node and the task embedding vectors of the related nodes of the predicted property task node;
[0029] According to the degree of dependence between different predicted property task nodes, the loss weight of each predicted property task node is determined, and the loss function of each predicted property task node is weighted calculated according to the loss weight of each predicted property task node, to obtain the loss function of the molecular property prediction model;
[0030] According to the loss function, the preset molecular property label corresponding to the sample molecule and the molecular vector, the molecular property prediction model is trained.
[0031] Optionally, the method further comprises:
[0032] obtaining a preset global graph prompt vector and a to-be-predicted molecule;
[0033] splicing the preset global graph prompt vector and the feature vector of the to-be-predicted molecule to obtain a to-be-predicted molecule vector;
[0034] According to the trained molecular property prediction model, the to-be-predicted molecule vector is subjected to molecular property prediction to obtain a molecular property prediction result of the to-be-predicted molecule.
[0035] According to another aspect of the present application, a molecular property prediction device is provided, the device comprising:
[0036] a data acquisition module configured to obtain a preset global graph prompt vector and a to-be-predicted molecule;
[0037] a data processing module configured to splice the preset global graph prompt vector and the feature vector of the to-be-predicted molecule to obtain a to-be-predicted molecule vector;
[0038] a property prediction module configured to perform molecular property prediction on the to-be-predicted molecule vector according to the trained molecular property prediction model to obtain a molecular property prediction result of the to-be-predicted molecule.
[0039] According to still another aspect of the present application, a storage medium having a computer program stored thereon is provided, the program being executed by a processor to implement the above-mentioned molecular property prediction model training method.
[0040] According to another aspect of the present application, a computer device is provided, comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein the processor implements the above-mentioned molecular property prediction model training method when executing the program.
[0041] By the above technical solution, the present application provides a molecular property prediction model training method, a storage medium, a computer device, and a molecular property prediction device. Through the above technical solution, the present application provides a molecular property prediction model training method, a storage medium, a computer device, and a molecular property prediction device. According to a preset candidate graph prompt extraction rule, a preset number of candidate graph prompts are extracted as target graph prompts. In the case where the preset number is given and unchanged, the extracted candidate graph prompts are as different as possible and representative. After determining the target graph prompts, the importance weights of each molecular graph are determined according to the target graph prompts. The multiple target graph prompts are fused and encoded into a global graph prompt vector. The global graph prompt vector is spliced with the feature vector of the sample molecule, respectively, to obtain the molecular vector of each sample molecule. Finally, according to the preset molecular property label corresponding to the sample molecule and the molecular vector, a molecular property prediction model is trained. The trained model can simultaneously predict multiple molecular properties, which can shorten the drug research and development cycle and improve the drug research and development efficiency.
[0042] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0043] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0044] Figure 1 A flowchart of a molecular property prediction model training method provided by an embodiment of the present application is shown;
[0045] Figure 2 A chemical molecular formula diagram provided by an embodiment of the present application is shown;
[0046] Figure 3 A molecular graph diagram of a chemical molecule provided by an embodiment of the present application is shown;
[0047] Figure 4 A flowchart of another molecular property prediction model training method provided by an embodiment of the present application is shown;
[0048] Figure 5 A flowchart of another molecular property prediction model training method provided by an embodiment of the present application is shown;
[0049] Figure 6 A flowchart of another molecular property prediction model training method provided by an embodiment of the application is shown.
[0050] Figure 7 A structural diagram of a molecular property prediction device provided by an embodiment of the application is shown. DETAILED DESCRIPTION
[0051] The application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the application and the features in the embodiments can be combined with each other without conflict.
[0052] A molecular property prediction model training method is provided in the embodiment, as shown in Figure 1 The method comprises the following steps.
[0053] In step 101, the molecular graphs of a plurality of sample molecules are obtained, and the molecular graph of each sample molecule is split to obtain a plurality of candidate graph prompts. In the candidate graph prompts, a preset number of candidate graph prompts are extracted as target graph prompts according to a preset candidate graph prompt extraction rule.
[0054] A graph neural network (GNN) can be used to model a molecule, that is, a molecular graph is constructed according to a chemical formula, and each atom in the chemical molecule is represented as a node in the molecular graph, and the force between the atoms is represented by the edge between the nodes. The node can carry different information to express different atomic symbols, and the edge can also carry different information to express different force modes, so a graph theory (molecular graph) in the computer is constructed. For example, the molecular formula of a chemical molecule is as shown in Figure 2 The molecular graph converted by the graph neural network is as shown in Figure 3 .
[0055] In the embodiment of the application, a plurality of chemical molecules are collected as sample molecules, and then each sample molecule is converted into a molecular graph according to the graph neural network to obtain the molecular graphs of the plurality of sample molecules. The molecular graph of each sample molecule is split to obtain a plurality of candidate graph prompts.
[0056] Figure hints are a special subgraph structure that can capture representative or indicative features in molecules. The BRICS algorithm (Breaking of Retrosynthetically Interesting Chemical Substructure) is used to split each molecule into several fragments of atoms or functional groups, and these fragments are used as candidate figure hints. The splitting process is as follows, assuming sample molecules are as follows:
[0057] Sample molecule 1: ethyl acetate, molecular formula C4H8O2, structure of a carbon atom connected to a carboxyl group (-COOH) and an ethyl group (-CH2CH3);
[0058] Sample molecule 2: benzoic acid, molecular formula C7H6O2, structure of a benzene ring (a cyclic structure of six carbon atoms and six hydrogen atoms) connected to a carboxyl group (-COOH);
[0059] Sample molecule 3: phenethyl alcohol, molecular formula C8H 10 O, structure of a benzene ring (a cyclic structure of six carbon atoms and six hydrogen atoms) connected to an ethanol group (-CH2CH2OH);
[0060] The three sample molecules are split into six fragments by the BRICS algorithm, and the six fragments are used as candidate figure hints, as follows:
[0061] Candidate figure hint 1: carboxyl group (-COOH), structure of a carbon atom connected to two oxygen atoms and one hydrogen atom;
[0062] Candidate figure hint 2: ethyl group (-CH2CH3), structure of a chain of two carbon atoms and five hydrogen atoms;
[0063] Candidate figure hint 3: benzene ring (a cyclic structure of six carbon atoms and six hydrogen atoms);
[0064] Candidate figure hint 4: ethanol group (-CH2CH2OH), structure of a chain of two carbon atoms and five hydrogen atoms, with one carbon atom also connected to an oxygen atom and a hydrogen atom;
[0065] Candidate figure hint 5: benzyl group (-C6H5CH2-), structure of a benzene ring connected to a methyl group (-CH2-);
[0066] Candidate figure hint 6: benzoic acid (C7H6O2), structure of a benzene ring connected to a carboxyl group (-COOH).
[0067] The candidate graph hint 5 and the candidate graph hint 6 are obtained by decomposing the sample molecule 2 and the sample molecule 3. The BRICS algorithm cuts a molecule according to some predefined chemical bonds, thereby obtaining some smaller fragments, which can be used as candidate graph hints or can be cut again until a certain size or complexity is reached.
[0068] The specific process of decomposing the candidate graph hint 5 and the candidate graph hint 6 from the sample molecule 2 and the sample molecule 3 by the BRICS algorithm is as follows:
[0069] The candidate graph hint 5: benzyl (-C6H5CH2-), which is cut from the molecule 3, and the cutting position is the bond between the benzene ring and the ethanol group;
[0070] The candidate graph hint 6: benzoic acid (C7H6O2), which is cut from the molecule 2, and the cutting position is the bond between the benzene ring and the carboxyl group.
[0071] Next, in the candidate graph hints, a preset number of candidate graph hints are extracted as target graph hints according to a preset candidate graph hint extraction rule. For example, there are 100 sample molecules, and 500 candidate graph hints are obtained after the sample molecules are split. The preset number is specified as 128, that is, 128 target graph hints are extracted from the aforementioned 500 candidate graph hints according to the preset candidate graph hint extraction rule. In particular, when the number of sample molecules is large, the preset number can be set to 128, and when the number of sample molecules is small, the preset number can be set to 32.
[0072] In order to generate high-quality and diversified target graph hints, a high-quality and diversified target graph hint set can be automatically generated based on a dynamic programming algorithm (DP) and a heuristic search algorithm (HS), so as to be able to cover different types and scales of molecular structure and attribute information.
[0073] Step 102, encode the target graph hints into target graph hint vectors, calculate the importance weights of each target graph hint vector to the molecular graph of each sample molecule, and calculate a global graph hint vector according to the target graph hint vectors and the importance weights.
[0074] Step 103, splice the global graph hint vector with the feature vector of each sample molecule respectively to obtain a molecular vector of each sample molecule;
[0075] Then, each candidate graph hint is converted into a fixed-length vector (target graph hint vector) according to the first GNN encoder, and then the importance weight of each candidate graph hint vector for each molecular graph is calculated according to the attention mechanism, and then all candidate graph hint vectors are fused into a global graph hint vector according to the weighted average method. The global graph hint vector is spliced with the feature vector of the molecular graph according to the second GNN encoder to obtain a molecular vector as the input of the molecular property prediction. At the same time, the molecular vector can also be restored into the original molecular graph according to the GNN decoder as a self-supervised learning target. Since the molecular vector combines the information of the candidate graph hint and the molecular graph, the decoding process can enhance the generalization ability of the model, so that the model can learn more rich and generalized molecular representation. By calculating the molecular vector z and the molecular graph obtained by inversely decoding the molecular vector, z and \hat{d} are used as a dataset to train the GNN encoder-decoder, which can make the process of encoding the molecular vector more accurate and enable the model to learn more molecular structure and attribute information.
[0076] Specifically, the similarity score s(m, f) between each molecular graph and each candidate graph hint can also be calculated, which represents the similarity between the molecular graph m and the candidate graph hint f. The similarity score can be calculated according to the similarity measurement method, such as Tanimoto coefficient or Dice coefficient, etc. Then, the weight w(m, f) between each molecular graph and each candidate graph hint is calculated according to the similarity score s(m, f) between each molecular graph and each candidate graph hint. The weight w represents the importance of the molecular m to the candidate graph hint f, and the weight calculation method is, for example, softmax function or sigmoid function, etc. Finally, each molecular graph and each candidate graph hint are combined by weighting according to the weight w(m, f) between each molecular graph and each candidate graph hint, to obtain a new molecular representation v(m) (molecular vector) representing the correlation between the molecular m and all candidate graph hints. The weighting combination method is, for example, linear combination or nonlinear combination, etc. Therefore, the new molecular representation v(m) of each molecule is obtained as the input of the model, so as to improve the performance and effect of the molecular property prediction.
[0077] Step 104, training a molecular property prediction model according to the preset molecular property label corresponding to the sample molecule and the molecular vector.
[0078] Finally, a molecular property prediction model is trained according to the preset molecular property label corresponding to the sample molecule and the molecular vector.
[0079] When pre-training with no-label molecular data, although the "pre-training, fine-tuning" paradigm can improve the generalization ability of GNNs to some extent, pre-training does not always bring significant improvement, especially when using random structure mask as a self-supervised signal, since random mask may destroy important information in the molecular structure, resulting in that the pre-trained model cannot fully utilize the patterns and rules in the molecular graph. At the same time, there is a domain shift problem between pre-training and fine-tuning, that is, the pre-trained model may not be well adapted to different molecular property tasks, because there may be structural and semantic differences between different tasks. In the way of pre-training on the pre-training-downstream task data and then fine-tuning, the multi-tasks are isolated and independent. For example, in Fine-tuning, the pre-trained language model "accommodates" various downstream tasks, which is specifically embodied by introducing various auxiliary task losses, adding them to the pre-trained model, and then continuing pre-training, so as to make it more adaptive to downstream tasks. Therefore, by constructing a unified framework, incrementally integrating data from different properties for continual learning, and through the "touching class by-passing" learning way of multi-task, the generalization ability of the model can be improved.
[0080] By applying the technical solutions of the embodiment, the graph prompt learning is a self-supervised learning framework, which guides GNNs to learn more rich and general molecular representations according to pre-defined or automatically generated graph prompts. At the same time, as an effective regularization means, the graph prompt can prevent the model from overfitting to a small amount of data, and can be used as a knowledge transfer mechanism to fuse cross-domain or cross-task information into the model. By constructing a multi-task molecular property prediction framework based on graph prompt learning, the graph prompt learning is embedded as a plug-in module into any GNNs, realizing the flexibility and compatibility of the model structure and parameters. A high-quality and diversified set of graph prompts is automatically generated by dynamic programming algorithm and heuristic search algorithm, covering different types and scales of molecular structure and attribute information. According to the attention mechanism and the method of multi-head autoencoder, different types and scales of graph prompts are encoded and decoded, so as to realize the adaptive selection and generation of graph prompts. Finally, according to the method of multi-task learning and graph convolutional network, the molecular properties of different tasks are predicted, the graph prompts are used as auxiliary information to enhance the generalization ability and transfer ability of the model, and through the encoder-decoder structure of the graph neural network, the candidate graph prompts and the molecular graph are encoded and decoded respectively, and by combining the target graph prompt vector with the molecular graph encoding, the molecular property prediction performance can be improved, and by decoding the molecular vector to restore the molecular graph, the generalization ability of the model can be enhanced.
[0081] Further, as a refinement and extension of the above embodiment, in order to fully describe the specific implementation process of the embodiment, another molecular property prediction model training method is provided, as shown in the figure, the method comprises: Figure 4
[0082] Step 201, obtaining the molecular graph of a plurality of sample molecules, splitting the molecular graph of each sample molecule to obtain a plurality of candidate graph prompts, and determining the number of intra-group graph prompts in the candidate graph prompts based on a preset number.
[0083] In the above embodiment of the present application, the molecular graph of a plurality of sample molecules is obtained, the molecular graph of each sample molecule is split to obtain a plurality of candidate graph prompts, and the number of intra-group graph prompts is determined based on a preset number, for example, assuming that the preset number is k, then the number of intra-group graph prompts is 1, 2, 3,..k respectively.
[0084] Step 202, for each intra-group graph prompt number, constructing a plurality of candidate graph prompt combinations, for each candidate graph prompt combination, calculating the quality score and diversity score of each candidate graph prompt in the candidate graph prompt combination, wherein each candidate graph prompt combination includes the intra-group graph prompt number of candidate graph prompts, the quality score of any candidate graph prompt is determined based on at least one of the appearance frequency of the candidate graph prompt in all candidate graph prompts corresponding to the sample molecule, the rarity of the candidate graph prompt and the complexity of the candidate graph prompt, and the diversity score of any candidate graph prompt is determined based on the structural difference and / or information complementarity between the candidate graph prompt and other candidate graph prompts corresponding to the sample molecule.
[0085] Next, for each intra-group graph prompt number, a plurality of candidate graph prompt combinations are constructed. For example, candidate graph prompt combinations containing 1 candidate graph prompt, candidate graph prompt combinations containing 2 candidate graph prompts,..., and candidate graph prompt combinations containing k candidate graph prompts are constructed respectively. At the same time, candidate graph prompt combinations that meet various combination modes of intra-group prompt numbers need to be constructed. Then, the quality score and diversity score of each candidate graph prompt in the candidate graph prompt combination are calculated. Specifically, first, define the evaluation function of the candidate graph prompt, which consists of two parts: one is the quality score of the candidate graph prompt, which reflects the appearance frequency, rarity, complexity and other indicators of the candidate graph prompt in all candidate graph prompts corresponding to the sample molecule; the other is the diversity score of the candidate graph prompt, which reflects the structural difference and information complementarity between the candidate graph prompt and other candidate graph prompts. By giving a candidate graph prompt set of a preset size, the evaluation function can be maximized to obtain the target graph prompt.
[0086] In particular, placing the candidate graph hints corresponding to multiple molecules in a set to calculate the evaluation value of each candidate graph hint can increase the number and diversity of candidate graph hints, thereby improving the coverage and representativeness of candidate graph hints. Since there may be some common or similar structural or attribute features between different molecules, these features can be used as useful graph hints to guide the model to learn better molecular representations. Therefore, placing the candidate graph hints split from multiple molecules together to calculate and optimize the evaluation value can obtain a high-quality and diversified target graph hint set.
[0087] In step 203, the evaluation value of the candidate graph hint is calculated based on the quality score and the diversity score, and the combination evaluation value of the candidate graph hint combination is calculated based on the evaluation value of each candidate graph hint in the candidate graph hint combination.
[0088] In step 204, for each number of intra-group graph hints, the candidate graph hint combination corresponding to the maximum value of the multiple combination evaluation values corresponding to the number of intra-group graph hints is taken as the target graph hint combination of the number of intra-group graph hints.
[0089] In step 205, for each target graph hint combination, the candidate graph hint with the maximum evaluation value in the target graph hint combination is taken as the target graph hint, and the target graph hint is encoded into a target graph hint vector.
[0090] Then, the evaluation value of the candidate graph hint is calculated according to the quality score and the diversity score, and the combination evaluation value of the candidate graph hint combination is calculated using a weighted average algorithm according to the evaluation value of each candidate graph hint in the candidate graph hint combination. For each number of intra-group graph hints, the candidate graph hint combination corresponding to the maximum value of the multiple combination evaluation values corresponding to the number of intra-group graph hints is taken as the target graph hint combination of the number of intra-group graph hints. For each target graph hint combination, the candidate graph hint with the maximum evaluation value in the target graph hint combination is taken as the target graph hint. To this end, the representative and diverse candidate graph hints are extracted for model training, which can improve the prediction performance of the model.
[0091] In particular, in order to extract the target graph hint satisfying the demand, the optimal evaluation function value of each candidate graph hint in the candidate graph hint set of different sizes can also be calculated according to the dynamic programming algorithm, and the corresponding predecessor graph hint is recorded, and the heuristic search algorithm is used to start from the candidate graph hint corresponding to the maximum evaluation function value, and gradually backtrack the predecessor graph hint until the target graph hint set of the preset size is reached. Finally, the target graph hint set is output as the final result. The aforementioned "different sizes" refers to the number of candidate graph hints in the candidate graph hint set. For example, if the preset size k = 10, the optimal evaluation function value of each candidate graph hint in the candidate graph hint set of size 1, 2,..., 10 needs to be calculated. The aforementioned "maximum evaluation function value" refers to the maximum value among the evaluation function values of all candidate graph hints in the candidate graph hint set of size k. The optimal evaluation function value of each candidate graph hint in the candidate graph hint set of different sizes is calculated by the dynamic programming algorithm, that is, the optimal evaluation function value of each candidate graph hint in the candidate graph hint set of different sizes is updated from small to large, until the size k is reached, and the optimal evaluation function value of each candidate graph hint in the candidate graph hint set of size k is obtained. Then, the maximum value is selected from the optimal evaluation function value as the starting point of the heuristic search algorithm, and the predecessor graph hint is gradually backtracked until the preset size k is reached. From the overall point of view, a total of (1 + 2 + 3 +... + k) optimal function values are calculated, which are composed of the optimal evaluation function values of all candidate graph hints in the graph hint set of different sizes. From the individual point of view, k optimal values are calculated for each candidate graph hint, which are composed of the optimal evaluation function values of a candidate graph hint in the candidate graph hint set of different sizes. Specifically, the aforementioned algorithm is, for example:
[0092] Input a sample molecule set D, a preset size k, and output a candidate graph hint set S. Initialize S as an empty set, and Q as a candidate graph hint queue.
[0093] For each sample molecule d e D, use the BRICS algorithm to split it into several fragments f, and add f to Q.
[0094] For each fragment f e Q, calculate its mass score q(f) and diversity score d(f, S) (initially S is an empty set).
[0095] For each fragment f e Q, initialize a one-dimensional array V(f) of size k + 1, and record a one-dimensional array P(f) of size k + 1, where,
[0096] V(f)[0] = 0, V(f)[i] = -∞ (i = 1,..., k), P(f)[0] = NULL.
[0097] For i = 1,..., k: take one fragment f from Q;
[0098] For j = i,..., 1: if V(f)[j-1] + q(f) + d(f, S) > V(f)[j], update V(f)[j] = V(f)[j-1] + q(f) + d(f, S), update P(f)[j] = f, and re-add f to Q;
[0099] Take one fragment f from Q such that V(f)[k] is maximum, and add it to S;
[0100] For i = k-1,..., 1: take one fragment f from P(f)[i+1] and add it to S;
[0101] Return S.
[0102] According to the above algorithm, firstly, a one-dimensional array V(f) of size k+1 is initialized, and a one-dimensional array P(f) of size k+1 is recorded, the two arrays are used to store the optimal evaluation function value of each candidate graph hint in different sizes of graph hint set and the corresponding predecessor graph hint, then a candidate graph hint f is taken out from the candidate graph hint queue Q, and the quality score q(f) and the diversity score d(f, S) are calculated, then starting from the size k and gradually decreasing to the size 1, for each size j, it is judged whether V(f)[j-1] + q(f) + d(f, S) > V(f)[j], if yes, it means that f is added to the graph hint set of size j-1, a larger evaluation function value can be obtained, then V(f)[j] = V(f)[j-1] + q(f) + d(f, S) and P(f)[j] = f are updated, and finally f is re-added to Q, until Q is empty or the preset number of times is reached, thus the optimal evaluation function value of each candidate graph hint in different sizes of graph hint set and the corresponding predecessor graph hint are obtained.
[0103] In particular, the candidate graph hint f is taken out from the candidate graph hint queue Q according to a preset order and rule, the preset order can be any order, as long as each candidate graph hint is taken out once, in the foregoing example, the candidate graph hints can be taken out in the order of candidate graph hint 1, candidate graph hint 2, candidate graph hint 3, and so on to candidate graph hint 6, or in other orders, for example, candidate graph hint 6, candidate graph hint 5, candidate graph hint 4, and so on to candidate graph hint 1, or candidate graph hint 3, candidate graph hint 1, candidate graph hint 5, candidate graph hint 4, and candidate graph hint 2, and so on. Different orders may affect the optimal evaluation function value of each candidate graph hint in different sizes of candidate graph hint set and the corresponding predecessor graph hint, but will not affect the final obtained target graph hint set of size k.
[0104] In a specific embodiment, assuming the preset size k = 2, i.e. a set containing two graph hints is needed, first, a one-dimensional array V(f) of size 3 is initialized, where V(f)[0] = 0, V(f)[1] = V(f)[2] = -∞, and a one-dimensional array P(f) of size 3 is recorded, where P(f)[0] = NULL. These two arrays are used to store the optimal evaluation function value of each candidate graph hint in the graph hint set of different sizes and the corresponding predecessor graph hint. Then the first candidate graph hint f1 is taken out from the candidate graph hint queue Q, and the quality score q(f1) and the diversity score d(f1, S) are calculated (S is empty at the beginning), assuming that q(f1) = 0.8 and d(f1, S) = 0.2. Then, starting from size 2 and decreasing to size 1, for each size j, it is determined whether V(f1)[j-1] + q(f1) + d(f1, S) > V(f1)[j]. If so, it means that f1 is added to the candidate graph hint set of size j-1, which can obtain a larger evaluation function value. Then V(f1)[j] = V(f1)[j-1] + q(f1) + d(f1, S) and P(f1)[j] = f1 are updated.
[0105] Specifically, when j = 2, V(f1)[1] + q(f1) + d(f1, S) = 0 + 0.8 + 0.2 = 1 > V(f1)[2] = -∞, so we update V(f1)[2] = 1 and P(f1)[2] = f1. When j = 1, V(f1)[0] + q(f1) + d(f1, S) = 0 + 0.8 + 0.2 = 1 > V(f1)[1] = -∞, then V(f1)[1] = 1 and P(f1)[1] = f1 are updated. Therefore, the calculation and update of the first candidate graph hint f1 are completed.
[0106] In step 205, the importance weight of each target graph hint vector to the molecular graph of each sample molecule is calculated, and the global graph hint vector is calculated according to the target graph hint vector and the importance weight.
[0107] Then, the importance weight of each target graph hint vector to the molecular graph of each sample molecule is calculated, and the target graph hint vector is fused to form a global graph hint vector. Combining the global graph hint vector with the sample molecule can enhance the generalization ability of the model.
[0108] Optionally, in step 205, the importance weight of each target graph hint vector to the molecular graph of each sample molecule is calculated, and the global graph hint vector is calculated according to the target graph hint vector and the importance weight, further comprising:
[0109] Step 205-1, according to the preset attention mechanism, the importance weight of each target graph prompt vector to the molecular graph of each sample molecule is calculated, and the average value of the importance weight of each target graph prompt vector to the molecular graph of different sample molecules is calculated.
[0110] Step 205-2, the average value of the importance weight of each target graph prompt vector is taken as the weighted weight value of each target graph prompt vector, and the global graph prompt vector is obtained by weighted fusion calculation of each target graph prompt based on the weighted weight value.
[0111] Specifically, the importance weight of each target graph prompt vector to the molecular graph of each sample molecule can also be calculated according to the preset attention mechanism, and the average value of the importance weight of each target graph prompt vector to the molecular graph of different sample molecules is calculated. The average value of the importance weight of each target graph prompt vector is taken as the weighted weight value of each target graph prompt vector, and the global graph prompt vector is obtained by weighted fusion calculation of each target graph prompt based on the weighted weight value.
[0112] Step 206, the global graph prompt vector is spliced with the feature vector of each sample molecule to obtain the molecular vector of each sample molecule, and the molecular property prediction model is trained according to the preset molecular property label corresponding to the sample molecule and the molecular vector.
[0113] Finally, the global graph prompt vector is spliced with the feature vector of each sample molecule to obtain the molecular vector of each sample molecule, and the molecular property prediction model is trained according to the preset molecular property label corresponding to the sample molecule and the molecular vector. The trained model can simultaneously predict multiple properties of a sample molecule, improve the efficiency of molecular property prediction, and shorten the drug development cycle.
[0114] By applying the technical solution of the embodiment, according to the graph prompt learning framework, flexibility and compatibility in model structure and parameters are realized, the model learning ability is increased by using the automatically generated high-quality and diversified graph prompt set, and the training method of the prompt mechanism is fused. Good precision results can be obtained with a small amount of labeled data. Compared with the manual parameter tuning method, the application embodiment can be applied to large-scale and replicable industrial expansion scenarios. By using machine learning method, the knowledge of multiple known property data is used to predict drug molecule properties, which promotes the drug development process, improves the accuracy of drug property prediction, shortens the drug development cycle, and improves the efficiency of drug development.
[0115] Further, as a refinement and expansion of the specific implementation of the above embodiment, in order to completely describe the specific implementation process of the embodiment, another molecular property prediction model training method is provided, such as Figure 5As shown, the method comprises:
[0116] In step 301, the molecular graphs of a plurality of sample molecules are obtained, the molecular graph of each sample molecule is split to obtain a plurality of candidate graph prompts, and a preset number of candidate graph prompts are extracted as target graph prompts according to a preset candidate graph prompt extraction rule in the candidate graph prompts.
[0117] In the above embodiments of the present application, the molecular graphs of a plurality of sample molecules are obtained, the molecular graph of each sample molecule is split to obtain a plurality of candidate graph prompts, and a preset number of candidate graph prompts are extracted as target graph prompts according to a preset candidate graph prompt extraction rule in the candidate graph prompts. By extracting high-quality and diversified target graph prompts, the generalization ability of the model can be improved.
[0118] In step 302, the target graph prompts are encoded into target graph prompt vectors, the importance weight of each target graph prompt vector to the molecular graph of each sample molecule is calculated, and a global graph prompt vector is calculated according to the target graph prompt vector and the importance weight.
[0119] Next, the target graph prompts are encoded into target graph prompt vectors, the importance weight of each target graph prompt vector to the molecular graph of each sample molecule is calculated, and a global graph prompt vector is calculated according to the target graph prompt vector and the importance weight. To this end, the target graph prompt vectors are fused into the global graph prompt vector according to the weight, which can cover rich information and improve the accuracy of model training
[0120] In step 303, for any sample molecule, a primary node feature matrix is constructed according to the atomic features of each atom in the molecular graph of the sample molecule, and an adjacency matrix is constructed according to the force features between adjacent atoms in the molecular graph.
[0121] Next, for any sample molecule, a primary node feature matrix is constructed according to the atomic features of each atom in the molecular graph of the sample molecule, and an adjacency matrix is constructed according to the force features between adjacent atoms in the molecular graph. The atomic features include the type of each atom (such as Oxygen, oxygen atom), the properties of the atom itself (Atom Properties), some features of the compound (Global Properties), etc. Each atom is regarded as a node, and the atomic bond is regarded as an edge. The edge also has corresponding features. Specifically, the primary node feature matrix is shown in Table 1, and the adjacency matrix is shown in Table 2:
[0122] Table 1
[0123]
[0124]
[0125] Table 2
[0126]
[0127] A graph neural network captures the dependencies of a graph through message passing between the nodes of a molecular graph. Unlike standard neural networks, graph neural networks preserve a state that can represent information from its neighborhood with arbitrary depth. Specifically, a graph neural network model updates the representation of a node by aggregating information from its neighboring nodes. Wherein the node label is repeatedly augmented by an ordered set of labels from neighboring nodes. The basic mechanism of this propagation is to first consider the neighborhood information as a graph substructure, and then model this substructure through a differentiable function by recursively projecting different substructures into different feature spaces. The information between neighbors and the center node is also called a message. The way the message is passed to the center node produces different propagation rules that characterize the network architecture.
[0128] The input of a graph neural network is the aforementioned molecular graph structure with node or edge attributes, that is, it includes the adjacency matrix A of the graph and the corresponding attribute information X. The graph neural network trains the implicit vector representation of each node in the graph according to the graph structure and the input node attribute, and the goal is to make the vector representation contain strong enough expression information, so that it can help each node to extract information, and finally obtain the information vector representation of the whole graph (such as extracting the molecular level information representation of the whole molecular compound through the features of the atomic nodes and the chemical bond information between atoms).
[0129] Wherein, if the learning of the graph network model is understood in the way of node message passing, it involves two processes, the message passing stage and the readout stage. The information passing stage is the forward propagation stage, which runs T steps in a loop and updates the node representation through the function M t Obtain information through the function U t Update the node, the equation of this stage is as follows,
[0130]
[0131]
[0132] Wherein, e vw represents the feature vector of the edge from node v to w.
[0133] The readout stage calculates a feature vector for the representation of the whole graph, which is realized by using the function R,
[0134]
[0135] Wherein, T represents the number of time steps, and the function M t , U tAnd R can use different model settings.
[0136] Step 304, the global graph prompt vector is spliced with the primary node feature matrix to obtain a secondary node feature matrix, and the secondary node feature matrix and the adjacency matrix are encoded into a molecular vector.
[0137] Next, the global graph prompt vector is spliced with the primary node feature matrix to obtain a secondary node feature matrix, and the secondary node feature matrix and the adjacency matrix are encoded into a molecular vector. To this end, the graph prompt is fused with the molecular graph.
[0138] Step 305, according to a plurality of predicted properties, a plurality of predicted property task nodes are constructed, a task relationship graph is constructed based on the correlation between the predicted property task nodes, and a task embedding vector of each predicted property task node is calculated.
[0139] Next, according to a plurality of predicted properties, a plurality of predicted property task nodes are constructed, a task relationship graph is constructed based on the correlation between the predicted property task nodes, and a task embedding vector of each predicted property task node is calculated.
[0140] After obtaining the molecular vector of each sample molecule, the corresponding molecular properties need to be predicted according to different tasks. Since different tasks may have different difficulties and correlations, a structured multi-task learning method is used to adjust the loss weight of each task using the relationship graph between tasks. First, a task relationship graph is constructed according to the Pearson correlation coefficient between tasks, wherein each node represents a task and each edge represents the correlation strength between two tasks. Then, the embedding vector of each task node is calculated based on the graph convolution network module, and the attention mechanism module is used to calculate the dependence of each task node on other task nodes. Finally, the adaptive threshold module is used to determine the loss weight of each task node, so that tasks with higher correlation or dependence have higher weights. The loss function of each task is weighted and averaged as the final optimization target of the model. For example:
[0141] Input molecular vector set Z, molecular property label set Y, output molecular property prediction set Initialization Empty set, T is a task set, and G is a task relationship graph.
[0142] For each molecular vector z∈Z, initialize y as the molecular property label vector corresponding to z, is the molecular property prediction vector corresponding to z.
[0143] For each task t∈T: use a fully connected layer F t Convert z into a scalar adding adding adding adding
[0144] For each task t e T: convert G into a node embedding matrix H = G1(G) using GCN module G1;
[0145] Compute the dependency weight w t for each node in H on t node using attention mechanism A2
[0146] Determine the loss weight a t for t node using adaptive threshold module A3 t ;
[0147] Compute the total loss function
[0148] Return
[0149] Step 306, for each predicted property task node, according to the task embedding vector of the predicted property task node and the task embedding vector of the related node of the predicted property task node, calculate the dependency degree between the predicted property task node and the related node.
[0150] Step 307, according to the dependency degree between different predicted property task nodes, determine the loss weight of each predicted property task node, and according to the loss weight of each predicted property task node, weight the loss function of each predicted property task node to obtain the loss function of the molecular property prediction model.
[0151] Specifically, input the candidate graph prompt set S, the sample molecule set D, and output the molecular vector set Z, initialize Z as an empty set;
[0152] For each sample molecule d e D: initialize the primary node feature matrix H(0) of d and the adjacency matrix A of d;
[0153] For each graph prompt s e S: convert s into a vector v s using GNN encoder E1 s ; Compute v s for d using attention mechanism A1 s ; s Fuse all v g into a global graph prompt vector v s∈S using weighted average method s ; s ;
[0154] v g is concatenated with H(0) to form a secondary node feature matrix H(1) = [H(0); v g ];
[0155] H(1) and A are taken as inputs using a GNN encoder E2 to obtain a target node feature matrix H(2) = E2(H(1), A);
[0156] H(2) is converted into a molecular vector z = P(H(2)) using a pooling operation P;
[0157] z is added to Z;
[0158] z is taken as input using a GNN decoder D to obtain a reconstructed molecular graph
[0159] The reconstruction loss between d and is calculated
[0160] Z is returned.
[0161] Step 308, according to the loss function, the preset molecular property label corresponding to the sample molecule and the molecular vector, training the molecular property prediction model.
[0162] Finally, according to the loss function, the preset molecular property label corresponding to the sample molecule and the molecular vector, training the molecular property prediction model.
[0163] In particular, the properties of small molecule drugs include the following: pharmacodynamics: pharmacodynamics refers to the biological activity of a compound on a specific target or disease, usually measured by half maximal inhibitory concentration (IC50), half maximal lethal dose (LD50) or other indicators. Pharmacodynamic prediction is a key task in drug discovery, which can help screen compounds with potential therapeutic value; pharmacokinetics: pharmacokinetics refers to the absorption, distribution, metabolism and excretion (ADME) process of a compound in the body, which affects the bioavailability, half-life and toxicity of the compound. Pharmacokinetic prediction can help optimize the physicochemical properties of the compound, improve its safety and effectiveness; toxicity: toxicity refers to the adverse effects of a compound on living organisms, including acute toxicity, chronic toxicity, carcinogenicity, mutagenicity, etc. Toxicity prediction can help assess the safety risk of the compound, reduce the failure rate and cost of drug development.
[0164] In the above embodiments of the present application, the disclosed dataset containing multiple drug properties is used for training: Therapeutics Data Commons (TDC), which collects datasets covering multiple drug discovery stages and tasks, with a total of 25 subcategories under the property prediction category, each subcategory corresponding to one or more specific datasets, and a total of 36 datasets. The datasets are all based on regression or classification tasks of small molecule structures, with the input being the SMILES string of the small molecule and the output being a continuous value or a discrete label. The datasets under the property prediction category are all collected or generated from public databases or literature, such as PubChem, ChEMBL, Tox21, etc. The following categories are included:
[0165] Molecular properties: This type of dataset contains the relationship between molecular structure and related properties, such as molecular weight, solubility, lipophilicity, etc.
[0166] Bioactivity: This type of dataset contains the relationship between molecular structure and specific target or cell line bioactivity, such as IC50, Ki, etc.
[0167] ADME: This type of dataset contains the relationship between molecular structure and ADME processes, such as plasma protein binding rate, liver microsomal stability, etc.
[0168] Toxicology: This type of dataset contains the relationship between molecular structure and different types of toxicity, such as acute oral toxicity, skin sensitization, etc.
[0169] Others: This type of dataset contains some other data related to drug discovery, such as antigen epitope prediction, antibody affinity prediction, etc.
[0170] By applying the technical solution of the present embodiment, each molecular vector v(m) is encoded by a shared encoder, and then each molecular property y(m) is decoded by multiple independent decoders, so that the encoder can learn a general molecular representation, and the decoder can learn a task-specific molecular property. To this end, the method of multi-task learning is used to improve the generalization ability and accuracy of the model.
[0171] Further, as a refinement and expansion of the specific implementation of the above embodiment, in order to fully describe the specific implementation process of the present embodiment, another molecular property prediction model training method is provided, as shown in Figure 6 The method comprises:
[0172] Step 401, obtaining a molecular graph of a plurality of sample molecules, splitting the molecular graph of each sample molecule to obtain a plurality of candidate graph prompts, and extracting a preset number of candidate graph prompts as target graph prompts according to a preset candidate graph prompt extraction rule in the candidate graph prompts.
[0173] Step 402, encode the target graph hints into target graph hint vectors, calculate the importance weight of each target graph hint vector to the molecular graph of each sample molecule, and calculate a global graph hint vector according to the target graph hint vectors and the importance weight.
[0174] Step 403, splice the global graph hint vector with the feature vector of each sample molecule respectively to obtain a molecular vector of each sample molecule.
[0175] Step 404, train a molecular property prediction model according to the preset molecular property label corresponding to the sample molecule and the molecular vector.
[0176] Step 405, obtain a preset global graph hint vector and a to-be-predicted molecule, splice the preset global graph hint vector with the feature vector of the to-be-predicted molecule to obtain a to-be-predicted molecular vector.
[0177] Step 406, perform molecular property prediction on the to-be-predicted molecular vector according to the trained molecular property prediction model to obtain a molecular property prediction result of the to-be-predicted molecule.
[0178] In the above embodiments of the present application, the molecular graph of a plurality of sample molecules is obtained, the molecular graph of each sample molecule is split to obtain a plurality of candidate graph hints, a preset number of candidate graph hints are extracted as target graph hints according to a preset candidate graph hint extraction rule in the candidate graph hints, the target graph hints are encoded into target graph hint vectors, the importance weight of each target graph hint vector to the molecular graph of each sample molecule is calculated, a global graph hint vector is calculated according to the target graph hint vectors and the importance weight, the global graph hint vector is spliced with the feature vector of each sample molecule respectively to obtain a molecular vector of each sample molecule, a molecular property prediction model is trained according to the preset molecular property label corresponding to the sample molecule and the molecular vector, a preset global graph hint vector and a to-be-predicted molecule are obtained, the preset global graph hint vector is spliced with the feature vector of the to-be-predicted molecule to obtain a to-be-predicted molecular vector, and molecular property prediction is performed on the to-be-predicted molecular vector according to the trained molecular property prediction model to obtain a molecular property prediction result of the to-be-predicted molecule. The trained model can simultaneously predict multiple properties of the to-be-predicted molecule, thereby improving the molecular property prediction efficiency.
[0179] Further, the present application provides a molecular property prediction device, as shown in Figure 7 The device comprises:
[0180] The data acquisition module 501 is configured to obtain a preset global graph hint vector and a to-be-predicted molecule.
[0181] The data processing module 502 is configured to splice the preset global graph hint vector with the feature vector of the to-be-predicted molecule to obtain a to-be-predicted molecular vector.
[0182] a property prediction module 503, configured to perform molecular property prediction on the to-be-predicted molecular vector according to the trained molecular property prediction model, to obtain a molecular property prediction result of the to-be-predicted molecule.
[0183] Based on the method as shown in Figure 1 , Figures 4 to 6 the embodiment of the present application also provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the molecular property prediction model training method as shown in Figure 1 , Figures 4 to 6 .
[0184] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each implementation scenario of the present application.
[0185] Based on the method as shown in Figure 1 , Figures 4 to 6 , and Figure 7 the virtual device embodiment, in order to achieve the above-mentioned purpose, the embodiment of the present application also provides a computer device, which can be a personal computer, a server, a network device, etc., and the computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the molecular property prediction model training method as shown in Figure 1 , Figures 4 to 6 .
[0186] Optionally, the computer device can also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen, an input unit such as a keyboard, etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.
[0187] Those skilled in the art can understand that the computer device structure provided by the embodiment of the present application does not constitute a limitation on the computer device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0188] The storage medium can further include an operating system and a network communication module. The operating system is a program for managing and saving hardware and software resources of a computer device and supports the running of an information processing program and other software and / or programs. The network communication module is used to realize communication between components in the storage medium and communication with other hardware and software in the entity device.
[0189] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and a necessary general hardware platform, or by hardware. The molecular graph of a plurality of sample molecules is obtained, each molecular graph is split to obtain a candidate graph hint, a preset number of candidate graph hints are extracted as target graph hints according to a preset candidate graph hint extraction rule, the target graph hints are encoded into target graph hint vectors, the importance weight of each target graph hint vector to each molecular graph is calculated, the global graph hint vector is calculated according to the target graph hint vector and the importance weight, the molecular vector of each sample molecule is obtained by splicing the global graph hint vector and the feature vector of each sample molecule, and the molecular property prediction model is trained according to the preset molecular property label and the molecular vector corresponding to the sample molecule. The trained model can simultaneously predict a plurality of molecular properties, thereby improving the drug research and development efficiency.
[0190] Those skilled in the art can understand that the drawings are only schematic diagrams of preferred implementation scenarios, and the modules or processes in the drawings are not necessarily required for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more devices different from the implementation scenario. The modules of the above implementation scenario can be combined into one module, or can be further split into a plurality of sub-modules.
[0191] The above application number is only for description, and does not represent the advantages and disadvantages of the implementation scenario. The above disclosure is only a few specific implementation scenarios of the present application, but the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A method for training a molecular property prediction model, characterized in that, The method includes: Molecular graphs of multiple sample molecules are obtained, and the molecular graphs of each sample molecule are split to obtain multiple candidate graph hints. Among the candidate graph hints, a preset number of candidate graph hints are extracted as target graph hints according to the preset candidate graph hint extraction rules. Each sample molecule corresponds to a molecular formula. The molecular graph is obtained by modeling the molecular formula through a graph neural network. The modeled molecular graph includes nodes and edges. Nodes represent atoms in the molecular formula, and edges between nodes represent the interaction forces between atoms. When modeling the molecular formula, the graph neural network uses the message passing mechanism between nodes in the molecular graph to capture graph dependencies and updates the node representation by aggregating the information of adjacent nodes. Node labels are recursively strengthened based on the ordered label set of adjacent nodes. The target map hints are encoded into target map hint vectors. The importance weight of each target map hint vector to the molecular map of each sample molecule is calculated. Based on the target map hint vectors and the importance weights, the global map hint vector is calculated. The global graph hint vector is concatenated with the feature vector of each sample molecule to obtain the molecular vector of each sample molecule. A molecular property prediction model is trained based on the preset molecular property labels and molecular vectors corresponding to the sample molecules.
2. The method according to claim 1, characterized in that, The step of extracting a preset number of candidate image hints as target image hints according to preset candidate image hint extraction rules includes: Based on the preset quantity, determine the number of prompts for multiple group images; For each group of image prompts, multiple candidate image prompt combinations are constructed, wherein each candidate image prompt combination includes the number of candidate image prompts for the group. For each candidate graph hint combination, calculate the evaluation value of each candidate graph hint in the candidate graph hint combination, and calculate the combined evaluation value of the candidate graph hint combination based on the evaluation values of each candidate graph hint; For each group of image prompts, the candidate image prompt combination corresponding to the maximum value among the multiple combination evaluation values corresponding to the group of image prompts is taken as the target image prompt combination for the group of image prompts. For each target graph hint combination, the candidate graph hint with the highest evaluation value in the target graph hint combination is selected as the target graph hint.
3. The method according to claim 2, characterized in that, The calculation of the evaluation value of each candidate map hint in the candidate map hint combination, and the calculation of the combined evaluation value of the candidate map hint combination based on the evaluation values of each candidate map hint, includes: Calculate the quality score and diversity score of each candidate graph hint in the candidate graph hint combination, wherein the quality score of any candidate graph hint is determined based on at least one of the occurrence frequency of the candidate graph hint in all candidate graph hints corresponding to the sample molecule, the rarity of the candidate graph hint, and the complexity of the candidate graph hint, and the diversity score of any candidate graph hint is determined based on the structural differences and / or information complementarity between the candidate graph hint and other candidate graph hints corresponding to the sample molecule; Based on the quality score and the diversity score, the evaluation value of the candidate map hints is calculated. Based on the evaluation values of each candidate map hint in the candidate map hint combination, the combined evaluation value of the candidate map hint combination is calculated.
4. The method according to claim 1, characterized in that, The calculation of the importance weight of each target map cue vector to the molecular map of each sample molecule, and the calculation of the global map cue vector based on the target map cue vector and the importance weight, includes: The importance weight of each target graph cue vector to the molecular graph of each sample molecule is calculated based on the preset attention mechanism, and the average importance weight of each target graph cue vector to the molecular graph of different sample molecules is calculated. The global graph cue vector is obtained by using the average importance weight of each target graph cue vector as the weighted weight value of each target graph cue vector, and then performing a weighted fusion calculation on each target graph cue based on the weighted weight value.
5. The method according to claim 1, characterized in that, The step of concatenating the global graph cue vector with the feature vectors of each sample molecule to obtain the molecular vector of each sample molecule includes: For any sample molecule, a primary node feature matrix is constructed based on the atomic features of each atom in the molecular graph of the sample molecule, and an adjacency matrix is constructed based on the interaction force features between adjacent atoms in the molecular graph. The global graph hint vector is concatenated with the primary node feature matrix to obtain the secondary node feature matrix. The secondary node feature matrix and the adjacency matrix are then encoded into a molecular vector.
6. The method according to claim 1, characterized in that, The step of training a molecular property prediction model based on the preset molecular property labels and molecular vectors corresponding to the sample molecules includes: Based on multiple predicted properties, multiple predicted property task nodes are constructed. Based on the correlation between the predicted property task nodes, a task relationship graph is constructed, and the task embedding vector of each predicted property task node is calculated. For each predicted task node, the degree of dependency between the predicted task node and the related nodes is calculated based on the task embedding vector of the predicted task node and the task embedding vector of the related nodes of the predicted task node. Based on the degree of dependence between task nodes with different predicted properties, the loss weight of each task node with predicted properties is determined, and the loss function of each task node with predicted properties is weighted according to the loss weight of each task node with predicted properties to obtain the loss function of the molecular property prediction model. A molecular property prediction model is trained based on the loss function, the preset molecular property labels and molecular vectors corresponding to the sample molecules.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the preset global graph hint vector and the molecule to be predicted; The preset global graph hint vector is concatenated with the feature vector of the molecule to be predicted to obtain the vector of the molecule to be predicted. Based on the trained molecular property prediction model, the molecular properties of the molecular vector to be predicted are predicted, and the molecular property prediction results of the molecular to be predicted are obtained.
8. A molecular property prediction device, characterized in that, The device includes: The data acquisition module is used to acquire the preset global graph hint vector and the molecule to be predicted; The data processing module is used to concatenate the preset global graph hint vector with the feature vector of the molecule to be predicted to obtain the vector of the molecule to be predicted. The property prediction module is used to predict the molecular properties of the molecular vector to be predicted based on the trained molecular property prediction model, and obtain the molecular property prediction results of the molecular vector to be predicted.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the molecular property prediction model training method according to any one of claims 1 to 7.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for training the molecular property prediction model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for training target image retrieval model
CN114565807A