Eutectic Prediction Method Based on Graph Representation Learning

By designing a eutectic prediction method based on graph representation learning, and using the CC-MPNN model to divide and score the compounds in substructure, the problems of low accuracy of eutectic prediction and insufficient interpretability in the prior art are solved, and higher prediction accuracy and interpretability are achieved.

CN115424681BActive Publication Date: 2025-05-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211029734.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-05-27
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

The prior art has limited accuracy in eutectic prediction and is difficult to explain compound embedded representations, which makes it difficult to clarify the mechanism of eutectic formation.

Method used

A eutectic prediction method based on graph representation learning is designed. By constructing a eutectic prediction model CC-MPNN, the neural network module and the co-attention module are used to perform dot product fusion, and the gated message delivery neural network is combined to divide and score compound molecules in substructure to improve the interpretability of prediction.

Benefits of technology

Improves the accuracy and interpretability of eutectic prediction, providing a computational prediction tool that can be used for drug discovery and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424681B_ABST
    Figure CN115424681B_ABST
Patent Text Reader

Abstract

The present invention discloses a eutectic prediction method based on graph representation learning, and proposes an interpretable model based on gated message passing neural network, namely CC-MPNN. The gated message passing neural network is used to obtain the substructures of compounds, and the co-attention mechanism is used to calculate the interaction scores between substructures, so as to design a eutectic prediction method based on graph representation learning and explore the correlation law between compound substructures and eutectic reactions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer-aided drug development, and in particular relates to a cocrystal prediction method based on graph representation learning. Background Art

[0002] Pharmaceutical co-crystals refer to crystals formed by the combination of active pharmaceutical ingredients (API) and co-crystal coformers under the action of hydrogen bonds or other non-covalent bonds. By forming co-crystals, it is possible to change the solubility, dissolution rate, solid-state stability, and hygroscopicity of drugs, thereby improving the physicochemical properties and drug performance such as bioavailability of drugs. Therefore, co-crystal engineering has become an effective design strategy in the fields of pharmaceuticals, chemistry, and materials.

[0003] In the pharmaceutical industry, cocrystals are often used to improve the physicochemical and biopharmaceutical properties of active pharmaceutical ingredients (APIs) without changing the chemical structure of the drug. Novel drug cocrystals with excellent quality have the potential to be patented and marketed as new drugs. Experimental methods for cocrystal preparation include solution crystallization, steam or solvent-assisted grinding, and reactive crystallization. Although the prospects of cocrystals are fascinating, they are expensive in terms of time, energy, and laboratory resources. Therefore, computer prediction models to help chemists conduct rational cocrystal design or cocrystal virtual screening have gradually become the focus of attention.

[0004] Recently, data-driven machine learning methods have become increasingly popular in the fields of chemistry and materials because their optimization strategies can automatically improve through empirical data obtained from a statistical perspective, thereby providing intelligent navigation in a nearly infinite chemical space. However, the use of machine learning methods in combination with molecular descriptors or molecular fingerprints only shows moderate accuracy in co-crystal prediction, and there are still considerable challenges in practical work, mainly in terms of lack of interpretability. The compound embedding representation learned by deep learning or graph representation is difficult to interpret, and the formation of co-crystals cannot be explained by chemical knowledge.

[0005] Therefore, considering the structural relationship between the cocrystal components, it is possible to clarify the mechanism of cocrystal formation and establish a prediction model to help chemists discover some new possible cocrystals. In view of this, it is necessary to design a new prediction method to predict drug and compound cocrystals. Summary of the invention

[0006] The purpose of the present invention is to solve the deficiencies of the prior art and to provide a eutectic prediction method based on graph representation learning.

[0007] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention is as follows:

[0008] The eutectic prediction method based on graph representation learning is special in that it includes the following steps:

[0009] 1) Constructing the eutectic prediction model CC-MPNN:

[0010] The co-crystal prediction model CC-MPNN is composed of a neural network module and a co-attention module, and the two are dot-product fused to predict labels;

[0011] The neural network modules include graph attention network, graph convolutional neural network and gated message passing neural network in sequence;

[0012] The co-attention module includes a multi-layer perceptron;

[0013] 2) Obtain sample data, train the eutectic prediction model CC-MPNN constructed in step 1), and obtain a trained eutectic prediction model; the specific training process is as follows:

[0014] 2.1) Collect sample data and build training and test data sets

[0015] The sample data is a compound pair that can or cannot form a co-crystal, which includes the SMILES sequence information of the compound molecule and a label indicating whether there is a co-crystal reaction between the compound and the API or ligand, and the label can be obtained from the Cambridge Crystal Structure Database;

[0016] 2.2) Using RDKit tool and one-hot encoding, the SMILES sequence information of the compound molecules involved in the data obtained in step 2.1) is converted into graph data;

[0017] 2.3) Using graph attention network and graph convolutional neural network (i.e. GAT+GCN) to preprocess the compound molecular graph data obtained in step 2.2), and obtain the atomic feature vectors m1, m2, …, mn of all compound molecules;

[0018] 2.4) using a gated message passing neural network (GMPNN) method to train the atomic feature vector obtained in step 2.3) in combination with the characteristics of each compound molecular bond to obtain the substructure characteristics of the compound molecule;

[0019] 2.5) Use the constructed co-attention layer to calculate the co-attention score between each substructure feature, that is, the atomic interaction score;

[0020] 2.6) performing a linear transformation on the substructure feature matrix of each compound in a pair of compounds, then multiplying the corresponding atomic feature points, and multiplying and summing the results with the interatomic interaction scores obtained in step 2.5) to finally obtain the scores of the corresponding compound pairs;

[0021] 2.7) Using the scores of compound pairs obtained in step 2.6), the model loss is calculated using the binary cross entropy loss function, and the network is trained through negative feedback adjustment based on the loss residual; in layman's terms, this step is to obtain the deviation between the network prediction value and the actual result through the calculated loss value, which is used to feedback and adjust the prediction result of the network;

[0022] 2.8) After the training is completed, the eutectic prediction model is finally obtained;

[0023] 3) Using the co-crystal prediction model trained in step 2), predict the co-crystal reaction of the compound pair.

[0024] Furthermore, the step 2.2) is specifically as follows:

[0025] First, the SMILES sequence is converted into the atomic feature matrix and the interatomic bond feature matrix (both belong to graph data) using the open source chemical toolbox RDKit and one-hot encoding. Each atom is a multi-dimensional (e.g. 41-dimensional) feature vector, which contains information about the type of atom, atomic valence, and number of free radical electrons. The feature matrix of each bond is a multi-dimensional (e.g. 6-dimensional) feature vector, which contains information about the bond type, whether it is conjugated, and whether it is cyclic.

[0026] Secondly, it is encapsulated with the compound CID and key index into dictionary type data to store each image data.

[0027] Furthermore, the step 2.3) is specifically as follows:

[0028] Use torch_geometric toolkit to import GAT and GCN network layers, input the graph data obtained in step 2.2) into GAT and GCN network layers, and obtain the feature matrix of each molecule after preprocessing. The obtained matrix contains the feature information of atoms in the molecular graph;

[0029] A compound is represented as G = (V, E), where V is a set of N nodes, E is a set of edges, and A∈R N×N is the adjacency matrix representing E;

[0030] GAT aggregates neighbor nodes through the attention mechanism, realizes the adaptive allocation of weights of different neighbors, converts the input node features of the graph into higher-level features, and performs linear transformation on each node with a weight matrix. Then perform self-attention-shared attention mechanism on the nodes

[0031]

[0032] Indicates the importance of the feature of node j to node i; then the softmax function is used to normalize the attention coefficient, and the output feature of the node is calculated as:

[0033]

[0034] Among them, σ(·) is a nonlinear activation function, α ij is the normalized attention coefficient;

[0035] The propagation of GCN is as follows:

[0036]

[0037] in, To add the adjacency matrix of the self-connected undirected graph, I N is the identity matrix, σ(·) is the activation function, and W (l) is a layer-specific trainable weight matrix.

[0038] Furthermore, MPNN is a framework for multi-layer spatial convolutional GNNs, each layer of which includes three main components, namely message passing, aggregation, and update:

[0039]

[0040]

[0041]

[0042] Since MPNN does not take into account the feature information of the edge, the step 2.4) is as follows: adopt the method of the MPNN variant gated message passing neural network GMPNN, and its propagation mode is:

[0043]

[0044]

[0045]

[0046] The gate is controlled by w j→i =δ(o j→i ) to achieve, h j→i is the feature of edge j→i, h j and h iis the node feature of j, i, and σ(·) is the sigma function.

[0047] Different from traditional MPNN, GMPNN does not perform any transformation on the feature vector during message passing, but only updates the key information to ensure that all nodes constituting the substructure remain in the same feature space.

[0048] Further, the step 2.5) is specifically as follows:

[0049]

[0050] in, is a characteristic of substructure i in compound x, It is the feature of substructure j in compound y. MLP is a nonlinear function. After normalization using the softmax function, the output is the common attention score of each pair of atoms between a pair of compounds.

[0051] Furthermore, in step 2.6), the prediction score of each compound pair is obtained in the following manner:

[0052]

[0053] in, is a characteristic of substructure i in compound x, is the characteristic of substructure j in compound y, T is the transpose, and δ is the sigmoid function.

[0054] Furthermore, the step 2.7) uses the binary cross entropy loss function to calculate the model loss, as follows:

[0055]

[0056] Where M is the number of compound pairs, p i is the probability that the sample is a positive sample, p i ′ is the probability that the sample is a negative sample; due to the imbalance of positive and negative samples in the data set, log(p i ) and log(1-p i ′) When summing, weighted average is used, and the weight is the ratio of positive and negative samples.

[0057] At the same time, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the special feature of the computer program is that when the computer program is executed by a processor, the steps of the above method are implemented.

[0058] An electronic device is special in that it includes a processor and a computer-readable storage medium; the computer-readable storage medium stores a computer program, and the computer program executes the steps of the above method when executed by the processor.

[0059] Principle of the present invention:

[0060] The present invention uses graph representation learning to design a gated message passing neural network based on substructure division, studies the co-crystal prediction method based on graph representation learning, and explores the relationship between compound substructure and co-crystal. An interpretable model based on compound substructure interaction, namely CC-MPNN, is proposed. Starting from the division of substructures, the CC-MPNN model first uses a message passing neural network to divide the compound molecules in each pair of co-crystals according to the acquired data to obtain the substructure; then a scoring network is constructed for each pair of substructures in the two compounds; finally, the substructure feature vector and the interaction score between substructures are combined to predict the formation of co-crystals.

[0061] The advantages of the present invention are:

[0062] The present invention proposes a prediction model based on graph representation learning, namely CC-MPNN, which solves the problem through a gated message passing neural network. CC-MPNN consists of a neural network module and a common attention mechanism module, which are dot-product fused and the parameters of the neural network are adjusted through a loss function to obtain a predicted score. This model can increase the interpretability of co-crystal formation by dividing substructures and calculating the interaction scores between substructures. The evaluation of CC-MPNN on the data set shows that CC-MPNN has good co-crystal prediction performance. The present invention can provide a computational prediction tool to promote drug discovery and development. The present invention not only improves the accuracy, but also provides a certain degree of interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is the overall architecture of the CC-MPNN method proposed in the present invention. DETAILED DESCRIPTION

[0064] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0065] The specific embodiments of the eutectic prediction method based on graph representation learning proposed by the present invention are as follows:

[0066] This example uses a co-crystal data set: the data set has 7884 pairs of compounds, including 3847 compounds, of which there are 6831 positive samples and 1054 negative samples. The data of compound pairs are divided into training set and test set in a ratio of 9:1.

[0067] For the SMILES sequence information of compound molecules, the SMILES sequence information of each compound is converted into graph data with the help of RDKit tool and one-hot encoding.

[0068] Each acquired graph data is processed through the graph attention network layer and the graph convolutional neural network layer to obtain the feature vector m1, m2,…, mn of each atomic node.

[0069] The gated message passing neural network (GMPNN) method is used to divide the processed graph data into substructures combined with the characteristics of the edges in the graph data, and the substructure features t1, t2, …, tn of each compound molecule are obtained.

[0070] Using the constructed joint attention mechanism, the attention score γ between the substructures of each pair of compounds that can form co-crystals is obtained. ij .

[0071] The final score of a pair of compounds is obtained by multiplying and summing all substructure feature pairs after linear transformation and the common attention score.

[0072] Using the scores of the output compound pairs and their original labels (i.e., actual results), the model loss is calculated using the binary cross entropy loss function, and the network is trained through negative feedback regulation based on the loss residual.

[0073] After the training is completed, a classification model for compound co-crystal prediction, namely, a co-crystal prediction model, is obtained.

[0074] In order to evaluate the prediction performance, the present invention selects accuracy, precision, recall and F1_score as basic evaluation indicators. The higher the values ​​of these indicators, the better the performance. These indicators are calculated using the scikit-learn package in python.

[0075] The trained prediction model is tested using the test set data. At the same time, the present invention is compared with other baseline basic methods in a unified data set. The test results are shown in Table 1:

[0076] Table 1 Performance of CC-MPNN

[0077]

[0078] It can be seen from Table 1 that the accuracy (Accuracy), precision (Precision), recall (Recall) and F1_score of the prediction method of the present invention are significantly higher than those of other baseline basic methods, and has a significant effect.

[0079] In summary, the present invention can be used to predict the eutectic reaction of compounds, and the known implementation methods and common knowledge of characteristics in the above-mentioned scheme are not described in detail here. It should be pointed out that for those skilled in the art, several improvements can be made without departing from the present invention, and these should also be regarded as the protection scope of the present invention, which will not affect the implementation effect of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of the claims, and the specific implementation methods and other records in the specification are used to interpret the content of the claims.

Claims

1. A eutectic prediction method based on graph representation learning, characterized in that, it includes the following steps: 1) Construct a eutectic prediction model CC-MPNN: The eutectic prediction model CC-MPNN is composed of a neural network module and a co-attention module, and the two are dot-product fused to predict the label; The neural network module sequentially includes a graph attention network, a graph convolutional neural network, and a gated message passing neural network; The co-attention module includes a multi-layer perceptron; 2) Obtain sample data, train the eutectic prediction model CC-MPNN constructed in step 1) to obtain a trained eutectic prediction model; the specific training process is as follows: 2.1) Collect sample data and construct a training data set and a test data set The sample data is a pair of compounds that can or cannot form a eutectic, which includes the SMILES sequence information of the compound molecules and the label of whether there is a eutectic reaction between the compound and the active pharmaceutical ingredient or ligand; 2.2) Use the RDKit tool and one-hot encoding to convert the SMILES sequence information of the compound molecules involved in the data obtained in step 2.1) into graph data; 2.3) Use the graph attention network and the graph convolutional neural network to preprocess the compound molecule graph data obtained in step 2.2) to obtain the atomic feature vectors m1, m2,..., mn of all compound molecules; 2.4) Use the method of the gated message passing neural network to train the atomic feature vectors obtained in step 2.3) combined with the characteristics of each compound molecule bond to obtain the substructure characteristics of the compound molecules; 2.5) Use the constructed co-attention layer to calculate the co-attention scores between each substructure feature, that is, the atomic interaction scores; 2.6) Linearly transform the substructure feature matrices of each compound in a pair of compounds, then dot-multiply the corresponding atomic features, multiply and sum them with the atomic interaction scores obtained in step 2.5), and finally obtain the scores of the corresponding compound pairs; 2.7) Use the scores of the compound pairs obtained in step 2.6), calculate the model loss using the binary cross-entropy loss function, and obtain the scores of the compound pairs calculated in step 2.6) by negatively feedback regulating the set network according to the loss residual; 2.8) After training is completed, a eutectic prediction model is finally obtained; 3) Use the eutectic prediction model trained in step 2) to predict the eutectic reaction of compound pairs.

2. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that, the specific step 2.2) is: First, use the open-source chemical toolbox RDKit and one-hot encoding to convert the SMILES sequence into an atomic feature matrix and an atomic bond feature matrix; among them, each atom is a multi-dimensional feature vector, which contains information such as the type of atom, the valence of the atom, and the number of radical electrons; each bond feature matrix is a multi-dimensional feature vector, which contains information such as the type of bond, whether it is conjugated, and whether it is cyclic; Secondly, then encapsulate it with the compound CID and the index of the bond into a dictionary-type data to store each piece of graph data.

3. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that, the specific operation of step 2.3) is as follows: Import the GAT and GCN network layers using the torch_geometric toolkit, and input the graph data obtained in step 2.2) into the GAT and GCN network layers to obtain the feature matrix of each molecule after preprocessing. The obtained matrix contains the feature information of atoms in the molecular graph; The compound is represented as G=(V, E), where V is a set of N nodes, E is a set of edges, and A∈R N×N is the adjacency matrix representing E; GAT aggregates neighbor nodes through an attention mechanism, realizes the adaptive allocation of different neighbor weights, transforms the input node features of the graph into higher-level features, and performs a linear transformation on each node with a weight matrix. Then, self-attention--shared attention mechanism a is executed on the nodes: represents the importance of the feature of node j to node i; then use the softmax function to normalize the attention coefficient, and calculate the output feature of the node as: Among them, σ(·) is a non-linear activation function, and α ij is a normalized attention coefficient; The propagation method of GCN is as follows: Among them, is the adjacency matrix of an undirected graph with self-connections added, I N is the identity matrix, σ(·) is the activation function, and W (l) is a specific layer of trainable weight matrix.

4. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that, the specific operation of step 2.4) is as follows: Adopt the method of the variant gated message passing neural network of MPNN, and its propagation method is: Among them, the gating is implemented by w j→i = δ(o j→i ) and h j→i is the feature of edge j → i, h j and h i are the node features of j and i, and σ(·) is the sigma function.

5. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that, the specific operation of step 2.5) is as follows: Among them, is the feature of substructure i in compound x, is the feature of substructure j in compound y. The MLP is a non-linear function, and the common attention scores of each pair of atoms between a pair of compounds are output after normalization using the softmax function.

6. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that, in step 2.6), the prediction score of each compound pair is obtained in the following manner: Among them, is the feature of substructure i in compound x, is the feature of substructure j in compound y, T is the transpose, and δ is the sigmoid function.

7. The eutectic prediction method based on graph representation learning according to claim 1, characterized in that: in step 2.7), the binary cross-entropy loss function is used to calculate the loss, specifically as follows: Among them, M is the number of compound pairs, and p i is the probability that the sample is a positive sample, and p' i is the probability that the sample is a negative sample; due to the imbalance between positive and negative samples in the dataset, when summing log(p i ) and log(1 - p'), i a weighted average is used, and the weights are the proportion of the number of positive and negative samples.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

9. An electronic device, characterized in that: comprises a processor and a computer-readable storage medium; a computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, the steps of the method according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Eutectic prediction method based on graph neural network and deep learning framework

    CN111882044A

  • Method for predicting crystal structures formed by two medicines

    CN112712858A