A synergistic drug prediction method based on fusion drug molecular graph structure and cellular gene expression profile

By embedding the gene characteristics of cells into the molecular structure diagram of the drug and extracting mixed characteristics using deep learning networks, the problem of failure to fully consider the association relationship between drugs and cell characteristics in the prior art is solved, and a higher prediction accuracy of drug synergistic combination is achieved.

CN118711673BActive Publication Date: 2025-05-13NANHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410877769.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2025-05-13
Estimated Expiration
2044-07-02

AI Technical Summary

Technical Problem

When predicting drug synergistic combinations, the prior art fails to fully consider the association between drug molecular characteristics and cell gene expression profile characteristics, resulting in insufficient prediction accuracy.

Method used

The prediction of synergistic drugs is achieved by embedding the gene features of cells into the molecular structure diagram of drugs, using topological adaptive graph convolutional networks and multi-layer perception networks to extract mixed features of drugs and cells, and input these features into the bidirectional long and short-term memory network and recurrent neural network.

Benefits of technology

By learning the intrinsic relationship between drug molecular structure and cellular characteristics, the prediction accuracy of drug synergistic combination is improved, the overfitting problem is alleviated, and the accuracy, efficiency and applicability of prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118711673B_ABST
    Figure CN118711673B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting synergistic drugs by integrating drug molecular graph structure and cell gene expression spectrum, and belongs to the technical field of drug synergistic combination prediction. The method comprises the following steps: embedding the gene characteristics of cells as nodes into the molecular structure graph of drugs to obtain a fused structure graph; extracting the mixed characteristics of drugs and cells from the fused structure graph through a topological adaptive graph convolution network and a multi-layer perception network; inputting the mixed characteristics into a bidirectional long short-term memory network and a recurrent neural network to achieve the prediction of synergistic drugs. Compared with the prior art, the present invention has the beneficial effect that the method provided by the present invention integrates drug molecular structure information and cell feature extraction, captures the intrinsic relationship between drug molecular structure and cell features, and thus improves the prediction accuracy of drug synergistic combination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug synergistic combination prediction, and specifically relates to a synergistic drug prediction method that combines drug molecular graph structure and cell gene expression spectrum. Background Art

[0002] In the treatment of diseases, combination therapy is often more effective than single drug therapy. For example, in traditional Chinese medicine, combinations of different herbs are often used to treat specific diseases. In addition, combination therapy has the additional advantage of reducing the dose and adverse reactions of each drug. Especially in the treatment of cancers that are more resistant to drugs, combination therapy has been an effective strategy for decades. Therefore, obtaining accurate synergistic drug combinations is an important task with significant clinical and economic significance. However, screening and obtaining all synergistic drug combinations is far from enough by clinical experiments and high-throughput screening, and the required workload is huge and costly. Therefore, there is an urgent need to design an effective and accurate computational method to achieve the prediction of synergistic drug combinations.

[0003] So far, many computational methods have been proposed to replace experiments to predict potential synergistic drug combinations. In particular, deep learning methods have gained great attention in predicting synergistic drug combinations. Various deep learning methods have been applied to predict synergistic drug combinations using datasets from high-throughput screening. These methods use different models to calculate the corresponding synergy values ​​from the datasets, including zero interaction potency (ZIP), Bliss independence model, and Loewe additivity model.

[0004] These methods and models are all within the deep learning modeling framework, relying only on the chemical fingerprints of drugs and the gene expression profiles of target cell lines to establish methods based on deep neural networks. They have not fully considered the correlation between the molecular characteristics of drugs and the characteristics of cellular gene expression profiles. Summary of the invention

[0005] The purpose of the embodiments of the present invention is to provide a collaborative drug prediction method and system for fusion drug molecular graph structure and cell gene expression spectrum, which establishes a correlation between drug molecular characteristics and cell gene expression spectrum characteristics, obtains larger and more characteristic information, improves prediction accuracy, and thus can solve at least one technical problem involved in the background technology.

[0006] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0007] The embodiment of the present invention provides a synergistic drug prediction method based on the molecular graph structure of a fusion drug and the cell gene expression profile, comprising the following steps:

[0008] Step S1, embedding the gene features of the cell as nodes into the molecular structure diagram of the drug to obtain a fused structure diagram;

[0009] Step S2, extracting mixed features of drugs and cells from the fused structure graph through a topology adaptive graph convolutional network and a multi-layer perception network;

[0010] Step S3, inputting the mixed features into a bidirectional long short-term memory network and a recurrent neural network to achieve prediction of synergistic drugs.

[0011] Optionally, in step S1, the gene characteristics of the cell are obtained by the following method:

[0012] 978 genes were filtered through the Landmark gene database to identify cells;

[0013] 22 genes without gene expression profile data were deleted from the 978 genes;

[0014] For each cell C k , by V k ∈R 956 Indicates its genetic characteristics.

[0015] Optionally, in step S1, the molecular structure diagram of the drug is obtained by the following method:

[0016] The RDKit tool is used to generate a two-dimensional undirected graph G = (V, E) corresponding to the drug molecule, that is, the molecular structure diagram of the drug, where V is the node set, E is the edge set, and N = |V| is defined to represent the total number of nodes.

[0017] Optionally, step S1 specifically includes:

[0018] The cell is reduced to 74 dimensions through an autoencoder and inserted into the molecular structure diagram of the corresponding drug in a node manner, so that all drug nodes are unidirectionally connected to the cell structure diagram to obtain a fused structure diagram.

[0019] Optionally, in step S2, the topology adaptive graph convolutional network is defined as follows:

[0020]

[0021] In the formula, the matrix H (l+1)corresponds to the features of the (l+1)th layer, with each row corresponding to the feature vector of each node; the variable K represents the upper limit of the neighbor order, the degree to which the node features are updated to a specific order based on the neighbor information; the degree matrix is ​​denoted by D, and is defined so that each diagonal element corresponds to the degree of a given node, indicating the number of its neighboring nodes; the matrix A represents the adjacency in the graph, which is characterized by non-zero elements indicating the existence of edges between nodes; the variable X represents the input feature matrix, with each row representing the input feature vector of a node; the variable θ(l)k refers to the kth weight matrix of the lth layer, which is used for linear transformation.

[0022] Optionally, in step S2, the multi-layer perceptron network comprises three hidden layers for extracting features from the 956-dimensional cell input, and in addition, a dropout layer is added between each layer to reduce network complexity and prevent overfitting.

[0023] Optionally, in step S3, after the mixed features are input into the bidirectional long short-term memory network and the recurrent neural network, a multi-head attention mechanism is used to determine the attention coefficient in the embedding space, and then self-enhanced contrast learning of the incoming short-term memory and long short-term memory features is used to alleviate transition smoothness, and the synergistic score is predicted by a multi-layer perceptron to achieve the prediction of synergistic drugs.

[0024] Optionally, after determining the attention coefficient in the embedding space, it is normalized by the softmax function, and the final embedding vector is defined as follows:

[0025]

[0026] Where r(i) is the global representation of the i-th graph; N i is the node count in the i-th graph; X k (i) is the feature of the kth node in the i-th graph.

[0027] Optionally, use SIMCLR as a contrastive learning method.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The method provided by the present invention integrates drug molecular structure information and cell feature extraction, captures the intrinsic relationship between drug molecular structure and cell features, learns the interaction of initial features, and alleviates over-smoothing by learning the representation of two different outputs, and solves the over-fitting problem caused by supervised learning alone, thereby improving the prediction accuracy of drug synergistic combination;

[0030] 2. The method provided by the present invention improves the accuracy, efficiency and applicability of the prediction of drug synergistic effects, and helps to promote scientific research and clinical practice in the medical field. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work, among which:

[0032] Figure 1 A schematic diagram of the model structure of the collaborative drug prediction method of the fusion drug molecular graph structure and cellular gene expression profile provided by the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0035] The present invention provides a synergistic drug prediction method based on the molecular graph structure of a fusion drug and the cellular gene expression profile, comprising the following steps:

[0036] Step S1, embedding the gene features of the cell as nodes into the molecular structure diagram of the drug to obtain a fused structure diagram; specifically comprising:

[0037] The cell is reduced to 74 dimensions through an autoencoder and inserted into the molecular structure diagram of the corresponding drug in a node manner, so that all drug nodes are unidirectionally connected to the cell structure diagram to obtain a fused structure diagram.

[0038] Step S2, extracting the mixed features of drugs and cells from the fused structure graph through a topology adaptive graph convolutional network (TAGCN) and a multi-layer perception network (MLP);

[0039] Step S3, inputting the mixed features into a bidirectional long short-term memory network and a recurrent neural network to achieve prediction of synergistic drugs.

[0040] In step S1, the gene characteristics of the cell are obtained by the following method:

[0041] 978 genes were filtered through the Landmark gene database to identify cells;

[0042] 22 genes without gene expression profile data were deleted from the 978 genes;

[0043] For each cell C k , by V k ∈R 956 Indicates its genetic characteristics.

[0044] In step S1, the molecular structure diagram of the drug is obtained by the following method:

[0045] The RDKit tool is used to generate a two-dimensional undirected graph G = (V, E) corresponding to the drug molecule, that is, the molecular structure diagram of the drug, where V is the node set, E is the edge set, and N = |V| is defined to represent the total number of nodes.

[0046] In step S2, the topology adaptive graph convolutional network is defined as follows:

[0047]

[0048] In the formula, the matrix H (l+1) corresponds to the features of the (1+1)th layer, with each row corresponding to the feature vector of each node; the variable K represents the upper limit of the neighbor order, the degree to which the node features are updated to a specific order based on the neighbor information; the degree matrix is ​​denoted by D, and is defined so that each diagonal element corresponds to the degree of a given node, indicating the number of its neighboring nodes; the matrix A represents the adjacency in the graph, which is characterized by non-zero elements indicating the existence of edges between nodes; the variable X represents the input feature matrix, with each row representing the input feature vector of a node; the variable θ(l)k refers to the kth weight matrix of the lth layer, which is used for linear transformation.

[0049] In step S2, the multi-layer perceptron network contains three hidden layers to extract features from the 956-dimensional cell input. In addition, a dropout layer is added between each layer to reduce the network complexity and prevent overfitting.

[0050] In step S3, after the mixed features are input into the bidirectional long short-term memory network and the recurrent neural network, a multi-head attention mechanism is used to determine the attention coefficient in the embedding space, and then self-enhanced contrastive learning of the short-term memory and long short-term memory features is performed, specifically using SIMCLR as a contrastive learning method to ease transition smoothing, and a multi-layer perceptron is used to predict the synergistic score to achieve the prediction of synergistic drugs.

[0051] After determining the attention coefficient in the embedding space, it is normalized by the softmax function, and the final embedding vector is defined as follows:

[0052]

[0053] Where r(i) is the global representation of the i-th graph; N i is the node count in the i-th graph; X k (i) is the feature of the kth node in the i-th graph.

[0054] The model of the synergistic drug prediction method of the fusion drug molecular graph structure and cell gene expression profile provided by the present invention is as follows Figure 1 As shown, it can be seen that it mainly includes an embedding module (EmbeddingModule) and a prediction module (Predictionmodule), wherein the embedding module first uses an autoencoder to reduce the 956-dimensional cell to 74 dimensions, and then connects the cell as a node to the drug molecular structure diagram in a unidirectional manner; secondly, the molecular structure diagram of the cell and the drug is transferred to the TAGCN network to extract the mixed features of the drug and the cell. In addition, the features of the 956-dimensional cell are cloned into the MLP to obtain the features of the cell. The prediction module transfers the features obtained from the embedding module to the LSTM to RNN contrast learning module, so the prediction module can learn the features of short-term memory and long short-term memory. Finally, a multi-layer fully connected network is used to predict the synergy score. Note that the SIMCLR method is used to calculate the loss of contrast learning.

[0055] The model provided by the present invention is compared with the existing model, and the comparison results are shown in Table 1 and Table 2.

[0056] Table 1 Regression performance of the model on the test set

[0057] Model Mse Pearson R2 MAE Prior art 77.597 0.813 0.659 5.855 The present invention 49 0.87 0.767 4.7

[0058] Table 2 Classification performance of the model on the test set

[0059] Model ROC-AUC ACC Precision Kappa This model 0.982 0.979 0.943 0.950 Prior art 0.960 0.968 0.925 0.914

[0060] It should be noted that, in the above table, ACC is the accuracy (Accuracy), ROC-AUC is the area under the ROC curve (Receiver Operating Characteristic Area under Curve), Precision is the accuracy evaluation standard, PRAUC is the area under the PR curve (Precision Recall Area under Curve), and Kappa is the kappa coefficient (Cohen's Kappa).

[0061] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0062] In addition, it should be noted that the scope of the methods and systems in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.

[0063] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A synergistic drug prediction method based on fusion drug molecular graph structure and cell gene expression profile, characterized in that: The steps include: Step S1, embedding the gene features of the cell as nodes into the molecular structure diagram of the drug to obtain a fused structure diagram; Step S2, extracting mixed features of drugs and cells from the fused structure graph through a topology adaptive graph convolutional network and a multi-layer perception network; Step S3, input the mixed features into the bidirectional long short-term memory network and the recurrent neural network, use the multi-head attention mechanism to determine the attention coefficient in the embedding space, and then alleviate the transition smoothness through the self-enhanced contrast learning of the short-term memory and long short-term memory features, and predict the synergistic score through the multi-layer perceptron to achieve the prediction of synergistic drugs.

2. The method according to claim 1, characterized in that: In step S1, the gene characteristics of the cell are obtained by the following method: 978 genes were filtered through the Landmark gene database to identify cells; 22 genes without gene expression profile data were deleted from the 978 genes; For each cell C k , by V k ∈R 956 Indicates its genetic characteristics.

3. The method according to claim 1, characterized in that In step S1, the molecular structure diagram of the drug is obtained by the following method: The RDKit tool is used to generate a two-dimensional undirected graph G = (V, E) corresponding to the drug molecule, that is, the molecular structure diagram of the drug, where V is the node set, E is the edge set, and N = |V| is defined to represent the total number of nodes.

4. The method according to claim 3, characterized in that: Step S1 specifically includes: The cell is reduced to 74 dimensions through an autoencoder and inserted into the molecular structure diagram of the corresponding drug in a node manner, so that all drug nodes are unidirectionally connected to the cell structure diagram to obtain a fused structure diagram.

5. The method according to claim 1, characterized in that In step S2, the topology adaptive graph convolutional network is defined as follows: In the formula, the matrix H (l+1) corresponds to the features of the (l+1)th layer, with each row corresponding to the feature vector of each node; the variable K represents the upper limit of the neighbor order, the degree to which the node features are updated to a specific order based on the neighbor information; the degree matrix is ​​denoted by D, and is defined so that each diagonal element corresponds to the degree of a given node, indicating the number of its neighboring nodes; the matrix A represents the adjacency in the graph, characterized by non-zero elements indicating the presence of edges between nodes; the variable X represents the input feature matrix, with each row representing the input feature vector of a node; the variable Refers to the kth weight matrix of the lth layer, which is used for linear transformation.

6. The method according to claim 3, characterized in that In step S2, the multi-layer perceptron network contains three hidden layers to extract features from the 956-dimensional cell input. In addition, a dropout layer is added between each layer to reduce the network complexity and prevent overfitting.

7. The method according to claim 1, characterized in that After determining the attention coefficient in the embedding space, it is normalized by the softmax function, and the final embedding vector is defined as follows: Where r(i) is the global representation of the i-th graph; N i is the node count in the i-th graph; x k (i) is the feature of the kth node in the i-th graph.

8. The method according to claim 1, characterized in that Use SIMCLR as a contrastive learning method.

Citation Information

Patent Citations

  • Neural network model prediction method for gene expression profile under drug exposure

    CN116092626A

  • Pharmacodynamic synergistic effect prediction method based on multi-modal data

    CN117953962A