Drug response prediction method based on gene relation network and drug substructure
By adopting deep learning methods such as enhanced graph attention networks and gated message delivery neural networks in cancer drug response prediction, multimodal features of cell lines and drugs are extracted, and the problem of processing high-dimensional data and 3D structural information in the prior art is solved, and more accurate drug response prediction and personalized treatment are achieved.
Patent Information
- Application Number
- CN202510230309.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively process high-dimensional data of cell lines and drug 3D structure information in cancer drug response prediction, resulting in large errors in prediction results and difficult to predict the response of unknown drugs.
The multimodal features of gene expression, copy number variation and mutations of cell lines were extracted using the enhanced graph attention network, and the substructure of the drug was obtained using the gabled message delivery neural network, and the drug 3D feature extraction was performed by constructing atom-bond map and bond-angle map, and drug response prediction was performed through the integration of features by multi-layer perceptrons.
It improves the accuracy of drug response prediction, can more accurately determine the IC50 value of the drug, reduce the risk of ineffective treatment, save medical costs, and promote the development of personalized medical care.
Smart Images

Figure BDA0005291284290000111 
Figure BDA0005291284290000131 
Figure BDA0005291284290000132
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drug prediction, and more particularly, to a method for predicting drug response based on a gene relationship network and a drug substructure. Background Art
[0002] Due to the heterogeneity of cancer, different patients often show different treatment outcomes even when suffering from the same cancer, which is also a major difficulty in cancer treatment. Cancer drug response prediction is to estimate the response of a patient's tumor cells to a specific anti-cancer drug, that is, the IC50 value of the drug, through a series of detection and analysis means before or at the beginning of cancer treatment. The IC50 value is an important indicator for measuring the drug-cancer cell line response. The size of the IC50 value directly reflects the strength of the effect of the drug on the cancer cell line. The smaller the IC50 value, the better the effect of the drug. By predicting the drug-cancer cell line response, it can help researchers treat different cancer patients more effectively, and also contribute to the discovery of potential candidate compounds for subsequent new drug research. Therefore, cancer drug response prediction is of great significance for the precise treatment of cancer.
[0003] Traditional biological experiments for predicting the IC50 value of drugs mainly include in vitro experiments and in vivo experiments. In vitro experiments are generally carried out in an artificially constructed simulation environment. A typical method is to use a cultured cell line, directly apply the drug to the cells, and then carefully observe the changes in various physiological indexes presented by the cells, and then calculate the IC50 value. The corresponding in vivo experiment requires injecting the drug into a living animal model, and then accurately determining the IC50 value of the drug by closely monitoring the physiological responses of the animal, including conventional indexes such as heart rate and blood pressure, biochemical indexes such as the content of specific substances in the blood, and observing the changes in the morphology and function of tissues and organs. Although the prediction of the IC50 value of drugs through biological experiments has high accuracy, in terms of time consumption, in vitro experiments require hundreds of experiments to determine the drugs that work on patients; in vivo experiments are even more time-consuming. From constructing an animal model to continuously monitoring the changes in the animal body, each stage consumes a large amount of time. At the same time, the consumable cost of in vitro experiments is quite high, and in vivo experiments not only have a high cost for purchasing animals, but also the costs of experimental equipment and professional personnel operation are not cheap. Therefore, it is very difficult for patients to receive personalized treatment through biological experiments. The development of high-throughput sequencing technology has enabled major research institutions to generate a large number of multi-omics databases of cancer cell lines and establish multiple data sets, which provide a solid foundation for cancer drug response models and make it possible to establish cancer drug response models for individualized treatment.
[0004] In recent years, scholars at home and abroad have proposed multiple cancer drug response prediction models. In the early stage, machine learning was mainly used to establish drug response prediction methods. Geeleher et al. proposed a method for predicting patients' chemotherapy responses, combining algorithms such as Ridge Regression with tools such as Principal Component Analysis (PCA) to form a method of training a model based on cell line data and applying it to clinical data prediction. This method captures the variability of patients' drug responses by constructing a statistical model and also uses whole-genome gene expression as a surrogate indicator for unmeasured phenotypes to predict drug responses. Ammad et al. proposed a new multi-task matrix factorization formula, introducing Component-wise Multiple Kernel Learning (MKL) into the Kernelized Bayesian Matrix Factorization (KBMF) method to form the cwKBMF method. This method solves the drug response prediction task by collecting evidence from multiple side data views and also uses known pathway information to learn drug response associations. However, the limitation of these machine learning-based drug prediction methods is that there is no way to handle the high-dimensional information in cell line omics data well, and it is prone to overfitting problems, resulting in large errors in the results. In addition, it is very difficult to predict unknown drugs, and only the IC50 values of known drugs can be predicted. Wang et al. proposed a new drug response prediction method, SRMF, introducing similarity regularization into the matrix factorization method to form the Similarity-Regularized Matrix Factorization (SRMF) method. This method solves the drug response prediction task by considering the similarity of drug chemical structures and the similarity of cell line gene expression profiles.
[0005] In contrast, researchers have gradually found that deep learning methods can effectively extract the characteristics of cell lines and drugs through neural networks, and cancer drug response prediction has gradually started to use deep learning methods. Et al. proposed a new drug response prediction method, Dr.VAE, which introduced Variational Autoencoders (VAE) into the drug response prediction model to form a deep learning architecture based on a probabilistic graphical model. This method solved the drug response prediction task by jointly modeling drug responses and transcriptomic perturbation effects. Liu et al. proposed a new cancer drug response prediction model, DeepCDR, which introduced a Graph Convolutional Network (GCN) into the cancer drug response prediction framework to form a hybrid graph convolutional network model. It solved the drug response prediction task by integrating multi-omics data of cancer cells and mining the intrinsic chemical structure information of drugs. Noghabi et al. proposed the MOLI model, which adopted a multi-omics late fusion strategy. It separately learned the features of somatic mutations, copy number variations, and gene expression data through specific encoding subnets, and then spliced and integrated them and input them into the classification subnet. During training, MOLI used a combined cost function that combined binary cross-entropy loss and triplet loss for end-to-end training, and used 5-fold cross-validation to tune hyperparameters. When predicting, the trained model was applied to the corresponding data to output the drug response prediction results.
[0006] However, many of the above methods only use single-omics data and do not consider multi-omics data. And among the methods that consider multi-omics data, the mutual relationships between genes in the omics are not considered, and the features of cell lines are not comprehensively extracted. For drug data, many do not consider the 3D structure information of drugs and the interaction information of drug substructures, and the feature extraction is not accurate.
[0007] In view of this, the present invention is specifically proposed. Summary of the Invention
[0008] In view of this, the present invention proposes a drug response prediction method based on a gene relationship network and drug substructures. It uses three different types of omics data with different focuses to extract cell line features, and uses multi-modal features such as drug molecular graphs, drug 3D structures, drug molecular fingerprints, and drug substructure relationships for drug feature extraction, improving the accuracy of feature extraction and being more capable of accurately determining the IC50 value of drugs.
[0009] Specifically, the present invention is implemented through the following technical solutions:
[0010] The present invention provides a drug response prediction method based on a gene relationship network and drug substructures, including the following steps:
[0011] Use an enhanced graph attention network to extract features from three multi-modal features of gene expression, copy number variation, and mutation of cell lines;
[0012] Use a gated message passing neural network to obtain the substructure of a drug, and extract features through an enhanced graph attention network;
[0013] Construct two complementary graph structures, an atom-bond graph and a bond-angle graph, and use a graph convolutional neural network to perform feature extraction and message passing on the two graphs, thereby extracting the 3D features of the drug;
[0014] Integrate the features of the drug and the cell line through a concat operation, and perform drug response prediction through an mlp.
[0015] In addition to providing a method for predicting drug response based on a gene relationship network and drug substructure, the present invention also provides a system for predicting drug response based on a gene relationship network and drug substructure, including:
[0016] Cell line feature extraction module: used to extract features of three multimodal features of gene expression, copy number variation, and mutation of the cell line by using an enhanced graph attention network;
[0017] Drug substructure extraction module: used to obtain the substructure of the drug by using a gated message passing neural network and extract features through an enhanced graph attention network;
[0018] Drug 3D feature extraction module: used to construct two complementary graph structures, an atom-bond graph and a bond-angle graph, and use a graph convolutional neural network to perform feature extraction and message passing on the two graphs, thereby extracting the 3D features of the drug;
[0019] Prediction module: used to integrate the features of the drug and the cell line through a concat operation for drug response prediction.
[0020] As is well known, cancer is a disease of abnormal cell proliferation that can occur in any tissue or organ of the human body. The occurrence of cancer usually involves abnormal cell proliferation, and these abnormal cells are usually not controlled by normal cell growth and death, resulting in the formation of tumors. When these malignant tumors grow large enough or spread to surrounding tissues, they will cause serious harm to the body. As one of the major diseases with a relatively high fatality rate, its diagnosis and treatment face huge challenges. Due to the heterogeneity of cancer, patients with the same cancer type may have different genomic characteristics. However, traditional cancer prediction is largely affected by tumor morphological evaluation, and for some tumors with similar histopathological manifestations, their clinical manifestations, courses of disease, and even treatment outcomes are very different. For cancer treatment, traditional methods are usually formulated based on the average characteristics and statistical data of the disease, ignoring the individual differences of each patient. This results in treatment plans that are sometimes not precise, and some patients may thus have adverse reactions or drug resistance to certain treatment methods.
[0021] Therefore, there is a need to design personalized pre - clinical prediction programs for cancer treatment to improve the effectiveness of disease treatment, which requires understanding the patient's response to various drugs. Although experimental measurement of drug response has high accuracy and reliability, it is time - consuming and costly. The development of high - throughput analysis techniques has made it realistic and feasible to study cancer diagnosis and treatment targeting individual differences based on data such as patients' genes and mutations. Predicting the sensitivity of patients to drugs based on the multi - omics characteristics of cancer cell lines measured from different patients plays an important role in achieving individualized treatment for patients.
[0022] In the past decade or so, high - throughput sequencing technology has led to the successive release of databases such as the Cancer Cell Line Encyclopedia (CCLE), the Genomics of Drug Sensitivity in Cancer (GDSC), and The Cancer Genome Atlas (TCGA). The emergence of such databases has provided a large amount of multi - omics data, drug data, clinical data, and drug response data of cancer cell lines. Among them, the GDSC database provides the drug - cell line half - maximal inhibitory concentration (IC50 value) of more than two hundred drugs and more than a thousand cell lines. In the field of biomedicine, multi - omics data refers to the integration of data at multiple different omics levels in cancer cell lines. For example, genomics, transcriptomics, copy number variation omics, and mutational omics are all one of the omics types of multi - omics data. These multi - omics substances jointly affect the phenotype of the life system. Specifically, cancer drug response prediction refers to predicting the patient's response to specific anti - cancer drugs by analyzing the patient's multi - omics data. This process can help doctors select the most effective treatment plan for patients, improve the treatment effect, and reduce the side effects brought by ineffective treatment. Therefore, developing cancer drug response prediction methods is of great significance. It not only helps improve the effectiveness and efficiency of cancer treatment but also promotes the development of personalized medicine, bringing more hope and possibilities to cancer patients.
[0023] Early cancer drug response prediction methods mainly rely on machine learning methods for drug prediction. Although these methods can predict cancer drug responses to a certain extent, they have many defects and limitations. For example, in the matrix factorization based method, it learns drug response associations by using known pathway information. Although it can effectively integrate multiple side data views to improve the accuracy of prediction, this method has overfitting problems for high-dimensional data such as omics data and needs to rely on more effective feature selection. The network-based method transforms the problem into a link prediction task by constructing a relationship network between cell lines and drugs, and uses the random walk method to predict cancer drug responses. However, this method is based on the assumption that similar cell lines and similar drugs have similar effects. In practical applications, this assumption has defects and ignores the small differences between similar cell lines and similar drugs, resulting in large errors in prediction results. In addition, early prediction methods require a large amount of known cell line data, drug data, and response data for support. However, in practical applications, the acquisition cost of these data is high and they are not necessarily completely accurate, which also limits the accuracy of prediction based on machine learning methods. At the same time, the drug response prediction method based on machine learning cannot predict unknown drugs, which also limits its application scenarios. Therefore, the early machine learning-based cancer drug response prediction methods cannot meet the actual needs in precision medicine, and new methods and technologies are urgently needed to solve these problems.
[0024] With the development of artificial intelligence technology, deep learning, a machine learning technology based on artificial neural networks, has achieved remarkable breakthroughs and applications in multiple fields and has also been widely applied in the field of precision medicine. In the formulation of personalized treatment plans, deep learning can predict drug responses based on multi-omics data such as the patient's genome and transcriptome, so as to formulate personalized treatment plans for patients. By analyzing the patient's gene mutations and drug molecular structures, deep learning models can predict the sensitivity of patients to specific drugs and help select the most effective drug combinations. In cancer prognosis analysis, deep learning can design cancer prognosis models, which can better predict the survival rate, recurrence rate, etc. of patients, which helps doctors adjust treatment strategies according to the patient's prognosis and avoid over-treatment or under-treatment. With the development of deep learning technology, deep learning-based cancer drug response prediction methods have gradually replaced traditional machine learning-based methods. Deep learning methods extract the biological data features of drugs and cancer cell lines, so as to effectively model large-scale drug data and cell line data, and then predict the responses of drugs in different cell lines. Compared with traditional machine learning methods, deep learning methods have better processing capabilities for high-dimensional data and can learn more complex multi-omics features of drugs and cell lines.
[0025] Among the above-mentioned many methods, the existing problems are as follows:
[0026] 1. Biological experimental methods: They are time-consuming and costly, and it is difficult for patients to receive personalized treatment through biological experiments.
[0027] 2. Machine learning-based methods: It is difficult to process high-dimensional cell line data and impossible to predict unknown drugs.
[0028] 3. Deep learning-based methods: Many of them only use single omics data and do not consider multi-omics data. Moreover, among the methods that consider multi-omics data, the mutual relationships between genes in omics are not considered, and the characteristics of cell lines are not comprehensively extracted; for drug data, many do not consider the 3D structure information of drugs and the interaction information of drug substructures, and the feature extraction is not accurate.
[0029] Especially in the deep learning methods, the main problems of the existing technologies are as follows:
[0030] Problem 1: Most of the existing methods for feature extraction of cancer cell line data only use one or two types of multi-omics data, and the mutual interactions between genes in multi-omics data are not considered.
[0031] Problem 2: For the current compound characterization, most of the existing methods do not consider the relationships between drug substructures in compounds.
[0032] Problem 3: Most of the existing CDR prediction methods only use drug molecular graphs or one-dimensional SMILES sequences of drugs to represent the molecular characteristics of drugs, without considering that drugs actually have 3D structures, and the roles played by different bond angles in the molecule may change.
[0033] After solving the above technologies by adopting the solution of the present invention, the solution of the present invention has the following beneficial effects:
[0034] (1) Currently, most methods use SMILES sequences, molecular graphs, and fingerprints as molecular characterizations, but do not consider the potential relationships between drug substructures, and drug molecules are actually 3D structures. In practical applications, different molecular bond angles have different therapeutic effects on cancer. Therefore, the present invention uses a Gated Message-Passing Neural Network to dynamically extract drug molecular substructures, and proposes a GALN module to integrate drug substructure features, uses GeoGNN to extract features of the 3D structure of drugs, and also uses the molecular fingerprints of drugs to obtain multi-level information of compounds in order to capture more comprehensive compound features, and finally uses them for the CDR prediction task.
[0035] (2) In the drug response prediction method, for the extraction of multi-omics features of cancer cell lines, most methods only consider one or two types of multi-omics features. At the same time, most of the methods for extracting multi-omics features do not consider the interaction information between genes in the omics information. Therefore, the present invention establishes a gene relationship network to extract the omics features of cell lines.
[0036] (3) The present invention adopts a 5-fold cross-validation method, which divides the entire data set into five subsets of equal size and conducts 5 iterations. In each iteration, one fold (subset) is selected as the test set, and the remaining 4 subsets are used as the training set. Finally, the average value of the performance metrics of the 5 iterations is used as the model evaluation result. Through 5-fold cross-validation, more stable and reliable performance metrics can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0038] Figure 1 is the overall architecture diagram of the drug response prediction method based on the gene relationship network and drug substructure of the present invention;
[0039] Figure 2 is the flowchart of the node aggregation method of the present invention;
[0040] Figure 3 is the flowchart of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0042] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0043] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information, depending on the context. As used herein, the word "if" may be interpreted as "when" or "while" or "in response to determining".
[0044] Embodiment
[0045] The present invention proposes a drug response prediction method based on a gene relationship network and a drug substructure, specifically including the following steps:
[0046] Use an enhanced graph attention network to extract features from three multi-modal features of gene expression, copy number variation, and mutation of cell lines;
[0047] Use a gated message passing neural network to obtain the substructure of the drug and extract features through an enhanced graph attention network;
[0048] Construct two complementary graph structures, an atom-bond graph and a bond-angle graph, and use a graph convolutional neural network to extract features and perform message passing on the two graphs, thereby extracting 3D features of the drug;
[0049] Integrate the features of the drug and the cell line through a concat operation, and perform drug response prediction through an mlp.
[0050] Specific implementation is carried out according to the following steps:
[0051] 1. Data source:
[0052] (1) GDSC database
[0053] In the cancer drug response prediction task, the GDSC (Genomics of Drug Sensitivity in Cancer) database plays a key role. This database focuses on the field of tumor drug sensitivity genomics, integrating data on the sensitivity of tumor cell lines to anticancer drugs with the genomic data of the cell lines, and is currently the largest public resource on cancer cell drug sensitivity and molecular markers of drug response. It contains 969 cell lines, 297 compounds, and 243,466 IC50 values between cell lines and drugs. Based on this, the data of the GDSC database is selected to construct a data set. For multi-omics data, this patent uses three omics information: gene expression genomics, copy number variation genomics, and mutation genomics.
[0054] A series of preprocessing operations were performed on the original data in GDSC. First, the IC50 value data was downloaded from the GDSC website. IC50 (half maximal inhibitory concentration) refers to the half inhibitory concentration of the measured antagonist. It can indicate the half amount of a certain drug or substance (inhibitor) in inhibiting certain biological programs (or certain substances included in this program, such as enzymes, cell receptors or microorganisms). The lower the IC50 value, the more sensitive the cells are to the drug, and the higher the possibility of the drug interacting with the cells. Secondly, those cell lines with missing part of the omics data were removed. At the same time, for the drug data, the parts with missing SMILES sequences and missing 80% of the IC50 values were removed. After this step, 734 cell lines and 177 drugs were finally retained.
[0055] (2) Pubchem database
[0056] The PubChem database is crucial in the field of drug research. Its data scale is extremely large, including the structural information of up to 118 million compounds, 318 million compound data, and 295 million compound bioactivity data. Moreover, it also covers 41 million literature abstract information, 5.1 million patent bibliographic information, as well as 113,242 target genes, 247,611 target proteins and 241,163 pathway information. The rich and comprehensive data resources make it an ideal source for obtaining drug data. Based on this, the SMILES sequences and 3D structure information of the 177 selected drugs were downloaded from the PubChem database.
[0057] (3) COSMIC database
[0058] For cell lines, considering the high-dimensional characteristics of omics data and to prevent overfitting problems, cancer-related genes were obtained from the COSMIC database to reduce the dimension of the cell line omics data.
[0059] 2 Data representation:
[0060] For drug data, SMILES sequences are used to characterize compounds. SMILES is a coding method that describes the chemical structure of compounds with short ASCII strings. With it, molecular graphs and 3D molecular graphs for drug characterization can be obtained. For cancer cell line data, to obtain features from multiple levels and perspectives, gene expression omics, copy number mutation omics, and mutation omics data are used as the input of the cell lines.
[0061] 1) Drug (compound) feature representation
[0062] 3D Molecular Graph Representation: Download the SDF format data containing the 3D structure of the compound from the Pubchem database and convert it to the PDB format data. For compounds lacking 3D structures, use the Merck Molecular Force Field (MMFF94) function in RDKit to calculate the simulated three-dimensional coordinates of the atoms in the molecule. Through the coordinate information, the positions of each atom in the molecule in space can be determined. Based on these coordinates, the geometric features of the molecule can be calculated, including bond lengths, bond angles, and atomic distance matrices, etc. In the present invention, the data is saved in the PDB format, and the composition information of the PDB file is shown in Table 1.
[0063] Table 1 Composition Information of the Drug PDB File
[0064]
[0065] Through the coordinate information, the positions of each atom in the molecule in space can be determined. Based on these coordinates, furthermore, the geometric features of the molecule can be calculated, including bond lengths, bond angles, and atomic distance matrices, etc.
[0066] Molecular Graph Representation: To further extract the substructure features of the drug, a molecular graph needs to be constructed. For the SMILES sequence of compound m, first parse the SMILES sequence into a molecular object n through MolFromSmiles in RDkit. The molecular object contains the basic information of atoms and chemical bonds. Then convert n into an undirected graph F, where each node represents an atom, and the edges between the nodes represent the chemical bonds between the atoms. According to the undirected graph F, an adjacency matrix D can be obtained, where when nodes i and j are connected, d ij = 1, otherwise it is 0.
[0067] Drug Molecular Fingerprints: GRSDRP uses three complementary molecular fingerprints to represent the physicochemical properties of compounds, including Pubchem molecular fingerprints, MACCS molecular fingerprints, and ECFP4 molecular fingerprints.
[0068] Pubchem Molecular Fingerprints: PubChem fingerprints are a widely adopted structure-critical molecular fingerprint, which consists of 881 binary digits. Each digit represents whether a specific substructure pattern exists in the molecule, and these substructure patterns are identified through SMARTS (Simple Molecular Input Line Entry System) queries.
[0069] MACCS Molecular Fingerprint: MACCS is a substructure-based fingerprint that uses an ordered list to record the presence of predefined functional groups, substructure motifs, and fragments, where 1 indicates the presence of a predefined chemical feature and 0 otherwise. The MACCS molecular fingerprint has versions of 166-bit binary strings and 960-dimensional strings. Since 166 dimensions already contain the vast majority of useful chemical fingerprints, 166-dimensional MACCS molecular fingerprints are used for representation in GRSPDR.
[0070] ECFP4 Molecular Fingerprint: ECFP4 is an extended connectivity fingerprint that captures circular substructures of different sizes around atoms through iterative hashing to accurately represent substructure features related to molecular properties. The generation process of the ECFP4 fingerprint involves hashing the neighborhood of each atom in the molecule to capture the local chemical environment and then performing global aggregation. It is not limited by activity cliffs. Compared with traditional SAR methods, the ECFP4 molecular fingerprint can better capture the biological activity of molecules.
[0071] The above molecular fingerprints are concatenated through the following formula to form FP (FingerPrint) to represent the physicochemical properties of compounds.
[0072] FP = (FP MACCS || FP ECFP4G || FP PubChem ) (1)
[0073] 2) Representation of Cell Line Characteristics
[0074] In the present invention, three types of multi-omics data, namely gene expression levels, copy number variations, and mutation data, are used as cell line representation data. Gene expression data is a one-dimensional sequence of decimals, where each value represents the gene expression level value of various genes in the cell, while copy number variations and mutation data are integer sequences of 0 and 1, with 1 indicating the occurrence of copy number variations or mutations and 0 indicating the non-occurrence of copy number variations or mutations. Due to the high-dimensionality of cell line omics data, relevant genes are downloaded from the COSMIC database, and non-cancer-related genes in the omics data are removed. Table 2 below shows some of the cell line genes used in the study.
[0075] Table 2 Some Cell Line Genes
[0076]
[0077] In addition, to establish a gene relationship graph, the cosine similarity is calculated to construct a gene relationship map, as shown in Formula 2:
[0078]
[0079] Among them, x i and x j represent gene i and gene j respectively, and S ij is the similarity between gene i and gene j. Since there are connections between each pair of genes, in order to establish a suitable graph to complete the feature extraction of cell line omics data, it is necessary to calculate the threshold:
[0080]
[0081] Among them, S = {S 12 , S 13 , S 14 ,..., S ij}, γ is a probability value representing the proportion of edges to be retained. By calculation, the threshold τ can be obtained. When S ij > τ, the value of the edge is set to 1, otherwise it is 0. After calculating the threshold of the relationship for each pair of genes, a gene relationship graph can be obtained, where each node represents a gene and the node feature is the gene feature.
[0082] 3. Introduction to the overall GRSDRP model framework of the present invention
[0083] Existing studies have shown that there are potential relationships between drug substructures, and at the cell line gene level, there are also associations between genes. However, existing methods rarely consider these two factors simultaneously. Therefore, the present invention proposes a drug response prediction method based on gene relationship network and substructure, named GRSDRP. This method realizes the accurate prediction of cancer drug response by integrating multi-omics features of cell lines and drug features. GRSDRP uses the SMILES sequence of drugs, 3D structure files, and gene expression omics, copy number variation omics, and methylation omics data of cell lines as input.
[0084] Specifically, first, the drug molecular graph and three types of molecular fingerprint features are extracted from the SMILES sequence of the drug, and at the same time, a gene relationship network is established from the multi-omics data of the cell line. GRSDRP consists of three parts: cell line feature extraction, drug feature extraction, and model prediction. Specifically, in the cell line feature extraction stage, GRSDRP uses the Enhanced Graph Attention Network v2 (EGATv2) to extract features from three types of omics data. In drug feature extraction, for the molecular graph, the Gated Message Passing Neural Network (GMPNN) is used to obtain the drug substructure, and then the Enhanced Graph Attention Network v2 (EGATv2) is used to extract the drug substructure features; for the molecular fingerprint features, a fully connected neural network (FCN) is used to obtain three different modal molecular fingerprint features; for the 3D structure of the drug, first, two complementary graph structures, the atom-bond graph and the bond-angle graph, are constructed according to the 3D structure file of the drug, and the Graph Convolutional Networks (GCN) is used to extract features and perform message passing on the two graphs, and then the 3D features of the drug are extracted. In model prediction, the three feature representations of the drug and the three feature representations of the cell line are integrated through the concat operation, and the Multilayer Perceptron (MLP) is used to complete the prediction of cancer drug response, that is, the prediction of the IC50 value. The overall architecture of the GRSDRP method is as Figure 1 shown.
[0085] 4. Enhanced Graph Attention Network v2 (EGATv2)
[0086] After establishing the gene relationship network and drug substructure extraction, in order to effectively extract the multi-omics features of the cell line and the drug substructure features, EGATv2 is designed, which can effectively learn feature representations from the gene relationship graph and the drug substructure. EGATv2 integrates GATv2 and Lipschitz Norm, extracts features through the graph attention network, and then optimizes the attention module through Lipschitz normalization, which can more effectively extract key features from the gene relationship graph and the drug substructure data.
[0087] Feature extraction: Given a graph \(G=(V, E)\), where \(V\) is the node attribute and \(E\) is the adjacency matrix of the nodes, the node embedding information of the \(i\)-th node at the \(k\)-th layer is updated from to by the following formula:
[0088]
[0089] where the attention coefficient \(\alpha\) ij reflects the importance of neighbor nodes. First, the feature information of nodes \(i\) and \(j\) is linearly transformed by the weight matrix, then processed by the LeakyReLU function to obtain the attention scores, and then normalized by the softmax function to obtain the attention coefficient \(\alpha\) ij . is a learnable weight matrix shared by all nodes. Through the aggregation of layers, the embedding representation of nodes can aggregate more global information. The node aggregation process is as shown in Figure 2 .
[0090] Feature normalization: Lipschitz Normalization (LN) is a normalization method for the self-attention mechanism. It enforces the Lipschitz continuity of the model by performing specific normalization operations on the attention scores. Since the graph attention network is prone to gradient explosion problems during training, affecting the training effect and model performance, LipschitzNorm is used to replace the original normalization layer to make the training process more stable and effectively improve the accuracy of feature extraction. The Lipschitz normalization formula is as follows:
[0091]
[0092] where \(\|W\|\) F is the Frobenius norm of the weight matrix \(W\), \(X = \{x 1 , x 2 , \ldots, x i \}\) is the attribute matrix of the nodes, and \(\|X T \|\) (∞,2) is the \((\infty, 2)\) norm of the matrix \(X T . Through Lipschitz normalization, the attention layer satisfies Lipschitz continuity and satisfies the following formula:
[0093]
[0094] where \(m\) and \(n\) are the numbers of input vectors and output vectors respectively.
[0095] 5 Gated Message Passing Neural Network (GMPNN)
[0096] Traditional drug representation methods are manually crafted and limited by existing human knowledge, making it difficult to discover new chemical substructures. Therefore, a gated message passing neural network is used to directly learn substructures of different sizes and shapes in molecules. Edges are regarded as gates with weights within [0,1] to control information flow, so as to generate substructures of different sizes and shapes. If the edge weight on a certain path is close to 0, the subsequent nodes on that path are cut off, thus defining substructures with irregular shapes.
[0097] Construct a drug molecule graph G=(V,E), where the nodes V represent atoms, the edges E are the bonds between atoms, and the feature vector x i of node v i ∈R d , and the feature vector of edge e ij is x ij ∈R d′ . Then convert the undirected graph into a directed graph, obtain e i→j and e j→i and assign them learnable weights in [0,1] respectively. Use the gated message passing neural network for message passing, aggregation, and update. Among them, the message passing from node v j to v i at the k-th layer is as follows:
[0098]
[0099] where represents the message passed from node v j to node v i at the k-th layer, and are the feature vectors of nodes v j and v i at the (k - 1)-th layer respectively.
[0100] The aggregation formula is as follows:
[0101]
[0102] where represents the aggregation of messages passed from all adjacent nodes of node v i at the k-th layer, represents the aggregated information of node v i at the k-th layer, and A (k) is the aggregation function at the k-th layer.
[0103] The message update formula is as follows:
[0104]
[0105] Among them, is the initial feature of the edge, is the aggregated information of the edge at the k-th layer.
[0106] 6 Geometry-based Graph Neural Network (GeoGNN)
[0107] The geometry-based graph neural network is a new type of graph neural network architecture, which is specifically used for molecular representation learning, especially excellent in the feature extraction of drug molecules. Compared with the traditional graph neural network (GNN), the geometry-based graph neural network not only considers the topological structure of the molecule (the relationship between atoms and chemical bonds), but also introduces the geometric information of the molecule (such as bond length, bond angle, etc.), so as to more comprehensively capture the physical and chemical properties of the molecule. The geometry-based graph neural network models the geometric information of the molecule through two graphs: the atom-bond graph (Atom-Bond Graph, G) and the bond-angle graph (Bond-Angle Graph, H). The atom-bond graph takes atoms as nodes and chemical bonds as edges; the bond-angle graph takes chemical bonds as nodes and bond angles as edges. This design enables it to learn the topological and geometric information of the molecule simultaneously.
[0108] Compared with the traditional method of drug feature extraction through molecular graphs, the present invention uses the 3D structure of drugs for feature extraction, which can more fully consider the geometric information of the molecule. In actual drug research and development, the differences in bonds and bond angles may change the physical and chemical properties of drugs. For example, lipoic acid is an antioxidant drug, and its dextrorotatory form has good antioxidant and anti-radiation activities, while the levorotatory form is basically inactive. Therefore, considering the actual 3D structure of drugs is crucial for extracting drug features.
[0109] The feature extraction process of this method includes three main steps: initialization, iterative update, and molecular representation generation. First, the features of atoms, chemical bonds, and bond angles are used as the initial input. These features include atom type, charge, bond type (single bond, double bond, etc.), and bond angle size, etc. The representation vectors of chemical bonds and atoms are updated through the iteration of the bond-angle graph H and the atom-bond graph G.
[0110] Update of bond representation: In the aggregation stage, GeoGNN will collect the information of the bonds adjacent to the current bond (i, j) and their bond angles. Specifically, for the bond (i, j), the model will consider all the bonds (i, m) connected to i and all the bonds (j, m) connected to j, and extract information from these bonds and their bond angles. These information are integrated through the aggregation function σ to form the aggregated information vector The specific formula is as follows:
[0111]
[0112] Among them, and respectively represent the sets of adjacent atoms of atoms i and j, and x mij and x ijm respectively represent the characteristics of the bond angles (m, i, j) and (i, j, m). The formula in the update stage is as follows:
[0113]
[0114] Among them, φ is the update function used to update the representation vector of the bond.
[0115] Update of atomic representation: In the atomic representation update stage, GeoGNN collects information on the atoms adjacent to the current atom u and their chemical bonds. Specifically, for atom i, the model considers all atoms j connected to i and the chemical bond (i, j) between them, and extracts information from these atoms and chemical bonds. This information is integrated through an aggregation function to form an aggregated information vector. The specific formula is as follows:
[0116]
[0117] Among them, represents the set of adjacent atoms of atom i, is the representation vector of atom j in the previous iteration, is the representation vector of the chemical bond (i, j) in the previous iteration. In the update stage, the model updates the representation vector of the current atom i according to the aggregated information vector . The formula in the update stage is as follows:
[0118]
[0119] Among them, θ is the update function used to update the representation vector of the atom.
[0120] Generation of molecular representation: After the iterative update is completed, GeoGNN integrates the representation vectors of all atoms through an aggregation function to generate the representation vector h G of the entire molecule. The specific formula is as follows:
[0121]
[0122] 7 Multilayer Perceptron (MLP)
[0123] Also known as the Fully Connected Network (FCN), it is a classic neural network structure where neurons in each layer are connected to all neurons in the next layer. An MLP is used to predict the drug response results, specifically using 3 hidden layers to predict the ln(IC50) value:
[0124]
[0125] where h i is the output of the hidden layer of the i-th layer, W i is the weight matrix, b i is the bias vector, is the predicted ln(IC50) value.
[0126] In addition, the present invention not only provides a method for predicting drug response, but also provides a drug response prediction system, specifically including:
[0127] Cell line feature extraction module: used to extract features of three multi-modal features of gene expression, copy number variation, and mutation of the cell line by using an enhanced graph attention network;
[0128] Drug substructure extraction module: used to obtain the substructure of the drug by using a gated message passing neural network and extract features through an enhanced graph attention network;
[0129] Drug 3D feature extraction module: used to construct two complementary graph structures of atom-bond graph and bond-angle graph, and use a graph convolutional neural network to extract features and message passing for the two graphs, and then extract drug 3D features;
[0130] Prediction module: used to integrate the features of the drug and the cell line through the concat operation and predict the drug response through an mlp.
[0131] In specific implementation, the above-mentioned modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above-mentioned units, reference can be made to the method embodiments described above, which will not be elaborated here.
[0132] In summary, the solution of the present invention has the following technical effects:
[0133] 1. Improve prediction accuracy: The technical solution considers the potential relationships between the substructures of the drug and between the genes in the cell line. Therefore, molecular substructures, molecular fingerprints, and molecular 3D results are used to extract compound features, a gene association network is established to extract multi-omics features of the cell line, and by fusing these features at different levels, the prediction model can more accurately simulate the interaction between the drug and the cell line, thereby improving the accuracy of drug response prediction.
[0134] 2. Reduce the risk of ineffective treatment: By means of the drug response prediction method in the research protocol, the actual application can reduce the patients' exposure to ineffective cancer drug treatments. This can not only reduce the physical burden and pain of patients caused by ineffective treatment, but also avoid the waste of medical resources, enabling patients to promptly switch to more effective treatment means.
[0135] 3. Save medical costs: Predicting cancer drug responses in the early stage of cancer treatment can avoid unnecessary drug trials. It reduces the usage of ineffective drugs and related testing costs, lowering the overall medical costs, and at the same time alleviating the economic pressure on patients' families.
[0136] 4. Facilitate new drug research and development: Provide key data support for the research and development of new cancer drugs. By analyzing the response data of a large number of patients to different drugs, researchers can deeply understand the drug action mechanism and targets, more specifically optimize the drug molecular structure, accelerate the new drug research and development process, and improve the research and development success rate.
[0137] 5. Promote the development of personalized medicine: The cancer drug response prediction method can fully consider the individual differences of each cancer patient, realizing personalized medicine in the true sense. Based on accurate drug response prediction, doctors can formulate personalized treatment plans according to the specific conditions of patients, improve the accuracy and effectiveness of treatment, and improve the prognosis of patients.
[0138] 6. Discover new drug uses: Through in-depth research on cancer drug responses, it is possible to discover new uses of existing drugs in treating other cancer types or specific cancer subtypes. This provides an opportunity for drug repositioning, enabling the full utilization of existing drug resources and developing new treatment strategies.
[0139] 7. Evaluate drug safety in advance: The prediction model can evaluate the possible adverse reactions of drugs in different patients, and identify potential safety risks in advance. This helps doctors take corresponding measures before medication, reduce the safety hazards during drug treatment, and improve the safety of patients' medication.
[0140] Figure 3 The structural schematic diagram of a computer device disclosed by the present invention. Refer to Figure 3 As shown, the computer device 400 includes at least a memory 402 and a processor 401; the memory 402 is connected to the processor through a communication bus 403, and is used to store computer instructions executable by the processor 401, and the processor 401 is used to read computer instructions from the memory 402 to implement the steps of the method described in any of the above embodiments.
[0141] For the above device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0142] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated into, dedicated logic circuitry.
[0143] Finally, it should be noted that: Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope claimed, but are mainly used to describe the features of specific embodiments of a particular invention. Certain features described in multiple embodiments in this specification can also be combined and implemented in a single embodiment. On the other hand, the various features described in a single embodiment can also be separately implemented in multiple embodiments or implemented in any suitable sub-combination. In addition, although the features may function in certain combinations as described above and are even initially claimed as such, one or more features from the claimed combination can in some cases be removed from the combination, and the claimed combination can be directed to a sub-combination or a variation of the sub-combination.
[0144] Similarly, although the operations are depicted in a specific order in the drawings, this should not be construed as requiring that the operations be performed in the specific order shown or sequentially, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0145] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0146] The foregoing is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the scope of protection of the present disclosure.
Claims
1. A drug response prediction method based on gene relationship network and drug substructure, characterized in that: The steps include: The enhanced graph attention network is used to extract three multimodal features of gene expression, copy number variation, and mutation of cell lines; The substructure of drugs is obtained using a gated message passing neural network and features are extracted using an enhanced graph attention network. Construct two complementary graph structures, the atom-bond graph and the bond-angle graph, and use graph convolutional neural networks to perform feature extraction and message passing on the two graphs to extract the 3D features of the drug; The characteristics of drugs and cell lines are integrated through the concat operation, and drug response prediction is performed through MLP.
2. The drug response prediction method according to claim 1, characterized in that: Before extracting the three multimodal features of gene expression, copy number variation, and mutation of the cell line, the representation methods of the cell line features include: After downloading the relevant genes, we construct a gene relationship map by calculating the cosine similarity: Among them, x i and x j represent gene i and gene j respectively, S ij For the similarity between gene i and gene j, the threshold is calculated: τ=P γ (S) Where S = {S 12 ,S 13 ,S 14 ,...,S ij }, γ is a probability value representing the proportion of retained edges. The threshold τ can be obtained by calculation. When S ij When τ > τ, the edge value is set to 1, otherwise it is set to 0. After calculating the threshold of the relationship for each pair of genes, the gene relationship graph is obtained.
3. The drug response prediction method according to claim 2, characterized in that: The method of obtaining the substructure of a drug using a gated message passing neural network includes the following steps: Establish a drug molecule graph G = (V, E), where node V represents an atom, edge E represents the bond between atoms, and node v i The eigenvector x i ∈R d , edge ij The eigenvector of ij ∈R d′ , and then convert the undirected graph into a directed graph to obtain e i→j and e j→i And assign a learnable weight of [0.1] respectively, and use a gated message passing neural network for message passing, aggregation, and updating; Among them, at the k-th layer node v j to v i The message passing formula is as follows: in, Indicates that at the kth level, from node v j Passed to node v i News, and They are node v j and v i At the k-1th layer, the aggregation formula is as follows: in, Represents node v i Aggregate the messages sent from all adjacent nodes at layer k. Represents node v i Aggregate information at the kth layer, A (k) is the aggregation function of the kth layer; The message update formula is as follows: in, is the initial feature of the edge, is the aggregate information of the edges at the kth layer.
4. The drug response prediction method according to claim 3, characterized in that: Methods for extracting 3D features of drugs include: The characteristics of atoms, chemical bonds, and bond angles are used as initial inputs, and the representation vectors of chemical bonds and atoms are updated through the iteration of the bond-bond angle graph H and the atom-bond graph G. The specific formula is as follows: in, and denotes the adjacent atomic sets of atom i and atom j respectively, and x mij and x ijm They represent the characteristics of the bond angles (m,i,j) and (i,j,m), respectively; The updated formula is as follows: Among them, φ is the update function, which is used to update the key representation vector; In the atomic representation update phase, this information is integrated through the aggregation function, the specific formula is as follows: in, represents the set of adjacent atoms of atom i, is the representation vector of atom j in the previous iteration, is the representation vector of the chemical bond (i, j) in the previous iteration; According to the aggregated information vector To update the representation vector of the current atom i, the update formula is as follows: Among them, θ is the update function, which is used to update the representation vector of the atom; After the iterative update is completed, the representation vectors of all atoms are integrated through the aggregation function to generate the representation vector h of the entire molecule G , the specific formula is as follows:
5. The drug response prediction method according to claim 4, characterized in that: The method for enhancing the feature extraction of the graph attention network includes: Given a graph G = (V, E), where V is the node attribute and E is the node adjacency matrix, the node embedding information of the i-th node in the k-th layer is transformed from Update to The formula is as follows: Among them, the attention coefficient α ij To reflect the importance of neighbor nodes, we first use the weight matrix to linearly transform the feature information of node i and node j, and then process it through the LeakeReLU function to get the attention score, and then normalize it through the softmax function to get the attention coefficient α ij , A learnable weight matrix shared by all nodes; The normalized formula is as follows: Among them, ‖W‖ F is the Frobenius norm of the weight matrix W, X = {x1, x2, ..., x i } is the attribute matrix of the node, ‖X T ‖ (∞,2) is the matrix X T The (∞,2) norm of is normalized by Lipschitz, so that the attention layer satisfies Lipschitz continuity and satisfies the following formula: Where m and n are the number of input vectors and output vectors respectively.
6. The drug response prediction method according to any one of claims 1 to 5, characterized in that: Use MLP to predict the ln(IC50) value of drug response results: X=[X cell ,X drug ] h1=ReLU(W1X+b1) h i =ReLU(W i h (i-1) +b i ) Among them, h i is the hidden layer output of the i-th layer, W i is the weight matrix, b i is the bias vector, is the predicted ln(IC50) value.
7. A drug reaction prediction system using the drug reaction prediction method according to any one of claims 1 to 6, characterized in that: include: Cell line feature extraction module: used to extract three multimodal features of gene expression, copy number variation, and mutation of cell lines using enhanced graph attention network; Drug substructure extraction module: used to obtain the substructure of drugs using a gated message passing neural network and extract features through an enhanced graph attention network; Drug 3D feature extraction module: used to construct two complementary graph structures, atom-bond graph and bond-angle graph, and use graph convolutional neural network to extract features and pass messages between the two graphs to extract drug 3D features; Prediction module: used to integrate the characteristics of drugs and cell lines through the concat operation and predict drug response through MLP.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the drug response prediction method according to any one of claims 1 to 6 are executed.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the drug response prediction method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Drug resistance prediction method and device, electronic equipment and storage medium
CN120432016A
Drug resistance prediction method, device, electronic device and storage medium
CN120432016B
Biomass pyrolytic reaction network prediction method based on graph theory and quantum chemistry
CN120613027A
A method for predicting biomass pyrolysis reaction networks based on graph theory and quantum chemistry
CN120613027B