A cancer driver gene identification method based on a graph neural network and negative sample generation
Patent Information
- Application Number
- CN202311715595.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-12-14
AI Technical Summary
尽管这些方法相较于传统方法有显著的进步,但癌症的发生往往涉及基因相互作用的不同方面,如代谢、激酶和调控等,而这些方法仅仅关注单个网络,这不可避免地忽视了相互作用中的不完整性和噪声
[0038] Considering the incompleteness of gene interactions in a single network, the method of this invention focuses on multi-layer biological networks. This method reduces the noise level of biological networks and represents nodes in weighted multi-layer networks as vectors, while effectively preserving structural information for identifying prognostic biomarkers.
Smart Images

Figure CN117831613B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method for identifying cancer driver genes. Background Technology
[0002] Cancer is characterized by the dysregulation of fundamental biological processes, including growth, proliferation, and cell death. The occurrence and progression of cancer are considered to be the result of the accumulation of alterations in driver genes, and accurately identifying cancer driver genes is crucial for understanding cancer pathogenesis and developing personalized cancer drugs. The availability of genomic, epigenomic, transcriptomic, and proteomic data from a large number of cancer patients provides valuable resources for researchers in this field. Furthermore, biological networks, as models describing biological entities and their interconnections, enable researchers to gain deeper insights into cancer from a biological systems perspective. Therefore, combining pan-cancer multi-omics data with biological networks has attracted widespread attention for identifying cancer driver genes. In recent years, deep learning has achieved certain successes in molecular biology and genomics, especially graph neural networks (GNNs), which have become a promising approach in this field. For example, the paper "Schulte-Sasse, R., Budach, S., Hnisz, D., et al.: Integration of multiomics data with graph convolutional networks to identify new cancer genes and their associated molecular mechanisms. Nature Machine Intelligence 3, 513–526 (2021)" describes EMOGI as a GCN-based cancer driver gene identification method that predicts cancer genes by combining protein-protein interaction (PPI) networks and pan-cancer multi-omics data. The paper "Peng, W., Tang, QR, Dai, W., et al.: Improving cancerdriver gene identification using multi-task learning on graph convolutional network. Briefings in bioinformatics (2021)" describes MTGCN as integrating PPI networks and multi-omics data, and improving the performance of cancer driver gene identification through a Chebyshev GCN-based multi-task learning method. While these methods represent a significant improvement over traditional approaches, cancer development often involves diverse aspects of gene interactions, such as metabolism, kinases, and regulation. These methods, focusing only on individual networks, inevitably overlook incompleteness and noise within these interactions. Furthermore, it's important to emphasize that most models neglect the effects of class imbalance, which can severely limit their predictive power.
[0003] Given the limitations of current research, it is necessary to provide a method that can efficiently identify cancer driver genes. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this invention provides a cancer driver gene identification method based on graph neural networks and negative sample generation. First, multiple biological networks and multi-omics data are selected, treating the multi-omics data as attributes of nodes in the network. A graph neural network method is used to obtain low-dimensional feature vectors of the nodes. Then, known cancer driver genes recorded in current databases are collected and labeled as positive samples. Anomaly detection algorithms are used to calculate the anomaly score of unlabeled genes relative to positive samples. Multiple databases are combined to filter unlabeled genes unrelated to cancer from largest to smallest as negative samples, with the number matching the number of positive samples to ensure a balance between positive and negative samples. Finally, the low-dimensional feature vectors of positive and negative samples are combined to form a standard dataset, which is then used to train a binary classifier. This binary classifier is applied to the remaining unlabeled genes to obtain the probability that these unlabeled genes are cancer driver genes. Finally, the nodes with the highest rankings are considered cancer driver genes and can be used for pan-cancer analysis.
[0005] The technical solution adopted by this invention to solve its technical problem includes the following steps:
[0006] Step 1: Select multiple biological networks to form a multi-dimensional network; the multiple biological networks are represented in the form of an adjacency matrix. R represents the number of biological networks; multi-omics data are collected as attributes of nodes in multi-dimensional networks. n represents the number of nodes in the multi-dimensional network, and f represents the dimension of the multi-omics data.
[0007] Step 2: By generating random node attributes, each network Both can generate a corresponding negative sample network. Each original network (the ground truth graph representation) and its negative example network (the perturbed graph representation) are simultaneously passed through a Graph Convolutional Neural Network (GCN) to obtain the encoding vector of each node in each network. This is then optimized using cross-entropy with logical values. Finally, the encoding vectors H of the nodes in the original network are... (r) With the encoding vector in the negative sample network The aggregated encoding vector H and the aggregated encoding vector obtained through the aggregation layer are respectively Finally, a regularizer is used to distinguish between the original network and the negative sample network to train the model and obtain the final feature representation vector Z of the gene;
[0008] Step 3: Based on the gene feature representation vector Z, collect known cancer genes from the database as positive samples, and the remaining genes are unlabeled genes; use an autoencoder to obtain the center c of a hypersphere from the positive samples, and then train a hypersphere through a neural network. Inputting unlabeled genes can obtain the abnormality scores of unlabeled genes compared with known cancer genes. Then sort the abnormality scores of the unlabeled genes from largest to smallest, and select genes that do not involve the "pathway in cancer" in the KEGG database, as well as genes that are not in the NCG database, the OMIM human Mendelian genetics database, and the COSMIC database, as negative samples. The number of negative samples is the same as the number of positive samples to achieve the purpose of balancing positive and negative samples.
[0009] Step 4: Finally, combine the positive and negative samples into a benchmark dataset, train an XGBoost classifier, and use it to predict the remaining unlabeled genes. Finally, obtain the probability that the remaining unlabeled genes are cancer driver genes.
[0010] Preferably, the biological network is a pathway network, a protein-complex network, or a kinase-substrate pairing network.
[0011] Preferably, the omics data are gene mutation omics, DNA methylation omics, and gene expression omics data.
[0012] Preferably, step 2 specifically comprises:
[0013] The DMGI algorithm is used to learn the feature representation vectors of nodes in a graph. First, GCN is introduced to generate the feature representation vectors of nodes in different networks. The encoding matrix H in (r) :
[0014]
[0015] Where A (r) Representing biological networks The adjacency matrix, yes The degree matrix, σ(·) represents the ReLU activation function, W (r) These are trainable parameters; similarly, the negative sample network... It can be obtained through GCN.
[0016] To obtain the graph-level node encoding vector representation, the Readout function is used:
[0017]
[0018] Where σ(·) represents the Sigmoid activation function. For node v i exist The encoded vector in;
[0019] Optimize the feature representation vector using a binary cross-entropy loss function with logical values.
[0020]
[0021] Where M (r) These are trainable parameters;
[0022] The global encoding vector H of the node is obtained by average pooling:
[0023]
[0024] Introducing a regularizer yields the feature representation vector Z of the node:
[0025]
[0026] The final objective function is expressed as:
[0027]
[0028] Where α and β are the regularizer and the l2 norm ||Θ||, respectively. 2 The coefficient.
[0029] Preferably, step 3 specifically comprises:
[0030] The DeepSVDD anomaly detection algorithm is used to generate anomaly scores for unlabeled genes, and negative samples are filtered out using a database. The goal of DeepSVDD is to learn a neural network, represented as... It maps the data to a hypersphere centered at c, in which... The goal of DeepSVDD is to ensure that normal points are contained within the hypersphere, while a small number of outliers are located outside the hypersphere, thereby achieving anomaly detection.
[0031]
[0032] Where pos represents a positive sample, i.e., a known cancer driver gene; c is obtained by pre-training an autoencoder.
[0033] After obtaining the trained hypersphere, for unlabeled genes, their anomaly scores s(z) unlabeled Calculate as follows:
[0034]
[0035] in These are the trained parameters.
[0036] Then, according to the abnormal score s(z) unlabeled Genes were selected as negative samples in descending order and recursively, provided that they did not involve the “pathway in cancer” pathway in the KEGG database, and were not in the NCG, OMIM, and COSMIC databases; the number of negative samples selected matched the number of positive samples, and the standard dataset consisted of positive and negative samples and their feature representation vectors.
[0037] The beneficial effects of this invention are as follows:
[0038] Considering the incompleteness of gene interactions in a single network, the method of this invention focuses on multi-layer biological networks. This method reduces the noise level of biological networks and represents nodes in weighted multi-layer networks as vectors, while effectively preserving structural information for identifying prognostic biomarkers. Attached Figure Description
[0039] Figure 1 A framework diagram of the method of this invention;
[0040] Figure 2 The AUROC and AUPRC values obtained from model training in this embodiment of the invention;
[0041] Figure 3 A comparison chart of AUROC and AUPRC before and after using the negative samples generated by the method of this invention.
[0042] Figure 4 This is a GO enrichment analysis diagram of the top 100 cancer driver genes predicted in an embodiment of the present invention.
[0043] Figure 5 (a) is a centrality comparison diagram of the top 100 cancer driver genes predicted by the embodiments of the present invention with known cancer genes and non-cancer genes.
[0044] Figure 5 (b)- Figure 5 (d) Cell line analysis diagrams for the ACTR2, ALDCA, and ANAPC2 genes, respectively. Detailed Implementation
[0045] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0046] The technical problem to be solved by the present invention is to provide a cancer driver gene identification method based on graph neural networks and negative sample generation, which addresses the shortcomings of existing technologies.
[0047] To solve the above-mentioned technical problems, the technical solution adopted in this invention is: integrating multi-source networks and pan-cancer multi-omics data, and using graph neural networks combined with negative sample generation to discover cancer driver genes, such as... Figure 1 As shown, it includes the following steps:
[0048] I. Data Collection
[0049] Multiple biological networks were selected to form a multivariate network for subsequent analysis. Classic biological networks were chosen: PPI networks, pathway networks, protein-complex networks, kinase-substrate pairing networks, metabolic networks, and regulatory networks; these biological networks can be represented in adjacency matrix form. R represents the number of biological networks. Multi-omics data is collected as attributes of nodes in multi-dimensional networks. n represents the number of nodes in the multi-omics network, and f represents the dimension of the multi-omics data. Multi-omics data includes omics data such as gene mutation omics, DNA methylation omics, and gene expression omics.
[0050] II. Learning the Feature Representation of Attributed Multivariate Networks
[0051] The DMGI (Multi-Network Embedding) algorithm is used to learn the feature representation vectors of nodes in a graph. Specifically, GCN (Generative Networking) is first introduced to generate feature representation vectors of nodes in different networks. The encoding matrix H in (r) :
[0052]
[0053] Where A (r) Representing biological networks The adjacency matrix, yes The degree matrix, σ(·) represents the ReLU activation function, W (r) These are trainable parameters. Similarly, the negative sample network... It can be obtained through GCN.
[0054] To obtain the graph-level node encoding vector representation, the Readout function is used:
[0055]
[0056] Where σ(·) represents the Sigmoid activation function. For node v i exist The encoded vector in.
[0057] The feature representation vector can be optimized using the binary cross-entropy loss function with logical values.
[0058]
[0059] Where σ(·) represents the Sigmoid activation function, M (r) These are trainable parameters.
[0060] The global encoding vector H of a node can be obtained using average pooling:
[0061]
[0062] Introducing a regularizer yields the feature representation vector Z of a node:
[0063]
[0064] The final objective function can be expressed as:
[0065]
[0066] Where α and β are the regularizer and the l2 norm ||Θ||, respectively. 2 The coefficient.
[0067] III. Generating Reliable Negative Samples
[0068] The DeepSVDD anomaly detection algorithm is used to generate anomaly scores for unlabeled genes, and reliable negative samples are selected using a database. The goal of DeepSVDD is to learn a neural network, represented as... It maps the data to a hypersphere centered at c. In this new space, The goal of DeepSVDD is to ensure that normal points are contained within the hypersphere, while a small number of outliers are located outside the hypersphere, thereby achieving anomaly detection.
[0069]
[0070] Where pos represents a positive sample, i.e., a known cancer driver gene, which can be collected from a database; c can be obtained by pre-training an autoencoder.
[0071] After obtaining the trained hypersphere, for unlabeled genes, their anomaly scores s(z) unlabeled It can be calculated as follows:
[0072]
[0073] in These are the trained parameters.
[0074] Then, according to the abnormal score s(z) unlabeled Genes were selected as negative samples in descending order and recursively, provided they did not involve the "pathway in cancer" pathway in the KEGG database, and were not found in the NCG, OMIM, and COSMIC databases. The number of selected negative samples matched the number of positive samples, forming a standard dataset consisting of positive and negative samples and their feature representation vectors.
[0075] IV. Cancer Driver Gene Prediction
[0076] XGBoost was trained using a standard dataset, and the trained XGBoost was used to predict the probability that the remaining unlabeled genes were cancer driver genes. A certain number of predicted high-probability cancer driver genes were selected for subsequent pan-cancer analysis.
[0077] Example:
[0078] To demonstrate the feasibility of the method of this invention, considering the complexity of cancer itself, six classic biological networks were selected and input into the algorithm model, including interaction networks, pathway networks, protein complex networks, kinase-substrate pairing networks, metabolic enzyme coupling interaction networks, and regulatory interaction networks compiled from the literature. For multi-omics data, gene expression data, gene variation data, and DNA methylation data were selected. Figure 2 The AUROC (0.9334) and AUPRC (0.9477) results after training the model and performing five-fold cross-validation are shown. The results demonstrate that the model proposed in this invention can effectively identify cancer driver genes.
[0079] (1) Comparison of cancer gene recognition performance
[0080] A comparative analysis was conducted on the method of this invention with other cancer driver gene identification methods, including MODIG and MTGCN, as well as baseline methods such as GCN, GAT, and DeepWalk integrated with XGBoost. The performance of the method of this invention and other methods was evaluated using five-fold cross-validation. It is worth emphasizing that, due to the involvement of multiple networks in the model of this invention, to ensure the consistency of input data, six networks, except for MODIG, were directly fused as input to the aforementioned comparative methods. For fair comparison, the positive and negative samples of the comparative methods were aligned with the model of this invention, and the parameters of these methods were set to default values. Table 1 shows that the model of this invention achieved the best performance among the comparative methods, indicating that this model can identify cancer driver genes from multi-omics data in complex multi-network relationships.
[0081] Table 1
[0082] AUROC 0.8568 0.9038 0.9060 0.8826 0.8757 0.9334 AUPRC 0.8525 0.9127 0.8925 0.8714 0.8892 0.9477
[0083] In addition, the performance of the comparison method was evaluated with and without using negative samples generated by the model of this invention, in order to assess the impact of sample balancing on performance. Figure 3 The results show that these contrastive methods improve performance in terms of AUC and AUPRC after introducing the generated negative samples. These results demonstrate that the negative samples generated by the model of this invention are reliable.
[0084] (2) The impact of multi-omics data and multiple networks on model capabilities
[0085] The impact of different networks and various omics data on model performance was analyzed, and the results are shown in Table 2. It was observed that integrating diverse networks significantly improved model performance, indicating that the model of this invention achieved the highest performance on a multi-gene network integrating all six biological networks. To evaluate the impact of multi-omics attributes on model performance, a comprehensive assessment was conducted, comparing model performance under different attribute combinations, including single attributes such as methylation, mutation, and gene expression, as well as attribute pairs such as expression and methylation, expression and mutation, methylation and mutation, and multi-omics attributes. Furthermore, the model performance without multi-omics attributes was significantly reduced, indicating that multi-omics data has a significant impact on improving the predictive ability of the model.
[0086] Table 2
[0087]
[0088] (3) Interpretability of predicted cancer driver genes
[0089] Two candidate cancer driver gene databases were used to analyze the top 100 predicted novel cancer driver genes (PCGs). One is Cancer Mine (bionlp.bcgsc.ca / cancermine / ), a database of driver genes, oncogenes, and tumor suppressor genes based on literature mining. The other is the Candidate Cancer Gene Database (umn.edu), which includes candidate driver genes and the precise genomic coordinates of all transposon-based screening experiments in the literature. It was observed that 81% of the PCGs had at least one cited study providing evidence of their status as a driver gene, demonstrating the effectiveness of the model presented in this invention.
[0090] The centrality score of each node in each network was calculated, and the average centrality score of each node was obtained. Figure 4The results show that PCGs exhibit significantly higher centrality values compared to non-cancer genes (NCGs), suggesting the importance of these predicted genes in the network. Meanwhile, a significant difference in centrality distributions exists between known cancer genes (KCGs) and NCGs, indicating the reliability of the negative samples generated by the model of this invention.
[0091] To further analyze PCGs from a functional perspective, gene ontology (GO) enrichment analysis was performed. These 100 genes were included in cancer-related GO entries such as "cell cycle," "mRNA destabilization," and "negative regulation of translation." Figure 4 Furthermore, based on the CRISPR Public 23Q2 dataset, a pan-cancer analysis of the dependency scores of target genes was calculated using DepMap (The Cancer Dependency Map Project at Broad Institute). A lower dependency score means that the gene is more essential for cell growth in a given cell line. 0 indicates that the gene is not essential, while -1 is roughly equivalent to the median of essential genes in pan-cancer studies (red line). Figure 5 This indicates that three of the first five PCGs (ACTR2, ALDCA, and ANAPC2) are essential genes. While the dependency scores of the remaining two genes were not statistically significant, they have been shown to be associated with cancer. For example, EFNA3 has been identified as a tumor suppressor gene in malignant peripheral nerve sheath tumor cells and is considered a driver gene in the progression of lung adenocarcinoma. CARD14 plays a crucial role in the regulation of breast cancer cell proliferation; its expression level was significantly increased in breast cancer samples compared to normal samples. Furthermore, high CARD14 expression is associated with decreased survival, invasiveness, and recurrence in human prostate cancer patients.
Claims
1. A method for identifying cancer driver genes based on graph neural networks and negative sample generation, characterized in that, Includes the following steps: Step 1: Select multiple biological networks to form a multi-dimensional network; the multiple biological networks are represented in the form of an adjacency matrix. , , ..., , Indicates the number of biological networks; Collecting multi-omics data as attributes of nodes in multi-dimensional networks , This indicates the number of nodes in a multi-element network. Representing the dimensions of multi-omics data; Step 2: By generating random node attributes, each network Both can generate a corresponding negative sample network. Each original network (the ground truth graph representation) and its negative example network (the perturbed graph representation) are simultaneously passed through a Graph Convolutional Neural Network (GCN) to obtain the encoding vector of each node in each network. This is then optimized using cross-entropy with logical values. Finally, the encoding vectors of the nodes in the original network are... With the encoding vector in the negative sample network The aggregated encoding vectors of the nodes in the original network are obtained through the aggregation layer. Aggregate encoding vectors in negative sample networks Finally, a regularizer is used to distinguish between the original network and the negative sample network to train the model and obtain the final feature representation vector of the gene. ; Step 3: Gene-based feature representation vector Known cancer genes were collected from the database as positive samples, while the remaining genes were unlabeled. The positive samples were then processed by an autoencoder to obtain the center of the hypersphere. Then, a hypersphere is trained through a neural network. Inputting unlabeled genes can obtain the abnormal scores of unlabeled genes compared with known cancer genes. The abnormal scores of unlabeled genes are then sorted from largest to smallest. Genes that do not involve the "pathway in cancer" pathway in the KEGG database, as well as genes that are not in the NCG database, the human Mendelian genetics database OMIM, and the COSMIC database, are selected as negative samples. The number of negative samples is the same as the number of positive samples to achieve the purpose of balancing positive and negative samples. Step 4: Finally, combine the positive and negative samples into a benchmark dataset, train an XGBoost classifier, and use it to predict the remaining unlabeled genes. Finally, obtain the probability that the remaining unlabeled genes are cancer driver genes.
2. The cancer driver gene identification method based on graph neural network and negative sample generation according to claim 1, characterized in that, The biological network is a pathway network, a protein-complex network, and a kinase-substrate pairing network.
3. The cancer driver gene identification method based on graph neural network and negative sample generation according to claim 1, characterized in that, The omics data include gene mutation omics, DNA methylation omics, and gene expression omics data.
4. The cancer driver gene identification method based on graph neural network and negative sample generation according to claim 1, characterized in that, Step 2 specifically involves: The DMGI algorithm is used to learn the feature representation vectors of nodes in a graph. First, GCN is introduced to generate the feature representation vectors of nodes in different networks. The encoding matrix in : in Representing biological networks The adjacency matrix, , yes The degree matrix, Represents the ReLU activation function. These are trainable parameters; similarly, the negative sample network... It can be obtained through GCN. ; To obtain the graph-level node encoding vector representation, the Readout function is used: in This represents the Sigmoid activation function. For nodes exist The encoded vector in; Optimize the feature representation vector using a binary cross-entropy loss function with logical values. : in These are trainable parameters; The average pooling method is used to obtain the aggregated encoding vector of the nodes in the original network. : Introducing a regularizer yields the feature representation vector of a node. : The final objective function is expressed as: in, and Regularizer and norm The coefficient.
5. The cancer driver gene identification method based on graph neural network and negative sample generation according to claim 4, characterized in that, Step 3 specifically involves: The DeepSVDD anomaly detection algorithm is used to generate anomaly scores for unlabeled genes, and negative samples are filtered out using a database. The goal of DeepSVDD is to learn a neural network, represented as... It maps data to a [agent / object] Within the hypersphere space centered on this new space... The goal of DeepSVDD is to ensure that normal points are contained within the hypersphere, while a small number of outliers are located outside the hypersphere, thereby achieving anomaly detection. in This represents a positive sample, i.e., a known cancer driver gene; , It is obtained by pre-training an autoencoder; After obtaining the trained hypersphere, the abnormal scores of unlabeled genes are... Calculate as follows: in These are pre-trained parameters; Then, according to the abnormal scores Genes were selected as negative samples in descending order and recursively, provided that they did not involve the "pathway in cancer" pathway in the KEGG database, and were not in the NCG, OMIM, and COSMIC databases; the number of negative samples selected matched the number of positive samples, and the standard dataset consisted of positive and negative samples and their feature representation vectors.
Citation Information
Patent Citations
Synthetic lethal gene prediction method and device based on graph neural network, terminal and medium
CN115240777A
Graph neural network method based on multi-scale graph contrast learning
CN115994560A