A method for identifying cancer driver genes based on multi-network graph convolution

Through the multi-network graph convolution method, the protein interaction network and biometric network are integrated, and combined with deep learning technology, the problem of low prediction accuracy of cancer-driven genes in the existing technology is solved, achieving higher recognition accuracy.

CN115019883BActive Publication Date: 2025-05-23KUNMING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210131034.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-13
Publication Date
2025-05-23
Estimated Expiration
2042-02-13

AI Technical Summary

Technical Problem

Existing graph convolution-based methods have low accuracy in predicting cancer-driven genes and fail to effectively exploit similarities between gene biometrics themselves.

Method used

The multi-network graph convolution method is adopted to build protein interaction networks and biometric networks, and combine deep learning technology to integrate multiomic data, which improves the prediction accuracy of cancer driver genes.

Benefits of technology

It significantly improves the recognition accuracy of cancer driver genes, provides more valuable reference information, and provides support for further research by biologists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019883B_ABST
    Figure CN115019883B_ABST
Patent Text Reader

Abstract

The present invention relates to a cancer driver gene identification method based on multi-network graph convolution, and belongs to the technical field of systems biology. The present invention first obtains a structural network based on a protein interaction network, then calculates enhanced features for each gene, obtains a feature network based on the similarity between the biological features of the gene, and then puts the structural network, feature network and biological features into a multi-network graph convolution model, trains the model, and uses the trained model to predict new cancer driver genes, and finally outputs a prediction score of whether each gene is a cancer driver gene. The present invention improves the accuracy of machine learning model prediction of cancer driver gene identification through a multi-network graph convolution recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cancer driver gene identification method based on multi-network graph convolution, belonging to the technical field of systems biology. Background Art

[0002] Mutations in a small number of genes can lead to the occurrence and development of cancer. Such genes are called cancer driver genes. Identifying cancer driver genes plays an important role in understanding the molecular mechanisms of cancer and developing precision therapeutic drugs. In the past decade, some large-scale cancer genomics projects have published genome, gene expression group, transcriptome and proteome data from thousands of cancer patients. These projects include The Cancer Genome Atlas (TCGA), the International Cancer Genome Consortium (ICGC) and the Catalog of Somatic Mutations in Cancer (COSMIC). These thousands of cancer genome data have promoted the development of computational methods to identify cancer driver genes.

[0003] These computational methods are designed based on one or more characteristics of cancer drivers. Early methods only used the mutation characteristics of genes, assuming that the more frequently mutated genes were more likely to be driver genes, such as MutSigCV. However, these methods often ignore driver genes with lower mutation frequencies. Since cancer driver genes often interact with each other to form protein complexes. Researchers introduce protein interaction (PPI) networks into the algorithm. For example, methods such as HotNet2 and MUFFINN predict driver genes by measuring the importance of genes in protein networks. In order to improve the prediction of cancer driver genes, some models reduce the unreliability of protein networks by introducing other types of data. For example, gene expression profiles, subcellular information, tissue-specific expression profiles, and gene function information. In addition to changes at the gene level, cancer can also occur through transcriptional and epigenetic dysregulation. For example, hypermethylation or hypomethylation of gene promoter regions can lead to the silencing of key tumor suppressor or oncogene activity, promoting cancer growth.

[0004] Biological networks are effective tools for describing the characteristics of biological entities and their relationships. In recent years, many methods for integrating multi-omics data based on biological networks have been developed. Network representation learning is an emerging technology that learns low-dimensional feature vectors for all network nodes while preserving the original network structure and node characteristics as much as possible. Therefore, it can effectively reduce the noise in the network and extract different types of node characteristics. Some methods based on deep neural networks, such as DeepWalk and node2vec, have been used to learn the characteristics of network nodes. Combining these characteristics with a classifier can predict cancer driver genes. Graph Convolutional Network (GCN) is a new network representation method that can learn the characteristics of network nodes by naturally integrating the network structure and node characteristics. EMOGI is a GCN-based method that uses biological omics data as gene characteristics and combines protein interaction networks to learn more complex gene characteristics. However, it does not consider the similarity between the gene biological characteristics themselves. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for identifying cancer driver genes based on multi-network graph convolution. Based on graph convolutional neural networks and integrating multi-omics data, it can effectively improve the accuracy of cancer driver gene identification, thus solving the problem of low accuracy in predicting cancer driver genes based on graph convolution in the prior art.

[0006] The technical solution of the present invention is: a method for identifying cancer driver genes based on multi-network graph convolution, and the specific steps are as follows:

[0007] Step1: Obtain a structure network according to the protein interaction network;

[0008] Step2: Calculate enhanced features for each gene;

[0009] Step3: Obtain a feature network according to the similarity between the biological characteristics of genes;

[0010] Step4: Put the structure network, the feature network, and the biological characteristics into a multi-network graph convolution model (MNGCN), train the model, and use the trained model to predict new cancer driver genes;

[0011] Step5: Output the prediction scores of whether each gene is a cancer driver gene.

[0012] The specific content of Step1 is: Delete the interaction data with a score less than 0.5 in the protein interaction data.

[0013] The specific content of Step2 is:

[0014] First, we calculate the differential methylation rate of the gene, which is the average value of the difference in methylation signals between cancer and matched normal samples in all samples of a cancer type:

[0015]

[0016] In the formula, represents the methylation value of gene i in cancer type c, and is the methylation signal in cancer and matched normal samples, S c A sample set representing cancer;

[0017] The differential expression rate of each gene was then obtained, which was measured by the log-fold change between its expression value in cancer and normal samples and then averaged across all samples;

[0018] Finally, the deepwalk algorithm is used in the structural network to obtain the network structural characteristics of each gene containing deep correlation relationships, determine the length of the network structural characteristics, and then concatenate them with the mutation rate, differential methylation rate, and differential expression rate of each gene in N types of cancer, and perform minimum and maximum normalization to obtain the enhanced characteristics of each gene.

[0019] The graph neural network-based method is limited by the problem of over-smoothing. The number of convolutional layers is often not suitable, but a small number of layers will cause the model to ignore the deep correlation between genes in the network. The multi-order neighbor information of genes in the interaction network is very helpful to improve the performance of predicting driver genes. Inspired by the previous random walk-based method, the present invention uses the DeepWalk algorithm in the interaction network to obtain the network structure characteristics of each gene containing deep correlation relationships, and then obtains enhanced features by splicing with biological features, directly using structural information to improve the accuracy of prediction.

[0020] The Step 3 is specifically as follows: according to the enhanced features obtained in Step 2, the cosine similarity between genes is calculated to obtain a cosine similarity matrix; for each gene, according to the cosine similarity matrix, 20 neighbors that are most similar to it are found and connected with edges to obtain a feature network.

[0021] The cosine similarity matrix is ​​an n*n matrix whose row and column names are genes, and each value in it is the similarity of the corresponding two genes. The values ​​in the matrix are between (0, 1), and the closer to 1, the more similar. The present invention finds the maximum 20 values ​​according to the cosine similarity matrix to determine the 20 most similar genes for each gene.

[0022] The multi-network graph convolution model used in the Step 4 is based on a graph convolution layer using the Chebyshev operator. The input is node features, structural networks, and feature networks. The node features, structural networks, node features, and feature networks are input into two sets of graph convolutions respectively, and the embedding matrices obtained by the two sets of graph convolutions are constrained for consistency. Finally, the prediction score is obtained and output through feature splicing and a fully connected layer.

[0023] The beneficial effects of the present invention are as follows: the present invention uses biological multi-omics data, obtains a structural network based on a protein interaction network, obtains a feature network based on the similarity between the biological characteristics of genes, and uses a multi-network graph convolution model to predict cancer driver genes.

[0024] The experimental results of the present invention show that compared with the existing methods for predicting driver genes based on machine learning and graph representation learning, the method proposed in the present invention can improve the accuracy of identifying cancer driver genes, and can provide valuable reference information for biologists to conduct experiments and further research on cancer driver gene identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flow chart of the MNGCN of the present invention;

[0026] Figure 2 It is a structural diagram of the multi-network graph convolution model used by the MNGCN of the present invention. DETAILED DESCRIPTION

[0027] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.

[0028] Example 1: The method described in the present invention is used to predict pan-cancer cancer driver genes, and the protein interaction data are from the Consensus Pathway Database (CPDB); the data of gene mutation, gene methylation and gene expression are from TCGA, involving more than 8,000 samples and 16 different cancer types (lung adenocarcinoma, lung squamous cell carcinoma, liver cancer, gastric cancer, colon cancer, esophageal cancer, rectal cancer, head and neck cancer, renal clear cell carcinoma, renal papillary cell carcinoma, thyroid cancer, endometrial cancer, cervical cancer, breast cancer, prostate cancer, bladder cancer).

[0029] The positive sample list was collected from NCG 6.0 labeled with this cancer type. For the negative sample list, starting from all genes, genes in NCG, COSMIC, OMIM databases and KEGG cancer pathways were recursively removed. Finally, the resulting benchmark dataset included 796 positive samples and 2187 negative samples.

[0030] like Figure 1 As shown in the figure, a cancer driver gene identification method based on multi-network graph convolution is shown in the figure. The specific steps are as follows:

[0031] Step 1: Obtain the structural network based on the protein interaction network;

[0032] After removing interactions with scores less than 0.5 in the protein interaction data, a structural network was obtained, which included 13,627 nodes and 504,378 edges.

[0033] Step 2: Calculate the enhancement features for each gene;

[0034] The mutation rate of a gene is defined as the average of single nucleotide variants (SNVs) and copy number aberrations (CNAs) across all samples in that cancer type.

[0035] The differential methylation rate of a gene is the average of the methylation signal differences between cancer and matched normal samples in all samples of a cancer type. It is defined as follows:

[0036]

[0037] in Represents the gene methylation value of gene i in cancer type c. and is the methylation signal in cancer and matched normal samples. S c A set of samples representing cancer.

[0038] The differential expression rate of each gene within a cancer type was measured by the log-fold change between its expression value in cancer and normal samples and then averaged across all samples.

[0039] The deepwalk algorithm was used in the structural network to obtain the network structure features of each gene containing deep correlation relationships, with a length of 16 dimensions, and then it was connected in series with the mutation rate, differential methylation rate, and differential expression rate of each gene in 16 types of cancer, and the minimum and maximum normalization was performed. In this way, each gene obtained a 64-dimensional enhanced feature.

[0040] Step 3: Obtain the feature network based on the similarity between the biological features of the genes;

[0041] According to the enhanced features obtained in step 2, the cosine similarity between genes is calculated to obtain the cosine similarity matrix. Each gene is connected to its 20 most similar neighbors with edges to obtain a feature network, including 13,627 nodes and 437,920 edges.

[0042] Step 4: Put the structural network, feature network and biological features into a multi-network graph convolutional model (MNGCN), train the model, and use the trained model to predict new cancer driver genes; the structure of the multi-network graph convolutional model is as follows Figure 2 shown.

[0043] The settings of Chebyshev graph convolution layers 1 and 2 are: input dimension 64, output dimension 300, Chebyshev filter size is set to 2, and the activation function is ReLU function. The settings of Chebyshev graph convolution layers 3 and 4 are: input dimension 300, output dimension 100, Chebyshev filter size is set to 2, and the activation function is ReLU function. The input dimension of the fully connected layer is 200, and the output dimension is 1. The dropout value during training is uniformly set to 0.5.

[0044] Compared with ordinary graph convolution layers, the Chebyshev graph convolution layer can more flexibly extract the features of multi-order neighbors and has a stronger ability to extract gene features and propagate information. Specifically:

[0045]

[0046] Z (k) This is calculated recursively by defining:

[0047] Z (1) =X

[0048]

[0049]

[0050] Z (k) It is an operator defined by bringing Chebyshev polynomials into graph neural networks, and X is the feature matrix.

[0051] Where H represents the genetic features learned by the Chebyshev graph convolutional layer. Θ∈R g×h is the weight parameter matrix of the neural network. h is the hidden layer dimension. k controls the size of the Chebyshev filter. Here k = 2. represents the scaled and normalized Laplacian matrix, Where I is the identity matrix, L is the regularized Laplacian matrix of the input network, and λ max is the largest eigenvalue of L. f is the activation function, such as ReLU.

[0052] The structural network, feature network and biological features are put into the multi-network graph convolutional model. The model predicts whether a gene is a cancer driver gene. It can be expressed as:

[0053]

[0054] in and is the embedding vector of gene i obtained by Chebyshev convolution layer 3 and 4 respectively, symbol represents the concatenation operation, w and b are learnable parameters.

[0055] The multi-network graph convolution model needs to minimize the loss function L during the training process. total as follows:

[0056] L total =L pre +αL con

[0057] Among them, L pre is the prediction loss, L con is a consistency constraint, α is a restriction L con We set the hyperparameter of , which we set to 0.0001.

[0058] Use the cross entropy loss function to calculate the prediction loss L pre :

[0059]

[0060] y i represents the known label (0 or 1) of gene i, n represents the number of genes in the training set, θ represents the parameters of the model. The smaller the prediction loss, the higher the final prediction accuracy of the model.

[0061] To further enhance the output embedding matrix H of Chebyshev convolutional layers 3 and 4 3 and H 4 commonality, design a consistency constraint L con .

[0062] First, use L 2 Normalization Normalize the embedding matrix to N 3 and N 4 .

[0063] Then, two normalized matrices are used to capture the similarity S of n nodes. 1 , S 2 , as follows:

[0064]

[0065]

[0066] Consistency means that the two similarity matrices should be similar, which results in the following constraints:

[0067]

[0068] Step 5: Output the prediction score of whether each gene is a cancer driver gene. The higher the prediction score, the higher the possibility that the gene is a cancer driver gene.

[0069] Example 2: The method described in the present invention is used to predict breast cancer driver genes, and the protein interaction data comes from the consensus pathway database (CPDB); the gene mutation, gene methylation and gene expression data come from TCGA.

[0070] The positive sample list was collected from NCG 6.0 labeled with that cancer type. For the negative sample list, we started with all genes and recursively removed genes from NCG, COSMIC, OMIM databases, and KEGG cancer pathways. Finally, our benchmark dataset for breast cancer includes 202 positive samples and 2187 negative samples.

[0071] like Figure 1 As shown in the figure, a cancer driver gene identification method based on multi-network graph convolution is shown in the figure. The specific steps are as follows:

[0072] Step 1: Obtain the structural network based on the protein interaction network;

[0073] After removing interactions with scores less than 0.5 in the protein interaction data, a structural network was obtained, which included 13,627 nodes and 504,378 edges.

[0074] Step 2: Calculate the enhancement features for each gene;

[0075] The mutation rate of a gene is defined as the average of single nucleotide variants (SNVs) and copy number aberrations (CNAs) across all samples in that cancer type.

[0076] The differential methylation rate of a gene is the average of the methylation signal differences between cancer and matched normal samples in all samples of a cancer type. It is defined as follows:

[0077]

[0078] in Represents the gene methylation value of gene i in cancer type c. and is the methylation signal in cancer and matched normal samples. S c A set of samples representing cancer.

[0079] The differential expression rate of each gene within a cancer type was measured by the log-fold change between its expression value in cancer and normal samples and then averaged across all samples.

[0080] The deepwalk algorithm was used in the structural network to obtain the network structure feature of each gene containing deep correlation relationships, with a length of 16 dimensions, which was then connected in series with the mutation rate, differential methylation rate, and differential expression rate of each gene in lung adenocarcinoma, and the minimum and maximum normalization was performed. In this way, each gene obtained a 19-dimensional enhanced feature.

[0081] Step 3: Obtain the feature network based on the similarity between the biological features of the genes;

[0082] According to the enhanced features obtained in step 2, the cosine similarity between genes is calculated to obtain the cosine similarity matrix. Each gene is connected to its 20 most similar neighbors with edges to obtain a feature network, including 13,627 nodes and 437,920 edges.

[0083] Step 4: Put the structural network, feature network and biological features into a multi-network graph convolutional model (MNGCN), train the model, and use the trained model to predict new cancer driver genes; the structure of the multi-network graph convolutional model is as follows Figure 2 shown.

[0084] The settings of Chebyshev graph convolution layers 1 and 2 are: input dimension 19, output dimension 150, Chebyshev filter size is set to 2, and the activation function is ReLU function. The settings of Chebyshev graph convolution layers 3 and 4 are: input dimension 150, output dimension 50, Chebyshev filter size is set to 2, and the activation function is ReLU function. The input dimension of the fully connected layer is 100, and the output dimension is 1. The dropout value during training is uniformly set to 0.5.

[0085] Compared with ordinary graph convolution layers, the Chebyshev graph convolution layer can more flexibly extract the features of multi-order neighbors and has a stronger ability to extract gene features and propagate information. Specifically:

[0086]

[0087] Z (k) This is calculated recursively by defining:

[0088] Z (1) =X

[0089]

[0090]

[0091] Where H represents the genetic features learned by the Chebyshev graph convolutional layer. Θ∈R g×h is the weight parameter matrix of the neural network. h is the hidden layer dimension. k controls the size of the Chebyshev filter. Here k = 2. represents the scaled and normalized Laplacian matrix, Where I is the identity matrix, L is the regularized Laplacian matrix of the input network, and λ max is the largest eigenvalue of L. f is the activation function, such as ReLU.

[0092] The structural network, feature network and biological features are put into the multi-network graph convolutional model. The model predicts whether a gene is a cancer driver gene. It can be expressed as:

[0093]

[0094] in and is the embedding vector of gene i obtained by Chebyshev convolution layer 3 and 4 respectively, symbol represents the concatenation operation, w and b are learnable parameters.

[0095] The multi-network graph convolution model needs to minimize the loss function L during the training process. total as follows:

[0096] L total =L pre +αL con

[0097] Among them, L pre is the prediction loss, L con is a consistency constraint, α is a restriction L con We set the hyperparameter of , which we set to 0.0001.

[0098] Use the cross entropy loss function to calculate the prediction loss L pre :

[0099]

[0100] y i represents the known label (0 or 1) of gene i, n represents the number of genes in the training set, θ represents the parameters of the model. The smaller the prediction loss, the higher the final prediction accuracy of the model.

[0101] To further enhance the output embedding matrix H of Chebyshev convolutional layers 3 and 4 3 and H 4 commonality, design a consistency constraint L con .

[0102] First, use L 2 Normalization Normalize the embedding matrix to N 3 and N 4 .

[0103] Then, two normalized matrices are used to capture the similarity S of n nodes. 1 , S 2 , as follows:

[0104]

[0105]

[0106] Consistency means that the two similarity matrices should be similar, which results in the following constraints:

[0107]

[0108] Step 5: Output the prediction score of whether each gene is a cancer driver gene. The higher the prediction score, the higher the possibility that the gene is a cancer driver gene.

[0109] Performance evaluation of cancer driver gene identification methods based on multi-network graph convolution:

[0110] In order to evaluate the performance of multi-network graph convolution (MNGCN), the proposed multi-network graph convolution method is compared with the traditional machine learning and graph convolution methods proposed previously. AUC, F1 score, precision, and recall are selected as evaluation indicators, and a comparative experiment between single-network graph convolution and multi-network graph convolution is conducted. All experimental results are averaged after ten 5-fold cross validations.

[0111] The experimental results for predicting pan-cancer cancer driver genes are shown in Table 1 , and the experimental results for predicting breast cancer driver genes are shown in Table 2 .

[0112]

[0113] Table 1

[0114]

[0115] Table 2

[0116] As can be seen from Table 1 and Table 2, the AUC and F1 score indicators of the MNGCN method are better than those of other prediction methods, and the precision and recall rates are excellent. It can be seen that the MNGCN method of the present invention has a very obvious advantage in predicting cancer driver genes compared with other existing methods.

[0117] Performance evaluation of multi-network strategies for cancer driver gene identification based on multi-network graph convolution:

[0118] In order to verify that the multi-network strategy of multi-network graph convolution (MNGCN) can improve performance, a comparative experiment was conducted between the graph convolution model using only the structure network (MNGCN-s) and the graph convolution model using only the feature network (MNGCN-f) and the multi-network graph convolution. The results are shown in Table 3.

[0119] Table 3

[0120]

[0121] As can be seen from Table 3, the AUC, F1 score, and recall rate indicators of the present invention are better than those of the method using a single network, proving that the multi-network strategy of multi-network graph convolution can improve the prediction performance.

[0122] In summary, after comparison with other prediction methods, the effectiveness of the cancer driver gene identification method based on multi-network graph convolution was demonstrated.

[0123] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A method for identifying cancer driver genes based on multi-network graph convolution, Features: Step 1: Obtain the structural network based on the protein interaction network; Step 2: Calculate the enhancement features for each gene; Step 3: Obtain the feature network based on the similarity between the biological features of the genes; Step 4: Put the structural network, feature network and biological features into the multi-network graph convolutional model, train the model, and use the trained model to predict new cancer driver genes; Step 5: Output the prediction score of whether each gene is a cancer driver gene; The Step 2 is specifically as follows: First, we calculate the differential methylation rate of the gene, which is the average value of the difference in methylation signals between cancer and matched normal samples in all samples of a cancer type: In the formula, represents the methylation value of gene i in cancer type c, and is the methylation signal in cancer and matched normal samples, S c A sample set representing cancer; The differential expression rate of each gene was then obtained, which was measured by the log-fold change between its expression value in cancer and normal samples and then averaged across all samples; Finally, the deepwalk algorithm is used in the structural network to obtain the network structure characteristics containing deep correlation relationships for each gene, which are then connected in series with the mutation rate, differential methylation rate, and differential expression rate of each gene in N types of cancer, and minimum and maximum normalization is performed to obtain the enhanced characteristics of each gene; The multi-network graph convolution model used in the Step 4 is based on a graph convolution layer using the Chebyshev operator. The input is node features, structural networks, and feature networks. The node features, structural networks, node features, and feature networks are input into two sets of graph convolutions respectively, and the embedding matrices obtained by the two sets of graph convolutions are constrained for consistency. Finally, the prediction score is obtained and output through feature splicing and a fully connected layer.

2. The method for identifying cancer driver genes based on multi-network graph convolution according to claim 1, It is characterized in that The Step 1 is specifically as follows: deleting the interaction data with a score less than 0.5 in the protein interaction data.

3. The method for identifying cancer driver genes based on multi-network graph convolution according to claim 1, It is characterized in that The Step 3 is specifically as follows: according to the enhanced features obtained in Step 2, the cosine similarity between genes is calculated to obtain a cosine similarity matrix; for each gene, according to the cosine similarity matrix, 20 neighbors that are most similar to it are found and connected with edges to obtain a feature network.

Citation Information

Patent Citations

  • Cancer driver gene prediction method and system based on local and global network centrality analysis

    CN113488104A