Multi-scale two-channel graph convolutional network for cancer driving gene identification

Through a multi-scale dual-channel graph convolution network, the personalized PageRank algorithm and enhanced dual-channel interaction module are used to solve the problem of multi-scale feature representation and global structural information processing in heterophilic biomolecular networks, and the recognition accuracy of cancer driver genes is improved.

CN120260675APending Publication Date: 2025-07-04SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325172.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing GCN-based methods deal with heterophilic molecular networks, it is difficult to effectively capture multi-scale feature representation and process global structural information, resulting in insufficient accuracy in cancer-driven gene recognition.

Method used

A multi-scale dual-channel graph convolution network is adopted to generate auxiliary networks through a personalized PageRank algorithm, combining multi-level feature extraction modules and enhanced dual-channel interactive modules to realize adaptive fusion of local and global features, and fusion of feature representations of each layer through dual residual connections and adaptive weights to improve recognition accuracy.

Benefits of technology

It significantly improves the accuracy of cancer-driven gene recognition in heterophilic biomolecular networks, improves AUROC and AUPRC performance, and overcomes the limitations of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260675A_ABST
    Figure CN120260675A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale two-channel graph convolutional network for cancer driven gene identification, and belongs to the technical field of biological information. The method comprises the following steps: firstly, inputting a heterophilic biomolecular network, gene node features in the network and an auxiliary network generated by using a personalized PageRank algorithm into a main module, and realizing self-adaptive fusion of local-global features through a multi-level feature extraction module; then, an improved graph attention convolutional network (GAT) is adopted to process different types of associations between genes, then dynamic information interaction between two channels is achieved through an enhanced dual-channel interaction module, and finally, multi-level features are fused through dual residual connection and self-adaptive weight. A probability score that comprehensively indicates that each gene is predicted as a cancer driver gene is formed, and the cancer driver gene is identified from the predicted score. According to the method, the problem that the heterophilic biomolecular network still has limitation in the aspects of capturing multi-scale feature representation and processing global structure information of the heterophilic network is solved, and the accuracy of cancer driver gene recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a multi-scale dual-channel graph convolutional network for cancer driver gene identification. Background Art

[0002] Cancer is a group of diseases characterized by abnormal cell proliferation caused by genetic mutations, which endow cancer cells with a selective growth advantage over surrounding normal cells. Genes associated with these driver mutations are called cancer driver genes. Accurately identifying cancer driver genes is of great significance for understanding the cancer pathogenesis, discovering biomarkers, and developing targeted treatment strategies.

[0003] In recent years, large-scale cancer research projects such as TCGA and ICGC have accumulated a large amount of multi-omics data, providing valuable resources for the computational identification of cancer driver genes. Based on these data, researchers have developed various computational methods for identifying cancer driver genes, including mutation frequency-based methods, network-based methods, and machine learning-based methods. Although these methods have made significant progress in identifying cancer driver genes, they still have limitations in dealing with the complex structure of biomolecular networks and capturing multi-scale feature representations.

[0004] Recently, GNN-based methods, especially GCN, have been applied to the identification of cancer driver genes, such as EMOGI, MTGCN, and HGDC, showing superior performance. These methods integrate multi-omics data and biomolecular networks to learn the low-dimensional representation of genes, thereby distinguishing cancer driver genes from non-driver genes. However, most of the existing GCN-based methods are based on the homophily assumption, that is, nodes connected in the network usually belong to the same category or have similar features. This assumption often does not hold in biomolecular networks because cancer driver genes tend to participate in different signaling pathways, providing a synergistic effect on cancer growth. In addition, the number of cancer driver genes (hundreds) is far less than the number of nodes in the biomolecular network (thousands), which leads to the heterophilic setting of the biomolecular network, that is, cancer driver genes are more likely to interact with other genes (even non-driver genes) rather than with other cancer driver genes.

[0005] Although methods such as HGDC have made progress in dealing with heterophilic biomolecular networks, they still have limitations in capturing multi-scale feature representations and processing the global structural information of heterophilic networks. The present invention proposes a multi-scale dual-channel graph convolutional network for cancer driver gene identification (MSDC), which overcomes the above limitations and improves the accuracy of cancer driver gene identification by introducing a multi-scale feature extraction mechanism and an enhanced dual-channel interaction module. Summary of the Invention

[0006] To address the above problems, the present invention provides a multi-scale two-channel graph convolutional network for cancer driver gene identification. It applies the personalized PageRank algorithm to enrich the original gene features, enabling the present invention to capture more structural information in the heterogeneous biomolecular network during calculation. Subsequently, a multi-level feature extraction mechanism and an enhanced two-channel interaction module are introduced in the present invention, solving the problem of still having limitations in capturing multi-scale feature representations and processing the global structural information of heterogeneous networks, which helps to improve the model performance. These modules enable the present invention to outperform other traditional methods in terms of the area under the receiver operating characteristic curve (AUROC) and the area under the precision-recall curve (AUPRC), and can accurately identify cancer driver genes in the heterogeneous biomolecular network.

[0007] To achieve the above object, the technical solution of the present invention is: to provide a multi-scale two-channel graph convolutional network for cancer driver gene identification. First, a personalized PageRank algorithm is used on the heterogeneous biomolecular network to generate an auxiliary network. Then, the heterogeneous biomolecular network, the auxiliary network, and the gene feature matrix are put into the main module. After training, new cancer driver genes are predicted, and the prediction scores of each gene in the heterogeneous biomolecular network are output, and cancer driver genes are identified according to the prediction scores.

[0008] The above multi-scale two-channel graph convolutional network for cancer driver gene identification specifically includes the following steps:

[0009] step1: Construct a heterogeneous biomolecular network and an auxiliary network, and obtain gene features as the input features of network nodes;

[0010] step2: Use the multi-scale feature extraction module to initially process the input features, and capture local and global feature information simultaneously;

[0011] step3: Extract features from the gene interaction network through an improved graph convolutional network, where the first channel processes the original biomolecular network and the second channel processes the auxiliary network. Then, an enhanced channel interaction module is used to achieve dynamic information exchange and fusion between the two channels;

[0012] step4: Adopt a hierarchical feature fusion mechanism to synthesize feature representations at different levels, and then perform final prediction by adaptively weighting the feature representations of multiple layers of the network. Cancer driver genes are identified according to the prediction scores.

[0013] Preferably, the heterophilic biomolecular network is a pathway network, a gene interaction network, or a protein interaction network.

[0014] Preferably, the auxiliary network is generated by inputting the heterophilic biomolecular network and the gene node features in the network into the personalized PageRank algorithm.

[0015] Preferably, step 2 is specifically as follows:

[0016] The core of this module is to balance the importance of local and global information through an adaptive attention mechanism. Its formal expression is as follows:

[0017] H locla = F elu (W transform x + b transform )

[0018] Where is the input feature, and are learnable transformation parameters, F elu is the exponential linear unit (ELU) activation function, represents the local transformation feature.

[0019] For the global feature, we obtain it through the average pooling operation of graph convolution:

[0020]

[0021] Where V represents the set of nodes, |V| is the number of nodes, and the global feature captures the statistical characteristics of the entire graph.

[0022] To adaptively fuse local and global information, we introduce an attention mechanism:

[0023] α = σ(W att [H local ∥ H global + b att )

[0024] Where and are the parameters of the attention network, σ is the Sigmoid activation function, ∥ represents the feature concatenation operation, and α ∈ [0, 1] represents the attention weight.

[0025] Finally, the multi-scale feature representation is obtained through weighted combination:

[0026] H multi = α · H local + (1 - α) · H global

[0027] Preferably, step 3 is specifically as follows:

[0028] To fully utilize the complementary information of the auxiliary network and the original biomolecular network, we designed an enhanced dual-channel interaction module to achieve dynamic information exchange between the two network channels. Its formal expression is as follows:

[0029] First, perform non-linear transformation on the two input features x1 and x2 (from the original network and the auxiliary network respectively):

[0030] t1 = F transform (x1) = W transform x1 + b transform

[0031] t2 = F transform (x2) = W transform x2 + btransform

[0032] Wherein, and are shared transformation parameters to ensure the consistency of the feature spaces of the two channels.

[0033] Then, we calculate the inter-channel attention and gating mechanism to control the information flow:

[0034] α att = σ(W att [t1 ∥ t2] + b att )

[0035] g1 = σ(W gate1 [t1 ∥ t2] + b gate1 )

[0036] g2 = σ(W gate2 [t1 ∥ t2] + b gate2 )

[0037] Wherein, W att 、W gate1 、 and are learnable parameters, σ is the Sigmoid activation function, att represents the inter-channel attention coefficient, and g1 and g2 are the gating values for controlling information exchange.

[0038] Finally, update the feature representation through the gating attention mechanism and residual connection:

[0039]

[0040] Preferably, step 4 is specifically as follows:

[0041] To comprehensively utilize information at different levels, we designed a hierarchical feature fusion mechanism that fuses representations of each layer through dual residual connections and adaptive weights. Its formal expression is as follows:

[0042] First, dual residual connections are applied after each convolutional layer:

[0043]

[0044] where and are the outputs of the l-th layer on the original network and the auxiliary network respectively, H enh is the initial enhanced feature, and are fusion parameters, is the representation of the l-th layer after fusion.

[0045] Finally, features of different layers are fused through adaptive weights:

[0046]

[0047] w = softmax([w0, w1, w2, w3])

[0048]

[0049] where and are output transformation parameters, w l is the learnable layer weight parameter, and the softmax function ensures that the weights sum to 1, is the final prediction score, and cancer driver genes are identified based on the prediction score.

[0050] The present invention has the following beneficial effects compared with the prior art:

[0051] (1) The multi-scale feature extraction mechanism proposed by the present invention can simultaneously capture local fine features and global context information, providing a more comprehensive gene representation, thereby improving the model's ability to understand complex biological relationships;

[0052] (2) The enhanced dual-channel interaction module designed by the present invention realizes adaptive information exchange between the two network channels through a dynamic gating mechanism and an attention mechanism, effectively integrating complementary information and enhancing the model's expressive ability;

[0053] (3) The hierarchical (adaptive) feature fusion mechanism introduced by the present invention fuses representations of each layer through dual residual connections and adaptive weights, alleviating the problems of feature over-smoothing and long-range dependence in deep graph neural networks, and ensuring the high-fidelity transmission of key biological signals. Description of the Drawings

[0054] By reading the following detailed description of the preferred embodiments, various other advantages and features will become more apparent. The accompanying drawings are only used to illustrate the preferred embodiments and are not intended to limit the present invention. In the drawings:

[0055] Figure 1 It is a schematic diagram of a multi-scale two-channel graph convolutional network architecture according to an embodiment of the present invention;

[0056] Figure 2 It shows a schematic diagram of a multi-scale feature extraction module according to an embodiment of the present invention;

[0057] Figure 3 It shows a schematic diagram of an enhanced two-channel interaction module according to an embodiment of the present invention. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] The present invention provides a multi-scale two-channel graph convolutional network for cancer driver gene identification, which overcomes the problem that the heterogeneous biological molecular network still has limitations in capturing multi-scale feature representations and processing the global structural information of the heterogeneous network, and improves the accuracy of cancer driver gene identification.

[0060] As Figure 1 shown, the present invention provides a multi-scale two-channel graph convolutional network for cancer driver gene identification, and the specific implementation process is as follows:

[0061] Step 1: Obtain the xenobiotic biomolecular network, including: an RNA interaction network (GGNet) with 11,183 nodes and 621,988 edges obtained from ENCORI (starBase v2.0); a pathway network (PathNet) with 7,695 nodes and 92,710 edges obtained by retaining KEGG and Reactome pathways from the human functional protein network; and a protein-protein interaction network (PPNet) with 11,395 nodes and 285,843 edges constructed based on interactions with a confidence level of the top 5% from STRING v11. The labels of genes (i.e., whether they are cancer driver genes) are determined based on two lists, one list containing the names of 796 genes identified as cancer driver genes and the other list containing the names of 2,187 genes identified as non-cancer driver genes. Genes identified as cancer driver genes are labeled as positive samples, non-cancer driver genes are labeled as negative samples, and genes not included in positive or negative samples are considered unlabeled data. The feature information of each gene consists of biomolecular features calculated from cancer-specific multi-omics data and system-level features calculated from global gene features. Specifically, the cancer-specific biomolecular features are obtained by processing gene mutation, DNA methylation, and gene expression data of 16 cancer types, totaling 48 dimensions; the system-level features include 10 features such as whether the gene is homologous to a cancer driver gene and the proportion of the gene in cancer cell lines. Finally, each gene is represented by a 58-dimensional feature vector.

[0062] Step 2: As Figure 2 shown, it is a schematic diagram of the multi-scale feature extraction module according to an embodiment of the present invention. First, a non-linear mapping is performed on the input features using a linear transformation and an ELU activation function to generate local feature vectors; subsequently, global average pooling is used to obtain the overall statistical features of the network, and the macroscopic structure information is restored to the node dimension through a batch mapping mechanism. The core innovation lies in designing a dynamic attention weight mechanism, constructing a joint representation of local and global features through a Sigmoid function, and generating an adaptive attention weight ranging from [0,1]. Finally, the module dynamically adjusts the importance of local and global features based on this weight and outputs a fused multi-scale representation vector. Its formal expression is as follows:

[0063] The extraction of local features is achieved through the following formula:

[0064] H local = F elu (W transform x + b transform )

[0065] where is the input feature, and are learnable transformation parameters, F elu is the exponential linear unit (ELU) activation function, represents local transformed features.

[0066] For global features, they are obtained through graph-level average pooling operation:

[0067]

[0068] where, V represents the set of nodes, |V| is the number of nodes, and the global feature captures the statistical characteristics of the entire graph.

[0069] To adaptively fuse local and global information, an attention mechanism is introduced:

[0070] α = σ(W att [H local ∥H global +b att )

[0071] where, and are the parameters of the attention network, σ is the Sigmoid activation function, ∥ represents the feature concatenation operation, and α ∈ [0, 1] represents the attention weight.

[0072] Finally, the multi-scale feature representation is obtained through weighted combination:

[0073] H multi = α · H local + (1 - α) · H global

[0074] Step 3: As Figure 3 shown, the schematic diagram of the enhanced dual-channel interaction module according to the embodiment of the present invention. First, the module maps the features of different network channels to the same feature space through linear transformation, eliminates semantic differences, and ensures the interactivity of features. Its core innovation lies in introducing a dynamic attention mechanism, using the Sigmoid function to accurately calculate the attention weights of cross-channel features, and adaptively adjusting the interaction intensity between network channels. Further, the bidirectional gating mechanism cleverly controls the information flow and fusion ratio of the features of each channel, realizing asymmetric and dynamic information transmission. In the feature update link, the core information of the original features is retained through residual connection, and incremental feature representation is carried out with the help of gating weights and attention coefficients, effectively alleviating the problem of gradient disappearance commonly found in deep networks. Its formal expression is as follows:

[0075] First, non-linear transformation is performed on two input features x1 and x2 (from the original network and the auxiliary network respectively):

[0076] t1 = F transform(x1) = W trabsform x1 + b transform

[0077] t2 = F transform (x2) = W transform x2 + b transform

[0078] Wherein, and are shared transformation parameters to ensure the consistency of the feature spaces of the two channels.

[0079] Then, channel attention and gating mechanisms are calculated to control the information flow:

[0080] α att = σ(W att [t1 ∥ t2] + b att )

[0081] g1 = σ(W gate1 [t1 ∥ t2] + b gate1 )

[0082] g2 = σ(W gate2 [t1 ∥ t2] + b gate2 )

[0083] Wherein, W att 、W gate1 、 and are learnable parameters, σ is the Sigmoid activation function, att represents the channel attention coefficient, and g1 and g2 are the gating values for controlling information exchange.

[0084] Finally, the feature representation is updated through the gating attention mechanism and residual connection:

[0085]

[0086] Step 4: In each layer of convolution in the network, the module first enhances the transitivity of the features through double residual connections to ensure that the key information at the bottom layer will not be lost during deep propagation. Specifically, in each layer of features of the original network and the auxiliary network, the ELU activation function and residual connection are applied, and the features of the two network channels are jointly processed through a linear fusion layer. During the fusion process, the model introduces an adaptive weight mechanism, and uses the Softmax function to dynamically adjust the feature weights of different layers, enabling the network to autonomously learn the importance of each layer of features according to the task characteristics. Finally, by performing a weighted sum of the feature representations from layer 0 to layer 3, the model generates a comprehensive and multi-scale feature representation. Its formal expression is as follows:

[0087] First, apply double residual connections after each convolution layer:

[0088]

[0089] Among them, and are the outputs of the l-th layer on the original network and the auxiliary network respectively, and H enh is the initial enhanced feature, and are fusion parameters, is the representation of the l-th layer after fusion.

[0090] Finally, fuse the features of different layers through adaptive weights:

[0091]

[0092] w = softmax([w0, w1, w2, w3])

[0093]

[0094] Among them, and are output transformation parameters, w l is the learnable layer weight parameter, and the softmax function ensures that the weights sum to 1. is the predicted score of each gene in the final output xenobiotic molecule network. Cancer driver genes are identified based on the predicted scores. The higher the gene prediction score, the greater the probability that it is a cancer driver gene.

[0095] To evaluate the performance of a multi-scale two-channel graph convolutional network provided by the present invention for cancer driver gene identification, the verification scheme adopts five-fold stratified cross-validation repeated ten times (abbreviation, ten-fold five-fold cross-validation method). This model is compared with deep learning methods (GAT, EMOGI, MTGCN, HGDC, etc.). It is set that all comparison methods use the same data set. Table 1 shows the AUROC and AUPRC performance of different methods on the xenobiotic molecule network.

[0096] Table 1

[0097]

[0098] As can be seen from Table 1, the performance of this model on the xenobiotic molecule network is better than that of existing methods. Compared with HGDC ranked second in performance, this model has an AUROC improvement of 0.0243 and an AUPRC improvement of 0.0644, that is, the accuracy of cancer driver gene identification in the xenobiotic molecule network is improved.

Claims

1. A multi-scale two-channel graph convolutional network for cancer driver gene identification, which specifically includes the following steps: Step 1: Construct a heterogeneous biomolecular network and an auxiliary network, and obtain gene features as the input features of network nodes; Step 2: Use a multi-scale feature extraction module to initially process the input features, and simultaneously capture local and global feature information; Step 3: Extract features from the gene interaction network through an improved graph convolutional network, where the first channel processes the original biomolecular network and the second channel processes the auxiliary network. Then, use an enhanced channel interaction module to achieve dynamic information exchange and fusion between the two channels; Step 4: Adopt a hierarchical feature fusion mechanism to synthesize feature representations at different levels, and then perform final prediction by adaptively weighting and fusing the feature representations of multiple layers of the network, and identify cancer driver genes according to the prediction scores.

2. A multi-scale dual-channel graph convolutional network for cancer driver gene identification according to claim 1, wherein, The heterogeneous biomolecular network is a pathway network, a gene interaction network, and a protein-protein interaction network.

3. A multi-scale dual-channel graph convolutional network for cancer driver gene identification according to claim 1, wherein, The auxiliary network is generated by inputting the heterogeneous biomolecular network into the personalized PageRank algorithm.

4. A multi-scale two-channel graph convolutional network for cancer driver gene identification according to claim 1, characterized in that, The specific content of step 2 is as follows: The core of this module is to balance the importance of local and global information through an adaptive attention mechanism, and its formal expression is as follows: H local = F elu (W transformx + b transform ) Among them, is the input feature, and are learnable transformation parameters, and F elu is the exponential linear unit (ELU) activation function, represents the local transformation feature. For global features, we obtain them through the average pooling operation of graph convolution: Among them, V represents the set of nodes, |V| is the number of nodes, and the global feature captures the statistical characteristics of the entire graph. To adaptively fuse local and global information, we introduce an attention mechanism: α = σ(W att [H local ||H global +b att ) Among them, and are the parameters of the attention network, σ is the Sigmoid activation function, ∥ represents the feature concatenation operation, and α ∈ [0, 1] represents the attention weight. Finally, the multi-scale feature representation is obtained through weighted combination: H multi = α·H local + (1 - α)·H global 。 5. A multi-scale dual-channel graph convolutional network for cancer driver gene identification according to claim 1, characterized in that, The specific content of step 3 is as follows: To make full use of the complementary information between the auxiliary network and the original biomolecular network, we design an enhanced two-channel interaction module to achieve dynamic information exchange between the two network channels. Its formal expression is as follows: First, perform non-linear transformation on the two input features x1 and x2 (from the original network and the auxiliary network respectively): t1 = F transform (x1) = W transform x1 + b transform t2 = F transform (x2) = W transform x2 + b transform Among them, and are shared conversion parameters to ensure the consistency of the feature spaces of the two channels. Then, we calculate the inter-channel attention and gating mechanism to control the information flow: α att = σ(W att [t1||t2]+b att ) g1 = σ(W gate1 [t1||t2] + b gate1 ) g2 = σ(W gate2 [t1||t2]+b gate2 ) Among them, and are learnable parameters, σ is the Sigmoid activation function, att represents the attention coefficient between channels, and g1 and g2 are gating values that control information exchange. Finally, update the feature representation through the gating attention mechanism and residual connection:

6. A multi-scale two-channel graph convolutional network for cancer driver gene identification according to claim 1, characterized in that, The specific content of step 4 is as follows: To comprehensively utilize information at different levels, we design a hierarchical feature fusion mechanism to fuse the representations of each layer through double residual connections and adaptive weights. Its formal expression is as follows: First, apply double residual connections after each convolutional layer: Among them, and are the outputs of the l-th layer on the original network and the auxiliary network respectively, and H enh is the initial enhanced feature, and are the fusion parameters, is the representation of the l-th layer after fusion. The advantages of this dual residual structure are as follows: The first-level residual (H enh ) preserves the information of the original enhanced features and alleviates the problem of information loss in deep networks; the second-level residual (R l-1 ) establishes direct connections between layers, promotes gradient flow, and retains the local structural information of the shallow layers. Finally, fuse the features of different layers through adaptive weights: Among them, and are output transformation parameters, where w l are learnable layer weight parameters, and the softmax function ensures that the weights sum to 1. are the final prediction scores, and cancer driver genes are identified based on the prediction scores.