Method for specific driver gene and common driver gene identification based on federated transfer learning and deep learning algorithm

By integrating multi-omics data through federated transfer learning and deep learning algorithms, and utilizing Chebyshev graph convolutional networks and multi-head attention mechanisms, cancer-specific and common driver genes across tumors are identified. This addresses the shortcomings of existing methods in identifying cancer-specific and common driver genes, achieving more efficient identification results.

CN120412748BActive Publication Date: 2025-11-04YUNNAN UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510460660.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-11-04
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing methods struggle to simultaneously identify both cancer-specific and common driver genes, and most are limited to analyzing single-genetic data, failing to fully capture the complex genetic interactions and cross-cancer associations in cancer.

Method used

We employ a method based on federated transfer learning and deep learning algorithms, utilizing Chebyshev graphical convolutional networks and multi-head attention mechanisms to integrate multi-omics data. We train global and local models through a federated transfer learning framework to identify cancer-specific and shared driver genes across tumors.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of cancer driver gene identification, providing valuable insights for targeted and personalized treatment, and outperforms existing methods in performance across multiple cancer types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412748B_ABST
    Figure CN120412748B_ABST
Patent Text Reader

Abstract

The present application relates to a specific driver gene and common driver gene identification method based on federal transfer learning and deep learning algorithm, comprising: based on the data of different cancer types, constructing a data set; based on the multi-head attention mechanism, preprocessing the data set; based on the preprocessed data set, training a neural network model to obtain a gene identification model; wherein the neural network model is constructed based on Chebyshev graph convolution network and graph convolution network, and the model training adopts the federal transfer learning method, and the training process is divided into server side and client side; based on the gene identification model, the cancer specific and common driver genes across tumors are identified. Compared with the existing method, the present application can more accurately and efficiently identify the specific and common driver genes across tumors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene identification, and particularly relates to a specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm. BACKGROUND

[0002] Cancer is a complex genetic disease and one of the leading causes of death worldwide. Identifying driver genes is crucial for deepening our understanding of cancer biology and developing targeted therapies. Mutations in these genes can lead to uncontrolled cell growth, so accurate detection of these genes is of great significance for elucidating the molecular mechanisms of cancer and guiding clinical diagnosis and treatment. The progress of high-throughput sequencing technology has made it possible to accumulate a large amount of genomic data, significantly improving our ability to identify cancer driver genes. Projects such as The Cancer Genome Atlas (TCGA), International Cancer Genome Consortium (ICGC), and Catalogue of Somatic Mutations in Cancer (COSMIC) have contributed genomic, transcriptomic, epigenomic, and proteomic data, which have played an important role in cancer research.

[0003] Using these multi-omics data, many computational methods have been developed to identify cancer driver genes. Traditional methods rely on statistical models to analyze genomic data by examining mutation frequency and distribution patterns to identify driver genes. For example, OncodriveCLUST detects spatial clustering of somatic mutations, while MutSigCV2 identifies significantly mutated cancer genes by assessing the relationship between somatic mutations and background mutation patterns and gene-specific factors. OncodriveFM identifies mutation genes with significant functional impact, while dNdScv reveals selection patterns in cancer genomes by analyzing non-synonymous / synonymous mutation ratios to detect cancer driver genes.

[0004] However, traditional statistical methods have not fully uncovered the complex genetic interactions that lead to cancer phenotypes, which often involve complex gene networks rather than individual genes. This limitation has prompted the development of network-based methods that map genes into biological networks to assess their interactions. These interactions are quantified through traditional machine learning or advanced deep learning techniques. Traditional machine learning focuses on creating hand-designed network-based features, such as centrality measures and dysregulation measures. Centrality measures indicate that the proteins encoded by driver genes are more central in the network, while dysregulation measures consider the impact of driver genes on connected gene expression. However, these methods struggle to uncover complex patterns in biological networks. Deep learning excels at managing these dynamics without relying on handcrafted feature engineering, leading to significant progress in driver gene identification. For example, deepDriver employs a convolutional neural network to analyze both mutation data and gene similarity networks simultaneously. Trans-Driver utilizes a Transformer network to integrate multi-omics data, improving driver gene identification by learning dataset differences and associations. The EMOGI model combines multi-dimensional multi-omics datasets with protein-protein interaction (PPI) networks, using graph convolutional networks (GCNs) to analyze complex relationships and gene features, thereby improving the accuracy of pan-cancer driver gene prediction. Building on this, the MTGCN model introduces a GCN-based multi-task learning framework that enhances gene features by integrating PPI network properties and balancing node and link prediction tasks. Recently, IMVRL-GCN uses GCNs from multiple-view omics data to identify cancer genes within a GNN framework. Similarly, MNGCL constructs different gene network views and uses contrastive learning to learn network-specific features. The HGDC model combines graph diffusion techniques and hierarchical attention mechanisms to improve prediction accuracy.

[0005] While the above methods have been successful in identifying cancer driver genes in cancer-specific or pan-cancer populations, they rarely address both issues simultaneously. Understanding the commonalities and differences of driver genes between different cancer types is crucial for deeper pathological insights and more efficient drug design. The TCGA pan-cancer project provides a large number of datasets from 12 cancer types, offering opportunities for comprehensive cancer research. For example, Zhang et al. developed optimization models ComMDP and SpeMDP to identify common and specific driver gene sets for 12 cancer types from sequencing data. Zhen et al. constructed co-expression networks for 16 cancer types using differentially expressed genes and extracted the largest connected components by integrating these networks to identify cancer-specific signatures. Similarly, Zhu et al. detected potential pan-cancer related genes by combining tissue-specific differential network analysis and constructing a pan-cancer network through network embedding. However, existing methods either separately identify pan-cancer (referred to as common) and cancer-specific driver genes or detect pan-cancer related genes by integrating cancer type-specific data. The former is difficult to capture the correlation of inter-cancer omics data due to different computational mechanisms for driver gene identification, while the latter requires advanced techniques to integrate cancer-specific data to address the heterogeneity of cross-cancer distribution. Moreover, most methods are currently limited to analyzing single omics data, which constitutes a significant limitation when considering the complexity of cancer reflected in multiple biological processes. This highlights the need for advanced methods to integrate multi-omics data from different cancers to more comprehensively discover the molecular landscape of cancer-specific and common driver genes. SUMMARY

[0006] An object of the present application is to propose a specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm to solve the above-mentioned problems existing in the prior art.

[0007] To achieve the above object, the present application provides the following scheme:

[0008] The specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm comprises:

[0009] Based on the data of different cancer types, a dataset is constructed; wherein the data of each cancer type comprises: DNA methylation data, copy number variation data, somatic mutation data;

[0010] Based on the multi-head attention mechanism, the dataset is preprocessed;

[0011] Based on the preprocessed dataset, a neural network model is trained to obtain a gene identification model; wherein the neural network model is constructed based on Chebyshev graph convolution network and graph convolution network, and the model training adopts a federated transfer learning method, and the training process is divided into server side and client side;

[0012] identify cancer-specific and shared driver genes across tumors based on the gene identification model.

[0013] Optionally, training the neural network model comprises:

[0014] randomly allocating data sets to the server side and the client side in advance;

[0015] the server side initializes global model parameters and sends the parameters to the client side to update local model parameters;

[0016] the server side collects and aggregates the updated local model parameters from each client side, and uses the neural network model to train the global model on the allocated preprocessed data set;

[0017] each client side uses the allocated preprocessed data set to download the global model parameters from the server side and uses the neural network model to train the local model;

[0018] the server side and the client side continuously exchange parameters until the model reaches a convergent state.

[0019] Optionally, constructing the data set based on data of different cancer types comprises:

[0020] extracting somatic mutation data, counting the number of various somatic mutations, and creating mutation-related features; wherein the somatic mutation data includes variation classification and variation type;

[0021] calculating the mean and variance of the copy number variation data to generate copy number variation data-related features;

[0022] processing the DNA methylation data of each gene to generate two values for each gene;

[0023] based on the results of processing the DNA methylation data, copy number variation data, and somatic mutation data, and in combination with gene expression values, PPI networks, and DNA replication timing features, a multi-omics data set is finally formed.

[0024] Optionally, preprocessing the data set based on the multi-head attention mechanism comprises:

[0025] performing linear transformation on the data in the data set, dividing pseudo queries, keys, and value matrices;

[0026] after linear transformation, each head independently uses scaled dot-product attention mechanism to calculate attention scores to measure the relationship between features;

[0027] multiply the attention scores by the corresponding value matrix to calculate attention weighting.

[0028] The outputs of all attention heads are spliced and further linearly converted to generate a comprehensive internal correlation feature encapsulating the complex mutual influence and dependence between features in the data set.

[0029] Optionally, each head independently uses a scaled dot-product attention mechanism to calculate the attention score, including:

[0030] Calculating the dot product of the query matrix and the key matrix;

[0031] Dividing the dot product of the query matrix and the key matrix by the square root of the dimension of the key vector to obtain a result score;

[0032] Normalizing the result score using a softmax function to obtain the attention score.

[0033] Optionally, the method of aggregating the updated local model parameters from each client includes:

[0034]

[0035] where ωt represents the global model parameters of the tth round, β is a hyperparameter controlling the update of the global model, t represents the local model parameters of client i in the tth round of training, R d represents a d-dimensional real space, N represents the number of clients, and i represents client i.

[0036] Optionally, training the global model includes:

[0037]

[0038] where ω' represents the global model parameters updated by the neural network model, DA represents the public data set, ChebGC represents the Chebyshev graph convolution model, and ω t+1 t+1 represents the global model parameters of the t+1th round.

[0039] Optionally, training the local model includes:

[0040]

[0041] where, represents the local model parameters of client i in the t+1th round of training, λ represents a regularization parameter, and ChebGC i represents the Chebyshev graph convolution model of client i, and DS i represents the private data of the client. Optionally, the neural network model includes a plurality of layers; each layer thereof performs the following operations: ​​

[0042]

[0043] wherein, H (l-1) denotes the feature matrix of the previous layer, W (l-1) denotes the weight matrix of the l-1th layer, and δ denotes a nonlinear activation function, denotes the kth order Chebyshev polynomial, denotes the learnable parameter of the kth order Chebyshev polynomial corresponding to the i-th layer of the neural network, K denotes the maximum order of the Chebyshev polynomial, and k denotes the current order of the Chebyshev polynomial;

[0044] Residual connections are introduced in each layer of the neural network model;

[0045] The output of the last layer of the neural network model is the ranking score of each gene, and the ranking score of each gene is further processed by a sigmoid activation function.

[0046] Optionally, training the neural network model further comprises;

[0047] A focal loss function is used to adaptively adjust the loss weight of the data set samples.

[0048] The present application has the following beneficial effects:

[0049] The present application integrates multi-omics data using a multi-head attention mechanism, and simultaneously trains global and local models using federated transfer learning combined with ChebNet GCNs, thereby helping to identify tumor-specific and shared driver genes. When the model FTDL-SCdriver of the present application is applied to a TCGA data set including 33 cancer types and 46 multi-omics features, the model of the present application exhibits superior performance compared to 9 similar methods. These findings provide valuable insights for developing targeted and personalized treatment strategies. BRIEF DESCRIPTION OF DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0051] Figure 1 A schematic diagram of the workflow of the FTDL-SCdriver of the embodiments of the present application;

[0052] Figure 2Process diagram for generating cross-omics features through multi-head attention mechanism for embodiments of the present application;

[0053] Figure 3 Process diagram for generating associated biological features through ChebNet GCNs for embodiments of the present application;

[0054] Figure 4 Optimal hyperparameter combination diagram for embodiments of the present application;

[0055] Figure 5 ROC curve and PR curve comparison diagram of FTDL-SCdriver and baseline model for embodiments of the present application;

[0056] Figure 6 Co-occurrence frequency statistical diagram of the top 200 driver genes predicted by the global model in 16 cancers for embodiments of the present application;

[0057] Figure 7 Pathway enrichment analysis heatmap of common driver genes not included in the NCG database for embodiments of the present application;

[0058] Figure 8 AUC and AUPRC comparison diagram of FTDL-SCdriver with and without FTL framework in 16 cancer types for embodiments of the present application;

[0059] Figure 9 Pathway enrichment analysis heatmap diagram of specific driver genes for embodiments of the present application;

[0060] Figure 10 Comparison diagram of FTL, CL and LL on ROC and AUPRC for embodiments of the present application;

[0061] Figure 11 Performance diagram of FTDL-SCdriver under different node numbers for embodiments of the present application;

[0062] Figure 12 Specific driver genes and common driver genes identification method flowchart based on federated transfer learning and deep learning algorithm for embodiments of the present application. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0064] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0065] Federal transfer learning (FTL) provides a promising framework to address these challenges in the background art, which can train server (global) and client (local) models simultaneously to capture cancer-specific and common driver genes. As a combination of federated learning (FL) and transfer learning, FTL leverages decentralized data from multiple domains while facilitating cross-task knowledge transfer. Within this framework, multi-omics data from different cancers are utilized, and the data sets are divided into global sets (multi-cancer) for server-side training and local sets (single cancer type) for client-side training. To model complex genetic interactions, the present embodiment employs ChebNet GCNs, which perform well in handling graph-structured data. Non-independent and identically distributed (Non-IID) heterogeneity across cancer types is alleviated by FedProx optimization, and a multi-head attention mechanism is used to extract associations across multi-omics data. Integrating these components, the present embodiment develops FTDL-SCdriver, a novel method that combines FTL, ChebNet GCNs, and multi-head attention. On TCGA data covering 33 cancer types and 46 features, FTDL-SCdriver outperforms 9 state-of-the-art methods. The identified driver genes are enriched in key pathways and show significant biological relevance, driving the development of targeted therapies and personalized treatment design.

[0066] As shown in Figure 12 The present embodiment proposes a specific driver gene and common driver gene identification method based on federal transfer learning and deep learning algorithm, mainly including the following steps:

[0067] Based on the data of different cancer types, a data set is constructed; wherein the data of each cancer type includes: DNA methylation data, copy number variation data, somatic mutation data;

[0068] Based on the multi-head attention mechanism, the data set is preprocessed;

[0069] Based on the preprocessed data set, a neural network model is trained to obtain a gene identification model; wherein the neural network model is constructed based on ChebNet GCNs and graph convolutional networks, and the model training adopts a federal transfer learning method, and the training process is divided into server side and client side;

[0070] Based on the gene identification model, cancer-specific and common driver genes across tumors are identified.

[0071] Specifically, in the present embodiment, a framework of federated transfer learning, FTDL-SCdriver, is adopted in the model training process, a neural network model (ChebNet GCNs) constructed based on Chebyshev graph convolutional network and graph convolutional network and a multi-head attention mechanism are used to identify cancer-specific and shared driver genes across tumors, and FedProx is used to solve the problem of data heterogeneity. The training process is divided into server side (global model trained on 17 cancer data) and client side (local model trained on 16 cancer data). Figure 1 The workflow of the FTDL-SCdriver framework is shown, which mainly includes two steps: data preprocessing and federated model training. Among them, the global model and the local model are both ChebNet GCNs. The global model is a shared ChebNet graph convolutional network trained on a mixed dataset of 17 cancers; at the same time, a special local model is constructed for each cancer type, and these models use independent data of the corresponding cancer type to train ChebNet GCN.

[0072] In the data preprocessing stage, data from different cancers is distributed and integrated to form the original feature set of each gene sample (hereinafter referred to as sample). Subsequently, a multi-head attention mechanism is applied to capture features across different omics data types and their internal correlations, thereby generating comprehensive cross-omics features. These features are combined with the original features to form the integrated features of each sample. Preprocessing is carried out uniformly on the client and server sides. In the model training stage, the server aggregates parameters from various local models and trains the global model on 17 cancer data using ChebNet GCNs. Each client downloads the global model parameters from the server and trains and optimizes the local model using ChebNet GCNs. This parameter exchange continues until the model reaches a state of convergence, indicating that the model has reached an optimal learning state.

[0073] Further, based on the data of different cancer types, the dataset is constructed to include:

[0074] The somatic mutation data is extracted, the number of various somatic mutations is counted, and mutation-related features are created; wherein the somatic mutation data includes: variation classification and variation type;

[0075] The mean and variance of the copy number variation data are calculated to generate features related to the copy number variation data;

[0076] The DNA methylation data of each gene is processed to generate two values for each gene: methylation mean and variance;

[0077] Based on the results of processing DNA methylation data, copy number variation data, and somatic mutation data, and combined with gene expression values, PPI network, and DNA replication timing features, a multi-omics dataset is finally formed.

[0078] Specifically, the omics dataset in this embodiment is derived from the data used by Trans_driver, while the protein-protein interaction (PPI) network is from the STRING database. The omics data is derived from the TCGA pan-cancer atlas, covering 33 cancer types, including DNA methylation, copy number variation (CNV), somatic mutation data, etc. Specifically, the somatic mutation data is extracted from TCGA, focusing on mutation classification (such as frameshift mutation, in-frame mutation, missense mutation, nonsense mutation, etc.) and mutation type (including deletion, insertion, and single nucleotide polymorphism, etc.). The number of various somatic mutations is counted, and 28 mutation-related features are created. CNV data generates 6 CNV-related features by calculating the mean and variance. Similarly, DNA methylation data for each gene is processed to generate two values for each gene. In addition, gene expression values, PPI networks, and DNA replication timing features are also integrated, finally forming a 46-dimensional multi-omics dataset. In the FTDL-SCdriver framework, 17 cancer types (Table 1) are assigned to the server side for training global models, and the remaining 16 cancer types (Table 2) are assigned to the client nodes for training local models. Tables 1-3 provide detailed information on the 33 cancer types. The PPI network is filtered to only retain interactions with a score higher than 0.5, resulting in a network containing 14,026 nodes and 557,950 edges.

[0079] Table 1: Statistics of node number, edge number, and positive and negative sample number for each cancer type on the server side.

[0080]

[0081]

[0082] Table 2: Statistics of node number, edge number, and positive and negative sample number for each cancer type on the client side.

[0083]

[0084] Table 3: Detailed information of 46 omics features.

[0085]

[0086]

[0087] For sample labels, the embodiment adopts a similar method as MTGCN. Client-side cancer type-specific positive samples come from NCG 6.0 and are labeled with corresponding cancer types and a pan-cancer set. Server-side positive samples include the pan-cancer types of NCG 6.0 and high-confidence genes in the literature. Negative samples are screened by excluding genes annotated in NCG, COSMIC, OMIM, and KEGG cancer pathways. Detailed information is listed in Tables 1 and 2.

[0088] Further, based on the multi-head attention mechanism, the preprocessing of the data set comprises:

[0089] Linear transformation is performed on the data in the data set, and pseudo query, key and value matrices are divided;

[0090] After linear transformation, each head independently uses a scaled dot-product attention mechanism to calculate an attention score to measure the relationship between features;

[0091] The attention score is multiplied by the corresponding value matrix to calculate the attention weight;

[0092] The outputs of all attention heads are spliced and further linearly converted to generate a comprehensive internal association feature that encapsulates the complex mutual influence and dependence relationships between features in the data set.

[0093] Specifically, the multi-head attention mechanism is used in the embodiment to extract gene cross-omics features, which comprises:

[0094] Gene features exhibit high complexity and diversity, influenced by multiple interacting factors. These factors include the regulation of gene expression by promoter methylation, variations in gene mutations across different cancer types, and the complex regulation of gene expression by gene interaction networks. To comprehensively elucidate these complex genomic relationships from multiple perspectives, the multi-head attention mechanism is adopted in the study. This mechanism is a key component of contemporary large models, particularly in the fields of natural language processing and deep learning. It calculates through multiple attention heads in parallel, with each head focusing on different molecular feature pairs, thereby capturing a wider range of interactions and dependencies.

[0095] As shown in Figure 2 , the feature vector of each sample is defined as X0={x1,x2,x3···,x n}, serving as input data for subsequent modules. In the multi-head attention mechanism, the input data X0is divided into pseudo query (Q i ), key (K i ) and value (V i ) matrices through linear transformation, and the calculation process of each attention head i is as follows:

[0096]

[0097] wherein, and are weight matrices, and are bias vectors, responsible for transforming the input into Q i ,K i , and V i , respectively.

[0098] After the linear transformation, each head independently computes attention scores using scaled dot-product attention mechanism to measure the relationship between features. The attention score of each head is computed by first calculating the dot product of the query matrix and the key matrix, and then dividing the dot product by the square root of the dimension of the key vector Finally, the softmax function is applied to normalize the scores. This process is crucial to prevent the problem of gradient vanishing and to ensure that the sum of attention weights for each query is 1. The formula for calculating the attention score is:

[0099]

[0100] Then, the attention weighted weights are calculated by multiplying the attention scores with the corresponding value matrix, denoted as:

[0101] Head i = Attention Score i · V i #(3)

[0102] Finally, the outputs of all attention heads are concatenated and further linearly transformed to generate a comprehensive internal association feature that encapsulates the complex mutual influence and dependence between features in the dataset. The final output of the multi-head attention mechanism is:

[0103] MulitHead(Q,K,V) = Concat(Head1, Head2…Head n )W O #(4)

[0104] where W O is the weight matrix of the final linear transformation, and Concat(·) represents the concatenation operation along the embedding dimension. This process significantly enhances the detection ability of complex genomic relationships and helps to extract internal biological features from omics data, which is crucial for driving gene differentiation.

[0105] Further, the neural network model comprises a plurality of layers;

[0106] Residual connections are introduced in each layer of the neural network model;

[0107] The output of the last layer of the neural network model is the ranking score of each gene, which is also processed by a sigmoid activation function.

[0108] Specifically, the present embodiment uses ChebNet GCNs to analyze the associated biological features in the PPI network:

[0109] Graph Convolutional Networks (GCNs) are a type of deep learning model specifically designed to handle graph-structured data. They learn node representations by propagating information between nodes, making them suitable for tasks such as node classification, graph classification, and link prediction. ChebNet GCNs extend the receptive field to include K-order neighbors by using K-order Chebyshev polynomials to approximate the graph convolution, thereby improving the flexibility of the model. In this study, the present embodiment employs ChebNet GCNs to analyze the intrinsic features of genes and their relationships in the PPI network. The implementation process of the present embodiment is shown in Figure 3 First, the present embodiment combines the original biological features with newly generated cross-omics features to create a comprehensive biological matrix. This matrix is represented as where N is the number of genes, and F is the dimension of gene features. The PPI network G is represented by its adjacency matrix A and degree matrix D. The Laplacian matrix L of the graph is defined as L = D - A. To normalize the Laplacian matrix so that its eigenvalues are in the range [-1, 1], the present embodiment uses the normalized Laplacian matrix ChebNet GCNs use Chebyshev polynomials T k to approximate the K-order polynomial expansion of , which is given by the following formula:

[0110]

[0111] where θ k is a learnable parameter, and T k is the k-th order Chebyshev polynomial.

[0112] For each layer, the ChebNet GCN performs the following operations:

[0113]

[0114] where H (l-1) is the feature matrix of the previous layer (when l = 0, it is the original input feature matrix X), W (l-1) is the weight matrix of the l-1 layer, used for linear transformation, and δ is the nonlinear activation function ReLU.

[0115] To alleviate the vanishing gradient problem and enhance the learning ability of the model, the embodiment introduces a residual connection in each layer of ChebNet. The function definition of the residual connection is as follows:

[0116] H (l+1) =H (l) +δ(W (l) X+b (l) )#(7)

[0117] Where H (l) is the output of the l-th layer, W (l) is the weight matrix of the fully connected layer, and b (l) is the bias vector.

[0118] After three layers of Chebyshev graph convolution layers, the ranking scores of each gene are obtained. These scores are then processed by a sigmoid activation function to ensure that the output values are limited to the range of 0 to 1. The closer the possibility score p i is to 1, the higher the likelihood that the gene is a driver gene.

[0119] In this embodiment, due to the imbalance between positive and negative samples and the existence of many difficult-to-classify samples, the focal loss function is used. The focal loss function adaptively adjusts the loss weight of the samples, making the model pay more attention to difficult-to-classify samples. The definition of the focal loss function is as follows:

[0120]

[0121] Where γ is a focusing parameter that controls the importance of difficult-to-classify samples, and α is a weight factor that balances the weights of positive and negative samples. The weight of positive samples is α, and the weight of negative samples is 1-α. After multiple iterations, the model can learn the characteristics of the gene itself and its connection relationship in the PPI network.

[0122] Further, training the neural network model comprises:

[0123] Randomly allocating data sets to the server side and the client side in advance;

[0124] The server side initializes the global model parameters and sends the parameters to the client side to update the local model parameters;

[0125] The server side collects and aggregates the updated local model parameters from each client, and uses the neural network model to train the global model on the allocated preprocessed data set;

[0126] Each client uses the allocated preprocessed data set to download the global model parameters from the server side and trains the local model using the neural network model.

[0127] The server and the clients continuously exchange parameters until the model reaches a state of convergence.

[0128] Specifically, in this embodiment, training local and global models in the FTL framework includes:

[0129] To identify cancer-specific and common driver genes in different cancers, this embodiment employs the FTL framework and combines the aforementioned ChebNet GCNs while developing both local and global models. This approach leverages the collective power of local models to build a comprehensive global model, while the global model in turn optimizes the capabilities of the local models. To address the issue of data heterogeneity in federated learning, this embodiment employs the FedProx algorithm, which mitigates the negative effects of statistical heterogeneity between different cancers by introducing a proximal term.

[0130] To this end, this embodiment divides the system into two main parts: the server side and the client side. This embodiment sets up N client nodes (E i ) and takes a trusted server (S) as the core. The server S is responsible for aggregating model parameters from clients and updating the global model using a comprehensive biological dataset (DA) that includes preprocessed comprehensive data for 17 cancers. The process starts with the server initializing global model parameters ω0, which are then sent to the clients. After collecting the updated local model parameters φ i from each client, the server aggregates these parameters using the following equation (9):

[0131]

[0132] where ω t denotes the global model parameters in the t-th round, β is a hyperparameter that controls the update of the global model, denotes the local model parameters of client i in the t-th round of training.

[0133] The server further trains the global model using the aggregated global parameters and updates the model parameters ω′ t+1 using ChebNet GCNs. Then, it sends the updated parameters ω′ t+1 to each client. The training process of the global model is described by the following function:

[0134]

[0135] Each client (E i ) provides its unique comprehensive dataset DS nContribute to the federated transfer learning process for specific cancer types. Use ChebNet GCNs and an additional regularization term for FedProx, train an independent local model for each client to mitigate the impact of uneven data distribution. The process starts with receiving global model parameters, then trains the local model in the following way:

[0136]

[0137] where λ is the regularization parameter, used to control the influence of global parameters ω′ t on the local model . After the local model training is completed, the local parameter updates are transmitted to the central server for aggregation.

[0138] By introducing a regularization term, the FedProx algorithm encourages each client's model to learn from local data while maintaining consistency with the global model. This two-level optimization process identifies the global model by aggregating local models and enhances local models by leveraging the knowledge embedded in the global model. This approach to some extent weakens the impact of statistical diversity among participants and improves the generalization ability of the model.

[0139] Algorithm 1 provides a pseudo-code summary of the FTDL-SCdriver.

[0140]

[0141]

[0142] To evaluate the FTDL-SCdriver, this embodiment compares the method of this embodiment with nine other methods, and uses 5-fold cross-validation to minimize variability. These methods are chosen because they share components or datasets with the model of this embodiment. Although existing methods can identify cancer or specific cancer type drivers, they often cannot capture these genes simultaneously and ignore the diversity distribution between different cancer types. To compare fairly, all methods are evaluated using the same input interaction network and feature extraction indicators, including 46-dimensional original biological features and 18-dimensional cross-omics features. In addition, all baseline convolutional neural networks (CNNs) are standardized to have the same number of layers and use their respective default parameters:

[0143] Trans-Driver: Use Transformer network with deep supervised learning to integrate multi-omics data to discover cancer driver genes, focusing on differences and associations between omics types;

[0144] EMOGI: A deep learning method based on graph convolutional networks that integrates multi-omic pan-cancer data and protein-protein interaction (PPI) networks to predict cancer genes;

[0145] MTGCN: A multi-task learning framework that enhances gene features by combining biological and network structure characteristics, and uses a Bayesian task weight learner to balance tasks in graph convolutional networks (GCN);

[0146] HGDC: A heterogeneous graph diffusion convolutional network method that integrates omic features and biological molecular networks, using a diffusion-generated auxiliary network to identify cancer driver genes;

[0147] IMVRL-GCN: A multi-view representation learning method that identifies cancer driver genes by integrating shared and specific cancer representations in multi-omic data, combining generative adversarial networks (GAN) and graph neural networks (GNN);

[0148] InDEP: An interpretable machine learning framework that uses cascading decision trees and Kernel SHAP modules for post-hoc interpretation to identify cancer driver genes;

[0149] GCN: A neural network model for processing graph data that aggregates feature information of nodes and their neighbors through graph convolution operations;

[0150] ChebNet: A variant of GCN that uses Chebyshev polynomials to efficiently capture graph structure information by approximating spectral filtering of the Laplacian matrix;

[0151] GAT: A graph attention network for processing graph data that introduces an attention mechanism to dynamically assign weights to nodes and their neighbors.

[0152] The parameters of this embodiment are set as follows:

[0153] This embodiment model is built in Python 3.7 environment, based on PyTorch 1.12.1 and PyTorchGeometric 2.3.1 implementation. The optimizer uses Adam, and Optuna is used for hyperparameter optimization. To ensure the robustness of the results, this embodiment performs 5-fold cross-validation in FTDL-SCdriver. Based on the AUC value after 500 rounds of learning, the optimal hyperparameters are selected, and the optimal median AUC value is shown in the block of Figure 4 .

[0154] This example compares the FTDL-SCdriver model with 9 baseline models to evaluate its performance in the cancer driver gene prediction task. The performance evaluation uses ROC curve and PR curve as indicators to better deal with the label imbalance problem. To ensure the fairness of the comparison, this example uses the fully trained global model of FTDL-SCdriver to predict the combined data of 16 different cancer datasets. Figure 5 It is shown that FTDL-SCdriver outperforms all baseline models in both indicators, especially in the PR curve, highlighting its excellent ability to accurately identify a small number of positive samples. This advantage may come from the ability of FTDL-SCdriver to learn both common and specific driver gene knowledge in multiple cancers, while other methods can only learn common knowledge. Among the baseline models, MTGCN ranks second, which may be because it effectively combines biological information and network structure information, thereby enhancing gene feature expression. In contrast, EMOGI and InDEP rank last two, both of which simply concatenate multiple omics data, lacking complex information integration, which shows that introducing more biological features into GCN can significantly improve prediction performance. While models such as HGDC, Trans-Driver, and IMVVRL-GCN rank in the middle, indicating that simply relying on network structure, multi-omics association information, or using multiple omics data independently is not enough to significantly improve prediction performance. In the analysis of the basic training components in this example, ChebNet GCN performs slightly better than GAT and significantly better than traditional GCN. This advantage may come from the fact that ChebNet GCN uses Chebyshev polynomials to approximate spectral graph convolution, allowing it to capture more complex patterns and more diverse graph structure information than other algorithms.

[0155] Ablation experiments:

[0156] The FTDL-SCdriver model combines federated transfer learning and the ChebNet GCN framework, and integrates original features and omics features through a multi-head attention mechanism to train local and global models to identify cancer driver genes. Model training includes four main components: (1) omics feature extraction of genes, (2) splicing of original features and omics features, (3) learning of associated biological features based on ChebNet GCN, and (4) training of the model using the FedProx algorithm under the federated learning framework. This example modifies the model components in three ways: (1) testing with and without federated transfer learning (No_FTL), (2) comparing the effects of FedAvg and FedProx aggregation methods, and (3) comparing the performance of ChebNet and traditional GCN in learning gene associations.

[0157] Table 4 shows the average AUC and AUPRC values of the global model on 16 cancer datasets, and compares the results with and without federated transfer learning under different input data and component combinations. The results show that the “FTL+FedProx+ChebGCN” combined model performs best among all combinations, highlighting the importance of each component. The highest AUC of the non-federated transfer learning (Non-FTL) model is only 0.892, while the FTL L-SCdriver reaches 0.953, fully demonstrating the advantages of federated collaborative training. The FedProx model is superior to the FedAvg model, possibly because FedProx performs better in handling Non-iid data problems. ChebNet GCN also outperforms traditional GCN in gene correlation learning. In addition, the dataset with aggregated features as input performs best, which shows that combining original multi-omics gene features with inter-omics features can significantly improve the identification of cancer driver genes. The results also show that the model trained on a single cancer dataset cannot learn common knowledge from other cancer datasets, resulting in a significant decrease in model performance, with an average decrease in AUC of 0.0416.

[0158] Table 4: Ablation experiments

[0159]

[0160]

[0161] Performance of FTL L-SCdriver in identifying common cancer driver genes:

[0162] Some driver genes are referred to as common driver genes, which often show abnormal changes in multiple cancer types. In this study, the global model trained in this embodiment uses multi-omics features and gene-gene correlation features of 33 types of cancer to predict these driver genes. This embodiment assumes that genes predicted as driver genes in multiple cancers by the global model are more likely to be common driver genes. Considering the potential incompleteness and false positives of the NCG (Network of Cancer Genes) database, this embodiment identifies the top 200 predicted driver genes based on each client dataset in the global model, and evaluates their overlap with the driver genes included in the NCG database. Figure 6Among the top 200 predicted driver genes, 92% of the genes were identified as driver genes in more than two cancers, only 2% and 6% of the genes were identified in two or one cancer, respectively. This indicates that the global model has strong consistency and can stably identify most of the driver genes in multiple cancers. In addition, most of the identified genes have been included in the NCG database, and only a small part has not been included in the NCG. This not only verifies the reliability of the NCG database, but also shows that the global model supplements a part of the potential driver genes not included in the NCG, reflecting the limitations of the NCG database in completeness and accuracy. In summary, the high consistency of the prediction results in multiple cancers helps the embodiment to obtain a more reliable set of candidate driver genes. Future research should focus on the functional verification of these genes, especially the genes not included in the NCG, to further explore their potential role in cancer occurrence and progression.

[0163] To explore the function of 43 common candidate driver genes not included in the NCG database, the embodiment uses the DAVID online tool for pathway enrichment analysis, and selects the top 20 enriched pathways according to the p value, as shown in Figure 7 Among them, GSK3B is a key regulatory factor in the "cancer pathway" and "PI3K-Akt signaling pathway", and is significantly related to the tumor progression and prognosis of part of ovarian cancer and pancreatic cancer, highlighting its core role in cancer occurrence and its value as a potential targeted therapy target. MAPK8, as a cancer driver gene, is enriched in the IL-17 signaling pathway, significantly affecting the tumor progression, immune escape, metastasis and invasion of multiple cancers such as breast cancer, non-small cell lung cancer and gastric cancer, and has important therapeutic potential. ROCK1, as a key gene for pan-cancer progression, regulates cytoskeletal remodeling and tumor microenvironment remodeling, and is enriched in cancer-related pathways such as "proteoglycans in cancer" and "local adhesion", further emphasizing its multi-level key role in tumor occurrence and development. In addition, HSP90AA1 regulates the classic PI3K-Akt signaling pathway and the androgen receptor signaling pathway, and plays a key role in the occurrence and progression of multiple cancers, highlighting its important value as a potential therapeutic target.

[0164] Performance of FTDL-SCdriver in identifying cancer-specific driver genes:

[0165] Studies have shown that "common driver genes" are genes shared by multiple cancer types, while "cancer-specific driver genes" are genes unique to a specific cancer. The FTDL-SCdriver model identifies common driver genes through a global model and identifies cancer-specific driver genes in combination with a local model. Such genes, due to their low frequency of occurrence in the overall cancer data, are often overlooked by pan-cancer analysis models based on the entire population. To verify the effectiveness of FTDL-SCdriver in identifying cancer-specific driver genes, the present embodiment first evaluates the AUC and AUPRC of each cancer type under the framework of federal transfer learning (FTL) (see Figure 8 ). Subsequently, the present embodiment systematically analyzes the top 200 genes predicted by the local model and excludes those that have been identified as common driver genes, thereby screening for cancer-specific driver genes. Finally, the present embodiment performs pathway enrichment analysis on these genes to reveal their potential role in the process of carcinogenesis.

[0166] As shown in Figure 8 , FTDL-SCdriver is significantly superior to the control method without FTL in identifying cancer driver genes, with an average AUC and AUPRC of 0.9697 and 0.9734, respectively, an increase of about 0.0659 and 0.0836 over the non-FTL model. Even for cancer types such as ESCA and SARC, which have fewer positive samples or extremely imbalanced label distribution, the performance of FTDL-SCdriver far exceeds that of the non-FTL model. This fully demonstrates that FTDL-SCdriver, through the federal transfer learning framework, makes full use of the shared information of multi-cancer data to enhance the analysis capability for a single cancer, reflecting the rationality and effectiveness of the model.

[0167] The detailed information of the driver genes predicted for each cancer type is shown in Table 5. Figure 9 The top 20 pathways in four selected cancer types (GBM, LUAD, OV, PRAD) are shown, sorted by P-value. The enrichment of other cancers is shown in the Figures 2-4The significant presence of cancer-specific driver genes in these pathways highlights the effectiveness of FTDL-SCdriver. For example, in cancer GBM, the neuroactive ligand-receptor interaction pathway was enriched, consistent with its neurogenic origin, and genes like PDGFRB and FLT3 are closely related to the occurrence of glioblastoma. In cancer LUAD, the MAPK and PI3K-Akt pathways were significantly enriched, highlighting their role in pathogenesis, and CCND1 and ARAF genes can be potential therapeutic targets. Ovarian cancer-specific genes, ATF2 and FGFR2, were enriched in the ErbB and NF-κB pathways, affecting tumor behavior and drug resistance. Prostate cancer-specific genes, including LYN and CDKN1A, were enriched in the PI3K / Akt, MAPK, and NF-κB pathways, which regulate key processes in tumor occurrence and progression.

[0168] Table 5: Detailed distribution of prediction results for each cancer type

[0169]

[0170] Robustness and stability test under different data distribution:

[0171] To evaluate the robustness and stability of the FTDL-SCdriver model under different data distributions, this embodiment refers to the comparative study of Olivia et al., which evaluates federated transfer learning (FTL) based on training accuracy, centralized learning (CL), and local learning (LL). In CL, this embodiment trains a ChebNet GCN model using a merged dataset from 16 cancer types. LL involves each client training a local model only on a single cancer data without data or parameter exchange. FTL, similar to FTDL-SCdriver, distributes the dataset to multiple clients and a central server and shares training parameters to optimize local and global models. Figure 10 The ROC and AUPRC value distributions under these methods are shown, which come from five-fold cross-validation of 10 iterations. This embodiment calculates the average performance of 16 cancers in each iteration of LL and FTL. The results show that FTL outperforms other methods in AUPRC and AUC indicators, highlighting its effectiveness in efficiently predicting and addressing non-iid data distribution using distributed data. The reason why FTL is better than CL is that it optimizes the local model by considering the diversity of data distribution through the knowledge of each participant. In contrast, CL aggregates the dataset without these considerations and lacks a global model covering 17 cancer types. Since LL only trains on a single cancer type, its sample size is limited, resulting in poor robustness in training and ranking last.

[0172] Regarding stability, CL shows the highest stability as it is trained on the combined dataset of 16 cancer types. FTL shows the second highest stability as it optimizes local models through transferred parameters while still relying on local datasets, leading to some fluctuations. FedProx in FTL effectively controls these fluctuations, keeping them within a narrow range: specifically, the fluctuations in AUC are 0.017, and the fluctuations in AUPRC are 0.01. In contrast, LL shows significant fluctuations as models trained on a single cancer type are applied to predict other 15 cancer types, leading to large performance fluctuations.

[0173] Scalability and robustness analysis under different numbers of clients:

[0174] To demonstrate the scalability and robustness of the FTDL-SCdriver model under different numbers of clients, this embodiment evaluates its AUPRC and ROC index performance under 1 to 33 clients. Figure 11 The AUPRC and ROC distributions obtained by the model in five-fold cross-validation over 10 iterations are shown. As the number of clients increases, the predictive performance of the model improves significantly, especially before eight clients. This indicates that FTDL-SCdriver can effectively leverage diverse information from different participants for joint training of driver gene prediction models, verifying the effectiveness and scalability of the method. Notably, the model remains highly stable and reliable with minimal fluctuations after more than eight clients. However, a slight decline occurs at 16 clients, which may be due to the small number of increased node samples. However, the model can still handle these nodes well, with a decline in ROC and AUPRC of only 0.0043 and 0.003. After that, the model remains stable. This indicates that a reasonable number of clients is sufficient for training, and excessive clients are unnecessary. The stability after eight clients further verifies the ability of FTDL-SCdriver to coordinate a large number of clients for federated model training.

[0175] Exploring the commonalities and differences of driver genes across different cancer types is crucial for advancing personalized medicine and targeted cancer treatment. In this study, this embodiment presents FTDL-SCdriver, a new method for identifying cancer-specific driver genes and common driver genes. The method integrates multi-omics data through a multi-head attention mechanism and applies federated transfer learning within the ChebNetGCN framework, effectively training global and local models using the FedProx algorithm to address different data distributions. According to the understanding of this embodiment, FTDL-SCdriver is the first framework to combine federated transfer learning with the identification of common and specific cancer driver genes across different cancer types.

[0176] Extensive testing and experiments were conducted to validate the effectiveness of the present embodiment. The results show that FTDL-SCdriver outperforms 9 state-of-the-art methods in predicting known cancer driver genes with significantly higher AUC and AUPR scores. The ablation experiments highlight the necessity and effectiveness of each component and data composition, with the aggregation method (Fedprox) having the most impact, followed by the FTL framework and the ChebGCN model. The concatenated data that integrates the raw data and learned cross-omic features performs best, as it comprehensively represents the biological processes. The global model consistently predicts a large fraction of driver genes in multiple cancer types, validating and complementing the NCG benchmark dataset. Meanwhile, the local model shows higher AUC and AUPRC in identifying cancer-specific genes and reveals pathways related to cancer development, such as the neuroactive ligand-receptor interaction pathway in GBM. The robustness, stability, and scalability of FTDL-SCdriver are crucial for its application under various data distributions and numbers of participants. It effectively utilizes distributed data to make efficient predictions under non-iid distributions, achieving the highest AUPRC and AUC values with minimal fluctuations. As the number of participants increases, the model significantly improves, especially before eight participants, after which it remains highly stable and reliable.

[0177] The above-described embodiments are merely intended to describe the preferred modes of the present application, and are not intended to limit the scope of the present application. Various modifications and improvements of the present application made by those skilled in the art without departing from the design spirit of the present application shall fall within the scope of the present application as defined by the appended claims.

Claims

1. A method for specific driver gene and common driver gene identification based on federated transfer learning and deep learning algorithm, characterized in that, The method comprises the following steps: Based on the data of different cancer types, a data set is constructed; wherein the data of each cancer type comprises: DNA methylation data, copy number variation data, somatic mutation data; Based on the multi-head attention mechanism, the data set is preprocessed; Based on the preprocessed data set, a neural network model is trained to obtain a gene recognition model; wherein the neural network model is constructed based on Chebyshev graph convolution network and graph convolution network, and the model training adopts federated transfer learning method, and the training process is divided into server side and client side; Based on the gene recognition model, cross-tumor cancer-specific and common driver genes are identified; Training the neural network model comprises: The server side initializes the global model parameters and sends the parameters to the client side to update the local model parameters; The server side collects and aggregates the updated local model parameters from each client, and trains the global model on the assigned preprocessed data set using the neural network model; Each client downloads the global model parameters from the server side using the assigned preprocessed data set, and trains the local model using the neural network model; The server side and the client side continuously exchange parameters until the model reaches a convergent state; The method of aggregating the updated local model parameters from each client comprises: Training the global model comprises: wherein, denotes the global model parameters at round t, β is a hyperparameter that controls the global model update, denotes the local model parameters of client i at round t, denotes the d-dimensional real space, denotes the number of clients, i denotes a client; Training the local model comprises: wherein, denote the updated global model parameters with the neural network model, denote the public dataset, denote the Chebyshev graph convolution model, denote the global model parameters of the t+1th round; The neural network model comprises several layers; each layer thereof performs the following operations: wherein, denotes the local model parameters of client i for the t+1th round of training, denotes a regularization parameter, denotes the Chebyshev graph convolutional model of client i, denotes the client private data; Residual connection is introduced in each layer of the neural network model; wherein, represents a feature matrix of a previous layer, represents a weight matrix of the layer, represents a non-linear activation function, represents a Chebyshev polynomial of order k, represents a learnable parameter of the Chebyshev polynomial of order k corresponding to the i-th layer of the neural network, represents a maximum order of the polynomial, represents a current order of the Chebyshev polynomial; The output of the last layer of the neural network model is the ranking score of each gene, and the ranking score of each gene is further processed by a sigmoid activation function. Based on the data of different cancer types, a data set is constructed, which comprises: 2.The specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm according to claim 1, wherein, Extract somatic mutation data, count the number of various somatic mutations, and create mutation-related features; wherein the somatic mutation data comprises: variation classification and variation type; Calculate the mean and variance of the copy number variation data to generate copy number variation data-related features; Process the DNA methylation data of each gene to generate two values for each gene; Based on the results of processing the DNA methylation data, copy number variation data and somatic mutation data, and combined with gene expression values, PPI network and DNA replication timing features, a multi-omics data set is finally formed. Based on the multi-head attention mechanism, the data set is preprocessed, which comprises: 3.The specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm of claim 1, wherein, Linear transformation is performed on the data in the data set to divide the pseudo-query, key and value matrices; After linear transformation, each head independently uses the scaled dot-product attention mechanism to calculate the attention score to measure the relationship between features; Multiply the attention score by the corresponding value matrix to calculate the attention weight; The outputs of all attention heads are spliced and further linearly converted to generate a comprehensive internal correlation feature, which encapsulates the complex mutual influence and dependence relationship between features in the data set. Each head independently uses the scaled dot-product attention mechanism to calculate the attention score, which comprises:

4. The method of identifying specific driver genes and shared driver genes based on federated transfer learning and deep learning algorithm according to claim 3, wherein, ​ Dot product of the query matrix and the key matrix is calculated; Divide the dot product of the query matrix and the key matrix by the square root of the dimension of the key vector to obtain a result score; The result score is normalized by using a softmax function to obtain an attention score. 5.The specific driver gene and common driver gene identification method based on federated transfer learning and deep learning algorithm of claim 1, wherein, Training the neural network model also includes: Adaptively adjusting the loss weight of the data set sample by using a focal loss function.

Citation Information

Patent Citations

  • Federal learning method and system for multi-source heterogeneous medical data

    CN117592555A

  • Cancer driver gene identification method based on heterograph diffusion convolutional network

    CN118280453A