Cancer driver gene prediction method and system based on graph neural network

Through the graph neural network-based method, a heterogeneous network introduced by multiomics data is constructed, which solves the problem of prediction accuracy deviation of cancer-driven genes in the prior art, achieves higher prediction accuracy and stability, and improves the interpretability and scalability of the model.

CN119993269APending Publication Date: 2025-05-13SHANDONG NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510086675.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art uses the accuracy bias in cancer-driven gene prediction, lacking a highly accurate and explainable solution.

Method used

A graph neural network-based method is adopted to obtain multiple protein interaction data, a gene-gene interaction network is constructed, and multiomics data is introduced into a heterogeneous network. Information transmission and feature updates are used for shared graph neural network and meta graph convolutional network, and predictions are finally made through multi-layer perception machines.

Benefits of technology

It significantly improves the accuracy of cancer-driven gene prediction, reduces the impact of data bias in a single network, improves the stability and reliability of prediction results, and enhances the interpretability and scalability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993269A_ABST
    Figure CN119993269A_ABST
Patent Text Reader

Abstract

The invention discloses a cancer driver gene prediction method and system based on a graph neural network, and the method comprises the steps: obtaining interaction data of M proteins; for each kind of protein interaction data, coding a relationship between proteins into a gene-gene interaction relationship, and finally obtaining M gene-gene interaction networks; configuring new nodes and new edges for the M genes and gene interaction networks respectively to obtain M heterogeneous networks; the new nodes are cancer-related multi-omics data, and the new edges are connection edges between the new nodes and original gene nodes; and inputting each heterogeneous network into a trained cancer driving gene prediction model to obtain a prediction result of whether each gene is a cancer driving gene or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics gene prediction technology, and in particular to a method and system for predicting cancer driver genes based on graph neural networks. Background Art

[0002] Cancer driver gene prediction is a very important task in bioinformatics. Mutations that occur in somatic cells and confer preferential growth advantages to tumor cells are called driver mutations, and the genes responsible for driving these mutations are called driver genes.

[0003] Finding cancer driver genes through traditional medical methods requires a lot of time and research costs. With the development of deep learning technology, compared with traditional network analysis, graph neural networks can obtain a variable number of neighbor node feature values ​​and local topological structures composed of neighbor nodes that are important for classification results, and train in protein interactions (PPI) in a semi-supervised manner to learn complex nonlinear structures to identify cancer driver genes, thereby better performing the classification task of gene nodes.

[0004] Existing studies based on protein interactions (PPIs) have constructed a heterogeneous network containing multi-omics data, nodes of multiple biological entity types, and multiple relationships, and applied advanced heterogeneous network characterization algorithms to the characterization mining of cancer-related genes. The main task is to classify and predict cancer driver genes in protein interactions (PPIs).

[0005] However, we found that there was a deviation in the accuracy of cancer driver gene prediction in models trained on different protein interactions (PPIs), indicating that the predicted cancer driver genes are different when different protein interactions (PPIs) are used. In summary, the existing technology for cancer driver gene prediction problems still lacks a highly accurate and interpretable solution. Summary of the invention

[0006] In order to solve the deficiencies of the prior art, the present invention provides a method and system for predicting cancer driver genes based on graph neural networks;

[0007] On the one hand, a method for predicting cancer driver genes based on graph neural networks is provided, including:

[0008] Obtain M kinds of protein interaction data; for each kind of protein interaction data, encode the relationship between proteins into a gene-gene interaction relationship, and finally obtain M gene-gene interaction networks;

[0009] For M gene-gene interaction networks, new nodes and new edges are respectively configured to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes;

[0010] Input each heterogeneous network into the trained cancer driver gene prediction model to obtain the prediction result of whether each gene is a cancer driver gene;

[0011] Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, and then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

[0012] On the other hand, a cancer driver gene prediction system based on graph neural network is provided, including:

[0013] An acquisition module is configured to: acquire M kinds of protein interaction data; for each kind of protein interaction data, encode the relationship between proteins into a gene-gene interaction relationship, and finally obtain M gene-gene interaction networks;

[0014] A configuration module is configured to: configure new nodes and new edges for M gene-gene interaction networks respectively, to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes;

[0015] A prediction module is configured to: input each heterogeneous network into the trained cancer driver gene prediction model to obtain a prediction result of whether each gene is a cancer driver gene;

[0016] Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, and then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

[0017] On the other hand, there is also provided an electronic device, comprising:

[0018] a memory for non-transitory storage of computer-readable instructions; and

[0019] a processor for executing the computer readable instructions,

[0020] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.

[0021] On the other hand, a storage medium is provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.

[0022] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is used to implement the method described in the first aspect when running on one or more processors.

[0023] The above technical solution has the following advantages or beneficial effects:

[0024] 1. In terms of prediction effect, this paper proposes for the first time a cancer driver gene prediction model based on a multi-input graph neural network. By integrating multi-omics data and protein interaction networks and using a heterogeneous graph neural network model, the accuracy of cancer driver gene prediction is significantly improved compared with the method based on a single-input graph neural network. In addition, the introduction of multiple protein interaction networks reduces the impact of single network data deviation and improves the stability and reliability of the prediction results.

[0025] 2. In terms of practicality and scalability, this method utilizes rich prior biological knowledge and integrates a variety of cancer-related omics data into the network, which improves the interpretability and practicality of the model. In addition, more types of omics data and new biological networks can be introduced as needed, which has good scalability and can adapt to different cancer types and diverse research needs.

[0026] 3. In terms of computational efficiency, the present invention not only reduces the redundancy of model parameters and improves computational efficiency by sharing the graph neural network module with heterogeneous genes, but also realizes information sharing and aggregation between different graphs by constructing a meta-graph and applying a meta-graph convolutional network, thereby improving overall computational efficiency.

[0027] 4. In terms of computing speed, the present invention improves the speed of training and testing by dividing the data into training sets and test sets and adopting a semi-supervised learning method. The present invention can utilize parallel computing technology to improve the speed of large-scale data processing and model training to meet practical application needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0029] Figure 1 It is a flow chart of a cancer driver gene prediction method based on a multi-input heterogeneous graph neural network of the present invention;

[0030] Figure 2 It is a type graph of nodes and edges in a heterogeneous network;

[0031] Figure 3 It is the data set partition diagram. DETAILED DESCRIPTION

[0032] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0033] Embodiment 1

[0034] This embodiment provides a method for predicting cancer driver genes based on graph neural networks;

[0035] Cancer driver gene prediction methods based on graph neural networks include:

[0036] S101: Obtain M types of protein interaction data; for each type of protein interaction data, encode the relationship between proteins into a gene-gene interaction relationship, and finally obtain M gene-gene interaction networks;

[0037] S102: configuring new nodes and new edges for the M gene-gene interaction networks respectively, to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes;

[0038] S103: Input each heterogeneous network into the trained cancer driver gene prediction model to obtain a prediction result of whether each gene is a cancer driver gene;

[0039] Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, and then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

[0040] Furthermore, the S101: obtain M types of protein interaction data, wherein the M types of protein interaction data include: CPDB, Multinet, PCNet, STRING-db, Iref and its latest version Iref.

[0041] Further, S101: for each protein interaction data, the relationship between proteins is encoded as a gene-gene interaction relationship, wherein the encoding process includes: converting the interaction relationship between proteins into the interaction relationship between genes through a protein-gene mapping table. If a protein corresponds to multiple genes, the main gene needs to be selected as a representative according to the weight or biological significance.

[0042] Furthermore, the M gene-gene interaction networks finally obtained include:

[0043] Genes are regarded as nodes of the network, and the relationships between genes are regarded as connecting edges between nodes, thus obtaining a gene-gene interaction network.

[0044] Further, S102: for the M gene-gene interaction networks, new nodes and new edges are respectively configured to obtain M heterogeneous networks; wherein the new nodes include: positional gene sets, pathways, MicroRNAs, computational gene sets, GeneOntology classification systems, and knowledge systems describing human disease-related phenotypes;

[0045] Positional gene sets are constructed based on the physical location of genes on chromosomes, helping to reveal the relationship between chromosome structure and function;

[0046] Pathways: a group of genes or proteins that interact with each other in a specific biological function, such as a signal transduction or metabolic pathway;

[0047] MicroRNA, small non-coding RNA, plays an important role in diseases by regulating gene expression;

[0048] Computational gene sets: functional gene sets predicted by algorithms, providing potential disease-related information;

[0049] GO, Gene Ontology classification system, is used to uniformly annotate gene functions, biological processes, and cellular components;

[0050] HPO, a knowledge system that describes human disease-related phenotypes and connects genes to disease characteristics;

[0051] Oncogenicity signatures, a collection of genes associated with core features of cancer, revealing the molecular mechanisms of tumorigenesis;

[0052] Cell type markers, a collection of key genes in a specific cell type, are used to explore the tissue-specific functions of genes;

[0053] The new edges refer to: the edges between gene nodes and positional gene set nodes, the edges between gene nodes and pathway nodes, the edges between gene nodes and MicroRNA nodes, the edges between gene nodes and computational gene set nodes, the edges between gene nodes and GO nodes, the edges between gene nodes and HPO nodes, the edges between gene nodes and carcinogenicity marker nodes, and the edges between gene nodes and cell type marker nodes.

[0054] Furthermore, the new node is multi-omics data related to cancer, including: positional gene sets, pathways, MicroRNAs, computational gene sets, GO, HPO, oncogenicity markers and cell type markers.

[0055] Furthermore, the S103: inputting each heterogeneous network into the trained cancer driver gene prediction model to obtain a prediction result of whether each gene is a cancer driver gene, the training process includes:

[0056] Constructing a training set and a test set; the training set and the test set are both heterogeneous networks of prediction results of whether they are known to be cancer driver genes;

[0057] The training set is input into the cancer driver gene prediction model to train the model; when the total loss function value of the model no longer decreases, the training is stopped to obtain the cancer driver gene prediction model after preliminary training;

[0058] The test set is input into the cancer driver gene prediction model after preliminary training to test the model. When the test evaluation index value exceeds the set threshold, it means that the training is completed and the current model is the final trained cancer driver gene prediction model.

[0059] Furthermore, the total loss function of the model is expressed as:

[0060] L total =L node +λL structure

[0061] L node is the node classification loss, L structure The loss is maintained for the structure and is balanced by the weight parameter λ.

[0062] Node classification loss is used to measure the accuracy of predicting gene driver mutations. Specifically, for each gene node i, its category label is yi , the probability distribution predicted by the model is The node classification loss uses the cross entropy loss function, which is defined as:

[0063]

[0064] Where N is the total number of gene nodes, C is the number of categories, and y i,c Indicates the true label of node i belonging to category C, represents the probability predicted by the model that node i belongs to category c.

[0065] The structure preservation loss is used to ensure that the model retains the topological structure of the original network during the learning process. Specifically, by measuring the similarity between neighbor nodes in the original network, it is defined as:

[0066]

[0067] Where |E| is the total number of edges in the heterogeneous network, (i,j) represents an edge in the network, and cos(h i ,h j ) represents the cosine similarity between the feature representations of node i and node j.

[0068] It should be understood that a protein interaction network is obtained; each protein interaction network is preprocessed, and various multi-omics data related to cancer are introduced as nodes, and the relationship between these newly added nodes and genes is regarded as an edge in a heterogeneous network to form a heterogeneous graph.

[0069] One of the graphs is selected as the test graph, and its data is divided into 75% training set and 25% test set. The data of all other graphs plus 75% of the data of this graph are used as the final training set input for training, and the optimal hyperparameter values ​​of the genetic heterogeneous shared graph neural network and meta-graph convolutional network are learned.

[0070] Similar to the training process, 25% of the selected test graph data is used as the test set input, and the semi-supervised node classification task is performed through the genetic heterogeneous shared graph neural network and meta-graph convolutional network to generate the final prediction results.

[0071] Furthermore, the trained cancer driver gene prediction model includes:

[0072] The shared graph neural network, meta-graph convolutional network and fully connected classification prediction network are connected in sequence, and the fully connected classification prediction network is implemented by a multi-layer perceptron.

[0073] Furthermore, a shared graph neural network is first used for each heterogeneous network to transfer information between nodes, including:

[0074] A shared graph neural network, consisting of several graph convolutional layers connected in sequence;

[0075] Each node in each heterogeneous network is updated through the graph convolution layer, and a new node representation is generated by aggregating the features of its neighboring nodes;

[0076] The feature update of a node includes two core steps: aggregation of neighbor features and updating of the node’s own features.

[0077] In the graph convolution layer, the neighbor feature aggregation of node u uses the following formula:

[0078]

[0079] Among them, Ν(u) represents the neighbor nodes of node u, Ν(u)∪{u} represents the neighbor set including the node itself, and h v represents the feature representation of node v, α uv represents the contribution weight of neighbor v to node u.

[0080] After feature aggregation, the feature representation of node u is updated to a new representation through linear transformation and nonlinear activation:

[0081] h′ u =σ(Wm u )

[0082] Where W is the learnable weight matrix of the graph convolutional layer, σ is the nonlinear activation function, and h′ u It is the output feature of node u in the current graph convolution layer.

[0083] Shared graph neural networks share the same graph convolutional layer parameters in all heterogeneous networks.

[0084] Furthermore, the specific formula of the shared graph neural network is as follows:

[0085]

[0086] Where m represents the number of current graph convolution layers, W is the shared graph convolution layer parameter (unified for all m), and N (m) (u) is the node u in the network G (m) The set of neighbor nodes in It's G (m) The weight between midpoints u and v.

[0087] Furthermore, the weight calculation of the graph convolution layer adopts a standard normalization method, and its specific formula is expressed as follows:

[0088]

[0089] in, and is the degree of nodes u and v.

[0090] Furthermore, the same gene node j of different heterogeneous networks is pointed to a meta-node to construct a meta-graph, and the meta-graph includes: a meta-node and several identical gene nodes j; there are interconnected edges between the meta-node and all gene nodes j, and the edges are directed edges from all gene nodes to the meta-node, and the weights of the edges are all set to 1.

[0091] Furthermore, the step of initializing the feature of the meta-node to the initial feature of the gene j includes: the initial features of all genes are the same, that is, the node feature x.

[0092] Furthermore, each meta-graph is input into a meta-graph convolutional network to obtain an updated feature representation of the meta-node, wherein the meta-graph convolutional network comprises: a plurality of sequentially connected meta-graph convolutional layers.

[0093] In the meta-graph convolutional layer, the neighbor feature aggregation of node u uses the following formula:

[0094]

[0095] Among them, Ν(u) represents the neighbor nodes of node u, Ν(u)∪{u} represents the neighbor set including the node itself, and h v represents the feature representation of node v, α uv represents the contribution weight of neighbor v to node u.

[0096] After feature aggregation, the feature representation of node u is updated to a new representation through linear transformation and nonlinear activation:

[0097] h′ u =σ(Wm u )

[0098] Where W is the learnable weight matrix of the meta-graph convolutional layer, σ is the nonlinear activation function, and h′ u It is the output feature of node u in the current graph convolution layer.

[0099] The meta-graph neural network shares the same graph convolutional layer parameters in all meta-graph networks.

[0100] Furthermore, the specific formula of the meta-graph convolutional network is as follows:

[0101]

[0102] Where m represents the number of convolutional layers in the current meta-graph, W is the shared convolutional layer parameter (unified for all m), and N (m) (u) is the node u in the network G(m) The set of neighbor nodes in It's G (m) The weight between midpoints u and v.

[0103] Furthermore, the updated feature representation of each meta-node is input into a multi-layer perceptron (MLP), which performs a series of linear transformations and nonlinear activation operations to obtain the predicted category of each meta-node, including:

[0104] Initialization: Input features

[0105] Hidden layer update: l=1,2,…,L-1;

[0106] Output layer:

[0107] in, is the weight matrix of the lth layer, is the bias vector of the lth layer, σ(·) is the activation function, and d l and d l-1 are the feature dimensions of the lth layer and the l-1th layer respectively, is the weight matrix of the output layer, C is the number of categories, is the bias vector of the output layer, is the probability distribution of meta-node u belonging to each category, normalized to probability using the softmax function.

[0108] The present invention uses the IG (Integrated Gradients) gradient integration module in the Captum toolkit to assign an importance score to each input feature. The input values ​​of the IG (Integrated Gradients) gradient integration module include the model, input features, reference values, target categories, and integration steps. The output value is the importance score of each input feature, which is used to explain the impact of the feature on the model output. IG is approximately the gradient integral of the model's output with respect to the input, and the gradient direction is along the straight line path from a specific baseline input to the current input. The most important node features in the network and the most critical edges in the network can be identified respectively.

[0109] like Figure 1 As shown, a cancer driver gene prediction method based on a multi-input heterogeneous graph neural network is provided. The method includes four parts: a multi-omics cancer gene heterogeneous network, a gene heterogeneous shared graph neural network, a meta-graph convolutional network, and a fully connected cancer gene prediction network. The loss function constrains the model training.

[0110] Multi-omics cancer gene heterogeneous networks (MCHN): In order to better utilize the rich prior biological knowledge and increase the interpretability of the model, in addition to the relationships between genes, various multi-omics data related to cancer from the MSigDB database are introduced as nodes in the graph, and the relationships between these newly added nodes and genes are regarded as edges in the heterogeneous network. The multi-omics cancer gene heterogeneous network (MCHN) is constructed by combining a variety of biological data and contains a heterogeneous network of different types of nodes and edges. Gene nodes represent genes involved in the cancer process; multi-omics nodes include positional gene sets, pathways, microRNAs, computational gene sets, etc. The characteristics of each node include cancer-related gene expression, copy number variation, DNA methylation and other information, and the characteristics of each edge represent the strength and type of interaction between nodes.

[0111] Gene heterogeneous shared graph neural network (GSGCN): For each input graph of multiple inputs, a graph neural network that performs message passing and updates the node representation matrix is ​​applied. The graph neural network is shared in all graphs. Gene heterogeneous shared graph neural network (GSGCN) consists of a series of shared graph convolutional layers. The graphs constructed from multi-omics cancer gene heterogeneous networks are respectively subjected to multiple shared graph convolution operations. The feature vector of each node is updated through the graph convolution layer, aggregating the features of its neighboring nodes to generate a new node representation.

[0112] Meta graph convolutions network (MGCN): In order to aggregate and share information between each graph, a meta graph is constructed, the same gene in all graphs is connected to a meta node, and a graph convolutional neural network is applied on each meta graph to update the information of the meta node. The meta graph convolution network (MGCN) consists of a series of meta graph convolutional layers to integrate information between different graphs. The gene node features output from the gene heterogeneous shared graph neural network (GSGCN) are subjected to meta graph construction and meta graph convolution operations to generate feature representations of meta nodes. The meta graph convolution layer performs convolution operations on the meta graph, aggregates the neighbor node features of each meta node, and updates the representation of the meta node.

[0113] Fully connected cancer gene prediction network (FCGN): The meta-nodes of the meta-graph convolutional network are input into a multi-layer perceptron for classification prediction to obtain the probability that the gene is a cancer driver gene. The fully connected cancer gene prediction network (FCGN), composed of a multi-layer perceptron, performs classification prediction to obtain the probability prediction result of the gene being a cancer driver gene.

[0114] Loss function: constrains the training of the network.

[0115] In order to represent the interaction between genes, the present invention uses 6 types of protein interaction (PPI) data and encodes the relationship between proteins as gene-gene interaction to form six basic homogeneous networks; then, through the construction of multi-omics cancer gene heterogeneous networks, multi-omics data is introduced to form six heterogeneous networks. Each heterogeneous network must pass through the gene heterogeneous shared graph neural network to transfer information between nodes, and then point the same gene node j in different graphs to a meta-node v j , construct j metagraphs G meta,j , and initialize the meta-node v corresponding to the initial features of gene j j Each meta-graph must pass through the meta-graph convolutional network to update the feature representation of the meta-node. At this stage, the model fuses and exchanges information between different networks. Finally, a multi-layer perceptron is used to predict the category of meta-node j.

[0116] The multi-omics cancer gene heterogeneous network (MCHN) is constructed by combining multiple biological data and contains heterogeneous networks with different types of nodes and edges. Gene nodes represent genes involved in the cancer process; multi-omics nodes include positional gene sets, pathways, microRNAs, computational gene sets, etc. These nodes are connected to gene nodes through different biological relationships to form edges in the heterogeneous network. The characteristics of each node include cancer-related gene expression, copy number variation, DNA methylation and other information, and the characteristics of each edge represent the strength and type of interaction between nodes. The types of nodes and edges in heterogeneous networks are as follows: Figure 2 shown.

[0117] The Gene Heterogeneous Shared Graph Neural Network (GSGCN) consists of a series of shared graph convolutional layers. The graphs constructed from the multi-omics cancer gene heterogeneous network are subjected to multiple shared graph convolution operations, and the feature vector of each node is updated through the graph convolution layer to aggregate the features of its neighboring nodes to generate a new node representation. The Gene Heterogeneous Shared Graph Neural Network shares the same graph convolutional layer parameters in all heterogeneous networks to ensure information transfer and sharing between different networks. This design allows us to process a variable number of graphs while keeping the number of trainable parameters fixed.

[0118] It should be noted that genetic heterogeneous shared graph neural networks can achieve efficient information transmission and sharing in multiple heterogeneous networks by sharing graph convolutional layer parameters, reducing the complexity and computational cost of model training.

[0119] The Meta-Graph Convolutional Network (MGCN) consists of a series of meta-graph convolutional layers to integrate information between different graphs. The gene node features output by the Gene Heterogeneous Shared Graph Neural Network (GSGCN) are transformed into meta-node feature representations through meta-graph construction and meta-graph convolution operations. In the meta-graph, all meta-nodes are connected by edges, and the weights of the edges are determined by the relationships between gene nodes in different networks. Each meta-node v j The initial features of are represented by the initial features of its corresponding gene node j. The meta-graph convolution layer performs convolution operations on the meta-graph, aggregates the features of the neighboring nodes of each meta-node, and updates the representation of the meta-node.

[0120] It is important to note that the meta-graph convolutional network effectively fuses and shares information from different networks by performing convolution operations on the meta-graph. The meta-graph convolutional network not only improves the prediction accuracy of the model, but also enhances its information integration ability and interpretability between different networks. In this way, the meta-graph convolutional network can more comprehensively capture the complex relationships between genes and improve the performance and reliability of cancer driver gene prediction.

[0121] Furthermore, the meta-nodes that have passed through the meta-graph convolutional network are input into a multi-layer perceptron for classification prediction to obtain the probability prediction result that the gene is a cancer driver gene.

[0122] In the method of the present invention, we obtain six protein interaction networks to train the proposed model and encode them as the relationship between genes. The six protein interaction networks are CPDB, Multinet, PCNet, STRING-db, Iref (2015) and its latest version Iref. The gene mutation frequency (MF), copy number aberration (CNA), DNA methylation (METH) and gene expression (GE) data of 29,446 samples from 16 different cancer types in the dataset TCGA are used as node features.

[0123] Depending on the data source, different confidence thresholds were used to filter out low-confidence interactions. Interactions with scores higher than 0.5 in CPDB and higher than 0.8 in STRINGdb were included in the network. Data for Multinet and IRefIndex (2015) were obtained from the Hotnet2 Git Hub repository. In the latest version of IRefIndex, the analysis was restricted to binary interactions between two human proteins. No further processing was done on PCNet.

[0124] In order to better utilize the rich prior biological knowledge and increase the interpretability of the model, we introduced a heterogeneous network, using various multi-omics data related to cancer from the MSigDB database as nodes, in addition to gene nodes. The present invention uses the latest version of the NCG database to define positive samples as cancer driver genes. Then, a set of potential cancer driver genes is obtained from the gene set supported by literature or research evidence in the CancerMine database. The genes remaining after removing the CancerMine set are considered negative samples.

[0125] The data is divided into training set and test set. One of the heterogeneous networks is selected as the test network, and its data is divided into 75% training set and 25% test set. The data of all other networks plus 75% of the data of this network are then divided into the final verification set and the final training set. The data set is divided as follows: Figure 3 During the training process, the validation set is used regularly to evaluate the model performance and adjust the hyperparameters to prevent overfitting.

[0126] The hyperparameters adjusted include learning rate, weight decay, number of hidden units, number of attention heads in GAT layers, dropout rate, and number of layers. The model was trained for 2000 epochs using the cross entropy loss function and the ADAM optimizer with a learning rate of 0.001. The initial graph neural network has 3 layers and a hidden dimension of 64, while the meta-graph convolutional network has a single layer and a hidden dimension of 64.

[0127] One of the graphs is selected as the test graph, and its data is divided into 75% training set and 25% test set. The data of all other graphs plus 75% of the data of this graph are used as the final training set input for training, and the optimal hyperparameter values ​​of the genetic heterogeneous shared graph neural network and meta-graph convolutional network are learned.

[0128] Similar to the training process, 25% of the selected test graph data is used as the test set input, and the semi-supervised node classification task is performed through the genetic heterogeneous shared graph neural network and meta-graph convolutional network to generate the final prediction results.

[0129] Through the construction of multi-omics cancer gene heterogeneous networks, multi-omics data is introduced to form six heterogeneous networks. Each heterogeneous network must pass through the gene heterogeneous shared graph neural network to transfer information between nodes, and then point the same gene node j in different graphs to a meta-node v j , construct j metagraphs G meta,j , and initialize the meta-node v corresponding to the initial features of gene j j Each meta-graph must pass through the meta-graph convolutional network to update the feature representation of the meta-node. At this stage, the model fuses and exchanges information between different networks. Finally, a multi-layer perceptron is used to predict the category of meta-node j. Compared with other cancer driver gene prediction methods based on deep learning, the present invention introduces multiple protein interaction networks, reduces the impact of single network data deviation, and improves the stability and reliability of the prediction results.

[0130] Embodiment 2

[0131] This embodiment provides a cancer driver gene prediction system based on graph neural network, including:

[0132] An acquisition module is configured to: acquire M kinds of protein interaction data; for each kind of protein interaction data, encode the relationship between proteins into a gene-gene interaction relationship, and finally obtain M gene-gene interaction networks;

[0133] A configuration module is configured to: configure new nodes and new edges for M gene-gene interaction networks respectively, to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes;

[0134] A prediction module is configured to: input each heterogeneous network into the trained cancer driver gene prediction model to obtain a prediction result of whether each gene is a cancer driver gene;

[0135] Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, and then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

[0136] It should be noted that the acquisition module, configuration module and prediction module described above correspond to steps S101 to S103 in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules described above as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0137] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0138] The proposed system can be implemented in other ways. For example, the system embodiment described above is only illustrative, and the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0139] Embodiment 3

[0140] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the above embodiment one.

[0141] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0142] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0143] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software.

[0144] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0145] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0146] Embodiment 4

[0147] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.

[0148] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A cancer driver gene prediction method based on graph neural network, characterized by: include: Obtain M kinds of protein interaction data; For each protein interaction data, the relationship between proteins is encoded as a gene-gene interaction relationship, and finally M gene-gene interaction networks are obtained; For M gene-gene interaction networks, new nodes and new edges are respectively configured to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes; Input each heterogeneous network into the trained cancer driver gene prediction model to obtain the prediction result of whether each gene is a cancer driver gene; Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; Finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

2. The method for predicting cancer driver genes based on graph neural network according to claim 1, characterized in that: Each heterogeneous network is input into the trained cancer driver gene prediction model to obtain the prediction result of whether each gene is a cancer driver gene. The training process includes: Constructing a training set and a test set; the training set and the test set are both heterogeneous networks of prediction results of whether they are known to be cancer driver genes; The training set is input into the cancer driver gene prediction model to train the model; when the total loss function value of the model no longer decreases, the training is stopped to obtain the cancer driver gene prediction model after preliminary training; The test set is input into the cancer driver gene prediction model after preliminary training to test the model. When the test evaluation index value exceeds the set threshold, it means that the training is completed and the current model is the cancer driver gene prediction model after final training. The total loss function of the model is expressed as: THE Total =L node +λL structure L node is the node classification loss, Ls tructure The loss is maintained for the structure and is balanced by the weight parameter λ; Node classification loss is used to measure the accuracy of predicting gene driver mutations; specifically, for each gene node i, its category label is y i , the probability distribution predicted by the model is The node classification loss uses the cross entropy loss function, which is defined as: Where N is the total number of gene nodes, C is the number of categories, and y i,c Indicates the true label of node i belonging to category C, represents the probability that the model predicts that node i belongs to category c; The structure preservation loss is used to ensure that the model retains the topological structure of the original network during the learning process; specifically, by measuring the similarity between neighbor nodes in the original network, it is defined as: Where |E| is the total number of edges in the heterogeneous network, (i, j) represents an edge in the network, and cos(h i ,h j ) represents the cosine similarity between the feature representations of node i and node j.

3. The method for predicting cancer driver genes based on graph neural network according to claim 1, characterized in that: For each heterogeneous network, a shared graph neural network is first used to transfer information between nodes, including: A shared graph neural network, consisting of several graph convolutional layers connected in sequence; Each node in each heterogeneous network is updated through the graph convolution layer, and a new node representation is generated by aggregating the features of its neighboring nodes; The feature update of a node includes two core steps: aggregation of neighbor features and update of the node’s own features; In the graph convolution layer, the neighbor feature aggregation of node u uses the following formula: Among them, N(u) represents the neighbor nodes of node u, N(u)∪{u) represents the neighbor set including the node itself, and h v represents the feature representation of node v, α uv represents the contribution weight of neighbor v to node u; After feature aggregation, the feature representation of node u is updated to a new representation through linear transformation and nonlinear activation: h′ u =σ(Wm u ) Where W is the learnable weight matrix of the graph convolutional layer, σ is the nonlinear activation function, and h′ u It is the output feature of node u in the current graph convolution layer.

4. The method for predicting cancer driver genes based on graph neural network according to claim 3, characterized in that: The specific formula of the shared graph neural network is: Among them, m represents the number of current graph convolution layers, W is the shared graph convolution layer parameter, and N (m) (u) is the node u in the network G (m) The set of neighbor nodes in It's G (m) The weight between midpoints u and v; The weight calculation of the graph convolution layer adopts a standard normalization method, and its specific formula is expressed as follows: in, and is the degree of nodes u and v.

5. The method for predicting cancer driver genes based on graph neural network according to claim 1, characterized in that: The same gene nodes j of different heterogeneous networks are pointed to a meta-node to construct a meta-graph, and the meta-graph includes: a meta-node and a plurality of the same gene nodes j; there are interconnected edges between the meta-node and all the gene nodes j, and the edges are directed edges from all the gene nodes to the meta-node, and the weights of the edges are all set to 1.

6. The method for predicting cancer driver genes based on graph neural network according to claim 1, characterized in that: The step of inputting each meta-graph into a meta-graph convolutional network to obtain updated feature representations of the meta-nodes, wherein the meta-graph convolutional network comprises: a plurality of sequentially connected meta-graph convolutional layers; In the meta-graph convolutional layer, the neighbor feature aggregation of node u uses the following formula: Among them, N(u) represents the neighbor nodes of node u, N(u)Y{u} represents the neighbor set including the node itself, and h v represents the feature representation of node v, α uv represents the contribution weight of neighbor v to node u; After feature aggregation, the feature representation of node u is updated to a new representation through linear transformation and nonlinear activation: h′ u =σ(Wm u ) Where W is the learnable weight matrix of the meta-graph convolutional layer, σ is the nonlinear activation function, and h′ u It is the output feature of node u in the current graph convolution layer; The meta-graph neural network shares the same graph convolutional layer parameters in all meta-graph networks.

7. The method for predicting cancer driver genes based on graph neural network according to claim 1, characterized in that: The specific formula of the meta-graph convolutional network is as follows: Among them, m represents the number of current meta-graph convolutional layers, W is the shared meta-graph convolutional layer parameter, and N (m) (u) is the node u in the network G (m) The set of neighbor nodes in It's G (m) The weight between midpoints u and v.

8. A cancer driver gene prediction system based on graph neural network, characterized by: include: An acquisition module is configured to: acquire M kinds of protein interaction data; For each protein interaction data, the relationship between proteins is encoded as a gene-gene interaction relationship, and finally M gene-gene interaction networks are obtained; A configuration module is configured to: configure new nodes and new edges for M gene-gene interaction networks respectively, to obtain M heterogeneous networks; the new nodes are multi-omics data related to cancer, and the new edges are connecting edges between the new nodes and the original gene nodes; A prediction module is configured to: input each heterogeneous network into the trained cancer driver gene prediction model to obtain a prediction result of whether each gene is a cancer driver gene; Among them, the trained cancer driver gene prediction model first uses a shared graph neural network to transfer information between nodes in each heterogeneous network, then points the same gene node j of different heterogeneous networks to a meta-node, constructs a meta-graph, and initializes the features of the meta-node to the initial features of gene j; N types of gene nodes, N meta-graphs are obtained; then each meta-graph is input into the meta-graph convolutional network to obtain the updated feature representation of the meta-node; Finally, the updated feature representation of each meta-node is input into the multi-layer perceptron to obtain the predicted category of each meta-node.

9. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; as well as a processor for executing the computer readable instructions, When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is executed.

10. A storage medium, characterized in that: The computer-readable instructions are non-transitory stored, wherein when the non-transitory computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Benign and malignant nodule grading evaluation system based on large model fusion ultrasonic imaging and thyroid gene marker

    CN120452757A

  • Cancer gene identification method based on multiplexing heterogeneous graph neural network

    CN120895103A