Glioblastoma prognosis prediction method based on expanded pathway
By adopting the deep neural network model based on expansion pathways and CGAN data augmentation technology in GBM prognostic prediction, the problem of HDLSS data and category imbalance is solved, the prediction accuracy is improved, and important prognostic factors and help in deep learning interpretability research.
Patent Information
- Application Number
- CN202411383510.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-06
AI Technical Summary
When using gene expression data for glioblastoma (GBM) prognosis prediction, the prior art faces the problems of high-dimensional, low sample size (HDLSS) data and category imbalance, resulting in poor accuracy of prediction results.
A deep neural network model based on expansion paths is adopted, combined with a conditional generation adversarial network (CGAN) for data augmentation. Through sparse encoding, a neural network model that can describe the expansion process of biological paths is constructed.
The problem of category imbalance in HDLSS data is effectively dealt with, improving the accuracy and reliability of GBM prognostic prediction, providing important prognostic factors, and providing assistance to deep learning interpretability research.
Smart Images

Figure CN119943362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of glioblastoma prognosis prediction, and in particular to a glioblastoma prognosis prediction method based on an extended pathway. Background Art
[0002] Glioblastoma (GBM), also known as glioblastoma multiforme, belongs to WHO grade 4 glioma. It is the most common primary malignant brain tumor and the most aggressive malignant brain tumor. In the United States alone, 12,120 patients were diagnosed with GBM in 2016, and the average survival rate of patients within 5 years was 5%. The peak incidence of GBM after age standardization is estimated to be 3.2 cases per 100,000 people. The incidence rate rises sharply after 54 years of age, reaching 15.24 cases per 100,000 people at the age of 75 to 84. In the past few decades, the median age of GBM has increased to 64 years, and the median survival of all GBM patients is only 8 months. Due to the short survival time of most patients, their prognosis is poor, and long-term survival of GBM patients is rare, of which more than 90% of patients die three years after diagnosis. Despite considerable efforts by medical workers, there are few reports on effective prolongation of survival of GBM patients. Although the treatment methods and technologies of neurosurgery, chemotherapy and radiotherapy have improved so far, the prognosis of GBM is still poor due to its high complexity and high mortality rate. Therefore, understanding the molecular mechanism of GBM and the development status of related biological pathways is of great significance to accelerate the progress of new treatments.
[0003] There are many methods that use gene expression data for prognosis prediction. Since the genomic data characteristics of most cancers are usually high-dimensional, GBM is no exception. High-dimensional, low-sample-size (HDLSS) data usually make prediction models sensitive to noise and false-positive associations, making prognosis prediction difficult. In addition, since cancer datasets are usually small and the number of samples in different categories may vary greatly, it has been shown in many classifiers that when the data is high-dimensional, the class imbalance problem is exacerbated, that is, even if there is no real difference between the classes, high dimensionality will cause the accuracy of the classification results to be biased towards the majority class. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a glioblastoma prognosis prediction method with a reasonable design, which improves GBM prognosis prediction based on expanded pathways and handles the HDLSS data category balance problem in response to the deficiencies of the existing technology.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for predicting the prognosis of glioblastoma based on an extended pathway, characterized in that the method comprises the following steps:
[0007] The first step is to obtain the GBM dataset from the Cancer Genome Atlas. The GBM dataset provides survival time and survival status. The data of surviving patients with a survival time of at least 24 months are retained, and the data of surviving patients with a survival time of less than 24 months are excluded, so that the data of the above-mentioned surviving patients are used as samples.
[0008] In the second step, the biological pathway database from the molecular signature database was used, in which the biological pathways of Reactome were extracted, and pathways containing at least 10 genes and genes using at least one pathway were selected from the biological pathways;
[0009] The third step is to build a deep neural network model, add the pathway selected in step 2 to the deep neural network model as an extended pathway layer, and introduce a data enhancement layer into the conditional generative adversarial network to enhance the data; input the GBM dataset from the input layer, and obtain more credible generated samples after being processed by the data enhancement layer; after synthesizing the samples, the data enhancement layer and the extended pathway layer are connected through sparse coding; in the hidden layer, the model captures the similarities and potential cross-regulations between gene pathways, and finally outputs the prediction results from the output layer.
[0010] The technical problem to be solved by the present invention can also be achieved through the following technical solutions. The processing process of the data enhancement layer described in step three is as follows: first, the data set is segmented into a training set and a test set, 80% of which is used for data balance and training models, and 20% is used for the final test; secondly, the data set is balanced by extracting category labels to automatically obtain the minority class, and based on the difference between the majority class and the minority class, a minority class of corresponding size is generated, and the CGAN balancing method is used to generate minority class samples in combination with specific category labels. The minority class samples and the training set constitute a synthetic data set, and finally, the synthetic data set is sent to the expansion path layer.
[0011] The technical problem to be solved by the present invention can also be achieved by the following technical solution: the CGAN is based on the GAN and combines the latent space z with additional information y as input, which is passed to the generator network;
[0012] The objective function of the CGAN model is:
[0013]
[0014] G is the generator, D is the discriminator, To minimize the operation of the generator G during the optimization process, is the maximization operation for the discriminator D in the optimization process. D(x|y) is the output probability of the discriminator D for the input sample x given the label y. G(z|y) is the sample generated from the latent space z by the generator G given the label y. V(D,G) represents the value or utility between the discriminator D and the generator G. represents the probability expectation of the real data distribution x, logS(x|y) represents the probability that the discriminator D predicts that the input X is a real sample when the label y is given, represents the probability expectation of the latent space z, where z is sampled from a pre-defined noise distribution, log(1-D(G(z|y))) represents the probability that the discriminator D predicts that the data generated by the generator G from the latent space z is fake given the label y; the above objective functions show the behavior identified by the adversarial process: one is related to better identifying samples belonging to the true distribution and the other is related to better identifying samples related to samples created by the generator.
[0015] The technical problem to be solved by the present invention can also be achieved through the following technical solutions. The expanded pathway layer described in step three expands the protein set of pathways and processes by utilizing the topological structure of the protein network, and adds new hypothetical genes to the pathways affecting GBM. The expansion process maps the proteins annotated in different cellular pathways into a large protein-protein interaction network, and expands these pathways by adding their closest network neighboring nodes.
[0016] The technical problem to be solved by the present invention can also be achieved by the following technical solution: the protein network is obtained from the STRING database, and is screened according to a standard threshold of 0.4 to establish connections between proteins; filtering rules are established according to graph theory to expand the acquired pathways; the proteins corresponding to the pathway genes are regarded as seed nodes and mapped to the protein network, and the direct neighbors of these seed nodes are regarded as candidate nodes for the expansion process and filtered according to the rules.
[0017] The technical problem to be solved by the present invention can also be achieved by the following technical solution. In the filtering step, the calculation formula of the rule is as follows:
[0018] (1) Node weight filtering, degree(v) is the number of direct connections of node v, which must be greater than 1:
[0019] degree(v)>1
[0020] (2) Direct path filtering: precesslinks(v,p) is the number of direct links from node v to other nodes in path p, and outsidelinks(v,p) is the number of direct links from node v to nodes outside path p. The quotient of the two is greater than the threshold T1, which is set to 1.0, corresponding to the condition for defining a “strong community” in a network in graph theory:
[0021]
[0022] (3) Path extension process filtering: trianglelinks(v,p) is the number of triangles that node v actually forms with nodes in path p and another candidate node. possibletriangles(v,p) is the number of triangles that all candidate nodes in path p can possibly form with node v. The quotient of the two is greater than the threshold T2, which is set to 0.1:
[0023]
[0024] (4) Path node coverage filtering, processlinks(v,p) is the number of direct connections from node v to other nodes in path p, processnodes(p) is the number of nodes in the entire path, and the quotient of the two is greater than the threshold T3, which is set to 0.3:
[0025]
[0026] A candidate node v must satisfy at least one of the rules (2)-(4) on the basis of satisfying rule (1) in order to be added to the path p.
[0027] The technical problem to be solved by the present invention can also be achieved by the following technical solution. The connection process of sparse coding in step 3 is to imitate the signal pathway in biology through sparse matrix control. The sparsity between the data enhancement layer and the extended pathway layer is defined by matrix A:
[0028] h (e) =a((W (d) ★A)h (d) +b (d) )
[0029] Where * represents the multiplication of two elements, a is the activation function, h (e) represents the output tensor of the extended path layer, h (d) Represents the output tensor on the data enhancement layer, W (d) and b (d) denote the weight matrix and bias vector respectively.
[0030] Compared with the prior art, the present invention uses the extended pathway and deep neural network EPDNN. EPDNN constructs a neural network model that can describe the biological pathway extension process, providing an important prognostic factor for accurately predicting patient survival. EPDNN will play an important role in the field of prognosis prediction and provide important help for deep learning interpretability research; at the same time, combined with the CGAN algorithm, it provides a reliable strategy for training deep neural network models for HDLSS data. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is an architecture diagram of the deep neural network model in the present invention;
[0032] Figure 2 Diagram of the process of using CGAN to enhance data for the data enhancement layer;
[0033] Figure 3 An expansion path process diagram for the expansion path layer;
[0034] Figure 4 This is a diagram of the operation process of sparse coding;
[0035] Figure 5 (A)-(E) is a performance comparison of different algorithms;
[0036] Figure 6 This is the ROC curve of the ablation experiment. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0038] A method for predicting the prognosis of glioblastoma based on an extended pathway, the method comprising the following steps:
[0039] 1.1 Data Acquisition
[0040] The data set used in the present invention is a GBM data set obtained from The Cancer Genome Atlas (TCGA). The TCGA database provides a huge amount of GBM clinical, sequencing and long-term follow-up data, which provides the possibility for genetic analysis of recurrent GBM and construction of prediction models.
[0041] The obtained GBM dataset contains 522 samples and gene expression data of 12,042 genes, and the dataset provides survival time and survival status. The experiment does not consider the current survival status, but only the survival time. Patients who have survived in the past 24 months are regarded as long-term survival (LTS), and patients who died in less than 24 months are regarded as short-term survival (STS). The data of surviving patients with a survival time of less than 24 months were excluded in the experiment because these data still have room for further development and do not meet the conditions of short-term survival. Finally, 99 LTS and 376 STS samples were screened and obtained. According to statistics, about 20% of the samples are LTS patients.
[0042] For pathway-based analysis, the study used a biological pathway database from the Molecular Signature Database (MSigDB). From MSigDB, we extracted the biological pathways of Reactome. Then, pathways containing less than 10 genes were excluded, because small pathways may often be redundant with larger pathways. As input features, we considered genes belonging to at least 1 pathway, because the pathway annotation of genes is crucial for constructing the mask matrix between the input layer and the pathway layer. Finally, 574 pathways and 4359 genes were considered in the experiment, and the gene expression data were standardized to have a mean of 0 and a standard deviation of 1.
[0043] 1.2 Overview of EPDNN
[0044] Reference Figure 1 , Figure 1 The overall architecture diagram of EPDNN. The expanded pathway is added to the model as an expanded pathway layer, and the conditional generative adversarial network introduces a data enhancement layer to enhance the data. The architecture of the deep neural network has five layers, consisting of an input layer, a data enhancement layer, an expanded pathway layer, a hidden layer, and an output layer. The data set is input from the input layer, and after being processed by the data enhancement layer, a more credible generated sample is obtained. After the sample is synthesized, the data enhancement layer and the expanded pathway layer are connected through a mask matrix, which represents the non-fully connected relationship between genes and pathways. In the hidden layer, the model can capture the similarities and potential cross-regulations between gene pathways, and finally output the prediction results from the output layer.
[0045] Note: (A) The pathway expansion module consists of three parts: extracting pathways, building gene networks, and expanding pathways. The gene network construction part requires mapping the protein network to a gene network, determining the connection relationship through support, and establishing rules based on graph theory during the expansion process; (B) The EPDNN deep learning framework consists of five layers, namely the input layer, data enhancement layer, pathway expansion layer, hidden layer, and output layer; (C) The data enhancement layer is embedded with a conditional generative adversarial network, which can better fit the original data.
[0046] 1.3 Data Enhancement Layer
[0047] Reference Figure 2 , the data enhancement layer is broken down in detail, and consists of four steps. First, split the data set into a training set and a test set, 80% of which is used for data balancing and training models, and 20% is used for the final test. Secondly, balance the data set. By extracting the category labels, the module automatically obtains the minority class, and generates a minority class of corresponding size based on the difference between the majority class and the minority class. The balancing method uses CGAN. Compared with SMOTE or GAN, CGAN can learn all categories of gene expression data, and then generate minority class samples in combination with specific category labels, making the generated samples more robust. Finally, the synthetic data set is sent to the pathway layer.
[0048] The standard GAN model structure consists of two neural networks, called the generator (G) and the discriminator (D), which are trained simultaneously to produce an adversarial process. For an input sample x, the goal of the discriminator is to predict the probability that the sample belongs to the real data distribution rather than being generated by the generator. At the same time, the task of the generator is to sample from a priori defined random noise distribution z and output synthetic samples that can mimic the real data distribution. CGAN is a variant of GAN. Based on GAN, the latent space z is combined with additional information y as input (y is the label in this invention) and passed to the generator network. The objective function of the CGAN model (Equation 1) shows the behavior identified by the adversarial process: one is related to better identifying samples belonging to the real distribution and the other is related to better identifying samples created by the generator:
[0049] 1.4 Expanding the access layer
[0050] Reference Figure 3 Biological pathways play a crucial role in the development of cancer. A study shows that by expanding the protein set of pathways and processes using the topology of protein networks, new putative genes can be added to pathways that affect GBM.
[17] The pathway expansion method refers to mapping proteins annotated in different cellular pathways into a large protein-protein interaction network and expanding these pathways by adding their closest network neighbors. The expansion rules are as follows: Figure 3 As shown in the figure, the black nodes represent the members that have been added to the path, and the colored nodes show the expansion process of the new nodes.
[0051] The protein interaction network (PPI) was obtained from the STRING database (https: / / cn.string-db.org) and was screened according to the standard threshold of 0.4 to establish the connection between proteins. The experiment established filtering rules based on graph theory and expanded the 574 pathways obtained. The proteins corresponding to the pathway genes were regarded as seed nodes and mapped to the protein network. The direct neighbors of these seed nodes were regarded as candidate nodes in the expansion process and filtered according to the rules. In the filtering step, the candidate node v must satisfy at least one of the rules (2)-(4) on the basis of satisfying rule (1) before it can be added to the path p. Figure 3 In the figure, the gray part indicates that the nodes do not meet rule (1) and are excluded, the red part indicates that the nodes meet (1) and (2) and are added to the path, the yellow part indicates that the nodes meet (1) and (3) and are added to the path, and the blue part indicates that the nodes meet (1) and (4) and are added to the path. The calculation formula is as follows:
[0052] (1) Node weight filtering, degree(v) is the number of direct connections of node v, which must be greater than 1:
[0053] degree(v)>1
[0054] (2) Direct path filtering: processlinks(v,p) is the number of direct links from node v to other nodes in path p, and outsidelinks(v,p) is the number of direct links from node v to nodes outside path p. The quotient of the two is greater than the threshold T1, which is set to 1.0, corresponding to the condition of defining a “strong community” in a network in graph theory.
[28] :
[0055]
[0056] (3) Path extension process filtering: trianglelinks(v,p) is the number of triangles actually formed by node v, nodes in path p and another candidate node. possibletriangles(v,p) is the number of triangles that all candidate nodes in path p can possibly form with node v. The quotient of the two is greater than the threshold T2, which is set to 0.1:
[0057]
[0058] (4) Path node coverage filtering, processlinks(v,p) is the number of direct connections from node v to other nodes in path p, processnodes(p) is the number of nodes in the entire path, and the quotient of the two is greater than the threshold T3, which is set to 0.3:
[0059]
[0060] 1.5 Sparse Coding
[0061] Between the data enhancement layer and the extended pathway layer, sparse coding is combined to determine the connection, and the extended pathway genes are encoded as sparse matrices. Through sparse matrix control, the signal pathway in biology is simulated, and the sparsity between the data enhancement layer and the extended pathway layer is defined by matrix A:
[0062] h (e) =a((W (d) ★A)h (d) +b (d) )
[0063] Where * represents the multiplication of two elements, and a is the activation function. (e) represents the output tensor of the extended path layer, h (d) Represents the output tensor on the data enhancement layer, W (d) and b (d) Represent the weight matrix and bias vector respectively. The element values of A are composed of 0 and 1, so that each node of the extended path layer focuses on the features determined by the path information. The actual operation process is as follows Figure 4 Indicated by the red arrow.
[0064] 2 Experiments and Results
[0065] 2.1 Parameter settings
[0066] All methods of the present invention are written based on Python 3.8 and Pytorch 2.1, and the code is debugged and run on PyCharm. The operating system uses Windows 11, the CPU is AMD R5-6600H, the GPU is NVIDIA GeForce RTX3050Ti, and the memory and video memory are 16GB and 12GB respectively. The detailed parameters of model training are shown in Table 1.
[0067] Table 1 Model training parameters
[0068]
[0069] 2.2 Evaluation Metrics
[0070] Assume that the samples belonging to the LTS sample type in the GBM test data set are called positive samples (Positive), and the samples belonging to the STS sample type are called negative samples (Negative). TP and FP represent the correctly and incorrectly classified positive samples, and TN and FN represent the correctly and incorrectly classified negative samples. The present invention selects the following five indicators to evaluate the multi-class classification problem.
[0071] 2.2.1 Accuracy
[0072] Accuracy is the most commonly used evaluation indicator for classification models, that is, the proportion of correctly classified samples in all samples. The higher the accuracy, the better the accuracy of the model in classifying samples as a whole. The calculation formula is as follows:
[0073]
[0074] 2.2.2 Area under the curve (AUC)
[0075] The area under the curve refers to the area under the ROC (Receiver Operating Characteristic curve), which is an indicator for evaluating the performance of a classification model. The ROC curve is plotted with the False Positive Rate (FPR) as the horizontal axis and the True Positive Rate (TPR) as the vertical axis, showing the performance changes of the model under different classification thresholds.
[0076] Arrange all (FPR, TPR) pairs in ascending order of FPR, and form a trapezoid between each pair of adjacent points (i, i+1). i The area calculation formula is as follows:
[0077]
[0078] Adding up the areas of all the trapezoids gives us an estimate of the AUC:
[0079]
[0080] 2.2.3 Precision
[0081] Precision measures the proportion of samples that are actually positive among all samples predicted by the model to be positive, which is very important when focusing on the accuracy of positive sample identification. The calculation formula is as follows:
[0082]
[0083] 2.2.4 Recall
[0084] Recall rate is an important indicator for evaluating the ability of a classification model to identify all actual positive samples. The higher the recall rate, the more comprehensive the model is in detecting positive samples, that is, it is less likely to miss positive samples. A high recall rate is particularly important for prognosis prediction, safety inspection, etc. The calculation formula is as follows:
[0085]
[0086] 2.2.5 F1-score
[0087] The F1 score is the harmonic mean of precision and recall. For a binary classification problem, the F1 score is calculated as follows:
[0088]
[0089] Precision and Recall are calculated by the formula of precision and recall respectively.
[0090] The evaluation content of the method described in the present invention includes the following two aspects:
[0091] First, the performance of the EPDNN is evaluated. In order to evaluate the performance of the EPDNN, the present invention uses classic machine learning models and the recent PASNet prognosis prediction model on the glioblastoma dataset to predict survival. Specifically, the study uses a five-fold cross-validation method repeated five times on the dataset, and compares the performance of each model by calculating the Accuracy, AUC, Precision, Recall and F1-score of five five-fold cross-validations on the test set, and taking the average of the results. These five indicators are commonly used indicators for evaluating the performance of binary classification tasks. For a fair comparison, before the dataset is input into other models, the same data processing and enhancement steps as EPDNN are performed. Figure 5(A)-Figure 5(E) As shown, the present invention compares EPDNN with logistic regression (LR), naive Bayes (NB), random forest (RF), and PASNet. It can be seen that EPDNN has achieved better results in the four indicators of Accuracy, AUC, Recall, and F1-score than the other four models, which are 0.8790, 0.9244, 0.8267, and 0.8671, respectively, which are all far behind the second place. In terms of Precision, it is not as good as 0.9826 of NB and 0.9423 of RF. The reason may be that the data set is too small and the amount of available data is small. The deep learning model may not be able to fully utilize its capacity, resulting in performance inferior to some traditional machine learning models with higher data efficiency. Among them, in Figure 5, (A) the average accuracy of the five models is compared; (B) the average area under the ROC curve of the five models is compared; (C) the average precision of the five models is compared; (D) the average recall of the five models is compared; (E) the average F1 score of the five models is compared.
[0092] Second, verify whether each module of EPDNN can help improve model performance. This study designed the following EPDNN variant models for comparison:
[0093] ① No Data Augmentation (NDA): This variant model removes the data augmentation layer for model training.
[0094] ② No Extended Pathways (NEP): This variant model replaces the extended pathway layer with a normal fully connected layer for model training.
[0095] ③ No data enhancement layer and extended path layer (NDA&NEP): This variant model removes the data enhancement layer and replaces the extended path layer with a normal fully connected layer for model training.
[0096] like Figure 6 As shown in the figure, compared with the complete EPDNN, the AUROC of NDA decreased by about 3.3% on average, and NDA&NEP decreased by about 12.61% on average compared with NEP. These results show that the performance of NDA and NDA&NEP is significantly worse and cannot adapt to all types of samples, indicating that the data enhancement module plays a certain role in improving the generalization ability of the model, and the improvement is more obvious for simple models. The difference of 3.4% between the performance of NEP and the AUROC of NDA shows that when the data set is relatively unbalanced, the effect of training the model by combining the initial data set with the extended pathway layer is better than that of the single enhanced data set. NEP is 6.61% lower than the complete model on average, which further highlights the importance of the extended pathway layer in improving the performance of EPDNN. It can be seen that each module of EPDNN has made a positive contribution to its optimal performance.
[0097] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for predicting the prognosis of glioblastoma based on an extended pathway, characterized in that: The Method The following steps are included: The first step is to obtain the GBM dataset from the Cancer Genome Atlas. The GBM dataset provides survival time and survival status. The data of surviving patients with a survival time of at least 24 months are retained, and the data of surviving patients with a survival time of less than 24 months are excluded, so that the data of the above-mentioned surviving patients are used as samples. In the second step, the biological pathway database from the molecular signature database was used, in which the biological pathways of Reactome were extracted, and pathways containing at least 10 genes and genes using at least one pathway were selected from the biological pathways; The third step is to build a deep neural network model, add the pathway selected in step 2 to the deep neural network model as an extended pathway layer, and introduce a data enhancement layer into the conditional generative adversarial network to enhance the data; The GBM dataset is input from the input layer and processed by the data enhancement layer to obtain more credible generated samples; After synthesizing the samples, the data enhancement layer and the extended pathway layer are connected through sparse coding; in the hidden layer, the model captures the similarities and potential cross-regulations between gene pathways, and finally outputs the prediction results from the output layer.
2. The method for predicting the prognosis of glioblastoma based on the extended pathway according to claim 1, characterized in that: The processing process of the data enhancement layer described in step three is as follows: first, split the data set into a training set and a test set, 80% of which is used for data balancing and training the model, and 20% is used for the final test; secondly, balance the data set by extracting the category label, automatically obtain the minority class, and generate a minority class of corresponding size based on the difference between the majority class and the minority class. Use the CGAN balancing method, and then combine the specific category label to generate minority class samples. The minority class samples and the training set constitute a synthetic data set. Finally, the synthetic data set is sent to the expansion path layer.
3. The method for predicting the prognosis of glioblastoma based on the extended pathway according to claim 2, characterized in that: The CGAN described above combines the latent space z with additional information y as input based on GAN and passes it to the generator network; The objective function of the CGAN model is: G is the generator, D is the discriminator, To minimize the operation of the generator G during the optimization process, is the maximization operation for the discriminator D in the optimization process. D(x|y) is the output probability of the discriminator D for the input sample x given the label y. G(z|y) is the sample generated from the latent space z by the generator G given the label y. V(D,G) represents the value or utility between the discriminator D and the generator G. represents the probability expectation of the real data distribution x, log D(x|y) represents the probability that the discriminator D predicts that the input X is a real sample when the label y is given, represents the probability expectation of the latent space z, where z is sampled from a predefined noise distribution, and log(1-D(G(z|y))) represents the probability that the discriminator D predicts that the data generated by the generator G from the latent space z is false when the label y is given; The above objective functions demonstrate the behavior of identification through the adversarial process: one is related to better identifying samples that belong to the true distribution and the other is related to better identifying samples created by the generator.
4. The method for predicting the prognosis of glioblastoma based on the extended pathway according to claim 1, characterized in that: The expanded pathway layer described in step 3 expands the protein set of pathways and processes by utilizing the topological structure of the protein network to add new putative genes to pathways affecting GBM. The expansion process is to map the proteins annotated to different cellular pathways into a large protein-protein interaction network and expand these pathways by adding their closest network neighbor nodes.
5. The method for predicting the prognosis of glioblastoma based on the extended pathway according to claim 4, characterized in that: The protein network is obtained from the STRING database, and is screened according to a standard threshold of 0.4 to establish connections between proteins; filtering rules are established based on graph theory to expand the obtained pathways; The proteins corresponding to the pathway genes are regarded as seed nodes and mapped to the protein network, and the direct neighbors of these seed nodes are regarded as candidate nodes for the expansion process and filtered according to the rules.
6. The method for predicting the prognosis of glioblastoma based on the extended pathway according to claim 5, characterized in that: In the filtering step, the rule is calculated as follows: (1) Node weight filtering, degree(v) is the number of direct connections of node v, which must be greater than 1: degree(v)>1 (2) Direct path filtering: processlinks(v,p) is the number of direct links from node v to other nodes in path p, and outsidelinks(v,p) is the number of direct links from node v to nodes outside path p. The quotient of the two is greater than the threshold T1, which is set to 1.0, corresponding to the condition of defining a "strong community" in a network in graph theory: (3) Path extension process filtering: trianglelinks(v,p) is the number of triangles actually formed by node v, nodes in path p and another candidate node. possibletriangles(v,p) is the number of triangles that all candidate nodes in path p can possibly form with node v. The quotient of the two is greater than the threshold T2, which is set to 0.1: (4) Path node coverage filtering, processlinks(v,p) is the number of direct connections from node v to other nodes in path p, processnodes(p) is the number of nodes in the entire path, and the quotient of the two is greater than the threshold T3, which is set to 0.3: A candidate node v must satisfy at least one of the rules (2)-(4) in addition to satisfying rule (1) in order to be added to path p.
7. The method for predicting the prognosis of glioblastoma based on the expanded pathway according to claim 1, characterized in that: The connection process of sparse coding in step 3 is controlled by sparse matrices, imitating the signal pathway in biology. The sparsity between the data enhancement layer and the expansion pathway layer is defined by matrix A: h (e) =a((W (d) *A)h (d) +b (d) ) Where ★ represents the multiplication of two elements, a is the activation function, h (e) represents the output tensor of the extended path layer, h (d) Represents the output tensor on the data enhancement layer, W (d) and b (d) denote the weight matrix and bias vector respectively.