Analysis method for prognosis gene characteristics of glioma based on machine learning
By constructing a machine learning-based prognostic scoring model and combining gene expression and pathological variable data of gliomas to generate nomograms, the problem of insufficient accuracy in prognostic assessment of gliomas is solved, and personalized prognostic risk analysis and decision support are realized.
Patent Information
- Application Number
- CN202610145226.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-12
AI Technical Summary
Current technologies lack methods to integrate gene expression information related to biological processes with clinicopathological variables, resulting in insufficient accuracy in prognostic assessment of gliomas and failing to meet the need for personalized and precise prognostic judgment.
A prognostic scoring model was constructed using machine learning methods. By combining prognostic gene expression data and pathological variable data, a correlation model was generated through multivariate regression analysis, and a nomogram was constructed to achieve prognostic risk analysis of glioma samples.
It improves the objectivity and accuracy of prognostic assessment for gliomas, provides more precise and individualized prognostic risk stratification results, and supports individualized decision-making.
Smart Images

Figure CN122024853A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biological detection, specifically to a method for analyzing the prognostic gene characteristics of glioma based on machine learning. Background Technology
[0002] Gliomas are the most common primary malignant tumors of the central nervous system, characterized by high heterogeneity and aggressiveness, and generally poor prognosis, especially high-grade gliomas. Currently, clinical prognostic assessment of gliomas mainly relies on histopathological grading and limited molecular markers (such as IDH mutations and 1p / 19q co-deletion status). However, these traditional indicators still struggle to achieve precise individualized prognostic stratification and cannot fully reflect the biological heterogeneity within the tumor and its complex interactions with the tumor microenvironment, thus requiring improvement in the accuracy of prognostic prediction.
[0003] The shortcomings and drawbacks of existing technologies lie in the lack of a method that can integrate gene expression information related to specific biological processes (such as copper death) with clinicopathological variables and automatically and quantitatively assess the prognosis of glioma samples based on data analysis models. This results in insufficient accuracy and objectivity in prognostic assessment, making it difficult to meet the clinical demand for personalized and precise prognostic judgments. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method for analyzing the prognostic gene characteristics of glioma based on machine learning, in order to solve the problems of insufficient accuracy and lack of quantitative analysis models in existing glioma prognostic assessment methods.
[0005] In a first aspect, embodiments of the present invention provide a method for analyzing the prognostic gene characteristics of gliomas based on machine learning, the method comprising: Obtain prognostic gene expression data and pathological variable data corresponding to glioma assessment samples; The prognostic gene score of the prognostic gene expression data is determined based on a pre-constructed prognostic scoring model. Based on the prognostic gene scores and the pathological variable data, the prognostic risk of the glioma assessment samples was analyzed to obtain the prognostic risk analysis results.
[0006] Furthermore, the prognostic risk analysis of the glioma assessment sample based on the prognostic gene score and the pathological variable data, to obtain the prognostic risk analysis results, includes: The risk level group to which the glioma assessment sample belongs is determined based on the prognostic gene score; Substituting the prognostic gene scores and the pathological variable data into a pre-constructed nomogram, the prognostic status value of the glioma assessment sample at a specified time point is obtained. The prognostic risk analysis results of the glioma assessment samples are generated based on the risk level grouping and the prognostic status value.
[0007] Furthermore, the method for constructing the pre-built nomogram includes: Obtain prognostic gene scores and corresponding pathological variable data for glioma training samples; Regression analysis was performed on the prognostic gene scores and the pathological variable data to generate a correlation model; Based on the preset weights of each variable in the association model, a nomogram is constructed that maps the prognostic gene score and the pathological variable data to the prognostic state value.
[0008] Furthermore, after determining the risk group to which the glioma assessment sample belongs based on the prognostic gene score, the method further includes: Obtain microenvironmental characteristic data and / or tumor mutation burden data for each of the aforementioned risk level groups; The microenvironmental characteristic data and / or the mutation load data are used as biological characteristic data to analyze the biological characteristic association between the biological characteristic data and the risk level grouping; The prognostic risk analysis results are generated based on the risk level grouping, the prognostic status value, and the correlation of the biological characteristics.
[0009] Furthermore, the method for constructing the prognostic scoring model includes: Obtain gene expression data from glioma training samples; Expression data of copper death-related genes were extracted from the gene expression data, and cluster analysis was performed based on the expression data of copper death-related genes to obtain multiple molecular subtypes; Differentially expressed genes among different molecular subtypes are identified, and the prognostic scoring model is constructed based on the differentially expressed genes.
[0010] Furthermore, the cluster analysis based on the expression data of the copper death-related genes yields multiple molecular subtypes, including: Consensus clustering analysis was performed on the expression data of the copper death-related genes to obtain the number of target clusters; The glioma training samples are divided into multiple molecular subtypes based on the target cluster number.
[0011] Furthermore, the construction of the prognostic scoring model based on the differentially expressed genes includes: From the differentially expressed genes, prognostic candidate genes related to the assessment time point of the glioma training samples were screened; Feature selection is performed on the candidate prognostic genes to obtain a set of prognostic feature genes; The prognostic scoring model is constructed based on the expression levels of each gene in the prognostic feature gene set.
[0012] Secondly, embodiments of the present invention provide a prognostic analysis device for glioma samples, the device comprising: The acquisition module is used to acquire prognostic gene expression data and pathological variable data corresponding to glioma assessment samples; A calculation module is used to calculate the prognostic gene score of the prognostic gene expression data based on a pre-constructed prognostic scoring model. The analysis module is used to analyze the prognostic risk of the glioma assessment sample based on the prognostic gene score and the pathological variable data, and obtain the prognostic risk analysis results.
[0013] Thirdly, embodiments of the present invention provide a computer device, including: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.
[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause a computer to perform the method described in the first aspect or any of its corresponding embodiments.
[0015] The method provided in this application has the following beneficial effects: The method provided in this application achieves systematic integration of multi-source and multi-dimensional information by acquiring prognostic gene expression data and pathological variable data corresponding to glioma assessment samples. This overcomes the limitations of traditional methods that rely on a single type of data for judgment, laying a data foundation for objective assessment. By determining the prognostic gene score based on a pre-constructed prognostic scoring model, complex gene expression profile information is transformed into a quantifiable comprehensive score. This process utilizes a trained mathematical model, significantly improving the objectivity and efficiency of the assessment and avoiding subjective biases from human experience-based judgments. Based on the prognostic gene score and pathological variable data, the prognostic risk of glioma assessment samples is analyzed to obtain prognostic risk analysis results. This achieves synergistic analysis and fusion of quantitative scores and key pathological features, thereby generating more accurate and individualized prognostic risk stratification results and providing data support with higher predictive value than traditional methods for decision-making. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a method for analyzing the prognostic gene characteristics of glioma based on machine learning, according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating another method for analyzing the prognostic gene characteristics of glioma based on machine learning, according to an embodiment of the present invention. Figure 3A This is a nomogram of predictions for 1-year, 3-year, and 5-year overall survival of patients with glioma according to an embodiment of the present invention; Figure 3B This is a schematic diagram of the ROC curve of the training set and the prediction of the total survival of 1 year, 3 years and 5 years according to an embodiment of the present invention. Figure 3C This is a schematic diagram of the ROC curve of the validation set and the prediction of the total survival of 1 year, 3 years and 5 years according to an embodiment of the present invention. Figure 3D This is a schematic diagram of the ROC curve of the test set and the prediction of the total survival of 1 year, 3 years and 5 years according to an embodiment of the present invention. Figure 4 This is a flowchart illustrating another method for analyzing the prognostic gene characteristics of glioma based on machine learning, according to an embodiment of the present invention. Figure 5 This is a flowchart illustrating another method for analyzing the prognostic gene characteristics of glioma based on machine learning, according to an embodiment of the present invention. Figure 6 This is a technical roadmap for multi-omics analysis and prognostic risk assessment of glioma according to embodiments of the present invention; Figure 7 This is a structural block diagram of a glioma sample prognostic analysis device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] According to embodiments of the present invention, a method for analyzing the prognostic gene characteristics of glioma based on machine learning is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0020] This embodiment provides a method for analyzing the prognostic gene characteristics of gliomas based on machine learning. Figure 1 This is a flowchart of a method for analyzing the prognostic gene features of glioma based on machine learning, according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps: Step S101: Obtain prognostic gene expression data and pathological variable data corresponding to the glioma assessment sample.
[0021] It should be noted that prognostic gene expression data refers to the expression levels of specific genes closely related to the prognosis of glioma, selected through screening. This group of genes originated from differentially expressed genes identified after in-depth analysis of copper death-related genes, and was further refined using machine learning algorithms. Their expression profiles are the core input for calculating the quantitative prognostic variance. Pathological variable data refers to pathological features associated with the glioma assessment samples, including but not limited to tumor grade, isocitrate dehydrogenase (IDH) gene mutation status, chromosome 1p / 19q co-deletion status, and O6-methylguanine-DNA methyltransferase (MGMT) promoter methylation status. These variables are key known factors influencing prognosis.
[0022] In this embodiment, the process is based on access to and standardization of public genomic databases. First, raw RNA sequencing (RNA-seq) data of glioma samples (including low-grade gliomas and glioblastomas) are obtained from authoritative databases (such as the Cancer Genome Atlas (TCGA)). Simultaneously, transcriptome data of normal brain tissue can be obtained from resources such as the Genotype-Tissue Expression (GTEx) project to provide controls or perform necessary calibration. Second, the obtained raw gene expression data undergoes rigorous preprocessing, including normalization using specialized tools (such as R packages) to eliminate batch effects, and precise matching and integration of gene expression data with detailed pathological annotation information of the samples to form a structured dataset. For prognostic gene expression data, the core is to extract the expression levels of a subset of specific prognostic characteristic genes determined by the model of this invention from the whole transcriptome data. For pathological variable data, it is obtained directly through an associated metadata table or after standardized interpretation based on raw detection results (such as sequencing and methylation microarray data).
[0023] As an example, during the data preprocessing stage, raw RNA-seq counts of 689 glioma samples, including low-grade gliomas (LGG) and glioblastomas (GBM), were obtained from the TCGA database, while RNA-seq data of 1141 normal brain samples were obtained from the GTEx database as a control. For missing values in the actual data, multiple imputations were performed using the R program to improve the dataset. To ensure the consistency of subsequent data analysis, the R package sva was used to normalize the integrated RNA-seq data, effectively removing technical variations introduced by different batches of experiments. Finally, by matching sample identifiers, a complete and high-quality training and validation dataset containing 504 LGG samples and 157 GBM samples was obtained.
[0024] By acquiring prognostic gene expression data and pathological variable data corresponding to glioma assessment samples, a systematic acquisition and integration of multi-source, heterogeneous biomedical data (genomic expression data and pathological data) was achieved, constructing a unified data foundation for subsequent analysis. This ensured the comprehensiveness and accuracy of the input information, including molecular features with specific prognostic indicative value extracted from massive genes (prognostic gene expression data) and validated key pathological indicators (pathological variable data), thus providing a foundation for constructing an information-rich and highly interpretable prognostic analysis framework.
[0025] Step S102: Determine the prognostic gene score of the prognostic gene expression data based on the pre-constructed prognostic scoring model.
[0026] It's important to note that the pre-built prognostic scoring model is a pre-trained mathematical model. Its core function is to convert specific gene expression levels (i.e., prognostic gene expression data) into a comprehensive quantitative prognostic gene score. This model was constructed through machine learning analysis of a large number of glioma training samples, and it internally includes mapping rules and parameters from gene expression levels to score values. The prognostic gene score is a continuous numerical result output by this model, comprehensively reflecting the prognostic risk information contained in the input gene expression profile. A higher score generally indicates a higher prognostic risk for the sample.
[0027] In this embodiment, the process involves calling and applying the pre-constructed prognostic scoring model for calculation. This prognostic scoring model is essentially a linear combination formula built based on multivariate regression analysis (e.g., Cox proportional hazards regression). The key parameters in the model are predetermined: specific prognostic feature genes (e.g., SOX3, MFAP2, CRLF1, etc.) and the corresponding regression coefficients (weights) for each gene. In application, the expression levels of these prognostic feature genes specified by the model are first extracted from the prognostic gene expression data. Then, the expression level of each gene is multiplied by its pre-trained regression coefficient, and the weighted results of all genes are summed to finally calculate a single prognostic gene score. This process is automated, requiring no manual intervention, ensuring the speed and consistency of the score calculation.
[0028] By determining the prognostic gene score based on pre-constructed prognostic scoring models, complex, high-dimensional gene expression profiles are compressed and transformed into a single, comparable quantitative indicator—the prognostic gene score. This process relies entirely on predefined mathematical models for calculation, completely eliminating subjective judgment differences that may exist in traditional pathological assessments. Furthermore, this score, based on a rigorously screened and validated gene set and its weights, can more accurately capture molecular biological characteristics closely related to disease progression and outcome, thus providing a stable quantitative basis for subsequent risk stratification and comprehensive analysis. Compared to qualitative or semi-quantitative judgments relying on single or a few biomarkers, this represents a significant improvement in the depth of information integration and the objectivity of prediction.
[0029] Step S103: Analyze the prognostic risk of glioma assessment samples based on prognostic gene scores and pathological variable data to obtain prognostic risk analysis results.
[0030] In this embodiment of the application, the prognostic risk of glioma assessment samples is analyzed based on prognostic gene scores and pathological variable data to obtain prognostic risk analysis results, including: determining the risk level group to which the glioma assessment samples belong based on the prognostic gene scores; substituting the prognostic gene scores and pathological variable data into a pre-constructed nomogram to obtain the prognostic status value of the glioma assessment samples at a specified time point; and generating the prognostic risk analysis results of the glioma assessment samples based on the risk level group and the prognostic status value.
[0031] It should be noted that risk grading refers to classifying glioma assessment samples into high-risk or low-risk categories based on a comparison between prognostic gene scores and a preset threshold (i.e., the optimal critical risk score). This preset threshold is determined in advance based on the state data of the training set samples using statistical methods (such as maximizing state differences). The pre-constructed nomogram is a graphical computational tool or mathematical model that integrates prognostic gene scores with multiple pathological variable data. Through preset contribution weights of each variable (derived from regression analysis), it can map specific input values to prognostic state values at specified time points (e.g., 1 year, 3 years, 5 years). The prognostic state value is a quantified probability value, representing the predicted prognosis of the glioma assessment sample at a specific future time point based on current input information (e.g., the probability of adverse events occurring). The prognostic risk analysis result is the final output; it is a structured conclusion that includes at least the risk grading of the samples and the quantified prognostic state value.
[0032] In this embodiment, firstly, the calculated prognostic gene score is compared with an optimal critical threshold (e.g., 0.18) determined beforehand through training set analysis. If the score is higher than this threshold, the assessed sample is classified as a high-risk sample; otherwise, it is classified as a low-risk sample. Secondly, the prognostic gene score of the same assessed sample and its corresponding pathological variable data (such as age, grade, IDH status, etc.) are used as input and substituted into a pre-constructed nomogram. The nomogram internally scores and summarizes each variable according to its preset weight, and outputs the prognostic status value of the sample at multiple specified time points (e.g., 1 year, 3 years, 5 years) through table lookup or calculation formula, i.e., the quantified risk probability. Finally, the risk level grouping (high risk / low risk) and the prognostic status values at each time point are integrated and formatted to generate easily interpretable prognostic risk analysis results.
[0033] By analyzing prognostic gene scores and pathological variables to assess the prognostic risk of glioma samples, this study represents a leap from simple quantitative scoring to comprehensive, structured decision-making information. Transforming continuous prognostic gene scores into discrete risk level groups (high / low risk) provides direct risk stratification, greatly facilitating the rapid understanding of core prognostic attributes. Introducing pre-constructed nomograms organically combines emerging molecular biomarker scores (prognostic gene scores) with classic pathological variables for synergistic analysis, thereby calculating more accurate prognostic status values (such as risk probabilities at multiple time points). This overcomes the limitations of relying solely on molecular models or traditional pathological indicators, enhancing predictive capabilities. The output prognostic risk analysis results are a comprehensive report integrating qualitative and quantitative, short-term and long-term predictive information, providing multi-dimensional data support for individualized response strategies and improving the practicality and accuracy of prognostic assessment.
[0034] In this embodiment of the application, the method for constructing the pre-built column chart includes: Step S201: Obtain the prognostic gene scores and corresponding pathological variable data of the glioma training samples.
[0035] It should be noted that glioma training samples refer to a set of glioma samples with known final states used to construct and train prognostic scoring and nomogram models. These samples are derived from historical cohorts (such as public databases), and their gene expression data and complete follow-up information (including overall state values) are known, used to learn patterns from the data and determine model parameters. Prognostic gene scoring here specifically refers to the score calculated for each sample in this batch of training samples using the pre-constructed prognostic scoring model. The corresponding pathological variable data refers to the pathological feature information matching the training samples.
[0036] In this embodiment, the acquisition and preprocessing of glioma training samples and their gene expression data and pathological variable data are similar to those for the evaluation samples, but the source is a specific training cohort (e.g., a training group randomly divided according to a certain proportion from glioma data in the TCGA database). Subsequently, the final prognostic scoring model is used to calculate the prognostic gene score for each sample in this training set. This process is retrospective, using the gene expression data of the training samples themselves and generating their corresponding scores using a predetermined model formula. Finally, the calculated prognostic gene score for each sample is associated and paired with the known pathological variable data (such as age, grade, IDH status, etc.) of that sample to form a structured dataset containing two columns of key information: a prognostic gene score column and a pathological variable data column (containing multiple variables). This dataset serves as the input data for constructing the nomogram.
[0037] Furthermore, by integrating CRPG scores with key clinicopathological variables (such as age, tumor grade, and IDH status), a visual nomogram was constructed using the rms software package. This tool can convert individual characteristics of samples into total scores, thereby directly predicting their status values at 1, 3, and 5 years. The nomogram demonstrated good calibration and predictive accuracy on both the internal test set and the external CGGA validation set (AUC values at all time points were above 0.70), proving its practical potential.
[0038] By acquiring prognostic gene scores and corresponding pathological variable data from glioma training samples, relevant input data was prepared for the subsequent construction of the nomogram model. This process ensured that the prognostic gene scores and pathological variable data used for modeling came from the same batch of samples, and that the scores were calculated based on the final determined prognostic model, guaranteeing the consistency of the data's internal logic. The paired dataset obtained in this way can realistically reflect the joint distribution relationship between molecular scores and pathological features in the training cohort, laying a solid and reliable data foundation for the next step of multivariate regression analysis to explore their common impact on prognostic status. This is an indispensable step in bridging the results of previous gene-level analysis with practical tools (nomograms).
[0039] Step S202: Perform regression analysis on prognostic gene scores and pathological variable data to generate a correlation model.
[0040] It should be noted that regression analysis is a multivariate statistical modeling method used to quantify the combined impact of prognostic gene scores and multiple pathological variables on the final state, and to establish mathematical correlations. For example, a Cox proportional hazards regression model can be used. This model can handle state data that includes time information, and its output is a mathematical model (i.e., a correlation model) containing regression coefficients (weights) for each input variable (prognostic gene score and each pathological variable), used to calculate the individual's hazard ratio or state value at a specific time point.
[0041] In this embodiment, the regression analysis is based on a prepared structured dataset. Each training sample in this dataset contains three parts of information: first, the prognostic gene score calculated by the prognostic scoring model; second, multiple pathological variable data matching the sample (such as age, WHO classification, IDH mutation status, 1p / 19q co-deletion status, MGMT promoter methylation status, etc.); and third, the sample's end-of-life status data, including follow-up time and event occurrence status (such as whether the person died). First, the prognostic gene score and all selected pathological variable data are used as independent variables (predictors), and the end-of-life status is used as the dependent variable to construct a Cox proportional hazards regression model. By fitting this model, the regression coefficients and their statistical significance (P-value) corresponding to each independent variable can be calculated. These coefficients reflect the magnitude and direction of the independent contribution of each variable to prognostic risk. The fitting process is usually completed using statistical computing software (such as the `survival` package in R). Finally, a mathematical equation that can be used for prediction, i.e., the correlation model, is obtained.
[0042] For example, the model can be represented as:
[0043] in, For risk function; Baseline risk function; ; For age; For classification; ; The regression coefficients for the age variable; The regression coefficients are for the tiered variables.
[0044] By performing regression analysis on prognostic gene scores and pathological variables, a correlation model was generated, enabling multi-dimensional integration and quantitative modeling of molecular markers (prognostic gene scores) and classic pathological factors. A rigorous statistical model (Cox regression) objectively quantified the independent prognostic impact of emerging gene features and traditional pathological variables, identifying variable combinations with significant predictive value. The generated correlation model is no longer a simple summation of single factors but considers potential synergistic or antagonistic effects among multiple factors, thus providing a more comprehensive and accurate characterization of the prognostic risk of gliomas. This model provides a solid mathematical foundation and specific weighting parameters for constructing user-friendly nomograms, making complex multivariate predictions visual and operational. This significantly improves the applicability and predictive accuracy of the prognostic assessment model, providing a reliable multi-factor decision-making basis for developing individualized response strategies.
[0045] Step S203: Based on the preset weights of each variable in the association model, construct a nomogram that maps prognostic gene scores and pathological variable data to prognostic state values.
[0046] It should be noted that the pre-set weights of each variable in the correlation model refer to the regression coefficients corresponding to each independent variable (i.e., prognostic gene scores and various pathological variable data) in the mathematical model generated through regression analysis (such as Cox regression). These coefficients are fixed after model training and become part of the model, used to quantify the contribution of each variable to prognostic risk. A nomogram is a graphical prediction tool built on an association model. It represents each variable in the model as a scaled line segment, assigns a score to each segment based on the user-inputted value, sums the scores of all variables to obtain a total score, and finally maps the total score directly to the prognostic state value (i.e., the predicted probability of the event occurring) at a specified time point (e.g., 1 year, 3 years, 5 years) using a total score-probability conversion scale.
[0047] In this embodiment of the application, the Cox proportional hazards regression model (i.e., the association model) obtained by fitting and the regression coefficients of its variables are used. The nomogram is constructed using specialized statistical plotting software packages (such as the rms package in R) with these coefficients as core parameters. The construction process includes: drawing independent axes for each predictor variable (prognostic gene score, age, tumor grade, etc.) included in the association model; determining the points corresponding to each value on the axis through linear transformation based on the variable's range and regression coefficient; variables with larger contribution weights (regression coefficients) have larger point changes per unit change; setting a total points axis in the nomogram, the range of which is the sum of all possible points for all variables; establishing the correspondence between the total points and prognostic state values at different time points (e.g., 1 year, 3 years, 5 years) based on the association model and a specified baseline state function, and plotting the corresponding probability scale. The final nomogram allows users to find the corresponding points on their respective axes based on the prognostic gene score and pathological variable data for a specific sample, sum them, locate them on the total points axis, and directly read the predicted probability (prognostic state value) of the sample at each specified time point from the probability scale.
[0048] As an example, a nomogram integrating CRPG scores and clinicopathological variables is shown in Figure 3 to estimate the 1-year, 3-year, and 5-year status values of glioma samples. Figure 3A The predictive performance of the nomogram in the training set, showing state values (e.g., AUC values) for 1 year, 3 years, and 5 years, are 0.778, 0.766, and 0.702, respectively. Figure 3B The test set values were 0.798, 0.787, and 0.737, respectively. Figure 3CFurthermore, the state values (such as AUC values) of the nomogram in the external validation set (CGGA) were 0.702, 0.749, and 0.756, respectively. Figure 3D The nomograms show accurate survival predictions. Calibration plots for the three cohorts (training, testing, and validation sets) demonstrate that the efficiency of the developed nomogram is comparable to that of the optimal model. Figure 3B -D).
[0049] By constructing a nomogram that maps prognostic gene scores and pathological variable data to prognostic status values, a complex multivariate statistical prediction model has been transformed into an intuitive and accurate decision support tool. Its core value lies in transforming abstract mathematical formulas and coefficients into clear graphs. Without needing to understand the underlying statistical principles or perform complex calculations, users can quickly obtain individualized quantitative prognostic predictions through simple graph lookup, improving the method's practicality and accessibility. This nomogram integrates prognostic gene scores representing emerging molecular characteristics with multiple key pathological variables within the same framework, visually demonstrating the relative contributions of each factor and ultimately outputting a personalized prognostic status value for a specific time point that integrates all information, achieving multidimensional individualized prognostic assessment. The graphical output facilitates subsequent human communication and provides a quantitative reference for developing individualized response and follow-up plans.
[0050] In this embodiment of the application, after determining the risk level group to which the glioma assessment sample belongs based on the prognostic gene score, the method further includes: Step S301: Obtain microenvironmental characteristic data and / or tumor mutation burden data for each risk level group.
[0051] It should be noted that tumor microenvironment (TME) characteristic data refers to a dataset that quantitatively describes the tumor microenvironment (TME) of glioma samples. It typically includes indicators inferred from gene expression data using computational biology algorithms (such as ESTIMATE and CIBERSORT), such as the Immune Score (reflecting the overall level of immune cell infiltration), the Stroma Score (reflecting the abundance of stromal components), and the relative infiltration proportions of various immune cells (such as CD8+ T cells, regulatory T cells, and tumor-associated macrophages). Tumor mutation burden data, on the other hand, is a quantitative indicator, usually referring to the Tumor Mutation Burden (TMB). It represents the total number of somatic mutations (including single nucleotide variants and small insertions / deletions) detected per million bases in the genome of a tumor sample, and is an important indicator for measuring tumor genomic instability and potential immunogenicity.
[0052] In this embodiment, after all samples have been divided into high-risk and low-risk groups based on prognostic gene scores, the microenvironmental characteristic data and tumor mutation burden data for each group are calculated separately. For the microenvironmental characteristic data, gene expression profiles of all samples within the group are first obtained. Then, the ESTIMATE algorithm is used to calculate the immune score and matrix score for each sample. Simultaneously, the CIBERSORT algorithm (or a similar deconvolution tool) is used to estimate the relative proportions of 22 (or more) immune cell subtypes in each sample based on a standardized gene expression feature matrix. By summarizing these indicators for all samples within a risk group, the microenvironmental characteristic profile of that group can be obtained. For the tumor mutation burden data, somatic mutation data (usually in mutation annotation format, MAF) of the samples within the group needs to be obtained. These data are processed using specialized analysis tools (such as the R package maftools), and the TMB value for each sample is calculated according to the standard formula (TMB = total number of mutations / total number of bases covered by sequencing × 10^6), thereby obtaining the TMB distribution of the group.
[0053] By acquiring microenvironmental characteristic data and / or tumor mutational burden data for each risk level group, and building upon prognostic risk stratification, this process further enables a deeper characterization and comparison of high-risk and low-risk sample groups at the fundamental level of tumor biology. This process links abstract risk level groupings with specific biological indicators (immune microenvironment, genomic mutations) reflecting the complex ecosystem within the tumor. By acquiring and comparing microenvironmental characteristic data and tumor mutational burden data from different risk groups, the potential biological mechanisms leading to prognostic differences can be revealed: for example, high-risk groups may exhibit immunosuppressive microenvironment characteristics (such as a higher proportion of regulatory T cells and a lower immune score) and higher genomic instability (higher TMB). This not only validates and enriches the biological significance of prognostic scores, making them more than just statistical predictive signals, but also provides direct and crucial data support and theoretical basis for subsequently developing differentiated response strategies (such as selecting immune strategies for specific microenvironments or screening samples that may benefit from immune checkpoint inhibitors based on TMB), enhancing the scientific depth and translational potential of the entire prognostic analysis system.
[0054] Step S302: Use microenvironmental characteristic data and / or mutation load data as biological characteristic data, and analyze the biological characteristic association between biological characteristic data and risk level grouping.
[0055] It should be noted that biological characteristic data specifically refers to the collective term for microenvironmental characteristic data and / or tumor mutation burden data. These data directly describe the immune status and genomic characteristics of tumors at the molecular level. Analyzing the associations of biological characteristics refers to using statistical methods to explore and quantify whether there is a significant and stable correspondence or dependency between the above-mentioned biological characteristic data and risk level groups (high-risk group and low-risk group) divided according to prognostic gene scores.
[0056] In this embodiment, firstly, biological characteristic data from all samples (including high-risk and low-risk groups) are integrated with corresponding risk level grouping labels. Then, each specific biological characteristic indicator (such as immune score, regulatory T cell ratio, TMB value, etc.) is analyzed separately. Common analytical methods include: for continuous data (such as immune score, TMB), nonparametric tests (such as the Mann-Whitney U test) are used to compare whether there is a statistically significant difference in the median of this indicator between the high-risk and low-risk groups. Similar comparisons can be performed for compositional data (such as the proportion of various immune cells). To quantify the association strength, the correlation coefficient (such as the Spearman rank correlation coefficient) between the prognostic gene score (as a continuous variable) and each biological characteristic data can be calculated, and its significance can be tested. Box plots are used to show the distribution differences of each indicator between the two groups, or scatter plots are drawn to show the trend relationship between the score and the biological characteristic. Through these analyses, it can be identified which biological characteristics show significant differences between the high- and low-risk groups, and the degree of their linear correlation with the prognostic score, thereby revealing the stable biological characteristic patterns behind the risk level grouping.
[0057] By analyzing the correlation between biological characteristics and risk stratification groups, a significant leap has been achieved from simple statistical risk prediction to mechanistic exploration and biological explanation. It goes beyond simply knowing whether a sample is high-risk or low-risk; it delves into the reasons for these risk differences. Establishing the correlation between biological characteristics and risk stratification groups reveals the specific tumor biological state corresponding to each prognostic risk stratification. For example, whether high risk is generally associated with an immunosuppressive microenvironment and higher genomic instability. This not only greatly enhances the biological credibility of the prognostic scoring model, grounding its conclusions in observable biology, but more importantly, it provides direct clues for developing targeted strategies. For instance, if high risk is clearly associated with specific immunosuppressive features, it suggests that this group may benefit from strategies to reverse immunosuppression, thus effectively guiding prognostic assessment results towards decision-making and achieving a closed loop from prediction to guidance.
[0058] Step S303: Generate prognostic risk analysis results based on risk level grouping, prognostic status values, and correlations of biological characteristics.
[0059] In this embodiment, firstly, the results generated for a specific sample are summarized: the risk level group of the sample and the prognostic status value at each time point. Then, based on the risk level group to which the sample belongs (e.g., high-risk group), typical biological feature descriptions of that risk level are retrieved and associated from the established biological feature association knowledge base. For example, this risk level is usually associated with an immunosuppressive microenvironment (high Treg cell proportion) and high tumor mutation burden (TMB). Finally, these three parts of information (grouping, status value, and feature association) are structurally integrated according to a preset template to generate standardized prognostic risk analysis results. These results can show that the sample is assessed as high-risk, with predicted probabilities of X%, Y%, and Z% at 1 year, 3 years, and 5 years, respectively. Furthermore, from a biological mechanism perspective, these high-risk samples typically exhibit specific microenvironmental or genomic characteristics, providing a reference for the selection of subsequent response strategies.
[0060] By generating prognostic risk analysis results based on risk level grouping, prognostic status values, and correlations with biological characteristics, a high-level synthesis and transformation of prognostic analysis information is achieved. Its core value lies in generating multi-dimensional decision support reports. Instead of providing fragmented risk scores or isolated probability values, it integrates qualitative risk stratification time-point probability predictions with explanatory biological mechanism clues into a logically rigorous and information-rich comprehensive analysis report, significantly enhancing the information content and practicality of the output results. The output prognostic risk analysis results not only answer the question of prognosis (risk level and status value) but also provide a preliminary explanation of the causes (correlation with biological characteristics). This combination allows for understanding the underlying biological basis while knowing the prediction results, thereby clarifying the direction of subsequent individualized responses (e.g., considering corresponding response strategies for indicated immunosuppression characteristics).
[0061] In this embodiment of the application, the method for constructing the prognostic scoring model includes: Step S401: Obtain gene expression data of glioma training samples.
[0062] It should be noted that glioma training samples refer to a set of glioma samples with known end states used to construct and train prognostic scoring models. These samples are derived from rigorously quality-controlled public genomic databases or research cohorts, characterized by the presence of complete gene expression data and corresponding long-term follow-up information (such as overall status values). Gene expression data here specifically refers to data obtained through high-throughput transcriptome sequencing technologies (such as RNA-seq) that quantify the transcriptional level of each gene in tumor samples, typically expressed as standardized gene expression values (such as transcripts per million reads, TPM), reflecting the biological state of the sample at the molecular level.
[0063] In this embodiment, firstly, authoritative public databases containing transcriptome data and information from glioma samples are identified and accessed, such as the Cancer Genome Atlas Database (TCGA) and the Chinese Glioma Genome Atlas Database (CGGA). Secondly, raw or processed RNA sequencing (RNA-seq) data files for specified sample sets, along with strictly matching annotation files, are downloaded in batches from these databases. For gene expression data, this typically involves downloading raw count data or expression matrices that have undergone standardization (e.g., TPM normalization). Next, the acquired raw data undergoes necessary preprocessing and quality control, including batch effect correction and normalization using specialized bioinformatics tools (e.g., the sva package in R) to ensure comparability of data from different batches or platforms. Finally, the processed gene expression data is precisely matched and integrated with actual data (e.g., sample IDs, state values, etc.) to form a training dataset, which serves as the input basis for all subsequent model construction.
[0064] By acquiring gene expression data from glioma training samples, a standardized data foundation was laid for the construction of a prognostic analysis system. This process ensured the scale and representativeness of the data: by utilizing large public databases, a sufficient number of glioma samples covering different subtypes and grades could be obtained, which enabled the subsequently trained model to have better statistical power and generalization ability; the rigorous download, preprocessing, and normalization process eliminated the interference of technical variations on the data, ensuring that the comparison of gene expression levels between different samples was reliable and accurate, which is a prerequisite for any subsequent bioinformatics analysis; by accurately matching gene expression profiles with detailed follow-up information, optimized materials were provided for subsequent supervised machine learning (such as state analysis and regression modeling), making it possible to analyze the correlation between gene expression patterns and prognosis from the data, which is the logical starting point from data to knowledge.
[0065] Step S402: Extract expression data of copper death-related genes from gene expression data, and perform cluster analysis based on the expression data of copper death-related genes to obtain multiple molecular subtypes.
[0066] In this embodiment of the application, cluster analysis is performed based on the expression data of copper death-related genes to obtain multiple molecular subtypes, including: performing consensus cluster analysis on the expression data of copper death-related genes to obtain the target number of clusters; and dividing the glioma training samples into multiple molecular subtypes according to the target number of clusters.
[0067] It should be noted that drug sensitivity analysis showed that the low CRPG score group had lower estimated half-maximal inhibitory concentrations (IC50) for chemotherapy drugs such as Elesclomol (a copper ion carrier), gefitinib, gemcitabine, shikonin, and camptothecin, suggesting greater sensitivity to these drugs. Meanwhile, the expression of immune checkpoint molecules (such as PD-1, PD-L1, and CTLA-4) was generally upregulated in the high-risk group, and this was associated with specific immunosuppressive microenvironment characteristics, providing a biological background for matching immune checkpoint inhibitor strategies to high-risk samples.
[0068] It should be noted that the expression data of copper death-related genes specifically refers to the expression levels of specific genes known to be involved in the copper-induced cell death (Cuproptosis) process in the training samples. This group of genes is based on predefined criteria from existing biological research (e.g., key genes such as FDX1, LIAS, LIPT1, DLAT, and DLD), and their expression levels reflect the activity status of the copper death pathway in the samples. Consensus clustering analysis is an unsupervised machine learning method that evaluates the stability of clustering results under different preset numbers of clusters (K values) by resampling and clustering the dataset multiple times, thereby determining the most reliable category division. Molecular subtypes refer to the multiple subsets of data with internal homogeneity and inter-group heterogeneity formed by dividing the glioma training samples based on their similarity in the expression patterns of copper death-related genes through the above clustering analysis.
[0069] Specifically, firstly, from the whole transcriptome expression matrix containing tens of thousands of genes, based on a known list of copper death-related genes (such as 10 core CRGs), the expression levels of these genes in all training samples are screened and extracted, forming a feature submatrix with significantly reduced dimensionality. Secondly, this feature submatrix is used as input, and a consensus clustering algorithm (e.g., implemented using the R package ConsensusClusterPlus) is applied. This process presets an exploration range for the K value (number of clusters) (e.g., 2 to 7). The algorithm iterates multiple times (e.g., 1000 times), randomly selecting a subset (e.g., 80%) from the samples each time for clustering (using a method such as PAM). Finally, by calculating the consensus matrix and evaluating metrics such as the consensus cumulative distribution function, the most reasonable target number of clusters (e.g., K=3) is determined. After determining this number, based on the copper death gene expression data of all training samples, the final clustering is performed, assigning all samples into the corresponding molecular subtypes (e.g., Cluster 1, Cluster 2, Cluster 3).
[0070] Cluster analysis based on expression data of copper death-related genes yielded multiple molecular subtypes, achieving efficient dimensionality reduction and pattern recognition for discovering biologically significant data from massive molecular datasets. By focusing on genes closely related to a specific biological process (copper death), a large amount of irrelevant transcriptomic noise was eliminated, achieving intelligent simplification of data dimensions and improving the efficiency and interpretability of subsequent analyses. Secondly, an unsupervised learning method, consensus clustering analysis, was employed to objectively classify seemingly homogeneous glioma samples into molecular subtypes with different intrinsic properties based on the similarity of gene expression patterns. This reveals potential heterogeneity at the level of copper death pathway activity, providing a new hierarchical perspective that traditional pathological classification cannot offer. This provides a clear comparative framework (i.e., comparisons between different subtypes) for identifying differentially expressed genes driving prognostic differences, and is the core of prognostic scoring model construction, moving from raw data to biological determination.
[0071] Step S403: Identify differentially expressed genes among different molecular subtypes and construct a prognostic scoring model based on the differentially expressed genes.
[0072] In this embodiment of the application, the prognostic scoring model based on differentially expressed genes includes: screening out prognostic candidate genes related to the assessment time point of the glioma training sample from differentially expressed genes; performing feature selection on the prognostic candidate genes to obtain a set of prognostic feature genes; and constructing a prognostic scoring model based on the expression level of each gene in the set of prognostic feature genes.
[0073] It should be noted that differentially expressed genes refer to genes whose expression levels differ statistically significantly among different molecular subtypes. These genes reflect key biological heterogeneity among subtypes. Prognostic candidate genes are genes whose expression levels are significantly correlated with the predicted state of the glioma training samples (reflected by assessment time-point data, such as the overall state value) after further screening from differentially expressed genes. The prognostic feature gene set is the core gene combination used to construct the final predictive model, obtained through secondary refinement from the prognostic candidate genes using machine learning algorithms; its number is far less than the initial differentially expressed genes. The prognostic scoring model is a calculation formula used to calculate the prognostic gene score, constructed based on the expression levels of each gene in this prognostic feature gene set and their predetermined weight coefficients through linear combination or other mathematical forms.
[0074] Specifically, firstly, based on the defined molecular subtypes, statistical methods (such as the limma R package) are used to systematically compare gene expression profiles among different subtypes. Using both fold change in expression level and statistical significance (p-value) as dual criteria, a set of genes exhibiting significant differences in expression among subtypes is selected. Secondly, prognostic candidate genes are screened: the expression data of the differentially expressed genes are combined with the state data (follow-up time and event state) of the training samples, and a univariate state analysis (such as Cox regression) is performed on each gene. Genes whose expression levels are significantly correlated with state values (p-value less than a threshold, such as 0.05) are retained, forming a list of prognostic candidate genes. Next, to avoid overfitting and identify the most informative genes, machine learning algorithms (such as random survival forest) are used to evaluate the importance of the prognostic candidate genes. Based on the evaluation results (such as variable importance scores), the few most important genes are selected to constitute the final set of prognostic feature genes (e.g., 11 genes). Finally, using this carefully selected set of prognostic feature genes as variables and their expression levels in the training set as input, multivariate regression analysis (such as multivariate Cox regression) is used to determine the independent regression coefficients (weights) of each gene. Based on these coefficients and gene expression levels, a linear combination formula is constructed, which is the final prognostic scoring model, and its calculation result is the prognostic gene score.
[0075] As an example, the prognostic value of 10 copper death-related genes (CRGs) was assessed before clustering. The TCGA cohort samples were divided into high- and low-expression groups based on the median expression of each CRG, and analysis using Kaplan-Meier state curves and log-rank tests revealed that, except for MTF1 and GLS, the expression levels of the remaining CRGs were significantly correlated with the total sample size. This preliminarily confirms the prognostic significance of CRGs in gliomas.
[0076] The three molecular subtypes (Cluster 1, 2, 3) identified through consensus clustering (using the ConsensusClusterPlus tool, parameters: maxK=7, 1000 replicates) exhibited significant biological and clinical differences. Cluster 2 showed the worst prognosis (median overall survival (OS) of 21.0 months) and was significantly associated with higher tumor grade, IDH wild-type status, MGMT promoter unmethylation, and non-co-deletion of 1p / 19q, among other adverse clinicopathological features. This reveals a close association between CRG-based molecular subtyping and aberrant aggressiveness.
[0077] In-depth bioinformatics analysis was performed on 1072 differentially expressed genes (DEGs) identified between Cluster 2 and other subtypes. Gene set enrichment analysis (GSVA) and KEGG pathway analysis revealed that Cluster 2 is significantly activated in immune regulation and inflammatory response pathways (such as cytokine-cytokine receptor interactions, antigen processing and presentation, and Toll-like receptor signaling pathways), while its basal energy metabolism pathways (such as the TCA cycle and oxidative phosphorylation) are inhibited. Protein-protein interaction (PPI) network analysis further confirmed the strong association between these DEGs and immune and inflammatory responses.
[0078] Univariate Cox regression analysis was performed on the above 1072 DEGs, and 90 prognostic candidate genes that were significantly associated with the overall state value were screened out (p<0.05).
[0079] The TCGA training set (n=465) and test set (n=196) were split in a 7:3 ratio. First, the random survival forest algorithm was used on the training set to evaluate the feature importance of 90 candidate genes, identifying 27 key genes with importance greater than 0. Then, through multivariate Cox proportional hazards regression analysis, 11 genes with independent prognostic value were finally identified from these 27 genes, forming the prognostic feature gene set: SOX3, MFAP2, CRLF1, NEFM, EN2, EGR1, ICAM5, MSMP, RGS4, SLC17A7, and NRGN. Based on the expression levels and regression coefficients of these genes, a linear CRPG prognostic scoring model formula was constructed.
[0080] In both the training set and the independent external validation set (CGGA cohort), the CRPG score demonstrated good prognostic discriminative ability. Kaplan-Meier analysis showed that the state values of the high-risk score group were significantly shorter than those of the low-risk group. Time-dependent ROC curve analysis revealed that the score had high AUC values in predicting state values at 1, 3, and 5 years (e.g., 0.785, 0.774, and 0.705 in the training set, respectively).
[0081] Analysis revealed that the CRPG high-risk score was significantly positively correlated with a higher tumor mutation burden (TMB) and negatively correlated with the cancer stem cell (CSC) index.
[0082] By constructing a prognostic scoring model based on differentially expressed genes, an automated process has been achieved to objectively extract and quantify key prognostic information from massive gene datasets. It no longer relies on prior knowledge or subjective selection of individual genes, but rather uses a progressive analysis—including differential comparison, state association, and machine learning screening—driven by the data itself, to gradually identify core gene combinations that can distinguish molecular subtypes, are strongly correlated with the final state, and have complementary predictive value. This process enhances the scientific rigor and robustness of the model construction. Ultimately, the prognostic scoring model constructed based on the set of prognostic feature genes can accurately transform high-dimensional and complex transcriptome information into quantitative scores with clear prognostic significance. This provides a core tool for subsequent applications that is both easy to operate (requiring only the detection of a small number of genes) and possesses high predictive performance.
[0083] As an example, such as Figure 6 As shown, starting with glioma RNA sequencing and somatic mutation data from the TCGA database, copy number variation (CNV) and somatic mutation analysis were first performed. Then, the samples were divided into three molecular subtypes—Cluster1, Cluster2, and Cluster3—using consistent clustering. For each subtype, gene set variation analysis (GSVA), CIBERSORT immune infiltration prediction, and ESTIMATE tumor microenvironment (TME) assessment were conducted. Differentially expressed genes were identified among the different subtypes, and gene set enrichment analysis (GSEA) was performed on these differentially expressed genes, including functional annotations from Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG). Next, key prognostic genes were screened using univariate Cox regression combined with random forest analysis. A prognostic model was constructed using multivariate Cox regression, and the samples were divided into high- and low-risk groups based on the optimal cutoff value. The model efficacy was further validated using receiver operating characteristic (ROC) curves and survival curves, and the mutation characteristics, immune status, and drug sensitivity of different risk groups were analyzed. Finally, a nomogram was constructed, and the model was validated using an internal validation set and an external CGGA test set.
[0084] This embodiment also provides a device for analyzing the prognostic gene characteristics of gliomas based on machine learning. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0085] This embodiment provides a device for analyzing the prognostic gene characteristics of gliomas based on machine learning, such as... Figure 7 As shown, it includes: The acquisition module 71 is used to acquire prognostic gene expression data and pathological variable data corresponding to the glioma assessment samples; The calculation module 72 is used to calculate the prognostic gene score of the prognostic gene expression data based on the pre-constructed prognostic scoring model. Analysis module 73 is used to analyze the prognostic risk of glioma assessment samples based on prognostic gene scores and pathological variable data, and obtain prognostic risk analysis results.
[0086] In this embodiment of the application, the analysis module 73 is specifically used to determine the risk level group to which the glioma assessment sample belongs based on the prognostic gene score; substitute the prognostic gene score and pathological variable data into a pre-constructed nomogram to obtain the prognostic status value of the glioma assessment sample at a specified time point; and generate the prognostic risk analysis result of the glioma assessment sample based on the risk level group and the prognostic status value.
[0087] In this embodiment of the application, the device further includes: a first construction module, used to acquire the prognostic gene score and corresponding pathological variable data of the glioma training sample; perform regression analysis on the prognostic gene score and pathological variable data to generate an association model; and construct a nomogram that maps the prognostic gene score and pathological variable data to prognostic state values according to the preset weights of each variable in the association model.
[0088] In this embodiment of the application, the device further includes: a generation module, used to acquire microenvironmental characteristic data and / or tumor mutation burden data for each risk level group; use the microenvironmental characteristic data and / or mutation burden data as biological characteristic data, analyze the biological characteristic association between the biological characteristic data and the risk level group; and generate prognostic risk analysis results based on the risk level group, prognostic status value, and biological characteristic association.
[0089] In this embodiment of the application, the device further includes: a second construction module, used to acquire gene expression data of glioma training samples; extract expression data of copper death-related genes from the gene expression data, and perform cluster analysis based on the expression data of copper death-related genes to obtain multiple molecular subtypes; identify differentially expressed genes among different molecular subtypes, and construct a prognostic scoring model based on the differentially expressed genes.
[0090] In this embodiment of the application, the second construction module is specifically used to perform consensus clustering analysis on the expression data of copper death-related genes to obtain the target number of clusters; and to divide the glioma training samples into multiple molecular subtypes according to the target number of clusters.
[0091] In this embodiment of the application, the second construction module is specifically used to screen out prognostic candidate genes related to the evaluation time point of the glioma training sample from differentially expressed genes; perform feature selection on the prognostic candidate genes to obtain a set of prognostic feature genes; and construct a prognostic scoring model based on the expression level of each gene in the set of prognostic feature genes.
[0092] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 8 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system).
[0093] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0094] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0095] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0096] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0097] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0098] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0099] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for analyzing the prognostic gene characteristics of glioma based on machine learning, characterized in that, The method includes: Obtain prognostic gene expression data and pathological variable data corresponding to glioma assessment samples; The prognostic gene score of the prognostic gene expression data is determined based on a pre-constructed prognostic scoring model. Based on the prognostic gene scores and the pathological variable data, the prognostic risk of the glioma assessment samples was analyzed to obtain the prognostic risk analysis results.
2. The method according to claim 1, characterized in that, The prognostic risk analysis of the glioma assessment sample based on the prognostic gene score and the pathological variable data yields the prognostic risk analysis results, including: The risk level group to which the glioma assessment sample belongs is determined based on the prognostic gene score; Substituting the prognostic gene scores and the pathological variable data into a pre-constructed nomogram, the prognostic status value of the glioma assessment sample at a specified time point is obtained. The prognostic risk analysis results of the glioma assessment samples are generated based on the risk level grouping and the prognostic status value.
3. The method according to claim 2, characterized in that, The method for constructing the pre-built nomogram includes: Obtain prognostic gene scores and corresponding pathological variable data for glioma training samples; Regression analysis was performed on the prognostic gene scores and the pathological variable data to generate a correlation model; Based on the preset weights of each variable in the association model, a nomogram is constructed that maps the prognostic gene score and the pathological variable data to the prognostic state value.
4. The method according to claim 1, characterized in that, After determining the risk group to which the glioma assessment sample belongs based on the prognostic gene score, the method further includes: Obtain microenvironmental characteristic data and / or tumor mutation burden data for each of the aforementioned risk level groups; The microenvironmental characteristic data and / or the mutation load data are used as biological characteristic data to analyze the biological characteristic association between the biological characteristic data and the risk level grouping; The prognostic risk analysis results are generated based on the risk level grouping, the prognostic status value, and the correlation of the biological characteristics.
5. The method according to claim 1, characterized in that, The method for constructing the prognostic scoring model includes: Obtain gene expression data from glioma training samples; Expression data of copper death-related genes were extracted from the gene expression data, and cluster analysis was performed based on the expression data of copper death-related genes to obtain multiple molecular subtypes; Differentially expressed genes among different molecular subtypes are identified, and the prognostic scoring model is constructed based on the differentially expressed genes.
6. The method according to claim 5, characterized in that, Cluster analysis based on the expression data of the copper death-related genes yielded multiple molecular subtypes, including: Consensus clustering analysis was performed on the expression data of the copper death-related genes to obtain the number of target clusters; The glioma training samples are divided into multiple molecular subtypes based on the target cluster number.
7. The method according to claim 5, characterized in that, The construction of the prognostic scoring model based on the differentially expressed genes includes: From the differentially expressed genes, prognostic candidate genes related to the assessment time point of the glioma training samples were screened; Feature selection is performed on the candidate prognostic genes to obtain a set of prognostic feature genes; The prognostic scoring model is constructed based on the expression levels of each gene in the prognostic feature gene set.
8. A device for analyzing the prognostic gene characteristics of glioma based on machine learning, characterized in that, The device includes: The acquisition module is used to acquire prognostic gene expression data and pathological variable data corresponding to glioma assessment samples; A calculation module is used to calculate the prognostic gene score of the prognostic gene expression data based on a pre-constructed prognostic scoring model. The analysis module is used to analyze the prognostic risk of the glioma assessment sample based on the prognostic gene score and the pathological variable data, and obtain the prognostic risk analysis results.
9. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.