Prediction model for temozolomide drug resistance of MGMT negative glioblastoma patient
By integrating cell line and clinical cohort data, genes such as OLR1, ULBP2, KCNK5, BDKRB2, and MUS81 were screened and a nomogram model was constructed. This solved the problem that a single biomarker could not accurately predict temozolomide resistance in MGMT-negative glioblastoma patients in existing technologies, and enabled more accurate resistance risk assessment and personalized treatment plans.
Patent Information
- Application Number
- CN202511176724.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-05
AI Technical Summary
Current technologies rely on a single biomarker, the methylation status of the MGMT promoter, which cannot accurately predict the risk of temozolomide resistance in MGMT-negative glioblastoma patients and cannot fully reflect the complexity of GBM resistance to TMZ.
By integrating in vitro drug resistance data from cell lines and survival prognostic information from clinical cohorts, candidate genes such as OLR1, ULBP2, KCNK5, BDKRB2, and MUS81 were screened using Lasso and multivariate Cox regression analysis. A multigene prediction model was constructed, and a nomogram model was formed to predict the drug resistance risk of patients.
It improves the predictive accuracy of temozolomide resistance risk in patients with MGMT-negative glioblastoma, provides an objective basis for early resistance risk stratification and personalized treatment plans, and the screened genes have potential drug target value.
Smart Images

Figure CN121075434A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tumor prognosis prediction technology, specifically a predictive model for temozolomide resistance in patients with MGMT-negative glioblastoma. Background Technology
[0002] Glioblastoma (GBM) is a highly malignant central nervous system tumor. Currently, surgical resection combined with temozolomide (TMZ) chemoradiotherapy is the standard treatment. Clinical practice has shown that the methylation status of the promoter region of the O-6-methylguanine-DNA methyltransferase (MGMT) gene is a core biomarker for predicting TMZ efficacy and patient prognosis. To clarify the terminology used in this invention, it should be noted that MGMT promoter methylation is a molecular-level cause that leads to transcriptional silencing of the MGMT gene, thereby inhibiting the expression of functional MGMT protein. In clinicopathological testing, the absence or significant downregulation of this protein expression is defined as MGMT negative. Therefore, in this invention and related technical fields, patients with MGMT promoter methylation correspond to MGMT negative patients. Typically, patients with MGMT promoter methylation exhibit higher drug sensitivity and longer survival because their MGMT protein expression is suppressed, resulting in a decreased ability of tumor cells to repair alkylation damage from TMZ.
[0003] However, in clinical practice, the strategy of relying solely on MGMT promoter methylation status for efficacy prediction has revealed significant limitations. A key clinical paradox is that a considerable proportion (approximately 30%-50%) of patients diagnosed with MGMT promoter methylation (i.e., MGMT-negative) should be potential beneficiaries of TMZ treatment, but instead develop acquired resistance during treatment, severely limiting their prognosis. This phenomenon clearly demonstrates that traditional predictive systems based on a single static biomarker cannot comprehensively reflect the complexity of GBM resistance to TMZ.
[0004] Further research has confirmed that TMZ resistance in GBM is not determined by a single gene or pathway, but rather by the dynamic action of a multi-level, multi-dimensional molecular regulatory network. Existing technologies have revealed that certain signaling pathways can bypass or compensate for the functional state of MGMT. For example, even with methylation of the MGMT promoter, signaling pathways such as Hedgehog can still directly activate MGMT gene expression through downstream transcription factors, thereby restoring the DNA damage repair capacity of tumor cells. Furthermore, the regulatory role of non-coding RNAs, such as certain long non-coding RNAs (lncRNAs), can indirectly upregulate MGMT protein levels by competitively binding to microRNA (miRNA)-associated proteins.
[0005] Tumor cells can also develop drug resistance mechanisms that are completely independent of MGMT. For example, by reprogramming lipid metabolism, tumor cells can produce specific antioxidant metabolites that can effectively neutralize TMZ-induced intracellular DNA damage, thereby reducing the killing effect of drugs. These findings all point to the conclusion that tumor drug resistance mechanisms are dynamic and heterogeneous, and that the gene expression profile at the transcriptome level contains key information revealing the core regulatory signals of drug resistance. However, most existing studies focus on a single isolated signaling pathway or metabolic pathway, lacking a systematic integrated analysis of multidimensional transcriptional regulatory networks, and therefore cannot construct predictive models that can accurately capture tumor heterogeneity and dynamic evolution during treatment.
[0006] Therefore, this invention proposes a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients to address the shortcomings of existing technologies. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients, solving the problem that relying on a single biomarker (such as the methylation status of the MGMT promoter) cannot accurately predict the risk of temozolomide resistance in MGMT-negative glioblastoma patients.
[0008] To address the aforementioned technical problems, this invention provides corresponding technical solutions.
[0009] The first aspect of this invention provides a method for constructing a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients, the method comprising the following steps:
[0010] S1. Obtain gene expression profile data of temozolomide-sensitive and resistant glioblastoma cell lines, and screen out differentially expressed genes through differential analysis.
[0011] S2. Obtain transcriptomic data and clinical prognostic information of the MGMT-negative glioblastoma patient cohort in the TCGA database, and select clinical samples from them that are MGMT-negative primary and have received standard TMZ chemotherapy and contain complete PFS and OS data, which are called the TCGA target cohort.
[0012] S3. Perform survival analysis on the differentially expressed genes screened in step S1 in the TCGA target cohort of step S2 to screen out core prognostic genes that are significantly associated with both progression-free survival (PFS) and overall survival (OS).
[0013] S4. Perform Lasso regression analysis and multivariate Cox regression analysis on the core prognostic genes screened in step S3 to determine candidate genes for constructing the prediction model.
[0014] S5. Based on the candidate genes determined in step S4 and their regression coefficients in multivariate Cox regression analysis, construct the prediction model.
[0015] In one specific embodiment, in step S1, the MGMT-negative glioblastoma cell line is the U251 cell line, and the gene expression profile data is derived from the GSE100736 dataset in the GEO database.
[0016] In one specific embodiment, in step S1, the differential analysis employs the limma algorithm, and the threshold for screening differentially expressed genes is:
[0017] P value < 0.05 and |Log FC| > 1;
[0018] Wherein, P value is the probability value used to determine the statistical significance of the difference in gene expression levels between susceptible and resistant strains; Log FC is the logarithm of the fold change in gene expression level between resistant and susceptible strains.
[0019] In one specific embodiment, the survival analysis in step S3 specifically includes:
[0020] S3-1. Perform univariate Cox regression analysis on the differentially expressed genes in relation to progression-free survival (PFS), and screen genes with a P-value < 0.05 as the threshold.
[0021] S3-2. For the genes screened in step S3-1, Kaplan-Meier analysis was performed based on their median expression level, which was associated with both progression-free survival (PFS) and overall survival (OS), to identify core prognostic genes with a P-value < 0.05 in both PFS and OS analyses.
[0022] In one specific embodiment, the candidate genes determined in step S4 are: OLR1, ULBP2, KCNK5, BDKRB2, and MUS81.
[0023] In one specific embodiment, step S5, which involves constructing the prediction model, includes:
[0024] Based on the expression levels of the identified candidate genes OLR1, ULBP2, KCNK5, BDKRB2, and MUS81 and their corresponding regression coefficients in multivariate Cox regression analysis, the risk score was calculated using the following formula:
[0025]
[0026] Where n is the number of candidate genes, i is the index of a single candidate gene, and β iExp represents the regression coefficient of candidate gene i in multivariate Cox regression analysis. i The expression level of candidate gene i.
[0027] Preferably, the prediction model is a nomogram model, which can predict the 1-year, 2-year, and 3-year progression-free survival rates of patients based on the risk score.
[0028] Preferably, the MGMT-negative glioblastoma patient is a primary glioblastoma patient with MGMT promoter methylation.
[0029] A second aspect of this invention provides a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients, obtained by any of the aforementioned construction methods, wherein the predictive model comprises a gene signature consisting of the following candidate genes:
[0030] OLR1, ULBP2, KCNK5, BDKRB2 and MUS81;
[0031] The prediction model is used to predict the risk of temozolomide resistance in patients based on the expression level of the gene signature in their tumor tissue.
[0032] In one specific embodiment, the model is a nomogram model, which converts the expression level of each gene in the gene signature into a risk score, and predicts the 1-year, 2-year and 3-year progression-free survival rates of patients based on the risk score.
[0033] This invention provides a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients.
[0034] It has the following beneficial effects:
[0035] 1. This invention integrates in vitro drug resistance data from cell lines and survival prognostic information from clinical cohorts, and utilizes Lasso and multivariate Cox regression analysis to determine a multi-gene predictive model consisting of OLR1, ULBP2, KCNK5, BDKRB2, and MUS81. Compared to existing technologies that rely solely on a single MGMT status for assessment, this invention comprehensively evaluates multiple genes closely related to drug resistance relapse, providing a more complete reflection of the biological mechanisms of drug resistance and thus improving the accuracy of predicting the risk of temozolomide resistance in MGMT-negative glioblastoma patients.
[0036] 2. The predictive model ultimately constructed in this invention, particularly in the form of a nomogram, can transform complex gene expression data into an intuitive risk score and survival probability associated with progression-free survival (e.g., 1-year, 2-year, and 3-year survival rates). This provides clinicians with an objective basis for stratifying the drug resistance risk of MGMT-negative patients early in treatment or before treatment, helping to plan alternative or adjuvant treatment options for patients predicted to have high drug resistance risk as early as possible, avoiding the delay in discovering drug resistance only in the later stages of traditional treatment.
[0037] 3. This invention not only provides a predictive tool, but the core candidate genes OLR1, ULBP2, KCNK5, BDKRB2, and MUS81 identified in the screening process themselves possess clear technical value. Since these genes were obtained through a screening process directly associated with temozolomide resistance and patient survival prognosis, they constitute potential drug targets. Developing combination therapies targeting these specific genes or their regulated signaling pathways provides a new technical approach and theoretical basis for overcoming temozolomide resistance. Attached Figure Description
[0038] Figure 1 Volcano plots and thermograms for differential analysis of the U251 temozolomide sensitive and resistant groups in GSE100736 of this invention;
[0039] Figure 2 This is a schematic diagram of the lasso regression analysis related to the PFS of the TCGA target queue in this invention;
[0040] Figure 3 This is a flowchart of the screening process for candidate genes related to TMZ resistance and relapse in MGMT-negative glioblastoma.
[0041] Figure 4 This is a schematic diagram of the Nomogram of the present invention;
[0042] Figure 5 This is a schematic diagram of the ROC curves for the 1-year and 3-year progression-free survival rates of patients according to the present invention. Detailed Implementation
[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] See attached document Figure 1 To be continued Figure 5This invention provides a specific implementation method for constructing a predictive model for temozolomide resistance in MGMT-negative glioblastoma patients. This implementation method acquires gene expression profile data from publicly available databases, and through a series of bioinformatics analyses and statistical screening steps, ultimately constructs a predictive model composed of specific candidate genes.
[0045] The method in this embodiment begins with obtaining gene expression profile data from publicly available databases. Specifically, the gene expression profile dataset GSE100736, containing temozolomide (TMZ)-sensitive and resistant strains of the glioblastoma cell line U251, is obtained from the GEO database. Transcriptome data and corresponding clinical prognostic information, including progression-free survival (PFS) and overall survival (OS), from a cohort of patients with primary glioblastoma who are methylated by the MGMT promoter, are obtained from the TCGA database. Clinical samples with complete PFS and / or OS data from MGMT-negative primary glioblastoma who have received standard TMZ chemotherapy are selected and referred to as the TCGA target cohort.
[0046] After obtaining the data, differential gene analysis was first performed on the cell line dataset GSE100736. (Refer to...) Figure 3 The first step of the process shown is to use the limma algorithm to compare the mRNA expression profiles of TMZ-sensitive and resistant strains, and to screen differentially expressed genes with a P value < 0.05 and |Log FC| > 1 as thresholds. Figure 1 The volcano map and heatmap of some differentially expressed genes obtained from this step are shown.
[0047] Next, the differentially expressed genes identified in the previous step were applied to the TCGA target cohort data for multiple rounds of survival analysis to screen for core prognostic genes. (Refer to...) Figure 3 The second step of the procedure is to first perform univariate Cox regression analysis on all differentially expressed genes related to PFS, and screen them with P value < 0.05 (P < 0.05) as the criterion; then, Kaplan-Meier analysis on PFS and OS is performed on the screened genes to screen out the core prognostic genes that are significantly associated with both PFS and OS.
[0048] To further streamline the number of genes and build a model, machine learning analysis was performed on the selected core prognostic genes. (Refer to...) Figure 3 The third step in the process shown is to first perform Lasso regression analysis. Figure 2 The analysis process was demonstrated, which used cross-validation to determine the optimal parameters and compressed the regression coefficients of some genes to zero, thereby screening out nine genes. Subsequently, these nine genes were included in a multivariate Cox regression analysis, which ultimately identified five independent candidate genes: OLR1, ULBP2, KCNK5, BDKRB2, and MUS81.
[0049] Finally, the risk scoring model was visualized as a nomogram, as shown in the attached figure. Figure 4 As shown in the figure, this nomogram integrates the expression levels of five candidate genes and calculates a total score to predict 1-year, 2-year, and 3-year survival rates associated with progression-free survival. To evaluate the predictive performance of the model, time-dependent ROC curves for its predictions of 1-year and 3-year survival rates were plotted, as shown in the attached figure. Figure 5 As shown.
[0050] To further illustrate the technical solution of the present invention, the specific construction process of the prediction model will be described in detail below.
[0051] Those skilled in the art will recognize that the analysis can be performed using statistical software. In this embodiment, all data analysis and statistical modeling were performed in an R language environment (version 4.0.0 or higher). The limma algorithm was implemented using the limma package, univariate and multivariate Cox regression analysis was implemented using the survival package, Lasso regression analysis was implemented using the glmnet package, nomogram construction was completed using the rms package, and ROC curve plotting and AUC value calculation were completed using the timeROC package. These software packages are all publicly available standard tools.
[0052] The initial step of this process is data acquisition and preprocessing. The data in this embodiment comes from two independent, publicly available databases to obtain gene expression information from in vitro drug-resistant cell lines and gene expression and prognostic information from clinical patient cohorts, respectively.
[0053] First, cell line gene expression profiles were obtained for initial differentially expressed genes screening. The dataset with accession number GSE100736 was retrieved from the Gene Expression Omnibus Database (GEO) of the National Center for Biotechnology Information (NCBI). This dataset contains gene expression profiles of the human glioblastoma cell line U251 before and after in vitro induction of temozolomide (TMZ) resistance. Specifically, the dataset includes two experimental groups: a TMZ-sensitive group (three biological replicates) and a TMZ-resistant group (three biological replicates). After obtaining the raw data, it was standardized and preprocessed to correct and eliminate non-biological technical biases between samples, ensuring the comparability of gene expression levels between different samples and providing a basis for subsequent differential expression analysis.
[0054] Secondly, clinical cohort data were acquired for model building and validation. Relevant data from the glioblastoma (TCGA-GBM) project were downloaded from the Cancer Genome Atlas (TCGA) database. To build an accurate model for a specific patient population, the acquired patient samples underwent rigorous screening, with the following inclusion criteria:
[0055] 1. Pathological diagnosis: primary glioblastoma;
[0056] 2. The test results show a clear MGMT promoter methylation status, and the result is promoter methylation (i.e., MGMT negative);
[0057] 3. Possesses complete transcriptome sequencing data;
[0058] 4. Possess complete and usable clinical follow-up information, which must include the time and event status of progression-free survival (PFS) and overall survival (OS).
[0059] After screening, patient data meeting all the above criteria were preprocessed. For clinical information, PFS and OS data were formatted into the standard format required for survival analysis. For transcriptome expression data, to meet the data distribution requirements of subsequent statistical analysis models, the quantitative gene expression values (such as TPM or FPKM values) were logarithmically transformed, for example, using log2(x+1), where x is the original quantitative expression value. Logarithmic transformation aims to make the skewed gene expression data closer to a normal distribution, satisfying the basic assumptions of various subsequent statistical models regarding data distribution, thereby ensuring the accuracy and reliability of the analysis results. After preprocessing, a standardized cohort dataset of MGMT-negative glioblastoma patients containing gene expression profiles and corresponding clinical prognostic information was obtained.
[0060] See attached document Figure 1 and appendix Figure 3 After data acquisition and preprocessing, the method in this embodiment proceeds to the differential gene screening step based on cell line gene expression profile data. The purpose of this step is to identify a set of genes directly associated with the formation of temozolomide (TMZ) resistance in vitro from the whole genome expression profile.
[0061] The input for this step is the gene expression matrix of the preprocessed GSE100736 dataset, which contains the standardized gene expression levels of three samples each from the TMZ-sensitive and TMZ-resistant groups. The analysis employs the limma (Linear-Models-for-Microarray-and-RNA-Seq-Data) algorithm based on linear models. This algorithm fits a linear model to the expression data of each gene to assess the expression differences between the two groups (sensitive and resistant strains). This method corrects for gene variance through an empirical Bayesian adjustment step, thereby obtaining more stable statistical inference results even with a small sample size.
[0062] Using the limma algorithm, two core statistical indicators were calculated for each gene: log fold change (Log FC) and statistical significance P-value.
[0063] Log FC is the base-2 logarithm of the fold change in gene expression level between TMZ-resistant and TMZ-sensitive strains, reflecting the magnitude and direction of the gene expression difference.
[0064] The P-value is the probability value used to determine whether the difference in gene expression levels between two groups is statistically significant. The smaller the value, the lower the probability that the difference is caused by random factors.
[0065] To screen for differentially expressed genes with both biological and statistical significance, this embodiment sets a dual screening threshold: P-value < 0.05 and |Log FC| > 1. A gene must meet both conditions simultaneously to be defined as a differentially expressed gene. The former ensures statistical significance of the gene expression change, while the latter ensures that the change in gene expression level is more than 2-fold. All genes meeting the conditions constitute a list of differentially expressed genes, which will serve as input for subsequent clinical cohort analyses.
[0066] Appendix Figure 1 The volcano plot and heatmap visually illustrate the results of this screening step. The volcano plot uses LogFC on the x-axis and -log10 (P value) on the y-axis; the points in the upper left and upper right corners represent the differentially expressed genes that were downregulated and upregulated. The heatmap shows the expression patterns of the selected differentially expressed genes across all six samples, clearly demonstrating significantly different gene expression profiles between the TMZ-sensitive and resistant groups. These results provide a data foundation for subsequent investigation of drug resistance biomarkers related to patient prognosis in clinical samples.
[0067] See attached document Figure 3 After obtaining the list of differentially expressed genes associated with in vitro drug resistance in the previous step, the method in this embodiment proceeds to the core prognostic gene screening step based on clinical cohort data. The purpose of this step is to test the genes screened in in vitro experiments in a real clinical setting, identifying genes that not only correlate with in vitro drug resistance but also show statistically significant associations with the actual survival prognosis (including progression-free survival (PFS) and overall survival (OS)) of patients with MGMT-negative glioblastoma. This screening process includes two consecutive statistical analysis phases.
[0068] The first stage was a univariate Cox regression analysis. This analysis, based on preprocessed TCGA patient cohort data, used the expression level of each gene in the differentially expressed gene list obtained in the previous step as an independent covariate, and performed association analysis with the patients' progression-free survival (PFS) data. For each gene, the univariate Cox proportional hazards model calculated a hazard ratio (HR) and its corresponding p-value. The HR value reflects the change in the risk of disease progression events for each unit increase in gene expression level; the p-value is used to assess the statistical significance of this association. In this embodiment, p < 0.05 was set as the screening threshold, meaning only genes whose expression levels were statistically significantly associated with PFS were retained, forming a preliminary list of prognostic-related genes.
[0069] The second phase was the Kaplan-Meier survival analysis. This phase further refined the list of prognostic-related genes selected in the first phase. For each gene in the list, patients were first divided into a high-expression group and a low-expression group based on its median expression level in the entire patient cohort. Subsequently, Kaplan-Meier survival curve analysis was performed on the high- and low-expression groups for the two clinical endpoints of progression-free survival (PFS) and overall survival (OS). This analysis visually illustrates the difference in survival probability between the two groups over time by plotting survival function curves. Simultaneously, the log-rank test was used to calculate the p-value to determine whether the difference between the two survival curves was statistically significant.
[0070] This embodiment sets strict screening criteria: a gene must simultaneously meet the following conditions: the log-rank test p-value for both its high and low expression groups in PFS analysis is less than 0.05, and the log-rank test p-value in OS analysis is also less than 0.05. Only genes that are significantly associated with both PFS and OS are defined as core prognostic genes and retained. Through this series of screenings, a significantly reduced set of core prognostic genes that are highly correlated with patient drug resistance relapse and survival outcomes is obtained, providing a foundation for the subsequent construction of robust predictive models.
[0071] See attached document Figure 2 and appendix Figure 3 After obtaining the core prognostic gene set that is significantly associated with patient survival, the method in this embodiment proceeds to the candidate gene identification and model core parameter construction steps based on machine learning. The purpose of this step is to reduce the dimensionality and select features of the core prognostic gene set through regularized regression to identify the gene combinations with the most predictive value and no collinearity, and to accurately calculate their weights in the model, thereby constructing a robust predictive model.
[0072] This step first employs Lasso (Least-Absolute-Shrinkage-and-Selection-Operator) regression analysis. This analysis uses the expression level matrix of core prognostic genes as input, patient progression-free survival (PFS) data as the response variable, and incorporates a Cox proportional hazards model. Lasso regression, by introducing an L1 penalty term into the model, can precisely compress the regression coefficients of some genes to zero while fitting the model, thus achieving variable selection. The optimal penalty parameter λ (lambda) value is determined by performing k-fold cross-validation (e.g., 10-fold cross-validation). In this embodiment, 10-fold cross-validation is performed, and the optimal penalty parameter λ value is determined based on the principle of minimizing partial-likelihood deviation. This value corresponds to a model with the smallest cross-validation error. (Appendix) Figure 2 This process is illustrated in the figure, which shows the trajectory of the regression coefficients of each gene as the log(λ) value changes, with the optimal λ value marked by a vertical line. At this optimal λ value, genes with non-zero regression coefficients are selected, forming an initial subset of candidate genes. In this embodiment, this step reduces the core prognostic gene set to 9 genes.
[0073] Next, to eliminate biased estimates of coefficients in the Lasso regression and further test the independent predictive value of the selected genes, these nine genes were included in a multivariate Cox proportional hazards regression model. This model uses the expression levels of these nine genes as covariates and analyzes their relationship with patient survival outcomes (PFS or OS). This model allows assessment of the independent contribution of each gene to patient survival risk after controlling for the effects of other gene expression levels. Finally, only genes with significant p-values (e.g., p < 0.05) in this multivariate analysis were identified as candidate genes for constructing the predictive model. In this embodiment, this step ultimately identified five candidate genes: OLR1, ULBP2, KCNK5, BDKRB2, and MUS81.
[0074] Based on these five finalized candidate genes and their respective regression coefficients obtained in multivariate Cox regression analysis, the following risk score calculation formula is established to quantify the risk of drug resistance relapse for each patient:
[0075]
[0076] In the formula: Risk Score represents the risk score calculated for a single patient; n represents the total number of final candidate genes, which is n = 5 in this embodiment; i represents the index of a single candidate gene, which ranges from 1 to n; β i Exp represents the regression coefficient value calculated for the i-th candidate gene in the above multivariate Cox regression analysis. This value is the core parameter for constructing this risk scoring model; i The value represents the gene expression level of the i-th candidate gene in the tumor tissue sample of the patient being tested. This value is a pre-processed (e.g., log2 conversion) quantitative transcriptome value.
[0077] See attached document Figure 4 and appendix Figure 5 After determining the risk score calculation formula, the method in this embodiment enters the final prediction model establishment and performance evaluation stage. The purpose of this stage is twofold: first, to transform the risk score model containing expression information of multiple candidate genes into an intuitive and operable clinical prediction tool; and second, to quantitatively evaluate the predictive performance of this tool.
[0078] To provide a graphical prediction tool that is easy to use in clinical practice, this embodiment will construct a nomogram based on the risk scoring models of five candidate genes: OLR1, ULBP2, KCNK5, BDKRB2, and MUS81. (See attached image) Figure 4 This is the nomogram constructed in this embodiment. Structurally, this nomogram integrates multiple predictor variables, presenting the results of a complex regression model graphically.
[0079] The nomogram consists of several axes: the upper part shows scores for multiple variables, including axes representing the expression levels of each candidate gene (OLR1 axis, ULBP2 axis, etc.) and a corresponding points axis. The middle part shows a total points axis, calculated by summing the points from each individual variable axis. The lower part shows multiple survival probability prediction axes for progression-free survival, including 1-year, 2-year, and 3-year survival axes. Users can locate the corresponding gene axis and read its point value based on the expression levels of each candidate gene in the patient's tumor sample. The points for all five genes are summed to obtain a total score. This total score is then located on the total points axis, and a vertical line is drawn downwards. The intersection of this line with the survival rate axis below represents the predicted 1-year, 2-year, and 3-year survival probabilities for the patient.
[0080] To objectively evaluate the predictive performance of the constructed nomogram model, this embodiment employs time-dependent receiver operating characteristic (ROC) curve analysis. The area under the ROC curve (AUC) is used as the core indicator for evaluating the model's discriminative ability, with a value ranging from 0.5 to 1.0. The closer the value is to 1.0, the stronger the model's predictive accuracy and discriminative ability.
[0081] Appendix Figure 5 The ROC curves for predicting 1-year and 3-year survival of patients in the TCGA cohort are shown. The horizontal axis represents the false positive rate (1-specificity), and the vertical axis represents the true positive rate (sensitivity). (See attached image.) Figure 5 As shown, the ROC curves predicting 1-year and 3-year survival rates are both located above the reference line (diagonal), indicating that the model has the ability to distinguish between different patient survival outcomes. By calculating the area under the curve, the quantitative AUC value of the model's predictive performance at specific time points (1 year and 3 years) can be obtained, which provides an objective measure of the model's predictive accuracy.
[0082] To verify the operability of the technical solution of the present invention in specific applications, this embodiment provides a specific process for applying a predictive model to conduct risk assessment.
[0083] See attached document Figure 4 This application example focuses on a patient with primary glioblastoma diagnosed with MGMT promoter methylation. First, a tumor tissue sample was obtained from the patient, and the expression levels of five candidate genes required for the model were determined using transcriptome sequencing. The standardized and logarithmically transformed gene expression level measurements are assumed to be as follows: OLR1 4.5, ULBP2 2.0, KCNK5 1.5, BDKRB2 2.5, and MUS81 5.6.
[0084] Next, use the appendix Figure 4 The nomogram shown is used to predict the survival prognosis of this patient. The specific steps are as follows:
[0085] First, find the point with a value of 4.5 on the OLR1 variable axis at the top of the nomogram, and draw a vertical line upward to the Points axis. Read the number of points corresponding to this vertical line as 30.
[0086] The second step is to find the point with a value of 2.0 on the ULBP2 variable axis and read its corresponding point number upwards, which is 20.
[0087] The third step is to find the point with a value of 1.5 on the KCNK5 variable axis and read its corresponding point number as 30.
[0088] Fourth step, find the point with a value of 2.5 on the BDKRB2 variable axis, and read its corresponding point number upwards to 50.
[0089] Fifth step, find the point with a value of 5.6 on the MUS81 variable axis and read its corresponding point number upwards as 10.
[0090] Then, the points corresponding to each of the above candidate genes are summed to calculate the patient's total points. The calculation process is: 30 + 20 + 30 + 50 + 10 = 140.
[0091] Finally, locate the 140-point mark on the Total-Points axis in the middle of the nomogram, and draw a perpendicular line downwards from that point, observing its intersection with the survival rate axes below. This perpendicular line intersects the 1-year survival rate axis at 0.2 and the 2-year survival rate axis at 0.2.
[0092] Therefore, based on the predictive model constructed in this embodiment, the predicted 1-year survival probability for this patient is 20%, and the predicted 2-year survival probability is also 20%. The predicted 3-year survival probability can also be obtained through the nomogram. These prediction results provide quantitative data for assessing the patient's temozolomide resistance risk and poor prognosis, and can serve as a reference for developing an individualized treatment plan.
[0093] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method of constructing a prediction model for temozolomide resistance in MGMT-negative glioblastoma patients, characterized by, The method comprises the following steps: S1, obtaining gene expression profile data of glioma cell lines sensitive to temozolomide and resistant to temozolomide, and screening differential genes through differential analysis; S2, obtaining transcriptome data and clinical prognosis information of MGMT-negative glioma patient cohort in TCGA database, and selecting clinical samples of MGMT-negative primary and receiving standard TMZ chemotherapy containing complete PFS and OS data, referred to as TCGA target cohort; S3, performing survival analysis on the screened differential genes in the TCGA target cohort to screen core prognostic genes significantly related to progression-free survival (PFS) and overall survival (OS); S4, performing Lasso regression analysis and multivariate Cox regression analysis on the screened core prognostic genes to determine candidate genes for constructing a prediction model; S5, constructing the prediction model based on the determined candidate genes and their regression coefficients in the multivariate Cox regression analysis. 2.The method for constructing a prediction model for temozolomide resistance of a MGMT-negative glioblastoma patient according to claim 1, characterized in that, In step S1, the step of obtaining gene expression profile data of glioma cell lines sensitive to temozolomide and resistant to temozolomide comprises: The MGMT-negative glioma cell line is the U251 cell line, and the gene expression profile data is derived from the GSE100736 dataset of the GEO database. 3.The method of claim 1, wherein the method is characterized by, In step S1, the step of screening differential genes through differential analysis comprises: The differential analysis uses the limma algorithm, and the threshold for screening differential genes is: P value<0.05 and |Log FC|>1; Wherein, P value is the probability value for determining the statistical significance of the difference in gene expression between the sensitive strain and the resistant strain; Log FC is the logarithmic value of the gene expression level change multiple of the resistant strain relative to the sensitive strain. 4.The method of claim 1, wherein the method is characterized by, In step S3, the step of performing survival analysis on the screened differential genes in the patient cohort comprises: S3-1, performing univariate Cox regression analysis of the differential genes related to progression-free survival (PFS) with a threshold of P value<0.05 to screen genes; S3-2, grouping the genes screened through univariate Cox regression analysis based on the median expression, performing Kaplan-Meier analysis related to progression-free survival (PFS) and overall survival (OS), and screening core prognostic genes with P value<0.05 in PFS analysis and OS analysis. 5.The method of claim 1, wherein the method is characterized by, The candidate genes include: OLR1, ULBP2, KCNK5, BDKRB2, and MUS81. 6.The method of claim 1, wherein the method is characterized by, In step S5, the step of constructing the prediction model comprises: Based on the expression levels of the candidate genes OLR1, ULBP2, KCNK5, BDKRB2, and MUS81 and their corresponding regression coefficients in the multivariate Cox regression analysis, the risk score is calculated by the following formula: where n is the number of candidate genes, i is the index of a single candidate gene, β i is the regression coefficient of candidate gene i in the multi-factor Cox regression analysis, Exp i is the expression level of candidate gene i. 7.The method of claim 1, wherein the method is characterized by, The prediction model is a nomogram model, and the nomogram model can predict the 1-year survival rate, 2-year survival rate, and 3-year survival rate of the patient's progression-free survival according to the risk score. 8.The method of claim 7, wherein the method is characterized by, The MGMT-negative glioma patient is an MGMT promoter methylation primary glioma patient.
9. A predictive model for temozolomide resistance in MGMT-negative glioblastoma patients, applied to the construction method according to any one of claims 1 to 8, characterized in that, The prediction model comprises a gene signature consisting of the following candidate genes: OLR1, ULBP2, KCNK5, BDKRB2 and MUS81. The prediction model is used to predict the risk of temozolomide drug resistance of a patient to be tested according to the expression level of the gene signature in the tumor tissue of the patient.
10. The predictive model of temozolomide resistance in a MGMT-negative glioblastoma patient according to claim 9, characterized in that, The model is a nomogram model, which converts the expression level of each gene in the gene signature into a risk score, and predicts the 1-year survival rate, 2-year survival rate and 3-year survival rate of the patient's progression-free survival according to the risk score.
Citation Information
Patent Citations
SE100736C1